Compare commits

...
824 Commits
Author SHA1 Message Date
Ryan Houdek 63ce78c41d Docs: Update for release FEX-2303 2023-03-06 08:50:41 -08:00
Mai fc38df2ff0 Merge pull request #2460 from Sonicadvance1/implement_memset
OpcodeDispatcher: Optimize REP STOS to MemSet operation
2023-03-04 11:31:58 -05:00
Mai 2fca207e14 Merge pull request #2462 from Sonicadvance1/fix_proton_2
FileManagement: Fixes Proton
2023-03-04 11:29:07 -05:00
Mai 308fa76aa3 Merge pull request #2463 from Sonicadvance1/update_rootfslinks
FEXRootFSFetcher: Update link to rootfs links file
2023-03-04 11:27:55 -05:00
Ryan Houdek b6ac26e0e9 FEXRootFSFetcher: Update link to rootfs links file
Switches to the new CDN which is significantly faster and has other
benefits.

In order to make sure we don't break old clients, switch to the new link
for a few months while leaving the old one operational.

The links file in the old CDN still points to the new rootfs links so
they get the performance improvement on old clients still.
2023-03-04 01:59:44 -08:00
Ryan Houdek ecd144de6a FileManagement: Fixes Proton
Need to ensure that dirfd is AT_FDCWD and also need to check flags
correctly.

Flags were incorrectly checking mode for O_WRONLY and also we should
check for O_APPEND. Split it out to a helper function just so it is
easier to see what is going on.

Fixes the issue of proton not finding `/lib64/ld-linux-x86-64.so.2`
2023-03-03 13:44:17 -08:00
Ryan Houdek 7e66508016 OpcodeDispatcher: Optimize REP STOS to MemSet operation
x86's REP STOS instruction is a memset (with element size!) with the
ability to choose a direction of execution.
Additionally it has a feature where if it faults part-way through the
copy, an application can catch the fault and continue afterwards to know
how many bytes got copied.

RCX is the counter which decrements for each element, and RDI is the
memory pointer. On fault these will reflect the last location that was
attempted to be written. FEX doesn't support this behaviour which makes
our lives easier.

Without supporting that feature, this turns in to a directional memset
by element size. Let's remove all the multiple blocks and just emit a
single IR operation to improve performance of the JIT.
Our generated code here was terrible, the IR was terrible, multiblock is
always slow with RA. Just a general overall improvement.

With profiling pressure-vessel this change deletes the hottest block that
appeared in the trace. This instruction is very commonly used for
memsetting a region to zero so it should be quite fast.

We can also optimize REP MOVS in the future with a Memcpy IR operation
in a similar fashion.

Additionally in the future these can be optimized to use ARM's new MOPS
instructions since the most common case is memset by byte. Which is when
we should expose the "Fast REP STOS" CPUID bit. Both setp/setm/sete and
cpyfp/cpyfm/cpyfe match `REP STOS` and `REP MOVS` respectively.
2023-03-03 09:16:28 -08:00
Ryan Houdek d6f50bf7b0 IR: Implement support for MemSet operation
This operation directly matches what the x86 STOS instruction does
without supporting its faulting behaviour.

STOS faulting behaviour is that RCX and RDI get updated to the last word
written. Which is something that FEX hasn't ever supported.
2023-03-03 09:16:28 -08:00
Ryan Houdek e7069f9f95 Merge pull request #2461 from lioncash/pair
ARMEmitter: Tidy up some assertion handling
2023-03-02 08:33:53 -08:00
Lioncache ea96ccb63d ARMEmitter: Add missing SVE floating-point compare vectors instructions
We're missing FACGE/FACGT and the aliases FACLE FACLT
2023-03-02 10:54:54 -05:00
Lioncache 4d4eac0987 ARMEmitter: Simplify SVE floating-point compare vectors
We can move the asserts into the helper function
2023-03-02 10:42:58 -05:00
Lioncache f0ee8a49b2 ARMEmitter: Simplify Emitter: SVE: SVE2 floating-point pairwise operations ops
We can centralize all of the assertion handling in the implementation
function.
2023-03-02 10:35:58 -05:00
Lioncache e0d8fc7c2b ARMEmitter: Simplify SVE2 integer halving add/subtract (predicated) ops
We can centralize all the assertion handling in the implementation
function.
2023-03-02 10:28:07 -05:00
Lioncache adf1de5562 ARMEmitter: Simplify SVE integer pairwise ops
We can centralize everything in the helper function itself, getting rid
of a few duplicated assertions.
2023-03-02 10:17:19 -05:00
Mai e310e29898 Merge pull request #2459 from Sonicadvance1/fix_pressure_vessel
FileManagement: Skip opening emulated writable files
2023-03-02 09:41:39 -05:00
Ryan Houdek 37ec68421c FileManagement: Skip opening emulated writable files
In the case that a file is getting opened to be created or writable then
skip EmuFD and rootfs searching for this file.
This fixes an edge case where if FEX was run with an unpacked rootfs
that was writable then pressure-vessel would break.

Fixes pressure-vessel with unpacked rootfs.
2023-03-02 00:26:34 -08:00
Ryan Houdek 41731e2680 Merge pull request #2458 from lioncash/pred
ARMEmitter: Remove predicate implicit conversion operators
2023-03-01 19:58:13 -08:00
Lioncache 5e6a3c6280 ARMEmitter: Remove predicate implicit conversion operators
Like with the vector registers, we can remove all implicit conversion
operators except the ones that convert down to the base PRegister class.

With this, all of the registers are now adequately constrained, so we
shouldn't have any wonky implicit conversions happening anymore.
2023-03-01 22:44:40 -05:00
Ryan Houdek e71e3ec930 Merge pull request #2457 from lioncash/sxtw
ARMEmitter: Make second sxtw parameter a WRegister
2023-03-01 19:35:31 -08:00
Lioncache 4cac100660 ARMEmitter: Make second sxtw parameter a WRegister
Matches the assembly use of it more closely.
2023-03-01 22:20:42 -05:00
Ryan Houdek 378e0692b9 Merge pull request #2456 from lioncash/reg
ARMEmitter: Remove implicit conversions from Register/XRegister/WRegister
2023-03-01 19:16:26 -08:00
Lioncache 678415c4c9 ARMEmitter: Remove implicit conversions from Register/XRegister/WRegister
Ensures that we're always explicit about the size of a register when
using APIs that enforce it.

The only implicit conversions we keep are conversions that convert down
to Register, but not anything that converts up the hierarchy or across
it.
2023-03-01 21:59:53 -05:00
Ryan Houdek e869b2fe67 Merge pull request #2455 from lioncash/comp
ARMEmitter: Remove predicate uint32_t conversion operators
2023-03-01 18:37:34 -08:00
Lioncache 2194a1027c ARMEmitter: Remove predicate uint32_t conversion operators
Now that we have dedicated comparison operators, we no longer need to
keep these implicit conversion operators around.
2023-03-01 21:16:55 -05:00
Lioncache 52b4378e49 ARMEmitter: Add comparison functions to register types
Gets rid of the need to compare indices directly in order to compare
register equality
2023-03-01 21:15:46 -05:00
Ryan Houdek 0f45318040 Merge pull request #2454 from lioncash/convert
ARMEmitter: Remove most implicit conversion operators for vector register types
2023-03-01 17:57:21 -08:00
Ryan Houdek 21fbcef0bd Merge pull request #2453 from lioncash/explicit
ARMEmitter: Make VRegister constructor explicit
2023-03-01 17:54:49 -08:00
Ryan Houdek ef02083767 Merge pull request #2452 from lioncash/consecutive
ARMEmitter: Handle sequential registers in lists nicer
2023-03-01 17:53:47 -08:00
Ryan Houdek 24904f48c4 Merge pull request #2451 from lioncash/saddl
ARMEmitter: Simplify size handling Advanced SIMD 3 different group
2023-03-01 17:45:42 -08:00
Lioncache 9461ab5094 ARMEmitter: Remove conversion operators for VRegister 2023-03-01 18:40:12 -05:00
Lioncache 66c8b14470 ARMEmitter: Remove conversion operators for QRegister 2023-03-01 18:29:11 -05:00
Lioncache 6ab78ca93b ARMEmitter: Remove conversion operators for DRegister 2023-03-01 18:20:32 -05:00
Lioncache 2fe808f5cd ARMEmitter: Remove conversion operators for SRegister 2023-03-01 18:02:31 -05:00
Lioncache 81a94b9ffe ARMEmitter: Remove conversion operators for BRegister 2023-03-01 18:00:24 -05:00
Lioncache 0d87ed46da ARMEmitter: Remove conversion operators for HRegister 2023-03-01 17:58:06 -05:00
Mai 545a216da6 Merge pull request #2448 from Sonicadvance1/optimize_openat
EmulatedFiles: Optimize openat handler
2023-03-01 17:20:14 -05:00
Lioncache 11f65df554 ARMEmitter: Make VRegister constructor explicit
All other parameter taking constructors for the other register types are
explicit, so this just makes behavior more consistent.
2023-03-01 16:48:10 -05:00
Lioncache 29ff642499 ARMEmitter: Make use of sequential register helper
Fixes assertion behavior on quite a bit of ASIMD load-store operations
as well as a few SVE ops as well
2023-03-01 15:25:46 -05:00
Lioncache d36517a9d3 ARMEmitter: Add helper for determining if vectors are sequential
A few vector instructions that take register lists often require
vector registers within the list to be sequential in the form of an
increasing list modulo the register file size.

For example:

v1,  v2, v3, v4
v31, v0, v1, v2

both fit these requirements.

This will be used to enforce this restriction within the asserts from a
single place.
2023-03-01 15:24:07 -05:00
Lioncache 83419b410d ARMEmitter: Simplify size handling Advanced SIMD 3 different group
A large amount of size handling in this category is just decrementing
the size by 1, so we can tidy up a bunch of conditionals by just doing
that instead.
2023-03-01 11:04:40 -05:00
Mai 77fad28b69 Merge pull request #2447 from Sonicadvance1/add_hypervisorbit_hide_option
CPUID: Adds an config option to hide hypervisor bit
2023-02-28 10:59:21 -05:00
Mai 70aefc9db2 Merge pull request #2450 from Sonicadvance1/fix_fexserver_zombie
FEXServerClient: Fixes instance where FEXServer can create a zombie
2023-02-28 10:58:11 -05:00
Mai d2e0adf540 Merge pull request #2449 from Sonicadvance1/fix_fexserver_daemon_systemd
FEXServer: Change systemd service environment variable key
2023-02-28 10:57:09 -05:00
Ryan Houdek 84060cd947 FEXServerClient: Fixes instance where FEXServer can create a zombie
When FEXServer is daemonizing through an instance of FEXLoader or
FEXInterpreter, it would leave a zombie process which was waiting for us
to read the process status.
Since we don't care about the child status and don't want to get blocked
by waitpid, just ignore the signal.

This tells the kernel that we don't care about the signal and will kill
the zombie process immediately.
Didn't notice this before since FEXServer started failing to daemonize.
2023-02-28 05:09:16 -08:00
Ryan Houdek aaf17b6d41 FEXServer: Change systemd service environment variable key
It looks like `SYSTEMD_EXEC_PID` can leak through to the executable
environment in regular situations. Instead let's key off of
`INVOCATATION_ID` which doesn't leak through.

Fixes an edge case behaviour where FEXServer wouldn't daemonize in some
systemd environments.
2023-02-28 05:07:02 -08:00
Ryan Houdek 8ded25ada7 EmulatedFiles: Optimize openat handler
Fixes #2443
I found out with some profiling that this we were spending a decent
amount of time with the `openat` syscall in heavily utilized situations.
While not super common in active gameplay situations, it matters
significantly in loading screens that this is fairly optimal.

The bulk of the time is spent in the emulated files handler to ensure
that whatever path we are given, we can capture file paths that we need
to emulate. The largest contributor being the std::filesystem::canonical
function call.

A couple of optimizations in place here.
1) Do a quick hashmap check right at the start to see if we exactly fit
2) Change from `std::fs::canonical` to `realpath`
3) Switch `GetEmulatedFDPath` to not use optional so it stops building
   on the stack

I'm still not super happy with the performance of `realpath` and also
not happy that we still need to use `lexically_normal` in one code path.
But short of writing a super hand-optimized `realpath` that fits our
constraints, I don't think we can do better.

Micro benchmark needs to test four different situations due to this
optimization.
1) Non-EmuFD path
2) Non-EmuFD path with dirfs
3) EmuFD path
4) EmuFD path with dirfs

And the performance improvement for each situation respectively
1) 12% performance improvement
  - 213413 openat syscalls/s -> 238999 syscalls/s
2) 17% performance improvement
  - 202085 openat syscalls/s -> 237309 syscalls/s
3) 17% performance improvement (/proc/cpuinfo)
  - 56616 openat syscalls/s -> 66231 syscalls/s
  - Includes overhead of generating temp FD and close syscall
4) 5% performance improvement (/proc/cpuinfo)
  - 51080 openat syscalls/s -> 53956 syscalls/s
  - Includes overhead of generating temp FD and close syscall

And for sake of comparison to the non-emulated system; My test system
can hit around 1-1.1 million openat syscalls per second in the same
microbench.

Nice little performance uplift.
2023-02-28 04:00:28 -08:00
Ryan Houdek 5b9fe8f26b CPUID: Adds an config option to hide hypervisor bit
This is known to cause issues in some cases. We hit the first game that
checks for this bit and early exits if it is found.

Lets the MMORPG Tibia run in non-VM situations.
Looks like they have more checks for VMs other than hypervisor bit, so
running under Parallels still won't work. Running on bare Linux is fine.
2023-02-27 23:11:23 -08:00
Ryan Houdek e65b429c83 Merge pull request #2446 from lioncash/cpy
ARMEmitter: Simplify advanced SIMD copy
2023-02-27 19:52:05 -08:00
Lioncache dd290f129f ARMEmitter: Simplify advanced SIMD copy
Same behavior, but collapses some if statements.
2023-02-27 22:32:39 -05:00
Ryan Houdek 1832cc80d6 Merge pull request #2445 from lioncash/unsigned
ARMEmitter: Centralize handling for unsigned offset load-stores
2023-02-27 18:18:30 -08:00
Ryan Houdek fe1faf9ebe Merge pull request #2444 from lioncash/scalar
ARMEmitter: Handle SVE Integer Compare - Scalars group
2023-02-27 18:16:31 -08:00
Lioncache 12d0a7fa98 ARMEmitter: Use constants for unsigned offset encoding limits
Allows us to give some names to these constants that are used in the
JIT instead of writing them by hand.
2023-02-27 17:44:17 -05:00
Lioncache c1b08079f3 ARMEmitter: Strengthen unsigned immediate load/store helper
Centralizes all the shifting behavior and whatnot into a single
function, making everything much more localized.

Also gets rid of a lot of magic constants related to the encoding limits
of immediates.
2023-02-27 17:44:13 -05:00
Mai d688026fe4 Merge pull request #2442 from Sonicadvance1/fix_misaligned_stack_signals
Dispatcher: Fixes crash with misalign stack returning from signal
2023-02-27 14:48:39 -05:00
Lioncache 1c388b455a ARMEmitter: Move missed SVE public helpers into private section 2023-02-27 14:45:59 -05:00
Lioncache 9426abc98d ARMEmitter: Handle SVE pointer conflict compare group 2023-02-27 14:27:29 -05:00
Lioncache 5ee2db34a7 ARMEmitter: Handle SVE conditionally terminate scalars group 2023-02-27 14:22:57 -05:00
Lioncache ad37c19043 ARMEmitter: Handle SVE integer compare scalar count and limit group 2023-02-27 14:09:40 -05:00
Ryan Houdek 0e6c5911b8 Dispatcher: Fixes crash with misalign stack returning from signal
When we were taking a signal that had a misaligned stack, we would store
the host stack at a weird offset.

After that point when we were trying to sigreturn we wouldn't know the
alignment of the stack coming back and we would try loading the host
stack from the wrong offset. Easy fix is to just align the host stack
location.

Fixes Ender Lilies, which was consistently crashing from a SIGCHLD due
to having a misaligned stack.

Side-change: Move the cookie check to the start of the restore. Doesn't
make sense to check the cookie after restoring state since it could be
quite wrong.
2023-02-26 19:58:22 -08:00
Ryan Houdek f2aa0026b5 Merge pull request #2439 from lioncash/log
Emitter/ALUOps: Fix typos in log messages
2023-02-23 16:44:19 -08:00
Lioncache 553efbeb29 Emitter/ALUOps: Fix typos in log messages
Fixes a few incorrect instruction names in the logs.
2023-02-23 19:14:45 -05:00
Ryan Houdek b39a882a2d Merge pull request #2438 from lioncash/restrict
OpcodeDispatcher: Restrict partial XMM stores to FPRs in StoreResult_WithOpSize
2023-02-23 16:06:53 -08:00
Lioncache 85f7f8e6c0 OpcodeDispatcher: Restrict partial XMM stores to FPRs in StoreResult_WithOpSize
As far as I know, nothing actually uses this path. Partially resolves
the TODO of dealing with partial writes.
2023-02-23 18:48:24 -05:00
Ryan Houdek 4d25de31de Merge pull request #2437 from lioncash/dup
OpcodeDispatcher: Remove now unused _VDupElement path in LoadSource_WithOpSize
2023-02-23 13:58:09 -08:00
Ryan Houdek 9e01730c6c Merge pull request #2436 from lioncash/builtin
Arm64Emitter: Use bit utils wrapper over __builtin_ffs
2023-02-23 13:51:24 -08:00
Lioncache fea3ee1298 OpcodeDispatcher: Remove now unused _VDupElement path in LoadSource_WithOpSize
XMM instances can't use high indices anymore, since we've gotten rid of
the only flag that allows this scenario to occur.
2023-02-23 16:09:03 -05:00
Lioncache e1c42315ed Arm64Emitter: Use bit utils wrapper over __builtin_ffs
Just keeps the use of builtins contained to one place.
2023-02-23 15:23:07 -05:00
Ryan Houdek 165db37c8d Merge pull request #2434 from lioncash/predmisc
ARMEmitter: Finish off SVE Predicate Misc group
2023-02-23 12:20:03 -08:00
Ryan Houdek 9b23ae9133 Merge pull request #2433 from lioncash/subsw
OpcodeDispatcher: Handle VPHSUBSW
2023-02-23 12:17:58 -08:00
Ryan Houdek f951a406e6 Merge pull request #2435 from lioncash/mov
OpcodeDispatcher: Share MOVHPD implementation with MOVHPS
2023-02-23 12:16:27 -08:00
Lioncache 497b5c0561 X86Tables: Reclaim FLAGS_SF_HIGH_XMM_REG as an unused flag
Now that we've moved MOVHPS over to sharing the implementation of
MOVHPD, the FLAGS_SF_HIGH_XMM_REG is now unused.

Since we're supporting AVX, this flag is kind of weird in terms of
behavior, since what determines the high part of a register is now
situationally different.

Also it's much more explicit to perform the insert directly in the
implementation of instructions, than relying on a flag to do it for us.

So, instead of keeping it around, we can reclaim it as unused for use
with any necessary behavior that we would require in the future.
2023-02-23 14:22:25 -05:00
Lioncache 95393b07fb OpcodeDispatcher: Share MOVHPD implementation with MOVHPS
These instructions essentially have the same behavior. This also allows
us to remove the only used instance of FLAGS_SF_HIGH_XMM_REG, which,
given that we now support AVX, has ambiguous use.

While we're at it, we can expand the tests to make use of the store to
memory variant.

Also removes an erroneous copy-pasted comment about ZEXTing. This is
from the MOVQ implementation function. MOVHPS/MOVHPD don't do any
ZEXTing, they either store to memory or insert into a register.
2023-02-23 14:07:22 -05:00
Lioncache 372da1b820 ARMEmitter: Handle PNEXT
Now, with the helper in place, we can implement PNEXT and finish off the
SVE Predicate Misc group.
2023-02-23 11:59:39 -05:00
Lioncache 61a59d0314 ARMEmitter: Unify SVE Predicate Misc group under single helper
Centralizes the implementations and also gets rid of some code in the
process.
2023-02-23 11:51:03 -05:00
Lioncache 1045e05870 OpcodeDispatcher: Handle VPHSUBSW 2023-02-23 10:55:57 -05:00
Lioncache 052872725c OpcodeDispatcher: Factor out PHSUBS implementation into helper
This will allow it to be shared in the AVX implementation.
2023-02-23 10:32:48 -05:00
Ryan Houdek 4d655218ab Merge pull request #2431 from lioncash/brk
ARMEmitter: Handle SVE partition break categories
2023-02-22 21:13:33 -08:00
Lioncache 78ba195b66 ARMEmitter: Handle SVE partition break condition category 2023-02-22 22:49:37 -05:00
Lioncache ae2b28716d ARMEmitter: Handle SVE propagate break to next partition category 2023-02-22 22:41:55 -05:00
Lioncache 326e5e8d57 ARMEmitter: Handle propagate break from previous partition category 2023-02-22 22:35:31 -05:00
Ryan Houdek 0a8fc2cbef Merge pull request #2430 from lioncash/assert
ARMEmitter: Handle SVE integer compare with wide elements category
2023-02-22 18:57:33 -08:00
Lioncache 552293b226 ARMEmitter: Handle SVE integer compare with wide elements category
We can piggy-back on top of the existing SVEIntegerCompareVector to make
these trivial to implement.
2023-02-22 21:03:35 -05:00
Lioncache 751a4c8019 ARMEmitter: Move assertion into SVEIntegerCompareVector
Same behavior, but centralizes the assertion. While we're at it, we can
also add another assert to ensure that only predicates p0-p7 are used.
2023-02-22 20:14:49 -05:00
Ryan Houdek 68b2072eab Merge pull request #2429 from lioncash/align
OpcodeDispatcher: Handle alignment for MOVAPS a little better
2023-02-22 14:40:28 -08:00
Lioncache e3cac40b1b OpcodeDispatcher: Fix SSE MOVAPS variants being treated as MOVUPS
0x10/0x11 in the two byte op table corresponds to MOVUPS
0x28/0x29 in the two byte op table corresponds to MOVAPS
2023-02-22 15:56:15 -05:00
Ryan Houdek 9b123353b3 Merge pull request #2428 from lioncash/hsub
OpcodeDispatcher: Handle VHSUBPD/VHSUBPS
2023-02-22 11:43:10 -08:00
Lioncache f2c0c55b9c OpcodeDispatcher: Handle VHSUBPS 2023-02-22 14:27:51 -05:00
Ryan Houdek a4c694ffc7 Merge pull request #2427 from lioncash/pred
ARMEmitter: Finish off SVE Permute Vector - Predicated group
2023-02-22 11:26:49 -08:00
Lioncache 1eb722dea7 OpcodeDispatcher: Handle VHSUBPD 2023-02-22 14:12:11 -05:00
Lioncache a6746988d7 x86_64/VectorOps: Fix behavior of UnZip2 with 64-bit element 256-bit vectors
The 256-bit variant of vshufpd uses extra immediate bits rather than the
same bits for the lower lane.
2023-02-22 14:12:11 -05:00
Lioncache 0a1707f1bd OpcodeDispatcher: Factor HSUBP implementation into helper
Will be used for implementing the AVX variants of the same instructions.
2023-02-22 12:12:55 -05:00
Lioncache 23b9d8e108 ARMEmitter: Add check for registers being consecutive in constructive SPLICE
Will catch cases where registers aren't consecutive in the constructive
variant. While we're at it, we can also amend EXT's similar but slightly wrong
consecutive check.

Also adds tests to ensure these corner-cases hold.
2023-02-22 11:59:58 -05:00
Lioncache 3fa44604ba ARMEmitter: Make SPLICE use SVEPermuteVectorPredicated
These are in the same instruction category, so we can use the helper to
simplify the implementation.
2023-02-22 11:48:11 -05:00
Lioncache d3bc0c084d ARMEmitter: Make CPY (SIMD&FP) and CPY (scalar) use SVEPermuteVectorPredicated
These fall under the same instruction category, so we can use the helper
to simplify the implementation.
2023-02-22 11:29:07 -05:00
Lioncache e206414919 ARMEmitter: Make COMPACT use SVEPermuteVectorPredicated
This falls under the same category of instructions, so we can use it to
simplify the implementation.
2023-02-22 11:23:37 -05:00
Lioncache e635cc5404 ARMEmitter: Use predicated helper with revb/revh/revw/rbit
Since these are under the same category, we can merge these and get rid
of a now unnecessary helper.
2023-02-22 11:17:12 -05:00
Lioncache 822d67467b ARMEmitter: Handle SVE conditionally extract element to GPR/scalar categories 2023-02-22 11:10:13 -05:00
Lioncache e043d2c0f5 ARMEmitter: Handle SVE conditionally broadcast element to vector category 2023-02-22 10:55:34 -05:00
Lioncache d58c4405f7 ARMEmitter: Handle extract element to general register/scalar categories 2023-02-22 10:47:49 -05:00
Mai 66d879f387 Merge pull request #2400 from Sonicadvance1/rip_reconstruct
Dispatcher: Support reconstructing RIP from block entry
2023-02-22 09:48:02 -05:00
Mai 55d3edb8e6 Merge pull request #2426 from Sonicadvance1/optimize_getemulatedpath
FileManagement: Optimize GetEmulatedFDPath with an FD!
2023-02-22 09:46:44 -05:00
Ryan Houdek 98f0f22f41 FileManagement: Optimize GetEmulatedFDPath with an FD!
Performance stats up front:
This improves pressure-vessel startup time on my test device by 10.1%
Improving the startup time from 9.71425 seconds to 8.7421 seconds.

Most filesystem based syscalls support a file descriptor version with an
*at suffix. This allows us to do these syscalls with pathnames that are
relative to the directory FD that is passed to the syscall.

This is pretty much exactly what we want when we are searching for files
inside of our rootfs. The only quirk ends up being that we are getting
passed absolute paths. This ends up being very simple to workaround by
stripping off the front '/' character. Doing this is just offsetting the
pointer passed to the syscall by one byte.

This does require having two temporary buffers of size PATH_MAX passed
to the handler since just like in the other implementation, we need to
keep the previous result around. The difference being now that we aren't
doing a bunch of std::string temporary manipulation and now we are
returning one of the passed in buffers back depending on the result.
2023-02-22 01:25:40 -08:00
Ryan Houdek 5f574fb935 Merge pull request #2425 from lioncash/xop
VEXTables: Remove VPERMIL2PD and VPERMIL2PS entries
2023-02-20 18:27:32 -08:00
Ryan Houdek 618f5bb869 Merge pull request #2424 from lioncash/permil
OpcodeDispatcher: Handle register variants of VPERMILPD/VPERMILPS
2023-02-20 18:02:14 -08:00
Lioncache b5ca5f173e VEXTables: Remove VPERMIL2PD and VPERMIL2PS entries
These are actually XOP instructions. That, despite being so, are encoded
using a VEX prefix.
2023-02-20 20:58:55 -05:00
Lioncache 5cf6a680bb OpcodeDispatcher: Handle register variants of VPERMILPD/VPERMILPS 2023-02-20 20:32:31 -05:00
Ryan Houdek 645f40bb96 Merge pull request #2423 from lioncash/permd
OpcodeDispatcher: Handle VPERMD/VPERMPS
2023-02-20 15:20:59 -08:00
Ryan Houdek 268deddd09 Merge pull request #2422 from lioncash/phadds
OpcodeDispatcher: Handle VPHADDSW
2023-02-20 15:20:20 -08:00
Ryan Houdek e4488b0cfc Merge pull request #2421 from lioncash/index
ARMEmitter: Handle SVE index generation category
2023-02-20 15:16:33 -08:00
Lioncache 65b9dcd20b OpcodeDispatcher: Handle VPERMPS
With the VPERMD work in place, this is trivial to support.
2023-02-20 17:00:39 -05:00
Lioncache b2c333c383 OpcodeDispatcher: Handle VPERMD 2023-02-20 17:00:35 -05:00
Lioncache 1ea53c65ab x86_64/VectorOps: Handle 8-bit VShlI IR op
Useful for handling VPERMD.
2023-02-20 16:45:13 -05:00
Lioncache 8beae0fce4 OpcodeDispatcher: Add VTrn/VTrn2 IR opcodes
Provides a convenient way to propogate indices at given intervals in
vectors. This makes permutation instructions a little less annoying to
implement.
2023-02-20 16:44:14 -05:00
Lioncache add775c5cd OpcodeDispatcher: Handle VPHADDSW 2023-02-20 12:25:28 -05:00
Lioncache f3e6f62356 OpcodeDispatcher: Factor PHADDS implementation into helper
This will be used to also handle the VEX variant of PHADDSW
2023-02-20 12:00:38 -05:00
Lioncache 59ab10f155 ARMEmitter: Move SVE instruction helpers into privare section
Moves some instruction helpers that existed outside of the private
section of the class back into them, so that we're not exposing
unnecessary things in the interface.
2023-02-20 11:39:21 -05:00
Lioncache 2e1bd4b32b ARMEmitter: Handle SVE index generation category 2023-02-20 11:30:09 -05:00
Mai f71f2445db Merge pull request #2389 from Sonicadvance1/remove_context_c_interface
FEXCore: Removes C wrapper interface
2023-02-20 10:08:07 -05:00
Mai f6e2fe1515 Merge pull request #2420 from Sonicadvance1/fix_syscall_race
Arm64: Fixes a race condition on syscall spilling SRA
2023-02-20 10:06:53 -05:00
Mai 11c8db5a14 Merge pull request #2419 from Sonicadvance1/cortex_c_classify
Scripts: Update fit_native script for X1C/A78C
2023-02-20 10:06:07 -05:00
Mai 65b2da20d6 Merge pull request #2418 from Sonicadvance1/optimize_aluop_dispatcher
OpcodeDispatcher: Optimize ALUOp handler
2023-02-20 10:05:47 -05:00
Ryan Houdek 273f5e1f26 Arm64: Fixes a race condition on syscall spilling SRA
When executing a non-inlined syscall, we spill all static registers.
We weren't storing in to the thread context that we have done this.
If a signal occured between FEX returning from the syscall (after the
blr) and before the `FillStaticRegs` then the signal handler would get
the incorrect register state.

This typically manifested as Steam getting a SIGCHLD, trying to recover
the guest stack pointer, and it that pointer would be zero or some other
corrupt value. Thus crashing inside of the signal handler.

Surprising that we hadn't hit this way more before this point, must have
needed hardware that tickled the race condition *just* right.
2023-02-19 16:12:06 -08:00
Ryan Houdek 35af4bd42a FEXCore: Removes C wrapper interface
This has been a long time coming. The C interface has been a thorn in
our side for no reason for a long time.

The purpose of this step is to remove the C interface without changing
behaviour as much as possible. This means that with this commit there
are still some bad practices but the remaining issues will be solved
with followup PRs.

Primarily, we still have a `DestroyContext(CTX)` static function which calls
the Context implementation's `DestroyContext` and does a raw C++ delete.

Follow up PR will remove that, but I didn't want to touch it yet since
it'll require checking to ensure the unique_ptr changes play nice with
our allocator hooking. Which this is already a huge PR without trying to
change behaviour.
2023-02-19 11:59:11 -08:00
Ryan Houdek 7f1464b135 Scripts: Update fit_native script for X1C/A78C
Cortex-X1C and A78C are relatively minor changes to their non-C
counterparts. Support classifying them in case clang understands them.

Fixes a minor perf regression noticed on the Lenovo X13s while testing.
2023-02-18 23:18:48 -08:00
Ryan Houdek e594b2c4c7 OpcodeDispatcher: Optimize ALUOp handler
Take a leaf from the Vector ops and have the jump entry choose the IR
op.
Also generate one atomic op and modify the IR type in the locked memory
type just like the non locked memory path.

This class of instructions in the number one instruction type percentage
wise, so making this more optimal will be a win.

It's a fairly minor optimization so it should be a small impact.
2023-02-18 03:32:43 -08:00
Ryan Houdek 2aead5aec2 Config: Removes the x86dec_SynchronizeRIPOnAllBlocks option
This is no longer necessary since we reconstruct up to block entry from
the previous commit.
2023-02-18 02:48:55 -08:00
Ryan Houdek c3f1f602fe Dispatcher: Support reconstructing RIP from block entry
This allows us to not update RIP on block entry, but still allow
reconstructing the RIP up until that point.

While still not full RIP reconstruction, this lets us update the signal
context's RIP just like the `x86dec_SynchronizeRIPOnAllBlocks` without
eating the cost of writing to RIP on block entry.
2023-02-18 02:48:55 -08:00
Ryan Houdek 9c256bfe96 Merge pull request #2413 from lioncash/unpred
ARMEmitter: Handle a few more vector permutation categories
2023-02-15 14:51:59 -08:00
Ryan Houdek b5bc8cd294 Merge pull request #2416 from lioncash/mov
VectorOps: Remove unnecessary mov in VUShrNI2/VSQXTN2/VSQXTUN2
2023-02-15 14:46:02 -08:00
Lioncache e78b573610 VectorOps: Remove unnecessary mov in VUShrNI2/VSQXTN2/VSQXTUN2
We can move the initial move down by SPLICE, which not only lets us turn
it into a MOVPRFX, but also we can safely move into the final
destination register directly, since we can be sure there's no
potential dependencies at this point
2023-02-15 17:14:26 -05:00
Mai 81a89ab747 Merge pull request #2415 from Sonicadvance1/spillsra_fix
Dispatcher: Fixes guest stack register usage
2023-02-15 15:51:49 -05:00
Ryan Houdek fd17a3de50 Dispatcher: Fixes guest stack register usage
Fixes #2410

We were pulling the guest RSP before spilling static registers back to
the state.
Move this to after we spill SRA state to fix this bug.

Thanks to @ifquant for diving in, identifying, and finding the exact bug.
2023-02-15 12:21:11 -08:00
Ryan Houdek 2f260ae6ad Merge pull request #2414 from lioncash/sve-ex
Arm64/VectorOps: Use SVE only with 256-bit op sizes
2023-02-15 12:11:05 -08:00
Lioncache bf7118fc85 Arm64/VectorOps: Use SVE only with 256-bit op sizes
Keeps all of the IR ops consistent with each other. Also removes some
redundant scalar checks that weren't really necessary.
2023-02-15 14:46:05 -05:00
Ryan Houdek a90f5363dd Merge pull request #2412 from lioncash/ptest
OpcodeDispatcher: Handle VPTEST
2023-02-15 10:31:12 -08:00
Ryan Houdek ab03e59500 Merge pull request #2411 from lioncash/zero
OpcodeDispatcher: Use VectorZero over VectorImm in InsertPSOpImpl
2023-02-15 10:30:23 -08:00
Lioncache a59d700bbe ARMEmitter: Handle SVE Permute Predicate category 2023-02-15 12:56:39 -05:00
Lioncache 7cf27a7c26 ARMEmitter: Handle SVE Permute Vector - Unpredicated category 2023-02-15 12:18:42 -05:00
Lioncache 14e1d16710 OpcodeDispatcher: Handle VPTEST 2023-02-15 11:24:37 -05:00
Lioncache 203f29a91f OpcodeDispatcher: Use VectorZero over VectorImm in InsertPSOpImpl
A little more straightforward than using VectorImm for the same purpose.
2023-02-15 09:37:54 -05:00
Ryan Houdek 25f0a03ceb Merge pull request #2407 from lioncash/mov
OpcodeDispatcher: Handle VMOVSD/VMOVSS
2023-02-14 22:36:13 -08:00
Lioncache 3ced41414e OpcodeDispatcher: Handle VMOVSD 2023-02-15 01:18:54 -05:00
Lioncache 1a64b26d03 OpcodeDispatcher: Handle VMOVSS 2023-02-15 01:18:15 -05:00
Ryan Houdek efafe0e6e9 Merge pull request #2408 from lioncash/pmaddwd
OpcodeDispatcher: Handle VPMADDWD
2023-02-14 17:52:57 -08:00
Ryan Houdek 35746c7669 Merge pull request #2406 from lioncash/shuffle
OpcodeDispatcher: Handle VSHUFPD/VSHUFPS
2023-02-14 17:47:39 -08:00
Lioncache 4a69b87cb9 OpcodeDispatcher: Handle VPMADDWD 2023-02-14 18:50:37 -05:00
Lioncache fb2de47e73 OpcodeDispatcher: Factor out PMADDWD implementation to helper
This will be used to centralize code to also implement the AVX variant.
2023-02-14 18:38:12 -05:00
Lioncache bcee3e9374 OpcodeDispatcher: Handle VSHUFPS 2023-02-14 16:46:03 -05:00
Lioncache 6d87154ac8 OpcodeDispatcher: Handle VSHUFPD 2023-02-14 16:46:03 -05:00
Lioncache c5d799df8c OpcodeDispatcher: Make SHUFOpImpl suitable for AVX
Drops in the AVX-specific bits into the helper in preparation for
implementing VSHUFPD and VSHUFPS
2023-02-14 16:45:32 -05:00
Lioncache 449645669a OpcodeDispatcher: Move SHUFOp implementation to helper function
Will be useful for handling both the SSE and AVX variants in the same
place.
2023-02-14 16:43:33 -05:00
Ryan Houdek 3ac7b2cddf Merge pull request #2405 from lioncash/shufw
OpcodeDispatcher: Handle VPSHUFD/VPSHUFHW/VPSHUFLW
2023-02-14 10:36:48 -08:00
Lioncache 504d409cf6 OpcodeDispatcher: Handle VPSHUFD 2023-02-14 13:13:09 -05:00
Lioncache 29a6d584a9 OpcodeDispatcher: Handle VPSHUFHW 2023-02-14 12:48:42 -05:00
Lioncache 310fcf969c OpcodeDispatcher: Handle VPSHUFLW 2023-02-14 12:32:47 -05:00
Ryan Houdek b329442c09 Merge pull request #2404 from lioncash/dup
IR: Add VDupFromGPR
2023-02-13 14:37:59 -08:00
Lioncache f4d799abdd OpcodeDispatcher: Make use of VDupFromGPR where applicable
Simplifies some of the IR usage.
2023-02-13 16:52:51 -05:00
Lioncache 4bb7f49c2a IR: Add VDupFromGPR
Allows broadcasting constants into vectors from GPRs. Resolves the only
remaining TODOs within our vector ops.
2023-02-13 16:52:47 -05:00
Ryan Houdek f7f2dc2210 Merge pull request #2403 from lioncash/err
ARMEmitter/ASIMDOps: Amend a few error logs
2023-02-13 12:31:52 -08:00
Ryan Houdek d40812929f Merge pull request #2402 from lioncash/shufb
OpcodeDispatcher: Handle VPSHUFB
2023-02-13 12:18:16 -08:00
Lioncache c381185a7d ARMEmitter/ASIMDOps: Amend a few error logs
A few were logging out the wrong instruction name on a precondition
failure.
2023-02-13 15:17:09 -05:00
Lioncache ac5d09885e OpcodeDispatcher: Handle VPSHUFB 2023-02-13 14:47:37 -05:00
Lioncache d9a505e22e OpcodeDispatcher: Factor PSHUFB implementation into helper
Will let us centralize the implementation for PSHUFB and VPSHUFB
2023-02-13 12:39:27 -05:00
Ryan Houdek a96ad0fc9d Merge pull request #2401 from lioncash/palign
OpcodeDispatcher: Handle VPALIGNR
2023-02-13 09:35:19 -08:00
Lioncache 9268a356f6 OpcodeDispatcher: Handle VPALIGNR 2023-02-13 10:53:02 -05:00
Lioncache 92141d3edc OpcodeDispatcher: Factor PALIGNR code into helper
Will allow us to centralize the implementation of PALIGNR and VPALIGNR.
2023-02-13 09:49:26 -05:00
Ryan Houdek 8c8b680640 Merge pull request #2398 from lioncash/sve2acc
ARMEmitter: Handle SVE2 Accumulate category
2023-02-10 23:42:43 -08:00
Lioncache 504be62a92 ARMEmitter: Handle SVE2 integer absolute difference and accumulate 2023-02-11 00:27:07 -05:00
Lioncache 62e2f1b45d ARMEmitter: Handle SVE2 bitwise shift and insert category 2023-02-11 00:27:04 -05:00
Lioncache 880cc72842 ARMEmitter: Handle SVE2 bitwise shift right and accumulate 2023-02-11 00:24:18 -05:00
Lioncache feacd897fc ARMEmitter: Handle SVE2 integer add/sub long with carry category 2023-02-11 00:24:18 -05:00
Lioncache fdd950e1d1 ARMEmitter: Handle SVE2 integer absolute difference and accumulate long category 2023-02-11 00:24:17 -05:00
Lioncache b02af95629 ARMEmitter: Handle SVE2 complex add category 2023-02-10 22:31:45 -05:00
Ryan Houdek 2bd64ad24e Merge pull request #2396 from lioncash/narrow
ARMEmitter: Finish off SVE Misc category
2023-02-09 10:28:52 -08:00
Lioncache fa95a823c9 ARMEmitter: Handle SVE2 bitwise shift left long category 2023-02-09 06:30:43 -05:00
Lioncache 399ed61380 ARMEmitter: Handle SVE2 integer add/sub interleaved long 2023-02-09 05:44:44 -05:00
Lioncache 71550e29eb ARMEmitter: Handle SVE integer matrix multiply accumulate 2023-02-09 05:34:11 -05:00
Lioncache 1b8d8f8280 ARMEmitter: Handle SVE2 interleaved XOR category 2023-02-09 05:16:59 -05:00
Lioncache ff6c70f5e1 ARMEmitter: Handle SVE2 bitwise permute category 2023-02-09 05:11:48 -05:00
Lioncache 13ee2b5ec4 ARMEmitter: Handle SVE2 add/sub narrow high part 2023-02-09 05:01:14 -05:00
Ryan Houdek 3c1ba846f7 Merge pull request #2394 from lioncash/cpy
ARMEmitter: Handle CPY (scalar) and CPY (SIMD&FP, scalar)
2023-02-09 01:18:17 -08:00
Lioncache 1b4488e7a3 ARMEmitter: Remove outdated histogram TODO
This was implemented along with histcnt
2023-02-09 04:04:39 -05:00
Lioncache 6b4df4c998 ARMEmitter: Handle CPY (SIMD&FP, scalar) 2023-02-09 03:53:31 -05:00
Lioncache 2a1ef0ba56 ARMEmitter: Handle CPY (scalar) 2023-02-09 03:46:07 -05:00
Ryan Houdek dd2e70e4aa Merge pull request #2393 from lioncash/wide2
ARMEmitter: Handle predicated wide shifts
2023-02-08 23:23:38 -08:00
Lioncache 915a8b23ae ARMEmitter: suffix unpredicated wide shifts
Keeps the naming convention consistent while avoiding clashing
overloads.
2023-02-09 02:00:28 -05:00
Lioncache 6eeafd0724 ARMEmitter: Handle predicated wide shifts 2023-02-09 01:58:23 -05:00
Ryan Houdek e6fc159d88 Merge pull request #2390 from lioncash/ext
OpcodeDispatcher: Handle VEXTRACTF128/VEXTRACTI128
2023-02-08 22:04:22 -08:00
Ryan Houdek 5fd68b6f07 Merge pull request #2392 from lioncash/wide
ARMEmitter: Handle unpredicated wide shifts and unpredicated shifts by immediates
2023-02-08 21:52:23 -08:00
Lioncache 5f80702cf1 ARMEmitter: Handle unpredicated bitwise shift by immediate 2023-02-09 00:15:35 -05:00
Lioncache b45b980b3a ARMEmitter: Handle unpredicated shifts by wide elements 2023-02-09 00:00:17 -05:00
Lioncache f341755e3b Externals: Update fex-gcc-target-test-bins
Allows filtering out the AVX-enabled tests on non-AVX capable systems.
2023-02-08 21:54:35 -05:00
Lioncache c53e7d759b guest_test_runner: Handle AVX-only binary tests
Because the binaries have no metadata, we allow a .json file to be
placed alongside a test indicating required features in a requirements
directory

We also check if the system itself supports those features and run tests
based off of that.
2023-02-08 21:42:04 -05:00
Mai 143ef57141 Merge pull request #2345 from Sonicadvance1/user_sigreturn
Support user supplied signal restorer.
2023-02-08 20:11:47 -05:00
Ryan Houdek 8689038533 Merge pull request #2391 from lioncash/aes
IR: Allow specifying register size for AES enc/dec ops and PCLMUL
2023-02-08 17:11:03 -08:00
Lioncache ade34eeda6 gcc tests: Handle pr57275 test
We now handle all instructions that this uses.
2023-02-08 17:49:29 -05:00
Lioncache 0218c966bd IR: Allow specifying register size for PCLMUL
This will allow us to support 256-bit vector operation in the future.
2023-02-08 16:35:20 -05:00
Lioncache ec5bc9cf3e IR: Allow specifying register sizes for AES enc/dec ops
This will allow us to support operating on 256-bit vectors.

Currently only sets up the bits and pieces on the x86-64 side, since
facilities for testing the 256-bit operations on ARM isn't set up yet.
2023-02-08 16:25:19 -05:00
Lioncache 63bf0d5826 OpcodeDispatcher: Handle VEXTRACTI128 2023-02-08 15:46:27 -05:00
Lioncache 2526fa8b6f OpcodeDispatcher: Handle VEXTRACTF128 2023-02-08 15:40:36 -05:00
Mai ef6f5d2003 Merge pull request #2388 from Sonicadvance1/move_fexbash
FEXBash: Move to Tools folder
2023-02-07 16:18:55 -05:00
Ryan Houdek be02cafb05 FEXBash: Move to Tools folder
Just a cleanup, no functional change.
2023-02-07 07:40:46 -08:00
Ryan Houdek e8fd8ef3b7 Merge pull request #2387 from lioncash/prfx
Arm64/VectorOps: Use movprfx with VBSL
2023-02-06 22:53:41 -08:00
Lioncache d7c6ed842d Arm64/VectorOps: Use movprfx with VBSL
We can use movprfx here to allow compressing the move and bsl operation
together on cpus that can handle it.
2023-02-07 00:44:08 -05:00
Ryan Houdek 86a6118b62 Merge pull request #2386 from lioncash/bsl
VectorOps: Only use VBSL 256-bit path if SVE is present
2023-02-06 21:40:15 -08:00
Lioncache 7a75e43125 VectorOps: Only use VBSL 256-bit path if SVE is present
With this in place, a _VMov isn't necessary for variable blends anymore,
since the vector upper lanes are guaranteed to be zeroed out in the 128-bit case.
2023-02-07 00:13:04 -05:00
Ryan Houdek 4ef3066b69 Merge pull request #2385 from lioncash/vblend
OpcodeDispatcher: Handle VPBLENDVB/VBLENDVPD/VBLENDVPS
2023-02-06 20:25:06 -08:00
Lioncache 88fee019a1 OpcodeDispatcher: Handle VPBLENDVB 2023-02-06 23:04:26 -05:00
Lioncache 5d3141dffc OpcodeDispatcher: Handle VBLENDVPD 2023-02-06 23:04:26 -05:00
Lioncache 94e91565b1 OpcodeDispatcher: Handle VBLENDVPS 2023-02-06 23:04:26 -05:00
Lioncache acbfee55b4 IR: Allow provising register size for VBSL
Necessary, since this will now be used with both 256-bit and 128-bit
registers, rather than just 128-bit.
2023-02-06 23:04:26 -05:00
Lioncache a2481d6892 OpcodeDispatcher: Add helper for AVX variable blends
These will be used by following instruction implementations.
2023-02-06 23:04:26 -05:00
Ryan Houdek e255f1cdef Merge pull request #2383 from lioncash/blend
OpcodeDispatcher: Handle VBLENDPD/VPBLENDW
2023-02-06 18:55:31 -08:00
Ryan Houdek cb3cfed9c2 Merge pull request #2384 from lioncash/sqadd
ARMEmitter: Handle SVE2 saturating add/subtract category
2023-02-06 18:55:23 -08:00
Lioncache 2b12a46d2d ARMEmitter: Handle UQSUBR 2023-02-06 21:29:10 -05:00
Lioncache 8f50109501 ARMEmitter: Handle SQSUBR 2023-02-06 21:29:10 -05:00
Lioncache 1d451b8df1 ARMEmitter: Handle USQADD 2023-02-06 21:29:10 -05:00
Lioncache 49772c6826 ARMEmitter: Handle SUQADD 2023-02-06 21:29:10 -05:00
Lioncache 535a2ab2ba ARMEmitter: Handle UQSUB (vectors, predicated) 2023-02-06 21:29:10 -05:00
Lioncache b8b212719f ARMEmitter: Handle SQSUB (vectors, predicated) 2023-02-06 21:29:10 -05:00
Lioncache be593d43ce ARMEmitter: Handle UQADD (vectors, predicated) 2023-02-06 21:29:10 -05:00
Lioncache a89b7c5dbb ARMEmitter: Handle SQADD (vectors, predicated) 2023-02-06 21:29:07 -05:00
Lioncache c682f51811 OpcodeDispatcher: Handle VPBLENDW 2023-02-06 21:23:30 -05:00
Lioncache 2c7562c54c OpcodeDispatcher: Handle VBLENDPD 2023-02-06 21:03:54 -05:00
Lioncache 8f5ec20cb7 OpcodeDispatcher: Add helper for AVX vector blends 2023-02-06 20:38:35 -05:00
Ryan Houdek 582108a68a Merge pull request #2382 from lioncash/dedup
ARMEmitter: Centralize instruction handling for a few categories
2023-02-06 17:30:09 -08:00
Lioncache 79abe2aa64 ARMEmitter: Simplify bitwise shift by immediate (predicated) category
Centralizes the immediate handling in the encoding helper function.

Lets us move all the asserts there as well.
2023-02-06 20:03:21 -05:00
Lioncache 9f3857b3b0 ARMEmitter: Simplify saturating extract narrow category
Centralizes the immediate handling in the encoding function.
2023-02-06 19:21:42 -05:00
Lioncache 9521638910 ARMEmitter: Simplify bitwise shift right narrow category
Centralizes the immediate handling in one place, making everything much
shorter.
2023-02-06 19:21:39 -05:00
Mai 60b76f53cf Merge pull request #2381 from Sonicadvance1/code_data_header
JIT: Adds a JIT data header and tail.
2023-02-06 18:07:20 -05:00
Mai 5da90aac46 Merge pull request #2378 from Sonicadvance1/fix_emitter_warnings
ARMEmitter: Fixes some warnings that cropped up.
2023-02-06 18:05:56 -05:00
Ryan Houdek ba5ad72ca2 JIT: Adds a JIT data header and tail.
This will be used to store various bits of data about the code going
forward.

Currently unused but that will change as we move forward.
2023-02-06 14:06:37 -08:00
Mai c7c47a827a Merge pull request #2377 from Sonicadvance1/code_data_support
Core: Support Data in JIT buffer header
2023-02-06 16:54:01 -05:00
Mai c4b66b41cd Merge pull request #2379 from Sonicadvance1/rename_fstatat64
Syscalls: Renamed fstatat64 to fstatat_64
2023-02-06 16:47:02 -05:00
Mai 6047ca9fe2 Merge pull request #2376 from Sonicadvance1/minor_flag_opt
Dispatcher: Minor flags optimization
2023-02-06 16:45:45 -05:00
Mai a5762b6faa Merge pull request #2375 from Sonicadvance1/inject_libsegfault
ELFCodeLoader: Adds an option to inject libSegFault
2023-02-06 16:44:52 -05:00
Ryan Houdek 3bc722ca69 Merge pull request #2380 from Joshua-Ashton/directfb_fix
Fix SDL2 directfb includes under Alpine Linux
2023-02-05 18:37:06 -08:00
Joshua Ashton d7d8a4e28a Fix SDL2 directfb includes under Alpine Linux 2023-02-06 02:14:36 +00:00
Ryan Houdek c54c568fef Syscalls: Renamed fstatat64 to fstatat_64
Similar to our other syscall conflicts, musl/Alpine Linux has a global
define that is conflicting with our name here
2023-02-05 18:13:45 -08:00
Ryan Houdek 37421d36e6 ARMEmitter: Fixes some warnings that cropped up. 2023-02-05 18:06:29 -08:00
Ryan Houdek bd86deb9ba Core: Support Data in JIT buffer header
Currently unused (The full data gets thrown away after CompileCode is
called), but allows us to separate code and data in what `CompileCode`
returns.

This will allow us put a header on JIT blocks which will fix a long
outstanding bug where RIP isn't always synchronized on block entry, but
since it only needs to synchronize on signal we can rebuild in the
handler. This future task will remove the `86dec_SynchronizeRIPOnAllBlocks`
config option, but the data will also end up being used for more things
in the future.
2023-02-05 17:55:31 -08:00
Ryan Houdek 2e701fc9e6 Dispatcher: Minor flags optimization
SelectCC shift wasn't necessary since we just need to ensure the final
result is zero when or'd together.

Also operations calculating SF can just use a BFE instead of a shifts
with a constant. BFE by immediate is more efficiently encoded in our IR.
2023-02-04 17:50:36 -08:00
Ryan Houdek 5b97e7f1a0 ELFCodeLoader: Adds an option to inject libSegFault
When used in conjuction with #2345 this is a useful way to enable
libSegFault in applications using application profiles.

Very useful for applications and games that use launcher scripts that
set LD_PRELOAD to nothing prior to launch.

A user was wanting this.
2023-02-04 11:17:26 -08:00
Ryan Houdek 844e27e9ad X86HelperGen: Support fallback sigreturn helpers
For the case that the 32-bit VDSO thunk library isn't available, have a
fallback that can work as well.
Otherwise 32-bit applications will just straight up crash on signal
return.
2023-02-04 10:54:30 -08:00
Ryan Houdek cf147e8ab2 github: Move install step to after the build
Also enable on all builders.
Some tests now rely on 32-bit thunks existing because we need VDSO.
2023-02-04 10:35:07 -08:00
Ryan Houdek e61132b481 VDSOEmu: Handle errors in VDSO
VDSO behaves like a raw syscall which doesn't set errno.
posix tests are testing that errno is set correctly.

Our VDSO handlers weren't wired up to return errors from VDSO correctly.
To handle this we need to have different handlers depending on if the
syscall being used comes from glibc or true VDSO.

This wasn't being uncovered previously since CI wasn't running with VDSO
thunks enabled, but now that it is this needs to be handled or CI will
fail.
2023-02-04 10:35:07 -08:00
Ryan Houdek c58e7a732e X86HelperGen: Remove now unused sigret codegen
This is no longer used so doesn't need to exist.
2023-02-04 10:35:07 -08:00
Ryan Houdek 0538574dd0 Dispatcher: Supports user provided signal restorer
This is required for backtrace to work correctly.
If we are using our custom instruction for returning from a signal, then
backtrace tries to read PC for the sigreturn code and finds our code,
breaking it.

Instead we now /correctly/ support using rt_sigreturn/sigreturn and the
restorer provided from the user.
To facilitate this, we now store a single 64-bit value on the stack to
return our host stack pointer to the correct location from before the
signal.
With cookie checking in place, we can know if an application betrays our
expectations and tries to pass its own signal frames.
If an application in the future /does/ try to pass its own signal
frames, that's unsafe and we cna deal with it then.
2023-02-04 10:35:07 -08:00
Ryan Houdek 1ed546d48f SignalDelegator: Reemit the default signal if it was caught
This fixes a bug where we are falling back down the default signal
delegator after a fatal error.
We need to reraise the event in the case that it didn't come from the
kernel.

Fixes backtrace crashing with incorrect signal when it tries to reraise
the signal that it handled using tgkill.
2023-02-04 10:35:07 -08:00
Ryan Houdek 5ba0053edc VDSOEmulation: Support parsing the 32-bit VDSO symbols
We need to extract the sigreturn handlers and pass them to the FEXCore
signal dispatcher.
2023-02-04 10:35:07 -08:00
Ryan Houdek abb8de0966 VDSO: Add sigreturn functions to VDSO
These need to be bit-exact following exactly what is shown in the
assembly.

libunwind parses where EIP is to see if it is in a stack frame.
Also needsto live in VDSO otherwise backtrace doesn't work.
2023-02-04 10:35:06 -08:00
Ryan Houdek abc596c634 IR: Removes SignalReturn op
This will no longer be used as we are swithing over to using the Linux
system call directly.
2023-02-04 10:35:06 -08:00
Ryan Houdek d107bc9a24 FEXCore: Adds handlers for signal handler returns
Lets the frontend syscall handlers for signal return call the JIT return
handlers directly.
2023-02-04 10:35:06 -08:00
Ryan Houdek a45047bc1e OpDispatcher: Removes SIGRET x86 instruction
We are switching over to syscalls.
2023-02-04 10:35:06 -08:00
Ryan Houdek 1089987a29 Merge pull request #2374 from lioncash/mul
ARMEmitter: Handle SVE SQDMULH/SQRDMULH (vector)
2023-02-04 02:23:59 -08:00
Lioncache 6522d3d6e4 ARMEmitter: Move 128-bit check into SVE2IntegerMultiplyVectors
Simplifies the amount of code needed. Also we can remove some
unnecessary namespacing to make these a little faster to grok when
looking at them.
2023-02-04 05:08:43 -05:00
Lioncache d65fcf7bb8 ARMEmitter: Handle SVE SQRDMULH (vectors) 2023-02-04 05:06:56 -05:00
Lioncache 16e0f628cd ARMEmitter: Handle SVE SQDMULH (vectors) 2023-02-04 05:05:27 -05:00
Ryan Houdek d81097482d Merge pull request #2373 from lioncash/vl
ARMEmitter: Handle ADDVL/ADDPL and RDVL
2023-02-04 01:34:14 -08:00
Lioncache 75bc997ab7 ARMEmitter: Handle RDVL 2023-02-04 01:11:57 -05:00
Lioncache 600e8749d7 ARMEmitter: Handle ADDPL 2023-02-04 01:05:25 -05:00
Lioncache fc3863f444 ARMEmitter: Handle ADDVL 2023-02-04 01:03:24 -05:00
Ryan Houdek 347abf09ef Merge pull request #2372 from lioncash/mla
ARMEmitter: Handle MLA/MLS (vector) and MAD/MSB
2023-02-03 21:09:26 -08:00
Lioncache 7442ef3a83 ARMEmitter: Handle SVE MSB 2023-02-03 23:14:10 -05:00
Lioncache 3783ad8dd1 ARMEmitter: Handle SVE MAD 2023-02-03 23:13:16 -05:00
Lioncache 65ad916984 ARMEmitter: Handle SVE MLS (vectors) 2023-02-03 23:06:42 -05:00
Lioncache d491ce7125 ARMEmitter: Handle SVE MLA (vectors) 2023-02-03 23:05:21 -05:00
Ryan Houdek c0bc5d9748 Merge pull request #2371 from lioncash/mul
ARMEmitter: Handle SVE predicated mul/div and finish off integer reduction category
2023-02-03 19:36:57 -08:00
Lioncache 3f6edf7b5a ARMEmitter: Clarify SVEReductionOperation as working on integer ops 2023-02-03 22:01:34 -05:00
Lioncache 5563a51b84 ARMEmitter: Allow 64-bit variants of min/max reduction
The instructions allow specifying 64-bit element sizes.

With this, we can also completely remove the size checking from the
functions, since the general SVE integer reduction operation already
checks for invalid sizes for us.
2023-02-03 22:00:56 -05:00
Lioncache dac075b871 ARMEmitter: Move min/max reduction over to generic reduction helper
Also enforces the use of a VRegister for the destination argument like
the manual.
2023-02-03 21:45:58 -05:00
Lioncache db8317caf8 ARMEmitter: Handle SVE ANDV (predicated) 2023-02-03 21:31:41 -05:00
Lioncache 3508f7a667 ARMEmitter: Handle SVE EORV (predicated) 2023-02-03 21:30:48 -05:00
Lioncache 02861f41eb ARMEmitter: Handle SVE ORV (predicated) 2023-02-03 21:26:06 -05:00
Lioncache 51c9f70904 ARMEmitter: Handle SVE UADDV (predicated) 2023-02-03 21:05:00 -05:00
Lioncache 4f9530cec3 ARMEmitter: Handle SVE SADDV (predicated) 2023-02-03 21:02:45 -05:00
Lioncache ecd711e691 ARMEmitter: Handle SVE UDIVR (predicated) 2023-02-03 20:49:57 -05:00
Lioncache 870115dd5d ARMEmitter: Handle SVE SDIVR (predicated) 2023-02-03 20:49:57 -05:00
Lioncache 848e5561ce ARMEmitter: Handle SVE UDIV (predicated) 2023-02-03 20:49:57 -05:00
Lioncache db8a9bb5cf ARMEmitter: Handle SVE SDIV (predicated) 2023-02-03 20:49:54 -05:00
Lioncache f8a1c43c06 ARMEmitter: Handle SVE UMULH (predicated) 2023-02-03 20:27:48 -05:00
Lioncache 3a98190119 ARMEmitter: Handle SVE SMULH (predicated) 2023-02-03 20:25:44 -05:00
Ryan Houdek 3d930ee4b8 Docs: Update for release FEX-2302 2023-02-03 17:24:08 -08:00
Lioncache 19ad19193e ARMEmitter: Handle SVE MUL (predicated) 2023-02-03 20:16:24 -05:00
Mai a7aeb4af7f Merge pull request #2368 from Sonicadvance1/fexrootfsfetcher_first_option
FEXRootFSFetcher: Support option to auto select first distro
2023-02-03 17:31:31 -05:00
Mai d2d528222c Merge pull request #2370 from Sonicadvance1/remove_pollremove
FEXServer: Remove POLLREMOVE usage
2023-02-03 17:30:45 -05:00
Ryan Houdek 6598eeee92 FEXServer: Remove POLLREMOVE usage
Fixes this file compiling on musl at least.

POLLREMOVE usage here is technically incorrect as it shouldn't be OR'd
with other flags.
But it is also additionally wrong here because the Linux kernel doesn't
even support this flag anymore, so it doesn't change behaviour.
2023-02-03 13:35:30 -08:00
Ryan Houdek c42fd4122b FEXRootFSFetcher: Support option to auto select first distro
Fixes #2356

In the case of the `-y` option being used, it will auto say "yes", but
when presented with the distro list this doesn't work. This happens when
used on a distro that doesn't have an exact match to what we provide.

Exposes a new option that when presented the distro list, auto select
the first option. Solving this issue when automating.
2023-02-03 10:48:23 -08:00
Ryan Houdek 9d33bba1c8 Merge pull request #2366 from lioncash/addsub
ARMEmitter: Handle integer add/subtract vectors (predicated) instruction class
2023-02-03 10:31:56 -08:00
Ryan Houdek a899f9f824 Merge pull request #2367 from lioncash/rmif
ARMEmitter: Handle RMIF, SETF8/SETF16
2023-02-02 20:55:54 -08:00
Lioncache 8c09356bd7 ARMEmitter: Handle SETF16 2023-02-02 23:27:40 -05:00
Lioncache 50bcc1b96f ARMEmitter: Handle SETF8 2023-02-02 23:26:07 -05:00
Lioncache 36831ebc37 ARMEmitter: Handle RMIF 2023-02-02 23:18:22 -05:00
Lioncache 44f5d788c8 ARMEmitter: Handle SUBR (vector, predicated) 2023-02-02 21:44:44 -05:00
Lioncache a42ae7d385 ARMEmitter: Handle SUB (vector, predicated) 2023-02-02 21:42:57 -05:00
Lioncache 5cf9bb2613 ARMEmitter: Handle ADD (vector, predicated) 2023-02-02 21:41:12 -05:00
Ryan Houdek 1cda029ed7 Merge pull request #2365 from lioncash/reduce
ARMEmitter: Handle SVE floating-point recursive reduction
2023-02-02 17:59:29 -08:00
Lioncache 4001dc1219 ARMEmitter: Handle SVE FMINV 2023-02-02 20:42:54 -05:00
Lioncache f77de7f283 ARMEmitter: Handle SVE FMAXV 2023-02-02 20:41:10 -05:00
Lioncache ac9f9d291b ARMEmitter: Handle SVE FMINNMV 2023-02-02 20:35:43 -05:00
Lioncache 25f97065df ARMEmitter: Handle SVE FMAXNMV 2023-02-02 20:33:37 -05:00
Lioncache 6fcbce0c52 ARMEmitter: Handle SVE FADDV 2023-02-02 20:28:11 -05:00
Ryan Houdek 2c9f99e5d6 Merge pull request #2364 from lioncash/hist
ARMEmitter: Add a few missing instructions
2023-02-02 13:31:08 -08:00
Lioncache 4c647a2e02 ARMEmitter: Handle NMATCH 2023-02-02 15:40:34 -05:00
Lioncache d0f00d53d3 ARMEmitter: Handle MATCH 2023-02-02 15:40:34 -05:00
Lioncache 6174437667 ARMEmitter: Handle SVE FCMLA 2023-02-02 15:40:34 -05:00
Lioncache 9c762861f6 ARMEmitter: Handle SVE FCADD 2023-02-02 15:40:34 -05:00
Lioncache 448785e693 ARMEmitter: Handle HISTSEG 2023-02-02 15:40:26 -05:00
Lioncache f8c68acc09 ARMEmitter: Handle HISTCNT 2023-02-02 15:40:18 -05:00
Ryan Houdek 65971effc7 Merge pull request #2363 from Sonicadvance1/fix_relative_execve
Config: Fix relative execve applications.
2023-02-02 05:01:32 -08:00
Ryan Houdek d5e7af5b96 Config: Fix relative execve applications.
I made the assumption from some bad historical knowledge that the kernel
will canonicalize relative filenames and symlinks for applications that
execute through execve.

This turns out to not be true. In fact it passes pathname untouched to
the interpreter. So we need to do an additional fix up on relative paths
to ensure glibc doesn't break.

Fixes a major bug that breaks a bunch of games.
2023-02-02 04:42:29 -08:00
Ryan Houdek 62e6ada112 Merge pull request #2362 from lioncash/blendd
OpcodeDispatcher: Handle VPBLENDD/VBLENDPS
2023-02-01 20:16:09 -08:00
Lioncache 2e232ac3dd OpcodeDispatcher: Handle VBLENDPS 2023-02-01 20:36:23 -05:00
Lioncache 88ff0db12a OpcodeDispatcher: Handle VPBLENDD 2023-02-01 20:36:15 -05:00
Ryan Houdek 9d35bc01c7 Merge pull request #2361 from Sonicadvance1/fix_global_symbol_overrides
Thunks: Fixes host symbol overrides
2023-02-01 07:09:25 -08:00
Ryan Houdek c8c1ebad01 unittests: Updates tests to have a dlsym_default function 2023-02-01 06:48:26 -08:00
Ryan Houdek cb5573995d Thunks: Fixes host symbol overrides
1) The host library needs to be loaded in the global namespace.

2) We need to use `RTLD_DEFAULT` instead of querying the object
   directly.

We need to load the host library in the global namespace so the symbols
end up in the global symbol table. This follows how all these symbols
/usually/ get loaded. Either by linking directly to the library or how
loaders will end up loading these.

We need to use RTLD_DEFAULT to follow symbol overriding rules that tend
to occur. For example, MangoHUD will LD_PRELOAD a library that provides
GLX and EGL symbols. Which FEX's thunk libraries need to pick up this
override.
If we are querying the host library directly then we fail to pickup
these overrides, thus breaking MangoHUD and other overlays.
2023-02-01 06:47:41 -08:00
Ryan Houdek fa1193f14c Merge pull request #2344 from Sonicadvance1/siginfo_32
FEXCore: Fixup 32-bit signal handling
2023-01-31 20:26:36 -08:00
Ryan Houdek 9a318cad95 Merge pull request #2360 from lioncash/ravd-adj
Arm64/VectorOps: Clamp shift amount to esize-1 for VSShr
2023-01-31 20:25:46 -08:00
Lioncache 4177d5c185 Arm64/VectorOps: Clamp shift amount to esize-1 for VSShr
Makes the behavior consistent with the x86 JIT.

We need to treat values larger than 31 as if they were 31 bit shifts in
order to handle sign-extending behavior properly.
2023-01-31 22:53:51 -05:00
Ryan Houdek fe79f61fc3 Merge pull request #2359 from lioncash/ravd
OpcodeDispatcher: Handle VPSRAVD
2023-01-31 18:32:48 -08:00
Lioncache d5316c8c7e OpcodeDispatcher: Handle VPSRAVD 2023-01-31 17:31:24 -05:00
Lioncache cc65f3e788 Arm64/VectorOps: Implement VSShr
This will be used for implementing VPSRAVD
2023-01-31 17:31:20 -05:00
Mai 787b6895e8 Merge pull request #2337 from Sonicadvance1/optimize_frontend
Frontend: Various optimizations
2023-01-31 14:46:11 +00:00
Mai 7be2e1ad34 Merge pull request #2330 from Sonicadvance1/implement_flushes
OpDispatcher: Adds support for CLWB and CLFLUSHOPT
2023-01-31 04:01:26 +00:00
Mai 9403c662a3 Merge pull request #2320 from Sonicadvance1/remove_numargs
IR: Removes NumArgs member from IR ops
2023-01-31 04:00:57 +00:00
Ryan Houdek 15f2b30a5b FEXLinuxTests: Adds 32-bit signal tests 2023-01-30 13:30:15 -08:00
Ryan Houdek d75e1f996f FEXCore: Fixup 32-bit signal handling
Follow-up to #2327.

Split off from #2176 and improved.

32-bit signals are a bit more complex than 64-bit due to behaviour
changing depending on if `rt_sigaction` and `sigaction` syscall is used
and if `SA_SIGINFO` is passed in to the flags.

With `SA_SIGINFO` used, both turn in to an `RT` frame, which is encoded
differently than without `SA_SIGINFO`.
Additionally 32-bit signals support both regular Linux stack ABI and
`regparm(3)` ABI.

Without `SA_SIGINFO` then `siginfo_t` is removed from the signal handler
arguments, but most of the rest still remains.
Also two of the arguments to the signal handler are forced to be nullptr
with `regparm(3)`.
2023-01-30 13:30:15 -08:00
Ryan Houdek 14fe95bd14 IR: Removes NumArgs member from IR ops
Split off from #2243 to remove each member individually.

Shaves 8-bits off of each IR op.
No need to cart around this data when it is constant for each operation.
Especially since most optimization passes don't need the data anyway.

Needed to add a new `GetRAArgs` to get the number of SSA arguments that
get RA versus `GetArgs` which returns all SSA arguments the IR operation
owns. This is what was causing #2243 to fail CI since it needs to know
the difference in some places.
2023-01-30 11:53:05 -08:00
Ryan Houdek 65b6b6d5dd Merge pull request #2355 from Sonicadvance1/siginfo_64
Dispatcher: Extract 64-bit signal frame save and restore
2023-01-30 11:50:03 -08:00
Mai f8e762fcfb Merge pull request #2319 from Sonicadvance1/remove_has_dest
IR: Remove HasDest member
2023-01-30 16:25:09 +00:00
Ryan Houdek 9cfd169fb8 Dispatcher: Extract 64-bit signal frame save and restore
Stripped from #2344 at request to ensure 64-bit code hasn't changed in a
meaningful way. So that PR can focus on 32-bit.
2023-01-27 01:17:32 -08:00
Ryan Houdek 3d29dac1b1 Merge pull request #2354 from neobrain/fix_single_line_shebang
Syscalls: Fix out-of-bounds read when handling single-line shebang files
2023-01-26 02:38:38 -08:00
Tony Wasserka 94ef3729bf Syscalls: Avoid unnecessary string copies and clean up error handling 2023-01-26 11:20:29 +01:00
Tony Wasserka 420c4ca08f Syscalls: Fix out-of-bounds read when handling single-line shebang files
string::find() returns npos (-1) if the given character was not found, so
it can't be used to construct a string like this. Luckily, the use of
std::span allows this code to be written such that it's both correct and
simpler than before.
2023-01-26 11:20:29 +01:00
Mai 477d4b6de8 Merge pull request #2353 from Sonicadvance1/fix_shebang_execve
Linux: Fixes shebang file execution
2023-01-26 05:19:45 +00:00
Ryan Houdek 2b318d276f Linux: Fixes shebang file execution
Somewhere during the refactoring/review process, failed to strip the
shebang prefix off of the arguments.
Causing shebang files to always fail as if the file never existed.

Fixes steam execution.
2023-01-25 20:25:51 -08:00
Mai da88c68e12 Merge pull request #2332 from Sonicadvance1/emitter_test_ci
Github: Add ARM emitter tests to CI
2023-01-25 23:47:39 +00:00
Mai 7f6a620c9e Merge pull request #2349 from Sonicadvance1/virtual_mem_size_32bit
Core: Adjust virtual memory size for 32-bit
2023-01-24 21:12:36 +00:00
Mai 1e90ebb400 Merge pull request #2323 from Sonicadvance1/pool_inline_constants
ConstProp: Pool inline constants
2023-01-24 21:11:56 +00:00
Ryan Houdek c6d46801ad ConstProp: Pool inline constants
In large blocks we can be generating a ton of inline constants. But in
most cases these end up being 0, 1, or (1 << N).
Add these to a map and reuse if possible. Makes some IR blocks
significantly smaller for later optimization passes.
2023-01-24 12:58:29 -08:00
Mai afaff9293b Merge pull request #2316 from Sonicadvance1/fix_negative_ficomi_f64
X87_F64: Fixes FICOM
2023-01-24 17:31:05 +00:00
Ryan Houdek dcce9add60 Merge pull request #2334 from Sonicadvance1/fix_execveat
FEXLoader: Adds support for execveat with AT_EMPTY_PATH
2023-01-23 02:32:15 -08:00
Ryan Houdek 472675d471 FEXLoader: Adds support for execveat with AT_EMPTY_PATH
Fixes #2136

This is a fairly tricky edge case to support with FEX.
If execveat is used with AT_EMPTY_PATH then the application can pass an
FD to execve instead of a filename. This includes FDs that have been
deleted from the disk so the child process can't open it by filename
anymore.

To work around this limitation, we need to pass the FD to the new FEX
process and open it directly, similar to how binfmt_misc works with FDs.
The FD will get passed through environment variables, which the new
process will check for and then remove the variable from the
environment.

Lots of prickly edge cases to support here.

Without binfmt_misc:
- Passes the FD to FEXLoader directly.
  - Requires duplicating the FD if it has O_CLOEXEC on the FD.

With binfmt_misc:
- Shebang file, pass directly to FEXLoader, just like without binfmt.
- x86 ELF Files, rely on the kernel's binfmt_misc support here.
- Unsupported ELF files, let kernel handle it through binfmt_misc

Argument handling:
- The application can pass in no arguments.
  - Means our application configurations were failing to find a config
  - Also various checks in the frontend were failing.
  - If opened through an FD, find the symlink for that FD for the
    application configuration instead.

Side note:
Fixed a performance issue in execve where when we were checking for file
format support. Either ELF or Shebang files, we were reading the /whole/
file upfront. We only need to read a header worth of ELF files, and only
257 bytes if it is potentially a shebang file. Should dramatically
reduce some application's execve times.
2023-01-23 02:06:50 -08:00
Mai a28039f7cd Merge pull request #2350 from Sonicadvance1/optimize_dispatcher_slightly
Arm64: Merge two loads in to an LDP
2023-01-23 08:35:13 +00:00
Mai f8d56a8170 Merge pull request #2351 from Sonicadvance1/support_long_address_generation
ARMEmitter: Support helper for long address generation
2023-01-23 02:05:06 +00:00
Ryan Houdek db4cb497e0 unittests: New emitter tests for LongAddressGen 2023-01-22 16:03:17 -08:00
Ryan Houdek a823d918c2 ARMEmitter: Support helper for long address generation
The current separated adr and adrp handlers are difficult to use if you
don't know if the resulting address is going to be within 1MB or 4GB.

Adds a `LongAddressGen` helper that will generate the various pieces of
code that will need to be emitted.

Backward labels:
 - Can generate three different code segments depending on distance to
   label
   - adr if label is within 1MB
   - adrp if label is 4K page aligned and within 4GB
   - adrp+add if label is within 4GB

Forward labels:
- Can generate three different code segments depending on distance to
  label
  - nop+adr if label is within 1MB
  - nop+adrp if label is 4K page aligned and within 4GB
  - adrp+add if label is within 4GB

There is still the limitation that this can't generate addresses to
labels that are >4GB away. Which is fine.
2023-01-22 16:03:17 -08:00
Ryan Houdek 79b8442dbc FHU: Add helpers for symlink checking 2023-01-22 14:23:55 -08:00
Ryan Houdek 7bf1742434 Arm64: Merge two loads in to an LDP
We can do a single LDP upfront when loading from the code cache, which
saves an instruction and one LDP costs the same as a single LDR.

Itty bitty optimization in the hot dispatcher.
2023-01-20 19:19:47 -08:00
Ryan Houdek 09f720d1cc Core: Adjust virtual memory size for 32-bit
We only need a 32-bit virtual memory size when running a 32-bit
application.

Just lowers some virtual memory space that we need to allocate.
2023-01-20 19:18:52 -08:00
Ryan Houdek 28dd94642a Merge pull request #2339 from Sonicadvance1/optimize_loadfile
FileLoading: Optimize FileLoad
2023-01-20 13:35:14 -08:00
Ryan Houdek 8dae785e9e Merge pull request #2327 from Sonicadvance1/siginfo
Dispatcher: Fixes x86-64 SA_SIGINFO generation
2023-01-20 13:34:56 -08:00
Ryan Houdek 7ef9189910 FEXLinuxTests: Fixup the tests
These were having some issues executing. 32-bit ones were getting
skipped even.
2023-01-20 12:55:38 -08:00
Ryan Houdek c5d0fe6999 FEXLinuxTests: Adds 64-bit siginfo test. 2023-01-20 11:08:57 -08:00
Ryan Houdek ac1bf0683d FEXLinuxTests: Adds support for 64-bit only tests 2023-01-20 11:08:57 -08:00
Ryan Houdek 7897803753 Dispatcher: Fixes x86-64 SA_SIGINFO generation
Pulled from #2176.

On x86-64 the SA_SIGINFO sa_flag is actually a no-op. It is always used
even if not set.

Ensure that we setup siginfo_t regardless of flag being set.

On 32-bit x86 this still needs to be adhered to.

Little side bits that don't change anything
- EFLAGS is passed in signfo correctly.
- User provided restorer usage locations is documented but not
  implemented.
2023-01-20 11:08:57 -08:00
Ryan Houdek 40e5690e3a Dispatcher: Encode eflags in uc_mcontext
We were missing this.
2023-01-20 11:08:57 -08:00
Ryan Houdek 8d0329ddaf Merge pull request #2348 from stevenvandenbrandenstift/fixupTgkill
fix ifdef to use HAS_SYSCALL_TGKILL for tgkill as it was intented
2023-01-19 13:42:28 -08:00
Steven Vanden Branden fd5bfd9e40 fix ifdef to use HAS_SYSCALL_TGKILL for tgkill as it was intented 2023-01-19 22:29:12 +01:00
Mai 5fd8fdbf5c Merge pull request #2346 from Sonicadvance1/armemitter_warnings
ARMEmitter: Removes some warnings that cropped up
2023-01-19 19:04:42 +00:00
Mai a486797e59 Merge pull request #2347 from Sonicadvance1/fix_jitsymbols
JitSymbols: Fixes file opening and writing
2023-01-19 19:04:22 +00:00
Ryan Houdek 6193bddaa5 JitSymbols: Fixes file opening and writing
We shouldn't use O_EXCL, since we need to overwrite previous entry PIDs
if they happen to exist. The kernel ensures that PIDs don't overlap, but
in some kernel configurations PIDs are aggressively reused, resulting
in O_EXCL quickly hitting an issue when writing stale files.

Additionally O_DIRECT, this doesn't allow us to write to files, so all
write functions were failing.

Additionally use O_APPEND, we are only ever appending, so let the kernel
know.

Additionally use O_TRUNC, in the case that a stale perf file exists,
this will immediately truncate the file to zero.
2023-01-19 01:13:44 -08:00
Ryan Houdek 87609b2938 CPUID: Only expose CLWB if supported 2023-01-18 17:56:21 -08:00
Ryan Houdek bd36bd55ca HostFeatures: Adds support for querying CLWB availability
The x86 runner doesn't support CLWB natively.
2023-01-18 17:56:21 -08:00
Ryan Houdek 965c6ff6cc unittests: Adds tests for CLWB and CLFLUSHOPT 2023-01-18 17:56:21 -08:00
Ryan Houdek 7450b5d406 OpDispatcher: Adds support for CLWB and CLFLUSHOPT
These fairly trivially map to AArch64 operations.

CLWB just maps to `dc cvac`
CLFLUSHOPT just maps to `dc civac` without the final `dsb`.

Also captures if something tries using XSAVEOPT without checking.
2023-01-18 17:56:21 -08:00
Ryan Houdek 4582c8d380 IR: Adds support for CacheLineClean and non-serializing clear
These will be used in the next commit.
2023-01-18 17:56:21 -08:00
Ryan Houdek b621b61d92 HostRunner: Fix for new xbyak 2023-01-18 17:53:52 -08:00
Ryan Houdek 7d3e7d2ab4 Externals: Update xbyak to v6.68 2023-01-18 17:53:52 -08:00
Ryan Houdek c4a1d7e0cd FileLoading: Optimize FileLoad
Optimize `FileLoad` by not using fstream.
For some reason fstream is just really bad.

Switching over to raw pread cuts the amount of time it takes to read
files by a quarter of CPU time.
2023-01-18 17:51:38 -08:00
Ryan Houdek bf4c5797db ARMEmitter: Removes some warnings that cropped up 2023-01-18 17:48:44 -08:00
Mai 4aa984aed9 Merge pull request #2322 from Sonicadvance1/opdispatcher_helpers
OpDispatcher: Fixes a few missing GPR/XMM helper usages
2023-01-19 00:05:11 +00:00
Mai 95e544c840 Merge pull request #2342 from Sonicadvance1/more_asimd_ops_pt2
ArmEmitter: Adds two more classes of ASIMD instructions
2023-01-18 20:33:40 +00:00
Mai 81e0ac7e0b Merge pull request #2331 from Sonicadvance1/more_asimd_ops
ArmEmitter: Adds three more classes of ASIMD instructions
2023-01-18 20:32:46 +00:00
Mai f8d92aa121 Merge pull request #2329 from Sonicadvance1/fix_cache_invalidation
Arm64: Fixes incorrect operation for CacheLineClear
2023-01-18 20:29:44 +00:00
Mai ee58c5de1d Merge pull request #2315 from Sonicadvance1/add_negative_unittests
unittests: Adds negative integer x87 tests
2023-01-18 20:27:59 +00:00
Mai 565ed450aa Merge pull request #2310 from Sonicadvance1/aarch64_move_to_switch
Arm64: Use switch statement for op handlers instead of jump table
2023-01-18 20:26:16 +00:00
Mai 90bcb8c70b Merge pull request #2309 from Sonicadvance1/remove_header
Emitter: Remove unused header
2023-01-18 20:25:07 +00:00
Ryan Houdek bbf9198cba Merge pull request #2324 from Sonicadvance1/jemalloc_disable_16k
External: Update JEMalloc to disable 16k pages
2023-01-17 12:56:23 -08:00
Ryan Houdek 9c93c6ffcd Merge pull request #2317 from Sonicadvance1/fix_spill_register
Arm64: Fix SpillRegister C&P error
2023-01-17 12:56:14 -08:00
Ryan Houdek 8974509c52 Merge pull request #2343 from Sonicadvance1/fexinterpreter_heartburn
FEXLoader: Build FEXInterpreter and FEXLoader independently
2023-01-17 02:25:54 -08:00
Ryan Houdek abc5aa6aa0 FEXLoader: Build FEXInterpreter and FEXLoader independently
This is causing some heartburn with the hardlink.

- Removes some termux cmake list hacking.
- Removes the need to do post-install packaging fixups when hardlinks get dropped.
- Removes a custom uninstall target that was necessary before.
2023-01-16 13:08:43 -08:00
Ryan Houdek a668c34dec unittests: Adds tests for two new subclasses 2023-01-15 20:16:10 -08:00
Ryan Houdek 0ef8574a56 ARMEmitter: Adds two instruction classes 2023-01-15 20:16:10 -08:00
Ryan Houdek e66ad12fa8 unittests: Adds unittests for added ops 2023-01-15 20:16:10 -08:00
Ryan Houdek f1e1eaa8e5 ArmEmitter: Adds four missing ASIMD Shift by Imm ops 2023-01-15 20:16:10 -08:00
Ryan Houdek f614fc6fac Merge pull request #2338 from Sonicadvance1/optimize_cpuid
CPUID: Optimize initialization
2023-01-14 14:15:35 -08:00
Ryan Houdek d9a1bb9c35 CPUID: Optimize initialization
Map lookup was quite expensive, switched over to three small vectors
that are constexpr instead.

Some file querying and parsing was fairly slow as well. Optimized to
make that CPU time to go away.

This improves initialization time of CPUIDEmu by 33%
2023-01-14 13:37:26 -08:00
Ryan Houdek f36bbf0f59 FileLoading: Add a very quick fixed size small file loading helper
If we have a fixed file to read, we can read it in three syscalls.
fstream is...weirdly slow in the other implementation.

Theoretically in the future this can be improved to a single syscall if
the `readfile` syscall ends up in upstream Linux.
2023-01-14 13:33:32 -08:00
Ryan Houdek f5e97f3542 Merge pull request #2333 from Sonicadvance1/use_rng_syscall
ELFCodeLoader: Don't use std::random_device for RNG
2023-01-14 12:27:05 -08:00
Ryan Houdek 7664359410 Merge pull request #2326 from Sonicadvance1/debug_cookie
MContext: Insert a stack cookie with assertions enabled
2023-01-14 12:13:54 -08:00
Ryan Houdek df8704215b Merge pull request #2328 from Sonicadvance1/update_install_script_links
Scripts: Update InstallFEX.py rootfs links
2023-01-14 12:13:40 -08:00
Ryan Houdek 34e1ba6129 Merge pull request #2318 from Sonicadvance1/fix_jit_symbol_crash
JitSymbols: Fixes a crash that can occur
2023-01-14 12:07:22 -08:00
Ryan Houdek fcddf86352 ELFCodeLoader: Don't use std::random_device for RNG
Fixes #2095

std::random_device can fail to find a random source and throw an assert
during initialization.

The constructor allows you to provide a token string to select an
explicit random source, but this is c++ library specific and will still
assert if the source isn't found. Also specific tokens are very much
target and library specific, so it is unsafe to use.

In the libstdc++ case, the default implementation will try to open
`/dev/urandom` which might not be available on all targets.

Instead of relying on this C++ object, use the `getrandom` syscall
directly to generate our RNG used in the ELFCodeLoader.
Resolves any sort of asserting case, blocks on RNG generation, and is
guaranteed to be available since this syscall has been available since a
very old kernel version.
2023-01-14 12:06:12 -08:00
Ryan Houdek d69afaf925 MContext: Insert a stack cookie with assertions enabled
Pulled from #2176.
Ensures that when we are handling signals we are actually restoring a
stack state that is what we expect..

While this could randomly intersect with other stack data, it is highly
unlikely and will still capture incorrect stack frames otherwise.

Keeps it out of release build to ensure we aren't sticking random data
in the stack when it wouldn't have even been checked.
2023-01-14 11:59:10 -08:00
Ryan Houdek dc5e739628 JitSymbols: Fixes a crash that can occur
When a process is in the process of forking and getting ready for
execve, it is common practice to do a `close_range` or close loop to
close all file descriptors before the execve.

This is a security/sanitization feature to ensure that FDs aren't leaked
to the child process. While it is more reliable to have these FDs opened
with O_CLOEXEC, people get it wrong all the time so this feature has
been put in place. Both python and glibc wrappers for launching
applications do this.

The problem with this for FEX is that we were using a FILE handle for
emitting JIT symbols to the perf file. When the underlying FD is ripped
out from under the FILE handle, it throws an assert that we can't
recover from.

Switching to a raw FD and checking to ensure the FD is still open on
writes means that we can safely stop JIT symbol logging when a process
is closing FDs under us.

Fixes a crash in Steam early startup where a python script is run for
checking if packages are installed.
2023-01-14 11:56:12 -08:00
Ryan Houdek b57a8ac086 IR: Update tests for new GPR offsets 2023-01-13 19:35:12 -08:00
Ryan Houdek 7696641b4c Frontend: Add some switch statement hints for most command paths. 2023-01-13 19:23:31 -08:00
Ryan Houdek b9fedfff7c Frontend: Optimize that Dst RAX/RCX and REX in byte is mutually exclusive. 2023-01-13 19:23:31 -08:00
Ryan Houdek 30dc92cdfd Frontend: Set Is8Bit{Dest,Src} immediately rather than query flags after the fact. 2023-01-13 19:23:31 -08:00
Ryan Houdek ed0d46e51a Frontend: Optimize NormalOp to most likely pick a NormalOp. It's the most common. 2023-01-13 19:23:30 -08:00
Ryan Houdek a065849e39 Frontend: Only add contained code pages at the end 2023-01-13 18:27:27 -08:00
Ryan Houdek c61ce1fe11 Frontend: Remove unnecessary checks 2023-01-13 18:26:51 -08:00
Ryan Houdek c98816b350 OpDispatcher: Make ReadByte check only happen with assertions enabled. 2023-01-13 18:26:14 -08:00
Ryan Houdek 2866dda73f Frontend: Optimize MapModRMToReg and MapVEXToReg 2023-01-13 18:24:36 -08:00
Ryan Houdek 47bd119c9c OpDispatcher: Fix MOVSeg from previous reordering 2023-01-13 18:24:20 -08:00
Ryan Houdek a5cc2536cb X86Enums: Sort GPRs by their encoding order. 2023-01-13 17:54:56 -08:00
Ryan Houdek 130dcb1704 HarnessHelper: Ensure placement of greg offsets 2023-01-13 17:54:18 -08:00
Ryan Houdek 777390c62a Github: Add ARM emitter tests to CI 2023-01-12 13:55:48 -08:00
Ryan Houdek fb3c8b3491 CMake: Add an option for compiling vixl disassembler 2023-01-12 13:55:00 -08:00
Ryan Houdek c229f906f8 External: Update vixl 2023-01-12 13:54:20 -08:00
Ryan Houdek d787a38744 ArmEmitter: Adds three more classes of ASIMD instructions
Adds three classes:
- Advanced SIMD three same (FP16)
- Advanced SIMD two-register miscellaneous (FP16)
- Advanced SIMD three-register extension

A handful of the three-register extension unit tests are disabled
because the vixl disassembler doesn't support them.

Only six more classes of ASIMD operations remaining once this is merged.
2023-01-12 13:38:48 -08:00
Ryan Houdek b2f7f526f8 Arm64: Fixes incorrect operation for CacheLineClear
CIVAU does Clean+Invalidate to `Point Of Unification`
CIVAC does Clean+Invalidate to `Point of Coherency`

`Point of Unification` means to L2/L3, so unification of core
visibility.

`Point of Coherency` means SLC/RAM, All cores, DNA engines, etc must be
coherent.
2023-01-11 19:53:33 -08:00
Ryan Houdek ab512b6ffa Scripts: Update InstallFEX.py rootfs links
This was never updated for 22.10, Updated list.
2023-01-10 23:44:34 -08:00
Ryan Houdek 1521e0a248 Merge pull request #2325 from cobalt2727/patch-1
fix tgkill
2023-01-10 20:07:50 -08:00
cobalt2727 0f131c4c1a fix tgkill
long time no see!
2023-01-10 22:21:25 -05:00
Ryan Houdek 716cafe6f7 External: Update JEMalloc to disable 16k pages
When tinkering I had enabled 16k page support in jemalloc.
This broke pressure-vessel/Proton executing. Back it back down to 4k page size
to fix this.

We'll need to come back to this to see if enabling this can be done
without breaking these projects.
2023-01-09 18:10:46 -08:00
Ryan Houdek bd55ed51b0 OpDispatcher: Moves a few missing XMM loadstores to helper usage
These were missed initially, these need to all be using the helper for
future optimizations.
2023-01-09 08:23:26 -08:00
Ryan Houdek 6d912be31e OpDispatcher: Moves a few missing GPR loadstores to helper usage
These were missed initially, these need to all be using the helper for
future optimizations.
2023-01-09 08:23:26 -08:00
Ryan Houdek 632add660c Merge pull request #2321 from CallumDev/f64-fprem-fix
Fix FPREM flags calculation in F64
2023-01-09 04:06:10 -08:00
CallumDev 806587d6ae Fix FPREM flags calculation in F64 2023-01-09 22:21:23 +10:30
Ryan Houdek 5c98db5f47 IR: Remove HasDest member
Split off from #2243 to remove each member individually.

IR ops are hardcoded by operation to have a destination or not.
No need to have each operation have a boolean for determining if the
operation has a destination or not.

The number of places things need to know if the operation has
destination or not is better served by using a lookup instead.
2023-01-08 17:57:47 -08:00
Ryan Houdek 4daf2f0793 Jit64: Fixes incorrect sign extension
We were accidentally zero extending.
2023-01-08 13:29:38 -08:00
Ryan Houdek 676cf59198 Jit64: Fixes incorrect sign extension
We were accidentally zero extending.
2023-01-08 13:28:17 -08:00
Ryan Houdek 5aacdd744c Arm64: Fixes incorrect sign extension
We were accidentally zero extending.
2023-01-08 13:25:25 -08:00
Ryan Houdek c3c68afc3c Arm64: Fixes incorrect sign extension
We were accidentally zero extending.
2023-01-08 13:23:40 -08:00
Ryan Houdek c7262120a6 Arm64: Fix SpillRegister C&P error
Was using the wrong sized registers in spill which was breaking Steam.
Oops.
2023-01-08 12:47:02 -08:00
Ryan Houdek f156615a3a unittests: Adds unittests to ensure FICOM works
Both x80 and x64 variants.
2023-01-08 11:04:28 -08:00
Ryan Houdek 6977ae6b79 X87_F64: Fixes FICOM
This was not correctly converting both 32-bit and 16-bit integers over
to 64-bit double.
2023-01-08 11:02:52 -08:00
Ryan Houdek 555d2e5b0b unittests: Adds negative integer x87 tests
All of these operations were only testing positive integers which is why
they didn't show 16-bit failures.

Adds a bunch of negative tests to each ones now that #2314 is merged,
which would have caught them.
2023-01-08 10:44:46 -08:00
Ryan Houdek c2325e1772 Merge pull request #2314 from CallumDev/f64-integer-fix
F64: Fix integer immediates for add,mul,div,sub
2023-01-08 10:42:21 -08:00
Ryan Houdek 9acb513393 Merge pull request #2313 from Sonicadvance1/fix_large_spills
Arm64: Fixes large offset spill slots
2023-01-08 08:47:14 -08:00
Ryan Houdek dfc3297192 Merge pull request #2311 from Sonicadvance1/optimize_struct_layout
X86Tables: Optimize struct layouts
2023-01-08 08:33:32 -08:00
Ryan Houdek 9322e55a3f Merge pull request #2312 from Sonicadvance1/update_jemalloc
Externals: Update jemalloc to 5.3.0
2023-01-08 08:33:19 -08:00
CallumDev 9373fa0c06 F64: Fix integer immediates for add,mul,div,sub 2023-01-09 01:29:17 +10:30
Ryan Houdek ca9400ba52 Arm64: Fixes large offset spill slots
Found an application today (hashtree tests) that causes us to spill a
large amount of values on to the stack.

We were encoding larger offsets than what unsigned offset load and store
can handle.

If the offset is too large for the loadstore, use a temporary to put the
offset in to first.
2023-01-07 18:03:24 -08:00
Ryan Houdek 2a6937fe59 Core: Fixes uninitialized ParentThread variable
Can cause crashes by not zero initializing. ParentThread isn't
initialized in the TestHarnessRunner when an unsupported test is ran.
2023-01-07 12:09:09 -08:00
Ryan Houdek 138752d512 Externals: Update jemalloc to 5.3.0
Apparently this has some tcache fixes and performance improvements
2023-01-07 11:58:48 -08:00
Ryan Houdek 6b63f9fa89 X86Tables: Optimize struct layouts
We were leaving some ugly holes in a couple of these structs.
Reorder them so they are packed more efficiently.
2023-01-06 18:48:23 -08:00
Ryan Houdek 9a748c020d Arm64: Use switch statement for op handlers instead of jump table
Removes some startup time where we are copying nearly a page worth of
16byte vtable pointers at startup.

Also allows the compiler to choose to inline functions if it wants to.
2023-01-06 17:44:28 -08:00
Ryan Houdek 842e36e9b2 Emitter: Remove unused header 2023-01-06 10:34:41 -08:00
Ryan Houdek 70d4a436cf Docs: Update for release FEX-2301 2023-01-06 07:53:17 -08:00
Ryan Houdek ec55ecdb31 Merge pull request #2290 from Sonicadvance1/new_arm_emitter
Create a new ARM64 Emitter and move JIT over to it.
2023-01-05 13:50:25 -08:00
Tony Wasserka 12b866c276 Merge pull request #2308 from neobrain/refactor_thunkdb_loading
ThunkDB: Clean up database loading
2023-01-05 14:56:35 +01:00
Tony Wasserka 1038ba7060 ThunksDB: Disable error message in 32-bit mode 2023-01-05 14:45:13 +01:00
Tony Wasserka 9bde513161 Thunks: Simplify state carried around while setting up ThunkOverlays 2023-01-05 12:15:09 +01:00
Tony Wasserka 2e3f77c43f ThunksDB: Clean up initial DB loading 2023-01-05 12:15:09 +01:00
Tony Wasserka 0d8de6463b Thunks: Move LoadThunkDatabase out of the header file 2023-01-05 12:15:08 +01:00
Tony Wasserka 98349ee485 ThunksDB: Clean up string replacement logic 2023-01-05 12:15:08 +01:00
Tony Wasserka e486833d75 Merge pull request #2307 from neobrain/fix_thunkdb_libnames
ThunksDB: Fix misspelt guest library names
2023-01-04 15:43:51 +01:00
Ryan Houdek 8a38999c7a RAData: Fixes uninitialized members 2023-01-04 05:30:01 -08:00
Ryan Houdek 2d0b61fd9d Emitter: Adds unit tests
Every operation that the emitter supports is tested in the unit tests.
For the most part uses vixl's dissassembler and a string comparison to
ensure the emitter is outputing what we expect.

For operations that are PC-relative we instead use a bit-exact test to
ensure it is outputting what we care about. This is because vixl
helpfully outputs the PC that the operation is acting on. Since these
locations aren't static in memory, the PC moves around per test.
2023-01-04 05:30:01 -08:00
Ryan Houdek ffd9bb547d Arm64: Convert ARM Emitter over to new emitter
Not yet complete. Missing a full SVE implementation and needs
testing/validation.
2023-01-04 05:30:01 -08:00
Ryan Houdek 7a6ef8821f Arm64: Adds new ARM emitter
Still needs more work.
Missing operations, cleanup, validation
Notably SVE is missing large chunks.
2023-01-04 05:30:01 -08:00
Tony Wasserka 64387bf00d Thunks: Make failure to find a guest thunk library a critical error 2023-01-03 18:15:50 +01:00
Tony Wasserka 8cc7e55394 ThunksDB: Fix misspelt guest library names 2023-01-03 18:15:50 +01:00
Ryan Houdek 874c1da1b5 External: Update vixl 2023-01-02 01:38:17 -08:00
Ryan Houdek 3904a5264f Merge pull request #2306 from lioncash/perm
OpcodeDispatcher: Handle immediate variants of VPERMILPD/VPERMILPS
2022-12-31 21:31:59 -08:00
lioncash b95c1719c3 OpcodeDispatcher: Handle VPERMILPS (immediate) 2023-01-01 05:15:25 +00:00
lioncash dfb3f31453 OpcodeDispatcher: Handle VPERMILPD (immediate) 2023-01-01 04:59:48 +00:00
Ryan Houdek c6297edac0 Merge pull request #2305 from lioncash/mask
OpcodeDispatcher: Handle VMASKMOVDQU
2022-12-31 20:31:20 -08:00
lioncash 8031f76642 OpcodeDispatcher: Handle VMASKMOVDQU 2023-01-01 04:18:46 +00:00
Ryan Houdek 4786ddc44c Merge pull request #2304 from lioncash/sub
OpcodeDispatcher: Handle VPHSUBD/VPHSUBW
2022-12-31 19:23:43 -08:00
lioncash ae8a5fa98d OpcodeDispatcher: Handle VPHSUBD 2023-01-01 03:10:21 +00:00
lioncash 6914598f9a OpcodeDispatcher: Handle VPHSUBW 2023-01-01 02:42:05 +00:00
lioncash 450aedc8b6 x86_64: Fix 256-bit UnZip/UnZip2
We weren't swapping the elements so that the operation acts like the two
vectors are concatenated.
2023-01-01 02:42:05 +00:00
lioncash 1e221210a2 OpcodeDispatcher: Move PHSUB impl to helper function
This will be used for the AVX variants.
2023-01-01 02:42:03 +00:00
Ryan Houdek 58ec2b2d7f Merge pull request #2303 from lioncash/swizz
OpcodeDispatcher: Zip elements instead of for loop insertion in PHSUB
2022-12-31 16:09:13 -08:00
lioncash 438adf2f45 x86_64/VectorOps: Handle OpSize==8 case in VUnZip/VUnZip2
Allows the x86 side of things to execute the new codepath in PHSUB
2022-12-31 23:30:59 +00:00
lioncash 9707e9a4df OpcodeDispatcher: Zip elements instead of for loop in PHSUB
Makes this much nicer for 128-bit and soon-to-be 256-bit vectors.
2022-12-31 23:30:38 +00:00
Ryan Houdek 9b8c92e275 Merge pull request #2302 from lioncash/dpp
OpcodeDispatcher: Handle VDPPD/VDPPS
2022-12-31 15:01:52 -08:00
Ryan Houdek 6caf764b7c Merge pull request #2301 from lioncash/insps
OpcodeDispatcher: Handle VINSERTPS
2022-12-31 14:59:11 -08:00
Ryan Houdek faa81f241b Merge pull request #2300 from lioncash/msk
OpcodeDispatcher: Handle VMOVMSKPD/VMOVMSKPS
2022-12-31 14:34:44 -08:00
lioncash 769c548ba4 OpcodeDispatcher: Handle VDPPD
x86 just doesn't have a 256-bit version of this op.
2022-12-31 21:13:20 +00:00
lioncash dae1676e4a OpcodeDispatcher: Handle VDPPS 2022-12-31 21:01:48 +00:00
lioncash 74526d1f02 OpcodeDispatcher: Move DPP op impl into helper function
This will be used to implement the AVX variants.
2022-12-31 21:01:26 +00:00
lioncash 8c005db81c OpcodeDispatcher: Handle VINSERTPS 2022-12-31 20:23:22 +00:00
lioncash 32b70c8590 OpcodeDispatcher: Move InsertPS impl to helper function
This will be reused for VINSERTPS
2022-12-31 20:23:18 +00:00
lioncash bdda14eb75 OpcodeDispatcher: Handle VMOVMSKPD 2022-12-31 19:24:14 +00:00
lioncash b8c0b0c267 OpcodeDispatcher: Handle VMOVMSKPS 2022-12-31 19:24:11 +00:00
lioncash ad7dc6ca0a OpcodeDispatcher: Move ADDSUB table entries
Makes them numerically ordered.
2022-12-31 19:07:19 +00:00
Ryan Houdek 64cd377e37 Merge pull request #2299 from lioncash/ph
OpcodeDispatcher: Handle VPUNPCKHBW/VPUNPCKHWD/VPUNPCKHDQ/VPUNPCKHQDQ
2022-12-30 17:18:14 -08:00
lioncash 45d7564716 OpcodeDispatcher: Handle VPUNPCKHBW 2022-12-30 14:07:30 +00:00
lioncash c585bae85d OpcodeDispatcher: Handle VPUNPCKHWD 2022-12-30 14:01:18 +00:00
lioncash d07383fa73 OpcodeDispatcher: Handle VPUNPCKHQDQ 2022-12-30 13:54:45 +00:00
lioncash e1cdcf0651 OpcodeDispatcher: Handle VPUNPCKHDQ 2022-12-30 13:46:53 +00:00
Ryan Houdek 138f1fc844 Merge pull request #2298 from lioncash/hps
OpcodeDispatcher: Handle VUNPCKHPD/VUNPCKHPS
2022-12-30 05:31:48 -08:00
lioncash ae69aa9a81 OpcodeDispatcher: Handle VUNPCKHPD 2022-12-30 13:18:33 +00:00
lioncash 6341ac6814 OpcodeDispatcher: Handle VUNPCKHPS 2022-12-30 13:05:01 +00:00
Ryan Houdek 6bc1c3fc30 Merge pull request #2297 from lioncash/lps
OpcodeDispatcher: Handle VPUNPCKLBW/VPUNPCKLWD/VPUNPCKLDQ/VPUNPCKLQDQ
2022-12-29 14:49:11 -08:00
lioncash d91f2ed6b0 OpcodeDispatcher: Handle VPUNPCKLBW 2022-12-29 16:36:45 +00:00
lioncash 7b30a241c0 OpcodeDispatcher: Handle VPUNPCKLWD 2022-12-29 16:30:12 +00:00
lioncash aaf8e3757d OpcodeDispatcher: Handle VPUNPCKLQDQ 2022-12-29 16:21:56 +00:00
lioncash d9c49c4ce1 OpcodeDispatcher: Handle VPUNPCKLDQ 2022-12-29 16:19:09 +00:00
Ryan Houdek 4560c5b73c Merge pull request #2296 from lioncash/lps
OpcodeDispatcher: Handle VUNPCKLPD/VUNPCKLPS
2022-12-29 08:01:23 -08:00
lioncash 9d05c8a67b OpcodeDispatcher: Handle VUNPCKLPD 2022-12-29 15:12:36 +00:00
lioncash be9578551a OpcodeDispatcher: Handle VUNPCKLPS 2022-12-29 15:00:13 +00:00
Ryan Houdek 4a884802f8 Merge pull request #2295 from lioncash/cvt
OpcodeDispatcher: Handle VCVTSS2SI/VCVTTSS2SI/VCVTSD2SI/VCVTTSD2SI
2022-12-29 03:14:54 -08:00
lioncash 75a01ed2b6 OpcodeDispatcher: Handle VCVTTSD2SI 2022-12-29 10:49:03 +00:00
lioncash 3ebe141032 OpcodeDispatcher: Handle VCVTSD2SI 2022-12-29 10:43:10 +00:00
lioncash 31e332bd61 OpcodeDispatcher: Handle VCVTTSS2SI 2022-12-29 10:26:19 +00:00
lioncash 764324d557 OpcodeDispatcher: Handle VCVTSS2SI 2022-12-29 10:16:12 +00:00
Ryan Houdek f37938576d Merge pull request #2292 from lioncash/cvt
OpcodeDispatcher: Handle VCVTPD2DQ/VCVTTPD2DQ/VCVTPS2DQ/VCVTTPS2DQ
2022-12-28 17:16:46 -08:00
Ryan Houdek 16969fcdad Merge pull request #2293 from neobrain/fix_thunks_ide_integration
Thunks: Fix IDE integration
2022-12-28 17:14:10 -08:00
Tony Wasserka 0dfe141d70 Thunks: Fix IDE integration 2022-12-28 17:45:05 +01:00
lioncash 2a7795fe2c OpcodeDispatcher: Handle VCVTTPD2DQ 2022-12-28 12:01:53 +00:00
lioncash b00b41b8fa OpcodeDispatcher: Handle VCVTPD2DQ 2022-12-28 11:54:28 +00:00
lioncash 39396789b1 Interpreter/ConversionOps: Prevent out of bounds/excessive element handling in Vector_FToF
Previously this could access more elements than it needs to.

In the case where we pass in a 128-bit vector and perform a 128-bit
operation, only two doubles should be handled, but in this case will
actually try and handle four double elements.

Consider converting doubles within a vector into floats:

e.g. _Vector_FToF(16, 4, Src, 8)

16 is our OpSize
4 is our DestElementSize
Src is our input vector
8 is the SrcElementSize

(OpSize << 1) / Op->SrcElementSize becomes:
(16 << 1) / 8 ->
32 / 8 ->
4

What we actually want it 2 here, we're indexing 128-bit vectors out of
its bounds (not a problem now, since we assume up to 256-bit in the
interpreter).

Similarly with 256-bit vectors

_Vector_FToF(32, 4, Src, 8)

would become:

64 / 8 -> 8

where we actually want 4, since we only have 4 64-bit elements in a
256-bit vector.

This corrects this in the interpreter by just special-casing the 64-bit
calculation.
2022-12-28 11:53:29 +00:00
lioncash 38a2886a59 OpcodeDispatcher: Handle VCVTTPS2DQ 2022-12-28 07:54:16 +00:00
lioncash bd8e1a80f6 OpcodeDispatcher: Handle VCVTPS2DQ 2022-12-28 07:45:37 +00:00
Ryan Houdek b7358b4926 Merge pull request #2261 from Sonicadvance1/optimize_lookup_pmr_map
LookupCache: Use a PMR map for our Blocklinks with monotonic allocator
2022-12-23 09:55:28 -08:00
Ryan Houdek 0e0f3f9290 LookupCache: Use a PMR map for our Blocklinks with monotonic allocator
We generate a /lot/ of block links which causes our shutdown time to
take a while for this map. For short running applications this shutdown
time can take a statistically significant amount of time, and this has
been haunting us for a while.

Using a monotonic buffer resource, we can cut the dellocation time down
to effectively zero on cache clear and shutdown.
For short running applications this basically means we clear their
shutdown time, for long running applications this means a cache clear is
less painful.

I've measured up to 2ms , but more often this takes ~0.5ms before
optimization. Now it is consistently ~0.25ms on a short running
application.
Times will vary /greatly/ depending on how full it has filled this map.
2022-12-23 09:41:54 -08:00
lioncash dfa113dcdb OpcodeDispatcher: Move Vector_CVT_Float_To_Int to helper function
This can be reused for the AVX variant.
2022-12-23 17:33:15 +00:00
Mai bf7d0f7ed9 Merge pull request #2287 from Sonicadvance1/rename_getreg
Arm64: Rename GetSrcPair, GetDst, and GetSrc
2022-12-22 08:40:10 +00:00
Ryan Houdek 82adc2f931 Merge pull request #2288 from lioncash/hrsw
OpcodeDispatcher: Handle VPMULHRSW
2022-12-22 00:38:45 -08:00
lioncash 94cb2ddae7 OpcodeDispatcher: Handle VPMULHRSW 2022-12-22 08:21:35 +00:00
lioncash 0496f8d5c1 OpcodeDispatcher: Move PMULHRSW impl to helper 2022-12-22 08:13:43 +00:00
Ryan Houdek 4a3af8d7f9 Merge pull request #2286 from lioncash/mulhw
OpcodeDispatcher: Handle VPMULHW/VPMULHUW
2022-12-22 00:06:53 -08:00
Ryan Houdek 37a9588855 Arm64: Rename GetSrcPair to GetRegPair
This matches the other handler's names and isn't only for source
registers, but also destination registers.
2022-12-22 00:03:41 -08:00
Ryan Houdek 24f72b5f30 Arm64: Rename GetDst & GetSrc to GetVReg
These only resolved to getting vector registers now and having
duplicating handlers for it was just confusing.
2022-12-21 23:59:21 -08:00
lioncash d927c4a903 OpcodeDispatcher: Handle VPMULHUW 2022-12-22 07:52:27 +00:00
lioncash 12afe95602 OpcodeDispatcher: Handle VPMULHW 2022-12-22 07:52:24 +00:00
lioncash 7e715b9e04 OpcodeDispatcher: Move PMULHW impl to helper
This can be used with AVX variants.
2022-12-22 07:33:26 +00:00
Ryan Houdek 9d58514f57 Merge pull request #2285 from lioncash/phmin
OpcodeDispatcher: Handle VPHMINPOSUW
2022-12-21 19:02:35 -08:00
lioncash 4f9402e5dd OpcodeDispatcher: Handle VPHMINPOSUW 2022-12-22 02:46:08 +00:00
lioncash 789093d158 OpcodeDispatcher: Move VPHMINPOSUW impl to separate function
This will be used with the AVX variant.
2022-12-22 02:38:31 +00:00
Ryan Houdek 33e8f21ac7 Merge pull request #2284 from lioncash/pmull
OpcodeDispatcher: Handle VPMULDQ/VPMULUDQ
2022-12-21 18:33:02 -08:00
lioncash d672528e62 OpcodeDispatcher: Handle VPMULUDQ 2022-12-22 02:18:37 +00:00
lioncash d7c959090d OpcodeDispatcher: Handle VPMULDQ 2022-12-22 02:08:53 +00:00
lioncash 0a4846a524 OpcodeDispatcher: Factor PMULLOp impl to regular function
This can be reused for the AVX implementations.
2022-12-22 01:30:06 +00:00
Ryan Houdek cecda7bbb6 Merge pull request #2283 from lioncash/cmpss
OpcodeDispatcher: Handle VCMPSD/VCMPSS
2022-12-21 16:58:01 -08:00
lioncash 2d9cb65d5c OpcodeDispatcher: Handle VCMPSD 2022-12-21 20:46:07 +00:00
lioncash e21002e0d7 OpcodeDispatcher: Handle VCMPSS 2022-12-21 20:39:13 +00:00
Ryan Houdek ce351282f2 Merge pull request #2282 from lioncash/debug
OpcodeDispatcher: Remove lingering debug log from VPFCMPOp
2022-12-21 00:53:13 -08:00
lioncash ab41856328 OpcodeDispatcher: Remove lingering debug log from VPFCMPOp 2022-12-21 08:40:27 +00:00
Ryan Houdek 345e9b97bf Merge pull request #2281 from lioncash/assert
OpcodeDispatcher: Convert runtime assert to static_assert in SHUFOps
2022-12-21 00:23:24 -08:00
Ryan Houdek 0c651dd5f8 Merge pull request #2280 from lioncash/cmp
OpcodeDispatcher: Handle VCMPPD/VCMPPS
2022-12-21 00:10:14 -08:00
lioncash f850a02d3f OpcodeDispatcher: Convert runtime assert to static_assert in SHUFOps 2022-12-21 08:05:32 +00:00
lioncash 983b53a0c2 OpcodeDispatcher: Handle VCMPPD 2022-12-21 06:12:14 +00:00
lioncash 10a6b5794b OpcodeDispatcher: Handle VCMPPS 2022-12-21 05:57:24 +00:00
lioncash 2c5aceb9b6 OpcodeDispatcher: Factor out VFCMPOp impl into a regular function
This can be used with AVX implementations.
2022-12-21 05:27:10 +00:00
Ryan Houdek 1668db046d Merge pull request #2279 from lioncash/psrldq
OpcodeDispatcher: Handle VPSRLDQ
2022-12-20 21:06:57 -08:00
lioncash 825e921940 OpcodeDispatcher: Handle VPSRLDQ 2022-12-21 04:51:49 +00:00
Ryan Houdek 515b3e485b Merge pull request #2278 from lioncash/shift
OpcodeDispatcher: Remove unnecessary usage of VMov in VPSLLDQOp
2022-12-20 20:30:08 -08:00
lioncash c06f0b7cb3 OpcodeDispatcher: Remove usage of VMov in VPSLLDQOp
The extract operation essentially does this for us.
2022-12-21 04:13:26 +00:00
Ryan Houdek 4aed60ee3d Merge pull request #2277 from lioncash/cvt
OpcodeDispatcher: Handle VCVTDQ2PD/VCVTDQ2PS
2022-12-20 20:02:12 -08:00
lioncash da9f7ec31f OpcodeDispatcher: Handle VCVTDQ2PD 2022-12-21 03:45:00 +00:00
lioncash 8ec932fc4c OpcodeDispatcher: Handle VCVTDQ2PS 2022-12-21 03:44:56 +00:00
lioncash 6f35a23161 x86_64/ConversionOps: Don't clear upper lane in Vector_SToF
If Dst and Vector are the same, this will obliterate the top lane data
that needs to be operated on in the 128-bit case.
2022-12-21 03:40:38 +00:00
lioncash bd70af9724 OpcodeDispatcher: Move Vector_CVT_Int_To_Float impl to regular function
This can be used with the AVX implementation as well.
2022-12-21 02:23:04 +00:00
lioncash bdba062f72 VEXTables: Amend VCVTTPS2DQ name 2022-12-21 02:17:31 +00:00
Ryan Houdek 60a2fb163c Merge pull request #2276 from lioncash/lldq
OpcodeDispatcher: Handle VPSLLDQ
2022-12-20 18:12:15 -08:00
lioncash d591b1ed8c Interpreter: Prevent overrun with 256-bit VExtr 2022-12-21 01:51:25 +00:00
lioncash 3bae4a225c OpcodeDispatcher: Handle VPSLLDQ 2022-12-21 01:46:17 +00:00
Ryan Houdek d0cb329608 Merge pull request #2275 from lioncash/sha1
OpcodeDispatcher: Simplify SHA1MSG1 implementation
2022-12-20 12:49:41 -08:00
lioncash bb80e7d45c OpcodeDispatcher: Simplify SHA1MSG1 implementation
We can just arrange the elements into a vector and XOR them all at once
2022-12-20 19:49:08 +00:00
Ryan Houdek 72a3b18279 Merge pull request #2274 from lioncash/right
OpcodeDispatcher: Handle immediate variants of VPSRLD/VPSRLQ/VPSRLW
2022-12-20 11:38:13 -08:00
lioncash 109ed7d112 OpcodeDispatcher: Handle VPSRLQ (immediate) 2022-12-20 19:24:58 +00:00
lioncash 133a644231 OpcodeDispatcher: Handle VPSRLD (immediate) 2022-12-20 19:12:16 +00:00
lioncash 666f8bfbd9 OpcodeDispatcher: Handle VPSRLW (immediate) 2022-12-20 19:07:47 +00:00
Ryan Houdek 1800451251 Merge pull request #2273 from lioncash/keygen
OpcodeDispatcher: Handle 128-bit AVX AES instructions
2022-12-20 10:52:53 -08:00
Ryan Houdek 1d9218224f Merge pull request #2272 from lioncash/psra
OpcodeDispatcher: Handle immediate variants of VPSRAD/VPSRAW
2022-12-20 10:50:22 -08:00
lioncash 7931bd1004 OpcodeDispatcher: Handle VAESDECLAST (128-bit) 2022-12-20 17:34:17 +00:00
lioncash 58978dd047 OpcodeDispatcher: Handle VAESDEC (128-bit) 2022-12-20 17:26:19 +00:00
lioncash 25fb243ac7 OpcodeDispatcher: Handle VAESENCLAST (128-bit) 2022-12-20 17:11:47 +00:00
lioncash 84f1e7ad4c OpcodeDispatcher: Handle VAESENC (128-bit)
Only 128-bit is required to be handled by base-level AVX.

The VAES feature flag indicates support for 256-bit VAESENC
2022-12-20 16:58:01 +00:00
lioncash 3fb5835453 OpcodeDispatcher: Handle VAESIMC
VAESIMC behaves exactly like AESIMC, except the upper lane of the vector
is always cleared.
2022-12-20 16:49:03 +00:00
lioncash bcb6726b22 OpcodeDispatcher: Handle VAESKEYGENASSIST
This does the exact same thing as AESKEYGENASSIST, except that the upper
lane gets cleared.
2022-12-20 16:15:42 +00:00
lioncash 2bed562eb6 OpcodeDispatcher: Extract AESKeyGenAssist impl to helper function
This can be reused for the AVX variant.
2022-12-20 16:01:25 +00:00
lioncash bae7209224 OpcodeDispatcher: Handle VPSRAD (immediate) 2022-12-20 15:49:38 +00:00
lioncash b53f8944ac OpcodeDispatcher: Handle VPSRAW (immediate) 2022-12-20 15:40:49 +00:00
Mai 03a061339a Merge pull request #2269 from Sonicadvance1/vixl_disassembler
Arm64: Enables debug option for disassembling the JIT code
2022-12-20 01:25:51 +00:00
Ryan Houdek 0d7c086b69 Arm64: Enables debug option for disassembling the JIT code
This is useful as a debug option and will be useful to have in upstream
while comparing output between current vixl emitter and the new emitter.

With this in place I can easily do binary comparisons to see where I
have mistakes in the new emitter.

We don't want this enabled in release builds as it is a debug feature.
This has already caught a bunch of mistakes, so make it easier by
upstreaming.
It'll likely be useful in the future as well when we are inspecting code
running in the vixl simulator.
2022-12-18 14:56:40 -08:00
Ryan Houdek b958fa39a5 External: Update vixl 2022-12-18 14:54:52 -08:00
Mai 2b6a020c4c Merge pull request #2260 from Sonicadvance1/optimize_lookup_map
LookupCache: Optimize cache clearing and allocation
2022-12-17 22:43:22 +00:00
Ryan Houdek 6e733bfc22 Merge pull request #2268 from lioncash/upack
OpcodeDispatcher: Handle VPACKUSDW/VPACKUSWB
2022-12-16 20:00:54 -08:00
lioncash 873d63002a OpcodeDispatcher: Handle VPACKUSDW 2022-12-17 03:43:06 +00:00
lioncash bb6a0f39f5 OpcodeDispatcher: Handle VPACKUSWB 2022-12-17 03:31:40 +00:00
lioncash 392e6ae424 OpcodeDispatcher: Factor out PACKUSOp impl into a regular function
We can use this for the AVX instructions too.
2022-12-17 03:19:28 +00:00
Ryan Houdek 01d22849cf Merge pull request #2267 from lioncash/pack
OpcodeDispatcher: Handle VPACKSSDW/VPACKSSWB
2022-12-16 19:16:28 -08:00
lioncash 0537f2d014 OpcodeDispatcher: Handle VPACKSSDW 2022-12-17 03:01:33 +00:00
lioncash f57debeb29 OpcodeDispatcher: Handle VPACKSSWB 2022-12-17 02:42:09 +00:00
lioncash 4ac031df59 OpcodeDispatcher: Move PACKSSOp impl to a regular function
We can reuse it with AVX versions.
2022-12-17 02:13:26 +00:00
Ryan Houdek 78b53bfa49 Merge pull request #2266 from lioncash/arith
OpcodeDispatcher: Handle vector versions of VPSRA{D, W}
2022-12-16 18:05:05 -08:00
lioncash c53fb7d697 OpcodeDispatcher: Handle VPSRAD (vector) 2022-12-17 01:51:42 +00:00
lioncash a1a52450cb OpcodeDispatcher: Handle VPSRAW (vector) 2022-12-17 01:40:25 +00:00
Ryan Houdek fabf453046 Merge pull request #2265 from lioncash/pextrw
OpcodeDispatcher: Handle remaining PEXTRW opcode
2022-12-16 17:35:33 -08:00
lioncash 68916ae2d9 OpcodeDispatcher: Move PSRAOp implementation to regular function
We can reuse this with the AVX variant.
2022-12-17 01:23:02 +00:00
lioncash bf56b7b2da OpcodeDispatcher: Handle remaining PEXTRW opcode 2022-12-17 01:14:22 +00:00
Ryan Houdek 905eb015c0 Merge pull request #2264 from lioncash/addsub
OpcodeHandler: Handle VADDSUBP{D, S}
2022-12-16 16:54:35 -08:00
lioncash 858f13e76a OpcodeDispatcher: Handle VADDSUBPD 2022-12-17 00:41:25 +00:00
lioncash 169d7bbf50 OpcodeDispatcher: Handle VADDSUBPS 2022-12-17 00:29:29 +00:00
lioncash 31c8d4acac OpcodeDispatcher: Factor out ADDSUB impl into regular function
We can reuse this with the AVX versions
2022-12-17 00:16:38 +00:00
lioncash 8291e600fa OpcodeDispatcher: Simplify ADDSUBPOp
Rather than looping vectors, we can interleave them together directly
with IR ops.
2022-12-17 00:11:51 +00:00
Ryan Houdek b26e4109fa Merge pull request #2263 from lioncash/mull
OpcodeDispatcher: Handle VPMULL{D, B}
2022-12-16 15:40:53 -08:00
lioncash dcc218a168 OpcodeDispatcher: Handle VPMULLD 2022-12-16 23:20:25 +00:00
lioncash 49b9b18b4a OpcodeDispatcher: Handle VPMULLW 2022-12-16 23:07:06 +00:00
Ryan Houdek ad3bf189c0 Merge pull request #2262 from lioncash/rlog
OpcodeDispatcher: Handle vector variants of VPSRL{D, Q, W}
2022-12-16 14:56:56 -08:00
lioncash 47b21fa758 OpcodeDispatcher: Handle VPSRLQ (vector)
Also mark VPMOVMSKB as UNDEC, since it's not implemented yet.
2022-12-16 22:18:40 +00:00
lioncash b6e82965df OpcodeDispatcher: Handle VPSRLD (vector) 2022-12-16 22:09:52 +00:00
lioncash 8dc8785340 OpcodeDispatcher: Handle VPSRLW (vector) 2022-12-16 22:00:53 +00:00
lioncash c710ab60b0 OpcodeDispatcher: Factor out PSRLDOp implementation to regular function
This will be used with the AVX variants of the shifts also
2022-12-16 21:43:11 +00:00
Ryan Houdek c86ba7646c Merge pull request #2259 from lioncash/pextr
OpcodeDispatcher: Handle VPEXTR{B, D, Q, W}/VEXTRACTPS
2022-12-16 11:25:38 -08:00
Ryan Houdek c1e301a5ed Merge pull request #2257 from lioncash/limm
OpcodeDispatcher: Handle immediate variants of VPSLL{D, Q, W}
2022-12-16 11:23:16 -08:00
Ryan Houdek cad0dc6848 LookupCache: Optimize cache clearing and allocation
Use one large allocation for all levels of the cache so they are
virtually contiguous.
This allows us to clear the cache entirely by using a single madvise
instead of three. Which ends up being quite a bit nicer.
2022-12-16 11:08:39 -08:00
Mai 0ebb15c732 Merge pull request #2258 from Sonicadvance1/fixed_syscall_spill
Arm64: Inline Syscall spill optimization
2022-12-16 18:48:08 +00:00
lioncash 37c743b616 OpcodeDispatcher: Handle VPSLLQ (immediate) 2022-12-16 18:37:27 +00:00
lioncash c810ae4018 OpcodeDispatcher: Handle VPSLLD (immediate) 2022-12-16 18:37:27 +00:00
lioncash d3481c8271 OpcodeDispatcher: Handle VPSLLW (immediate) 2022-12-16 18:37:27 +00:00
lioncash 7c1e152441 OpcodeDispatcher: Extract PSLLI impl to regular function
This will be reused for the AVX variants.
2022-12-16 18:37:20 +00:00
lioncash f11ac8674d OpcodeDispatcher: Handle VEXTRACTPS 2022-12-16 18:13:55 +00:00
Ryan Houdek 1fecf89bfc Arm64: Inline Syscall spill optimization
This was likely an issue with signals racing to the spill handler, which
we have fixed bugs with over the past few months.

This means we don't need to spill all SRA GPR registers anymore, at most
we need to spill three registers that intersect with syscall arguments.
2022-12-16 10:04:16 -08:00
lioncash 21ad0fa334 OpcodeDispatcher: Handle VPEXTRQ
VPEXTRQ uses VEX.W to handle size differencing, since it shares an
encoding spot with VPEXTRD, so we need to handle that a little
differently.
2022-12-16 18:02:24 +00:00
lioncash 3429815103 OpcodeDispatcher: Handle VPEXTRD 2022-12-16 17:33:01 +00:00
lioncash 559ff1582e OpcodeDispatcher: Handle VPEXTRW 2022-12-16 17:29:16 +00:00
lioncash 2e973ae079 OpcodeDispatcher: Handle VPEXTRB 2022-12-16 14:37:47 +00:00
Mai 1ab4471ef9 Merge pull request #2255 from Sonicadvance1/optimize_sve_spillfill
Arm64: Optimize SVE register spilling and filling
2022-12-16 13:19:05 +00:00
Ryan Houdek 40e073c8b2 Arm64: Optimize SVE register spilling and filling
Causes the dispatcher to drop from 4476 bytes down to 3900 for
SVE-256bit supporting targets.

This is done by significantly reducing SVE loadstore ops. Going from 8
instructions per 4 registers, down to 2 instructions.

This is done by switching from 1 register loadstore instructions up to 4
register loadstore instructions. Which should significantly improve
performance on future SVE platforms.

Filling and Spilling to the context is still using the old code path
because SVE doesn't offer non-interleaving loadstores.
Spilling and filling on the stack is fine because we don't need to match
context state.
2022-12-16 00:25:50 -08:00
Ryan Houdek 58fab721b3 Merge pull request #2254 from lioncash/logical
OpcodeDispatcher: Handle vector variants of VPSLL{D, Q, W}
2022-12-15 22:52:05 -08:00
lioncash 8fac21e43f OpcodeDispatcher: Handle VPSLLQ (vector) 2022-12-16 06:34:00 +00:00
lioncash d9a1e97bc1 OpcodeDispatcher: Handle VPSLLD (vector) 2022-12-16 06:34:00 +00:00
lioncash 848f1a2f78 OpcodeDispatcher: Handle VPSLLW (vector) 2022-12-16 06:34:00 +00:00
lioncash 7b8a46d934 OpcodeDispatcher: Move PSLL impl into a regular function 2022-12-16 06:33:58 +00:00
Mai 9a8852f9b6 Merge pull request #2250 from Sonicadvance1/optimize_spilling_filling
Arm64: Optimizing spilling and filling
2022-12-16 04:47:22 +00:00
Mai 65e8bf9d72 Merge pull request #2253 from Sonicadvance1/single_page_dispatcher
Arm64: Reduce dispatcher to 1 page
2022-12-16 04:44:55 +00:00
Ryan Houdek 344ec33ba5 Merge pull request #2252 from lioncash/fadd
Arm64/VectorOps: Simplify FADDP result merging
2022-12-15 20:37:11 -08:00
Ryan Houdek 5dc7dfacb3 Arm64: Reduce dispatcher to 1 page
We currently only use 2236 bytes, no need for two pages.
Once #2250 is merged we will use 1716 bytes
2022-12-15 20:33:28 -08:00
lioncash 122aa8a69a Arm64/VectorOps: Simplify FADDP result merging
Keeps the implementation similarly in sync with VAddP.
2022-12-16 04:19:46 +00:00
Ryan Houdek 8ce6c08152 Merge pull request #2251 from lioncash/hadd
OpcodeDispatcher: Handle VPHADDW/VPHADDD
2022-12-15 20:11:07 -08:00
Ryan Houdek 1beb791d52 Arm64: Optimizing spilling and filling
Just makes these a little more optimal when jumping out of the JIT.

Noticed these while working on the new emitter.
2022-12-15 20:04:16 -08:00
lioncash 27c0d4a9f5 OpcodeDispatcher: Handle VPHADDD 2022-12-16 03:28:57 +00:00
lioncash dd4ba7562f OpcodeDispatcher: Handle VPHADDW 2022-12-16 03:28:57 +00:00
lioncash bd9d8e8fe5 x86_64: Correct handling for 128-bit/256-bit VAddP
Makes the behavior consistent with the ARM JIT.
2022-12-16 03:28:57 +00:00
lioncash c7ac204322 Arm64/VectorOps: Simplify VAddP merging operation
We can just merge the two results together instead of shifting to the
left and then ORing together.
2022-12-16 03:28:52 +00:00
Ryan Houdek 4c013c867f Merge pull request #2249 from lioncash/clear
Crypto: Explicitly clear upper lane with VPCLMULQDQ
2022-12-15 17:33:57 -08:00
lioncash 5e634fcbc9 Crypto: Explicitly clear upper lane with VPCLMULQDQ
Ensures the 128-bit case will be handled when extending for 256-bit
2022-12-16 01:08:48 +00:00
Ryan Houdek 91c00d2cb6 Merge pull request #2248 from lioncash/acc
X86Tables: Restrict CVTDQ2PD and CVTTSD2SI to 64-bit memory accesses
2022-12-15 16:48:42 -08:00
lioncash e985dcdb22 X86Tables: Restrict CVTTSD2SI src to 64 bit
When accessing memory, this should only be doing a 64-bit access, rather
than a 128-bit one.
2022-12-15 23:59:15 +00:00
lioncash ee9778480d X86Tables: Restrict CVTDQ2PD src to 64 bit
When accessing memory, this should only be doing a 64-bit access, rather
than a 128-bit one.
2022-12-15 23:46:48 +00:00
Mai 048daa4579 Merge pull request #2244 from Sonicadvance1/move_to_header
ARM64: Moves RA functions to header
2022-12-15 23:13:32 +00:00
Ryan Houdek 6ae8a1e55f ARM64: Moves RA functions to header
These are just some basic address calculations and a load, we want these
to be inlined as much as possible.
2022-12-15 15:00:33 -08:00
Ryan Houdek dc2eaf6511 Merge pull request #2246 from lioncash/extend
OpcodeDispatcher: Handle VPMOVSXB{D, W, Q}/VPMOVSXW{D, Q}/VPMOVSXDQ/VPMOVZXB{D, W, Q}/VPMOVZXW{D, Q}/VPMOVZXDQ
2022-12-15 14:19:28 -08:00
Ryan Houdek 0e233a96f0 Merge pull request #2247 from lioncash/roundacc
OpcodeDispatcher: Narrow memory access with scalar rounding operations
2022-12-15 14:17:49 -08:00
lioncash ba5fafcd7f OpcodeDispatcher: Narrow memory access with scalar rounding operations
These should only be accessing a 32-bit or 64-bit portion of memory
depending on single or double precision variants are used. Previously
we'd be doing a full 128-bit load.
2022-12-15 19:42:37 +00:00
lioncash b12503fe32 OpcodeDispatcher: Handle VPMOVSXDQ 2022-12-15 18:10:38 +00:00
lioncash aa63c7b94d OpcodeDispatcher: Handle VPMOVSXWQ 2022-12-15 18:08:00 +00:00
lioncash cccbb7f595 OpcodeDispatcher: Handle VPMOVSXWD 2022-12-15 18:01:43 +00:00
lioncash ce12ed60ae OpcodeDispatcher: Handle VPMOVSXBQ 2022-12-15 17:58:42 +00:00
lioncash d7eab5f787 OpcodeDispatcher: Handle VPMOVSXBD 2022-12-15 17:54:51 +00:00
lioncash 21537a3636 OpcodeDispatcher: Handle VPMOVSXBW 2022-12-15 17:50:58 +00:00
lioncash 588a2611a7 OpcodeDispatcher: Handle VPMOVZXDQ 2022-12-15 17:45:14 +00:00
lioncash 2895a09101 OpcodeDispatcher: Handle VPMOVZXWQ 2022-12-15 17:41:15 +00:00
lioncash 5c8d40d9be OpcodeDispatcher: Handle VPMOVZXWD 2022-12-15 17:37:51 +00:00
lioncash b4079cfea3 OpcodeDispatcher: Handle VPMOVZXBQ 2022-12-15 17:32:18 +00:00
lioncash 2b5570a910 OpcodeDispatcher: Handle VPMOVZXBD 2022-12-15 17:28:35 +00:00
lioncash 6bb0c5b24c OpcodeDispatcher: Handle VPMOVZXBW 2022-12-15 17:18:49 +00:00
lioncash bc31f98f16 OpcodeDispatcher: Move ExtendVectorElements impl to regular function
This can be reused for the AVX versions.
2022-12-15 17:11:02 +00:00
Ryan Houdek 4b891d6147 Merge pull request #2245 from lioncash/split
OpcodeDispatcher: Move template impl to regular function where applicable
2022-12-14 18:18:43 -08:00
lioncash 58c3e20bd1 OpcodeDispatcher: Move template impl to regular function where applicable
Reduces the amount of code size generated by the specializations.

Only targets ones that are heavily templated like the generic op helper
functions.
2022-12-15 01:54:12 +00:00
Ryan Houdek d5f3a091d0 Merge pull request #2216 from Sonicadvance1/32bit_host_thunk_support
Initial 32-bit host thunk feature support
2022-12-14 12:05:37 -08:00
Ryan Houdek a14e03f35d Update guest thunk lib register usage comment 2022-12-14 11:40:33 -08:00
Ryan Houdek 5c1789952e GuestThunks: Disable stack protector on 32-bit 2022-12-14 11:29:19 -08:00
Ryan Houdek f5809f24f7 GuestLibs: Fixes accidental guest lib setting 2022-12-14 11:29:19 -08:00
Ryan Houdek 122a9114a3 Thunks: 32-bit host library support 2022-12-14 11:29:19 -08:00
Ryan Houdek d8f226b460 Support 32-bit thunks ABI 2022-12-14 11:29:19 -08:00
Ryan Houdek 7171c5ae39 Support 32-bit thunksdb 2022-12-14 11:29:19 -08:00
Ryan Houdek 798a78534a Support Indirect thunk callback with mm0 as custom ABI 2022-12-14 11:24:18 -08:00
Ryan Houdek ae4a04b560 Fix incorrect THUNK_ABI prefix 2022-12-14 11:24:18 -08:00
Ryan Houdek 1971c8d505 32bit host thunk lib config path support 2022-12-14 11:24:18 -08:00
Ryan Houdek 1ca356371d Merge pull request #2242 from lioncash/round
OpcodeDispatcher: Handle VROUNDS{D, S}/VROUNDP{D, S}
2022-12-13 23:00:51 -08:00
lioncash 27ea6096a2 OpcodeDispatcher: Handle VROUNDSD 2022-12-14 06:41:36 +00:00
lioncash 2244dd9847 OpcodeDispatcher: Handle VROUNDSS 2022-12-14 06:34:58 +00:00
lioncash ca2f4bd468 OpcodeDispatcher: Handle VROUNDPD 2022-12-14 06:28:17 +00:00
lioncash 6b5c94be23 OpcodeDispatcher: Handle VROUNDPS 2022-12-14 06:27:59 +00:00
lioncash 779dc48d8d OpcodeDispatcher: Factor out VectorRound into VectorRoundImpl
This will be used in following commits for the AVX versions that use
this.
2022-12-14 05:52:16 +00:00
Ryan Houdek 4b2164768f Merge pull request #2241 from lioncash/ins
OpcodeDispatcher: Handle VINSERTF128/VINSERTI128
2022-12-13 20:45:44 -08:00
lioncash 90828aeb11 OpcodeDispatcher: Handle VINSERTI128 2022-12-14 04:26:42 +00:00
lioncash fe7c6da1e2 OpcodeDispatcher: Handle VINSERTF128 2022-12-14 04:24:04 +00:00
Ryan Houdek f3d0fa6f60 Merge pull request #2240 from lioncash/perm2
OpcodeDispatcher: Handle VPERM2F128/VPERM2I128
2022-12-13 19:57:31 -08:00
lioncash a9ad0d081c OpcodeDispatcher: Handle VPERM2I128 2022-12-14 03:41:29 +00:00
lioncash 54885bec32 OpcodeDispatcher: Handle VPERM2F128 2022-12-14 03:41:22 +00:00
Ryan Houdek e8aa79bea9 Merge pull request #2239 from lioncash/dec
Frontend: Handle 256-bit destination sizes directly
2022-12-13 19:01:18 -08:00
Ryan Houdek 60a45615df Merge pull request #2238 from lioncash/permq
OpcodeDispatcher: Handle VPERMQ/VPERMPD
2022-12-13 17:54:21 -08:00
lioncash 8a961bfcc5 VEXTables: Specify VPERMQ/VPERMPD as 256-bit
The AVX versions of these operands only operate on 256-bit ymm
registers, so we can specify this directly to be a little more
self-documenting.
2022-12-14 01:51:36 +00:00
lioncash d6ab7a4f97 Frontend: Handle 256-bit destination sizes directly
Previously the only time we'd promote to a 256-bit size is if the VEX.L
bit was set in the 128-bit path.

Allow specifying 256-bit sizes directly.
2022-12-14 01:50:05 +00:00
lioncash 7114fb3293 OpcodeDispatcher: Handle VPERMPD 2022-12-14 01:37:27 +00:00
lioncash 8a87aff730 OpcodeDispatcher: Handle VPERMQ 2022-12-14 01:30:55 +00:00
Ryan Houdek ded257c92f Merge pull request #2237 from lioncash/hadd
OpcodeDispatcher: Handle VHADDP{D, S}
2022-12-13 16:00:13 -08:00
lioncash c5b4719793 OpcodeDispatcher: Handle VHADDPD 2022-12-13 23:45:42 +00:00
lioncash 0f6201108f OpcodeDispatcher: Handle VHADDPS 2022-12-13 23:45:38 +00:00
lioncash b589dce7f5 x86_64/VectorOps: Make VFADDP behavior consistent with ARMv8
We need to swap the second and third results to be consistent with ARM.
Also fixes the mistake where I used vhaddpd instead of vhaddps in the
single-precision 256-bit case.
2022-12-13 23:34:12 +00:00
Ryan Houdek 9de5840f7a Merge pull request #2236 from lioncash/max
OpcodeDispatcher: Handle VPMAXS{B, D, W}/VPMAXU{B, D, W}
2022-12-13 14:11:17 -08:00
lioncash c98fffd33d OpcodeDispatcher: Handle VPMAXSD 2022-12-13 21:49:40 +00:00
lioncash de3777cc78 OpcodeDispatcher: Handle VPMAXSW 2022-12-13 21:47:17 +00:00
lioncash dd640e7a3d OpcodeDispatcher: Handle VPMAXSB 2022-12-13 21:44:46 +00:00
lioncash d53ddb73bf OpcodeDispatcher: Handle VPMAXUD 2022-12-13 21:40:29 +00:00
lioncash 25e9333abb OpcodeDispatcher: Handle VPMAXUW 2022-12-13 21:38:52 +00:00
lioncash 85766dd074 OpcodeDispatcher: Handle VPMAXUB 2022-12-13 21:36:47 +00:00
Ryan Houdek 40bab6b58e Merge pull request #2235 from lioncash/min
OpcodeDispatcher: Handle VPMINS{B, D, W}/VPMINU{B, D, W}
2022-12-13 13:31:14 -08:00
lioncash b0e0a2a165 OpcodeDispatcher: Handle VPMINSD 2022-12-13 20:51:14 +00:00
lioncash 0efcb912b5 OpcodeDispatcher: Handle VPMINSW 2022-12-13 20:49:08 +00:00
lioncash 9d0cc58737 OpcodeDispatcher: Handle VPMINSB 2022-12-13 20:47:00 +00:00
lioncash b671ed57ef OpcodeDispatcher: Handle VPMINUD 2022-12-13 20:41:53 +00:00
lioncash a2a44d188a OpcodeDispatcher: Handle VPMINUW 2022-12-13 20:39:55 +00:00
lioncash 364064536b OpcodeDispatcher: Handle VPMINUB 2022-12-13 20:36:51 +00:00
Mai 98a454169d Merge pull request #2234 from lioncash/sadd
OpcodeDispatcher: Handle VPADDS{B, W}/VPSUBS{B, W}
2022-12-13 20:25:25 +00:00
lioncash 6d44370f28 OpcodeDispatcher: Handle VPSUBSW 2022-12-13 18:52:55 +00:00
lioncash 273e2977a8 OpcodeDispatcher: Handle VPSUBSB 2022-12-13 18:50:02 +00:00
lioncash 7264b07d4f OpcodeDispatcher: Handle VPADDSW 2022-12-13 18:42:11 +00:00
lioncash 92351e7f33 OpcodeDispatcher: Handle VPADDSB 2022-12-13 18:40:08 +00:00
Ryan Houdek 757602bb1e Merge pull request #2233 from lioncash/uadd
OpcodeDispatcher: Handle VPADDUS{B, W}/VPSUBUS{B, W}
2022-12-13 10:25:59 -08:00
lioncash 6aaffec67f OpcodeDispatcher: Handle VPSUBUSW 2022-12-13 18:05:46 +00:00
lioncash 6f474cedd3 OpcodeDispatcher: Handle VPSUBUSB 2022-12-13 18:03:13 +00:00
lioncash d9176114c5 OpcodeDispatcher: Handle VPADDUSW 2022-12-13 18:03:07 +00:00
lioncash 287cee5b41 OpcodeDispatcher: Handle VPADDUSB 2022-12-13 18:02:55 +00:00
Mai a90067fb1e Merge pull request #2232 from lioncash/psub
OpcodeDispatcher: Handle VPSUB{B, D, Q, W}
2022-12-13 17:41:49 +00:00
lioncash 90d23098db OpcoodeDispatcher: Handle VPSUBQ 2022-12-13 06:01:19 +00:00
lioncash 384a09bbf1 OpcoodeDispatcher: Handle VPSUBD 2022-12-13 05:58:37 +00:00
lioncash 2fd29c47d4 OpcoodeDispatcher: Handle VPSUBW 2022-12-13 05:58:34 +00:00
lioncash e8aa8d89ec OpcoodeDispatcher: Handle VPSUBB 2022-12-13 05:47:50 +00:00
Ryan Houdek 1bc013d5f0 Merge pull request #2231 from lioncash/psign
OpcodeDispatcher: Handle VPSIGN{B, D, W}
2022-12-12 21:43:59 -08:00
lioncash 469ff91311 OpcodeDispatcher: Handle VPSIGND 2022-12-13 05:24:35 +00:00
lioncash ef14c411ce OpcodeDispatcher: Handle VPSIGNW 2022-12-13 05:20:12 +00:00
lioncash c2c5d176e4 OpcodeDispatcher: Handle VPSIGNB 2022-12-13 05:13:26 +00:00
lioncash 2328430d2e OpcodeDispatcher: Factor PSIGN handling into helper
This will allow us to use this with VEX and non-VEX variants without
needing to insert the upper-lane clearing for 128-bit variants into the
non-VEX path.

That, and this also allows us to not need to add additional template
arguments
2022-12-13 05:13:17 +00:00
Ryan Houdek a07a533640 Merge pull request #2230 from lioncash/div
OpcodeDispatcher: Handle VDIVP{D, S}/VDIVS{D, S}
2022-12-12 20:40:26 -08:00
lioncash 8c59e3e9e2 OpcodeDispatcher: Handle VDIVSD 2022-12-13 04:28:19 +00:00
lioncash fed861fa6b OpcodeDispatcher: Handle VDIVSS 2022-12-13 04:22:00 +00:00
lioncash ce9969ee8f OpcodeDispatcher: Handle VDIVPD 2022-12-13 04:16:33 +00:00
lioncash 9330ca41ea OpcodeDispatcher: Handle VDIVPS 2022-12-13 04:12:32 +00:00
Ryan Houdek eefcea49f4 Merge pull request #2229 from lioncash/mul
OpcodeDispatcher: Handle VMULP{D, S}/VMULS{D, S}
2022-12-12 20:03:43 -08:00
lioncash ed1b060494 OpcodeDispatcher: Handle VMULSD 2022-12-13 03:46:23 +00:00
lioncash 6db165e24a OpcodeDispatcher: Handle VMULSS 2022-12-13 03:42:08 +00:00
lioncash 58d20f199e OpcodeDispatcher: Handle VMULPD 2022-12-13 03:34:40 +00:00
lioncash 437ab47ae7 OpcodeDispatcher: Handle VMULPS 2022-12-13 03:29:35 +00:00
Ryan Houdek d6b137e6b7 Merge pull request #2228 from lioncash/min
OpcodeDispatcher: Handle VMAXP{D, S}/VMAXS{D, S}/VMINP{D, S}/VMINS{D, S}
2022-12-12 19:23:02 -08:00
lioncash e1de89af79 OpcodeDispatcher: Handle VMAXSD 2022-12-13 03:07:55 +00:00
lioncash 42d24c21e1 OpcodeDispatcher: Handle VMAXSS 2022-12-13 03:07:55 +00:00
lioncash 3590f090c7 OpcodeDispatcher: Handle VMAXPD 2022-12-13 03:07:55 +00:00
lioncash b7e177c11c OpcodeDispatcher: Handle VMAXPS 2022-12-13 03:07:55 +00:00
lioncash 92f92ddbbe OpcodeDispatcher: Handle VMINSD 2022-12-13 03:07:55 +00:00
lioncash f8d851b9b5 OpcodeDispatcher: Handle VMINSS 2022-12-13 03:07:55 +00:00
lioncash 1689742e96 OpcodeDispatcher: Handle VMINPD 2022-12-13 03:07:52 +00:00
lioncash 462b6b8c1c OpcodeDispatcher: Handle VMINPS 2022-12-13 01:48:57 +00:00
Ryan Houdek db90390179 Merge pull request #2227 from lioncash/sub
OpcodeDispatcher: Handle VSUBP{D, S}/ VSUBS{D, S}
2022-12-12 13:50:31 -08:00
lioncash e15fa66225 OpcodeDispatcher: Handle VSUBSD 2022-12-12 21:37:41 +00:00
lioncash 04a1fa6dc2 OpcodeDispatcher: Handle VSUBSS 2022-12-12 21:23:37 +00:00
lioncash f5a337a142 OpcodeDispatcher: Handle VSUBPD 2022-12-12 21:09:53 +00:00
lioncash 2b9d0314ce OpcodeDispatcher: Handle VSUBPS 2022-12-12 21:05:30 +00:00
Ryan Houdek 293734408c Merge pull request #2225 from lioncash/rcps
OpcodeDispatcher: Handle VRCPPS/VRCPSS
2022-12-12 12:04:46 -08:00
Ryan Houdek 03fbb923b3 Merge pull request #2226 from lioncash/lddqu
OpcodeDispatcher: Handle VLDDQU
2022-12-12 12:03:53 -08:00
lioncash 39f0c8542f OpcodeDispatcher: Handle VRCPSS 2022-12-12 19:48:21 +00:00
lioncash 6877d5b3ec OpcodeDispatcher: Handle VRCPPS 2022-12-12 19:48:21 +00:00
lioncash a6c30b35dc OpcodeDispatcher: Handle VLDDQU 2022-12-12 19:41:29 +00:00
Ryan Houdek a57f3a6264 Merge pull request #2224 from lioncash/abs
OpcodeDispatcher: Handle VPABS{B, D, W}
2022-12-12 11:28:07 -08:00
lioncash fa0ff71ddf OpcodeDispatcher: Handle VPABSD 2022-12-12 18:47:10 +00:00
lioncash c91ccf2cbe OpcodeDispatcher: Handle VPABSW 2022-12-12 18:47:10 +00:00
lioncash 41df5f816d OpcodeDispatcher: Handle VPABSB 2022-12-12 18:47:10 +00:00
Ryan Houdek 573896d0b7 Merge pull request #2223 from lioncash/pcmp
OpcodeDispatcher: Handle VPCMPEQ{B, D, Q, W}/VPCMPGT{B, D, Q, W}
2022-12-12 10:32:40 -08:00
lioncash 36a6264571 OpcodeDispatcher: Handle VPCMPEQQ 2022-12-12 18:00:38 +00:00
lioncash 12f01bc93a OpcodeDispatcher: Handle VPCMPEQD 2022-12-12 17:55:07 +00:00
lioncash 777b2c7966 OpcodeDispatcher: Handle VPCMPEQW 2022-12-12 17:51:36 +00:00
lioncash f0141f124d OpcodeDispatcher: Handle VPCMPEQB 2022-12-12 17:42:29 +00:00
lioncash 283b178285 OpcodeDispatcher: Handle VPCMPGTQ 2022-12-12 17:29:36 +00:00
lioncash d3a5eef08a OpcodeDispatcher: Handle VPCMPGTD 2022-12-12 17:13:46 +00:00
lioncash 1bac33ff44 OpcodeDispatcher: Handle VPCMPGTW 2022-12-12 17:13:46 +00:00
lioncash 327a6f52fd OpcodeDispatcher: Handle VPCMPGTB 2022-12-12 17:13:46 +00:00
Ryan Houdek 4f313f5d40 Merge pull request #2219 from Sonicadvance1/handle_pf_write
Dispatcher: Calculate REG_ERR correctly using ARM ESR_EL1
2022-12-12 09:02:27 -08:00
Ryan Houdek b42b4e03a4 Merge pull request #2218 from Sonicadvance1/GOT_optimization
OpCodeDispatcher: Optimize a case of GOT calculation
2022-12-12 09:02:18 -08:00
Ryan Houdek ab14375a03 Merge pull request #2222 from lioncash/rsqrt
OpcodeDispatcher: Handle VRSQRTSS/VRSQRTPS
2022-12-12 09:02:04 -08:00
Ryan Houdek ace90aac95 Merge pull request #2221 from lioncash/pbroad
OpcodeDispatcher: Handle VPBROADCAST{B, D, Q, W}/VBROADCASTI128
2022-12-12 09:01:56 -08:00
lioncash c4c93f5bfe OpcodeDispatcher: Handle VRSQRTSS 2022-12-12 16:30:34 +00:00
lioncash 3504ba068e OpcodeDispatcher: Handle VRSQRTPS 2022-12-12 16:11:16 +00:00
lioncash 88b88c9cd3 OpcodeDispatcher: Handle VBROADCASTI128 2022-12-12 15:51:12 +00:00
lioncash e99928990e OpcodeDispatcher: Handle VPBROADCASTQ 2022-12-12 15:41:41 +00:00
lioncash 6733f83471 OpcodeDispatcher: Handle VPBROADCASTD 2022-12-12 15:37:59 +00:00
lioncash a14cce27a4 OpcodeDispatcher: Handle VPBROADCASTW 2022-12-12 15:34:17 +00:00
lioncash 04d5b53389 OpcodeDispatcher: Handle VPBROADCASTB 2022-12-12 15:31:17 +00:00
Ryan Houdek a6b0181cd4 OpCodeDispatcher: Optimize a case of GOT calculation
32-bit GOT calculation needs to do a call+pop to do get the EIP on
32-bit. LEA doesn't work because it there is no EIP relative ops like on
x86-64.

This causes a terrible block split on every GOT calculation without the
optimization in place.

Now the block can continue through this weird GOT calculation.

This will be worthwhile for our 32-bit thunks where for some reason the
GOT calculation can't be removed. The GOT is calculated even though it
isn't used.
2022-12-10 02:50:48 -08:00
Ryan Houdek 3afd5691a4 unittests: Adds unit test to test ERR 2022-12-09 15:44:22 -08:00
Ryan Houdek 82ad26307c Dispatcher: Calculate REG_ERR correctly using ARM ESR_EL1
On Fault then ARM will return information about the fault in ESR_EL1 to
the user.

We need to decode what ESR_EL1 means in the context of the fault to get
the flags we care about.

The flags we care about specifically are PF_USER and PF_WRITE.
PF_PROT would have been interesting but I didn't see when this gets
returned to the user. ARM makes the difference if the page is unmapped
or "mapped" with PROT_NONE. x86 doesn't make the distinction here.

This should fix an issue that a user was hitting.
2022-12-09 15:44:11 -08:00
Mai dc9737a394 Merge pull request #2217 from Sonicadvance1/fix_global_app_config
Config: Fixes global application configs
2022-12-09 18:02:08 +00:00
Ryan Houdek eaef06d14e Config: Fixes global application configs
Accidentally was checking for SteamID layer types twice, rather than
global.

Fixes steamwebhelper config not getting loaded from global config.
2022-12-09 08:26:51 -08:00
Ryan Houdek 2123868a42 Merge pull request #2215 from lioncash/broadcast
OpcodeDispatcher: Handle VBROADCASTSD/VBROADCASTSD/VBROADCASTF128
2022-12-07 19:50:19 -08:00
lioncash b891999a7f OpcodeDispatcher: Handle VBROADCASTF128 2022-12-08 03:18:58 +00:00
lioncash a53fd07bda OpcodeDispatcher: Handle VBROADCASTSD 2022-12-08 02:58:12 +00:00
lioncash 8f213b75be OpcodeDispatcher: Handle VBROADCASTSS 2022-12-08 02:40:36 +00:00
Ryan Houdek b73aeb8902 Merge pull request #2214 from lioncash/stmxcsr
OpcodeDispatcher: Handle VLDMXCSR/VSTMXCSR
2022-12-07 17:07:07 -08:00
lioncash d642c1a646 OpcodeDispatcher: Handle VLDMXCSR/VSTMXCSR 2022-12-08 00:42:13 +00:00
Ryan Houdek d965ae03c4 Merge pull request #2210 from lioncash/sqrt
OpcodeDispatcher: Handle VSQRTPD/VSQRTPS/VSQRTSD/VSQRTSS
2022-12-07 14:53:16 -08:00
lioncash e42de0b645 OpcodeDispatcher: Handle VSQRTSD 2022-12-07 22:42:16 +00:00
lioncash 9ef5247dd7 OpcodeDispatcher: Handle VSQRTSS 2022-12-07 22:42:16 +00:00
lioncash 25428cb28c OpcodeDispatcher: Handle VSQRTPD 2022-12-07 22:42:16 +00:00
lioncash 2125949d6d OpcodeDispatcher: Handle VSQRTPS 2022-12-07 22:42:16 +00:00
Ryan Houdek 7ac21e794d Merge pull request #2212 from lioncash/comiss
OpcodeDispatcher: Handle VCOMISD/VCOMISS/VUCOMISD/VUCOMISS
2022-12-07 14:38:49 -08:00
lioncash 4aa0f3d0a4 OpcodeDispatcher: Handle VCOMISD 2022-12-07 22:06:25 +00:00
lioncash 740c983f65 OpcodeDispatcher: Handle VCOMISS 2022-12-07 22:06:25 +00:00
lioncash 83bccc0032 OpcodeDispatcher: Handle VUCOMISD 2022-12-07 22:06:22 +00:00
Ryan Houdek a98920d4e4 Merge pull request #2211 from lioncash/avg
OpcodeDispatcher: Handle VPAVGB/VPAVGW
2022-12-07 13:37:58 -08:00
lioncash d1ab636df1 OpcodeDispatcher: Handle VUCOMISS 2022-12-07 21:22:19 +00:00
lioncash 26b629833e OpcodeDispatcher: Handle VPAVGW 2022-12-07 21:07:43 +00:00
lioncash 95964f8dd8 OpcodeDispatcher: Handle VPAVGB 2022-12-07 21:07:40 +00:00
Ryan Houdek 94ae2e3a9c Merge pull request #2209 from lioncash/ins
IR: Handle 128-bit VInsElement with SVE
2022-12-07 11:08:47 -08:00
lioncash eae33b0c50 IR: Handle 128-bit VInsElement with SVE
Currently VDupElement allows duplicating 128-bit elements in 256-bit
vectors with SVE, so we can extend VInsElement to have similar behavior.
2022-12-07 18:41:45 +00:00
Ryan Houdek e9aa368a62 Merge pull request #2208 from lioncash/zero
OpcodeDispatcher: Explicitly zero upper lanes
2022-12-07 10:13:53 -08:00
lioncash 5a37786da7 OpcodeDispatcher: Explicitly zero upper lanes
Makes our intent to zero-extend the upper lanes explicit and lets us
remove a special case in the LoadSource implementation.

This also makes things a little nicer since we're not hardcoding 32 byte
stores.
2022-12-07 17:22:20 +00:00
Ryan Houdek 5ac44baa2c Merge pull request #2207 from lioncash/adds
OpcodeDispatcher: Handle VADDSD/VADDSS
2022-12-07 09:19:59 -08:00
lioncash 4cf3805950 OpcodeDispatcher: Handle VADDSD 2022-12-07 16:43:47 +00:00
lioncash 1f5a1826a6 OpcodeDispatcher: Handle VADDSS 2022-12-07 16:43:44 +00:00
Ryan Houdek bf86df7a66 Merge pull request #2206 from lioncash/haddp
OpcodeDispatcher: Merge HADDP/PHADD into VectorALUOp
2022-12-07 07:39:19 -08:00
lioncash 6a63ae2d9c OpcodeDispatcher: Merge PHADD into VectorALUOp 2022-12-07 15:18:19 +00:00
lioncash a7a1e2abd3 OpcodeDispatcher: Merge HADDP into VectorALUOp 2022-12-07 15:14:10 +00:00
Ryan Houdek 3322f8b890 Merge pull request #2205 from lioncash/pavg
OpcodeDispatcher: Merge PAVGOp with VectorALUOp
2022-12-07 07:08:48 -08:00
lioncash 047ae13c98 OpcodeDispatcher: Merge PAVGOp with VectorALUOp
This can be merged into it, considering it only has one IR op.
2022-12-07 14:51:08 +00:00
Ryan Houdek 9eaa45f922 Merge pull request #2202 from lioncash/addmerge
OpcodeDispatcher: Merge PADDQOp, PSUBQOp, PADDSOp, PSUBSOp with VectorALUOp
2022-12-05 17:48:59 -08:00
lioncash 3d5e0c5832 OpcodeDispatcher: Merge PSUBSOp with VectorALUOp 2022-12-06 01:31:26 +00:00
lioncash 71a36763df OpcodeDispatcher: Merge PADDSOp with VectorALUOp 2022-12-06 01:27:31 +00:00
lioncash 07bd3137ef OpcodeDispatcher: Merge PSUBQOp with VectorALUOp 2022-12-06 01:21:33 +00:00
lioncash c2b40b4dd3 OpcodeDispather: Merge PADDQOp with VectorALUOp 2022-12-06 01:13:12 +00:00
Ryan Houdek 4b16718602 Merge pull request #2201 from lioncash/bic
OpcodeDispatcher: Merge ANDNOp with VectorALUROp
2022-12-05 15:15:31 -08:00
Ryan Houdek 2bf7e09862 Merge pull request #2200 from lioncash/bic
OpcodeDispatcher: Simplify VANDN
2022-12-05 14:31:33 -08:00
lioncash db74e46fc3 OpcodeDispatcher: Merge ANDNOp with VectorALUROp
Now that ANDNOp is reduced to one IR op, we can merge it with
VectorALUROp
2022-12-05 22:28:39 +00:00
lioncash 13bba59ba6 OpcodeDispatcher: Simplify ANDNOp and VANDNOp
We can just use VBic here to simplify everything.
2022-12-05 22:15:15 +00:00
498 changed files with 57970 additions and 11259 deletions

No files matched your search

+16 -6
View File
@@ -72,6 +72,11 @@ jobs:
# Execute the build. You can specify a specific target with "--target <NAME>"
run: cmake --build . --config $BUILD_TYPE
- name: Install
working-directory: ${{runner.workspace}}/build
shell: bash
run: cmake --build . --config $BUILD_TYPE --target install
- name: ASM Tests
working-directory: ${{runner.workspace}}/build
shell: bash
@@ -166,6 +171,17 @@ jobs:
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_APITests.log || true
- name: ARMEmitter tests
working-directory: ${{runner.workspace}}/build
shell: bash
run: cmake --build . --config $BUILD_TYPE --target emitter_tests
- name: ARMEmitter Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_ARMEmitterTests.log || true
- name: FEXLinuxTests
working-directory: ${{runner.workspace}}/build
shell: bash
@@ -188,12 +204,6 @@ jobs:
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_ThunkgenTests.log || true
- name: Install
if: matrix.arch[1] == 'x64'
working-directory: ${{runner.workspace}}/build
shell: bash
run: cmake --build . --config $BUILD_TYPE --target install
- name: Test GL No-Thunks
if: matrix.arch[1] == 'x64'
working-directory: ${{runner.workspace}}/build
+7
View File
@@ -29,6 +29,8 @@ option(ENABLE_INTERPRETER "Enables FEX's Interpreter" FALSE)
option(ENABLE_CCACHE "Enables ccache for compile caching" TRUE)
option(ENABLE_TERMUX_BUILD "Forces building for Termux on a non-Termux build machine" FALSE)
option(ENABLE_VIXL_SIMULATOR "Forces the FEX JIT to use the VIXL simulator" FALSE)
option(ENABLE_VIXL_DISASSEMBLER "Enables debug disassembler output with VIXL" FALSE)
option(COMPILE_VIXL_DISASSEMBLER "Compiles the vixl disassembler in to vixl" FALSE)
option(ENABLE_FEXCORE_PROFILER "Enables use of the FEXCore timeline profiling capabilities" FALSE)
set (FEXCORE_PROFILER_BACKEND "gpuvis" CACHE STRING "Set which backend you want to use for the FEXCore profiler")
@@ -196,6 +198,11 @@ set (CMAKE_LINKER_FLAGS_RELEASE "${CMAKE_LINKER_FLAGS_RELEASE} -fomit-frame-poin
include_directories(External/robin-map/include/)
if (BUILD_TESTS)
# Enable vixl disassembler if tests are enabled.
set(COMPILE_VIXL_DISASSEMBLER TRUE)
endif()
add_subdirectory(External/vixl/)
include_directories(External/vixl/src/)
-5
View File
@@ -1,5 +0,0 @@
{
"Config": {
"x86dec_SynchronizeRIPOnAllBlocks": "1"
}
}
-5
View File
@@ -1,5 +0,0 @@
{
"Config": {
"x86dec_SynchronizeRIPOnAllBlocks": "1"
}
}
+5
View File
@@ -0,0 +1,5 @@
{
"Config": {
"HideHypervisorBit": "1"
}
}
-5
View File
@@ -1,5 +0,0 @@
{
"Config": {
"x86dec_SynchronizeRIPOnAllBlocks": "1"
}
}
-5
View File
@@ -1,5 +0,0 @@
{
"Config": {
"x86dec_SynchronizeRIPOnAllBlocks": "1"
}
}
+69 -69
View File
@@ -6,10 +6,10 @@
"X11"
],
"Overlay": [
"@PREFIX_LIB@/x86_64-linux-gnu/libGL.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libGL.so.1",
"@PREFIX_LIB@/x86_64-linux-gnu/libGL.so.1.2.0",
"@PREFIX_LIB@/x86_64-linux-gnu/libGL.so.1.7.0"
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libGL.so",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libGL.so.1",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libGL.so.1.2.0",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libGL.so.1.7.0"
]
},
"GLESv2": {
@@ -18,17 +18,17 @@
"X11"
],
"Overlay": [
"@PREFIX_LIB@/x86_64-linux-gnu/libGLESv2.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libGLESv2.so.2",
"@PREFIX_LIB@/x86_64-linux-gnu/libGLESv2.so.2.0.0"
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libGLESv2.so",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libGLESv2.so.2",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libGLESv2.so.2.0.0"
]
},
"X11": {
"Library": "libX11-guest.so",
"Overlay": [
"@PREFIX_LIB@/x86_64-linux-gnu/libX11.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libX11.so.6",
"@PREFIX_LIB@/x86_64-linux-gnu/libX11.so.6.4.0"
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libX11.so",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libX11.so.6",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libX11.so.6.4.0"
]
},
"Vulkan": {
@@ -37,8 +37,8 @@
"xcb"
],
"Overlay": [
"@PREFIX_LIB@/x86_64-linux-gnu/libvulkan.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libvulkan.so.1",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libvulkan.so",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libvulkan.so.1",
"@HOME@/.local/share/Steam/ubuntu12_32/steam-runtime/pinned_libs_64/libvulkan.so.1"
],
"Comment": [
@@ -48,129 +48,129 @@
"xcb": {
"Library": "libxcb-guest.so",
"Overlay": [
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb.so.1",
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb.so.1.1.0"
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxcb.so",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxcb.so.1",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxcb.so.1.1.0"
]
},
"xcb-dri2": {
"Library": "libxcb_dri2-guest.so",
"Library": "libxcb-dri2-guest.so",
"Overlay": [
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-dri2.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-dri2.so.0",
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-dri2.so.0.0.0"
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxcb-dri2.so",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxcb-dri2.so.0",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxcb-dri2.so.0.0.0"
]
},
"xcb-dri3": {
"Library": "libxcb_dri3-guest.so",
"Library": "libxcb-dri3-guest.so",
"Overlay": [
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-dri3.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-dri3.so.0",
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-dri3.so.0.0.0"
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxcb-dri3.so",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxcb-dri3.so.0",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxcb-dri3.so.0.0.0"
]
},
"xcb-xfixes": {
"Library": "libxcb_xfixes-guest.so",
"Library": "libxcb-xfixes-guest.so",
"Overlay": [
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-xfixes.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-xfixes.so.0",
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-xfixes.so.0.0.0"
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxcb-xfixes.so",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxcb-xfixes.so.0",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxcb-xfixes.so.0.0.0"
]
},
"xcb-shm": {
"Library": "libxcb_shm-guest.so",
"Library": "libxcb-shm-guest.so",
"Overlay": [
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-shm.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-shm.so.0",
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-shm.so.0.0.0"
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxcb-shm.so",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxcb-shm.so.0",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxcb-shm.so.0.0.0"
]
},
"xcb-sync": {
"Library": "libxcb_sync-guest.so",
"Library": "libxcb-sync-guest.so",
"Overlay": [
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-sync.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-sync.so.1",
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-sync.so.1.0.0"
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxcb-sync.so",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxcb-sync.so.1",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxcb-sync.so.1.0.0"
]
},
"xcb-randr": {
"Library": "libxcb_randr-guest.so",
"Library": "libxcb-randr-guest.so",
"Overlay": [
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-randr.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-randr.so.0",
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-randr.so.0.1.0"
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxcb-randr.so",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxcb-randr.so.0",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxcb-randr.so.0.1.0"
]
},
"xcb-present": {
"Library": "libxcb_present-guest.so",
"Library": "libxcb-present-guest.so",
"Overlay": [
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-present.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-present.so.0",
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-present.so.0.0.0"
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxcb-present.so",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxcb-present.so.0",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxcb-present.so.0.0.0"
]
},
"xcb-glx": {
"Library": "libxcb_glx-guest.so",
"Library": "libxcb-glx-guest.so",
"Overlay": [
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-glx.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-glx.so.0",
"@PREFIX_LIB@/x86_64-linux-gnu/libxcb-glx.so.0.0.0"
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxcb-glx.so",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxcb-glx.so.0",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxcb-glx.so.0.0.0"
]
},
"xshmfence": {
"Library": "libshmfence-guest.so",
"Library": "libxshmfence-guest.so",
"Overlay": [
"@PREFIX_LIB@/x86_64-linux-gnu/libxshmfence.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libxshmfence.so.1",
"@PREFIX_LIB@/x86_64-linux-gnu/libxshmfence.so.1.0.0"
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxshmfence.so",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxshmfence.so.1",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxshmfence.so.1.0.0"
]
},
"drm": {
"Library": "libdrm-guest.so",
"Overlay": [
"@PREFIX_LIB@/x86_64-linux-gnu/libdrm.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libdrm.so.2",
"@PREFIX_LIB@/x86_64-linux-gnu/libdrm.so.2.4.0"
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libdrm.so",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libdrm.so.2",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libdrm.so.2.4.0"
]
},
"asound": {
"Library": "libasound-guest.so",
"Overlay": [
"@PREFIX_LIB@/x86_64-linux-gnu/libasound.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libasound.so.2",
"@PREFIX_LIB@/x86_64-linux-gnu/libasound.so.2.0.0"
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libasound.so",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libasound.so.2",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libasound.so.2.0.0"
]
},
"Xrender": {
"Library": "libXrender-guest.so",
"Overlay": [
"@PREFIX_LIB@/x86_64-linux-gnu/libXrender.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libXrender.so.1",
"@PREFIX_LIB@/x86_64-linux-gnu/libXrender.so.1.3.0"
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libXrender.so",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libXrender.so.1",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libXrender.so.1.3.0"
]
},
"Xext": {
"Library": "libXext-guest.so",
"Overlay": [
"@PREFIX_LIB@/x86_64-linux-gnu/libXext.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libXext.so.6",
"@PREFIX_LIB@/x86_64-linux-gnu/libXext.so.6.4.0"
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libXext.so",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libXext.so.6",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libXext.so.6.4.0"
]
},
"Xfixes": {
"Library": "libXfixes-guest.so",
"Overlay": [
"@PREFIX_LIB@/x86_64-linux-gnu/libXfixes.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libXfixes.so.3",
"@PREFIX_LIB@/x86_64-linux-gnu/libXfixes.so.3.1.0"
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libXfixes.so",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libXfixes.so.3",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libXfixes.so.3.1.0"
]
},
"OpenCL": {
"Library" : "libOpenCL-guest.so",
"Overlay": [
"@PREFIX_LIB@/x86_64-linux-gnu/libOpenCL.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libOpenCL.so.1",
"@PREFIX_LIB@/x86_64-linux-gnu/libOpenCL.so.1.0.0"
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libOpenCL.so",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libOpenCL.so.1",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libOpenCL.so.1.0.0"
]
},
"":{}
+4
View File
@@ -82,3 +82,7 @@ add_subdirectory(Source/)
install (DIRECTORY include/FEXCore ${CMAKE_BINARY_DIR}/include/FEXCore
DESTINATION include
COMPONENT Development)
if (BUILD_TESTS)
add_subdirectory(unittests/)
endif()
+38 -11
View File
@@ -281,9 +281,7 @@ def print_ir_structs(defines):
output_file.write("\tvoid* Data[0];\n")
output_file.write("\tIROps Op;\n\n")
output_file.write("\tuint8_t Size;\n")
output_file.write("\tuint8_t NumArgs;\n")
output_file.write("\tuint8_t ElementSize : 7;\n")
output_file.write("\tbool HasDest : 1;\n")
output_file.write("\tuint8_t ElementSize;\n")
output_file.write("\ttemplate<typename T>\n")
output_file.write("\tT const* C() const { return reinterpret_cast<T const*>(Data); }\n")
@@ -358,8 +356,10 @@ def print_ir_sizes():
output_file.write("[[nodiscard, gnu::const, gnu::visibility(\"default\")]] std::string_view const& GetName(IROps Op);\n")
output_file.write("[[nodiscard, gnu::const, gnu::visibility(\"default\")]] uint8_t GetArgs(IROps Op);\n")
output_file.write("[[nodiscard, gnu::const, gnu::visibility(\"default\")]] uint8_t GetRAArgs(IROps Op);\n")
output_file.write("[[nodiscard, gnu::const, gnu::visibility(\"default\")]] FEXCore::IR::RegisterClassType GetRegClass(IROps Op);\n\n")
output_file.write("[[nodiscard, gnu::const, gnu::visibility(\"default\")]] bool HasSideEffects(IROps Op);\n")
output_file.write("[[nodiscard, gnu::const, gnu::visibility(\"default\")]] bool GetHasDest(IROps Op);\n")
output_file.write("#undef IROP_SIZES\n")
output_file.write("#endif\n\n")
@@ -417,7 +417,7 @@ def print_ir_getname():
def print_ir_getraargs():
output_file.write("#ifdef IROP_GETRAARGS_IMPL\n")
output_file.write("constexpr std::array<uint8_t, OP_LAST + 1> IRArgs = {\n")
output_file.write("constexpr std::array<uint8_t, OP_LAST + 1> IRRAArgs = {\n")
for op in IROps:
SSAArgs = op.SSAArgNum
@@ -430,6 +430,18 @@ def print_ir_getraargs():
output_file.write("};\n\n")
output_file.write("constexpr std::array<uint8_t, OP_LAST + 1> IRArgs = {\n")
for op in IROps:
SSAArgs = op.SSAArgNum
output_file.write("\t{},\n".format(SSAArgs))
output_file.write("};\n\n")
output_file.write("uint8_t GetRAArgs(IROps Op) {\n")
output_file.write(" return IRRAArgs[Op];\n")
output_file.write("}\n")
output_file.write("uint8_t GetArgs(IROps Op) {\n")
output_file.write(" return IRArgs[Op];\n")
output_file.write("}\n")
@@ -453,6 +465,25 @@ def print_ir_hassideeffects():
output_file.write("#undef IROP_HASSIDEEFFECTS_IMPL\n")
output_file.write("#endif\n\n")
def print_ir_gethasdest():
output_file.write("#ifdef IROP_GETHASDEST_IMPL\n")
output_file.write("constexpr std::array<bool, OP_LAST + 1> IRDest = {\n")
for op in IROps:
if op.HasDest:
output_file.write("\ttrue,\n")
else:
output_file.write("\tfalse,\n")
output_file.write("};\n\n")
output_file.write("bool GetHasDest(IROps Op) {\n")
output_file.write(" return IRDest[Op];\n")
output_file.write("}\n")
output_file.write("#undef IROP_GETHASDEST_IMPL\n")
output_file.write("#endif\n\n")
# Print out IR argument printing
def print_ir_arg_printer():
output_file.write("#ifdef IROP_ARGPRINTER_HELPER\n")
@@ -547,13 +578,13 @@ def print_ir_allocator_helpers():
output_file.write("\tuint8_t GetOpElements(const OrderedNode *Op) const {\n")
output_file.write("\t\tauto HeaderOp = Op->Header.Value.GetNode(DualListData.DataBegin());\n")
output_file.write("\t\tLOGMAN_THROW_A_FMT(HeaderOp->HasDest, \"Op {} has no dest\\n\", GetName(HeaderOp->Op));\n")
output_file.write("\t\tLOGMAN_THROW_A_FMT(OpHasDest(Op), \"Op {} has no dest\\n\", GetName(HeaderOp->Op));\n")
output_file.write("\t\treturn HeaderOp->Size / HeaderOp->ElementSize;\n")
output_file.write("\t}\n\n")
output_file.write("\tbool OpHasDest(const OrderedNode *Op) const {\n")
output_file.write("\t\tauto HeaderOp = Op->Header.Value.GetNode(DualListData.DataBegin());\n")
output_file.write("\t\treturn HeaderOp->HasDest;\n")
output_file.write("\t\treturn GetHasDest(HeaderOp->Op);\n")
output_file.write("\t}\n\n")
output_file.write("\tIROps GetOpType(const OrderedNode *Op) const {\n")
@@ -631,8 +662,6 @@ def print_ir_allocator_helpers():
output_file.write("\t\tOp.first->Header.Size = InferSize;\n")
output_file.write("\t\tOp.first->Header.NumArgs = {};\n".format(op.SSAArgNum))
# Some ops without a destination still need an operating size
# Effectively reusing the destination size value for operation size
if op.DestSize != None:
@@ -643,9 +672,6 @@ def print_ir_allocator_helpers():
else:
output_file.write("\t\tOp.first->Header.ElementSize = Op.first->Header.Size / ({});\n".format(op.NumElements))
if (op.HasDest):
output_file.write("\t\tOp.first->Header.HasDest = true;\n")
# Insert validation here
if op.EmitValidation != None:
output_file.write("\t\t#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED\n")
@@ -733,6 +759,7 @@ print_ir_reg_classes()
print_ir_getname()
print_ir_getraargs()
print_ir_hassideeffects()
print_ir_gethasdest()
print_ir_arg_printer()
print_ir_allocator_helpers()
print_ir_parser_switch_helper()
+4
View File
@@ -182,6 +182,10 @@ if (ENABLE_VIXL_SIMULATOR)
list(APPEND DEFINES -DVIXL_SIMULATOR=1 -DVIXL_INCLUDE_SIMULATOR_AARCH64=1)
endif()
if (ENABLE_VIXL_DISASSEMBLER)
list(APPEND DEFINES -DVIXL_DISASSEMBLER=1)
endif()
if (ENABLE_JIT_X86_64)
list(APPEND SRCS
Interface/Core/JIT/x86_64/JIT.cpp
+42 -20
View File
@@ -1,64 +1,86 @@
#include "Common/JitSymbols.h"
#include <fcntl.h>
#include <string>
#include <unistd.h>
#include <fmt/format.h>
namespace FEXCore {
JITSymbols::JITSymbols() : fp{nullptr, std::fclose} {
JITSymbols::JITSymbols() {
}
JITSymbols::~JITSymbols() = default;
void JITSymbols::InitFile() {
const auto PerfMap = fmt::format("/tmp/perf-{}.map", getpid());
fp.reset(fopen(PerfMap.c_str(), "wb"));
if (fp) {
// Disable buffering on this file
setvbuf(fp.get(), nullptr, _IONBF, 0);
JITSymbols::~JITSymbols() {
if (fd != -1) {
close(fd);
}
}
void JITSymbols::InitFile() {
// We can't use FILE here since we must be robust against forking processes closing our FD from under us.
const auto PerfMap = fmt::format("/tmp/perf-{}.map", getpid());
fd = open(PerfMap.c_str(), O_CREAT | O_TRUNC | O_WRONLY | O_APPEND, 0644);
}
void JITSymbols::Register(const void *HostAddr, uint64_t GuestAddr, uint32_t CodeSize) {
if (!fp) return;
if (fd == -1) return;
// Linux perf format is very straightforward
// `<HostPtr> <Size> <Name>\n`
fmt::print(fp.get(), "{} {:x} JIT_0x{:x}_{}\n", HostAddr, CodeSize, GuestAddr, HostAddr);
const auto Buffer = fmt::format("{} {:x} JIT_0x{:x}_{}\n", HostAddr, CodeSize, GuestAddr, HostAddr);
auto Result = write(fd, Buffer.c_str(), Buffer.size());
if (Result == -1 && errno == EBADF) {
fd = -1;
}
}
void JITSymbols::Register(const void *HostAddr, uint32_t CodeSize, std::string_view Name) {
if (!fp) return;
if (fd == -1) return;
// Linux perf format is very straightforward
// `<HostPtr> <Size> <Name>\n`
fmt::print(fp.get(), "{} {:x} {}_{}\n", HostAddr, CodeSize, Name, HostAddr);
const auto Buffer = fmt::format("{} {:x} {}_{}\n", HostAddr, CodeSize, Name, HostAddr);
auto Result = write(fd, Buffer.c_str(), Buffer.size());
if (Result == -1 && errno == EBADF) {
fd = -1;
}
}
void JITSymbols::Register(const void *HostAddr, uint32_t CodeSize, std::string_view Name, uintptr_t Offset) {
if (!fp) return;
if (fd == -1) return;
// Linux perf format is very straightforward
// `<HostPtr> <Size> <Name>\n`
fmt::print(fp.get(), "{} {:x} {}+0x{:x} ({})\n", HostAddr, CodeSize, Name, Offset, HostAddr);
const auto Buffer = fmt::format("{} {:x} {}+0x{:x} ({})\n", HostAddr, CodeSize, Name, Offset, HostAddr);
auto Result = write(fd, Buffer.c_str(), Buffer.size());
if (Result == -1 && errno == EBADF) {
fd = -1;
}
}
void JITSymbols::RegisterNamedRegion(const void *HostAddr, uint32_t CodeSize, std::string_view Name) {
if (!fp) return;
if (fd == -1) return;
// Linux perf format is very straightforward
// `<HostPtr> <Size> <Name>\n`
fmt::print(fp.get(), "{} {:x} {}\n", HostAddr, CodeSize, Name);
const auto Buffer = fmt::format("{} {:x} {}\n", HostAddr, CodeSize, Name);
auto Result = write(fd, Buffer.c_str(), Buffer.size());
if (Result == -1 && errno == EBADF) {
fd = -1;
}
}
void JITSymbols::RegisterJITSpace(const void *HostAddr, uint32_t CodeSize) {
if (!fp) return;
if (fd == -1) return;
// Linux perf format is very straightforward
// `<HostPtr> <Size> <Name>\n`
fmt::print(fp.get(), "{} {:x} FEXJIT\n", HostAddr, CodeSize);
const auto Buffer = fmt::format("{} {:x} FEXJIT\n", HostAddr, CodeSize);
auto Result = write(fd, Buffer.c_str(), Buffer.size());
if (Result == -1 && errno == EBADF) {
fd = -1;
}
}
} // namespace FEXCore
+1 -3
View File
@@ -19,8 +19,6 @@ public:
void RegisterJITSpace(const void *HostAddr, uint32_t CodeSize);
private:
using FILEPtr = std::unique_ptr<FILE, decltype(&std::fclose)>;
FILEPtr fp;
int fd{-1};
};
}
+2 -2
View File
@@ -30,7 +30,7 @@
#include <tiny-json.h>
namespace FEXCore::Context {
struct Context;
class Context;
}
namespace FEXCore::Config {
@@ -686,7 +686,7 @@ namespace JSON {
AppLoader::AppLoader(const std::string& Filename, FEXCore::Config::LayerType Type)
: FEXCore::Config::OptionMapper(Type) {
const bool Global = Type == FEXCore::Config::LayerType::LAYER_GLOBAL_STEAM_APP ||
Type == FEXCore::Config::LayerType::LAYER_LOCAL_STEAM_APP;
Type == FEXCore::Config::LayerType::LAYER_GLOBAL_APP;
Config = FEXCore::Config::GetApplicationConfig(Filename, Global);
// Immediately load so we can reload the meta layer
+21 -5
View File
@@ -91,6 +91,13 @@
"Folder to find the guest-side thunking libraries."
]
},
"ThunkHostLibs32": {
"Type": "str",
"Default": "@CMAKE_INSTALL_PREFIX@/lib/fex-emu/HostThunks_32/",
"Desc": [
"Folder to find the 32-bit host-side thunking libraries."
]
},
"ThunkGuestLibs32": {
"Type": "str",
"Default": "@CMAKE_INSTALL_PREFIX@/share/fex-emu/GuestThunks_32/",
@@ -233,6 +240,17 @@
"Also needs x86_64-linux-gnu-objdump in PATH.",
"Can be very slow."
]
},
"InjectLibSegFault": {
"Type": "bool",
"Default": "false",
"Desc": [
"Sets the environment variable LD_PRELOAD=libSegFault.so",
"This allows the user to very easily enable libSegFault without dealing with environment variables",
"Very useful for applications that have launch scripts that set the variable to nothing at launch",
"Set this in an application configuration for injecting in to only specific applications.",
"\tNote: If x86/x86_64 libSegFault.so isn't installed then this option won't work."
]
}
},
"Logging": {
@@ -325,14 +343,12 @@
"Useful for a process that keeps restarting and doesn't work"
]
},
"x86dec_SynchronizeRIPOnAllBlocks": {
"HideHypervisorBit": {
"Type": "bool",
"Default": "false",
"Desc": [
"An application that uses try-catch or longjump extensively needs the ability to do context aware state flushing",
"In the case of FEX's block-linking, it won't always ensure that RIP is synchronized.",
"If an exception occurs and RIP isn't synchronized, then FEX's exception stack restore may not long jump as expected",
"Can be useful for Wine applications that rely on stack unwinding"
"Hides the hypervisor CPUID bit when set.",
"Should only be used for applications that have issues with this set."
]
}
},
+57 -152
View File
@@ -28,203 +28,108 @@ namespace FEXCore::Context {
FEXCore::Paths::ShutdownPaths();
}
FEXCore::Context::Context *CreateNewContext() {
return new FEXCore::Context::Context{};
FEXCore::Context::Context *FEXCore::Context::Context::CreateNewContext() {
return new FEXCore::Context::ContextImpl{};
}
bool InitializeContext(FEXCore::Context::Context *CTX) {
return FEXCore::CPU::CreateCPUCore(CTX);
}
void DestroyContext(FEXCore::Context::Context *CTX) {
if (CTX->ParentThread) {
CTX->DestroyThread(CTX->ParentThread);
}
void FEXCore::Context::Context::DestroyContext(FEXCore::Context::Context *CTX) {
CTX->DestroyContext();
delete CTX;
}
FEXCore::Core::InternalThreadState* InitCore(FEXCore::Context::Context *CTX, uint64_t InitialRIP, uint64_t StackPointer) {
return CTX->InitCore(InitialRIP, StackPointer);
bool FEXCore::Context::ContextImpl::InitializeContext() {
return FEXCore::CPU::CreateCPUCore(this);
}
void SetExitHandler(FEXCore::Context::Context *CTX, ExitHandler handler) {
CTX->CustomExitHandler = std::move(handler);
void FEXCore::Context::ContextImpl::DestroyContext() {
if (ParentThread) {
DestroyThread(ParentThread);
}
}
ExitHandler GetExitHandler(const FEXCore::Context::Context *CTX) {
return CTX->CustomExitHandler;
void FEXCore::Context::ContextImpl::SetExitHandler(ExitHandler handler) {
CustomExitHandler = std::move(handler);
}
void Run(FEXCore::Context::Context *CTX) {
CTX->Run();
ExitHandler FEXCore::Context::ContextImpl::GetExitHandler() const {
return CustomExitHandler;
}
void Step(FEXCore::Context::Context *CTX) {
CTX->Step();
void FEXCore::Context::ContextImpl::Stop() {
Stop(false);
}
void CompileRIP(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP) {
Thread->CTX->CompileBlock(Thread->CurrentFrame, GuestRIP);
void FEXCore::Context::ContextImpl::CompileRIP(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP) {
CompileBlock(Thread->CurrentFrame, GuestRIP);
}
FEXCore::Context::ExitReason RunUntilExit(FEXCore::Context::Context *CTX) {
return CTX->RunUntilExit();
FEXCore::Context::ExitReason FEXCore::Context::ContextImpl::GetExitReason() {
return ParentThread->ExitReason;
}
int GetProgramStatus(const FEXCore::Context::Context *CTX) {
return CTX->GetProgramStatus();
bool FEXCore::Context::ContextImpl::IsDone() const {
return IsPaused();
}
FEXCore::Context::ExitReason GetExitReason(const FEXCore::Context::Context *CTX) {
return CTX->ParentThread->ExitReason;
void FEXCore::Context::ContextImpl::GetCPUState(FEXCore::Core::CPUState *State) const {
memcpy(State, ParentThread->CurrentFrame, sizeof(FEXCore::Core::CPUState));
}
bool IsDone(const FEXCore::Context::Context *CTX) {
return CTX->IsPaused();
void FEXCore::Context::ContextImpl::SetCPUState(const FEXCore::Core::CPUState *State) {
memcpy(ParentThread->CurrentFrame, State, sizeof(FEXCore::Core::CPUState));
}
void GetCPUState(const FEXCore::Context::Context *CTX, FEXCore::Core::CPUState *State) {
memcpy(State, CTX->ParentThread->CurrentFrame, sizeof(FEXCore::Core::CPUState));
void FEXCore::Context::ContextImpl::SetCustomCPUBackendFactory(CustomCPUFactoryType Factory) {
CustomCPUFactory = std::move(Factory);
}
void SetCPUState(FEXCore::Context::Context *CTX, const FEXCore::Core::CPUState *State) {
memcpy(CTX->ParentThread->CurrentFrame, State, sizeof(FEXCore::Core::CPUState));
}
void Pause(FEXCore::Context::Context *CTX) {
CTX->Pause();
}
void Stop(FEXCore::Context::Context *CTX) {
CTX->Stop(false);
}
void SetCustomCPUBackendFactory(FEXCore::Context::Context *CTX, CustomCPUFactoryType Factory) {
CTX->CustomCPUFactory = std::move(Factory);
}
bool AddVirtualMemoryMapping([[maybe_unused]] FEXCore::Context::Context *CTX, [[maybe_unused]] uint64_t VirtualAddress, [[maybe_unused]] uint64_t PhysicalAddress, [[maybe_unused]] uint64_t Size) {
bool FEXCore::Context::ContextImpl::AddVirtualMemoryMapping([[maybe_unused]] uint64_t VirtualAddress, [[maybe_unused]] uint64_t PhysicalAddress, [[maybe_unused]] uint64_t Size) {
return false;
}
void RegisterExternalSyscallVisitor(FEXCore::Context::Context *CTX, [[maybe_unused]] uint64_t Syscall, [[maybe_unused]] FEXCore::HLE::SyscallVisitor *Visitor) {
HostFeatures FEXCore::Context::ContextImpl::GetHostFeatures() const {
return HostFeatures;
}
HostFeatures GetHostFeatures(const FEXCore::Context::Context *CTX) {
return CTX->HostFeatures;
void FEXCore::Context::ContextImpl::SetSignalDelegator(FEXCore::SignalDelegator *_SignalDelegation) {
SignalDelegation = _SignalDelegation;
}
void HandleCallback(FEXCore::Context::Context *CTX, FEXCore::Core::InternalThreadState *Thread, uint64_t RIP) {
CTX->HandleCallback(Thread, RIP);
void FEXCore::Context::ContextImpl::SetSyscallHandler(FEXCore::HLE::SyscallHandler *Handler) {
SyscallHandler = Handler;
SourcecodeResolver = Handler->GetSourcecodeResolver();
}
void RegisterHostSignalHandler(FEXCore::Context::Context *CTX, int Signal, HostSignalDelegatorFunction Func, bool Required) {
CTX->RegisterHostSignalHandler(Signal, std::move(Func), Required);
FEXCore::CPUID::FunctionResults FEXCore::Context::ContextImpl::RunCPUIDFunction(uint32_t Function, uint32_t Leaf) {
return CPUID.RunFunction(Function, Leaf);
}
void RegisterFrontendHostSignalHandler(FEXCore::Context::Context *CTX, int Signal, HostSignalDelegatorFunction Func, bool Required) {
CTX->RegisterFrontendHostSignalHandler(Signal, std::move(Func), Required);
FEXCore::CPUID::FunctionResults FEXCore::Context::ContextImpl::RunCPUIDFunctionName(uint32_t Function, uint32_t Leaf, uint32_t CPU) {
return CPUID.RunFunctionName(Function, Leaf, CPU);
}
FEXCore::Core::InternalThreadState* CreateThread(FEXCore::Context::Context *CTX, FEXCore::Core::CPUState *NewThreadState, uint64_t ParentTID) {
return CTX->CreateThread(NewThreadState, ParentTID);
}
void ExecutionThread(FEXCore::Context::Context *CTX, FEXCore::Core::InternalThreadState *Thread) {
return CTX->ExecutionThread(Thread);
}
void InitializeThread(FEXCore::Context::Context *CTX, FEXCore::Core::InternalThreadState *Thread) {
return CTX->InitializeThread(Thread);
}
void RunThread(FEXCore::Context::Context *CTX, FEXCore::Core::InternalThreadState *Thread) {
CTX->RunThread(Thread);
}
void StopThread(FEXCore::Context::Context *CTX, FEXCore::Core::InternalThreadState *Thread) {
CTX->StopThread(Thread);
}
void DestroyThread(FEXCore::Context::Context *CTX, FEXCore::Core::InternalThreadState *Thread) {
CTX->DestroyThread(Thread);
}
void CleanupAfterFork(FEXCore::Context::Context *CTX, FEXCore::Core::InternalThreadState *Thread) {
CTX->CleanupAfterFork(Thread);
}
void SetSignalDelegator(FEXCore::Context::Context *CTX, FEXCore::SignalDelegator *SignalDelegation) {
CTX->SignalDelegation = SignalDelegation;
}
void SetSyscallHandler(FEXCore::Context::Context *CTX, FEXCore::HLE::SyscallHandler *Handler) {
CTX->SyscallHandler = Handler;
CTX->SourcecodeResolver = Handler->GetSourcecodeResolver();
}
FEXCore::CPUID::FunctionResults RunCPUIDFunction(FEXCore::Context::Context *CTX, uint32_t Function, uint32_t Leaf) {
return CTX->CPUID.RunFunction(Function, Leaf);
}
FEX_DEFAULT_VISIBILITY FEXCore::CPUID::FunctionResults RunCPUIDFunctionName(FEXCore::Context::Context *CTX, uint32_t Function, uint32_t Leaf, uint32_t CPU) {
return CTX->CPUID.RunFunctionName(Function, Leaf, CPU);
}
void SetAOTIRLoader(FEXCore::Context::Context *CTX, std::function<int(const std::string&)> CacheReader) {
CTX->SetAOTIRLoader(CacheReader);
}
void SetAOTIRWriter(FEXCore::Context::Context *CTX, std::function<std::unique_ptr<std::ofstream>(const std::string&)> CacheWriter) {
CTX->SetAOTIRWriter(CacheWriter);
}
void SetAOTIRRenamer(FEXCore::Context::Context *CTX, std::function<void(const std::string&)> CacheRenamer) {
CTX->SetAOTIRRenamer(CacheRenamer);
}
void FinalizeAOTIRCache(FEXCore::Context::Context *CTX) {
CTX->FinalizeAOTIRCache();
}
void WriteFilesWithCode(FEXCore::Context::Context *CTX, std::function<void(const std::string& fileid, const std::string& filename)> Writer) {
CTX->WriteFilesWithCode(Writer);
}
IR::AOTIRCacheEntry *LoadAOTIRCacheEntry(FEXCore::Context::Context *CTX, const std::string &Name) {
return CTX->LoadAOTIRCacheEntry(Name);
}
void UnloadAOTIRCacheEntry(FEXCore::Context::Context *CTX, IR::AOTIRCacheEntry *Entry) {
return CTX->UnloadAOTIRCacheEntry(Entry);
}
CustomIRResult AddCustomIREntrypoint(FEXCore::Context::Context *CTX, uintptr_t Entrypoint, std::function<void(uintptr_t Entrypoint, FEXCore::IR::IREmitter *)> Handler, void *Creator, void *Data) {
return CTX->AddCustomIREntrypoint(Entrypoint, Handler, Creator, Data);
}
void AppendThunkDefinitions(FEXCore::Context::Context *CTX, std::vector<FEXCore::IR::ThunkDefinition> const& Definitions) {
CTX->AppendThunkDefinitions(Definitions);
void SetVDSOSigReturn(FEXCore::Context::Context *CTX, const VDSOSigReturn &Pointers) {
CTX->SetVDSOSigReturn(Pointers);
}
namespace Debug {
void CompileRIP(FEXCore::Context::Context *CTX, uint64_t RIP) {
CTX->CompileRIP(CTX->ParentThread, RIP);
}
uint64_t GetThreadCount(FEXCore::Context::Context *CTX) {
return CTX->GetThreadCount();
}
//void CompileRIP(FEXCore::Context::Context *CTX, uint64_t RIP) {
// CTX->CompileRIP(CTX->ParentThread, RIP);
//}
//uint64_t GetThreadCount(FEXCore::Context::Context *CTX) {
// return CTX->GetThreadCount();
//}
FEXCore::Core::RuntimeStats *GetRuntimeStatsForThread(FEXCore::Context::Context *CTX, uint64_t Thread) {
return CTX->GetRuntimeStatsForThread(Thread);
}
//FEXCore::Core::RuntimeStats *GetRuntimeStatsForThread(FEXCore::Context::Context *CTX, uint64_t Thread) {
// return CTX->GetRuntimeStatsForThread(Thread);
//}
bool GetDebugDataForRIP(FEXCore::Context::Context *CTX, uint64_t RIP, FEXCore::Core::DebugData *Data) {
return CTX->GetDebugDataForRIP(RIP, Data);
}
//bool GetDebugDataForRIP(FEXCore::Context::Context *CTX, uint64_t RIP, FEXCore::Core::DebugData *Data) {
// return CTX->GetDebugDataForRIP(RIP, Data);
//}
bool FindHostCodeForRIP(FEXCore::Context::Context *CTX, uint64_t RIP, uint8_t **Code) {
return CTX->FindHostCodeForRIP(RIP, Code);
}
//bool FindHostCodeForRIP(FEXCore::Context::Context *CTX, uint64_t RIP, uint8_t **Code) {
// return CTX->FindHostCodeForRIP(RIP, Code);
//}
// XXX:
// bool FindIRForRIP(FEXCore::Context::Context *CTX, uint64_t RIP, FEXCore::IR::IntrusiveIRList **ir) {
+140 -96
View File
@@ -70,7 +70,130 @@ namespace FEXCore::Context {
MODE_SINGLESTEP = 1,
};
struct Context {
class ContextImpl final : public FEXCore::Context::Context {
public:
// Context base class implementation.
bool InitializeContext() override;
void DestroyContext() override;
FEXCore::Core::InternalThreadState* InitCore(uint64_t InitialRIP, uint64_t StackPointer) override;
void SetExitHandler(ExitHandler handler) override;
ExitHandler GetExitHandler() const override;
void Pause() override;
void Run() override;
void Stop() override;
void Step() override;
ExitReason RunUntilExit() override;
void CompileRIP(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP) override;
int GetProgramStatus() const override;
ExitReason GetExitReason() override;
bool IsDone() const override;
void GetCPUState(FEXCore::Core::CPUState *State) const override;
void SetCPUState(const FEXCore::Core::CPUState *State) override;
void SetCustomCPUBackendFactory(CustomCPUFactoryType Factory) override;
bool AddVirtualMemoryMapping(uint64_t VirtualAddress, uint64_t PhysicalAddress, uint64_t Size) override;
HostFeatures GetHostFeatures() const override;
void HandleCallback(FEXCore::Core::InternalThreadState *Thread, uint64_t RIP) override;
void RegisterHostSignalHandler(int Signal, HostSignalDelegatorFunction Func, bool Required) override;
[[noreturn]] void HandleSignalHandlerReturn(bool RT) override ;
void RegisterFrontendHostSignalHandler(int Signal, HostSignalDelegatorFunction Func, bool Required) override;
/**
* @brief Used to create FEX thread objects in preparation for creating a true OS thread. Does set a TID or PID.
*
* @param NewThreadState The initial thread state to setup for our state
* @param ParentTID The PID that was the parent thread that created this
*
* @return The InternalThreadState object that tracks all of the emulated thread's state
*
* Usecases:
* OS thread Creation:
* - Thread = CreateThread(NewState, PPID);
* - InitializeThread(Thread);
* OS fork (New thread created with a clone of thread state):
* - clone{2, 3}
* - Thread = CreateThread(CopyOfThreadState, PPID);
* - ExecutionThread(Thread); // Starts executing without creating another host thread
* Thunk callback executing guest code from native host thread
* - Thread = CreateThread(NewState, PPID);
* - InitializeThreadTLSData(Thread);
* - HandleCallback(Thread, RIP);
*/
FEXCore::Core::InternalThreadState* CreateThread(FEXCore::Core::CPUState *NewThreadState, uint64_t ParentTID) override;
// Public for threading
void ExecutionThread(FEXCore::Core::InternalThreadState *Thread) override;
/**
* @brief Initializes the OS thread object and prepares to start executing on that new OS thread
*
* @param Thread The internal FEX thread state object
*
* The OS thread will wait until RunThread is executed
*/
void InitializeThread(FEXCore::Core::InternalThreadState *Thread) override;
/**
* @brief Starts the OS thread object to start executing guest code
*
* @param Thread The internal FEX thread state object
*/
void RunThread(FEXCore::Core::InternalThreadState *Thread) override;
void StopThread(FEXCore::Core::InternalThreadState *Thread) override;
/**
* @brief Destroys this FEX thread object and stops tracking it internally
*
* @param Thread The internal FEX thread state object
*/
void DestroyThread(FEXCore::Core::InternalThreadState *Thread) override;
void CleanupAfterFork(FEXCore::Core::InternalThreadState *Thread) override;
void SetSignalDelegator(FEXCore::SignalDelegator *SignalDelegation) override;
void SetSyscallHandler(FEXCore::HLE::SyscallHandler *Handler) override;
FEXCore::CPUID::FunctionResults RunCPUIDFunction(uint32_t Function, uint32_t Leaf) override;
FEXCore::CPUID::FunctionResults RunCPUIDFunctionName(uint32_t Function, uint32_t Leaf, uint32_t CPU) override;
FEXCore::IR::AOTIRCacheEntry *LoadAOTIRCacheEntry(const std::string& Name) override;
void UnloadAOTIRCacheEntry(FEXCore::IR::AOTIRCacheEntry *Entry) override;
void SetAOTIRLoader(std::function<int(const std::string&)> CacheReader) override {
IRCaptureCache.SetAOTIRLoader(CacheReader);
}
void SetAOTIRWriter(std::function<std::unique_ptr<std::ofstream>(const std::string&)> CacheWriter) override {
IRCaptureCache.SetAOTIRWriter(CacheWriter);
}
void SetAOTIRRenamer(std::function<void(const std::string&)> CacheRenamer) override {
IRCaptureCache.SetAOTIRRenamer(CacheRenamer);
}
void FinalizeAOTIRCache() override {
IRCaptureCache.FinalizeAOTIRCache();
}
void WriteFilesWithCode(std::function<void(const std::string& fileid, const std::string& filename)> Writer) override {
IRCaptureCache.WriteFilesWithCode(Writer);
}
void InvalidateGuestCodeRange(uint64_t Start, uint64_t Length) override;
void InvalidateGuestCodeRange(uint64_t Start, uint64_t Length, std::function<void(uint64_t start, uint64_t Length)> callback) override;
void MarkMemoryShared() override;
void ConfigureAOTGen(FEXCore::Core::InternalThreadState *Thread, std::set<uint64_t> *ExternalBranches, uint64_t SectionMaxAddress) override;
// returns false if a handler was already registered
CustomIRResult AddCustomIREntrypoint(uintptr_t Entrypoint, std::function<void(uintptr_t Entrypoint, FEXCore::IR::IREmitter *)> Handler, void *Creator = nullptr, void *Data = nullptr) override;
void AppendThunkDefinitions(std::vector<FEXCore::IR::ThunkDefinition> const& Definitions) override;
public:
friend class FEXCore::HLE::SyscallHandler;
#ifdef JIT_ARM64
friend class FEXCore::CPU::Arm64JITCore;
@@ -105,6 +228,7 @@ namespace FEXCore::Context {
FEX_CONFIG_OPT(MaxInstPerBlock, MAXINST);
FEX_CONFIG_OPT(RootFSPath, ROOTFS);
FEX_CONFIG_OPT(ThunkHostLibsPath, THUNKHOSTLIBS);
FEX_CONFIG_OPT(ThunkHostLibsPath32, THUNKHOSTLIBS32);
FEX_CONFIG_OPT(ThunkConfigFile, THUNKCONFIG);
FEX_CONFIG_OPT(DumpIR, DUMPIR);
FEX_CONFIG_OPT(StaticRegisterAllocation, SRA);
@@ -115,14 +239,13 @@ namespace FEXCore::Context {
FEX_CONFIG_OPT(ParanoidTSO, PARANOIDTSO);
FEX_CONFIG_OPT(CacheObjectCodeCompilation, CACHEOBJECTCODECOMPILATION);
FEX_CONFIG_OPT(x87ReducedPrecision, X87REDUCEDPRECISION);
FEX_CONFIG_OPT(x86dec_SynchronizeRIPOnAllBlocks, X86DEC_SYNCHRONIZERIPONALLBLOCKS);
FEX_CONFIG_OPT(EnableAVX, ENABLEAVX);
} Config;
FEXCore::HostFeatures HostFeatures;
std::mutex ThreadCreationMutex;
FEXCore::Core::InternalThreadState* ParentThread;
FEXCore::Core::InternalThreadState* ParentThread{};
std::vector<FEXCore::Core::InternalThreadState*> Threads;
std::atomic_bool CoreShuttingDown{false};
bool NeedToCheckXID{true};
@@ -151,36 +274,27 @@ namespace FEXCore::Context {
SignalDelegator *SignalDelegation{};
X86GeneratedCode X86CodeGen;
VDSOSigReturn VDSOPointers{};
Context();
~Context();
ContextImpl();
~ContextImpl();
FEXCore::Core::InternalThreadState* InitCore(uint64_t InitialRIP, uint64_t StackPointer);
FEXCore::Context::ExitReason RunUntilExit();
int GetProgramStatus() const;
bool IsPaused() const { return !Running; }
void Pause();
void Run();
void WaitForThreadsToRun();
void Step();
void Stop(bool IgnoreCurrentThread);
void WaitForIdle();
void StopThread(FEXCore::Core::InternalThreadState *Thread);
void SignalThread(FEXCore::Core::InternalThreadState *Thread, FEXCore::Core::SignalEvent Event);
bool GetGdbServerStatus() const { return DebugServer != nullptr; }
void StartGdbServer();
void StopGdbServer();
void HandleCallback(FEXCore::Core::InternalThreadState *Thread, uint64_t RIP);
void RegisterHostSignalHandler(int Signal, HostSignalDelegatorFunction Func, bool Required);
void RegisterFrontendHostSignalHandler(int Signal, HostSignalDelegatorFunction Func, bool Required);
static void ThreadRemoveCodeEntry(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP);
static void ThreadAddBlockLink(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestDestination, uintptr_t HostLink, const std::function<void()> &delinker);
template<auto Fn>
static uint64_t ThreadExitFunctionLink(FEXCore::Core::CpuStateFrame *Frame, uint64_t *record) {
FHU::ScopedSignalMaskWithSharedLock lk(Frame->Thread->CTX->CodeInvalidationMutex);
FHU::ScopedSignalMaskWithSharedLock lk(static_cast<ContextImpl*>(Frame->Thread->CTX)->CodeInvalidationMutex);
return Fn(Frame, record);
}
@@ -189,21 +303,17 @@ namespace FEXCore::Context {
// Must be called from owning thread
static void ThreadRemoveCodeEntryFromJit(FEXCore::Core::CpuStateFrame *Frame, uint64_t GuestRIP) {
auto Thread = Frame->Thread;
LogMan::Throw::AFmt(Thread->ThreadManager.GetTID() == FHU::Syscalls::gettid(), "Must be called from owning thread {}, not {}", Thread->ThreadManager.GetTID(), FHU::Syscalls::gettid());
FHU::ScopedSignalMaskWithUniqueLock lk(Thread->CTX->CodeInvalidationMutex);
FHU::ScopedSignalMaskWithUniqueLock lk(static_cast<ContextImpl*>(Thread->CTX)->CodeInvalidationMutex);
ThreadRemoveCodeEntry(Thread, GuestRIP);
}
// returns false if a handler was already registered
CustomIRResult AddCustomIREntrypoint(uintptr_t Entrypoint, std::function<void(uintptr_t Entrypoint, FEXCore::IR::IREmitter *)> Handler, void *Creator, void *Data);
void RemoveCustomIREntrypoint(uintptr_t Entrypoint);
// Debugger interface
void CompileRIP(FEXCore::Core::InternalThreadState *Thread, uint64_t RIP);
uint64_t GetThreadCount() const;
FEXCore::Core::RuntimeStats *GetRuntimeStatsForThread(uint64_t Thread);
bool GetDebugDataForRIP(uint64_t RIP, FEXCore::Core::DebugData *Data);
@@ -235,29 +345,6 @@ namespace FEXCore::Context {
void CompileBlockJit(FEXCore::Core::CpuStateFrame *Frame, uint64_t GuestRIP);
// Used for thread creation from syscalls
/**
* @brief Used to create FEX thread objects in preparation for creating a true OS thread. Does set a TID or PID.
*
* @param NewThreadState The initial thread state to setup for our state
* @param ParentTID The PID that was the parent thread that created this
*
* @return The InternalThreadState object that tracks all of the emulated thread's state
*
* Usecases:
* OS thread Creation:
* - Thread = CreateThread(NewState, PPID);
* - InitializeThread(Thread);
* OS fork (New thread created with a clone of thread state):
* - clone{2, 3}
* - Thread = CreateThread(CopyOfThreadState, PPID);
* - ExecutionThread(Thread); // Starts executing without creating another host thread
* Thunk callback executing guest code from native host thread
* - Thread = CreateThread(NewState, PPID);
* - InitializeThreadTLSData(Thread);
* - HandleCallback(Thread, RIP);
*/
FEXCore::Core::InternalThreadState* CreateThread(FEXCore::Core::CPUState *NewThreadState, uint64_t ParentTID);
/**
* @brief Initializes TID, PID and TLS data for a thread
*
@@ -265,71 +352,28 @@ namespace FEXCore::Context {
*/
void InitializeThreadTLSData(FEXCore::Core::InternalThreadState *Thread);
/**
* @brief Initializes the OS thread object and prepares to start executing on that new OS thread
*
* @param Thread The internal FEX thread state object
*
* The OS thread will wait until RunThread is executed
*/
void InitializeThread(FEXCore::Core::InternalThreadState *Thread);
/**
* @brief Starts the OS thread object to start executing guest code
*
* @param Thread The internal FEX thread state object
*/
void RunThread(FEXCore::Core::InternalThreadState *Thread);
/**
* @brief Destroys this FEX thread object and stops tracking it internally
*
* @param Thread The internal FEX thread state object
*/
void DestroyThread(FEXCore::Core::InternalThreadState *Thread);
void CopyMemoryMapping(FEXCore::Core::InternalThreadState *ParentThread, FEXCore::Core::InternalThreadState *ChildThread);
void CleanupAfterFork(FEXCore::Core::InternalThreadState *ExceptForThread);
std::vector<FEXCore::Core::InternalThreadState*>* GetThreads() { return &Threads; }
uint8_t GetGPRSize() const { return Config.Is64BitMode ? 8 : 4; }
IR::AOTIRCacheEntry *LoadAOTIRCacheEntry(const std::string &filename);
void UnloadAOTIRCacheEntry(IR::AOTIRCacheEntry *Entry);
FEXCore::JITSymbols Symbols;
// Public for threading
void ExecutionThread(FEXCore::Core::InternalThreadState *Thread);
void SetVDSOSigReturn(const VDSOSigReturn &Pointers) override {
VDSOPointers = Pointers;
if (VDSOPointers.VDSO_kernel_sigreturn == nullptr) {
VDSOPointers.VDSO_kernel_sigreturn = reinterpret_cast<void*>(X86CodeGen.sigreturn_32);
}
void FinalizeAOTIRCache() {
IRCaptureCache.FinalizeAOTIRCache();
if (VDSOPointers.VDSO_kernel_rt_sigreturn == nullptr) {
VDSOPointers.VDSO_kernel_rt_sigreturn = reinterpret_cast<void*>(X86CodeGen.rt_sigreturn_32);
}
}
void WriteFilesWithCode(std::function<void(const std::string& fileid, const std::string& filename)> Writer) {
IRCaptureCache.WriteFilesWithCode(Writer);
}
void SetAOTIRLoader(std::function<int(const std::string&)> CacheReader) {
IRCaptureCache.SetAOTIRLoader(CacheReader);
}
void SetAOTIRWriter(std::function<std::unique_ptr<std::ofstream>(const std::string&)> CacheWriter) {
IRCaptureCache.SetAOTIRWriter(CacheWriter);
}
void SetAOTIRRenamer(std::function<void(const std::string&)> CacheRenamer) {
IRCaptureCache.SetAOTIRRenamer(CacheRenamer);
}
void AppendThunkDefinitions(std::vector<FEXCore::IR::ThunkDefinition> const& Definitions);
FEXCore::Utils::PooledAllocatorMMap OpDispatcherAllocator;
FEXCore::Utils::PooledAllocatorMMap FrontendAllocator;
void MarkMemoryShared();
bool IsTSOEnabled() { return (IsMemoryShared || !Config.TSOAutoMigration) && Config.TSOEnabled; }
protected:
@@ -1,7 +1,6 @@
#include "Interface/Core/ArchHelpers/Arm64.h"
#include "Interface/Core/ArchHelpers/MContext.h"
#include <aarch64/cpu-aarch64.h>
#include "Interface/Core/ArchHelpers/CodeEmitter/Buffer.h"
#include <FEXCore/Utils/EnumUtils.h>
#include <FEXCore/Utils/LogManager.h>
@@ -572,7 +571,7 @@ bool HandleAtomicVectorStore(void *_ucontext, void *_info, uint32_t Instr) {
PC[1] = STP;
PC[2] = DMB;
// Back up one instruction and have another go
vixl::aarch64::CPU::EnsureIAndDCacheCoherency(&PC[0], 16);
FEXCore::ARMEmitter::Buffer::ClearICache(&PC[0], 16);
return true;
}
}
@@ -2311,7 +2310,7 @@ bool HandleSIGBUS(bool ParanoidTSO, int Signal, void *info, void *ucontext) {
return false;
}
vixl::aarch64::CPU::EnsureIAndDCacheCoherency(&PC[-1], 16);
FEXCore::ARMEmitter::Buffer::ClearICache(&PC[-1], 16);
return true;
}
return false;
@@ -1,9 +1,11 @@
#include "Interface/Core/ArchHelpers/Arm64Emitter.h"
#include "Interface/Core/ArchHelpers/CodeEmitter/Emitter.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Context/Context.h"
#include "Interface/HLE/Thunks/Thunks.h"
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Utils/BitUtils.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/MathUtils.h>
@@ -19,39 +21,25 @@
namespace FEXCore::CPU {
// We want vixl to not allocate a default buffer. Jit and dispatcher will manually create one.
Arm64Emitter::Arm64Emitter(FEXCore::Context::Context *ctx, size_t size)
: vixl::aarch64::Assembler(size ? (byte*)FEXCore::Allocator::mmap(nullptr, size, PROT_READ | PROT_WRITE | PROT_EXEC, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0) : reinterpret_cast<byte*>(~0ULL),
size,
vixl::aarch64::PositionDependentCode)
Arm64Emitter::Arm64Emitter(FEXCore::Context::ContextImpl *ctx, size_t size)
: Emitter(size ? (uint8_t*)FEXCore::Allocator::mmap(nullptr, size, PROT_READ | PROT_WRITE | PROT_EXEC, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0) : nullptr, size)
, EmitterCTX {ctx} {
CPU.SetUp();
#ifdef VIXL_SIMULATOR
auto Features = vixl::CPUFeatures::All();
#else
auto Features = vixl::CPUFeatures::InferFromOS();
if (ctx->HostFeatures.SupportsAtomics) {
// Hypervisor can hide this on the c630?
Features.Combine(vixl::CPUFeatures::Feature::kLORegions);
}
#endif
SetCPUFeatures(Features);
}
Arm64Emitter::~Arm64Emitter() {
auto CodeBuffer = GetBuffer();
if (CodeBuffer->GetCapacity()) {
FEXCore::Allocator::munmap(CodeBuffer->GetStartAddress<void*>(), CodeBuffer->GetCapacity());
auto BufferSize = GetBufferSize();
if (BufferSize) {
FEXCore::Allocator::munmap(GetBufferBase(), BufferSize);
}
}
void Arm64Emitter::LoadConstant(vixl::aarch64::Register Reg, uint64_t Constant, bool NOPPad) {
bool Is64Bit = Reg.IsX();
void Arm64Emitter::LoadConstant(ARMEmitter::Size s, ARMEmitter::Register Reg, uint64_t Constant, bool NOPPad) {
bool Is64Bit = s == ARMEmitter::Size::i64Bit;
int Segments = Is64Bit ? 4 : 2;
if (Is64Bit && ((~Constant)>> 16) == 0) {
movn(Reg, (~Constant) & 0xFFFF);
movn(s, Reg, (~Constant) & 0xFFFF);
if (NOPPad) {
nop(); nop(); nop();
@@ -98,17 +86,17 @@ void Arm64Emitter::LoadConstant(vixl::aarch64::Register Reg, uint64_t Constant,
else {
// Need to use ADRP + ADD
adrp(Reg, AlignedOffset >> 12);
add(Reg, Reg, Constant & 0xFFF);
add(s, Reg, Reg, Constant & 0xFFF);
NumMoves = 2;
}
}
}
else {
movz(Reg, (Constant) & 0xFFFF, 0);
movz(s, Reg, (Constant) & 0xFFFF, 0);
for (int i = 1; i < Segments; ++i) {
uint16_t Part = (Constant >> (i * 16)) & 0xFFFF;
if (Part) {
movk(Reg, Part, i * 16);
movk(s, Reg, Part, i * 16);
++NumMoves;
}
}
@@ -124,126 +112,143 @@ void Arm64Emitter::LoadConstant(vixl::aarch64::Register Reg, uint64_t Constant,
void Arm64Emitter::PushCalleeSavedRegisters() {
// We need to save pairs of registers
// We save r19-r30
MemOperand PairOffset(sp, -16, PreIndex);
const std::array<std::pair<vixl::aarch64::Register, vixl::aarch64::Register>, 6> CalleeSaved = {{
{x19, x20},
{x21, x22},
{x23, x24},
{x25, x26},
{x27, x28},
{x29, x30},
const std::array<std::pair<ARMEmitter::XRegister, ARMEmitter::XRegister>, 6> CalleeSaved = {{
{ARMEmitter::XReg::x19, ARMEmitter::XReg::x20},
{ARMEmitter::XReg::x21, ARMEmitter::XReg::x22},
{ARMEmitter::XReg::x23, ARMEmitter::XReg::x24},
{ARMEmitter::XReg::x25, ARMEmitter::XReg::x26},
{ARMEmitter::XReg::x27, ARMEmitter::XReg::x28},
{ARMEmitter::XReg::x29, ARMEmitter::XReg::x30},
}};
for (auto &RegPair : CalleeSaved) {
stp(RegPair.first, RegPair.second, PairOffset);
stp<ARMEmitter::IndexType::PRE>(RegPair.first, RegPair.second, ARMEmitter::Reg::rsp, -16);
}
// Additionally we need to store the lower 64bits of v8-v15
// Here's a fun thing, we can use two ST4 instructions to store everything
// We just need a single sub to sp before that
const std::array<
std::tuple<vixl::aarch64::VRegister,
vixl::aarch64::VRegister,
vixl::aarch64::VRegister,
vixl::aarch64::VRegister>, 2> FPRs = {{
{v8, v9, v10, v11},
{v12, v13, v14, v15},
std::tuple<ARMEmitter::DRegister,
ARMEmitter::DRegister,
ARMEmitter::DRegister,
ARMEmitter::DRegister>, 2> FPRs = {{
{ARMEmitter::DReg::d8, ARMEmitter::DReg::d9, ARMEmitter::DReg::d10, ARMEmitter::DReg::d11},
{ARMEmitter::DReg::d12, ARMEmitter::DReg::d13, ARMEmitter::DReg::d14, ARMEmitter::DReg::d15},
}};
uint32_t VectorSaveSize = sizeof(uint64_t) * 8;
sub(sp, sp, VectorSaveSize);
sub(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, ARMEmitter::Reg::rsp, VectorSaveSize);
// SP supporting move
// We just saved x19 so it is safe
add(x19, sp, 0);
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r19, ARMEmitter::Reg::rsp, 0);
MemOperand QuadOffset(x19, 32, PostIndex);
for (auto &RegQuad : FPRs) {
st4(std::get<0>(RegQuad).D(),
std::get<1>(RegQuad).D(),
std::get<2>(RegQuad).D(),
std::get<3>(RegQuad).D(),
st4(ARMEmitter::SubRegSize::i64Bit,
std::get<0>(RegQuad),
std::get<1>(RegQuad),
std::get<2>(RegQuad),
std::get<3>(RegQuad),
0,
QuadOffset);
ARMEmitter::Reg::r19,
32);
}
}
void Arm64Emitter::PopCalleeSavedRegisters() {
const std::array<
std::tuple<vixl::aarch64::VRegister,
vixl::aarch64::VRegister,
vixl::aarch64::VRegister,
vixl::aarch64::VRegister>, 2> FPRs = {{
{v12, v13, v14, v15},
{v8, v9, v10, v11},
std::tuple<ARMEmitter::DRegister,
ARMEmitter::DRegister,
ARMEmitter::DRegister,
ARMEmitter::DRegister>, 2> FPRs = {{
{ARMEmitter::DReg::d12, ARMEmitter::DReg::d13, ARMEmitter::DReg::d14, ARMEmitter::DReg::d15},
{ARMEmitter::DReg::d8, ARMEmitter::DReg::d9, ARMEmitter::DReg::d10, ARMEmitter::DReg::d11},
}};
MemOperand QuadOffset(sp, 32, PostIndex);
for (auto &RegQuad : FPRs) {
ld4(std::get<0>(RegQuad).D(),
std::get<1>(RegQuad).D(),
std::get<2>(RegQuad).D(),
std::get<3>(RegQuad).D(),
ld4(ARMEmitter::SubRegSize::i64Bit,
std::get<0>(RegQuad),
std::get<1>(RegQuad),
std::get<2>(RegQuad),
std::get<3>(RegQuad),
0,
QuadOffset);
ARMEmitter::Reg::rsp,
32);
}
MemOperand PairOffset(sp, 16, PostIndex);
const std::array<std::pair<vixl::aarch64::Register, vixl::aarch64::Register>, 6> CalleeSaved = {{
{x29, x30},
{x27, x28},
{x25, x26},
{x23, x24},
{x21, x22},
{x19, x20},
const std::array<std::pair<ARMEmitter::XRegister, ARMEmitter::XRegister>, 6> CalleeSaved = {{
{ARMEmitter::XReg::x29, ARMEmitter::XReg::x30},
{ARMEmitter::XReg::x27, ARMEmitter::XReg::x28},
{ARMEmitter::XReg::x25, ARMEmitter::XReg::x26},
{ARMEmitter::XReg::x23, ARMEmitter::XReg::x24},
{ARMEmitter::XReg::x21, ARMEmitter::XReg::x22},
{ARMEmitter::XReg::x19, ARMEmitter::XReg::x20},
}};
for (auto &RegPair : CalleeSaved) {
ldp(RegPair.first, RegPair.second, PairOffset);
ldp<ARMEmitter::IndexType::POST>(RegPair.first, RegPair.second, ARMEmitter::Reg::rsp, 16);
}
}
void Arm64Emitter::SpillStaticRegs(bool FPRs, uint32_t GPRSpillMask, uint32_t FPRSpillMask) {
if (StaticRegisterAllocation()) {
for (size_t i = 0; i < SRA64.size(); i+=2) {
auto Reg1 = SRA64[i];
auto Reg2 = SRA64[i+1];
if (((1U << Reg1.GetCode()) & GPRSpillMask) &&
((1U << Reg2.GetCode()) & GPRSpillMask)) {
stp(Reg1, Reg2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i])));
}
else if (((1U << Reg1.GetCode()) & GPRSpillMask)) {
str(Reg1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i])));
}
else if (((1U << Reg2.GetCode()) & GPRSpillMask)) {
str(Reg2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i+1])));
}
if (!StaticRegisterAllocation()) {
return;
}
for (size_t i = 0; i < SRA64.size(); i+=2) {
auto Reg1 = SRA64[i];
auto Reg2 = SRA64[i+1];
if (((1U << Reg1.Idx()) & GPRSpillMask) &&
((1U << Reg2.Idx()) & GPRSpillMask)) {
stp<ARMEmitter::IndexType::OFFSET>(Reg1.X(), Reg2.X(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i]));
}
else if (((1U << Reg1.Idx()) & GPRSpillMask)) {
str(Reg1.X(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i]));
}
else if (((1U << Reg2.Idx()) & GPRSpillMask)) {
str(Reg2.X(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i+1]));
}
}
if (FPRs) {
if (EmitterCTX->HostFeatures.SupportsAVX) {
for (size_t i = 0; i < SRAFPR.size(); i++) {
const auto Reg = SRAFPR[i];
if (FPRs) {
if (EmitterCTX->HostFeatures.SupportsAVX) {
for (size_t i = 0; i < SRAFPR.size(); i++) {
const auto Reg = SRAFPR[i];
if (((1U << Reg.GetCode()) & FPRSpillMask) != 0) {
mov(TMP4, offsetof(Core::CpuStateFrame, State.xmm.avx.data[i][0]));
st1b(Reg.Z().VnB(), PRED_TMP_32B, SVEMemOperand(STATE, TMP4));
}
if (((1U << Reg.Idx()) & FPRSpillMask) != 0) {
mov(ARMEmitter::Size::i64Bit, TMP4.R(), offsetof(Core::CpuStateFrame, State.xmm.avx.data[i][0]));
st1b<ARMEmitter::SubRegSize::i8Bit>(Reg.Z(), PRED_TMP_32B, STATE.R(), TMP4.R());
}
} else {
}
} else {
if (GPRSpillMask && FPRSpillMask == ~0U) {
// Optimize the common case where we can spill four registers per instruction
auto TmpReg = SRA64[FindFirstSetBit(GPRSpillMask)];
// Load the sse offset in to the temporary register
add(ARMEmitter::Size::i64Bit, TmpReg, STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[0][0]));
for (size_t i = 0; i < SRAFPR.size(); i += 4) {
const auto Reg1 = SRAFPR[i];
const auto Reg2 = SRAFPR[i + 1];
const auto Reg3 = SRAFPR[i + 2];
const auto Reg4 = SRAFPR[i + 3];
st1<ARMEmitter::SubRegSize::i64Bit>(Reg1.Q(), Reg2.Q(), Reg3.Q(), Reg4.Q(), TmpReg, 64);
}
}
else {
for (size_t i = 0; i < SRAFPR.size(); i += 2) {
const auto Reg1 = SRAFPR[i];
const auto Reg2 = SRAFPR[i + 1];
if (((1U << Reg1.GetCode()) & FPRSpillMask) &&
((1U << Reg2.GetCode()) & FPRSpillMask)) {
stp(Reg1.Q(), Reg2.Q(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[i][0])));
if (((1U << Reg1.Idx()) & FPRSpillMask) &&
((1U << Reg2.Idx()) & FPRSpillMask)) {
stp<ARMEmitter::IndexType::OFFSET>(Reg1.Q(), Reg2.Q(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[i][0]));
}
else if (((1U << Reg1.GetCode()) & FPRSpillMask)) {
str(Reg1.Q(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[i][0])));
else if (((1U << Reg1.Idx()) & FPRSpillMask)) {
str(Reg1.Q(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[i][0]));
}
else if (((1U << Reg2.GetCode()) & FPRSpillMask)) {
str(Reg2.Q(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[i+1][0])));
else if (((1U << Reg2.Idx()) & FPRSpillMask)) {
str(Reg2.Q(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[i+1][0]));
}
}
}
@@ -252,126 +257,137 @@ void Arm64Emitter::SpillStaticRegs(bool FPRs, uint32_t GPRSpillMask, uint32_t FP
}
void Arm64Emitter::FillStaticRegs(bool FPRs, uint32_t GPRFillMask, uint32_t FPRFillMask) {
if (StaticRegisterAllocation()) {
for (size_t i = 0; i < SRA64.size(); i+=2) {
auto Reg1 = SRA64[i];
auto Reg2 = SRA64[i+1];
if (((1U << Reg1.GetCode()) & GPRFillMask) &&
((1U << Reg2.GetCode()) & GPRFillMask)) {
ldp(Reg1, Reg2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i])));
}
else if (((1U << Reg1.GetCode()) & GPRFillMask)) {
ldr(Reg1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i])));
}
else if (((1U << Reg2.GetCode()) & GPRFillMask)) {
ldr(Reg2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i+1])));
}
}
if (!StaticRegisterAllocation()) {
return;
}
if (FPRs) {
if (EmitterCTX->HostFeatures.SupportsAVX) {
// Set up predicate registers.
// We don't bother spilling these in SpillStaticRegs,
// since all that matters is we restore them on a fill.
// It's not a concern if they get trounced by something else.
ptrue(PRED_TMP_16B.VnB(), SVE_VL16);
ptrue(PRED_TMP_32B.VnB(), SVE_VL32);
if (FPRs) {
if (EmitterCTX->HostFeatures.SupportsAVX) {
// Set up predicate registers.
// We don't bother spilling these in SpillStaticRegs,
// since all that matters is we restore them on a fill.
// It's not a concern if they get trounced by something else.
ptrue<ARMEmitter::SubRegSize::i8Bit>(PRED_TMP_16B, ARMEmitter::PredicatePattern::SVE_VL16);
ptrue<ARMEmitter::SubRegSize::i8Bit>(PRED_TMP_32B, ARMEmitter::PredicatePattern::SVE_VL32);
for (size_t i = 0; i < SRAFPR.size(); i++) {
const auto Reg = SRAFPR[i];
if (((1U << Reg.GetCode()) & FPRFillMask) != 0) {
mov(TMP4, offsetof(Core::CpuStateFrame, State.xmm.avx.data[i][0]));
ld1b(Reg.Z().VnB(), PRED_TMP_32B.Zeroing(), SVEMemOperand(STATE, TMP4));
}
for (size_t i = 0; i < SRAFPR.size(); i++) {
const auto Reg = SRAFPR[i];
if (((1U << Reg.Idx()) & FPRFillMask) != 0) {
mov(ARMEmitter::Size::i64Bit, TMP4.R(), offsetof(Core::CpuStateFrame, State.xmm.avx.data[i][0]));
ld1b<ARMEmitter::SubRegSize::i8Bit>(Reg.Z(), PRED_TMP_32B.Zeroing(), STATE.R(), TMP4.R());
}
} else {
}
} else {
if (GPRFillMask && FPRFillMask == ~0U) {
// Optimize the common case where we can fill four registers per instruction.
// Use one of the filling static registers before we fill it.
auto TmpReg = SRA64[FindFirstSetBit(GPRFillMask)];
// Load the sse offset in to the temporary register
add(ARMEmitter::Size::i64Bit, TmpReg, STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[0][0]));
for (size_t i = 0; i < SRAFPR.size(); i += 4) {
const auto Reg1 = SRAFPR[i];
const auto Reg2 = SRAFPR[i + 1];
const auto Reg3 = SRAFPR[i + 2];
const auto Reg4 = SRAFPR[i + 3];
ld1<ARMEmitter::SubRegSize::i64Bit>(Reg1.Q(), Reg2.Q(), Reg3.Q(), Reg4.Q(), TmpReg, 64);
}
}
else {
for (size_t i = 0; i < SRAFPR.size(); i += 2) {
const auto Reg1 = SRAFPR[i];
const auto Reg2 = SRAFPR[i + 1];
if (((1U << Reg1.GetCode()) & FPRFillMask) &&
((1U << Reg2.GetCode()) & FPRFillMask)) {
ldp(Reg1.Q(), Reg2.Q(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[i][0])));
if (((1U << Reg1.Idx()) & FPRFillMask) &&
((1U << Reg2.Idx()) & FPRFillMask)) {
ldp<ARMEmitter::IndexType::OFFSET>(Reg1.Q(), Reg2.Q(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[i][0]));
}
else if (((1U << Reg1.GetCode()) & FPRFillMask)) {
ldr(Reg1.Q(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[i][0])));
else if (((1U << Reg1.Idx()) & FPRFillMask)) {
ldr(Reg1.Q(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[i][0]));
}
else if (((1U << Reg2.GetCode()) & FPRFillMask)) {
ldr(Reg2.Q(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[i+1][0])));
else if (((1U << Reg2.Idx()) & FPRFillMask)) {
ldr(Reg2.Q(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[i+1][0]));
}
}
}
}
}
for (size_t i = 0; i < SRA64.size(); i+=2) {
auto Reg1 = SRA64[i];
auto Reg2 = SRA64[i+1];
if (((1U << Reg1.Idx()) & GPRFillMask) &&
((1U << Reg2.Idx()) & GPRFillMask)) {
ldp<ARMEmitter::IndexType::OFFSET>(Reg1.X(), Reg2.X(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i]));
}
else if ((1U << Reg1.Idx()) & GPRFillMask) {
ldr(Reg1.X(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i]));
}
else if ((1U << Reg2.Idx()) & GPRFillMask) {
ldr(Reg2.X(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i+1]));
}
}
}
void Arm64Emitter::PushDynamicRegsAndLR() {
void Arm64Emitter::PushDynamicRegsAndLR(FEXCore::ARMEmitter::Register TmpReg) {
const auto CanUseSVE = EmitterCTX->HostFeatures.SupportsAVX;
const auto GPRSize = (RA64.size() + 1) * Core::CPUState::GPR_REG_SIZE;
const auto GPRSize = 1 * Core::CPUState::GPR_REG_SIZE;
const auto FPRRegSize = CanUseSVE ? Core::CPUState::XMM_AVX_REG_SIZE
: Core::CPUState::XMM_SSE_REG_SIZE;
const auto FPRSize = RAFPR.size() * FPRRegSize;
const uint64_t SPOffset = AlignUp(GPRSize + FPRSize, 16);
sub(sp, sp, SPOffset);
int i = 0;
sub(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, ARMEmitter::Reg::rsp, SPOffset);
// rsp capable move
add(ARMEmitter::Size::i64Bit, TmpReg, ARMEmitter::Reg::rsp, 0);
if (CanUseSVE) {
for (const auto& RA : RAFPR) {
mov(TMP4, i * 8);
st1b(RA.Z().VnB(), PRED_TMP_32B, SVEMemOperand(sp, TMP4));
i += 4;
for (size_t i = 0; i < RAFPR.size(); i += 4) {
const auto Reg1 = RAFPR[i];
const auto Reg2 = RAFPR[i + 1];
const auto Reg3 = RAFPR[i + 2];
const auto Reg4 = RAFPR[i + 3];
st4b(Reg1.Z(), Reg2.Z(), Reg3.Z(), Reg4.Z(), PRED_TMP_32B, TmpReg, 0);
add(ARMEmitter::Size::i64Bit, TmpReg, TmpReg, 32 * 4);
}
} else {
for (const auto& RA : RAFPR) {
str(RA.Q(), MemOperand(sp, i * 8));
i += 2;
static_assert(RAFPR.size() % 4 == 0, "Needs to have multiple of 4 FPRs for RA");
for (size_t i = 0; i < RAFPR.size(); i += 4) {
const auto Reg1 = RAFPR[i];
const auto Reg2 = RAFPR[i + 1];
const auto Reg3 = RAFPR[i + 2];
const auto Reg4 = RAFPR[i + 3];
st1<ARMEmitter::SubRegSize::i64Bit>(Reg1.Q(), Reg2.Q(), Reg3.Q(), Reg4.Q(), TmpReg, 64);
}
}
#if 0 // All GPRs should be caller saved
for (const auto& RA : RA64) {
str(RA, MemOperand(sp, i * 8));
i++;
}
#endif
str(lr, MemOperand(sp, i * 8));
str(ARMEmitter::XReg::lr, TmpReg, 0);
}
void Arm64Emitter::PopDynamicRegsAndLR() {
const auto CanUseSVE = EmitterCTX->HostFeatures.SupportsAVX;
const auto GPRSize = (RA64.size() + 1) * Core::CPUState::GPR_REG_SIZE;
const auto FPRRegSize = CanUseSVE ? Core::CPUState::XMM_AVX_REG_SIZE
: Core::CPUState::XMM_SSE_REG_SIZE;
const auto FPRSize = RAFPR.size() * FPRRegSize;
const uint64_t SPOffset = AlignUp(GPRSize + FPRSize, 16);
int i = 0;
if (CanUseSVE) {
for (const auto& RA : RAFPR) {
mov(TMP4, i * 8);
ld1b(RA.Z().VnB(), PRED_TMP_32B.Zeroing(), SVEMemOperand(sp, TMP4));
i += 4;
for (size_t i = 0; i < RAFPR.size(); i += 4) {
const auto Reg1 = RAFPR[i];
const auto Reg2 = RAFPR[i + 1];
const auto Reg3 = RAFPR[i + 2];
const auto Reg4 = RAFPR[i + 3];
ld4b(Reg1.Z(), Reg2.Z(), Reg3.Z(), Reg4.Z(), PRED_TMP_32B.Zeroing(), ARMEmitter::Reg::rsp);
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, ARMEmitter::Reg::rsp, 32 * 4);
}
} else {
for (const auto& RA : RAFPR) {
ldr(RA.Q(), MemOperand(sp, i * 8));
i += 2;
for (size_t i = 0; i < RAFPR.size(); i += 4) {
const auto Reg1 = RAFPR[i];
const auto Reg2 = RAFPR[i + 1];
const auto Reg3 = RAFPR[i + 2];
const auto Reg4 = RAFPR[i + 3];
ld1<ARMEmitter::SubRegSize::i64Bit>(Reg1.Q(), Reg2.Q(), Reg3.Q(), Reg4.Q(), ARMEmitter::Reg::rsp, 64);
}
}
#if 0 // All GPRs should be caller saved
for (const auto& RA : RA64) {
ldr(RA, MemOperand(sp, i * 8));
i++;
}
#endif
ldr(lr, MemOperand(sp, i * 8));
add(sp, sp, SPOffset);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
}
void Arm64Emitter::Align16B() {
@@ -1,5 +1,9 @@
#pragma once
#include "FEXCore/Utils/EnumUtils.h"
#include "Interface/Core/ArchHelpers/CodeEmitter/Emitter.h"
#include "Interface/Core/ArchHelpers/CodeEmitter/Registers.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Core/ObjectCache/Relocations.h"
@@ -8,6 +12,9 @@
#include <aarch64/cpu-aarch64.h>
#include <aarch64/operands-aarch64.h>
#include <platform-vixl.h>
#ifdef VIXL_DISASSEMBLER
#include <aarch64/disasm-aarch64.h>
#endif
#ifdef VIXL_SIMULATOR
#include <aarch64/simulator-aarch64.h>
#include <aarch64/simulator-constants-aarch64.h>
@@ -21,78 +28,70 @@
#include <utility>
namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
// All but x29 are caller saved
const std::array<aarch64::Register, 16> SRA64 = {
x4, x5, x6, x7, x8, x9, x10, x11,
x12, x18, x17, x16, x15, x14, x13, x29
constexpr std::array<FEXCore::ARMEmitter::Register, 16> SRA64 = {
FEXCore::ARMEmitter::Reg::r4, FEXCore::ARMEmitter::Reg::r5, FEXCore::ARMEmitter::Reg::r6, FEXCore::ARMEmitter::Reg::r7, FEXCore::ARMEmitter::Reg::r8, FEXCore::ARMEmitter::Reg::r9, FEXCore::ARMEmitter::Reg::r10, FEXCore::ARMEmitter::Reg::r11,
FEXCore::ARMEmitter::Reg::r12, FEXCore::ARMEmitter::Reg::r18, FEXCore::ARMEmitter::Reg::r17, FEXCore::ARMEmitter::Reg::r16, FEXCore::ARMEmitter::Reg::r15, FEXCore::ARMEmitter::Reg::r14, FEXCore::ARMEmitter::Reg::r13, FEXCore::ARMEmitter::Reg::r29
};
// All are callee saved
const std::array<aarch64::Register, 9> RA64 = {
x20, x21, x22, x23, x24, x25, x26, x27,
x19
constexpr std::array<FEXCore::ARMEmitter::Register, 9> RA64 = {
FEXCore::ARMEmitter::Reg::r20, FEXCore::ARMEmitter::Reg::r21, FEXCore::ARMEmitter::Reg::r22, FEXCore::ARMEmitter::Reg::r23, FEXCore::ARMEmitter::Reg::r24, FEXCore::ARMEmitter::Reg::r25, FEXCore::ARMEmitter::Reg::r26, FEXCore::ARMEmitter::Reg::r27,
FEXCore::ARMEmitter::Reg::r19
};
const std::array<std::pair<aarch64::Register, aarch64::Register>, 4> RA64Pair = {{
{x20, x21},
{x22, x23},
{x24, x25},
{x26, x27},
}};
const std::array<std::pair<aarch64::Register, aarch64::Register>, 4> RA32Pair = {{
{w20, w21},
{w22, w23},
{w24, w25},
{w26, w27},
constexpr std::array<std::pair<FEXCore::ARMEmitter::Register, FEXCore::ARMEmitter::Register>, 4> RA64Pair = {{
{FEXCore::ARMEmitter::Reg::r20, FEXCore::ARMEmitter::Reg::r21},
{FEXCore::ARMEmitter::Reg::r22, FEXCore::ARMEmitter::Reg::r23},
{FEXCore::ARMEmitter::Reg::r24, FEXCore::ARMEmitter::Reg::r25},
{FEXCore::ARMEmitter::Reg::r26, FEXCore::ARMEmitter::Reg::r27},
}};
// All are caller saved
const std::array<aarch64::VRegister, 16> SRAFPR = {
v16, v17, v18, v19, v20, v21, v22, v23,
v24, v25, v26, v27, v28, v29, v30, v31
constexpr std::array<FEXCore::ARMEmitter::VRegister, 16> SRAFPR = {
FEXCore::ARMEmitter::VReg::v16, FEXCore::ARMEmitter::VReg::v17, FEXCore::ARMEmitter::VReg::v18, FEXCore::ARMEmitter::VReg::v19, FEXCore::ARMEmitter::VReg::v20, FEXCore::ARMEmitter::VReg::v21, FEXCore::ARMEmitter::VReg::v22, FEXCore::ARMEmitter::VReg::v23,
FEXCore::ARMEmitter::VReg::v24, FEXCore::ARMEmitter::VReg::v25, FEXCore::ARMEmitter::VReg::v26, FEXCore::ARMEmitter::VReg::v27, FEXCore::ARMEmitter::VReg::v28, FEXCore::ARMEmitter::VReg::v29, FEXCore::ARMEmitter::VReg::v30, FEXCore::ARMEmitter::VReg::v31
};
// v8..v15 = (lower 64bits) Callee saved
const std::array<aarch64::VRegister, 12> RAFPR = {
/*v0, v1, v2, v3,*/v4, v5, v6, v7, // v0 ~ v3 are used as temps
v8, v9, v10, v11, v12, v13, v14, v15
constexpr std::array<FEXCore::ARMEmitter::VRegister, 12> RAFPR = {
/*FEXCore::ARMEmitter::VReg::v0, FEXCore::ARMEmitter::VReg::v1, FEXCore::ARMEmitter::VReg::v2, FEXCore::ARMEmitter::VReg::v3,*/FEXCore::ARMEmitter::VReg::v4, FEXCore::ARMEmitter::VReg::v5, FEXCore::ARMEmitter::VReg::v6, FEXCore::ARMEmitter::VReg::v7, // FEXCore::ARMEmitter::VReg::v0 ~ FEXCore::ARMEmitter::VReg::v3 are used as temps
FEXCore::ARMEmitter::VReg::v8, FEXCore::ARMEmitter::VReg::v9, FEXCore::ARMEmitter::VReg::v10, FEXCore::ARMEmitter::VReg::v11, FEXCore::ARMEmitter::VReg::v12, FEXCore::ARMEmitter::VReg::v13, FEXCore::ARMEmitter::VReg::v14, FEXCore::ARMEmitter::VReg::v15
};
// Contains the address to the currently available CPU state
#define STATE x28
constexpr auto STATE = FEXCore::ARMEmitter::XReg::x28;
// GPR temporaries. Only x3 can be used across spill boundaries
// so if these ever need to change, be very careful about that.
#define TMP1 x0
#define TMP2 x1
#define TMP3 x2
#define TMP4 x3
constexpr auto TMP1 = FEXCore::ARMEmitter::XReg::x0;
constexpr auto TMP2 = FEXCore::ARMEmitter::XReg::x1;
constexpr auto TMP3 = FEXCore::ARMEmitter::XReg::x2;
constexpr auto TMP4 = FEXCore::ARMEmitter::XReg::x3;
// Vector temporaries
#define VTMP1 v1
#define VTMP2 v2
#define VTMP3 v3
constexpr auto VTMP1 = FEXCore::ARMEmitter::VReg::v0;
constexpr auto VTMP2 = FEXCore::ARMEmitter::VReg::v1;
constexpr auto VTMP3 = FEXCore::ARMEmitter::VReg::v2;
constexpr auto VTMP4 = FEXCore::ARMEmitter::VReg::v3;
// Predicate register temporaries (used when AVX support is enabled)
// PRED_TMP_16B indicates a predicate register that indicates the first 16 bytes set to 1.
// PRED_TMP_32B indicates a predicate register that indicates the first 32 bytes set to 1.
#define PRED_TMP_16B p6
#define PRED_TMP_32B p7
constexpr FEXCore::ARMEmitter::PRegister PRED_TMP_16B = FEXCore::ARMEmitter::PReg::p6;
constexpr FEXCore::ARMEmitter::PRegister PRED_TMP_32B = FEXCore::ARMEmitter::PReg::p7;
// This class contains common emitter utility functions that can
// be used by both Arm64 JIT and ARM64 Dispatcher
class Arm64Emitter : public vixl::aarch64::Assembler {
class Arm64Emitter : public FEXCore::ARMEmitter::Emitter {
protected:
Arm64Emitter(FEXCore::Context::Context *ctx, size_t size);
Arm64Emitter(FEXCore::Context::ContextImpl *ctx, size_t size);
~Arm64Emitter();
FEXCore::Context::Context *EmitterCTX;
FEXCore::Context::ContextImpl *EmitterCTX;
vixl::aarch64::CPU CPU;
void LoadConstant(vixl::aarch64::Register Reg, uint64_t Constant, bool NOPPad = false);
void LoadConstant(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register Reg, uint64_t Constant, bool NOPPad = false);
// NOTE: These functions WILL clobber the register TMP4 if AVX support is enabled
// and FPRs are being spilled or filled. If only GPRs are spilled/filled, then
@@ -106,13 +105,14 @@ protected:
// We can't guarantee only the lower 64bits are used so flush everything
static constexpr uint32_t CALLER_FPR_MASK = ~0U;
void PushDynamicRegsAndLR();
void PushDynamicRegsAndLR(FEXCore::ARMEmitter::Register TmpReg);
void PopDynamicRegsAndLR();
void PushCalleeSavedRegisters();
void PopCalleeSavedRegisters();
void Align16B();
#ifdef VIXL_SIMULATOR
// Generates a vixl simulator runtime call.
//
@@ -124,61 +124,64 @@ protected:
// 2) Simulator wrapper handler
// 3) Function to call
// 4) Style of the function call (Call versus tail-call)
template<typename R, typename... P>
void GenerateRuntimeCall(R (*Function)(P...)) {
uintptr_t SimulatorWrapperAddress = reinterpret_cast<uintptr_t>(
&(Simulator::RuntimeCallStructHelper<R, P...>::Wrapper));
&(vixl::aarch64::Simulator::RuntimeCallStructHelper<R, P...>::Wrapper));
uintptr_t FunctionAddress = reinterpret_cast<uintptr_t>(Function);
hlt(kRuntimeCallOpcode);
hlt(vixl::aarch64::kRuntimeCallOpcode);
// Simulator wrapper address pointer.
dc(SimulatorWrapperAddress);
dc64(SimulatorWrapperAddress);
// Runtime function address to call
dc(FunctionAddress);
dc64(FunctionAddress);
// Call type
dc32(kCallRuntime);
dc32(vixl::aarch64::kCallRuntime);
}
template<typename R, typename... P>
void GenerateIndirectRuntimeCall(vixl::aarch64::Register Reg) {
void GenerateIndirectRuntimeCall(ARMEmitter::Register Reg) {
uintptr_t SimulatorWrapperAddress = reinterpret_cast<uintptr_t>(
&(Simulator::RuntimeCallStructHelper<R, P...>::Wrapper));
&(vixl::aarch64::Simulator::RuntimeCallStructHelper<R, P...>::Wrapper));
hlt(kIndirectRuntimeCallOpcode);
hlt(vixl::aarch64::kIndirectRuntimeCallOpcode);
// Simulator wrapper address pointer.
dc(SimulatorWrapperAddress);
dc64(SimulatorWrapperAddress);
// Register that contains the function to call
dc(Reg.GetCode());
dc32(Reg.Idx());
// Call type
dc32(kCallRuntime);
dc32(vixl::aarch64::kCallRuntime);
}
template<>
void GenerateIndirectRuntimeCall<float, __uint128_t>(vixl::aarch64::Register Reg) {
void GenerateIndirectRuntimeCall<float, __uint128_t>(ARMEmitter::Register Reg) {
uintptr_t SimulatorWrapperAddress = reinterpret_cast<uintptr_t>(
&(Simulator::RuntimeCallStructHelper<float, __uint128_t>::Wrapper));
&(vixl::aarch64::Simulator::RuntimeCallStructHelper<float, __uint128_t>::Wrapper));
hlt(kIndirectRuntimeCallOpcode);
hlt(vixl::aarch64::kIndirectRuntimeCallOpcode);
// Simulator wrapper address pointer.
dc(SimulatorWrapperAddress);
dc64(SimulatorWrapperAddress);
// Register that contains the function to call
dc(Reg.GetCode());
dc32(Reg.Idx());
// Call type
dc32(kCallRuntime);
dc32(vixl::aarch64::kCallRuntime);
}
#endif
#ifdef VIXL_DISASSEMBLER
vixl::aarch64::PrintDisassembler Disasm {stderr};
#endif
FEX_CONFIG_OPT(StaticRegisterAllocation, SRA);
};
@@ -0,0 +1,991 @@
/* ALU instruction emitters.
*
* Almost all of these operations have `ARMEmitter::Size` as their first argument.
* This allows both 32-bit and 64-bit selection of how that instruction is going to operate.
*
* Some emitter operations explicitly use `XRegister` or `WRegister`.
* This is usually due to the instruction only supporting one operating size.
* Although in some cases is a minor convenience without any performance implications.
*
* FEX-Emu ALU operations usually have a 32-bit or 64-bit operating size encoded in the IR operation,
* This allows FEX to use a single helper function which decodes to both handlers.
*/
private:
static bool IsADRRange(int64_t Imm) {
return Imm >= -1048576 && Imm <= 1048575;
}
static bool IsADRPRange(int64_t Imm) {
return Imm >= -4294967296 && Imm <= 4294963200;
}
static bool IsADRPAligned(int64_t Imm) {
return (Imm & 0xFFF) == 0;
}
public:
// PC relative
void adr(FEXCore::ARMEmitter::Register rd, uint32_t Imm) {
constexpr uint32_t Op = 0b0001'0000 << 24;
DataProcessing_PCRel_Imm(Op, rd, Imm);
}
void adr(FEXCore::ARMEmitter::Register rd, BackwardLabel const* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(IsADRRange(Imm), "Unscaled offset too large");
constexpr uint32_t Op = 0b0001'0000 << 24;
DataProcessing_PCRel_Imm(Op, rd, Imm);
}
void adr(FEXCore::ARMEmitter::Register rd, ForwardLabel *Label) {
Label->Insts.emplace_back(ForwardLabel::Instructions{ .Location = GetCursorAddress<uint8_t*>(), .Type = ForwardLabel::Instructions::InstType::ADR });
constexpr uint32_t Op = 0b0001'0000 << 24;
DataProcessing_PCRel_Imm(Op, rd, 0);
}
void adr(FEXCore::ARMEmitter::Register rd, BiDirectionalLabel *Label) {
if (Label->Backward.Location) {
adr(rd, &Label->Backward);
}
else {
adr(rd, &Label->Forward);
}
}
void adrp(FEXCore::ARMEmitter::Register rd, uint32_t Imm) {
constexpr uint32_t Op = 0b1001'0000 << 24;
DataProcessing_PCRel_Imm(Op, rd, Imm);
}
void adrp(FEXCore::ARMEmitter::Register rd, BackwardLabel const* Label) {
int64_t Imm = reinterpret_cast<int64_t>(Label->Location) - (GetCursorAddress<int64_t>() & ~0xFFFLL);
LOGMAN_THROW_A_FMT(IsADRPRange(Imm) && IsADRPAligned(Imm), "Unscaled offset too large");
constexpr uint32_t Op = 0b1001'0000 << 24;
DataProcessing_PCRel_Imm(Op, rd, Imm);
}
void adrp(FEXCore::ARMEmitter::Register rd, ForwardLabel *Label) {
Label->Insts.emplace_back(ForwardLabel::Instructions{ .Location = GetCursorAddress<uint8_t*>(), .Type = ForwardLabel::Instructions::InstType::ADRP });
constexpr uint32_t Op = 0b1001'0000 << 24;
DataProcessing_PCRel_Imm(Op, rd, 0);
}
void adrp(FEXCore::ARMEmitter::Register rd, BiDirectionalLabel *Label) {
if (Label->Backward.Location) {
adrp(rd, &Label->Backward);
}
else {
adrp(rd, &Label->Forward);
}
}
void LongAddressGen(FEXCore::ARMEmitter::Register rd, BackwardLabel const* Label) {
int64_t Imm = reinterpret_cast<int64_t>(Label->Location) - (GetCursorAddress<int64_t>());
if (IsADRRange(Imm)) {
// If the range is in ADR range then we can just use ADR.
adr(rd, Label);
}
else if (IsADRPRange(Imm)) {
int64_t ADRPImm = (reinterpret_cast<int64_t>(Label->Location) & ~0xFFFLL)
- (GetCursorAddress<int64_t>() & ~0xFFFLL);
// If the range is in the ADRP range then we can use ADRP.
bool NeedsOffset = !IsADRPAligned(reinterpret_cast<uint64_t>(Label->Location));
uint64_t AlignedOffset = reinterpret_cast<uint64_t>(Label->Location) & 0xFFFULL;
// First emit ADRP
adrp(rd, ADRPImm >> 12);
if (NeedsOffset) {
// Now even an add
add(ARMEmitter::Size::i64Bit, rd, rd, AlignedOffset);
}
}
else {
LOGMAN_MSG_A_FMT("Unscaled offset too large");
FEX_UNREACHABLE;
}
}
void LongAddressGen(FEXCore::ARMEmitter::Register rd, ForwardLabel* Label) {
Label->Insts.emplace_back(ForwardLabel::Instructions{ .Location = GetCursorAddress<uint8_t*>(), .Type = ForwardLabel::Instructions::InstType::LONG_ADDRESS_GEN });
// Emit a register index and a nop. These will be backpatched.
dc32(rd.Idx());
nop();
}
void LongAddressGen(FEXCore::ARMEmitter::Register rd, BiDirectionalLabel *Label) {
if (Label->Backward.Location) {
LongAddressGen(rd, &Label->Backward);
}
else {
LongAddressGen(rd, &Label->Forward);
}
}
// Add/subtract immediate
void add(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, uint32_t Imm, bool LSL12 = false) {
constexpr uint32_t Op = 0b0001'0001'0 << 23;
DataProcessing_AddSub_Imm(Op, s, rd, rn, Imm, LSL12);
}
void adds(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, uint32_t Imm, bool LSL12 = false) {
constexpr uint32_t Op = 0b0011'0001'0 << 23;
DataProcessing_AddSub_Imm(Op, s, rd, rn, Imm, LSL12);
}
void sub(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, uint32_t Imm, bool LSL12 = false) {
constexpr uint32_t Op = 0b0101'0001'0 << 23;
DataProcessing_AddSub_Imm(Op, s, rd, rn, Imm, LSL12);
}
void cmp(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rn, uint32_t Imm, bool LSL12 = false) {
constexpr uint32_t Op = 0b0111'0001'0 << 23;
DataProcessing_AddSub_Imm(Op, s, FEXCore::ARMEmitter::Reg::rsp, rn, Imm, LSL12);
}
void subs(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, uint32_t Imm, bool LSL12 = false) {
constexpr uint32_t Op = 0b0111'0001'0 << 23;
DataProcessing_AddSub_Imm(Op, s, rd, rn, Imm, LSL12);
}
// Logical immediate
void and_(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, uint64_t Imm) {
uint32_t n, immr, imms;
[[maybe_unused]] const auto IsImm = vixl::aarch64::Assembler::IsImmLogical(Imm,
RegSizeInBits(s),
&n,
&imms,
&immr);
LOGMAN_THROW_A_FMT(IsImm, "Couldn't encode immediate to logical op");
and_(s, rd, rn, n, immr, imms);
}
void bic(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, uint64_t Imm) {
and_(s, rd, rn, ~Imm);
}
void ands(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, uint64_t Imm) {
uint32_t n, immr, imms;
[[maybe_unused]] const auto IsImm = vixl::aarch64::Assembler::IsImmLogical(Imm,
RegSizeInBits(s),
&n,
&imms,
&immr);
LOGMAN_THROW_A_FMT(IsImm, "Couldn't encode immediate to logical op");
ands(s, rd, rn, n, immr, imms);
}
void bics(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, uint64_t Imm) {
ands(s, rd, rn, ~Imm);
}
void orr(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, uint64_t Imm) {
uint32_t n, immr, imms;
[[maybe_unused]] const auto IsImm = vixl::aarch64::Assembler::IsImmLogical(Imm,
RegSizeInBits(s),
&n,
&imms,
&immr);
LOGMAN_THROW_A_FMT(IsImm, "Couldn't encode immediate to logical op");
orr(s, rd, rn, n, immr, imms);
}
void eor(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, uint64_t Imm) {
uint32_t n, immr, imms;
[[maybe_unused]] const auto IsImm = vixl::aarch64::Assembler::IsImmLogical(Imm,
RegSizeInBits(s),
&n,
&imms,
&immr);
LOGMAN_THROW_A_FMT(IsImm, "Couldn't encode immediate to logical op");
eor(s, rd, rn, n, immr, imms);
}
// Move wide immediate
void movn(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, uint32_t Imm, uint32_t Offset = 0) {
LOGMAN_THROW_A_FMT((Imm & 0xFFFF0000U) == 0, "Upper bits of move wide not valid");
LOGMAN_THROW_A_FMT((Offset % 16) == 0, "Offset must be 16bit aligned");
constexpr uint32_t Op = 0b001'0010'100 << 21;
DataProcessing_MoveWide(Op, s, rd, Imm, Offset >> 4);
}
void mov(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, uint32_t Imm) {
movz(s, rd, Imm, 0);
}
void mov(FEXCore::ARMEmitter::XRegister rd, uint32_t Imm) {
movz(FEXCore::ARMEmitter::Size::i64Bit, rd.R(), Imm, 0);
}
void mov(FEXCore::ARMEmitter::WRegister rd, uint32_t Imm) {
movz(FEXCore::ARMEmitter::Size::i32Bit, rd.R(), Imm, 0);
}
void movz(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, uint32_t Imm, uint32_t Offset = 0) {
LOGMAN_THROW_A_FMT((Imm & 0xFFFF0000U) == 0, "Upper bits of move wide not valid");
LOGMAN_THROW_A_FMT((Offset % 16) == 0, "Offset must be 16bit aligned");
constexpr uint32_t Op = 0b101'0010'100 << 21;
DataProcessing_MoveWide(Op, s, rd, Imm, Offset >> 4);
}
void movk(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, uint32_t Imm, uint32_t Offset = 0) {
LOGMAN_THROW_A_FMT((Imm & 0xFFFF0000U) == 0, "Upper bits of move wide not valid");
LOGMAN_THROW_A_FMT((Offset % 16) == 0, "Offset must be 16bit aligned");
constexpr uint32_t Op = 0b111'0010'100 << 21;
DataProcessing_MoveWide(Op, s, rd, Imm, Offset >> 4);
}
void movn(FEXCore::ARMEmitter::XRegister rd, uint32_t Imm, uint32_t Offset = 0) {
movn(FEXCore::ARMEmitter::Size::i64Bit, rd.R(), Imm, Offset);
}
void movz(FEXCore::ARMEmitter::XRegister rd, uint32_t Imm, uint32_t Offset = 0) {
movz(FEXCore::ARMEmitter::Size::i64Bit, rd.R(), Imm, Offset);
}
void movk(FEXCore::ARMEmitter::XRegister rd, uint32_t Imm, uint32_t Offset = 0) {
movk(FEXCore::ARMEmitter::Size::i64Bit, rd.R(), Imm, Offset);
}
void movn(FEXCore::ARMEmitter::WRegister rd, uint32_t Imm, uint32_t Offset = 0) {
movn(FEXCore::ARMEmitter::Size::i32Bit, rd.R(), Imm, Offset);
}
void movz(FEXCore::ARMEmitter::WRegister rd, uint32_t Imm, uint32_t Offset = 0) {
movz(FEXCore::ARMEmitter::Size::i32Bit, rd.R(), Imm, Offset);
}
void movk(FEXCore::ARMEmitter::WRegister rd, uint32_t Imm, uint32_t Offset = 0) {
movk(FEXCore::ARMEmitter::Size::i32Bit, rd.R(), Imm, Offset);
}
// Bitfield
void sxtb(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn) {
sbfm(s, rd, rn, 0, 7);
}
void sxth(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn) {
sbfm(s, rd, rn, 0, 15);
}
void sxtw(FEXCore::ARMEmitter::XRegister rd, FEXCore::ARMEmitter::WRegister rn) {
sbfm(ARMEmitter::Size::i64Bit, rd, rn.X(), 0, 31);
}
void sbfx(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, uint32_t lsb, uint32_t width) {
LOGMAN_THROW_A_FMT(width > 0, "sbfx needs width > 0");
LOGMAN_THROW_A_FMT((lsb + width) <= RegSizeInBits(s), "Tried to sbfx a region larger than the register");
sbfm(s, rd, rn, lsb, lsb + width - 1);
}
void asr(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, uint32_t shift) {
LOGMAN_THROW_A_FMT(shift <= RegSizeInBits(s), "Tried to asr a region larger than the register");
sbfm(s, rd, rn, shift, RegSizeInBits(s) - 1);
}
void uxtb(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn) {
ubfm(s, rd, rn, 0, 7);
}
void uxth(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn) {
ubfm(s, rd, rn, 0, 15);
}
void uxtw(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn) {
ubfm(s, rd, rn, 0, 31);
}
void ubfm(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, uint32_t immr, uint32_t imms) {
constexpr uint32_t Op = 0b0101'0011'00 << 22;
DataProcessing_Logical_Imm(Op, s, rd, rn, s == ARMEmitter::Size::i64Bit, immr, imms);
}
void lsl(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, uint32_t shift) {
const auto RegSize = RegSizeInBits(s);
LOGMAN_THROW_A_FMT(shift < RegSize, "Tried to lsl a region larger than the register");
ubfm(s, rd, rn, (RegSize - shift) % RegSize, RegSize - shift - 1);
}
void lsr(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, uint32_t shift) {
const auto RegSize = RegSizeInBits(s);
LOGMAN_THROW_A_FMT(shift < RegSize, "Tried to lsr a region larger than the register");
ubfm(s, rd, rn, shift, RegSize - 1);
}
void ubfx(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, uint32_t lsb, uint32_t width) {
LOGMAN_THROW_A_FMT(width > 0, "ubfx needs width > 0");
LOGMAN_THROW_A_FMT((lsb + width) <= RegSizeInBits(s), "Tried to ubfx a region larger than the register");
ubfm(s, rd, rn, lsb, lsb + width - 1);
}
void bfi(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, uint32_t lsb, uint32_t width) {
const auto RegSize = RegSizeInBits(s);
LOGMAN_THROW_A_FMT(width > 0, "bfi needs width > 0");
LOGMAN_THROW_A_FMT((lsb + width) <= RegSize, "Tried to bfi a region larger than the register");
bfm(s, rd, rn, (RegSize - lsb) & (RegSize - 1), width - 1);
}
// Extract
void extr(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, FEXCore::ARMEmitter::Register rm, uint32_t Imm) {
constexpr uint32_t Op = 0b001'0011'100 << 21;
LOGMAN_THROW_A_FMT(Imm < RegSizeInBits(s), "Tried to extr a region larger than the register");
DataProcessing_Extract(Op, s, rd, rn, rm, Imm);
}
void ror(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, uint32_t Imm) {
extr(s, rd, rn, rn, Imm);
}
// Data processing - 2 source
void udiv(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, FEXCore::ARMEmitter::Register rm) {
constexpr uint32_t Op = (0b001'1010'110U << 21) |
(0b0000'10U << 10);
DataProcessing_2Source(Op, s, rd, rn, rm);
}
void sdiv(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, FEXCore::ARMEmitter::Register rm) {
constexpr uint32_t Op = (0b001'1010'110U << 21) |
(0b0000'11U << 10);
DataProcessing_2Source(Op, s, rd, rn, rm);
}
void lslv(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, FEXCore::ARMEmitter::Register rm) {
constexpr uint32_t Op = (0b001'1010'110U << 21) |
(0b0010'00U << 10);
DataProcessing_2Source(Op, s, rd, rn, rm);
}
void lsrv(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, FEXCore::ARMEmitter::Register rm) {
constexpr uint32_t Op = (0b001'1010'110U << 21) |
(0b0010'01U << 10);
DataProcessing_2Source(Op, s, rd, rn, rm);
}
void asrv(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, FEXCore::ARMEmitter::Register rm) {
constexpr uint32_t Op = (0b001'1010'110U << 21) |
(0b0010'10U << 10);
DataProcessing_2Source(Op, s, rd, rn, rm);
}
void rorv(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, FEXCore::ARMEmitter::Register rm) {
constexpr uint32_t Op = (0b001'1010'110U << 21) |
(0b0010'11U << 10);
DataProcessing_2Source(Op, s, rd, rn, rm);
}
void crc32b(FEXCore::ARMEmitter::WRegister rd, FEXCore::ARMEmitter::WRegister rn, FEXCore::ARMEmitter::WRegister rm) {
constexpr uint32_t Op = (0b001'1010'110U << 21) |
(0b0100'00U << 10);
DataProcessing_2Source(Op, ARMEmitter::Size::i32Bit, rd, rn, rm);
}
void crc32h(FEXCore::ARMEmitter::WRegister rd, FEXCore::ARMEmitter::WRegister rn, FEXCore::ARMEmitter::WRegister rm) {
constexpr uint32_t Op = (0b001'1010'110U << 21) |
(0b0100'01U << 10);
DataProcessing_2Source(Op, ARMEmitter::Size::i32Bit, rd, rn, rm);
}
void crc32w(FEXCore::ARMEmitter::WRegister rd, FEXCore::ARMEmitter::WRegister rn, FEXCore::ARMEmitter::WRegister rm) {
constexpr uint32_t Op = (0b001'1010'110U << 21) |
(0b0100'10U << 10);
DataProcessing_2Source(Op, ARMEmitter::Size::i32Bit, rd, rn, rm);
}
void crc32cb(FEXCore::ARMEmitter::WRegister rd, FEXCore::ARMEmitter::WRegister rn, FEXCore::ARMEmitter::WRegister rm) {
constexpr uint32_t Op = (0b001'1010'110U << 21) |
(0b0101'00U << 10);
DataProcessing_2Source(Op, ARMEmitter::Size::i32Bit, rd, rn, rm);
}
void crc32ch(FEXCore::ARMEmitter::WRegister rd, FEXCore::ARMEmitter::WRegister rn, FEXCore::ARMEmitter::WRegister rm) {
constexpr uint32_t Op = (0b001'1010'110U << 21) |
(0b0101'01U << 10);
DataProcessing_2Source(Op, ARMEmitter::Size::i32Bit, rd, rn, rm);
}
void crc32cw(FEXCore::ARMEmitter::WRegister rd, FEXCore::ARMEmitter::WRegister rn, FEXCore::ARMEmitter::WRegister rm) {
constexpr uint32_t Op = (0b001'1010'110U << 21) |
(0b0101'10U << 10);
DataProcessing_2Source(Op, ARMEmitter::Size::i32Bit, rd, rn, rm);
}
void subp(FEXCore::ARMEmitter::XRegister rd, FEXCore::ARMEmitter::XRegister rn, FEXCore::ARMEmitter::XRegister rm) {
constexpr uint32_t Op = (0b001'1010'110U << 21) |
(0b0000'00U << 10);
DataProcessing_2Source(Op, FEXCore::ARMEmitter::Size::i64Bit, rd, rn, rm);
}
void irg(FEXCore::ARMEmitter::XRegister rd, FEXCore::ARMEmitter::XRegister rn, FEXCore::ARMEmitter::XRegister rm) {
constexpr uint32_t Op = (0b001'1010'110U << 21) |
(0b0001'00U << 10);
DataProcessing_2Source(Op, FEXCore::ARMEmitter::Size::i64Bit, rd, rn, rm);
}
void gmi(FEXCore::ARMEmitter::XRegister rd, FEXCore::ARMEmitter::XRegister rn, FEXCore::ARMEmitter::XRegister rm) {
constexpr uint32_t Op = (0b001'1010'110U << 21) |
(0b0001'01U << 10);
DataProcessing_2Source(Op, FEXCore::ARMEmitter::Size::i64Bit, rd, rn, rm);
}
void pacga(FEXCore::ARMEmitter::XRegister rd, FEXCore::ARMEmitter::XRegister rn, FEXCore::ARMEmitter::XRegister rm) {
constexpr uint32_t Op = (0b001'1010'110U << 21) |
(0b0011'00U << 10);
DataProcessing_2Source(Op, FEXCore::ARMEmitter::Size::i64Bit, rd, rn, rm);
}
void crc32x(FEXCore::ARMEmitter::XRegister rd, FEXCore::ARMEmitter::XRegister rn, FEXCore::ARMEmitter::XRegister rm) {
constexpr uint32_t Op = (0b001'1010'110U << 21) |
(0b0100'11U << 10);
DataProcessing_2Source(Op, FEXCore::ARMEmitter::Size::i64Bit, rd, rn, rm);
}
void crc32cx(FEXCore::ARMEmitter::XRegister rd, FEXCore::ARMEmitter::XRegister rn, FEXCore::ARMEmitter::XRegister rm) {
constexpr uint32_t Op = (0b001'1010'110U << 21) |
(0b0101'11U << 10);
DataProcessing_2Source(Op, FEXCore::ARMEmitter::Size::i64Bit, rd, rn, rm);
}
void subps(FEXCore::ARMEmitter::XRegister rd, FEXCore::ARMEmitter::XRegister rn, FEXCore::ARMEmitter::XRegister rm) {
constexpr uint32_t Op = (0b011'1010'110U << 21) |
(0b0000'00U << 10);
DataProcessing_2Source(Op, FEXCore::ARMEmitter::Size::i64Bit, rd, rn, rm);
}
// Data processing - 1 source
void rbit(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn) {
constexpr uint32_t Op = (0b101'1010'110U << 21) |
(0b0'0000U << 16) |
(0b0000'00U << 10);
DataProcessing_1Source(Op, s, rd, rn);
}
void rev16(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn) {
constexpr uint32_t Op = (0b101'1010'110U << 21) |
(0b0'0000U << 16) |
(0b0000'01U << 10);
DataProcessing_1Source(Op, s, rd, rn);
}
void rev(FEXCore::ARMEmitter::WRegister rd, FEXCore::ARMEmitter::WRegister rn) {
constexpr uint32_t Op = (0b101'1010'110U << 21) |
(0b0'0000U << 16) |
(0b0000'10U << 10);
DataProcessing_1Source(Op, FEXCore::ARMEmitter::Size::i32Bit, rd, rn);
}
void rev32(FEXCore::ARMEmitter::XRegister rd, FEXCore::ARMEmitter::XRegister rn) {
constexpr uint32_t Op = (0b101'1010'110U << 21) |
(0b0'0000U << 16) |
(0b0000'10U << 10);
DataProcessing_1Source(Op, FEXCore::ARMEmitter::Size::i64Bit, rd, rn);
}
void clz(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn) {
constexpr uint32_t Op = (0b101'1010'110U << 21) |
(0b0'0000U << 16) |
(0b0001'00U << 10);
DataProcessing_1Source(Op, s, rd, rn);
}
void cls(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn) {
constexpr uint32_t Op = (0b101'1010'110U << 21) |
(0b0'0000U << 16) |
(0b0001'01U << 10);
DataProcessing_1Source(Op, s, rd, rn);
}
void rev(FEXCore::ARMEmitter::XRegister rd, FEXCore::ARMEmitter::XRegister rn) {
constexpr uint32_t Op = (0b101'1010'110U << 21) |
(0b0'0000U << 16) |
(0b0000'11U << 10);
DataProcessing_1Source(Op, FEXCore::ARMEmitter::Size::i64Bit, rd, rn);
}
void rev(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn) {
uint32_t Op = (0b101'1010'110U << 21) |
(0b0'0000U << 16) |
(0b0000'10U << 10) |
(s == ARMEmitter::Size::i64Bit ? (1U << 10) : 0);
DataProcessing_1Source(Op, s, rd, rn);
}
// TODO: PAUTH
// Logical - shifted register
void mov(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn) {
orr(s, rd, FEXCore::ARMEmitter::Reg::zr, rn, ARMEmitter::ShiftType::LSL, 0);
}
void mov(FEXCore::ARMEmitter::XRegister rd, FEXCore::ARMEmitter::XRegister rn) {
orr(FEXCore::ARMEmitter::Size::i64Bit, rd.R(), FEXCore::ARMEmitter::Reg::zr, rn.R(), ARMEmitter::ShiftType::LSL, 0);
}
void mov(FEXCore::ARMEmitter::WRegister rd, FEXCore::ARMEmitter::WRegister rn) {
orr(FEXCore::ARMEmitter::Size::i32Bit, rd.R(), FEXCore::ARMEmitter::Reg::zr, rn.R(), ARMEmitter::ShiftType::LSL, 0);
}
void mvn(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, FEXCore::ARMEmitter::ShiftType Shift = FEXCore::ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
orn(s, rd, FEXCore::ARMEmitter::Reg::zr, rn, Shift, amt);
}
void and_(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, FEXCore::ARMEmitter::Register rm, FEXCore::ARMEmitter::ShiftType Shift = FEXCore::ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
constexpr uint32_t Op = 0b000'1010'000U << 21;
DataProcessing_Shifted_Reg(Op, s, rd, rn, rm, Shift, amt);
}
void ands(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, FEXCore::ARMEmitter::Register rm, FEXCore::ARMEmitter::ShiftType Shift = FEXCore::ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
constexpr uint32_t Op = 0b110'1010'000U << 21;
DataProcessing_Shifted_Reg(Op, s, rd, rn, rm, Shift, amt);
}
void bic(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, FEXCore::ARMEmitter::Register rm, FEXCore::ARMEmitter::ShiftType Shift = FEXCore::ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
constexpr uint32_t Op = 0b000'1010'001U << 21;
DataProcessing_Shifted_Reg(Op, s, rd, rn, rm, Shift, amt);
}
void bics(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, FEXCore::ARMEmitter::Register rm, FEXCore::ARMEmitter::ShiftType Shift = FEXCore::ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
constexpr uint32_t Op = 0b110'1010'001U << 21;
DataProcessing_Shifted_Reg(Op, s, rd, rn, rm, Shift, amt);
}
void orr(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, FEXCore::ARMEmitter::Register rm, FEXCore::ARMEmitter::ShiftType Shift = FEXCore::ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
constexpr uint32_t Op = 0b010'1010'000U << 21;
DataProcessing_Shifted_Reg(Op, s, rd, rn, rm, Shift, amt);
}
void orn(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, FEXCore::ARMEmitter::Register rm, FEXCore::ARMEmitter::ShiftType Shift = FEXCore::ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
constexpr uint32_t Op = 0b010'1010'001U << 21;
DataProcessing_Shifted_Reg(Op, s, rd, rn, rm, Shift, amt);
}
void eor(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, FEXCore::ARMEmitter::Register rm, FEXCore::ARMEmitter::ShiftType Shift = FEXCore::ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
constexpr uint32_t Op = 0b100'1010'000U << 21;
DataProcessing_Shifted_Reg(Op, s, rd, rn, rm, Shift, amt);
}
void eon(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, FEXCore::ARMEmitter::Register rm, FEXCore::ARMEmitter::ShiftType Shift = FEXCore::ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
constexpr uint32_t Op = 0b100'1010'001U << 21;
DataProcessing_Shifted_Reg(Op, s, rd, rn, rm, Shift, amt);
}
// AddSub - shifted register
void add(FEXCore::ARMEmitter::XRegister rd, FEXCore::ARMEmitter::XRegister rn, FEXCore::ARMEmitter::XRegister rm, FEXCore::ARMEmitter::ShiftType Shift = FEXCore::ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
add(ARMEmitter::Size::i64Bit, rd.R(), rn.R(), rm.R(), Shift, amt);
}
void adds(FEXCore::ARMEmitter::XRegister rd, FEXCore::ARMEmitter::XRegister rn, FEXCore::ARMEmitter::XRegister rm, FEXCore::ARMEmitter::ShiftType Shift = FEXCore::ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
adds(ARMEmitter::Size::i64Bit, rd.R(), rn.R(), rm.R(), Shift, amt);
}
void sub(FEXCore::ARMEmitter::XRegister rd, FEXCore::ARMEmitter::XRegister rn, FEXCore::ARMEmitter::XRegister rm, FEXCore::ARMEmitter::ShiftType Shift = FEXCore::ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
sub(ARMEmitter::Size::i64Bit, rd.R(), rn.R(), rm.R(), Shift, amt);
}
void neg(FEXCore::ARMEmitter::XRegister rd, FEXCore::ARMEmitter::XRegister rm, FEXCore::ARMEmitter::ShiftType Shift = FEXCore::ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
sub(rd, FEXCore::ARMEmitter::XReg::zr, rm, Shift, amt);
}
void cmp(FEXCore::ARMEmitter::XRegister rn, FEXCore::ARMEmitter::XRegister rm, FEXCore::ARMEmitter::ShiftType Shift = FEXCore::ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
subs(ARMEmitter::Size::i64Bit, FEXCore::ARMEmitter::Reg::rsp, rn.R(), rm.R(), Shift, amt);
}
void subs(FEXCore::ARMEmitter::XRegister rd, FEXCore::ARMEmitter::XRegister rn, FEXCore::ARMEmitter::XRegister rm, FEXCore::ARMEmitter::ShiftType Shift = FEXCore::ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
subs(ARMEmitter::Size::i64Bit, rd.R(), rn.R(), rm.R(), Shift, amt);
}
void negs(FEXCore::ARMEmitter::XRegister rd, FEXCore::ARMEmitter::XRegister rm, FEXCore::ARMEmitter::ShiftType Shift = FEXCore::ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
subs(rd, FEXCore::ARMEmitter::XReg::zr, rm, Shift, amt);
}
void add(FEXCore::ARMEmitter::WRegister rd, FEXCore::ARMEmitter::WRegister rn, FEXCore::ARMEmitter::WRegister rm, FEXCore::ARMEmitter::ShiftType Shift = FEXCore::ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
add(ARMEmitter::Size::i32Bit, rd.R(), rn.R(), rm.R(), Shift, amt);
}
void adds(FEXCore::ARMEmitter::WRegister rd, FEXCore::ARMEmitter::WRegister rn, FEXCore::ARMEmitter::WRegister rm, FEXCore::ARMEmitter::ShiftType Shift = FEXCore::ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
adds(ARMEmitter::Size::i32Bit, rd.R(), rn.R(), rm.R(), Shift, amt);
}
void sub(FEXCore::ARMEmitter::WRegister rd, FEXCore::ARMEmitter::WRegister rn, FEXCore::ARMEmitter::WRegister rm, FEXCore::ARMEmitter::ShiftType Shift = FEXCore::ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
sub(ARMEmitter::Size::i32Bit, rd.R(), rn.R(), rm.R(), Shift, amt);
}
void neg(FEXCore::ARMEmitter::WRegister rd, FEXCore::ARMEmitter::WRegister rm, FEXCore::ARMEmitter::ShiftType Shift = FEXCore::ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
sub(rd, FEXCore::ARMEmitter::WReg::zr, rm, Shift, amt);
}
void cmp(FEXCore::ARMEmitter::WRegister rn, FEXCore::ARMEmitter::WRegister rm, FEXCore::ARMEmitter::ShiftType Shift = FEXCore::ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
subs(ARMEmitter::Size::i32Bit, FEXCore::ARMEmitter::Reg::rsp, rn.R(), rm.R(), Shift, amt);
}
void subs(FEXCore::ARMEmitter::WRegister rd, FEXCore::ARMEmitter::WRegister rn, FEXCore::ARMEmitter::WRegister rm, FEXCore::ARMEmitter::ShiftType Shift = FEXCore::ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
subs(ARMEmitter::Size::i32Bit, rd.R(), rn.R(), rm.R(), Shift, amt);
}
void negs(FEXCore::ARMEmitter::WRegister rd, FEXCore::ARMEmitter::WRegister rm, FEXCore::ARMEmitter::ShiftType Shift = FEXCore::ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
subs(rd, FEXCore::ARMEmitter::WReg::zr, rm, Shift, amt);
}
void add(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, FEXCore::ARMEmitter::Register rm, FEXCore::ARMEmitter::ShiftType Shift = FEXCore::ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
LOGMAN_THROW_AA_FMT(Shift != FEXCore::ARMEmitter::ShiftType::ROR, "Doesn't support ROR");
constexpr uint32_t Op = 0b000'1011'000U << 21;
DataProcessing_Shifted_Reg(Op, s, rd, rn, rm, Shift, amt);
}
void adds(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, FEXCore::ARMEmitter::Register rm, FEXCore::ARMEmitter::ShiftType Shift = FEXCore::ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
LOGMAN_THROW_AA_FMT(Shift != FEXCore::ARMEmitter::ShiftType::ROR, "Doesn't support ROR");
constexpr uint32_t Op = 0b010'1011'000U << 21;
DataProcessing_Shifted_Reg(Op, s, rd, rn, rm, Shift, amt);
}
void sub(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, FEXCore::ARMEmitter::Register rm, FEXCore::ARMEmitter::ShiftType Shift = FEXCore::ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
LOGMAN_THROW_AA_FMT(Shift != FEXCore::ARMEmitter::ShiftType::ROR, "Doesn't support ROR");
constexpr uint32_t Op = 0b100'1011'000U << 21;
DataProcessing_Shifted_Reg(Op, s, rd, rn, rm, Shift, amt);
}
void neg(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rm, FEXCore::ARMEmitter::ShiftType Shift = FEXCore::ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
sub(s, rd, FEXCore::ARMEmitter::Reg::zr, rm, Shift, amt);
}
void cmp(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rn, FEXCore::ARMEmitter::Register rm, FEXCore::ARMEmitter::ShiftType Shift = FEXCore::ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
subs(s, FEXCore::ARMEmitter::Reg::zr, rn, rm, Shift, amt);
}
void subs(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, FEXCore::ARMEmitter::Register rm, FEXCore::ARMEmitter::ShiftType Shift = FEXCore::ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
LOGMAN_THROW_AA_FMT(Shift != FEXCore::ARMEmitter::ShiftType::ROR, "Doesn't support ROR");
constexpr uint32_t Op = 0b110'1011'000U << 21;
DataProcessing_Shifted_Reg(Op, s, rd, rn, rm, Shift, amt);
}
void negs(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rm, FEXCore::ARMEmitter::ShiftType Shift = FEXCore::ARMEmitter::ShiftType::LSL, uint32_t amt = 0) {
subs(s, rd, FEXCore::ARMEmitter::Reg::zr, rm, Shift, amt);
}
// AddSub - extended register
void add(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, FEXCore::ARMEmitter::Register rm, FEXCore::ARMEmitter::ExtendedType Option, uint32_t Shift = 0) {
LOGMAN_THROW_AA_FMT(Shift <= 4, "Shift amount is too large");
constexpr uint32_t Op = 0b000'1011'001U << 21;
DataProcessing_Extended_Reg(Op, s, rd, rn, rm, Option, Shift);
}
void adds(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, FEXCore::ARMEmitter::Register rm, FEXCore::ARMEmitter::ExtendedType Option, uint32_t Shift = 0) {
constexpr uint32_t Op = 0b010'1011'001U << 21;
DataProcessing_Extended_Reg(Op, s, rd, rn, rm, Option, Shift);
}
void sub(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, FEXCore::ARMEmitter::Register rm, FEXCore::ARMEmitter::ExtendedType Option, uint32_t Shift = 0) {
constexpr uint32_t Op = 0b100'1011'001U << 21;
DataProcessing_Extended_Reg(Op, s, rd, rn, rm, Option, Shift);
}
void subs(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, FEXCore::ARMEmitter::Register rm, FEXCore::ARMEmitter::ExtendedType Option, uint32_t Shift = 0) {
constexpr uint32_t Op = 0b110'1011'001U << 21;
DataProcessing_Extended_Reg(Op, s, rd, rn, rm, Option, Shift);
}
void cmp(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rn, FEXCore::ARMEmitter::Register rm, FEXCore::ARMEmitter::ExtendedType Option, uint32_t Shift = 0) {
constexpr uint32_t Op = 0b110'1011'001U << 21;
DataProcessing_Extended_Reg(Op, s, FEXCore::ARMEmitter::Reg::zr, rn, rm, Option, Shift);
}
// AddSub - with carry
void adc(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, FEXCore::ARMEmitter::Register rm) {
constexpr uint32_t Op = 0b0001'1010'000U << 21;
DataProcessing_Extended_Reg(Op, s, rd, rn, rm, FEXCore::ARMEmitter::ExtendedType::UXTB, 0);
}
void adcs(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, FEXCore::ARMEmitter::Register rm) {
constexpr uint32_t Op = 0b0011'1010'000U << 21;
DataProcessing_Extended_Reg(Op, s, rd, rn, rm, FEXCore::ARMEmitter::ExtendedType::UXTB, 0);
}
void sbc(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, FEXCore::ARMEmitter::Register rm) {
constexpr uint32_t Op = 0b0101'1010'000U << 21;
DataProcessing_Extended_Reg(Op, s, rd, rn, rm, FEXCore::ARMEmitter::ExtendedType::UXTB, 0);
}
void sbcs(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, FEXCore::ARMEmitter::Register rm) {
constexpr uint32_t Op = 0b0111'1010'000U << 21;
DataProcessing_Extended_Reg(Op, s, rd, rn, rm, FEXCore::ARMEmitter::ExtendedType::UXTB, 0);
}
// Rotate right into flags
void rmif(XRegister rn, uint32_t shift, uint32_t mask) {
LOGMAN_THROW_AA_FMT(shift <= 63, "Shift must be within 0-63. Shift: {}", shift);
LOGMAN_THROW_AA_FMT(mask <= 15, "Mask must be within 0-15. Mask: {}", mask);
uint32_t Op = 0b1011'1010'0000'0000'0000'0100'0000'0000;
Op |= rn.Idx() << 5;
Op |= shift << 15;
Op |= mask;
dc32(Op);
}
// Evaluate into flags
void setf8(WRegister rn) {
constexpr uint32_t Op = 0b0011'1010'0000'0000'0000'1000'0000'1101;
EvaluateIntoFlags(Op, 0, rn);
}
void setf16(WRegister rn) {
constexpr uint32_t Op = 0b0011'1010'0000'0000'0000'1000'0000'1101;
EvaluateIntoFlags(Op, 1, rn);
}
// Conditional compare - register
void ccmn(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rn, FEXCore::ARMEmitter::Register rm, FEXCore::ARMEmitter::StatusFlags flags, FEXCore::ARMEmitter::Condition Cond) {
constexpr uint32_t Op = 0b0011'1010'010 << 21;
ConditionalCompare(Op, 0, 0b00, 0, s, rn, rm, flags, Cond);
}
void ccmp(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rn, FEXCore::ARMEmitter::Register rm, FEXCore::ARMEmitter::StatusFlags flags, FEXCore::ARMEmitter::Condition Cond) {
constexpr uint32_t Op = 0b0011'1010'010 << 21;
ConditionalCompare(Op, 1, 0b00, 0, s, rn, rm, flags, Cond);
}
// Conditional compare - immediate
void ccmn(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rn, uint32_t rm, FEXCore::ARMEmitter::StatusFlags flags, FEXCore::ARMEmitter::Condition Cond) {
LOGMAN_THROW_A_FMT((rm & ~0b1'1111) == 0, "Comparison imm too large");
constexpr uint32_t Op = 0b0011'1010'010 << 21;
ConditionalCompare(Op, 0, 0b10, 0, s, rn, rm, flags, Cond);
}
void ccmp(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rn, uint32_t rm, FEXCore::ARMEmitter::StatusFlags flags, FEXCore::ARMEmitter::Condition Cond) {
LOGMAN_THROW_A_FMT((rm & ~0b1'1111) == 0, "Comparison imm too large");
constexpr uint32_t Op = 0b0011'1010'010 << 21;
ConditionalCompare(Op, 1, 0b10, 0, s, rn, rm, flags, Cond);
}
// Conditional select
void csel(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, FEXCore::ARMEmitter::Register rm, FEXCore::ARMEmitter::Condition Cond) {
constexpr uint32_t Op = 0b0001'1010'100 << 21;
ConditionalCompare(Op, 0, 0b00, s, rd, rn, rm, Cond);
}
void cset(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Condition Cond) {
constexpr uint32_t Op = 0b0001'1010'100 << 21;
ConditionalCompare(Op, 0, 0b01, s, rd, FEXCore::ARMEmitter::Reg::zr, FEXCore::ARMEmitter::Reg::zr, static_cast<FEXCore::ARMEmitter::Condition>(FEXCore::ToUnderlying(Cond) ^ FEXCore::ToUnderlying(FEXCore::ARMEmitter::Condition::CC_NE)));
}
void csinc(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, FEXCore::ARMEmitter::Register rm, FEXCore::ARMEmitter::Condition Cond) {
constexpr uint32_t Op = 0b0001'1010'100 << 21;
ConditionalCompare(Op, 0, 0b01, s, rd, rn, rm, Cond);
}
void csinv(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, FEXCore::ARMEmitter::Register rm, FEXCore::ARMEmitter::Condition Cond) {
constexpr uint32_t Op = 0b0001'1010'100 << 21;
ConditionalCompare(Op, 1, 0b00, s, rd, rn, rm, Cond);
}
void csneg(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, FEXCore::ARMEmitter::Register rm, FEXCore::ARMEmitter::Condition Cond) {
constexpr uint32_t Op = 0b0001'1010'100 << 21;
ConditionalCompare(Op, 1, 0b01, s, rd, rn, rm, Cond);
}
// Data processing - 3 source
void madd(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, FEXCore::ARMEmitter::Register rm, FEXCore::ARMEmitter::Register ra) {
constexpr uint32_t Op = 0b001'1011'000U << 21;
DataProcessing_3Source(Op, 0, s, rd, rn, rm, ra);
}
void mul(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, FEXCore::ARMEmitter::Register rm) {
madd(s, rd, rn, rm, XReg::zr);
}
void msub(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, FEXCore::ARMEmitter::Register rm, FEXCore::ARMEmitter::Register ra) {
constexpr uint32_t Op = 0b001'1011'000U << 21;
DataProcessing_3Source(Op, 1, s, rd, rn, rm, ra);
}
void mneg(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, FEXCore::ARMEmitter::Register rm) {
msub(s, rd, rn, rm, XReg::zr);
}
void smaddl(FEXCore::ARMEmitter::XRegister rd, FEXCore::ARMEmitter::WRegister rn, FEXCore::ARMEmitter::WRegister rm, FEXCore::ARMEmitter::XRegister ra) {
constexpr uint32_t Op = 0b001'1011'001U << 21;
DataProcessing_3Source(Op, 0, FEXCore::ARMEmitter::Size::i64Bit, rd, rn, rm, ra);
}
void smull(FEXCore::ARMEmitter::XRegister rd, FEXCore::ARMEmitter::WRegister rn, FEXCore::ARMEmitter::WRegister rm) {
smaddl(rd, rn, rm, XReg::zr);
}
void smsubl(FEXCore::ARMEmitter::XRegister rd, FEXCore::ARMEmitter::WRegister rn, FEXCore::ARMEmitter::WRegister rm, FEXCore::ARMEmitter::XRegister ra) {
constexpr uint32_t Op = 0b001'1011'001U << 21;
DataProcessing_3Source(Op, 1, FEXCore::ARMEmitter::Size::i64Bit, rd, rn, rm, ra);
}
void smnegl(FEXCore::ARMEmitter::XRegister rd, FEXCore::ARMEmitter::WRegister rn, FEXCore::ARMEmitter::WRegister rm) {
smsubl(rd, rn, rm, XReg::zr);
}
void smulh(FEXCore::ARMEmitter::XRegister rd, FEXCore::ARMEmitter::XRegister rn, FEXCore::ARMEmitter::XRegister rm) {
constexpr uint32_t Op = 0b001'1011'010U << 21;
DataProcessing_3Source(Op, 0, FEXCore::ARMEmitter::Size::i64Bit, rd, rn, rm, FEXCore::ARMEmitter::Reg::zr);
}
void umaddl(FEXCore::ARMEmitter::XRegister rd, FEXCore::ARMEmitter::WRegister rn, FEXCore::ARMEmitter::WRegister rm, FEXCore::ARMEmitter::XRegister ra) {
constexpr uint32_t Op = 0b001'1011'101U << 21;
DataProcessing_3Source(Op, 0, FEXCore::ARMEmitter::Size::i64Bit, rd, rn, rm, ra);
}
void umull(FEXCore::ARMEmitter::XRegister rd, FEXCore::ARMEmitter::WRegister rn, FEXCore::ARMEmitter::WRegister rm) {
umaddl(rd, rn, rm, XReg::zr);
}
void umsubl(FEXCore::ARMEmitter::XRegister rd, FEXCore::ARMEmitter::WRegister rn, FEXCore::ARMEmitter::WRegister rm, FEXCore::ARMEmitter::XRegister ra) {
constexpr uint32_t Op = 0b001'1011'101U << 21;
DataProcessing_3Source(Op, 1, FEXCore::ARMEmitter::Size::i64Bit, rd, rn, rm, ra);
}
void umnegl(FEXCore::ARMEmitter::XRegister rd, FEXCore::ARMEmitter::WRegister rn, FEXCore::ARMEmitter::WRegister rm) {
umsubl(rd, rn, rm, XReg::zr);
}
void umulh(FEXCore::ARMEmitter::XRegister rd, FEXCore::ARMEmitter::XRegister rn, FEXCore::ARMEmitter::XRegister rm) {
constexpr uint32_t Op = 0b001'1011'110U << 21;
DataProcessing_3Source(Op, 0, FEXCore::ARMEmitter::Size::i64Bit, rd, rn, rm, FEXCore::ARMEmitter::Reg::zr);
}
private:
void and_(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, uint32_t n, uint32_t immr, uint32_t imms) {
constexpr uint32_t Op = 0b001'0010'00 << 22;
DataProcessing_Logical_Imm(Op, s, rd, rn, n, immr, imms);
}
void ands(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, uint32_t n, uint32_t immr, uint32_t imms) {
constexpr uint32_t Op = 0b111'0010'00 << 22;
DataProcessing_Logical_Imm(Op, s, rd, rn, n, immr, imms);
}
void orr(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, uint32_t n, uint32_t immr, uint32_t imms) {
constexpr uint32_t Op = 0b011'0010'00 << 22;
DataProcessing_Logical_Imm(Op, s, rd, rn, n, immr, imms);
}
void eor(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, uint32_t n, uint32_t immr, uint32_t imms) {
constexpr uint32_t Op = 0b101'0010'00 << 22;
DataProcessing_Logical_Imm(Op, s, rd, rn, n, immr, imms);
}
void sbfm(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, uint32_t immr, uint32_t imms) {
constexpr uint32_t Op = 0b0001'0011'00 << 22;
DataProcessing_Logical_Imm(Op, s, rd, rn, s == ARMEmitter::Size::i64Bit, immr, imms);
}
void bfm(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, uint32_t immr, uint32_t imms) {
constexpr uint32_t Op = 0b0011'0011'00 << 22;
DataProcessing_Logical_Imm(Op, s, rd, rn, s == ARMEmitter::Size::i64Bit, immr, imms);
}
// 4.1.64 - Data processing - Immediate
void DataProcessing_PCRel_Imm(uint32_t Op, FEXCore::ARMEmitter::Register rd, uint32_t Imm) {
// Ensure the immediate is masked.
Imm &= 0b1'1111'1111'1111'1111'1111U;
uint32_t Instr = Op;
Instr |= (Imm & 0b11) << 29;
Instr |= (Imm >> 2) << 5;
Instr |= Encode_rd(rd);
dc32(Instr);
}
void DataProcessing_AddSub_Imm(uint32_t Op, FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, uint32_t Imm, bool LSL12) {
bool TooLarge = (Imm & ~0b1111'1111'1111U) != 0;
if (TooLarge && !LSL12 && ((Imm >> 12) & ~0b1111'1111'1111U) == 0) {
// We can convert an immediate
TooLarge = false;
LSL12 = true;
Imm >>= 12;
}
LOGMAN_THROW_AA_FMT(TooLarge == false, "Imm amount too large: 0x{:x}", Imm);
const uint32_t SF = s == FEXCore::ARMEmitter::Size::i64Bit ? (1U << 31) : 0;
uint32_t Instr = Op;
Instr |= SF;
Instr |= LSL12 << 22;
Instr |= Imm << 10;
Instr |= Encode_rn(rn);
Instr |= Encode_rd(rd);
dc32(Instr);
}
// Move Wide
void DataProcessing_MoveWide(uint32_t Op, FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, uint32_t Imm, uint32_t Offset) {
const uint32_t SF = s == FEXCore::ARMEmitter::Size::i64Bit ? (1U << 31) : 0;
uint32_t Instr = Op;
Instr |= SF;
Instr |= Imm << 5;
Instr |= Offset << 21;
Instr |= Encode_rd(rd);
dc32(Instr);
}
// Logical immediate
void DataProcessing_Logical_Imm(uint32_t Op, FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, uint32_t n, uint32_t immr, uint32_t imms) {
const uint32_t SF = s == FEXCore::ARMEmitter::Size::i64Bit ? (1U << 31) : 0;
uint32_t Instr = Op;
Instr |= SF;
Instr |= n << 22;
Instr |= immr << 16;
Instr |= imms << 10;
Instr |= Encode_rn(rn);
Instr |= Encode_rd(rd);
dc32(Instr);
}
void DataProcessing_Extract(uint32_t Op, FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, FEXCore::ARMEmitter::Register rm, uint32_t Imm) {
const uint32_t SF = s == FEXCore::ARMEmitter::Size::i64Bit ? (1U << 31) : 0;
// Current ARMv8 spec hardcodes SF == N for this class of instructions.
// Anythign else is undefined behaviour.
const uint32_t N = s == FEXCore::ARMEmitter::Size::i64Bit ? (1U << 22) : 0;
uint32_t Instr = Op;
Instr |= SF;
Instr |= N;
Instr |= Encode_rm(rm);
Instr |= Imm << 10;
Instr |= Encode_rn(rn);
Instr |= Encode_rd(rd);
dc32(Instr);
}
// Data-processing - 2 source
void DataProcessing_2Source(uint32_t Op, FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, FEXCore::ARMEmitter::Register rm) {
const uint32_t SF = s == FEXCore::ARMEmitter::Size::i64Bit ? (1U << 31) : 0;
uint32_t Instr = Op;
Instr |= SF;
Instr |= Encode_rm(rm);
Instr |= Encode_rn(rn);
Instr |= Encode_rd(rd);
dc32(Instr);
}
// Data processing - 1 source
template<typename T>
void DataProcessing_1Source(uint32_t Op, FEXCore::ARMEmitter::Size s, T rd, T rn) {
const uint32_t SF = s == FEXCore::ARMEmitter::Size::i64Bit ? (1U << 31) : 0;
uint32_t Instr = Op;
Instr |= SF;
Instr |= Encode_rn(rn);
Instr |= Encode_rd(rd);
dc32(Instr);
}
// AddSub - shifted register
void DataProcessing_Shifted_Reg(uint32_t Op, FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, FEXCore::ARMEmitter::Register rm, FEXCore::ARMEmitter::ShiftType Shift, uint32_t amt) {
LOGMAN_THROW_AA_FMT((amt & ~0b11'1111U) == 0, "Shift amount too large");
const uint32_t SF = s == FEXCore::ARMEmitter::Size::i64Bit ? (1U << 31) : 0;
uint32_t Instr = Op;
Instr |= SF;
Instr |= FEXCore::ToUnderlying(Shift) << 22;
Instr |= Encode_rm(rm);
Instr |= static_cast<uint32_t>(amt) << 10;
Instr |= Encode_rn(rn);
Instr |= Encode_rd(rd);
dc32(Instr);
}
// AddSub - extended register
void DataProcessing_Extended_Reg(uint32_t Op, FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, FEXCore::ARMEmitter::Register rm, FEXCore::ARMEmitter::ExtendedType Option, uint32_t Shift) {
const uint32_t SF = s == FEXCore::ARMEmitter::Size::i64Bit ? (1U << 31) : 0;
uint32_t Instr = Op;
Instr |= SF;
Instr |= Encode_rm(rm);
Instr |= FEXCore::ToUnderlying(Option) << 13;
Instr |= static_cast<uint32_t>(Shift) << 10;
Instr |= Encode_rn(rn);
Instr |= Encode_rd(rd);
dc32(Instr);
}
// Conditional compare - register
template<typename T>
void ConditionalCompare(uint32_t Op, uint32_t o1, uint32_t o2, uint32_t o3, FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rn, T rm, FEXCore::ARMEmitter::StatusFlags flags, FEXCore::ARMEmitter::Condition Cond) {
const uint32_t SF = s == FEXCore::ARMEmitter::Size::i64Bit ? (1U << 31) : 0;
uint32_t Instr = Op;
Instr |= SF;
Instr |= o1 << 30;
Instr |= Encode_rm(rm);
Instr |= FEXCore::ToUnderlying(Cond) << 12;
Instr |= o2 << 10;
Instr |= Encode_rn(rn);
Instr |= o3 << 4;
Instr |= FEXCore::ToUnderlying(flags);
dc32(Instr);
}
template<typename T>
void ConditionalCompare(uint32_t Op, uint32_t o1, uint32_t o2, FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, T rm, FEXCore::ARMEmitter::Condition Cond) {
const uint32_t SF = s == FEXCore::ARMEmitter::Size::i64Bit ? (1U << 31) : 0;
uint32_t Instr = Op;
Instr |= SF;
Instr |= o1 << 30;
Instr |= Encode_rm(rm);
Instr |= FEXCore::ToUnderlying(Cond) << 12;
Instr |= o2 << 10;
Instr |= Encode_rn(rn);
Instr |= Encode_rd(rd);
dc32(Instr);
}
// Data-processing - 3 source
void DataProcessing_3Source(uint32_t Op, uint32_t Op0, FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, FEXCore::ARMEmitter::Register rm, FEXCore::ARMEmitter::Register ra) {
const uint32_t SF = s == FEXCore::ARMEmitter::Size::i64Bit ? (1U << 31) : 0;
uint32_t Instr = Op;
Instr |= SF;
Instr |= Encode_rm(rm);
Instr |= Op0 << 15;
Instr |= Encode_ra(ra);
Instr |= Encode_rn(rn);
Instr |= Encode_rd(rd);
dc32(Instr);
}
void EvaluateIntoFlags(uint32_t op, uint32_t size, WRegister rn) {
uint32_t Instr = op;
Instr |= size << 14;
Instr |= rn.Idx() << 5;
dc32(Instr);
}
File diff suppressed because it is too large. Load diff
@@ -0,0 +1,322 @@
/* Branch instruction emitters.
*
* Most of these instructions will use `BackwardLabel`, `ForwardLabel`, or `BiDirectionLabel` to determine where a branch targets.
*/
public:
// Branches, Exception Generating and System instructions
public:
// Conditional branch immediate
///< Branch conditional
void b(FEXCore::ARMEmitter::Condition Cond, uint32_t Imm) {
constexpr uint32_t Op = 0b0101'010 << 25;
Branch_Conditional(Op, 0, 0, Cond, Imm);
}
void b(FEXCore::ARMEmitter::Condition Cond, BackwardLabel const* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0), "Unscaled offset too large");
constexpr uint32_t Op = 0b0101'010 << 25;
Branch_Conditional(Op, 0, 0, Cond, Imm >> 2);
}
void b(FEXCore::ARMEmitter::Condition Cond, ForwardLabel *Label) {
Label->Insts.emplace_back(ForwardLabel::Instructions{ .Location = GetCursorAddress<uint8_t*>(), .Type = ForwardLabel::Instructions::InstType::BC });
constexpr uint32_t Op = 0b0101'010 << 25;
Branch_Conditional(Op, 0, 0, Cond, 0);
}
void b(FEXCore::ARMEmitter::Condition Cond, BiDirectionalLabel *Label) {
if (Label->Backward.Location) {
b(Cond, &Label->Backward);
}
else {
b(Cond, &Label->Forward);
}
}
///< Branch consistent conditional
void bc(FEXCore::ARMEmitter::Condition Cond, uint32_t Imm) {
constexpr uint32_t Op = 0b0101'010 << 25;
Branch_Conditional(Op, 0, 1, Cond, Imm);
}
void bc(FEXCore::ARMEmitter::Condition Cond, BackwardLabel const* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0), "Unscaled offset too large");
constexpr uint32_t Op = 0b0101'010 << 25;
Branch_Conditional(Op, 0, 1, Cond, Imm >> 2);
}
void bc(FEXCore::ARMEmitter::Condition Cond, ForwardLabel *Label) {
Label->Insts.emplace_back(ForwardLabel::Instructions{ .Location = GetCursorAddress<uint8_t*>(), .Type = ForwardLabel::Instructions::InstType::BC });
constexpr uint32_t Op = 0b0101'010 << 25;
Branch_Conditional(Op, 0, 1, Cond, 0);
}
void bc(FEXCore::ARMEmitter::Condition Cond, BiDirectionalLabel *Label) {
if (Label->Backward.Location) {
bc(Cond, &Label->Backward);
}
else {
bc(Cond, &Label->Forward);
}
}
// Unconditional branch register
void br(FEXCore::ARMEmitter::Register rn) {
constexpr uint32_t Op = 0b1101011 << 25 |
0b0'000 << 21 | // opc
0b1'1111 << 16 | // op2
0b0000'00 << 10 | // op3
0b0'0000; // op4
UnconditionalBranch(Op, rn);
}
void blr(FEXCore::ARMEmitter::Register rn) {
constexpr uint32_t Op = 0b1101011 << 25 |
0b0'001 << 21 | // opc
0b1'1111 << 16 | // op2
0b0000'00 << 10 | // op3
0b0'0000; // op4
UnconditionalBranch(Op, rn);
}
void ret(FEXCore::ARMEmitter::Register rn = FEXCore::ARMEmitter::Reg::r30) {
constexpr uint32_t Op = 0b1101011 << 25 |
0b0'010 << 21 | // opc
0b1'1111 << 16 | // op2
0b0000'00 << 10 | // op3
0b0'0000; // op4
UnconditionalBranch(Op, rn);
}
// Unconditional branch immediate
void b(uint32_t Imm) {
constexpr uint32_t Op = 0b0001'01 << 26;
UnconditionalBranch(Op, Imm);
}
void b(BackwardLabel const* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(Imm >= -134217728 && Imm <= 134217724 && ((Imm & 0b11) == 0), "Unscaled offset too large");
constexpr uint32_t Op = 0b0001'01 << 26;
UnconditionalBranch(Op, Imm >> 2);
}
void b(ForwardLabel *Label) {
Label->Insts.emplace_back(ForwardLabel::Instructions{ .Location = GetCursorAddress<uint8_t*>(), .Type = ForwardLabel::Instructions::InstType::B });
constexpr uint32_t Op = 0b0001'01 << 26;
UnconditionalBranch(Op, 0);
}
void b(BiDirectionalLabel *Label) {
if (Label->Backward.Location) {
b(&Label->Backward);
}
else {
b(&Label->Forward);
}
}
void bl(uint32_t Imm) {
constexpr uint32_t Op = 0b1001'01 << 26;
UnconditionalBranch(Op, Imm);
}
void bl(BackwardLabel const* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(Imm >= -134217728 && Imm <= 134217724 && ((Imm & 0b11) == 0), "Unscaled offset too large");
constexpr uint32_t Op = 0b1001'01 << 26;
UnconditionalBranch(Op, Imm >> 2);
}
void bl(ForwardLabel *Label) {
Label->Insts.emplace_back(ForwardLabel::Instructions{ .Location = GetCursorAddress<uint8_t*>(), .Type = ForwardLabel::Instructions::InstType::B });
constexpr uint32_t Op = 0b1001'01 << 26;
UnconditionalBranch(Op, 0);
}
void bl(BiDirectionalLabel *Label) {
if (Label->Backward.Location) {
bl(&Label->Backward);
}
else {
bl(&Label->Forward);
}
}
// Compare and branch
void cbz(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rt, uint32_t Imm) {
constexpr uint32_t Op = 0b0011'0100 << 24;
CompareAndBranch(Op, s, rt, Imm);
}
void cbz(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rt, BackwardLabel const* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0), "Unscaled offset too large");
constexpr uint32_t Op = 0b0011'0100 << 24;
CompareAndBranch(Op, s, rt, Imm >> 2);
}
void cbz(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rt, ForwardLabel *Label) {
Label->Insts.emplace_back(ForwardLabel::Instructions{ .Location = GetCursorAddress<uint8_t*>(), .Type = ForwardLabel::Instructions::InstType::BC });
constexpr uint32_t Op = 0b0011'0100 << 24;
CompareAndBranch(Op, s, rt, 0);
}
void cbz(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rt, BiDirectionalLabel *Label) {
if (Label->Backward.Location) {
cbz(s, rt, &Label->Backward);
}
else {
cbz(s, rt, &Label->Forward);
}
}
void cbnz(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rt, uint32_t Imm) {
constexpr uint32_t Op = 0b0011'0101 << 24;
CompareAndBranch(Op, s, rt, Imm);
}
void cbnz(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rt, BackwardLabel const* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0), "Unscaled offset too large");
constexpr uint32_t Op = 0b0011'0101 << 24;
CompareAndBranch(Op, s, rt, Imm >> 2);
}
void cbnz(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rt, ForwardLabel *Label) {
Label->Insts.emplace_back(ForwardLabel::Instructions{ .Location = GetCursorAddress<uint8_t*>(), .Type = ForwardLabel::Instructions::InstType::BC });
constexpr uint32_t Op = 0b0011'0101 << 24;
CompareAndBranch(Op, s, rt, 0);
}
void cbnz(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rt, BiDirectionalLabel *Label) {
if (Label->Backward.Location) {
cbnz(s, rt, &Label->Backward);
}
else {
cbnz(s, rt, &Label->Forward);
}
}
// Test and branch immediate
void tbz(FEXCore::ARMEmitter::Register rt, uint32_t Bit, uint32_t Imm) {
constexpr uint32_t Op = 0b0011'0110 << 24;
TestAndBranch(Op, rt, Bit, Imm);
}
void tbz(FEXCore::ARMEmitter::Register rt, uint32_t Bit, BackwardLabel const* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(Imm >= -32768 && Imm <= 32764 && ((Imm & 0b11) == 0), "Unscaled offset too large");
constexpr uint32_t Op = 0b0011'0110 << 24;
TestAndBranch(Op, rt, Bit, Imm >> 2);
}
void tbz(FEXCore::ARMEmitter::Register rt, uint32_t Bit, ForwardLabel *Label) {
Label->Insts.emplace_back(ForwardLabel::Instructions{ .Location = GetCursorAddress<uint8_t*>(), .Type = ForwardLabel::Instructions::InstType::TEST_BRANCH });
constexpr uint32_t Op = 0b0011'0110 << 24;
TestAndBranch(Op, rt, Bit, 0);
}
void tbz(FEXCore::ARMEmitter::Register rt, uint32_t Bit, BiDirectionalLabel *Label) {
if (Label->Backward.Location) {
tbz(rt, Bit, &Label->Backward);
}
else {
tbz(rt, Bit, &Label->Forward);
}
}
void tbnz(FEXCore::ARMEmitter::Register rt, uint32_t Bit, uint32_t Imm) {
constexpr uint32_t Op = 0b0011'0111 << 24;
TestAndBranch(Op, rt, Bit, Imm);
}
void tbnz(FEXCore::ARMEmitter::Register rt, uint32_t Bit, BackwardLabel const* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(Imm >= -32768 && Imm <= 32764 && ((Imm & 0b11) == 0), "Unscaled offset too large");
constexpr uint32_t Op = 0b0011'0111 << 24;
TestAndBranch(Op, rt, Bit, Imm >> 2);
}
void tbnz(FEXCore::ARMEmitter::Register rt, uint32_t Bit, ForwardLabel *Label) {
Label->Insts.emplace_back(ForwardLabel::Instructions{ .Location = GetCursorAddress<uint8_t*>(), .Type = ForwardLabel::Instructions::InstType::TEST_BRANCH });
constexpr uint32_t Op = 0b0011'0111 << 24;
TestAndBranch(Op, rt, Bit, 0);
}
void tbnz(FEXCore::ARMEmitter::Register rt, uint32_t Bit, BiDirectionalLabel *Label) {
if (Label->Backward.Location) {
tbnz(rt, Bit, &Label->Backward);
}
else {
tbnz(rt, Bit, &Label->Forward);
}
}
private:
// Conditional branch immediate
void Branch_Conditional(uint32_t Op, uint32_t Op1, uint32_t Op0, FEXCore::ARMEmitter::Condition Cond, uint32_t Imm) {
uint32_t Instr = Op;
Instr |= Op1 << 24;
Instr |= (Imm & 0x7'FFFF) << 5;
Instr |= Op0 << 4;
Instr |= FEXCore::ToUnderlying(Cond);
dc32(Instr);
}
// Unconditional branch register
void UnconditionalBranch(uint32_t Op, FEXCore::ARMEmitter::Register rn) {
uint32_t Instr = Op;
Instr |= Encode_rn(rn);
dc32(Instr);
}
// Unconditional branch - immediate
void UnconditionalBranch(uint32_t Op, uint32_t Imm) {
uint32_t Instr = Op;
Instr |= Imm & 0x3FF'FFFF;
dc32(Instr);
}
// Compare and branch
void CompareAndBranch(uint32_t Op, FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rt, uint32_t Imm) {
const uint32_t SF = s == FEXCore::ARMEmitter::Size::i64Bit ? (1U << 31) : 0;
uint32_t Instr = Op;
Instr |= SF;
Instr |= (Imm & 0x7'FFFF) << 5;
Instr |= Encode_rt(rt);
dc32(Instr);
}
// Test and branch - immediate
void TestAndBranch(uint32_t Op, FEXCore::ARMEmitter::Register rt, uint32_t Bit, uint32_t Imm) {
uint32_t Instr = Op;
Instr |= (Bit >> 5) << 31;
Instr |= (Bit & 0b1'1111) << 19;
Instr |= (Imm & 0x3FFF) << 5;
Instr |= Encode_rt(rt);
dc32(Instr);
}
@@ -0,0 +1,105 @@
#pragma once
#include <cstddef>
#include <cstdint>
#include <cstring>
namespace FEXCore::ARMEmitter {
class Buffer {
public:
Buffer() {
SetBuffer(nullptr, 0);
}
Buffer(uint8_t* Base, uint64_t BaseSize) {
SetBuffer(Base, BaseSize);
}
void SetBuffer(uint8_t* Base, uint64_t BaseSize) {
BufferBase = Base;
CurrentOffset = BufferBase;
Size = BaseSize;
}
void dc8(uint8_t Data) {
decltype(Data) *Memory = reinterpret_cast<decltype(Data)*>(CurrentOffset);
*Memory = Data;
CurrentOffset += sizeof(Data);
}
void dc16(uint16_t Data) {
decltype(Data) *Memory = reinterpret_cast<decltype(Data)*>(CurrentOffset);
*Memory = Data;
CurrentOffset += sizeof(Data);
}
void dc32(uint32_t Data) {
decltype(Data) *Memory = reinterpret_cast<decltype(Data)*>(CurrentOffset);
*Memory = Data;
CurrentOffset += sizeof(Data);
}
void dc64(uint64_t Data) {
decltype(Data) *Memory = reinterpret_cast<decltype(Data)*>(CurrentOffset);
*Memory = Data;
CurrentOffset += sizeof(Data);
}
void EmitString(const char *String) {
const auto StringLength = strlen(String);
memcpy(CurrentOffset, String, StringLength);
CurrentOffset += StringLength;
}
void Align() {
// Align the buffer to instruction size
auto CurrentAlignment = reinterpret_cast<uint64_t>(CurrentOffset) & 0b11;
if (!CurrentAlignment) {
return;
}
CurrentOffset += 4 - CurrentAlignment;
}
template<typename T>
T GetCursorAddress() const {
return reinterpret_cast<T>(CurrentOffset);
}
static void ClearICache(void* Begin, std::size_t Length) {
__builtin___clear_cache(static_cast<char*>(Begin), static_cast<char*>(Begin) + Length);
}
size_t GetCursorOffset() const {
return static_cast<size_t>(CurrentOffset - BufferBase);
}
uint8_t *GetBufferBase() const {
return BufferBase;
}
void CursorIncrement(size_t Size) {
CurrentOffset += Size;
}
void SetCursorOffset(size_t Offset) {
CurrentOffset = BufferBase + Offset;
}
uint64_t GetBufferSize() const {
return Size;
}
template<typename T>
size_t GetCursorOffsetFromAddress(const T* Address) const {
return static_cast<size_t>(reinterpret_cast<const uint8_t*>(Address) - BufferBase);
}
protected:
void ResetBuffer() {
CurrentOffset = BufferBase;
}
uint8_t* BufferBase;
uint8_t* CurrentOffset;
uint64_t Size;
};
}
@@ -0,0 +1,801 @@
#pragma once
#include "Interface/Core/ArchHelpers/CodeEmitter/Buffer.h"
#include "Interface/Core/ArchHelpers/CodeEmitter/Registers.h"
#include <FEXCore/Utils/CompilerDefs.h>
#include <FEXCore/Utils/EnumUtils.h>
#include <FEXCore/Utils/LogManager.h>
#include <aarch64/assembler-aarch64.h>
#include <cstdint>
#include <utility>
#include <type_traits>
#include <vector>
/*
* Welcome to FEX-Emu's custom AArch64 emitter.
* This was written specifically to avoid the performance cost of the vixl emitter.
*
* There are some specific design constraints in this design to target a couple features:
* - High performance
* - Low CPU cache performance hit
* - Significantly reduced code footprint
* - Low number of branches
*
* These requirements are mostly achieved by removing a bunch of developer conveniences
* that vixl provides. The developer needs to take a lot of care to not shoot themselves in the foot.
*
* Misc design decisions:
* - Registers are encoded as basic uint32_t enums.
* - Converting between different registers is zero-cost.
* - Passing around as arguments are as cheap as registers
* - Contrast to vixl where every register requires living on the stack.
* - Registers can get encoded in to instructions with a simple `BFM` instruction.
*
* - Instructions are very simply emitted, allowing direct inlining most of the time.
* - These are simple enough that multiple back-to-back instructions get optimized to 128-bit load-store operations.
* - Contrast to vixl where pretty much no instruction emitter gets inlined.
*
* - Instruction emitters are /mostly/ unsized. Most instructions take a size argument first, which gets encoded
* directly in to the instruction.
* - Contrast to vixl where the register arguments are how the instructions determine operating size.
* - Size argument allows FEX to use `CSEL` to select a size at runtime, instead of branching.
* - Some instructions are explicitly sized based on register type. Read comments in the respective `inl` files to
* see why.
* Some scalar/vector operations are an example of this.
*
* - Almost zero helper functions.
* - Primary exception to this rule is load-store operations. These will use a helper to make
* it easier to select the correct load-store instruction. Mostly because these are a nightmare selecting
* the right instruction.
*/
namespace FEXCore::ARMEmitter {
/*
* This `Size` enum is used for most ALU operations.
* These follow the AArch64 encoding style in most cases.
*/
enum class Size : uint32_t {
i32Bit = 0,
i64Bit,
};
// This allows us to get the `Size` enum in bits.
template<Size size>
constexpr size_t RegSizeInBits() {
constexpr size_t RegSize[] = {
32, 64, 128,
};
return RegSize[FEXCore::ToUnderlying(size)];
}
[[maybe_unused]]
static inline size_t RegSizeInBits(Size size) {
constexpr size_t RegSize[] = {
32, 64, 128,
};
return RegSize[FEXCore::ToUnderlying(size)];
}
/* This `SubRegSize` enum is used for most ASIMD operations.
* These follow the AArch64 encoding style in most cases.
*/
enum class SubRegSize : uint32_t {
i8Bit = 0b00,
i16Bit = 0b01,
i32Bit = 0b10,
i64Bit = 0b11,
i128Bit = 0b100,
};
// This allows us to get the `SubRegSize` in bits.
template<SubRegSize size>
constexpr size_t SubRegSizeInBits() {
return (1 << FEXCore::ToUnderlying(size)) * 8;
}
[[maybe_unused]]
static inline size_t SubRegSizeInBits(SubRegSize size) {
return (1 << FEXCore::ToUnderlying(size)) * 8;
}
/* This `ScalarRegSize` enum is used for most scalar float
* operations.
*
* This is specifically duplicated from `SubRegSize` to have strongly
* typed functions.
*
* `ScalarRegSize` specifically doesn't have `i128Bit` because scalar operations
* can't operate at 128-bit.
*/
enum class ScalarRegSize : uint32_t {
i8Bit = 0b00,
i16Bit = 0b01,
i32Bit = 0b10,
i64Bit = 0b11,
};
// This allows us to get the `ScalarRegSize` in bits.
template<ScalarRegSize size>
constexpr size_t ScalarRegSizeInBits() {
return (1 << FEXCore::ToUnderlying(size)) * 8;
}
[[maybe_unused]]
static inline size_t ScalarRegSizeInBits(ScalarRegSize size) {
return (1 << FEXCore::ToUnderlying(size)) * 8;
}
/* This `VectorRegSizePair` union allows us to have an overlapping type
* to select a scalar operation or a vector depending on which operation
* we pass in.
* Useful in FEX's vector operations that behave as scalar or vector
* depending on various factors. But since the operation will have the sa,e
* element size, we want to choose the operation more easily
*/
union VectorRegSizePair {
ScalarRegSize Scalar;
SubRegSize Vector;
};
// This allows us to create a `VectorRegSizePair` union.
[[maybe_unused]]
static inline VectorRegSizePair ToVectorSizePair(SubRegSize size) {
return VectorRegSizePair {.Vector = size};
}
[[maybe_unused]]
static inline VectorRegSizePair ToVectorSizePair(ScalarRegSize size) {
return VectorRegSizePair {.Scalar = size};
}
// This `ShiftType` enum is used for ALU shift-register encoded instructions.
enum class ShiftType : uint32_t {
LSL = 0,
LSR,
ASR,
ROR,
};
// This `ExtendedType` enum is used for ALU extended-register encoded instructions.
enum class ExtendedType : uint32_t {
UXTB = 0b000,
UXTH = 0b001,
UXTW = 0b010,
UXTX = 0b011,
SXTB = 0b100,
SXTH = 0b101,
SXTW = 0b110,
SXTX = 0b111,
LSL_32 = UXTW,
LSL_64 = UXTX,
};
// This `Condition` enum is used for various conditional instructions.
enum class Condition : uint32_t {
// Meaning: Int - Float
CC_EQ = 0, // Equal - Equal
CC_NE, // Not Eq - Not Eq or unordered
CC_CS, // Carry set - Greater than, equal, or unordered
CC_CC, // Carry clear - Less than
CC_MI, // Minus/Negative - Less than
CC_PL, // Plus, positive or zero - GT, equal, or unordered
CC_VS, // Overflow - Unordered
CC_VC, // No Overflow - Ordered
CC_HI, // Unsigned higher - GT, or unordered
CC_LS, // Unsigned lower or same - LT or EQ
CC_GE, // Signed GT or EQ - GT or EQ
CC_LT, // Signed LT - LT or Unordered
CC_GT, // Signed GT - GT
CC_LE, // Signed LT or EQ - LT, EQ, or Unordered
CC_AL, // Always - Always
CC_NV, // Always - Always
// Aliases
CC_HS = CC_CS,
CC_LO = CC_CC,
};
/*
* This `StatusFlags` enum is used for conditional compare encoded instructions.
* These directly encode to the `nzcv` flags.
*/
enum class StatusFlags : uint32_t {
None = 0,
Flag_V = 0b0001,
Flag_C = 0b0010,
Flag_Z = 0b0100,
Flag_N = 0b1000,
Flag_NZCV = Flag_N | Flag_Z | Flag_C | Flag_V,
};
/*
* This `IndexType` enum is used for load-store instructions.
* Not all load-store instructions use this, so the user needs to be careful.
*/
enum class IndexType {
POST,
OFFSET,
PRE,
UNPRIVILEGED,
};
/* This `SVEMemOperand` class is used for the helper SVE load-store instructions.
* Load-store instructions are quite expressive, so having a helper that handles these differences is worth it.
*/
class SVEMemOperand final {
public:
SVEMemOperand(XRegister rn, XRegister rm = XReg::zr)
: rn {rn}
, MetaType {
.ScalarScalarType {
.Header = { .MemType = TYPE_SCALAR_SCALAR },
.rm = rm,
}
} {}
SVEMemOperand(XRegister rn, int32_t imm = 0)
: rn {rn}
, MetaType {
.ScalarImmType {
.Header = { .MemType = TYPE_SCALAR_IMM },
.Imm = imm,
}
} {}
Register rn;
enum Type {
TYPE_SCALAR_SCALAR,
TYPE_SCALAR_IMM,
TYPE_SCALAR_VECTOR,
TYPE_VECTOR_IMM,
};
struct HeaderStruct {
Type MemType;
};
union {
HeaderStruct Header;
struct {
HeaderStruct Header;
Register rm;
} ScalarScalarType;
struct {
HeaderStruct Header;
int32_t Imm;
} ScalarImmType;
struct {
HeaderStruct Header;
ZRegister zm;
// TODO: Implement support for modifier
} ScalarVectorType;
struct {
HeaderStruct Header;
// rn will be a ZRegister
int32_t Imm;
} VectorImmType;
} MetaType;
};
/* This `ExtendedMemOperand` class is used for the helper load-store instructions.
* Load-store instructions are quite expressive, so having a helper that handles these differences is worth it.
*/
class ExtendedMemOperand final {
public:
ExtendedMemOperand(XRegister rn, XRegister rm = XReg::zr, ExtendedType Option = ExtendedType::LSL_64, uint32_t Shift = 0)
: rn {rn}
, MetaType {
.ExtendedType {
.Header = { .MemType = TYPE_EXTENDED },
.rm = rm,
.Option = Option,
.Shift = Shift,
}
} {}
ExtendedMemOperand(XRegister rn, IndexType Index = IndexType::OFFSET, int32_t Imm = 0)
: rn {rn}
, MetaType {
.ImmType {
.Header = { .MemType = TYPE_IMM },
.Index = Index,
.Imm = Imm,
}
} {}
Register rn;
enum Type {
TYPE_EXTENDED,
TYPE_IMM,
};
struct HeaderStruct {
Type MemType;
};
union {
HeaderStruct Header;
struct {
HeaderStruct Header;
Register rm;
ExtendedType Option;
uint32_t Shift;
} ExtendedType;
struct {
HeaderStruct Header;
IndexType Index;
int32_t Imm;
} ImmType;
} MetaType;
};
template<uint32_t op0, uint32_t op1, uint32_t CRn, uint32_t CRm, uint32_t op2>
constexpr uint32_t GenSystemReg() {
return op0 << 19 |
op1 << 16 |
CRn << 12 |
CRm << 8 |
op2 << 5;
};
// This `SystemRegister` enum is used for the mrs/msr instructions.
enum class SystemRegister : uint32_t {
CTR_EL0 = GenSystemReg<0b11, 0b011, 0b0000, 0b0000, 0b001>(),
DCZID_EL0 = GenSystemReg<0b11, 0b011, 0b0000, 0b0000, 0b111>(),
TPIDR_EL0 = GenSystemReg<0b11, 0b011, 0b1101, 0b0000, 0b010>(),
RNDR = GenSystemReg<0b11, 0b011, 0b0010, 0b0100, 0b000>(),
RNDRRS = GenSystemReg<0b11, 0b011, 0b0010, 0b0100, 0b001>(),
NZCV = GenSystemReg<0b11, 0b011, 0b0100, 0b0010, 0b000>(),
FPCR = GenSystemReg<0b11, 0b011, 0b0100, 0b0100, 0b000>(),
CNTFRQ_EL0 = GenSystemReg<0b11, 0b011, 0b1110, 0b0000, 0b000>(),
CNTVCT_EL0 = GenSystemReg<0b11, 0b011, 0b1110, 0b0000, 0b010>(),
};
template<uint32_t op1, uint32_t CRm, uint32_t op2>
constexpr uint32_t GenDCReg() {
return op1 << 16 |
CRm << 8 |
op2 << 5;
};
// This `DataCacheOperation` enum is used for the dc instruction.
enum class DataCacheOperation : uint32_t {
IVAC = GenDCReg<0b000, 0b0110, 0b001>(),
ISW = GenDCReg<0b000, 0b0110, 0b010>(),
CSW = GenDCReg<0b000, 0b1010, 0b010>(),
CISW = GenDCReg<0b000, 0b1110, 0b010>(),
ZVA = GenDCReg<0b011, 0b0100, 0b001>(),
CVAC = GenDCReg<0b011, 0b1010, 0b001>(),
CVAU = GenDCReg<0b011, 0b1011, 0b001>(),
CIVAC = GenDCReg<0b011, 0b1110, 0b001>(),
// MTE2
IGVAC = GenDCReg<0b000, 0b0110, 0b011>(),
IGSW = GenDCReg<0b000, 0b0110, 0b100>(),
IGDVAC = GenDCReg<0b000, 0b0110, 0b101>(),
IGDSW = GenDCReg<0b000, 0b0110, 0b110>(),
CGSW = GenDCReg<0b000, 0b1010, 0b100>(),
CGDSW = GenDCReg<0b000, 0b1010, 0b110>(),
CIGSW = GenDCReg<0b000, 0b1110, 0b100>(),
CIGDSW = GenDCReg<0b000, 0b1110, 0b110>(),
// MTE
GVA = GenDCReg<0b011, 0b0100, 0b011>(),
GZVA = GenDCReg<0b011, 0b0100, 0b100>(),
CGVAC = GenDCReg<0b011, 0b1010, 0b011>(),
CGDVAC = GenDCReg<0b011, 0b1010, 0b101>(),
CGVAP = GenDCReg<0b011, 0b1100, 0b011>(),
CGDVAP = GenDCReg<0b011, 0b1100, 0b101>(),
CGVADP = GenDCReg<0b011, 0b1101, 0b011>(),
CGDVADP = GenDCReg<0b011, 0b1101, 0b101>(),
CIGVAC = GenDCReg<0b011, 0b1110, 0b011>(),
CIGDVAC = GenDCReg<0b011, 0b1110, 0b101>(),
// DPB
CVAP = GenDCReg<0b011, 0b1100, 0b001>(),
// DPB2
CVADP = GenDCReg<0b011, 0b1101, 0b001>(),
};
template<uint32_t CRm, uint32_t op2>
constexpr uint32_t GenHintBarrierReg() {
return CRm << 8 |
op2 << 5;
}
// This `HintRegister` enum is used for the hint instruction.
enum class HintRegister : uint32_t {
NOP = GenHintBarrierReg<0b0000, 0b000>(),
YIELD = GenHintBarrierReg<0b0000, 0b001>(),
WFE = GenHintBarrierReg<0b0000, 0b010>(),
WFI = GenHintBarrierReg<0b0000, 0b011>(),
SEV = GenHintBarrierReg<0b0000, 0b100>(),
SEVL = GenHintBarrierReg<0b0000, 0b101>(),
DGH = GenHintBarrierReg<0b0000, 0b110>(),
CSDB = GenHintBarrierReg<0b0010, 0b100>(),
};
// This `BarrierRegister` enum is used for the various barrier instructions.
enum class BarrierRegister : uint32_t {
CLREX = GenHintBarrierReg<0b0000, 0b010>(),
TCOMMIT = GenHintBarrierReg<0b0000, 0b011>(),
DSB = GenHintBarrierReg<0b0000, 0b100>(),
DMB = GenHintBarrierReg<0b0000, 0b101>(),
ISB = GenHintBarrierReg<0b0000, 0b110>(),
SB = GenHintBarrierReg<0b0000, 0b111>(),
};
// This `BarrierScope` enum is used for the dsb/dmb instructions.
enum class BarrierScope : uint32_t {
// Outer shareable
OSHLD = 0b0001,
OSHST = 0b0010,
OSH = 0b0011,
// Non shareable
NSHLD = 0b0101,
NSHST = 0b0110,
NSH = 0b0111,
// Inner shareable
ISHLD = 0b1001,
ISHST = 0b1010,
ISH = 0b1011,
// Full System visibility
LD = 0b1101,
ST = 0b1110,
SY = 0b1111,
};
// This `Prefetch` enum is used for prefetch instructions.
enum class Prefetch : uint32_t {
// Prefetch for load
PLDL1KEEP = 0b00000,
PLDL1STRM = 0b00001,
PLDL2KEEP = 0b00010,
PLDL2STRM = 0b00011,
PLDL3KEEP = 0b00100,
PLDL3STRM = 0b00101,
// Preload instructions
PLIL1KEEP = 0b01000,
PLIL1STRM = 0b01001,
PLIL2KEEP = 0b01010,
PLIL2STRM = 0b01011,
PLIL3KEEP = 0b01100,
PLIL3STRM = 0b01101,
// Preload for store
PSTL1KEEP = 0b10000,
PSTL1STRM = 0b10001,
PSTL2KEEP = 0b10010,
PSTL2STRM = 0b10011,
PSTL3KEEP = 0b10100,
PSTL3STRM = 0b10101,
};
// This `PredicatePattern` enun is used for some SVE instructions.
enum class PredicatePattern : uint32_t {
SVE_POW2 = 0b00000,
SVE_VL1 = 0b00001,
SVE_VL2 = 0b00010,
SVE_VL3 = 0b00011,
SVE_VL4 = 0b00100,
SVE_VL5 = 0b00101,
SVE_VL6 = 0b00110,
SVE_VL7 = 0b00111,
SVE_VL8 = 0b01000,
SVE_VL16 = 0b01001,
SVE_VL32 = 0b01010,
SVE_VL64 = 0b01011,
SVE_VL128 = 0b01100,
SVE_VL256 = 0b01101,
SVE_MUL4 = 0b11101,
SVE_MUL3 = 0b11110,
SVE_ALL = 0b11111,
};
/* This `BackwardLabel` struct used for retaining a location for PC-Relative instructions.
* This is specifically a label for a target that is logically `below` an instruction that uses it.
* Which means that a branch would jump backwards.
*/
struct BackwardLabel {
uint8_t *Location{};
};
/* This `ForwardLabel` struct used for retaining a location for PC-Relative instructions.
* This is specifically a label for a target that is logically `above` an instruction that uses it.
* Which means that a branch would jump forwards.
*
* This can be bound to multiple instructions, so it needs a vector for each bind instruction type.
*/
struct ForwardLabel {
struct Instructions {
enum class InstType {
ADR,
ADRP,
B,
BC,
TEST_BRANCH,
RELATIVE_LOAD,
LONG_ADDRESS_GEN,
};
uint8_t *Location{};
InstType Type;
};
std::vector<Instructions> Insts{};
};
/* This `BiDirectionalLabel` struct used for retaining a location for PC-Relative instructions.
* This is specifically a label for a target that is in either direction of an instruction that uses it.
* Which means a branch could jump backwards or forwards depending on situation.
*/
struct BiDirectionalLabel {
BackwardLabel Backward;
ForwardLabel Forward;
};
// Some FCMA ASIMD instructions support a rotation argument.
enum class Rotation : uint32_t {
ROTATE_0 = 0b00,
ROTATE_90 = 0b01,
ROTATE_180 = 0b10,
ROTATE_270 = 0b11,
};
// Concept for contraining some instructions to accept only an XRegister or WRegister.
// Particularly for operations that differ encodings depending on which one is used.
template <typename T>
concept IsXOrWRegister = std::is_same_v<T, XRegister> || std::is_same_v<T, WRegister>;
// Whether or not a given set of vector registers are sequential
// in increasing order as far as the register file is concerned (modulo its size)
//
// For example, a set of registers like:
//
// v1, v2, v3 and
// v31, v0, v1
//
// would both be considered sequential sequences, and some instructions in particular
// limit register lists to these kind of sequences.
//
template <typename T, typename... Args>
constexpr bool AreVectorsSequential(T first, const Args&... args) {
// Ensure we always have a pair of registers to compare against.
static_assert(sizeof...(args) >= 1, "Number of arguments must be greater than 1");
const auto fn = [](auto& lhs, const auto& rhs) {
const auto result = ((lhs.Idx() + 1) % 32) == rhs.Idx();
lhs = rhs;
return result;
};
return (fn(first, args) && ...);
}
// This is an emitter that is designed around the smallest code bloat as possible.
// Eschewing most developer convenience in order to keep code as small as possible.
// Choices:
// - Size of ops passed as an argument rather than template to let the compiler use csel instead of branching.
// - Registers are unsized so they can be passed in a GPR and not need conversion operations
class Emitter : public FEXCore::ARMEmitter::Buffer {
public:
Emitter() = default;
Emitter(uint8_t* Base, uint64_t BaseSize)
: Buffer (Base, BaseSize) {
}
// Bind a backward label to an address.
// Address that is bound is the current emitter location.
void Bind(BackwardLabel *Label) {
LOGMAN_THROW_AA_FMT(Label->Location == nullptr, "Trying to bind a label twice");
Label->Location = GetCursorAddress<uint8_t*>();
}
// Bind a forward label to a location.
// This walks all the instructions in the label's vector.
// Then backpatching all instructions that have used the label.
template<bool WarnAboutEmpty = false>
void Bind(ForwardLabel *Label) {
if constexpr (WarnAboutEmpty) {
LOGMAN_THROW_A_FMT(Label->Insts.empty() == false, "Binding forward label that didn't have any instructions using it");
}
uint8_t *CurrentAddress = GetCursorAddress<uint8_t*>();
for (const auto &Inst : Label->Insts) {
// Patch up the instructions
switch (Inst.Type) {
case ForwardLabel::Instructions::InstType::ADR: {
uint32_t *Instruction = reinterpret_cast<uint32_t*>(Inst.Location);
int64_t Imm = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(Instruction);
LOGMAN_THROW_A_FMT(IsADRRange(Imm), "Unscaled offset too large");
uint32_t InstMask = 0b11 << 29 | 0b1111'1111'1111'1111'111 << 5;
uint32_t Offset = static_cast<uint32_t>(Imm) & 0x3F'FFFF;
uint32_t Inst = *Instruction & ~InstMask;
Inst |= (Offset & 0b11) << 29;
Inst |= (Offset >> 2) << 5;
*Instruction = Inst;
break;
}
case ForwardLabel::Instructions::InstType::ADRP: {
uint32_t *Instruction = reinterpret_cast<uint32_t*>(Inst.Location);
int64_t Imm = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(Instruction);
LOGMAN_THROW_A_FMT(IsADRPRange(Imm) && IsADRPAligned(Imm), "Unscaled offset too large");
Imm >>= 12;
uint32_t InstMask = 0b11 << 29 | 0b1111'1111'1111'1111'111 << 5;
uint32_t Offset = static_cast<uint32_t>(Imm) & 0x3F'FFFF;
uint32_t Inst = *Instruction & ~InstMask;
Inst |= (Offset & 0b11) << 29;
Inst |= (Offset >> 2) << 5;
*Instruction = Inst;
break;
}
case ForwardLabel::Instructions::InstType::B: {
uint32_t *Instruction = reinterpret_cast<uint32_t*>(Inst.Location);
int64_t Imm = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(Instruction);
LOGMAN_THROW_A_FMT(Imm >= -134217728 && Imm <= 134217724 && ((Imm & 0b11) == 0), "Unscaled offset too large");
Imm >>= 2;
uint32_t InstMask = 0x3FF'FFFF;
uint32_t Offset = static_cast<uint32_t>(Imm) & InstMask;
uint32_t Inst = *Instruction & ~InstMask;
Inst |= Offset;
*Instruction = Inst;
break;
}
case ForwardLabel::Instructions::InstType::TEST_BRANCH: {
uint32_t *Instruction = reinterpret_cast<uint32_t*>(Inst.Location);
int64_t Imm = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(Instruction);
LOGMAN_THROW_A_FMT(Imm >= -32768 && Imm <= 32764 && ((Imm & 0b11) == 0), "Unscaled offset too large");
Imm >>= 2;
uint32_t InstMask = 0x3FFF;
uint32_t Offset = static_cast<uint32_t>(Imm) & InstMask;
uint32_t Inst = *Instruction & ~(InstMask << 5);
Inst |= Offset << 5;
*Instruction = Inst;
break;
}
case ForwardLabel::Instructions::InstType::BC:
case ForwardLabel::Instructions::InstType::RELATIVE_LOAD: {
uint32_t *Instruction = reinterpret_cast<uint32_t*>(Inst.Location);
int64_t Imm = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(Instruction);
LOGMAN_THROW_A_FMT(Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0), "Unscaled offset too large");
Imm >>= 2;
uint32_t InstMask = 0x7'FFFF;
uint32_t Offset = static_cast<uint32_t>(Imm) & InstMask;
uint32_t Inst = *Instruction & ~(InstMask << 5);
Inst |= Offset << 5;
*Instruction = Inst;
break;
}
case ForwardLabel::Instructions::InstType::LONG_ADDRESS_GEN: {
uint32_t *Instructions = reinterpret_cast<uint32_t*>(Inst.Location);
int64_t ImmInstOne = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(&Instructions[0]);
int64_t ImmInstTwo = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(&Instructions[1]);
auto OriginalOffset = GetCursorOffset();
auto InstOffset = GetCursorOffsetFromAddress(Instructions);
SetCursorOffset(InstOffset);
// We encoded the destination register in to the first instruction space.
// Read it back.
ARMEmitter::Register DestReg(Instructions[0]);
if (IsADRRange(ImmInstTwo)) {
// If within ADR range from the second instruction, then we can emit NOP+ADR
nop();
adr(DestReg, static_cast<uint32_t>(ImmInstTwo) & 0x7FFF);
}
else if (IsADRPRange(ImmInstOne)) {
// If within ADRP range from the first instruction, then we are /definitely/ in range for the second instruction.
// First check if we are in non-offset range for second instruction.
if (IsADRPAligned(reinterpret_cast<uint64_t>(CurrentAddress))) {
// We can emit nop + adrp
nop();
adrp(DestReg, static_cast<uint32_t>(ImmInstTwo >> 12) & 0x7FFF);
}
else {
// Not aligned, need adrp + add
adrp(DestReg, static_cast<uint32_t>(ImmInstOne >> 12) & 0x7FFF);
add(ARMEmitter::Size::i64Bit, DestReg, DestReg, ImmInstOne & 0xFFF);
}
}
else {
LOGMAN_MSG_A_FMT("Unscaled offset is too large");
FEX_UNREACHABLE;
}
SetCursorOffset(OriginalOffset);
break;
}
default: LOGMAN_MSG_A_FMT("Unexpected inst type in label fixup");
}
}
}
// Bind a bidirectional location to a location.
// Binds both forwards and backwards depending on how the label was used.
void Bind(BiDirectionalLabel *Label) {
if (!Label->Backward.Location) {
Bind(&Label->Backward);
}
Bind<false>(&Label->Forward);
}
public:
// TODO: Implement SME when it matters.
#include "Interface/Core/ArchHelpers/CodeEmitter/ALUOps.inl"
#include "Interface/Core/ArchHelpers/CodeEmitter/BranchOps.inl"
#include "Interface/Core/ArchHelpers/CodeEmitter/LoadstoreOps.inl"
#include "Interface/Core/ArchHelpers/CodeEmitter/SystemOps.inl"
#include "Interface/Core/ArchHelpers/CodeEmitter/ScalarOps.inl"
#include "Interface/Core/ArchHelpers/CodeEmitter/ASIMDOps.inl"
#include "Interface/Core/ArchHelpers/CodeEmitter/SVEOps.inl"
private:
template<typename T>
uint32_t Encode_ra(T Reg) const {
return Reg.Idx() << 10;
}
uint32_t Encode_ra(uint32_t Reg) const {
return Reg << 10;
}
template<typename T>
uint32_t Encode_rt2(T Reg) const {
return Reg.Idx() << 10;
}
template<>
uint32_t Encode_rt2(uint32_t Reg) const {
return Reg << 10;
}
template<typename T>
uint32_t Encode_rm(T Reg) const {
return Reg.Idx() << 16;
}
uint32_t Encode_rm(uint32_t Reg) const {
return Reg << 16;
}
template<typename T>
uint32_t Encode_rs(T Reg) const {
return Reg.Idx() << 16;
}
uint32_t Encode_rs(uint32_t Reg) const {
return Reg << 16;
}
template<typename T>
uint32_t Encode_rn(T Reg) const {
return Reg.Idx() << 5;
}
uint32_t Encode_rn(uint32_t Reg) const {
return Reg << 5;
}
template<typename T>
uint32_t Encode_rd(T Reg) const {
return Reg.Idx();
}
uint32_t Encode_rd(uint32_t Reg) const {
return Reg;
}
template<typename T>
uint32_t Encode_rt(T Reg) const {
return Reg.Idx();
}
template<>
uint32_t Encode_rt(Prefetch Reg) const {
return FEXCore::ToUnderlying(Reg);
}
uint32_t Encode_rt(uint32_t Reg) const {
return Reg;
}
template<typename T>
uint32_t Encode_pd(T Reg) const {
return FEXCore::ToUnderlying(Reg);
}
};
}
File diff suppressed because it is too large. Load diff
File diff suppressed because it is too large. Load diff
File diff suppressed because it is too large. Load diff
File diff suppressed because it is too large. Load diff
@@ -0,0 +1,175 @@
/* System instruction emitters.
*
* This is mostly a mashup of various instruction types.
* Nothing follows an explicit pattern since they are mostly different.
*/
public:
// System with result
// TODO: SYSL
// System Instruction
// TODO: AT
// TODO: CFP
// TODO: CPP
void dc(FEXCore::ARMEmitter::DataCacheOperation DCOp, FEXCore::ARMEmitter::Register rt) {
constexpr uint32_t Op = 0b1101'0101'0000'1000'0111 << 12;
SystemInstruction(Op, 0, FEXCore::ToUnderlying(DCOp), rt);
}
// TODO: DVP
// TODO: IC
// TODO: TLBI
// Exception generation
void svc(uint32_t Imm) {
ExceptionGeneration(0b000, 0b000, 0b01, Imm);
}
void hvc(uint32_t Imm) {
ExceptionGeneration(0b000, 0b000, 0b10, Imm);
}
void smc(uint32_t Imm) {
ExceptionGeneration(0b000, 0b000, 0b11, Imm);
}
void brk(uint32_t Imm) {
ExceptionGeneration(0b001, 0b000, 0b00, Imm);
}
void hlt(uint32_t Imm) {
ExceptionGeneration(0b010, 0b000, 0b00, Imm);
}
void tcancel(uint32_t Imm) {
ExceptionGeneration(0b011, 0b000, 0b00, Imm);
}
void dcps1(uint32_t Imm) {
ExceptionGeneration(0b101, 0b000, 0b01, Imm);
}
void dcps2(uint32_t Imm) {
ExceptionGeneration(0b101, 0b000, 0b10, Imm);
}
void dcps3(uint32_t Imm) {
ExceptionGeneration(0b101, 0b000, 0b11, Imm);
}
// System instructions with register argument
void wfet(FEXCore::ARMEmitter::Register rt) {
SystemInstructionWithReg(0b0000, 0b000, rt);
}
void wfit(FEXCore::ARMEmitter::Register rt) {
SystemInstructionWithReg(0b0000, 0b001, rt);
}
// Hints
void nop() {
Hint(FEXCore::ARMEmitter::HintRegister::NOP);
}
void yield() {
Hint(FEXCore::ARMEmitter::HintRegister::YIELD);
}
void wfe() {
Hint(FEXCore::ARMEmitter::HintRegister::WFE);
}
void wfi() {
Hint(FEXCore::ARMEmitter::HintRegister::WFI);
}
void sev() {
Hint(FEXCore::ARMEmitter::HintRegister::SEV);
}
void sevl() {
Hint(FEXCore::ARMEmitter::HintRegister::SEVL);
}
void dgh() {
Hint(FEXCore::ARMEmitter::HintRegister::DGH);
}
void csdb() {
Hint(FEXCore::ARMEmitter::HintRegister::CSDB);
}
// Barriers
void clrex(uint32_t imm = 15) {
LOGMAN_THROW_AA_FMT(imm < 16, "Immediate out of range");
Barrier(FEXCore::ARMEmitter::BarrierRegister::CLREX, imm);
}
void dsb(FEXCore::ARMEmitter::BarrierScope Scope) {
Barrier(FEXCore::ARMEmitter::BarrierRegister::DSB, FEXCore::ToUnderlying(Scope));
}
void dmb(FEXCore::ARMEmitter::BarrierScope Scope) {
Barrier(FEXCore::ARMEmitter::BarrierRegister::DMB, FEXCore::ToUnderlying(Scope));
}
void isb() {
Barrier(FEXCore::ARMEmitter::BarrierRegister::ISB, FEXCore::ToUnderlying(FEXCore::ARMEmitter::BarrierScope::SY));
}
void sb() {
Barrier(FEXCore::ARMEmitter::BarrierRegister::SB, 0);
}
void tcommit() {
Barrier(FEXCore::ARMEmitter::BarrierRegister::TCOMMIT, 0);
}
// System register move
void msr(FEXCore::ARMEmitter::SystemRegister reg, FEXCore::ARMEmitter::Register rt) {
constexpr uint32_t Op = 0b1101'0101'0001 << 20;
SystemRegisterMove(Op, rt, reg);
}
void mrs(FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::SystemRegister reg) {
constexpr uint32_t Op = 0b1101'0101'0011 << 20;
SystemRegisterMove(Op, rd, reg);
}
private:
// Exception Generation
void ExceptionGeneration(uint32_t opc, uint32_t op2, uint32_t LL, uint32_t Imm) {
LOGMAN_THROW_AA_FMT((Imm & 0xFFFF'0000) == 0, "Imm amount too large");
uint32_t Instr = 0b1101'0100 << 24;
Instr |= opc << 21;
Instr |= Imm << 5;
Instr |= op2 << 2;
Instr |= LL;
dc32(Instr);
}
// System instructions with register argument
void SystemInstructionWithReg(uint32_t CRm, uint32_t op2, FEXCore::ARMEmitter::Register rt) {
uint32_t Instr = 0b1101'0101'0000'0011'0001 << 12;
Instr |= CRm << 8;
Instr |= op2 << 5;
Instr |= Encode_rt(rt);
dc32(Instr);
}
// Hints
void Hint(FEXCore::ARMEmitter::HintRegister Reg) {
uint32_t Instr = 0b1101'0101'0000'0011'0010'0000'0001'1111U;
Instr |= FEXCore::ToUnderlying(Reg);
dc32(Instr);
}
// Barriers
void Barrier(FEXCore::ARMEmitter::BarrierRegister Reg, uint32_t CRm) {
uint32_t Instr = 0b1101'0101'0000'0011'0011'0000'0001'1111U;
Instr |= CRm << 8;
Instr |= FEXCore::ToUnderlying(Reg);
dc32(Instr);
}
// System Instruction
void SystemInstruction(uint32_t Op, uint32_t L, uint32_t SubOp, FEXCore::ARMEmitter::Register rt) {
uint32_t Instr = Op;
Instr |= L << 21;
Instr |= SubOp;
Instr |= Encode_rt(rt);
dc32(Instr);
}
// System register move
void SystemRegisterMove(uint32_t Op, FEXCore::ARMEmitter::Register rt, FEXCore::ARMEmitter::SystemRegister reg) {
uint32_t Instr = Op;
Instr |= FEXCore::ToUnderlying(reg);
Instr |= Encode_rt(rt);
dc32(Instr);
}
@@ -3,6 +3,7 @@
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Core/UContext.h>
#include <FEXCore/Core/X86Enums.h>
#include <signal.h>
#include <string.h>
@@ -10,16 +11,25 @@
#include <stdint.h>
#include <type_traits>
namespace FEXCore::ArchHelpers::Context {
enum ContextFlags : uint32_t {
CONTEXT_FLAG_INJIT = (1U << 0),
CONTEXT_FLAG_32BIT = (1U << 1),
};
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
constexpr uint64_t STACK_COOKIE_MAGIC = 0x4142434445464748ULL;
#endif
struct X86ContextBackup {
// Host State
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
// During debug builds, insert a cookie on the stack.
// This is useful for validation that the stack is trying to be restored from the correct location.
// During stack restore, we ensure this is set to the value we expect.
// If given an incorrect stack location, or corrupted stack then this cookie will be wrong.
uint64_t StackCookie;
#endif
// RIP and RSP is stored in GPRs here
uint64_t GPRs[23];
FEXCore::x86_64::_libc_fpstate FPRState;
@@ -39,6 +49,9 @@ struct X86ContextBackup {
struct ArmContextBackup {
// Host State
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
uint64_t StackCookie;
#endif
uint64_t GPRs[31];
uint64_t PrevSP;
uint64_t PrevPC;
@@ -76,6 +89,7 @@ static inline mcontext_t* GetMContext(void* ucontext) {
#ifdef _M_ARM_64
constexpr uint32_t FPR_MAGIC = 0x46508001U;
constexpr uint32_t ESR1_MAGIC = 0x45535201U;
struct HostCTXHeader {
uint32_t Magic;
@@ -89,6 +103,11 @@ struct HostFPRState {
__uint128_t FPRs[32];
};
struct HostESRState {
HostCTXHeader Head;
uint64_t ESR;
};
static inline uint64_t GetSp(void* ucontext) {
return GetMContext(ucontext)->sp;
}
@@ -129,6 +148,61 @@ static inline __uint128_t GetArmFPR(void* ucontext, uint32_t id) {
return HostState->FPRs[id];
}
static inline uint64_t GetArmESR(void* ucontext) {
auto MContext = GetMContext(ucontext);
size_t i = 0;
auto HostState = reinterpret_cast<HostCTXHeader*>(&MContext->__reserved[i]);
do {
if (HostState->Magic == ESR1_MAGIC) {
auto ESR = reinterpret_cast<HostESRState*>(HostState);
return ESR->ESR;
}
i += HostState->Size;
HostState = reinterpret_cast<HostCTXHeader*>(&MContext->__reserved[i]);
} while (HostState->Size != 0);
return 0;
}
constexpr static uint64_t ESR1_EC = 0b111111U << 26;
constexpr static uint64_t ESR1_EC_DataAbort = 0b100100U << 26;
// Write-Not-Read flag
// When set - Abort is due to a write
constexpr static uint64_t ESR1_WNR = 1 << 6;
// DFSC - Default Status Code
// Translation fault - No page mapped
// Permissions fault - Page mapped but with incorrect permission from access.
constexpr static uint64_t ESR1_DataAbort_DFSC = 0b111111;
constexpr static uint64_t ESR1_DataAbort_TranslationFault_EL0 = 0b000111;
constexpr static uint64_t ESR1_DataAbort_PermissionFault_EL0 = 0b001111;
constexpr static uint64_t ESR1_DataAbort_Level = 0b11;
constexpr static uint64_t ESR1_DataAbort_Level_EL3 = 0b00;
constexpr static uint64_t ESR1_DataAbort_Level_EL2 = 0b01;
constexpr static uint64_t ESR1_DataAbort_Level_EL1 = 0b10;
constexpr static uint64_t ESR1_DataAbort_Level_EL0 = 0b11;
static inline uint32_t GetProtectFlags(void* ucontext) {
uint64_t ESR = GetArmESR(ucontext);
LOGMAN_THROW_A_FMT((ESR & ESR1_EC) == ESR1_EC_DataAbort, "Unknown ESR1 EC type: 0x{:x} != 0x{:x}", ESR & ESR1_EC, ESR1_EC_DataAbort);
uint32_t ProtectFlags{};
if ((ESR & ESR1_DataAbort_Level) == ESR1_DataAbort_Level_EL0) {
// Always a user error for us.
ProtectFlags |= X86State::X86_PF_USER;
}
if (ESR & ESR1_WNR) {
// Fault was due to a write
ProtectFlags |= X86State::X86_PF_WRITE;
}
// PF_PROT is not returned to user on x86, so don't return the difference between permission fault and translation fault.
return ProtectFlags;
}
using ContextBackup = ArmContextBackup;
template <typename T>
static inline void BackupContext(void* ucontext, T *Backup) {
@@ -150,6 +224,10 @@ static inline void BackupContext(void* ucontext, T *Backup) {
// Save the signal mask so we can restore it
memcpy(&Backup->sa_mask, &_ucontext->uc_sigmask, sizeof(uint64_t));
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
Backup->StackCookie = STACK_COOKIE_MAGIC;
#endif
} else {
// This must be a runtime error
ERROR_AND_DIE_FMT("Wrong context type");
@@ -159,8 +237,10 @@ static inline void BackupContext(void* ucontext, T *Backup) {
template <typename T>
static inline void RestoreContext(void* ucontext, T *Backup) {
if constexpr (std::is_same<T, ArmContextBackup>::value) {
LOGMAN_THROW_A_FMT(Backup->StackCookie == STACK_COOKIE_MAGIC, "Stack cookie didn't match! 0x{:x}", Backup->StackCookie);
auto _ucontext = GetUContext(ucontext);
auto _mcontext = GetMContext(ucontext);
auto _mcontext = GetMContext(ucontext);
HostFPRState *HostState = reinterpret_cast<HostFPRState*>(&_mcontext->__reserved[0]);
LOGMAN_THROW_AA_FMT(HostState->Head.Magic == FPR_MAGIC, "Wrong FPR Magic: 0x{:08x}", HostState->Head.Magic);
@@ -222,6 +302,10 @@ static inline __uint128_t GetArmFPR(void* ucontext, uint32_t id) {
ERROR_AND_DIE_FMT("Not implemented for x86 host");
}
static inline uint32_t GetProtectFlags(void* ucontext) {
return GetMContext(ucontext)->gregs[REG_ERR];
}
using ContextBackup = X86ContextBackup;
template <typename T>
static inline void BackupContext(void* ucontext, T *Backup) {
@@ -237,6 +321,10 @@ static inline void BackupContext(void* ucontext, T *Backup) {
// Save the signal mask so we can restore it
memcpy(&Backup->sa_mask, &_ucontext->uc_sigmask, sizeof(uint64_t));
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
Backup->StackCookie = STACK_COOKIE_MAGIC;
#endif
} else {
// This must be a runtime error
ERROR_AND_DIE_FMT("Wrong context type");
@@ -246,6 +334,8 @@ static inline void BackupContext(void* ucontext, T *Backup) {
template <typename T>
static inline void RestoreContext(void* ucontext, T *Backup) {
if constexpr (std::is_same<T, X86ContextBackup>::value) {
LOGMAN_THROW_A_FMT(Backup->StackCookie == STACK_COOKIE_MAGIC, "Stack cookie didn't match! 0x{:x}", Backup->StackCookie);
auto _ucontext = GetUContext(ucontext);
auto _mcontext = GetMContext(ucontext);
+2 -2
View File
@@ -60,8 +60,8 @@ auto CPUBackend::AllocateNewCodeBuffer(size_t Size) -> CodeBuffer {
FEXCore::Allocator::mmap(nullptr, Buffer.Size, PROT_READ | PROT_WRITE | PROT_EXEC, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0));
LOGMAN_THROW_AA_FMT(!!Buffer.Ptr, "Couldn't allocate code buffer");
if (ThreadState->CTX->Config.GlobalJITNaming()) {
ThreadState->CTX->Symbols.RegisterJITSpace(Buffer.Ptr, Buffer.Size);
if (static_cast<Context::ContextImpl*>(ThreadState->CTX)->Config.GlobalJITNaming()) {
static_cast<Context::ContextImpl*>(ThreadState->CTX)->Symbols.RegisterJITSpace(Buffer.Ptr, Buffer.Size);
}
return Buffer;
}
+23 -100
View File
@@ -88,7 +88,6 @@ static uint32_t CalculateNumberOfCPUs() {
// when AVX implementations are further along.
constexpr uint32_t SUPPORTS_AVX = 0;
// #define CPUID_AMD
#ifdef CPUID_AMD
constexpr uint32_t FAMILY_IDENTIFIER =
0 | // Stepping
@@ -122,25 +121,24 @@ void CPUIDEmu::SetupHostHybridFlag() {
uint64_t MIDR{};
for (size_t i = 0; i < CPUs; ++i) {
std::error_code ec{};
std::string MIDRPath = "/sys/devices/system/cpu/cpu" + std::to_string(i) + "/regs/identification/midr_el1";
if (std::filesystem::exists(MIDRPath, ec)) {
std::vector<char> Data{};
// Needs to be a fixed size since depending on kernel it will try to read a full page of data and fail
// Only read 18 bytes for a 64bit value prefixed with 0x
if (FEXCore::FileLoading::LoadFile(Data, MIDRPath, 18)) {
uint64_t NewMIDR{};
std::string_view MIDRView(&Data.at(0), 18);
if (FEXCore::StrConv::Conv(MIDRView, &NewMIDR)) {
if (MIDR != 0 && MIDR != NewMIDR) {
// CPU mismatch, claim hybrid
Hybrid = true;
}
std::string MIDRPath = fmt::format("/sys/devices/system/cpu/cpu{}/regs/identification/midr_el1", i);
// Truncate to 32-bits, top 32-bits are all reserved in MIDR
PerCPUData[i].ProductName = ProductNames::ARM_UNKNOWN;
PerCPUData[i].MIDR = NewMIDR;
MIDR = NewMIDR;
std::array<char, 18> Data;
// Needs to be a fixed size since depending on kernel it will try to read a full page of data and fail
// Only read 18 bytes for a 64bit value prefixed with 0x
if (FEXCore::FileLoading::LoadFileToBuffer(MIDRPath, Data) == sizeof(Data)) {
uint64_t NewMIDR{};
std::string_view MIDRView(Data.data(), sizeof(Data));
if (FEXCore::StrConv::Conv(MIDRView, &NewMIDR)) {
if (MIDR != 0 && MIDR != NewMIDR) {
// CPU mismatch, claim hybrid
Hybrid = true;
}
// Truncate to 32-bits, top 32-bits are all reserved in MIDR
PerCPUData[i].ProductName = ProductNames::ARM_UNKNOWN;
PerCPUData[i].MIDR = NewMIDR;
MIDR = NewMIDR;
}
}
}
@@ -412,6 +410,9 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_01h(uint32_t Leaf) {
// XXX: Enable once the rest of the SSE4.2 instructions are emulated
uint32_t SupportsSSE42 = CTX->HostFeatures.SupportsCRC && false ? 1 : 0;
// Hypervisor bit is normally set but some applications have issues with it.
uint32_t Hypervisor = HideHypervisorBit() ? 0 : 1;
Res.eax = FAMILY_IDENTIFIER;
Res.ebx = 0 | // Brand index
@@ -451,7 +452,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_01h(uint32_t Leaf) {
(SUPPORTS_AVX << 28) | // AVX
(0 << 29) | // F16C
(CTX->HostFeatures.SupportsRAND << 30) | // RDRAND
(1 << 31); // Hypervisor always returns one
(Hypervisor << 31);
Res.edx =
(1 << 0) | // FPU
@@ -657,8 +658,8 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_07h(uint32_t Leaf) {
(0 << 20) | // SMAP Supervisor mode access prevention and CLAC/STAC instructions
(0 << 21) | // Reserved
(0 << 22) | // Reserved
(0 << 23) | // CLFLUSHOPT instruction
(0 << 24) | // CLWB instruction
(1 << 23) | // CLFLUSHOPT instruction
(CTX->HostFeatures.SupportsCLWB << 24) | // CLWB instruction
(0 << 25) | // Intel processor trace
(0 << 26) | // Reserved
(0 << 27) | // Reserved
@@ -1212,87 +1213,9 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_Reserved(uint32_t Leaf) {
return Res;
}
void CPUIDEmu::Init(FEXCore::Context::Context *ctx) {
void CPUIDEmu::Init(FEXCore::Context::ContextImpl *ctx) {
CTX = ctx;
RegisterFunction(0, &CPUIDEmu::Function_0h);
RegisterFunction(1, &CPUIDEmu::Function_01h);
RegisterFunction(2, &CPUIDEmu::Function_02h);
// 3: Serial Number(previously), now reserved
#ifndef CPUID_AMD
// Deterministic cache parameters for each level
RegisterFunction(0x4, &CPUIDEmu::Function_04h);
#endif
// 5: Monitor/mwait
// Thermal and power management
RegisterFunction(6, &CPUIDEmu::Function_06h);
// Extended feature flags
RegisterFunction(7, &CPUIDEmu::Function_07h);
// 9: Direct Cache Access information
// 0x0A: Architectural performance monitoring
// 0x0B: Extended topology enumeration
// 0x0D: Processor extended state enumeration
RegisterFunction(0x0D, &CPUIDEmu::Function_0Dh);
// 0x0F: Intel RDT monitoring
// 0x10: Intel RDT allocation enumeration
// 0x12: Intel SGX capability enumeration
// 0x13: Reserved
// 0x14: Intel Processor trace
#ifndef CPUID_AMD
// Timestamp counter information
// Doesn't exist on AMD hardware
RegisterFunction(0x15, &CPUIDEmu::Function_15h);
#endif
// 0x16: Processor frequency information
// 0x17: SoC vendor attribute enumeration
// 0x1A: Hybrid Information Sub-leaf
#ifndef CPUID_AMD
RegisterFunction(0x1A, &CPUIDEmu::Function_1Ah);
#endif
// Hypervisor CPUID information leaf
RegisterFunction(0x4000'0000, &CPUIDEmu::Function_4000_0000h);
RegisterFunction(0x4000'0001, &CPUIDEmu::Function_4000_0001h);
// Largest extended function number
RegisterFunction(0x8000'0000, &CPUIDEmu::Function_8000_0000h);
// Processor vendor
RegisterFunction(0x8000'0001, &CPUIDEmu::Function_8000_0001h);
// Processor brand string
RegisterFunction(0x8000'0002, &CPUIDEmu::Function_8000_0002h);
// Processor brand string continued
RegisterFunction(0x8000'0003, &CPUIDEmu::Function_8000_0003h);
// Processor brand string continued
RegisterFunction(0x8000'0004, &CPUIDEmu::Function_8000_0004h);
// 0x8000'0005: L1 Cache and TLB identifiers
#ifdef CPUID_AMD
RegisterFunction(0x8000'0005, &CPUIDEmu::Function_8000_0005h);
#else
// This is full reserved on Intel platforms
RegisterFunction(0x8000'0005, &CPUIDEmu::Function_Reserved);
#endif
// 0x8000'0006: L2 Cache identifiers
RegisterFunction(0x8000'0006, &CPUIDEmu::Function_8000_0006h);
// Advanced power management information
RegisterFunction(0x8000'0007, &CPUIDEmu::Function_8000_0007h);
// Virtual and physical address sizes
RegisterFunction(0x8000'0008, &CPUIDEmu::Function_8000_0008h);
// 0x8000'000A: SVM Revision
// TLB 1GB page identifiers
RegisterFunction(0x8000'0019, &CPUIDEmu::Function_8000_0019h);
// 0x8000'001A: Performance optimization identifiers
// 0x8000'001B: Instruction based sampling identifiers
// 0x8000'001C: Lightweight profiling capabilities
// 0x8000'001D: Cache properties
#ifdef CPUID_AMD
// Deterministic cache parameters for each level
RegisterFunction(0x8000'001D, &CPUIDEmu::Function_8000_001Dh);
#endif
// 0x8000'001E: Extended APIC ID
// 0x8000'001F: AMD Secure Encryption
// Setup some state tracking
SetupHostHybridFlag();
}
+173 -14
View File
@@ -10,9 +10,12 @@
namespace FEXCore {
namespace Context {
struct Context;
class ContextImpl;
}
// Debugging define to switch what family of CPU we execute as.
// Might be useful if an application makes an assumption about a CPU.
// #define CPUID_AMD
class CPUIDEmu final {
private:
constexpr static uint32_t CPUID_VENDOR_INTEL1 = 0x756E6547; // "Genu"
@@ -28,16 +31,27 @@ public:
// if we report anything differently then applications are likely to break
constexpr static uint64_t CACHELINE_SIZE = 64;
void Init(FEXCore::Context::Context *ctx);
void Init(FEXCore::Context::ContextImpl *ctx);
FEXCore::CPUID::FunctionResults RunFunction(uint32_t Function, uint32_t Leaf) {
const auto Handler = FunctionHandlers.find(Function);
if (Handler == FunctionHandlers.end()) {
return Function_Reserved(Leaf);
if (Function < Primary.size()) {
const auto Handler = Primary[Function];
return (this->*Handler)(Leaf);
}
return (this->*Handler->second)(Leaf);
constexpr uint32_t HypervisorBase = 0x4000'0000;
if (Function >= HypervisorBase && Function < (HypervisorBase + Hypervisor.size())) {
const auto Handler = Hypervisor[Function - HypervisorBase];
return (this->*Handler)(Leaf);
}
constexpr uint32_t ExtendedBase = 0x8000'0000;
if (Function >= ExtendedBase && Function < (ExtendedBase + Extended.size())) {
const auto Handler = Extended[Function - ExtendedBase];
return (this->*Handler)(Leaf);
}
return Function_Reserved(Leaf);
}
FEXCore::CPUID::FunctionResults RunFunctionName(uint32_t Function, uint32_t Leaf, uint32_t CPU) {
@@ -50,16 +64,12 @@ public:
}
private:
FEXCore::Context::Context *CTX;
FEXCore::Context::ContextImpl *CTX;
bool Hybrid{};
FEX_CONFIG_OPT(Cores, THREADS);
FEX_CONFIG_OPT(HideHypervisorBit, HIDEHYPERVISORBIT);
using FunctionHandler = FEXCore::CPUID::FunctionResults (CPUIDEmu::*)(uint32_t Leaf);
void RegisterFunction(uint32_t Function, FunctionHandler Handler) {
FunctionHandlers.insert_or_assign(Function, Handler);
}
std::unordered_map<uint32_t, FunctionHandler> FunctionHandlers;
struct CPUData {
const char *ProductName{};
#ifdef _M_ARM_64
@@ -95,12 +105,161 @@ private:
FEXCore::CPUID::FunctionResults Function_8000_0006h(uint32_t Leaf);
FEXCore::CPUID::FunctionResults Function_8000_0007h(uint32_t Leaf);
FEXCore::CPUID::FunctionResults Function_8000_0008h(uint32_t Leaf);
FEXCore::CPUID::FunctionResults Function_8000_0009h(uint32_t Leaf);
FEXCore::CPUID::FunctionResults Function_8000_0019h(uint32_t Leaf);
FEXCore::CPUID::FunctionResults Function_8000_001Dh(uint32_t Leaf);
FEXCore::CPUID::FunctionResults Function_Reserved(uint32_t Leaf);
void SetupHostHybridFlag();
static constexpr std::array<FunctionHandler, 27> Primary = {
// 0: Highest function parameter and ID
&CPUIDEmu::Function_0h,
// 1: Processor info
&CPUIDEmu::Function_01h,
// 2: Cache and TLB info
&CPUIDEmu::Function_02h,
// 3: Serial Number(previously), now reserved
&CPUIDEmu::Function_Reserved,
#ifndef CPUID_AMD
// 4: Deterministic cache parameters for each level
&CPUIDEmu::Function_04h,
#else
&CPUIDEmu::Function_Reserved,
#endif
// 5: Monitor/mwait
&CPUIDEmu::Function_Reserved,
// 6: Thermal and power management
&CPUIDEmu::Function_06h,
// 7: Extended feature flags
&CPUIDEmu::Function_07h,
// 0x08: Reserved?
&CPUIDEmu::Function_Reserved,
// 9: Direct Cache Access information
&CPUIDEmu::Function_Reserved,
// 0x0A: Architectural performance monitoring
&CPUIDEmu::Function_Reserved,
// 0x0B: Extended topology enumeration
&CPUIDEmu::Function_Reserved,
// 0x0C: Reserved?
&CPUIDEmu::Function_Reserved,
// 0x0D: Processor extended state enumeration
&CPUIDEmu::Function_0Dh,
// 0x0E: Reserved?
&CPUIDEmu::Function_Reserved,
// 0x0F: Intel RDT monitoring
&CPUIDEmu::Function_Reserved,
// 0x10: Intel RDT allocation enumeration
&CPUIDEmu::Function_Reserved,
// 0x12: Reserved?
&CPUIDEmu::Function_Reserved,
// 0x12: Intel SGX capability enumeration
&CPUIDEmu::Function_Reserved,
// 0x13: Reserved
&CPUIDEmu::Function_Reserved,
// 0x14: Intel Processor trace
&CPUIDEmu::Function_Reserved,
#ifndef CPUID_AMD
// Timestamp counter information
// Doesn't exist on AMD hardware
&CPUIDEmu::Function_15h,
#else
&CPUIDEmu::Function_Reserved,
#endif
// 0x16: Processor frequency information
&CPUIDEmu::Function_Reserved,
// 0x17: SoC vendor attribute enumeration
&CPUIDEmu::Function_Reserved,
// 0x18: Reserved?
&CPUIDEmu::Function_Reserved,
// 0x19: Reserved?
&CPUIDEmu::Function_Reserved,
#ifndef CPUID_AMD
// 0x1A: Hybrid Information Sub-leaf
&CPUIDEmu::Function_1Ah,
#else
&CPUIDEmu::Function_Reserved,
#endif
};
static constexpr std::array<FunctionHandler, 2> Hypervisor = {
// Hypervisor CPUID information leaf
&CPUIDEmu::Function_4000_0000h,
// FEX-Emu specific leaf
&CPUIDEmu::Function_4000_0001h,
};
static constexpr std::array<FunctionHandler, 32> Extended = {
// Largest extended function number
&CPUIDEmu::Function_8000_0000h,
// Processor vendor
&CPUIDEmu::Function_8000_0001h,
// Processor brand string
&CPUIDEmu::Function_8000_0002h,
// Processor brand string continued
&CPUIDEmu::Function_8000_0003h,
// Processor brand string continued
&CPUIDEmu::Function_8000_0004h,
#ifdef CPUID_AMD
// 0x8000'0005: L1 Cache and TLB identifiers
&CPUIDEmu::Function_8000_0005h,
#else
&CPUIDEmu::Function_Reserved,
#endif
// 0x8000'0006: L2 Cache identifiers
&CPUIDEmu::Function_8000_0006h,
// 0x8000'0007: Advanced power management information
&CPUIDEmu::Function_8000_0007h,
// 0x8000'0008: Virtual and physical address sizes
&CPUIDEmu::Function_8000_0008h,
// 0x8000'0009: Reserved?
&CPUIDEmu::Function_Reserved,
// 0x8000'000A: SVM Revision
&CPUIDEmu::Function_Reserved,
// 0x8000'000B: Reserved?
&CPUIDEmu::Function_Reserved,
// 0x8000'000C: Reserved?
&CPUIDEmu::Function_Reserved,
// 0x8000'000D: Reserved?
&CPUIDEmu::Function_Reserved,
// 0x8000'000E: Reserved?
&CPUIDEmu::Function_Reserved,
// 0x8000'000F: Reserved?
&CPUIDEmu::Function_Reserved,
// 0x8000'0010: Reserved?
&CPUIDEmu::Function_Reserved,
// 0x8000'0011: Reserved?
&CPUIDEmu::Function_Reserved,
// 0x8000'0012: Reserved?
&CPUIDEmu::Function_Reserved,
// 0x8000'0013: Reserved?
&CPUIDEmu::Function_Reserved,
// 0x8000'0014: Reserved?
&CPUIDEmu::Function_Reserved,
// 0x8000'0015: Reserved?
&CPUIDEmu::Function_Reserved,
// 0x8000'0016: Reserved?
&CPUIDEmu::Function_Reserved,
// 0x8000'0017: Reserved?
&CPUIDEmu::Function_Reserved,
// 0x8000'0018: Reserved?
&CPUIDEmu::Function_Reserved,
// 0x8000'0019: TLB 1GB page identifiers
&CPUIDEmu::Function_8000_0019h,
// 0x8000'001A: Performance optimization identifiers
&CPUIDEmu::Function_Reserved,
// 0x8000'001B: Instruction based sampling identifiers
&CPUIDEmu::Function_Reserved,
// 0x8000'001C: Lightweight profiling capabilities
&CPUIDEmu::Function_Reserved,
#ifdef CPUID_AMD
// 0x8000'001D: Cache properties
&CPUIDEmu::Function_8000_001Dh,
#else
&CPUIDEmu::Function_Reserved,
#endif
// 0x8000'001E: Extended APIC ID
&CPUIDEmu::Function_Reserved,
// 0x8000'001F: AMD Secure Encryption
&CPUIDEmu::Function_Reserved,
};
};
}
+118 -102
View File
@@ -79,7 +79,7 @@ $end_info$
namespace FEXCore::CPU {
bool CreateCPUCore(FEXCore::Context::Context *CTX) {
bool CreateCPUCore(Context::ContextImpl *CTX) {
// This should be used for generating things that are shared between threads
CTX->CPUID.Init(CTX);
return true;
@@ -147,7 +147,7 @@ std::string_view const& GetGRegName(unsigned Reg) {
} // namespace FEXCore::Core
namespace FEXCore::Context {
Context::Context()
ContextImpl::ContextImpl()
: IRCaptureCache {this} {
#ifdef BLOCKSTATS
BlockData = std::make_unique<FEXCore::BlockSamplingData>();
@@ -159,6 +159,11 @@ namespace FEXCore::Context {
HostFeatures.SupportsAVX = false;
}
if (!Config.Is64BitMode()) {
// When operating in 32-bit mode, the virtual memory we care about is only the lower 32-bits.
Config.VirtualMemSize = 1ULL << 32;
}
if (Config.BlockJITNaming() ||
Config.GlobalJITNaming() ||
Config.LibraryJITNaming()) {
@@ -167,7 +172,7 @@ namespace FEXCore::Context {
}
}
Context::~Context() {
ContextImpl::~ContextImpl() {
{
if (CodeObjectCacheService) {
CodeObjectCacheService->Shutdown();
@@ -209,7 +214,7 @@ namespace FEXCore::Context {
return NewThreadState;
}
FEXCore::Core::InternalThreadState* Context::InitCore(uint64_t InitialRIP, uint64_t StackPointer) {
FEXCore::Core::InternalThreadState* ContextImpl::InitCore(uint64_t InitialRIP, uint64_t StackPointer) {
// Initialize the CPU core signal handlers & DispatcherConfig
switch (Config.Core) {
#ifdef INTERPRETER_ENABLED
@@ -250,13 +255,13 @@ namespace FEXCore::Context {
// Initialize common signal handlers
auto PauseHandler = [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
return Thread->CTX->Dispatcher->HandleSignalPause(Thread, Signal, info, ucontext);
return static_cast<ContextImpl*>(Thread->CTX)->Dispatcher->HandleSignalPause(Thread, Signal, info, ucontext);
};
SignalDelegation->RegisterHostSignalHandler(SignalDelegator::SIGNAL_FOR_PAUSE, PauseHandler, true);
auto GuestSignalHandler = [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack) -> bool {
return Thread->CTX->Dispatcher->HandleGuestSignal(Thread, Signal, info, ucontext, GuestAction, GuestStack);
return static_cast<ContextImpl*>(Thread->CTX)->Dispatcher->HandleGuestSignal(Thread, Signal, info, ucontext, GuestAction, GuestStack);
};
for (uint32_t Signal = 0; Signal <= SignalDelegator::MAX_SIGNALS; ++Signal) {
@@ -290,30 +295,45 @@ namespace FEXCore::Context {
return Thread;
}
void Context::StartGdbServer() {
void ContextImpl::StartGdbServer() {
if (!DebugServer) {
DebugServer = std::make_unique<GdbServer>(this);
StartPaused = true;
}
}
void Context::StopGdbServer() {
void ContextImpl::StopGdbServer() {
DebugServer.reset();
}
void Context::HandleCallback(FEXCore::Core::InternalThreadState *Thread, uint64_t RIP) {
Thread->CTX->Dispatcher->ExecuteJITCallback(Thread->CurrentFrame, RIP);
void ContextImpl::HandleCallback(FEXCore::Core::InternalThreadState *Thread, uint64_t RIP) {
static_cast<ContextImpl*>(Thread->CTX)->Dispatcher->ExecuteJITCallback(Thread->CurrentFrame, RIP);
}
void Context::RegisterHostSignalHandler(int Signal, HostSignalDelegatorFunction Func, bool Required) {
void ContextImpl::HandleSignalHandlerReturn(bool RT) {
using SignalHandlerReturnFunc = void(*)();
SignalHandlerReturnFunc SignalHandlerReturn{};
if (RT) {
SignalHandlerReturn = reinterpret_cast<SignalHandlerReturnFunc>(Dispatcher->SignalHandlerReturnAddressRT);
}
else {
SignalHandlerReturn = reinterpret_cast<SignalHandlerReturnFunc>(Dispatcher->SignalHandlerReturnAddress);
}
SignalHandlerReturn();
FEX_UNREACHABLE;
}
void ContextImpl::RegisterHostSignalHandler(int Signal, HostSignalDelegatorFunction Func, bool Required) {
SignalDelegation->RegisterHostSignalHandler(Signal, Func, Required);
}
void Context::RegisterFrontendHostSignalHandler(int Signal, HostSignalDelegatorFunction Func, bool Required) {
void ContextImpl::RegisterFrontendHostSignalHandler(int Signal, HostSignalDelegatorFunction Func, bool Required) {
SignalDelegation->RegisterFrontendHostSignalHandler(Signal, Func, Required);
}
void Context::WaitForIdle() {
void ContextImpl::WaitForIdle() {
std::unique_lock<std::mutex> lk(IdleWaitMutex);
IdleWaitCV.wait(lk, [this] {
return IdleWaitRefCount.load() == 0;
@@ -322,7 +342,7 @@ namespace FEXCore::Context {
Running = false;
}
void Context::WaitForIdleWithTimeout() {
void ContextImpl::WaitForIdleWithTimeout() {
std::unique_lock<std::mutex> lk(IdleWaitMutex);
bool WaitResult = IdleWaitCV.wait_for(lk, std::chrono::milliseconds(1500),
[this] {
@@ -340,7 +360,7 @@ namespace FEXCore::Context {
WaitForIdle();
}
void Context::NotifyPause() {
void ContextImpl::NotifyPause() {
// Tell all the threads that they should pause
std::lock_guard<std::mutex> lk(ThreadCreationMutex);
@@ -353,7 +373,7 @@ namespace FEXCore::Context {
}
}
void Context::Pause() {
void ContextImpl::Pause() {
// If we aren't running, WaitForIdle will never compete.
if (Running) {
NotifyPause();
@@ -362,7 +382,7 @@ namespace FEXCore::Context {
}
}
void Context::Run() {
void ContextImpl::Run() {
// Spin up all the threads
std::lock_guard<std::mutex> lk(ThreadCreationMutex);
for (auto &Thread : Threads) {
@@ -374,7 +394,7 @@ namespace FEXCore::Context {
}
}
void Context::WaitForThreadsToRun() {
void ContextImpl::WaitForThreadsToRun() {
size_t NumThreads{};
{
std::lock_guard<std::mutex> lk(ThreadCreationMutex);
@@ -390,7 +410,7 @@ namespace FEXCore::Context {
Running = true;
}
void Context::Step() {
void ContextImpl::Step() {
{
std::lock_guard<std::mutex> lk(ThreadCreationMutex);
// Walk the threads and tell them to clear their caches
@@ -410,7 +430,7 @@ namespace FEXCore::Context {
this->Config.MaxInstPerBlock = PreviousMaxIntPerBlock;
}
void Context::Stop(bool IgnoreCurrentThread) {
void ContextImpl::Stop(bool IgnoreCurrentThread) {
pid_t tid = FHU::Syscalls::gettid();
FEXCore::Core::InternalThreadState* CurrentThread{};
@@ -448,21 +468,21 @@ namespace FEXCore::Context {
}
}
void Context::StopThread(FEXCore::Core::InternalThreadState *Thread) {
void ContextImpl::StopThread(FEXCore::Core::InternalThreadState *Thread) {
if (Thread->RunningEvents.Running.exchange(false)) {
Thread->SignalReason.store(FEXCore::Core::SignalEvent::Stop);
FHU::Syscalls::tgkill(Thread->ThreadManager.PID, Thread->ThreadManager.TID, SignalDelegator::SIGNAL_FOR_PAUSE);
}
}
void Context::SignalThread(FEXCore::Core::InternalThreadState *Thread, FEXCore::Core::SignalEvent Event) {
void ContextImpl::SignalThread(FEXCore::Core::InternalThreadState *Thread, FEXCore::Core::SignalEvent Event) {
if (Thread->RunningEvents.Running.load()) {
Thread->SignalReason.store(Event);
FHU::Syscalls::tgkill(Thread->ThreadManager.PID, Thread->ThreadManager.TID, SignalDelegator::SIGNAL_FOR_PAUSE);
}
}
FEXCore::Context::ExitReason Context::RunUntilExit() {
FEXCore::Context::ExitReason ContextImpl::RunUntilExit() {
if(!StartPaused) {
// We will only have one thread at this point, but just in case run notify everything
std::lock_guard lk(ThreadCreationMutex);
@@ -483,16 +503,16 @@ namespace FEXCore::Context {
}
}
int Context::GetProgramStatus() const {
int ContextImpl::GetProgramStatus() const {
return ParentThread->StatusCode;
}
void Context::InitializeThreadData(FEXCore::Core::InternalThreadState *Thread) {
void ContextImpl::InitializeThreadData(FEXCore::Core::InternalThreadState *Thread) {
Thread->CPUBackend->Initialize();
}
struct ExecutionThreadHandler {
FEXCore::Context::Context *This;
ContextImpl *This;
FEXCore::Core::InternalThreadState *Thread;
};
@@ -503,7 +523,7 @@ namespace FEXCore::Context {
return nullptr;
}
void Context::InitializeThread(FEXCore::Core::InternalThreadState *Thread) {
void ContextImpl::InitializeThread(FEXCore::Core::InternalThreadState *Thread) {
// This will create the execution thread but it won't actually start executing
ExecutionThreadHandler *Arg = reinterpret_cast<ExecutionThreadHandler*>(FEXCore::Allocator::malloc(sizeof(ExecutionThreadHandler)));
Arg->This = this;
@@ -525,7 +545,7 @@ namespace FEXCore::Context {
}
}
void Context::InitializeThreadTLSData(FEXCore::Core::InternalThreadState *Thread) {
void ContextImpl::InitializeThreadTLSData(FEXCore::Core::InternalThreadState *Thread) {
// Let's do some initial bookkeeping here
Thread->ThreadManager.TID = FHU::Syscalls::gettid();
Thread->ThreadManager.PID = ::getpid();
@@ -533,12 +553,12 @@ namespace FEXCore::Context {
ThunkHandler->RegisterTLSState(Thread);
}
void Context::RunThread(FEXCore::Core::InternalThreadState *Thread) {
void ContextImpl::RunThread(FEXCore::Core::InternalThreadState *Thread) {
// Tell the thread to start executing
Thread->StartRunning.NotifyAll();
}
void Context::InitializeCompiler(FEXCore::Core::InternalThreadState* Thread) {
void ContextImpl::InitializeCompiler(FEXCore::Core::InternalThreadState* Thread) {
Thread->OpDispatcher = std::make_unique<FEXCore::IR::OpDispatchBuilder>(this);
Thread->OpDispatcher->SetMultiblock(Config.Multiblock);
Thread->LookupCache = std::make_unique<FEXCore::LookupCache>(this);
@@ -590,7 +610,7 @@ namespace FEXCore::Context {
}
}
FEXCore::Core::InternalThreadState* Context::CreateThread(FEXCore::Core::CPUState *NewThreadState, uint64_t ParentTID) {
FEXCore::Core::InternalThreadState* ContextImpl::CreateThread(FEXCore::Core::CPUState *NewThreadState, uint64_t ParentTID) {
FEXCore::Core::InternalThreadState *Thread = new FEXCore::Core::InternalThreadState{};
// Copy over the new thread state to the new object
@@ -612,7 +632,7 @@ namespace FEXCore::Context {
return Thread;
}
void Context::DestroyThread(FEXCore::Core::InternalThreadState *Thread) {
void ContextImpl::DestroyThread(FEXCore::Core::InternalThreadState *Thread) {
// remove new thread object
{
std::lock_guard lk(ThreadCreationMutex);
@@ -631,7 +651,7 @@ namespace FEXCore::Context {
delete Thread;
}
void Context::CleanupAfterFork(FEXCore::Core::InternalThreadState *LiveThread) {
void ContextImpl::CleanupAfterFork(FEXCore::Core::InternalThreadState *LiveThread) {
// This function is called after fork
// We need to cleanup some of the thread data that is dead
for (auto &DeadThread : Threads) {
@@ -670,11 +690,11 @@ namespace FEXCore::Context {
FEXCore::Threads::Thread::CleanupAfterFork();
}
void Context::AddBlockMapping(FEXCore::Core::InternalThreadState *Thread, uint64_t Address, void *Ptr) {
void ContextImpl::AddBlockMapping(FEXCore::Core::InternalThreadState *Thread, uint64_t Address, void *Ptr) {
Thread->LookupCache->AddBlockMapping(Address, Ptr);
}
void Context::ClearCodeCache(FEXCore::Core::InternalThreadState *Thread) {
void ContextImpl::ClearCodeCache(FEXCore::Core::InternalThreadState *Thread) {
FEXCORE_PROFILE_INSTANT("ClearCodeCache");
{
@@ -692,7 +712,7 @@ namespace FEXCore::Context {
static void IRDumper(FEXCore::Core::InternalThreadState *Thread, IR::IREmitter *IREmitter, uint64_t GuestRIP, IR::RegisterAllocationData* RA) {
FILE* f = nullptr;
bool CloseAfter = false;
const auto DumpIRStr = Thread->CTX->Config.DumpIR();
const auto DumpIRStr = static_cast<ContextImpl*>(Thread->CTX)->Config.DumpIR();
// DumpIRStr might be no if not dumping but ShouldDump is set in OpDisp
if (DumpIRStr =="stderr" || DumpIRStr =="no") {
@@ -719,7 +739,7 @@ namespace FEXCore::Context {
}
};
static void ValidateIR(FEXCore::Context::Context *ctx, IR::IREmitter *IREmitter) {
static void ValidateIR(ContextImpl *ctx, IR::IREmitter *IREmitter) {
// Convert to text, Parse, Convert to text again and make sure the texts match
std::stringstream out;
static auto compaction = IR::CreateIRCompaction(ctx->OpDispatcherAllocator);
@@ -743,7 +763,7 @@ namespace FEXCore::Context {
}
}
Context::GenerateIRResult Context::GenerateIR(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP, bool ExtendedDebugInfo) {
ContextImpl::GenerateIRResult ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP, bool ExtendedDebugInfo) {
FEXCORE_PROFILE_SCOPED("GenerateIR");
Thread->OpDispatcher->ReownOrClaimBuffer();
@@ -770,7 +790,7 @@ namespace FEXCore::Context {
Thread->FrontendDecoder->DecodeInstructionsAtEntry(GuestCode, GuestRIP, [Thread](uint64_t BlockEntry, uint64_t Start, uint64_t Length) {
if (Thread->LookupCache->AddBlockExecutableRange(BlockEntry, Start, Length)) {
Thread->CTX->SyscallHandler->MarkGuestExecutableRange(Start, Length);
static_cast<ContextImpl*>(Thread->CTX)->SyscallHandler->MarkGuestExecutableRange(Start, Length);
}
});
@@ -790,13 +810,6 @@ namespace FEXCore::Context {
// Reset any block-specific state
Thread->OpDispatcher->StartNewBlock();
if (Config.x86dec_SynchronizeRIPOnAllBlocks) {
// Ensure the RIP is synchronized to the context on block entry.
// In the case of block linking, the RIP may not have synchronized.
auto NewRIP = Thread->OpDispatcher->_EntrypointOffset(Block.Entry - GuestRIP, GPRSize);
Thread->OpDispatcher->_StoreContext(GPRSize, IR::GPRClass, NewRIP, offsetof(FEXCore::Core::CPUState, rip));
}
uint64_t InstsInBlock = Block.NumInstructions;
for (size_t i = 0; i < InstsInBlock; ++i) {
@@ -856,22 +869,25 @@ namespace FEXCore::Context {
Thread->OpDispatcher->_ExitFunction(Thread->OpDispatcher->_EntrypointOffset(Block.Entry - GuestRIP, GPRSize));
}
// If we had a dispatch error then leave early
if (HadDispatchError) {
if (TotalInstructions == 0) {
// Couldn't handle any instruction in op dispatcher
Thread->OpDispatcher->ResetWorkingList();
return { nullptr, nullptr, 0, 0, 0, 0 };
}
else {
const uint8_t GPRSize = GetGPRSize();
const bool NeedsBlockEnd = (HadDispatchError && TotalInstructions > 0) ||
(Thread->OpDispatcher->NeedsBlockEnder() && i + 1 == InstsInBlock);
// We had some instructions. Early exit
Thread->OpDispatcher->_ExitFunction(Thread->OpDispatcher->_EntrypointOffset(Block.Entry + BlockInstructionsLength - GuestRIP, GPRSize));
break;
}
// If we had a dispatch error then leave early
if (HadDispatchError && TotalInstructions == 0) {
// Couldn't handle any instruction in op dispatcher
Thread->OpDispatcher->ResetWorkingList();
return { nullptr, nullptr, 0, 0, 0, 0 };
}
if (NeedsBlockEnd) {
const uint8_t GPRSize = GetGPRSize();
// We had some instructions. Early exit
Thread->OpDispatcher->_ExitFunction(Thread->OpDispatcher->_EntrypointOffset(Block.Entry + BlockInstructionsLength - GuestRIP, GPRSize));
break;
}
if (Thread->OpDispatcher->FinishOp(DecodedInfo->PC + DecodedInfo->InstSize, i + 1 == InstsInBlock)) {
break;
}
@@ -885,14 +901,14 @@ namespace FEXCore::Context {
IR::IREmitter *IREmitter = Thread->OpDispatcher.get();
auto ShouldDump = Thread->CTX->Config.DumpIR() != "no" || Thread->OpDispatcher->ShouldDump;
auto ShouldDump = static_cast<ContextImpl*>(Thread->CTX)->Config.DumpIR() != "no" || Thread->OpDispatcher->ShouldDump;
// Debug
{
if (ShouldDump) {
IRDumper(Thread, IREmitter, GuestRIP, nullptr);
}
if (Thread->CTX->Config.ValidateIRarser) {
if (static_cast<ContextImpl*>(Thread->CTX)->Config.ValidateIRarser) {
ValidateIR(this, IREmitter);
}
}
@@ -922,7 +938,7 @@ namespace FEXCore::Context {
};
}
Context::CompileCodeResult Context::CompileCode(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP) {
ContextImpl::CompileCodeResult ContextImpl::CompileCode(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP) {
FEXCore::IR::IRListView *IRList {};
FEXCore::Core::DebugData *DebugData {};
FEXCore::IR::RegisterAllocationData::UniquePtr RAData {};
@@ -994,7 +1010,10 @@ namespace FEXCore::Context {
}
// Attempt to get the CPU backend to compile this code
return {
.CompiledCode = Thread->CPUBackend->CompileCode(GuestRIP, IRList, DebugData, RAData.get(), GetGdbServerStatus()),
// FEX currently throws away the CPUBackend::CompiledCode object other than the entrypoint
// In the future with code caching getting wired up, we will pass the rest of the data forward.
// TODO: Pass the data forward when code caching is wired up to this.
.CompiledCode = Thread->CPUBackend->CompileCode(GuestRIP, IRList, DebugData, RAData.get(), GetGdbServerStatus()).BlockEntry,
.IRData = IRList,
.DebugData = DebugData,
.RAData = std::move(RAData),
@@ -1004,7 +1023,7 @@ namespace FEXCore::Context {
};
}
void Context::CompileBlockJit(FEXCore::Core::CpuStateFrame *Frame, uint64_t GuestRIP) {
void ContextImpl::CompileBlockJit(FEXCore::Core::CpuStateFrame *Frame, uint64_t GuestRIP) {
auto NewBlock = CompileBlock(Frame, GuestRIP);
if (NewBlock == 0) {
@@ -1015,7 +1034,7 @@ namespace FEXCore::Context {
}
}
uintptr_t Context::CompileBlock(FEXCore::Core::CpuStateFrame *Frame, uint64_t GuestRIP) {
uintptr_t ContextImpl::CompileBlock(FEXCore::Core::CpuStateFrame *Frame, uint64_t GuestRIP) {
FEXCORE_PROFILE_SCOPED("CompileBlock");
auto Thread = Frame->Thread;
@@ -1115,7 +1134,7 @@ namespace FEXCore::Context {
return (uintptr_t)CodePtr;
}
void Context::ExecutionThread(FEXCore::Core::InternalThreadState *Thread) {
void ContextImpl::ExecutionThread(FEXCore::Core::InternalThreadState *Thread) {
Core::ThreadData.Thread = Thread;
Thread->ExitReason = FEXCore::Context::ExitReason::EXIT_WAITING;
@@ -1126,7 +1145,7 @@ namespace FEXCore::Context {
// Now notify the thread that we are initialized
Thread->ThreadWaiting.NotifyAll();
if (Thread != Thread->CTX->ParentThread || StartPaused || Thread->StartPaused) {
if (Thread != static_cast<ContextImpl*>(Thread->CTX)->ParentThread || StartPaused || Thread->StartPaused) {
// Parent thread doesn't need to wait to run
Thread->StartRunning.Wait();
}
@@ -1138,7 +1157,7 @@ namespace FEXCore::Context {
Thread->RunningEvents.Running = true;
Thread->CTX->Dispatcher->ExecuteDispatch(Thread->CurrentFrame);
static_cast<ContextImpl*>(Thread->CTX)->Dispatcher->ExecuteDispatch(Thread->CurrentFrame);
Thread->RunningEvents.Running = false;
}
@@ -1167,7 +1186,7 @@ namespace FEXCore::Context {
SignalDelegation->UninstallTLSState(Thread);
// If the parent thread is waiting to join, then we can't destroy our thread object
if (!Thread->DestroyedByParent && Thread != Thread->CTX->ParentThread) {
if (!Thread->DestroyedByParent && Thread != static_cast<ContextImpl*>(Thread->CTX)->ParentThread) {
Thread->CTX->DestroyThread(Thread);
}
}
@@ -1180,34 +1199,34 @@ namespace FEXCore::Context {
for (auto it = lower; it != upper; it++) {
for (auto Address: it->second) {
Context::ThreadRemoveCodeEntry(Thread, Address);
ContextImpl::ThreadRemoveCodeEntry(Thread, Address);
}
it->second.clear();
}
}
static void InvalidateGuestCodeRangeInternal(FEXCore::Context::Context *CTX, uint64_t Start, uint64_t Length) {
std::lock_guard lk(CTX->ThreadCreationMutex);
static void InvalidateGuestCodeRangeInternal(ContextImpl *CTX, uint64_t Start, uint64_t Length) {
std::lock_guard lk(static_cast<ContextImpl*>(CTX)->ThreadCreationMutex);
for (auto &Thread : CTX->Threads) {
for (auto &Thread : static_cast<ContextImpl*>(CTX)->Threads) {
InvalidateGuestThreadCodeRange(Thread, Start, Length);
}
}
void InvalidateGuestCodeRange(FEXCore::Context::Context *CTX, uint64_t Start, uint64_t Length) {
FHU::ScopedSignalMaskWithUniqueLock CodeInvalidationLock(CTX->CodeInvalidationMutex);
void ContextImpl::InvalidateGuestCodeRange(uint64_t Start, uint64_t Length) {
FHU::ScopedSignalMaskWithUniqueLock CodeInvalidationLock(CodeInvalidationMutex);
InvalidateGuestCodeRangeInternal(CTX, Start, Length);
InvalidateGuestCodeRangeInternal(this, Start, Length);
}
void InvalidateGuestCodeRange(FEXCore::Context::Context *CTX, uint64_t Start, uint64_t Length, std::function<void(uint64_t start, uint64_t Length)> CallAfter) {
FHU::ScopedSignalMaskWithUniqueLock CodeInvalidationLock(CTX->CodeInvalidationMutex);
void ContextImpl::InvalidateGuestCodeRange(uint64_t Start, uint64_t Length, std::function<void(uint64_t start, uint64_t Length)> CallAfter) {
FHU::ScopedSignalMaskWithUniqueLock CodeInvalidationLock(CodeInvalidationMutex);
InvalidateGuestCodeRangeInternal(CTX, Start, Length);
InvalidateGuestCodeRangeInternal(this, Start, Length);
CallAfter(Start, Length);
}
void Context::MarkMemoryShared() {
void ContextImpl::MarkMemoryShared() {
if (!IsMemoryShared) {
IsMemoryShared = true;
@@ -1227,18 +1246,14 @@ namespace FEXCore::Context {
}
}
void MarkMemoryShared(FEXCore::Context::Context *CTX) {
CTX->MarkMemoryShared();
}
void Context::ThreadAddBlockLink(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestDestination, uintptr_t HostLink, const std::function<void()> &delinker) {
std::shared_lock lk(Thread->CTX->CodeInvalidationMutex);
void ContextImpl::ThreadAddBlockLink(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestDestination, uintptr_t HostLink, const std::function<void()> &delinker) {
std::shared_lock lk(static_cast<ContextImpl*>(Thread->CTX)->CodeInvalidationMutex);
Thread->LookupCache->AddBlockLink(GuestDestination, HostLink, delinker);
}
void Context::ThreadRemoveCodeEntry(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP) {
LogMan::Throw::AFmt(Thread->CTX->CodeInvalidationMutex.try_lock() == false, "CodeInvalidationMutex needs to be unique_locked here");
void ContextImpl::ThreadRemoveCodeEntry(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP) {
LogMan::Throw::AFmt(static_cast<ContextImpl*>(Thread->CTX)->CodeInvalidationMutex.try_lock() == false, "CodeInvalidationMutex needs to be unique_locked here");
std::lock_guard<std::recursive_mutex> lk(Thread->LookupCache->WriteLock);
@@ -1246,7 +1261,7 @@ namespace FEXCore::Context {
Thread->LookupCache->Erase(GuestRIP);
}
CustomIRResult Context::AddCustomIREntrypoint(uintptr_t Entrypoint, std::function<void(uintptr_t Entrypoint, FEXCore::IR::IREmitter *)> Handler, void *Creator, void *Data) {
CustomIRResult ContextImpl::AddCustomIREntrypoint(uintptr_t Entrypoint, std::function<void(uintptr_t Entrypoint, FEXCore::IR::IREmitter *)> Handler, void *Creator, void *Data) {
LOGMAN_THROW_A_FMT(Config.Is64BitMode || !(Entrypoint >> 32), "64-bit Entrypoint in 32-bit mode {:x}", Entrypoint);
std::unique_lock lk(CustomIRMutex);
@@ -1262,12 +1277,12 @@ namespace FEXCore::Context {
}
}
void Context::RemoveCustomIREntrypoint(uintptr_t Entrypoint) {
void ContextImpl::RemoveCustomIREntrypoint(uintptr_t Entrypoint) {
LOGMAN_THROW_A_FMT(Config.Is64BitMode || !(Entrypoint >> 32), "64-bit Entrypoint in 32-bit mode {:x}", Entrypoint);
std::scoped_lock lk(CustomIRMutex);
InvalidateGuestCodeRange(this, Entrypoint, 1, [this](uint64_t Entrypoint, uint64_t) {
InvalidateGuestCodeRange(Entrypoint, 1, [this](uint64_t Entrypoint, uint64_t) {
CustomIRHandlers.erase(Entrypoint);
});
}
@@ -1277,24 +1292,26 @@ namespace FEXCore::Context {
uint64_t RIPBackup = Thread->CurrentFrame->State.rip;
Thread->CurrentFrame->State.rip = RIP;
auto CTX = static_cast<ContextImpl*>(Thread->CTX);
// Erase the RIP from all the storage backings if it exists
ThreadRemoveCodeEntry(Thread, RIP);
CTX->ThreadRemoveCodeEntry(Thread, RIP);
// We don't care if compilation passes or not
CompileBlock(Thread->CurrentFrame, RIP);
CTX->CompileBlock(Thread->CurrentFrame, RIP);
Thread->CurrentFrame->State.rip = RIPBackup;
}
uint64_t Context::GetThreadCount() const {
uint64_t ContextImpl::GetThreadCount() const {
return Threads.size();
}
FEXCore::Core::RuntimeStats *Context::GetRuntimeStatsForThread(uint64_t Thread) {
FEXCore::Core::RuntimeStats *ContextImpl::GetRuntimeStatsForThread(uint64_t Thread) {
return &Threads[Thread]->Stats;
}
bool Context::GetDebugDataForRIP(uint64_t RIP, FEXCore::Core::DebugData *Data) {
bool ContextImpl::GetDebugDataForRIP(uint64_t RIP, FEXCore::Core::DebugData *Data) {
std::lock_guard<std::recursive_mutex> lk(ParentThread->LookupCache->WriteLock);
auto it = ParentThread->DebugStore.find(RIP);
if (it == ParentThread->DebugStore.end()) {
@@ -1305,7 +1322,7 @@ namespace FEXCore::Context {
return true;
}
bool Context::FindHostCodeForRIP(uint64_t RIP, uint8_t **Code) {
bool ContextImpl::FindHostCodeForRIP(uint64_t RIP, uint8_t **Code) {
uintptr_t HostCode = ParentThread->LookupCache->FindBlock(RIP);
if (!HostCode) {
return false;
@@ -1321,7 +1338,7 @@ namespace FEXCore::Context {
return Result;
}
IR::AOTIRCacheEntry *Context::LoadAOTIRCacheEntry(const std::string &filename) {
IR::AOTIRCacheEntry *ContextImpl::LoadAOTIRCacheEntry(const std::string &filename) {
auto rv = IRCaptureCache.LoadAOTIRCacheEntry(filename);
if (DebugServer) {
DebugServer->AlertLibrariesChanged();
@@ -1329,19 +1346,18 @@ namespace FEXCore::Context {
return rv;
}
void Context::UnloadAOTIRCacheEntry(IR::AOTIRCacheEntry *Entry) {
void ContextImpl::UnloadAOTIRCacheEntry(IR::AOTIRCacheEntry *Entry) {
IRCaptureCache.UnloadAOTIRCacheEntry(Entry);
if (DebugServer) {
DebugServer->AlertLibrariesChanged();
}
}
void Context::AppendThunkDefinitions(std::vector<FEXCore::IR::ThunkDefinition> const& Definitions) {
void ContextImpl::AppendThunkDefinitions(std::vector<FEXCore::IR::ThunkDefinition> const& Definitions) {
ThunkHandler->AppendThunkDefinitions(Definitions);
}
void ConfigureAOTGen(FEXCore::Core::InternalThreadState *Thread, std::set<uint64_t> *ExternalBranches, uint64_t SectionMaxAddress) {
void ContextImpl::ConfigureAOTGen(FEXCore::Core::InternalThreadState *Thread, std::set<uint64_t> *ExternalBranches, uint64_t SectionMaxAddress) {
Thread->FrontendDecoder->SetExternalBranches(ExternalBranches);
Thread->FrontendDecoder->SetSectionMaxAddress(SectionMaxAddress);
}
+3 -3
View File
@@ -5,7 +5,7 @@ namespace FEXCore {
}
namespace FEXCore::Context {
struct Context;
class ContextImpl;
}
namespace FEXCore::CPU {
@@ -17,7 +17,7 @@ namespace FEXCore::CPU {
*
* @return true if core was able to be create
*/
bool CreateCPUCore(FEXCore::Context::Context *CTX);
bool CreateCPUCore(FEXCore::Context::ContextImpl *CTX);
bool LoadCode(FEXCore::Context::Context *CTX, FEXCore::CodeLoader *Loader);
bool LoadCode(FEXCore::Context::ContextImpl *CTX, FEXCore::CodeLoader *Loader);
}
@@ -1,3 +1,4 @@
#include "Interface/Core/ArchHelpers/CodeEmitter/Emitter.h"
#include "Interface/Core/LookupCache.h"
#include "Interface/Core/ArchHelpers/MContext.h"
@@ -30,29 +31,28 @@
#include <sys/syscall.h>
#include <unistd.h>
#define STATE_PTR(STATE_TYPE, FIELD) \
MemOperand(STATE, offsetof(FEXCore::Core::STATE_TYPE, FIELD))
namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
constexpr size_t MAX_DISPATCHER_CODE_SIZE = 4096;
constexpr size_t MAX_DISPATCHER_CODE_SIZE = 8192;
Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, const DispatcherConfig &config)
Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::ContextImpl *ctx, const DispatcherConfig &config)
: FEXCore::CPU::Dispatcher(ctx, config), Arm64Emitter(ctx, MAX_DISPATCHER_CODE_SIZE)
#ifdef VIXL_SIMULATOR
, Simulator {&Decoder}
#endif
{
#ifdef VIXL_SIMULATOR
// Hardcode a 256-bit vector width if we are running in the simulator.
Simulator.SetVectorLengthInBits(256);
#endif
SetAllowAssembler(true);
EmitDispatcher();
}
void Arm64Dispatcher::EmitDispatcher() {
#ifdef VIXL_DISASSEMBLER
const auto DisasmBegin = GetCursorAddress<const vixl::aarch64::Instruction*>();
#endif
DispatchPtr = GetCursorAddress<AsmDispatch>();
@@ -64,9 +64,9 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, const Dispatche
// Ptr();
// }
Literal l_CTX {reinterpret_cast<uintptr_t>(CTX)};
Literal l_Sleep {reinterpret_cast<uint64_t>(SleepThread)};
Literal l_CompileBlock {GetCompileBlockPtr()};
ARMEmitter::ForwardLabel l_CTX;
ARMEmitter::ForwardLabel l_Sleep;
ARMEmitter::ForwardLabel l_CompileBlock;
// Push all the register we need to save
PushCalleeSavedRegisters();
@@ -74,12 +74,12 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, const Dispatche
// Push our memory base to the correct register
// Move our thread pointer to the correct register
// This is passed in to parameter 0 (x0)
mov(STATE, x0);
mov(STATE, ARMEmitter::XReg::x0);
// Save this stack pointer so we can cleanly shutdown the emulation with a long jump
// regardless of where we were in the stack
add(x0, sp, 0);
str(x0, STATE_PTR(CpuStateFrame, ReturningStackLocation));
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, ARMEmitter::Reg::rsp, 0);
str(ARMEmitter::XReg::x0, STATE_PTR(CpuStateFrame, ReturningStackLocation));
AbsoluteLoopTopAddressFillSRA = GetCursorAddress<uint64_t>();
@@ -89,91 +89,99 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, const Dispatche
// We want to ensure that we are 16 byte aligned at the top of this loop
Align16B();
aarch64::Label FullLookup{};
aarch64::Label CallBlock{};
aarch64::Label LoopTop{};
aarch64::Label ExitSpillSRA{};
aarch64::Label ThreadPauseHandler{};
ARMEmitter::BiDirectionalLabel FullLookup{};
ARMEmitter::BiDirectionalLabel CallBlock{};
ARMEmitter::BackwardLabel LoopTop{};
bind(&LoopTop);
AbsoluteLoopTopAddress = GetLabelAddress<uint64_t>(&LoopTop);
Bind(&LoopTop);
AbsoluteLoopTopAddress = GetCursorAddress<uint64_t>();
// Load in our RIP
// Don't modify x2 since it contains our RIP once the block doesn't exist
ldr(x2, STATE_PTR(CpuStateFrame, State.rip));
auto RipReg = x2;
auto RipReg = ARMEmitter::XReg::x2;
ldr(RipReg, STATE_PTR(CpuStateFrame, State.rip));
// L1 Cache
ldr(x0, STATE_PTR(CpuStateFrame, Pointers.Common.L1Pointer));
ldr(ARMEmitter::XReg::x0, STATE_PTR(CpuStateFrame, Pointers.Common.L1Pointer));
and_(x3, RipReg, LookupCache::L1_ENTRIES_MASK);
add(x0, x0, Operand(x3, Shift::LSL, 4));
ldp(x3, x0, MemOperand(x0));
cmp(x0, RipReg);
b(&FullLookup, Condition::ne);
and_(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r3, RipReg.R(), LookupCache::L1_ENTRIES_MASK);
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, ARMEmitter::Reg::r0, ARMEmitter::Reg::r3, ARMEmitter::ShiftType::LSL , 4);
ldp<ARMEmitter::IndexType::OFFSET>(ARMEmitter::XReg::x3, ARMEmitter::XReg::x0, ARMEmitter::Reg::r0, 0);
cmp(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, RipReg.R());
b(ARMEmitter::Condition::CC_NE, &FullLookup);
br(x3);
br(ARMEmitter::Reg::r3);
// L1C check failed, do a full lookup
bind(&FullLookup);
Bind(&FullLookup);
// This is the block cache lookup routine
// It matches what is going on it LookupCache.h::FindBlock
ldr(x0, STATE_PTR(CpuStateFrame, Pointers.Common.L2Pointer));
ldr(ARMEmitter::XReg::x0, STATE_PTR(CpuStateFrame, Pointers.Common.L2Pointer));
// Mask the address by the virtual address size so we can check for aliases
uint64_t VirtualMemorySize = CTX->Config.VirtualMemSize;
if (std::popcount(VirtualMemorySize) == 1) {
and_(x3, RipReg, VirtualMemorySize - 1);
and_(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r3, RipReg.R(), VirtualMemorySize - 1);
}
else {
LoadConstant(x3, VirtualMemorySize);
and_(x3, RipReg, x3);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r3, VirtualMemorySize);
and_(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r3, RipReg.R(), ARMEmitter::Reg::r3);
}
aarch64::Label NoBlock;
#ifdef VIXL_SIMULATOR
// VIXL simulator can't run syscalls.
constexpr bool SignalSafeCompile = false;
#else
constexpr bool SignalSafeCompile = true;
#endif
ARMEmitter::ForwardLabel NoBlock;
{
// Offset the address and add to our page pointer
lsr(x1, x3, 12);
lsr(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r1, ARMEmitter::Reg::r3, 12);
// Load the pointer from the offset
ldr(x0, MemOperand(x0, x1, Shift::LSL, 3));
ldr(ARMEmitter::XReg::x0, ARMEmitter::Reg::r0, ARMEmitter::Reg::r1, ARMEmitter::ExtendedType::LSL_64, 3);
// If page pointer is zero then we have no block
cbz(x0, &NoBlock);
cbz(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, &NoBlock);
// Steal the page offset
and_(x1, x3, 0x0FFF);
and_(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r1, ARMEmitter::Reg::r3, 0x0FFF);
// Shift the offset by the size of the block cache entry
add(x0, x0, Operand(x1, Shift::LSL, (int)log2(sizeof(FEXCore::LookupCache::LookupCacheEntry))));
add(ARMEmitter::XReg::x0, ARMEmitter::XReg::x0, ARMEmitter::XReg::x1, ARMEmitter::ShiftType::LSL, (int)log2(sizeof(FEXCore::LookupCache::LookupCacheEntry)));
// Load the guest address first to ensure it maps to the address we are currently at
// The the full LookupCacheEntry with a single LDP.
// Check the guest address first to ensure it maps to the address we are currently at.
// This fixes aliasing problems
ldr(x1, MemOperand(x0, offsetof(FEXCore::LookupCache::LookupCacheEntry, GuestCode)));
cmp(x1, RipReg);
b(&NoBlock, Condition::ne);
ldp<ARMEmitter::IndexType::OFFSET>(ARMEmitter::XReg::x3, ARMEmitter::XReg::x1, ARMEmitter::Reg::r0, 0);
// Now load the actual host block to execute if we can
ldr(x3, MemOperand(x0, offsetof(FEXCore::LookupCache::LookupCacheEntry, HostCode)));
cbz(x3, &NoBlock);
// If the guest address doesn't match, Compile the block.
cmp(ARMEmitter::XReg::x1, RipReg);
b(ARMEmitter::Condition::CC_NE, &NoBlock);
// Check the host address to see if it matches, else compile the block.
cbz(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r3, &NoBlock);
// If we've made it here then we have a real compiled block
{
// update L1 cache
ldr(x0, STATE_PTR(CpuStateFrame, Pointers.Common.L1Pointer));
ldr(ARMEmitter::XReg::x0, STATE_PTR(CpuStateFrame, Pointers.Common.L1Pointer));
and_(x1, RipReg, LookupCache::L1_ENTRIES_MASK);
add(x0, x0, Operand(x1, Shift::LSL, 4));
stp(x3, x2, MemOperand(x0));
and_(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r1, RipReg.R(), LookupCache::L1_ENTRIES_MASK);
add(ARMEmitter::XReg::x0, ARMEmitter::XReg::x0, ARMEmitter::XReg::x1, ARMEmitter::ShiftType::LSL, 4);
stp<ARMEmitter::IndexType::OFFSET>(ARMEmitter::XReg::x3, ARMEmitter::XReg::x2, ARMEmitter::Reg::r0);
// Jump to the block
br(x3);
br(ARMEmitter::Reg::r3);
}
}
{
bind(&ExitSpillSRA);
ThreadStopHandlerAddressSpillSRA = GetCursorAddress<uint64_t>();
if (config.StaticRegisterAllocation)
SpillStaticRegs();
@@ -187,12 +195,6 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, const Dispatche
ret();
}
#ifdef VIXL_SIMULATOR
// VIXL simulator can't run syscalls.
constexpr bool SignalSafeCompile = false;
#else
constexpr bool SignalSafeCompile = true;
#endif
{
ExitFunctionLinkerAddress = GetCursorAddress<uint64_t>();
if (config.StaticRegisterAllocation)
@@ -207,52 +209,52 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, const Dispatche
// X3: Size of mask, sizeof(uint64_t)
// X8: Syscall
LoadConstant(x0, ~0ULL);
stp(x0, x0, MemOperand(sp, -16, PreIndex));
LoadConstant(x0, SIG_SETMASK);
add(x1, sp, 0);
add(x2, sp, 0);
LoadConstant(x3, 8);
LoadConstant(x8, SYS_rt_sigprocmask);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, ~0ULL);
stp<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::x0, ARMEmitter::XReg::x0, ARMEmitter::Reg::rsp, -16);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, SIG_SETMASK);
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r1, ARMEmitter::Reg::rsp, 0);
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r2, ARMEmitter::Reg::rsp, 0);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r3, 8);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r8, SYS_rt_sigprocmask);
svc(0);
}
mov(x0, STATE);
mov(x1, lr);
mov(ARMEmitter::XReg::x0, STATE);
mov(ARMEmitter::XReg::x1, ARMEmitter::XReg::lr);
ldr(x2, STATE_PTR(CpuStateFrame, Pointers.Common.ExitFunctionLink));
ldr(ARMEmitter::XReg::x2, STATE_PTR(CpuStateFrame, Pointers.Common.ExitFunctionLink));
#ifdef VIXL_SIMULATOR
GenerateIndirectRuntimeCall<uintptr_t, void *, void *>(x2);
GenerateIndirectRuntimeCall<uintptr_t, void *, void *>(ARMEmitter::Reg::r2);
#else
blr(x2);
blr(ARMEmitter::Reg::r2);
#endif
if (SignalSafeCompile) {
// Now restore the signal mask
// Living in the same location
mov(x4, x0);
LoadConstant(x0, SIG_SETMASK);
add(x1, sp, 0);
LoadConstant(x2, 0);
LoadConstant(x3, 8);
LoadConstant(x8, SYS_rt_sigprocmask);
mov(ARMEmitter::XReg::x4, ARMEmitter::XReg::x0);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, SIG_SETMASK);
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r1, ARMEmitter::Reg::rsp, 0);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r2, 0);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r3, 8);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r8, SYS_rt_sigprocmask);
svc(0);
// Bring stack back
add(sp, sp, 16);
mov(x0, x4);
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, ARMEmitter::Reg::rsp, 16);
mov(ARMEmitter::XReg::x0, ARMEmitter::XReg::x4);
}
if (config.StaticRegisterAllocation)
FillStaticRegs();
br(x0);
br(ARMEmitter::Reg::r0);
}
// Need to create the block
{
bind(&NoBlock);
Bind(&NoBlock);
if (config.StaticRegisterAllocation)
SpillStaticRegs();
@@ -266,42 +268,42 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, const Dispatche
// X3: Size of mask, sizeof(uint64_t)
// X8: Syscall
LoadConstant(x0, ~0ULL);
stp(x0, x2, MemOperand(sp, -16, PreIndex));
LoadConstant(x0, SIG_SETMASK);
add(x1, sp, 0);
add(x2, sp, 0);
LoadConstant(x3, 8);
LoadConstant(x8, SYS_rt_sigprocmask);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, ~0ULL);
stp<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::x0, ARMEmitter::XReg::x2, ARMEmitter::Reg::rsp, -16);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, SIG_SETMASK);
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r1, ARMEmitter::Reg::rsp, 0);
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r2, ARMEmitter::Reg::rsp, 0);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r3, 8);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r8, SYS_rt_sigprocmask);
svc(0);
// Reload x2 to bring back RIP
ldr(x2, MemOperand(sp, 8, Offset));
ldr(ARMEmitter::XReg::x2, ARMEmitter::Reg::rsp, 8);
}
ldr(x0, &l_CTX);
mov(x1, STATE);
ldr(x3, &l_CompileBlock);
ldr(ARMEmitter::XReg::x0, &l_CTX);
mov(ARMEmitter::XReg::x1, STATE);
ldr(ARMEmitter::XReg::x3, &l_CompileBlock);
// X2 contains our guest RIP
#ifdef VIXL_SIMULATOR
GenerateIndirectRuntimeCall<void, void *, uint64_t, void *>(x3);
GenerateIndirectRuntimeCall<void, void *, uint64_t, void *>(ARMEmitter::Reg::r3);
#else
blr(x3); // { CTX, Frame, RIP}
blr(ARMEmitter::Reg::r3); // { CTX, Frame, RIP}
#endif
if (SignalSafeCompile) {
// Now restore the signal mask
// Living in the same location
LoadConstant(x0, SIG_SETMASK);
add(x1, sp, 0);
LoadConstant(x2, 0);
LoadConstant(x3, 8);
LoadConstant(x8, SYS_rt_sigprocmask);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, SIG_SETMASK);
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r1, ARMEmitter::Reg::rsp, 0);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r2, 0);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r3, 8);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r8, SYS_rt_sigprocmask);
svc(0);
// Bring stack back
add(sp, sp, 16);
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, ARMEmitter::Reg::rsp, 16);
}
if (config.StaticRegisterAllocation)
@@ -318,10 +320,18 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, const Dispatche
hlt(0);
}
{
SignalHandlerReturnAddressRT = GetCursorAddress<uint64_t>();
// Now to get back to our old location we need to do a fault dance
// We can't use SIGTRAP here since gdb catches it and never gives it to the application!
hlt(0);
}
{
// Guest SIGILL handler
// Needs to be distinct from the SignalHandlerReturnAddress
GuestSignal_SIGILL = GetCursorAddress<uint64_t>();
GuestSignal_SIGILL = GetCursorAddress<uint64_t>();
if (config.StaticRegisterAllocation)
SpillStaticRegs();
@@ -352,8 +362,8 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, const Dispatche
// brk = SIGTRAP
// ??? = SIGSEGV
// Force a SIGSEGV by loading zero
LoadConstant(x1, 0);
ldr(x1, MemOperand(x1));
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r1, 0);
ldr(ARMEmitter::XReg::x1, ARMEmitter::Reg::r1);
}
{
@@ -361,19 +371,18 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, const Dispatche
if (config.StaticRegisterAllocation)
SpillStaticRegs();
bind(&ThreadPauseHandler);
ThreadPauseHandlerAddress = GetCursorAddress<uint64_t>();
// We are pausing, this means the frontend should be waiting for this thread to idle
// We will have faulted and jumped to this location at this point
// Call our sleep handler
ldr(x0, &l_CTX);
mov(x1, STATE);
ldr(x2, &l_Sleep);
ldr(ARMEmitter::XReg::x0, &l_CTX);
mov(ARMEmitter::XReg::x1, STATE);
ldr(ARMEmitter::XReg::x2, &l_Sleep);
#ifdef VIXL_SIMULATOR
GenerateIndirectRuntimeCall<void, void *, void *>(x2);
GenerateIndirectRuntimeCall<void, void *, void *>(ARMEmitter::Reg::r2);
#else
blr(x2);
blr(ARMEmitter::Reg::r2);
#endif
PauseReturnInstruction = GetCursorAddress<uint64_t>();
@@ -403,27 +412,27 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, const Dispatche
PushCalleeSavedRegisters();
// First thing we need to move the thread state pointer back in to our register
mov(STATE, x0);
mov(STATE, ARMEmitter::XReg::x0);
// Make sure to adjust the refcounter so we don't clear the cache now
ldr(w2, STATE_PTR(CpuStateFrame, SignalHandlerRefCounter));
add(w2, w2, 1);
str(w2, STATE_PTR(CpuStateFrame, SignalHandlerRefCounter));
ldr(ARMEmitter::WReg::w2, STATE_PTR(CpuStateFrame, SignalHandlerRefCounter));
add(ARMEmitter::Size::i32Bit, ARMEmitter::Reg::r2, ARMEmitter::Reg::r2, 1);
str(ARMEmitter::WReg::w2, STATE_PTR(CpuStateFrame, SignalHandlerRefCounter));
// Now push the callback return trampoline to the guest stack
// Guest will be misaligned because calling a thunk won't correct the guest's stack once we call the callback from the host
LoadConstant(x0, CTX->X86CodeGen.CallbackReturn);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, CTX->X86CodeGen.CallbackReturn);
ldr(x2, STATE_PTR(CpuStateFrame, State.gregs[X86State::REG_RSP]));
sub(x2, x2, 16);
str(x2, STATE_PTR(CpuStateFrame, State.gregs[X86State::REG_RSP]));
ldr(ARMEmitter::XReg::x2, STATE_PTR(CpuStateFrame, State.gregs[X86State::REG_RSP]));
sub(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r2, ARMEmitter::Reg::r2, 16);
str(ARMEmitter::XReg::x2, STATE_PTR(CpuStateFrame, State.gregs[X86State::REG_RSP]));
// Store the trampoline to the guest stack
// Guest stack is now correctly misaligned after a regular call instruction
str(x0, MemOperand(x2));
str(ARMEmitter::XReg::x0, ARMEmitter::Reg::r2, 0);
// Store RIP to the context state
str(x1, STATE_PTR(CpuStateFrame, State.rip));
str(ARMEmitter::XReg::x1, STATE_PTR(CpuStateFrame, State.rip));
// load static regs
if (config.StaticRegisterAllocation)
@@ -436,14 +445,14 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, const Dispatche
{
LUDIVHandlerAddress = GetCursorAddress<uint64_t>();
PushDynamicRegsAndLR();
PushDynamicRegsAndLR(ARMEmitter::Reg::r3);
SpillStaticRegs();
ldr(x3, STATE_PTR(CpuStateFrame, Pointers.AArch64.LUDIV));
ldr(ARMEmitter::XReg::x3, STATE_PTR(CpuStateFrame, Pointers.AArch64.LUDIV));
#ifdef VIXL_SIMULATOR
GenerateIndirectRuntimeCall<uint64_t, uint64_t, uint64_t, uint64_t>(x3);
GenerateIndirectRuntimeCall<uint64_t, uint64_t, uint64_t, uint64_t>(ARMEmitter::Reg::r3);
#else
blr(x3);
blr(ARMEmitter::Reg::r3);
#endif
FillStaticRegs();
@@ -458,14 +467,14 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, const Dispatche
{
LDIVHandlerAddress = GetCursorAddress<uint64_t>();
PushDynamicRegsAndLR();
PushDynamicRegsAndLR(ARMEmitter::Reg::r3);
SpillStaticRegs();
ldr(x3, STATE_PTR(CpuStateFrame, Pointers.AArch64.LDIV));
ldr(ARMEmitter::XReg::x3, STATE_PTR(CpuStateFrame, Pointers.AArch64.LDIV));
#ifdef VIXL_SIMULATOR
GenerateIndirectRuntimeCall<uint64_t, uint64_t, uint64_t, uint64_t>(x3);
GenerateIndirectRuntimeCall<uint64_t, uint64_t, uint64_t, uint64_t>(ARMEmitter::Reg::r3);
#else
blr(x3);
blr(ARMEmitter::Reg::r3);
#endif
FillStaticRegs();
@@ -480,14 +489,14 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, const Dispatche
{
LUREMHandlerAddress = GetCursorAddress<uint64_t>();
PushDynamicRegsAndLR();
PushDynamicRegsAndLR(ARMEmitter::Reg::r3);
SpillStaticRegs();
ldr(x3, STATE_PTR(CpuStateFrame, Pointers.AArch64.LUREM));
ldr(ARMEmitter::XReg::x3, STATE_PTR(CpuStateFrame, Pointers.AArch64.LUREM));
#ifdef VIXL_SIMULATOR
GenerateIndirectRuntimeCall<uint64_t, uint64_t, uint64_t, uint64_t>(x3);
GenerateIndirectRuntimeCall<uint64_t, uint64_t, uint64_t, uint64_t>(ARMEmitter::Reg::r3);
#else
blr(x3);
blr(ARMEmitter::Reg::r3);
#endif
FillStaticRegs();
@@ -502,14 +511,15 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, const Dispatche
{
LREMHandlerAddress = GetCursorAddress<uint64_t>();
PushDynamicRegsAndLR();
PushDynamicRegsAndLR(ARMEmitter::Reg::r3);
SpillStaticRegs();
ldr(x3, STATE_PTR(CpuStateFrame, Pointers.AArch64.LREM));
ldr(ARMEmitter::XReg::x3, STATE_PTR(CpuStateFrame, Pointers.AArch64.LREM));
#ifdef VIXL_SIMULATOR
GenerateIndirectRuntimeCall<uint64_t, uint64_t, uint64_t, uint64_t>(x3);
GenerateIndirectRuntimeCall<uint64_t, uint64_t, uint64_t, uint64_t>(ARMEmitter::Reg::r3);
#else
blr(x3);
blr(ARMEmitter::Reg::r3);
#endif
FillStaticRegs();
@@ -521,16 +531,16 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, const Dispatche
ret();
}
place(&l_CTX);
place(&l_Sleep);
place(&l_CompileBlock);
Bind(&l_CTX);
dc64(reinterpret_cast<uintptr_t>(CTX));
Bind(&l_Sleep);
dc64(reinterpret_cast<uint64_t>(SleepThread));
Bind(&l_CompileBlock);
dc64(GetCompileBlockPtr());
FinalizeCode();
Start = reinterpret_cast<uint64_t>(DispatchPtr);
End = GetCursorAddress<uint64_t>();
vixl::aarch64::CPU::EnsureIAndDCacheCoherency(reinterpret_cast<void*>(DispatchPtr), End - reinterpret_cast<uint64_t>(DispatchPtr));
GetBuffer()->SetExecutable();
ClearICache(reinterpret_cast<void*>(DispatchPtr), End - reinterpret_cast<uint64_t>(DispatchPtr));
if (CTX->Config.BlockJITNaming()) {
std::string Name = "Dispatch_" + std::to_string(FHU::Syscalls::gettid());
@@ -539,106 +549,100 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, const Dispatche
if (CTX->Config.GlobalJITNaming()) {
CTX->Symbols.RegisterJITSpace(reinterpret_cast<void*>(DispatchPtr), End - reinterpret_cast<uint64_t>(DispatchPtr));
}
#ifdef VIXL_DISASSEMBLER
const auto DisasmEnd = GetCursorAddress<const vixl::aarch64::Instruction*>();
Disasm.DisassembleBuffer(DisasmBegin, DisasmEnd);
#endif
}
#ifdef VIXL_SIMULATOR
void Arm64Dispatcher::ExecuteDispatch(FEXCore::Core::CpuStateFrame *Frame) {
Simulator.WriteXRegister(0, reinterpret_cast<int64_t>(Frame));
Simulator.RunFrom(reinterpret_cast<Instruction const*>(DispatchPtr));
Simulator.RunFrom(reinterpret_cast<vixl::aarch64::Instruction const*>(DispatchPtr));
}
void Arm64Dispatcher::ExecuteJITCallback(FEXCore::Core::CpuStateFrame *Frame, uint64_t RIP) {
Simulator.WriteXRegister(0, reinterpret_cast<int64_t>(Frame));
Simulator.WriteXRegister(1, RIP);
Simulator.RunFrom(reinterpret_cast<Instruction const*>(CallbackPtr));
Simulator.RunFrom(reinterpret_cast<vixl::aarch64::Instruction const*>(CallbackPtr));
}
#endif
// Used by GenerateGDBPauseCheck, GenerateInterpreterTrampoline, destination buffer is set before use
static thread_local vixl::aarch64::Assembler emit((uint8_t*)&emit, 1);
size_t Arm64Dispatcher::GenerateGDBPauseCheck(uint8_t *CodeBuffer, uint64_t GuestRIP) {
FEXCore::ARMEmitter::Emitter emit{CodeBuffer, MaxGDBPauseCheckSize};
*emit.GetBuffer() = vixl::CodeBuffer(CodeBuffer, MaxGDBPauseCheckSize);
vixl::CodeBufferCheckScope scope(&emit, MaxGDBPauseCheckSize, vixl::CodeBufferCheckScope::kDontReserveBufferSpace, vixl::CodeBufferCheckScope::kNoAssert);
aarch64::Label RunBlock;
ARMEmitter::ForwardLabel RunBlock;
// If we have a gdb server running then run in a less efficient mode that checks if we need to exit
// This happens when single stepping
static_assert(sizeof(FEXCore::Context::Context::Config.RunningMode) == 4, "This is expected to be size of 4");
emit.ldr(x0, STATE_PTR(CpuStateFrame, Thread)); // Get thread
emit.ldr(x0, MemOperand(x0, offsetof(FEXCore::Core::InternalThreadState, CTX))); // Get Context
emit.ldr(w0, MemOperand(x0, offsetof(FEXCore::Context::Context, Config.RunningMode)));
static_assert(sizeof(FEXCore::Context::ContextImpl::Config.RunningMode) == 4, "This is expected to be size of 4");
emit.ldr(ARMEmitter::XReg::x0, STATE_PTR(CpuStateFrame, Thread));
emit.ldr(ARMEmitter::XReg::x0, ARMEmitter::Reg::r0, offsetof(FEXCore::Core::InternalThreadState, CTX)); // Get Context
emit.ldr(ARMEmitter::WReg::w0, ARMEmitter::Reg::r0, offsetof(FEXCore::Context::ContextImpl, Config.RunningMode));
// If the value == 0 then we don't need to stop
emit.cbz(w0, &RunBlock);
emit.cbz(ARMEmitter::Size::i32Bit, ARMEmitter::Reg::r0, &RunBlock);
{
Literal l_GuestRIP {GuestRIP};
ARMEmitter::ForwardLabel l_GuestRIP;
// Make sure RIP is syncronized to the context
emit.ldr(x0, &l_GuestRIP);
emit.str(x0, STATE_PTR(CpuStateFrame, State.rip));
emit.ldr(ARMEmitter::XReg::x0, &l_GuestRIP);
emit.str(ARMEmitter::XReg::x0, STATE_PTR(CpuStateFrame, State.rip));
// Stop the thread
emit.ldr(x0, STATE_PTR(CpuStateFrame, Pointers.Common.ThreadPauseHandlerSpillSRA));
emit.br(x0);
emit.place(&l_GuestRIP);
emit.ldr(ARMEmitter::XReg::x0, STATE_PTR(CpuStateFrame, Pointers.Common.ThreadPauseHandlerSpillSRA));
emit.br(ARMEmitter::Reg::r0);
emit.Bind(&l_GuestRIP);
emit.dc64(GuestRIP);
}
emit.bind(&RunBlock);
emit.FinalizeCode();
emit.Bind(&RunBlock);
auto UsedBytes = emit.GetBuffer()->GetCursorOffset();
vixl::aarch64::CPU::EnsureIAndDCacheCoherency(CodeBuffer, UsedBytes);
auto UsedBytes = emit.GetCursorOffset();
emit.ClearICache(CodeBuffer, UsedBytes);
return UsedBytes;
}
size_t Arm64Dispatcher::GenerateInterpreterTrampoline(uint8_t *CodeBuffer) {
LOGMAN_THROW_AA_FMT(!config.StaticRegisterAllocation, "GenerateInterpreterTrampoline dispatcher does not support SRA");
*emit.GetBuffer() = vixl::CodeBuffer(CodeBuffer, MaxInterpreterTrampolineSize);
FEXCore::ARMEmitter::Emitter emit{CodeBuffer, MaxInterpreterTrampolineSize};
ARMEmitter::ForwardLabel InlineIRData;
vixl::CodeBufferCheckScope scope(&emit, MaxInterpreterTrampolineSize, vixl::CodeBufferCheckScope::kDontReserveBufferSpace, vixl::CodeBufferCheckScope::kNoAssert);
emit.mov(ARMEmitter::XReg::x0, STATE);
emit.adr(ARMEmitter::Reg::r1, &InlineIRData);
aarch64::Label InlineIRData;
emit.ldr(ARMEmitter::XReg::x3, STATE_PTR(CpuStateFrame, Pointers.Interpreter.FragmentExecuter));
emit.blr(ARMEmitter::Reg::r3);
emit.mov(x0, STATE);
emit.adr(x1, &InlineIRData);
emit.ldr(ARMEmitter::XReg::x0, STATE_PTR(CpuStateFrame, Pointers.Common.DispatcherLoopTop));
emit.br(ARMEmitter::Reg::r0);
emit.ldr(x3, STATE_PTR(CpuStateFrame, Pointers.Interpreter.FragmentExecuter));
emit.blr(x3);
emit.Bind(&InlineIRData);
emit.ldr(x0, STATE_PTR(CpuStateFrame, Pointers.Common.DispatcherLoopTop));
emit.br(x0);
emit.bind(&InlineIRData);
emit.FinalizeCode();
auto UsedBytes = emit.GetBuffer()->GetCursorOffset();
vixl::aarch64::CPU::EnsureIAndDCacheCoherency(CodeBuffer, UsedBytes);
auto UsedBytes = emit.GetCursorOffset();
emit.ClearICache(CodeBuffer, UsedBytes);
return UsedBytes;
}
void Arm64Dispatcher::SpillSRA(FEXCore::Core::InternalThreadState *Thread, void *ucontext, uint32_t IgnoreMask) {
for (size_t i = 0; i < SRA64.size(); i++) {
if (IgnoreMask & (1U << SRA64[i].GetCode())) {
if (IgnoreMask & (1U << SRA64[i].Idx())) {
// Skip this one, it's already spilled
continue;
}
Thread->CurrentFrame->State.gregs[i] = ArchHelpers::Context::GetArmReg(ucontext, SRA64[i].GetCode());
Thread->CurrentFrame->State.gregs[i] = ArchHelpers::Context::GetArmReg(ucontext, SRA64[i].Idx());
}
if (EmitterCTX->HostFeatures.SupportsAVX) {
for (size_t i = 0; i < SRAFPR.size(); i++) {
auto FPR = ArchHelpers::Context::GetArmFPR(ucontext, SRAFPR[i].GetCode());
auto FPR = ArchHelpers::Context::GetArmFPR(ucontext, SRAFPR[i].Idx());
memcpy(&Thread->CurrentFrame->State.xmm.avx.data[i][0], &FPR, sizeof(__uint128_t));
}
} else {
for (size_t i = 0; i < SRAFPR.size(); i++) {
auto FPR = ArchHelpers::Context::GetArmFPR(ucontext, SRAFPR[i].GetCode());
auto FPR = ArchHelpers::Context::GetArmFPR(ucontext, SRAFPR[i].Idx());
memcpy(&Thread->CurrentFrame->State.xmm.sse.data[i][0], &FPR, sizeof(__uint128_t));
}
}
@@ -658,6 +662,7 @@ void Arm64Dispatcher::InitThreadPointers(FEXCore::Core::InternalThreadState *Thr
Common.GuestSignal_SIGTRAP = GuestSignal_SIGTRAP;
Common.GuestSignal_SIGSEGV = GuestSignal_SIGSEGV;
Common.SignalReturnHandler = SignalHandlerReturnAddress;
Common.SignalReturnHandlerRT = SignalHandlerReturnAddressRT;
auto &AArch64 = Thread->CurrentFrame->Pointers.AArch64;
AArch64.LUDIVHandler = LUDIVHandlerAddress;
@@ -667,7 +672,7 @@ void Arm64Dispatcher::InitThreadPointers(FEXCore::Core::InternalThreadState *Thr
}
}
std::unique_ptr<Dispatcher> Dispatcher::CreateArm64(FEXCore::Context::Context *CTX, const DispatcherConfig &Config) {
std::unique_ptr<Dispatcher> Dispatcher::CreateArm64(FEXCore::Context::ContextImpl *CTX, const DispatcherConfig &Config) {
return std::make_unique<Arm64Dispatcher>(CTX, Config);
}
@@ -7,19 +7,18 @@
#include <aarch64/simulator-aarch64.h>
#endif
namespace FEXCore::Context {
struct Context;
}
namespace FEXCore::Core {
struct InternalThreadState;
}
#define STATE_PTR(STATE_TYPE, FIELD) \
STATE.R(), offsetof(FEXCore::Core::STATE_TYPE, FIELD)
namespace FEXCore::CPU {
class Arm64Dispatcher final : public Dispatcher, public Arm64Emitter {
public:
Arm64Dispatcher(FEXCore::Context::Context *ctx, const DispatcherConfig &config);
Arm64Dispatcher(FEXCore::Context::ContextImpl *ctx, const DispatcherConfig &config);
void InitThreadPointers(FEXCore::Core::InternalThreadState *Thread) override;
size_t GenerateGDBPauseCheck(uint8_t *CodeBuffer, uint64_t GuestRIP) override;
size_t GenerateInterpreterTrampoline(uint8_t *CodeBuffer) override;
@@ -29,6 +28,8 @@ class Arm64Dispatcher final : public Dispatcher, public Arm64Emitter {
void ExecuteJITCallback(FEXCore::Core::CpuStateFrame *Frame, uint64_t RIP) override;
#endif
void EmitDispatcher();
protected:
void SpillSRA(FEXCore::Core::InternalThreadState *Thread, void *ucontext, uint32_t IgnoreMask) override;
File diff suppressed because it is too large. Load diff
@@ -20,7 +20,7 @@ struct InternalThreadState;
}
namespace FEXCore::Context {
struct Context;
class ContextImpl;
}
namespace FEXCore::CPU {
@@ -44,6 +44,7 @@ public:
uint64_t ThreadPauseHandlerAddressSpillSRA{};
uint64_t ExitFunctionLinkerAddress{};
uint64_t SignalHandlerReturnAddress{};
uint64_t SignalHandlerReturnAddressRT{};
uint64_t GuestSignal_SIGILL{};
uint64_t GuestSignal_SIGTRAP{};
uint64_t GuestSignal_SIGSEGV{};
@@ -73,8 +74,8 @@ public:
virtual size_t GenerateGDBPauseCheck(uint8_t *CodeBuffer, uint64_t GuestRIP) = 0;
virtual size_t GenerateInterpreterTrampoline(uint8_t *CodeBuffer) = 0;
static std::unique_ptr<Dispatcher> CreateX86(FEXCore::Context::Context *CTX, const DispatcherConfig &Config);
static std::unique_ptr<Dispatcher> CreateArm64(FEXCore::Context::Context *CTX, const DispatcherConfig &Config);
static std::unique_ptr<Dispatcher> CreateX86(FEXCore::Context::ContextImpl *CTX, const DispatcherConfig &Config);
static std::unique_ptr<Dispatcher> CreateArm64(FEXCore::Context::ContextImpl *CTX, const DispatcherConfig &Config);
virtual void ExecuteDispatch(FEXCore::Core::CpuStateFrame *Frame) {
DispatchPtr(Frame);
@@ -85,21 +86,83 @@ public:
}
protected:
Dispatcher(FEXCore::Context::Context *ctx, const DispatcherConfig &Config)
Dispatcher(FEXCore::Context::ContextImpl *ctx, const DispatcherConfig &Config)
: CTX {ctx}
, config {Config}
{}
uint64_t ReconstructRIPFromContext(FEXCore::Core::CpuStateFrame *Frame, void *ucontext) const;
void RestoreFrame_x64(ArchHelpers::Context::ContextBackup* Context, FEXCore::Core::CpuStateFrame *Frame, void *ucontext);
void RestoreFrame_ia32(ArchHelpers::Context::ContextBackup* Context, FEXCore::Core::CpuStateFrame *Frame, void *ucontext);
void RestoreRTFrame_ia32(ArchHelpers::Context::ContextBackup* Context, FEXCore::Core::CpuStateFrame *Frame, void *ucontext);
///< Setup the signal frame for x64.
uint64_t SetupFrame_x64(FEXCore::Core::InternalThreadState *Thread, ArchHelpers::Context::ContextBackup* ContextBackup, FEXCore::Core::CpuStateFrame *Frame,
int Signal, siginfo_t *HostSigInfo, void *ucontext,
GuestSigAction *GuestAction, stack_t *GuestStack,
uint64_t NewGuestSP, const uint32_t eflags);
///< Setup the signal frame for a 32-bit signal without SA_SIGINFO.
uint64_t SetupFrame_ia32(ArchHelpers::Context::ContextBackup* ContextBackup, FEXCore::Core::CpuStateFrame *Frame,
int Signal, siginfo_t *HostSigInfo, void *ucontext,
GuestSigAction *GuestAction, stack_t *GuestStack,
uint64_t NewGuestSP, const uint32_t eflags);
///< Setup the signal frame for a 32-bit signal with SA_SIGINFO.
uint64_t SetupRTFrame_ia32(ArchHelpers::Context::ContextBackup* ContextBackup, FEXCore::Core::CpuStateFrame *Frame,
int Signal, siginfo_t *HostSigInfo, void *ucontext,
GuestSigAction *GuestAction, stack_t *GuestStack,
uint64_t NewGuestSP, const uint32_t eflags);
ArchHelpers::Context::ContextBackup* StoreThreadState(FEXCore::Core::InternalThreadState *Thread, int Signal, void *ucontext);
void RestoreThreadState(FEXCore::Core::InternalThreadState *Thread, void *ucontext);
enum class RestoreType {
TYPE_REALTIME, ///< Signal restore type is from a `realtime` signal.
TYPE_NONREALTIME, ///< Signal restore type is from a `non-realtime` signal.
TYPE_PAUSE, ///< Signal restore type is from a GDB pause event.
};
/*
* Signal frames on 32-bit architecture needs to match exactly how the kernel generates the frame.
* This is because large parts of the signal frame definition is part of the UAPI.
* This means that when FEX sets up the signal frame, it needs to match the UAPI stack setup.
*
* The two signal stack frame types below describe the two different 32-bit frame types.
*/
// The 32-bit non-realtime signal frame.
// This frame type is used when the guest signal is used without the `SA_SIGINFO` flag.
struct SigFrame_i32 {
uint32_t pretcode; ///< sigreturn return branch point.
int32_t Signal; ///< The signal hit.
FEXCore::x86::sigcontext sc; ///< The signal context.
x86::_libc_fpstate fpstate_unused; ///< Unused fpstate. Retained for backwards compatibility.
uint32_t extramask[1]; ///< Upper 32-bits of the signal mask. Lower 32-bits is in the sigcontext.
char retcode[8]; ///< Unused but needs to be filled. GDB seemingly uses as a debug marker.
///< FP state now follows after this.
};
// The 32-bit realtime signal frame.
// This frame type is used when the guest signal is used with the `SA_SIGINFO` flag.
struct RTSigFrame_i32 {
uint32_t pretcode; ///< sigreturn return branch point.
int32_t Signal; ///< The signal hit.
uint32_t pinfo; ///< Pointer to siginfo_t
uint32_t puc; ///< Pointer to ucontext_t
FEXCore::x86::siginfo_t info;
FEXCore::x86::ucontext_t uc;
char retcode[8]; ///< Unused but needs to be filled. GDB seemingly uses as a debug marker.
///< FP state now follows after this.
};
void RestoreThreadState(FEXCore::Core::InternalThreadState *Thread, void *ucontext, RestoreType Type);
std::stack<uint64_t, std::vector<uint64_t>> SignalFrames;
virtual void SpillSRA(FEXCore::Core::InternalThreadState *Thread, void *ucontext, uint32_t IgnoreMask) {}
FEXCore::Context::Context *CTX;
FEXCore::Context::ContextImpl *CTX;
DispatcherConfig config;
static void SleepThread(FEXCore::Context::Context *ctx, FEXCore::Core::CpuStateFrame *Frame);
static void SleepThread(FEXCore::Context::ContextImpl *ctx, FEXCore::Core::CpuStateFrame *Frame);
static uint64_t GetCompileBlockPtr();
@@ -27,7 +27,7 @@ namespace FEXCore::CPU {
static constexpr size_t MAX_DISPATCHER_CODE_SIZE = 4096;
#define STATE r14
X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, const DispatcherConfig &config)
X86Dispatcher::X86Dispatcher(FEXCore::Context::ContextImpl *ctx, const DispatcherConfig &config)
: Dispatcher(ctx, config)
, Xbyak::CodeGenerator(MAX_DISPATCHER_CODE_SIZE,
FEXCore::Allocator::mmap(nullptr, MAX_DISPATCHER_CODE_SIZE, PROT_READ | PROT_WRITE | PROT_EXEC, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0),
@@ -344,6 +344,12 @@ X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, const DispatcherCon
ud2();
}
{
// RT Signal return handler
SignalHandlerReturnAddressRT = getCurr<uint64_t>();
ud2();
}
{
// Guest SIGILL handler
// Needs to be distinct from the SignalHandlerReturnAddress
@@ -427,7 +433,7 @@ size_t X86Dispatcher::GenerateGDBPauseCheck(uint8_t *CodeBuffer, uint64_t GuestR
emit.mov(rax, reinterpret_cast<uint64_t>(CTX));
// If the value == 0 then we don't need to stop
emit.cmp(dword [rax + (offsetof(FEXCore::Context::Context, Config.RunningMode))], 0);
emit.cmp(dword [rax + (offsetof(FEXCore::Context::ContextImpl, Config.RunningMode))], 0);
emit.je(RunBlock);
{
// Make sure RIP is syncronized to the context
@@ -486,13 +492,14 @@ void X86Dispatcher::InitThreadPointers(FEXCore::Core::InternalThreadState *Threa
Common.GuestSignal_SIGTRAP = GuestSignal_SIGTRAP;
Common.GuestSignal_SIGSEGV = GuestSignal_SIGSEGV;
Common.SignalReturnHandler = SignalHandlerReturnAddress;
Common.SignalReturnHandlerRT = SignalHandlerReturnAddressRT;
auto &Interpreter = Thread->CurrentFrame->Pointers.Interpreter;
(uintptr_t&)Interpreter.CallbackReturn = IntCallbackReturnAddress;
}
}
std::unique_ptr<Dispatcher> Dispatcher::CreateX86(FEXCore::Context::Context *CTX, const DispatcherConfig &Config) {
std::unique_ptr<Dispatcher> Dispatcher::CreateX86(FEXCore::Context::ContextImpl *CTX, const DispatcherConfig &Config) {
return std::make_unique<X86Dispatcher>(CTX, Config);
}
@@ -5,10 +5,6 @@
#define XBYAK64
#include <xbyak/xbyak.h>
namespace FEXCore::Context {
struct Context;
}
namespace FEXCore::Core {
struct InternalThreadState;
}
@@ -17,7 +13,7 @@ namespace FEXCore::CPU {
class X86Dispatcher final : public Dispatcher, public Xbyak::CodeGenerator {
public:
X86Dispatcher(FEXCore::Context::Context *ctx, const DispatcherConfig &config);
X86Dispatcher(FEXCore::Context::ContextImpl *ctx, const DispatcherConfig &config);
void InitThreadPointers(FEXCore::Core::InternalThreadState *Thread) override;
size_t GenerateGDBPauseCheck(uint8_t *CodeBuffer, uint64_t GuestRIP) override;
size_t GenerateInterpreterTrampoline(uint8_t *CodeBuffer) override;
+172 -247
View File
@@ -32,26 +32,6 @@ using namespace FEXCore::X86Tables;
static uint32_t MapModRMToReg(uint8_t REX, uint8_t bits, bool HighBits, bool HasREX, bool HasXMM, bool HasMM, uint8_t InvalidOffset = 16) {
using GPRArray = std::array<uint32_t, 16>;
static constexpr GPRArray GPRIndexes = {
// Classical ordering?
FEXCore::X86State::REG_RAX,
FEXCore::X86State::REG_RCX,
FEXCore::X86State::REG_RDX,
FEXCore::X86State::REG_RBX,
FEXCore::X86State::REG_RSP,
FEXCore::X86State::REG_RBP,
FEXCore::X86State::REG_RSI,
FEXCore::X86State::REG_RDI,
FEXCore::X86State::REG_R8,
FEXCore::X86State::REG_R9,
FEXCore::X86State::REG_R10,
FEXCore::X86State::REG_R11,
FEXCore::X86State::REG_R12,
FEXCore::X86State::REG_R13,
FEXCore::X86State::REG_R14,
FEXCore::X86State::REG_R15,
};
static constexpr GPRArray GPR8BitHighIndexes = {
// Classical ordering?
FEXCore::X86State::REG_RAX,
@@ -72,112 +52,34 @@ static uint32_t MapModRMToReg(uint8_t REX, uint8_t bits, bool HighBits, bool Has
FEXCore::X86State::REG_R15,
};
static constexpr GPRArray XMMIndexes = {
FEXCore::X86State::REG_XMM_0,
FEXCore::X86State::REG_XMM_1,
FEXCore::X86State::REG_XMM_2,
FEXCore::X86State::REG_XMM_3,
FEXCore::X86State::REG_XMM_4,
FEXCore::X86State::REG_XMM_5,
FEXCore::X86State::REG_XMM_6,
FEXCore::X86State::REG_XMM_7,
FEXCore::X86State::REG_XMM_8,
FEXCore::X86State::REG_XMM_9,
FEXCore::X86State::REG_XMM_10,
FEXCore::X86State::REG_XMM_11,
FEXCore::X86State::REG_XMM_12,
FEXCore::X86State::REG_XMM_13,
FEXCore::X86State::REG_XMM_14,
FEXCore::X86State::REG_XMM_15,
};
static constexpr GPRArray MMIndexes = {
FEXCore::X86State::REG_MM_0,
FEXCore::X86State::REG_MM_1,
FEXCore::X86State::REG_MM_2,
FEXCore::X86State::REG_MM_3,
FEXCore::X86State::REG_MM_4,
FEXCore::X86State::REG_MM_5,
FEXCore::X86State::REG_MM_6,
FEXCore::X86State::REG_MM_7,
FEXCore::X86State::REG_INVALID,
FEXCore::X86State::REG_INVALID,
FEXCore::X86State::REG_INVALID,
FEXCore::X86State::REG_INVALID,
FEXCore::X86State::REG_INVALID,
FEXCore::X86State::REG_INVALID,
FEXCore::X86State::REG_INVALID,
FEXCore::X86State::REG_INVALID
};
const GPRArray *GPRs = &GPRIndexes;
if (HasXMM) {
GPRs = &XMMIndexes;
}
else if (HasMM) {
GPRs = &MMIndexes;
}
else if (HighBits && !HasREX) {
GPRs = &GPR8BitHighIndexes;
}
uint8_t Offset = (REX << 3) | bits;
if (Offset == InvalidOffset) {
return FEXCore::X86State::REG_INVALID;
}
return (*GPRs)[(REX << 3) | bits];
if (HasXMM) {
return FEXCore::X86State::REG_XMM_0 + Offset;
}
else if (HasMM) {
return FEXCore::X86State::REG_MM_0 + Offset;
}
else if (!(HighBits && !HasREX)) {
return FEXCore::X86State::REG_RAX + Offset;
}
return GPR8BitHighIndexes[Offset];
}
static uint32_t MapVEXToReg(uint8_t vvvv, bool HasXMM) {
using GPRArray = std::array<uint32_t, 16>;
static constexpr GPRArray GPRIndexes = {
FEXCore::X86State::REG_RAX,
FEXCore::X86State::REG_RCX,
FEXCore::X86State::REG_RDX,
FEXCore::X86State::REG_RBX,
FEXCore::X86State::REG_RSP,
FEXCore::X86State::REG_RBP,
FEXCore::X86State::REG_RSI,
FEXCore::X86State::REG_RDI,
FEXCore::X86State::REG_R8,
FEXCore::X86State::REG_R9,
FEXCore::X86State::REG_R10,
FEXCore::X86State::REG_R11,
FEXCore::X86State::REG_R12,
FEXCore::X86State::REG_R13,
FEXCore::X86State::REG_R14,
FEXCore::X86State::REG_R15,
};
static constexpr GPRArray XMMIndexes = {
FEXCore::X86State::REG_XMM_0,
FEXCore::X86State::REG_XMM_1,
FEXCore::X86State::REG_XMM_2,
FEXCore::X86State::REG_XMM_3,
FEXCore::X86State::REG_XMM_4,
FEXCore::X86State::REG_XMM_5,
FEXCore::X86State::REG_XMM_6,
FEXCore::X86State::REG_XMM_7,
FEXCore::X86State::REG_XMM_8,
FEXCore::X86State::REG_XMM_9,
FEXCore::X86State::REG_XMM_10,
FEXCore::X86State::REG_XMM_11,
FEXCore::X86State::REG_XMM_12,
FEXCore::X86State::REG_XMM_13,
FEXCore::X86State::REG_XMM_14,
FEXCore::X86State::REG_XMM_15,
};
if (HasXMM) {
return XMMIndexes[vvvv];
return FEXCore::X86State::REG_XMM_0 + vvvv;
} else {
return GPRIndexes[vvvv];
return FEXCore::X86State::REG_RAX + vvvv;
}
}
Decoder::Decoder(FEXCore::Context::Context *ctx)
Decoder::Decoder(FEXCore::Context::ContextImpl *ctx)
: CTX {ctx}
, OSABI { ctx->SyscallHandler ? ctx->SyscallHandler->GetOSABI() : FEXCore::HLE::SyscallOSABI::OS_UNKNOWN }
, PoolObject {ctx->FrontendAllocator, sizeof(FEXCore::X86Tables::DecodedInst) * DefaultDecodedBufferSize} {
@@ -206,7 +108,7 @@ uint64_t Decoder::ReadData(uint8_t Size) {
uint64_t Res = 0;
std::memcpy(&Res, &InstStream[InstructionSize], Size);
#ifndef NDEBUG
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
for(size_t i = 0; i < Size; ++i) {
ReadByte();
}
@@ -384,12 +286,6 @@ bool Decoder::NormalOp(FEXCore::X86Tables::X86InstInfo const *Info, uint16_t Op,
DecodeInst->OP = Op;
DecodeInst->TableInfo = Info;
// XXX: Once we support 32bit x86 then this will be necessary to support
if (Info->Type == FEXCore::X86Tables::TYPE_LEGACY_PREFIX) {
LogMan::Msg::DFmt("Legacy Prefix");
return false;
}
if (Info->Type == FEXCore::X86Tables::TYPE_UNKNOWN) {
LogMan::Msg::DFmt("Unknown instruction: {} 0x{:04x} 0x{:x}", Info->Name ?: "UND", Op, DecodeInst->PC);
return false;
@@ -425,10 +321,17 @@ bool Decoder::NormalOp(FEXCore::X86Tables::X86InstInfo const *Info, uint16_t Op,
const bool HasMODRM = !!(Info->Flags & FEXCore::X86Tables::InstFlags::FLAGS_MODRM);
const bool HasREX = !!(DecodeInst->Flags & DecodeFlags::FLAG_REX_PREFIX);
const bool HasHighXMM = HAS_XMM_SUBFLAG(Info->Flags, FEXCore::X86Tables::InstFlags::FLAGS_SF_HIGH_XMM_REG);
const bool Has16BitAddressing = !CTX->Config.Is64BitMode &&
DecodeInst->Flags & DecodeFlags::FLAG_ADDRESS_SIZE;
// This is used for ModRM register modification
// For both modrm.reg and modrm.rm(when mod == 0b11) when value is >= 0b100
// then it changes from expected registers to the high 8bits of the lower registers
// Bit annoying to support
// In the case of no modrm (REX in byte situation) then it is unaffected
bool Is8BitSrc{};
bool Is8BitDest{};
// If we require ModRM and haven't decoded it yet, do it now
// Some instructions have to read modrm upfront, others do it later
if (HasMODRM && !DecodeInst->DecodedModRM) {
@@ -445,6 +348,7 @@ bool Decoder::NormalOp(FEXCore::X86Tables::X86InstInfo const *Info, uint16_t Op,
if (DstSizeFlag == FEXCore::X86Tables::InstFlags::SIZE_8BIT) {
DecodeInst->Flags |= DecodeFlags::GenSizeDstSize(DecodeFlags::SIZE_8BIT);
DestSize = 1;
Is8BitDest = true;
}
else if (DstSizeFlag == FEXCore::X86Tables::InstFlags::SIZE_16BIT) {
DecodeInst->Flags |= DecodeFlags::GenSizeDstSize(DecodeFlags::SIZE_16BIT);
@@ -459,6 +363,10 @@ bool Decoder::NormalOp(FEXCore::X86Tables::X86InstInfo const *Info, uint16_t Op,
DestSize = 16;
}
}
else if (DstSizeFlag == FEXCore::X86Tables::InstFlags::SIZE_256BIT) {
DecodeInst->Flags |= DecodeFlags::GenSizeDstSize(DecodeFlags::SIZE_256BIT);
DestSize = 32;
}
else if (HasNarrowingDisplacement &&
(DstSizeFlag == FEXCore::X86Tables::InstFlags::SIZE_DEF ||
DstSizeFlag == FEXCore::X86Tables::InstFlags::SIZE_64BITDEF)) {
@@ -483,6 +391,7 @@ bool Decoder::NormalOp(FEXCore::X86Tables::X86InstInfo const *Info, uint16_t Op,
// Decode sources
if (SrcSizeFlag == FEXCore::X86Tables::InstFlags::SIZE_8BIT) {
DecodeInst->Flags |= DecodeFlags::GenSizeSrcSize(DecodeFlags::SIZE_8BIT);
Is8BitSrc = true;
}
else if (SrcSizeFlag == FEXCore::X86Tables::InstFlags::SIZE_16BIT) {
DecodeInst->Flags |= DecodeFlags::GenSizeSrcSize(DecodeFlags::SIZE_16BIT);
@@ -516,14 +425,6 @@ bool Decoder::NormalOp(FEXCore::X86Tables::X86InstInfo const *Info, uint16_t Op,
}
}
// This is used for ModRM register modification
// For both modrm.reg and modrm.rm(when mod == 0b11) when value is >= 0b100
// then it changes from expected registers to the high 8bits of the lower registers
// Bit annoying to support
// In the case of no modrm (REX in byte situation) then it is unaffected
const bool Is8BitSrc = (DecodeFlags::GetSizeSrcFlags(DecodeInst->Flags) == DecodeFlags::SIZE_8BIT);
const bool Is8BitDest = (DecodeFlags::GetSizeDstFlags(DecodeInst->Flags) == DecodeFlags::SIZE_8BIT);
auto *CurrentDest = &DecodeInst->Dest;
if (HAS_NON_XMM_SUBFLAG(Info->Flags, FEXCore::X86Tables::InstFlags::FLAGS_SF_DST_RAX) ||
@@ -534,8 +435,7 @@ bool Decoder::NormalOp(FEXCore::X86Tables::X86InstInfo const *Info, uint16_t Op,
CurrentDest->Data.GPR.GPR = HAS_NON_XMM_SUBFLAG(Info->Flags, FEXCore::X86Tables::InstFlags::FLAGS_SF_DST_RAX) ? FEXCore::X86State::REG_RAX : FEXCore::X86State::REG_RDX;
CurrentDest = &DecodeInst->Src[0];
}
if (HAS_NON_XMM_SUBFLAG(Info->Flags, FEXCore::X86Tables::InstFlags::FLAGS_SF_REX_IN_BYTE)) {
else if (HAS_NON_XMM_SUBFLAG(Info->Flags, FEXCore::X86Tables::InstFlags::FLAGS_SF_REX_IN_BYTE)) {
LOGMAN_THROW_AA_FMT(!HasMODRM, "This instruction shouldn't have ModRM!");
// If the REX is in the byte that means the lower nibble of the OP contains the destination GPR
@@ -543,7 +443,7 @@ bool Decoder::NormalOp(FEXCore::X86Tables::X86InstInfo const *Info, uint16_t Op,
// ADDITIONALLY:
// If there is a REX prefix then that allows extended GPR usage
CurrentDest->Type = DecodedOperand::OpType::GPR;
DecodeInst->Dest.Data.GPR.HighBits = (Is8BitDest && !HasREX && (Op & 0b111) >= 0b100) || HasHighXMM;
DecodeInst->Dest.Data.GPR.HighBits = (Is8BitDest && !HasREX && (Op & 0b111) >= 0b100);
CurrentDest->Data.GPR.GPR = MapModRMToReg(DecodeInst->Flags & DecodeFlags::FLAG_REX_XGPR_B ? 1 : 0, Op & 0b111, Is8BitDest, HasREX, false, false);
if (CurrentDest->Data.GPR.GPR == FEXCore::X86State::REG_INVALID)
@@ -571,7 +471,7 @@ bool Decoder::NormalOp(FEXCore::X86Tables::X86InstInfo const *Info, uint16_t Op,
// Decode the GPR source first
GPR.Type = DecodedOperand::OpType::GPR;
GPR.Data.GPR.HighBits = (GPR8Bit && ModRM.reg >= 0b100 && !HasREX) || HasHighXMM;
GPR.Data.GPR.HighBits = (GPR8Bit && ModRM.reg >= 0b100 && !HasREX);
GPR.Data.GPR.GPR = MapModRMToReg(DecodeInst->Flags & DecodeFlags::FLAG_REX_XGPR_R ? 1 : 0, ModRM.reg, GPR8Bit, HasREX, HasXMMGPR, HasMMGPR);
if (GPR.Data.GPR.GPR == FEXCore::X86State::REG_INVALID)
@@ -581,7 +481,7 @@ bool Decoder::NormalOp(FEXCore::X86Tables::X86InstInfo const *Info, uint16_t Op,
// ModRM.Mod != 0b11 == Register-direct addressing
if (ModRM.mod == 0b11) {
NonGPR.Type = DecodedOperand::OpType::GPR;
NonGPR.Data.GPR.HighBits = (NonGPR8Bit && ModRM.rm >= 0b100 && !HasREX) || HasHighXMM;
NonGPR.Data.GPR.HighBits = (NonGPR8Bit && ModRM.rm >= 0b100 && !HasREX);
NonGPR.Data.GPR.GPR = MapModRMToReg(DecodeInst->Flags & DecodeFlags::FLAG_REX_XGPR_B ? 1 : 0, ModRM.rm, NonGPR8Bit, HasREX, HasXMMNonGPR, HasMMNonGPR);
if (NonGPR.Data.GPR.GPR == FEXCore::X86State::REG_INVALID)
return false;
@@ -680,12 +580,6 @@ bool Decoder::NormalOpHeader(FEXCore::X86Tables::X86InstInfo const *Info, uint16
DecodeInst->OP = Op;
DecodeInst->TableInfo = Info;
// XXX: Once we support 32bit x86 then this will be necessary to support
if (Info->Type == FEXCore::X86Tables::TYPE_LEGACY_PREFIX) {
LogMan::Msg::DFmt("Legacy Prefix");
return false;
}
if (Info->Type == FEXCore::X86Tables::TYPE_UNKNOWN) {
LogMan::Msg::DFmt("Unknown instruction: {} 0x{:04x} 0x{:x}", Info->Name ?: "UND", Op, DecodeInst->PC);
return false;
@@ -699,7 +593,11 @@ bool Decoder::NormalOpHeader(FEXCore::X86Tables::X86InstInfo const *Info, uint16
LOGMAN_THROW_AA_FMT(Info->Type != FEXCore::X86Tables::TYPE_REX_PREFIX,
"REX PREFIX should have been decoded before this!");
if (Info->Type >= FEXCore::X86Tables::TYPE_GROUP_1 &&
// A normal instruction is the most likely.
if (Info->Type == FEXCore::X86Tables::TYPE_INST) [[likely]] {
return NormalOp(Info, Op);
}
else if (Info->Type >= FEXCore::X86Tables::TYPE_GROUP_1 &&
Info->Type <= FEXCore::X86Tables::TYPE_GROUP_11) {
uint8_t ModRMByte = ReadByte();
DecodeInst->ModRM = ModRMByte;
@@ -847,7 +745,8 @@ bool Decoder::NormalOpHeader(FEXCore::X86Tables::X86InstInfo const *Info, uint16
return NormalOp(&EVEXTableOps[EVEXOp], EVEXOp);
}
return NormalOp(Info, Op);
LOGMAN_MSG_A_FMT("Invalid instruction decoding type");
FEX_UNREACHABLE;
}
bool Decoder::DecodeInstruction(uint64_t PC) {
@@ -866,105 +765,106 @@ bool Decoder::DecodeInstruction(uint64_t PC) {
case 0x0F: {// Escape Op
uint8_t EscapeOp = ReadByte();
switch (EscapeOp) {
case 0x0F: [[unlikely]] { // 3DNow!
// 3DNow! Instruction Encoding: 0F 0F [ModRM] [SIB] [Displacement] [Opcode]
// Decode ModRM
uint8_t ModRMByte = ReadByte();
DecodeInst->ModRM = ModRMByte;
DecodeInst->DecodedModRM = true;
case 0x0F: [[unlikely]] { // 3DNow!
// 3DNow! Instruction Encoding: 0F 0F [ModRM] [SIB] [Displacement] [Opcode]
// Decode ModRM
uint8_t ModRMByte = ReadByte();
DecodeInst->ModRM = ModRMByte;
DecodeInst->DecodedModRM = true;
FEXCore::X86Tables::ModRMDecoded ModRM;
ModRM.Hex = DecodeInst->ModRM;
FEXCore::X86Tables::ModRMDecoded ModRM;
ModRM.Hex = DecodeInst->ModRM;
const bool Has16BitAddressing = !CTX->Config.Is64BitMode &&
DecodeInst->Flags & DecodeFlags::FLAG_ADDRESS_SIZE;
const bool Has16BitAddressing = !CTX->Config.Is64BitMode &&
DecodeInst->Flags & DecodeFlags::FLAG_ADDRESS_SIZE;
// All 3DNow! instructions have the second argument as the rm handler
// We need to decode it upfront to get the displacement out of the way
if (ModRM.mod != 0b11) {
auto Disp = DecodeModRMs_Disp[Has16BitAddressing];
(this->*Disp)(&DecodeInst->Src[0], ModRM);
// All 3DNow! instructions have the second argument as the rm handler
// We need to decode it upfront to get the displacement out of the way
if (ModRM.mod != 0b11) {
auto Disp = DecodeModRMs_Disp[Has16BitAddressing];
(this->*Disp)(&DecodeInst->Src[0], ModRM);
}
// Take a peek at the op just past the displacement
uint8_t LocalOp = ReadByte();
return NormalOpHeader(&FEXCore::X86Tables::DDDNowOps[LocalOp], LocalOp);
break;
}
case 0x38: { // F38 Table!
constexpr uint16_t PF_38_NONE = 0;
constexpr uint16_t PF_38_66 = (1U << 0);
constexpr uint16_t PF_38_F2 = (1U << 1);
constexpr uint16_t PF_38_F3 = (1U << 2);
// Take a peek at the op just past the displacement
uint8_t LocalOp = ReadByte();
return NormalOpHeader(&FEXCore::X86Tables::DDDNowOps[LocalOp], LocalOp);
break;
}
case 0x38: { // F38 Table!
constexpr uint16_t PF_38_NONE = 0;
constexpr uint16_t PF_38_66 = (1U << 0);
constexpr uint16_t PF_38_F2 = (1U << 1);
constexpr uint16_t PF_38_F3 = (1U << 2);
uint16_t Prefix = PF_38_NONE;
if (DecodeInst->Flags & DecodeFlags::FLAG_OPERAND_SIZE) {
Prefix |= PF_38_66;
}
if (DecodeInst->Flags & DecodeFlags::FLAG_REPNE_PREFIX) {
Prefix |= PF_38_F2;
}
if (DecodeInst->Flags & DecodeFlags::FLAG_REP_PREFIX) {
Prefix |= PF_38_F3;
}
uint16_t Prefix = PF_38_NONE;
if (DecodeInst->Flags & DecodeFlags::FLAG_OPERAND_SIZE) {
Prefix |= PF_38_66;
}
if (DecodeInst->Flags & DecodeFlags::FLAG_REPNE_PREFIX) {
Prefix |= PF_38_F2;
}
if (DecodeInst->Flags & DecodeFlags::FLAG_REP_PREFIX) {
Prefix |= PF_38_F3;
uint16_t LocalOp = (Prefix << 8) | ReadByte();
return NormalOpHeader(&FEXCore::X86Tables::H0F38TableOps[LocalOp], LocalOp);
break;
}
case 0x3A: { // F3A Table!
constexpr uint16_t PF_3A_NONE = 0;
constexpr uint16_t PF_3A_66 = (1 << 0);
constexpr uint16_t PF_3A_REX = (1 << 1);
uint16_t LocalOp = (Prefix << 8) | ReadByte();
return NormalOpHeader(&FEXCore::X86Tables::H0F38TableOps[LocalOp], LocalOp);
break;
}
case 0x3A: { // F3A Table!
constexpr uint16_t PF_3A_NONE = 0;
constexpr uint16_t PF_3A_66 = (1 << 0);
constexpr uint16_t PF_3A_REX = (1 << 1);
uint16_t Prefix = PF_3A_NONE;
if (DecodeInst->LastEscapePrefix == 0x66) // Operand Size
Prefix = PF_3A_66;
uint16_t Prefix = PF_3A_NONE;
if (DecodeInst->LastEscapePrefix == 0x66) // Operand Size
Prefix = PF_3A_66;
if (DecodeInst->Flags & DecodeFlags::FLAG_REX_WIDENING)
Prefix |= PF_3A_REX;
if (DecodeInst->Flags & DecodeFlags::FLAG_REX_WIDENING)
Prefix |= PF_3A_REX;
uint16_t LocalOp = (Prefix << 8) | ReadByte();
return NormalOpHeader(&FEXCore::X86Tables::H0F3ATableOps[LocalOp], LocalOp);
break;
}
default: [[likely]] { // Two byte table!
// x86-64 abuses three legacy prefixes to extend the table encodings
// 0x66 - Operand Size prefix
// 0xF2 - REPNE prefix
// 0xF3 - REP prefix
// If any of these three prefixes are used then it falls down the subtable
// Additionally: If you hit repeat of differnt prefixes then only the LAST one before this one works for subtable selection
uint16_t LocalOp = (Prefix << 8) | ReadByte();
return NormalOpHeader(&FEXCore::X86Tables::H0F3ATableOps[LocalOp], LocalOp);
break;
}
default: // Two byte table!
// x86-64 abuses three legacy prefixes to extend the table encodings
// 0x66 - Operand Size prefix
// 0xF2 - REPNE prefix
// 0xF3 - REP prefix
// If any of these three prefixes are used then it falls down the subtable
// Additionally: If you hit repeat of differnt prefixes then only the LAST one before this one works for subtable selection
bool NoOverlay = (FEXCore::X86Tables::SecondBaseOps[EscapeOp].Flags & InstFlags::FLAGS_NO_OVERLAY) != 0;
bool NoOverlay66 = (FEXCore::X86Tables::SecondBaseOps[EscapeOp].Flags & InstFlags::FLAGS_NO_OVERLAY66) != 0;
bool NoOverlay = (FEXCore::X86Tables::SecondBaseOps[EscapeOp].Flags & InstFlags::FLAGS_NO_OVERLAY) != 0;
bool NoOverlay66 = (FEXCore::X86Tables::SecondBaseOps[EscapeOp].Flags & InstFlags::FLAGS_NO_OVERLAY66) != 0;
if (NoOverlay) { // This section of the table ignores prefix extention
return NormalOpHeader(&FEXCore::X86Tables::SecondBaseOps[EscapeOp], EscapeOp);
if (NoOverlay) { // This section of the table ignores prefix extention
return NormalOpHeader(&FEXCore::X86Tables::SecondBaseOps[EscapeOp], EscapeOp);
}
else if (DecodeInst->LastEscapePrefix == 0xF3) { // REP
// Remove prefix so it doesn't effect calculations.
// This is only an escape prefix rather tan modifier now
DecodeInst->Flags &= ~DecodeFlags::FLAG_REP_PREFIX;
return NormalOpHeader(&FEXCore::X86Tables::RepModOps[EscapeOp], EscapeOp);
}
else if (DecodeInst->LastEscapePrefix == 0xF2) { // REPNE
// Remove prefix so it doesn't effect calculations.
// This is only an escape prefix rather tan modifier now
DecodeInst->Flags &= ~DecodeFlags::FLAG_REPNE_PREFIX;
return NormalOpHeader(&FEXCore::X86Tables::RepNEModOps[EscapeOp], EscapeOp);
}
else if (DecodeInst->LastEscapePrefix == 0x66 && !NoOverlay66) { // Operand Size
// Remove prefix so it doesn't effect calculations.
// This is only an escape prefix rather tan modifier now
DecodeInst->Flags &= ~DecodeFlags::FLAG_OPERAND_SIZE;
DecodeFlags::PopOpAddrIf(&DecodeInst->Flags, DecodeFlags::FLAG_OPERAND_SIZE_LAST);
return NormalOpHeader(&FEXCore::X86Tables::OpSizeModOps[EscapeOp], EscapeOp);
}
else {
return NormalOpHeader(&FEXCore::X86Tables::SecondBaseOps[EscapeOp], EscapeOp);
}
break;
}
else if (DecodeInst->LastEscapePrefix == 0xF3) { // REP
// Remove prefix so it doesn't effect calculations.
// This is only an escape prefix rather tan modifier now
DecodeInst->Flags &= ~DecodeFlags::FLAG_REP_PREFIX;
return NormalOpHeader(&FEXCore::X86Tables::RepModOps[EscapeOp], EscapeOp);
}
else if (DecodeInst->LastEscapePrefix == 0xF2) { // REPNE
// Remove prefix so it doesn't effect calculations.
// This is only an escape prefix rather tan modifier now
DecodeInst->Flags &= ~DecodeFlags::FLAG_REPNE_PREFIX;
return NormalOpHeader(&FEXCore::X86Tables::RepNEModOps[EscapeOp], EscapeOp);
}
else if (DecodeInst->LastEscapePrefix == 0x66 && !NoOverlay66) { // Operand Size
// Remove prefix so it doesn't effect calculations.
// This is only an escape prefix rather tan modifier now
DecodeInst->Flags &= ~DecodeFlags::FLAG_OPERAND_SIZE;
DecodeFlags::PopOpAddrIf(&DecodeInst->Flags, DecodeFlags::FLAG_OPERAND_SIZE_LAST);
return NormalOpHeader(&FEXCore::X86Tables::OpSizeModOps[EscapeOp], EscapeOp);
}
else {
return NormalOpHeader(&FEXCore::X86Tables::SecondBaseOps[EscapeOp], EscapeOp);
}
break;
}
break;
}
@@ -1017,7 +917,7 @@ bool Decoder::DecodeInstruction(uint64_t PC) {
case 0x65: // GS prefix
DecodeInst->Flags |= DecodeFlags::FLAG_GS_PREFIX;
break;
default: { // Default base table
default: [[likely]] { // Default base table
auto Info = &FEXCore::X86Tables::BaseOps[Op];
if (Info->Type == FEXCore::X86Tables::TYPE_REX_PREFIX) {
@@ -1127,6 +1027,35 @@ void Decoder::BranchTargetInMultiblockRange() {
}
}
bool Decoder::BranchTargetCanContinue(bool FinalInstruction) const {
if (FinalInstruction) {
return false;
}
uint64_t TargetRIP = 0;
const uint8_t GPRSize = CTX->GetGPRSize();
if (DecodeInst->OP == 0xE8) { // Call - immediate target
const uint64_t NextRIP = DecodeInst->PC + DecodeInst->InstSize;
LOGMAN_THROW_A_FMT(DecodeInst->Src[0].IsLiteral(), "Had wrong operand type");
TargetRIP = DecodeInst->PC + DecodeInst->InstSize + DecodeInst->Src[0].Data.Literal.Value;
if (GPRSize == 4) {
// If we are running a 32bit guest then wrap around addresses that go above 32bit
TargetRIP &= 0xFFFFFFFFU;
}
if (TargetRIP == NextRIP) {
// Optimize the case that the instruction is jumping just after itself.
// This is a GOT calculation which we can optimize out.
// Optimization occurs inside of the OpDispatcher implementation
return true;
}
}
return false;
}
const uint8_t *Decoder::AdjustAddrForSpecialRegion(uint8_t const* _InstStream, uint64_t EntryPoint, uint64_t RIP) {
constexpr uint64_t VSyscall_Base = 0xFFFF'FFFF'FF60'0000ULL;
constexpr uint64_t VSyscall_End = VSyscall_Base + 0x1000;
@@ -1207,24 +1136,19 @@ void Decoder::DecodeInstructionsAtEntry(uint8_t const* _InstStream, uint64_t PC,
auto OpMinPage = OpMinAddress & FHU::FEX_PAGE_MASK;
auto OpMaxPage = OpMaxAddress & FHU::FEX_PAGE_MASK;
if (OpMinPage != CurrentCodePage) {
CurrentCodePage = OpMinPage;
if (CodePages.insert(CurrentCodePage).second) {
AddContainedCodePage(PC, CurrentCodePage, FHU::FEX_PAGE_SIZE);
}
CodePages.insert(CurrentCodePage);
}
if (OpMaxPage != CurrentCodePage) {
CurrentCodePage = OpMaxPage;
if (CodePages.insert(CurrentCodePage).second) {
AddContainedCodePage(PC, CurrentCodePage, FHU::FEX_PAGE_SIZE);
}
CodePages.insert(CurrentCodePage);
}
bool ErrorDuringDecoding = !DecodeInstruction(RIPToDecode + PCOffset);
if (ErrorDuringDecoding) {
if (ErrorDuringDecoding) [[unlikely]] {
LogMan::Msg::DFmt("Couldn't Decode something at 0x{:x}, Started at 0x{:x}", RIPToDecode + PCOffset, PC);
// Put an invalid instruction in the stream so the core can raise SIGILL if hit
CurrentBlockDecoding.HasInvalidInstruction = true;
@@ -1251,23 +1175,21 @@ void Decoder::DecodeInstructionsAtEntry(uint8_t const* _InstStream, uint64_t PC,
CanContinue = true;
}
bool FinalInstruction = DecodedSize >= CTX->Config.MaxInstPerBlock ||
DecodedSize >= DefaultDecodedBufferSize ||
TotalInstructions >= CTX->Config.MaxInstPerBlock;
if (DecodeInst->TableInfo->Flags & FEXCore::X86Tables::InstFlags::FLAGS_SETS_RIP) {
// If we have multiblock enabled
// If the branch target is within our multiblock range then we can keep going on
// We don't want to short circuit this since we want to calculate our ranges still
BranchTargetInMultiblockRange();
// Bypass branches if we can continue through them in some cases.
CanContinue |= BranchTargetCanContinue(FinalInstruction);
}
if (!CanContinue) {
break;
}
if (DecodedSize >= CTX->Config.MaxInstPerBlock ||
DecodedSize >= DefaultDecodedBufferSize) {
break;
}
if (TotalInstructions >= CTX->Config.MaxInstPerBlock) {
if (FinalInstruction || !CanContinue) {
break;
}
@@ -1283,6 +1205,9 @@ void Decoder::DecodeInstructionsAtEntry(uint8_t const* _InstStream, uint64_t PC,
CurrentBlockDecoding.DecodedInstructions = &DecodedBuffer[BlockStartOffset];
}
for (auto CodePage : CodePages) {
AddContainedCodePage(PC, CodePage, FHU::FEX_PAGE_SIZE);
}
// sort for better branching
std::sort(Blocks.begin(), Blocks.end(), [](const FEXCore::Frontend::Decoder::DecodedBlocks& a, const FEXCore::Frontend::Decoder::DecodedBlocks& b) {
+4 -3
View File
@@ -11,7 +11,7 @@
#include <vector>
namespace FEXCore::Context {
struct Context;
class ContextImpl;
}
namespace FEXCore::Frontend {
@@ -25,7 +25,7 @@ public:
bool HasInvalidInstruction{};
};
Decoder(FEXCore::Context::Context *ctx);
Decoder(FEXCore::Context::ContextImpl *ctx);
~Decoder();
void DecodeInstructionsAtEntry(uint8_t const* InstStream, uint64_t PC, std::function<void(uint64_t BlockEntry, uint64_t Start, uint64_t Length)> AddContainedCodePage);
@@ -52,12 +52,13 @@ private:
bool L; // VEX.L bit (if set then 256 bit operation, if unset then scalar or 128-bit operation)
};
FEXCore::Context::Context *CTX;
FEXCore::Context::ContextImpl *CTX;
const FEXCore::HLE::SyscallOSABI OSABI{};
bool DecodeInstruction(uint64_t PC);
void BranchTargetInMultiblockRange();
bool BranchTargetCanContinue(bool FinalInstruction) const;
uint8_t ReadByte();
uint8_t PeekByte(uint8_t Offset) const;
+2 -2
View File
@@ -68,11 +68,11 @@ void GdbServer::WaitForThreadWakeup() {
ThreadBreakEvent.Wait();
}
GdbServer::GdbServer(FEXCore::Context::Context *ctx) : CTX(ctx) {
GdbServer::GdbServer(FEXCore::Context::ContextImpl *ctx) : CTX(ctx) {
// Pass all signals by default
std::fill(PassSignals.begin(), PassSignals.end(), true);
Context::SetExitHandler(ctx, [this](uint64_t ThreadId, FEXCore::Context::ExitReason ExitReason) {
ctx->SetExitHandler([this](uint64_t ThreadId, FEXCore::Context::ExitReason ExitReason) {
if (ExitReason == FEXCore::Context::ExitReason::EXIT_DEBUG) {
this->Break(SIGTRAP);
}
+3 -3
View File
@@ -19,12 +19,12 @@ $end_info$
namespace FEXCore {
namespace Context {
struct Context;
class ContextImpl;
}
class GdbServer {
public:
GdbServer(FEXCore::Context::Context *ctx);
GdbServer(FEXCore::Context::ContextImpl *ctx);
// Public for threading
void GdbServerLoop();
@@ -75,7 +75,7 @@ private:
std::string readRegs();
HandledPacketType readReg(const std::string& packet);
FEXCore::Context::Context *CTX;
FEXCore::Context::ContextImpl *CTX;
std::unique_ptr<FEXCore::Threads::Thread> gdbServerThread;
std::unique_ptr<std::iostream> CommsStream;
std::mutex sendMutex;
@@ -79,6 +79,7 @@ HostFeatures::HostFeatures() {
SupportsSHA = true;
SupportsBMI1 = true;
SupportsBMI2 = true;
SupportsCLWB = true;
if (!SupportsAtomics) {
WARN_ONCE_FMT("Host CPU doesn't support atomics. Expect bad performance");
@@ -128,6 +129,7 @@ HostFeatures::HostFeatures() {
SupportsSHA = Features.has(Xbyak::util::Cpu::tSHA);
SupportsBMI1 = Features.has(Xbyak::util::Cpu::tBMI1);
SupportsBMI2 = Features.has(Xbyak::util::Cpu::tBMI2);
SupportsBMI2 = Features.has(Xbyak::util::Cpu::tCLWB);
SupportsPMULL_128Bit = Features.has(Xbyak::util::Cpu::tPCLMULQDQ);
// xbyak doesn't know how to check for CLZero
@@ -17,20 +17,8 @@ $end_info$
#include <unistd.h>
namespace FEXCore::CPU {
[[noreturn]]
static void SignalReturn(FEXCore::Core::InternalThreadState *Thread) {
Thread->CTX->SignalThread(Thread, FEXCore::Core::SignalEvent::Return);
LOGMAN_MSG_A_FMT("unreachable");
FEX_UNREACHABLE;
}
#define DEF_OP(x) void InterpreterOps::Op_##x(IR::IROp_Header *IROp, IROpData *Data, IR::NodeID Node)
DEF_OP(SignalReturn) {
SignalReturn(Data->State);
}
DEF_OP(CallbackReturn) {
Data->State->CurrentFrame->Pointers.Interpreter.CallbackReturn(Data->State, Data->StackEntry);
}
@@ -91,7 +79,7 @@ DEF_OP(Syscall) {
Args.Argument[j] = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[j]);
}
uint64_t Res = FEXCore::Context::HandleSyscall(Data->State->CTX->SyscallHandler, Data->State->CurrentFrame, &Args);
uint64_t Res = FEXCore::Context::HandleSyscall(static_cast<Context::ContextImpl*>(Data->State->CTX)->SyscallHandler, Data->State->CurrentFrame, &Args);
GD = Res;
}
@@ -126,7 +114,7 @@ DEF_OP(InlineSyscall) {
DEF_OP(Thunk) {
auto Op = IROp->C<IR::IROp_Thunk>();
auto thunkFn = Data->State->CTX->ThunkHandler->LookupThunk(Op->ThunkNameHash);
auto thunkFn = static_cast<Context::ContextImpl*>(Data->State->CTX)->ThunkHandler->LookupThunk(Op->ThunkNameHash);
thunkFn(*GetSrc<void**>(Data->SSAData, Op->ArgPtr));
}
@@ -142,7 +130,7 @@ DEF_OP(ValidateCode) {
}
DEF_OP(ThreadRemoveCodeEntry) {
Data->State->CTX->ThreadRemoveCodeEntryFromJit(Data->State->CurrentFrame, Data->CurrentEntry);
static_cast<Context::ContextImpl*>(Data->State->CTX)->ThreadRemoveCodeEntryFromJit(Data->State->CurrentFrame, Data->CurrentEntry);
}
DEF_OP(CPUID) {
@@ -151,7 +139,7 @@ DEF_OP(CPUID) {
const uint64_t Arg = *GetSrc<uint64_t*>(Data->SSAData, Op->Function);
const uint64_t Leaf = *GetSrc<uint64_t*>(Data->SSAData, Op->Leaf);
auto Results = Data->State->CTX->CPUID.RunFunction(Arg, Leaf);
auto Results = Data->State->CTX->RunCPUIDFunction(Arg, Leaf);
memcpy(DstPtr, &Results, sizeof(uint32_t) * 4);
}
@@ -62,6 +62,23 @@ DEF_OP(VCastFromGPR) {
memcpy(GDP, GetSrc<void*>(Data->SSAData, Op->Src), Op->Header.ElementSize);
}
DEF_OP(VDupFromGPR) {
const auto Op = IROp->C<IR::IROp_VDupFromGPR>();
const auto OpSize = IROp->Size;
const auto ElementSize = IROp->ElementSize;
const auto NumElements = OpSize / IROp->ElementSize;
uint8_t Tmp[Core::CPUState::XMM_AVX_REG_SIZE]{};
const auto *Src = GetSrc<void*>(Data->SSAData, Op->Src);
for (size_t i = 0; i < NumElements; i++) {
memcpy(Tmp + (i * ElementSize), Src, ElementSize);
}
memcpy(GDP, Tmp, sizeof(Tmp));
}
DEF_OP(Float_FromGPR_S) {
auto Op = IROp->C<IR::IROp_Float_FromGPR_S>();
@@ -195,7 +212,7 @@ DEF_OP(Vector_FToF) {
// Sometimes is used to convert from a 128bit vector register
// in to a 64bit vector register with different sized elements
// eg: %ssa5 i32v2 = Vector_FToF %ssa4 i128, #0x8
uint8_t Elements = (OpSize << 1) / Op->SrcElementSize;
uint8_t Elements = OpSize == 8 ? 2 : OpSize / Op->SrcElementSize;
DO_VECTOR_1SRC_2TYPE_OP_NOSIZE(float, double, Func, 0, 0)
break;
}
@@ -26,7 +26,7 @@ public:
[[nodiscard]] std::string GetName() override { return "Interpreter"; }
[[nodiscard]] void *CompileCode(uint64_t Entry,
[[nodiscard]] CPUBackend::CompiledCode CompileCode(uint64_t Entry,
FEXCore::IR::IRListView const *IR,
FEXCore::Core::DebugData *DebugData,
FEXCore::IR::RegisterAllocationData *RAData, bool GDBEnabled) override;
@@ -35,7 +35,7 @@ public:
[[nodiscard]] bool NeedsOpDispatch() override { return true; }
static void InitializeSignalHandlers(FEXCore::Context::Context *CTX);
static void InitializeSignalHandlers(FEXCore::Context::ContextImpl *CTX);
void ClearCache() override;
@@ -49,7 +49,10 @@ InterpreterCore::InterpreterCore(Dispatcher *Dispatcher, FEXCore::Core::Internal
ClearCache();
}
void InterpreterCore::InitializeSignalHandlers(FEXCore::Context::Context *CTX) {
void InterpreterCore::InitializeSignalHandlers(FEXCore::Context::ContextImpl *CTX) {
CTX->SignalDelegation->RegisterHostSignalHandler(SIGILL, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
return reinterpret_cast<Context::ContextImpl*>(Thread->CTX)->Dispatcher->HandleSIGILL(Thread, Signal, info, ucontext);
}, true);
#ifdef _M_ARM_64
CTX->SignalDelegation->RegisterHostSignalHandler(SIGBUS, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
@@ -58,18 +61,21 @@ void InterpreterCore::InitializeSignalHandlers(FEXCore::Context::Context *CTX) {
#endif
}
void *InterpreterCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IRListView const *IR, [[maybe_unused]] FEXCore::Core::DebugData *DebugData, FEXCore::IR::RegisterAllocationData *RAData, bool GDBEnabled) {
CPUBackend::CompiledCode InterpreterCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IRListView const *IR, [[maybe_unused]] FEXCore::Core::DebugData *DebugData, FEXCore::IR::RegisterAllocationData *RAData, bool GDBEnabled) {
const auto IRSize = AlignUp(IR->GetInlineSize(), 16);
const auto MaxSize = IRSize + Dispatcher::MaxInterpreterTrampolineSize + GDBEnabled * Dispatcher::MaxGDBPauseCheckSize;
if ((BufferUsed + MaxSize) > CurrentCodeBuffer->Size) {
ThreadState->CTX->ClearCodeCache(ThreadState);
static_cast<Context::ContextImpl*>(ThreadState->CTX)->ClearCodeCache(ThreadState);
}
const auto BufferStart = CurrentCodeBuffer->Ptr + BufferUsed;
CPUBackend::CompiledCode CodeData{};
auto DestBuffer = BufferStart;
const auto BufferStartOffset = BufferUsed;
CodeData.BlockBegin = CodeData.BlockEntry = CurrentCodeBuffer->Ptr + BufferStartOffset;
auto DestBuffer = CodeData.BlockBegin;
if (GDBEnabled) {
const auto GDBSize = Dispatch->GenerateGDBPauseCheck(DestBuffer, Entry);
@@ -86,7 +92,9 @@ void *InterpreterCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR:
DestBuffer += IRSize;
BufferUsed += IRSize;
return BufferStart;
CodeData.Size = BufferUsed - BufferStartOffset;
return CodeData;
}
void InterpreterCore::ClearCache() {
@@ -95,11 +103,11 @@ void InterpreterCore::ClearCache() {
BufferUsed = 0;
}
std::unique_ptr<CPUBackend> CreateInterpreterCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread) {
std::unique_ptr<CPUBackend> CreateInterpreterCore(FEXCore::Context::ContextImpl *ctx, FEXCore::Core::InternalThreadState *Thread) {
return std::make_unique<InterpreterCore>(ctx->Dispatcher.get(), Thread);
}
void InitializeInterpreterSignalHandlers(FEXCore::Context::Context *CTX) {
void InitializeInterpreterSignalHandlers(FEXCore::Context::ContextImpl *CTX) {
InterpreterCore::InitializeSignalHandlers(CTX);
}
@@ -3,7 +3,7 @@
#include <memory>
namespace FEXCore::Context {
struct Context;
class ContextImpl;
}
namespace FEXCore::Core {
@@ -14,9 +14,9 @@ namespace FEXCore::CPU {
class CPUBackend;
struct DispatcherConfig;
[[nodiscard]] std::unique_ptr<CPUBackend> CreateInterpreterCore(FEXCore::Context::Context *ctx,
[[nodiscard]] std::unique_ptr<CPUBackend> CreateInterpreterCore(FEXCore::Context::ContextImpl *ctx,
FEXCore::Core::InternalThreadState *Thread);
void InitializeInterpreterSignalHandlers(FEXCore::Context::Context *CTX);
void InitializeInterpreterSignalHandlers(FEXCore::Context::ContextImpl *CTX);
CPUBackendFeatures GetInterpreterBackendFeatures();
} // namespace FEXCore::CPU
@@ -113,7 +113,6 @@ constexpr OpHandlerArray InterpreterOpHandlers = [] {
REGISTER_OP(ATOMICFETCHNEG, AtomicFetchNeg);
// Branch ops
REGISTER_OP(SIGNALRETURN, SignalReturn);
REGISTER_OP(CALLBACKRETURN, CallbackReturn);
REGISTER_OP(EXITFUNCTION, ExitFunction);
REGISTER_OP(JUMP, Jump);
@@ -128,6 +127,7 @@ constexpr OpHandlerArray InterpreterOpHandlers = [] {
// Conversion ops
REGISTER_OP(VINSGPR, VInsGPR);
REGISTER_OP(VCASTFROMGPR, VCastFromGPR);
REGISTER_OP(VDUPFROMGPR, VDupFromGPR);
REGISTER_OP(FLOAT_FROMGPR_S, Float_FromGPR_S);
REGISTER_OP(FLOAT_FTOF, Float_FToF);
REGISTER_OP(VECTOR_STOF, Vector_SToF);
@@ -154,7 +154,9 @@ constexpr OpHandlerArray InterpreterOpHandlers = [] {
REGISTER_OP(STOREMEM, StoreMem);
REGISTER_OP(LOADMEMTSO, LoadMem);
REGISTER_OP(STOREMEMTSO, StoreMem);
REGISTER_OP(MEMSET, MemSet);
REGISTER_OP(CACHELINECLEAR, CacheLineClear);
REGISTER_OP(CACHELINECLEAN, CacheLineClean);
REGISTER_OP(CACHELINEZERO, CacheLineZero);
// Misc ops
@@ -221,6 +223,8 @@ constexpr OpHandlerArray InterpreterOpHandlers = [] {
REGISTER_OP(VZIP2, VZip);
REGISTER_OP(VUNZIP, VUnZip);
REGISTER_OP(VUNZIP2, VUnZip);
REGISTER_OP(VTRN, VTrn);
REGISTER_OP(VTRN2, VTrn);
REGISTER_OP(VBSL, VBSL);
REGISTER_OP(VCMPEQ, VCMPEQ);
REGISTER_OP(VCMPEQZ, VCMPEQZ);
@@ -329,7 +333,6 @@ void InterpreterOps::InterpretIR(FEXCore::Core::CpuStateFrame *Frame, FEXCore::I
const uintptr_t ListSize = CurrentIR->GetSSACount();
static_assert(sizeof(FEXCore::IR::IROp_Header) == 4);
static_assert(sizeof(FEXCore::IR::OrderedNode) == 16);
auto BlockEnd = CurrentIR->GetBlocks().end();
@@ -142,7 +142,6 @@ namespace FEXCore::CPU {
DEF_OP(AtomicFetchNeg);
///< Branch ops
DEF_OP(SignalReturn);
DEF_OP(CallbackReturn);
DEF_OP(ExitFunction);
DEF_OP(Jump);
@@ -157,6 +156,7 @@ namespace FEXCore::CPU {
///< Conversion ops
DEF_OP(VInsGPR);
DEF_OP(VCastFromGPR);
DEF_OP(VDupFromGPR);
DEF_OP(Float_FromGPR_S);
DEF_OP(Float_FToF);
DEF_OP(Vector_SToF);
@@ -181,7 +181,9 @@ namespace FEXCore::CPU {
DEF_OP(StoreFlag);
DEF_OP(LoadMem);
DEF_OP(StoreMem);
DEF_OP(MemSet);
DEF_OP(CacheLineClear);
DEF_OP(CacheLineClean);
DEF_OP(CacheLineZero);
///< Misc ops
@@ -241,6 +243,7 @@ namespace FEXCore::CPU {
DEF_OP(VSMax);
DEF_OP(VZip);
DEF_OP(VUnZip);
DEF_OP(VTrn);
DEF_OP(VBSL);
DEF_OP(VCMPEQ);
DEF_OP(VCMPEQZ);
@@ -23,6 +23,22 @@ static inline void CacheLineFlush(char *Addr) {
#endif
}
static inline void CacheLineClean(char *Addr) {
#ifdef _M_X86_64
__asm volatile (
"clwb (%[Addr]);"
:: [Addr] "r" (Addr)
: "memory");
#elif _M_ARM_64
__asm volatile (
"dc cvac, %[Addr]"
:: [Addr] "r" (Addr)
: "memory");
#else
LOGMAN_THROW_A_FMT("Unsupported architecture with cacheline clean");
#endif
}
#define DEF_OP(x) void InterpreterOps::Op_##x(IR::IROp_Header *IROp, IROpData *Data, IR::NodeID Node)
DEF_OP(LoadContext) {
const auto Op = IROp->C<IR::IROp_LoadContext>();
@@ -272,6 +288,111 @@ DEF_OP(StoreMem) {
}
}
DEF_OP(MemSet) {
const auto Op = IROp->C<IR::IROp_MemSet>();
const int32_t Size = Op->Size;
char *MemData = *GetSrc<char **>(Data->SSAData, Op->Addr);
const auto Value = *GetSrc<uint64_t*>(Data->SSAData, Op->Value);
const auto Length = *GetSrc<uint64_t*>(Data->SSAData, Op->Length);
const auto Direction = *GetSrc<uint8_t*>(Data->SSAData, Op->Direction);
auto MemSetElements = [](auto* Memory, uint64_t Value, size_t Length) {
for (size_t i = 0; i < Length; ++i) {
Memory[i] = Value;
}
};
auto MemSetElementsInverse = [](auto* Memory, uint64_t Value, size_t Length) {
for (size_t i = 0; i < Length; ++i) {
Memory[-i] = Value;
}
};
if (Direction == 0) { // Forward
if (Op->IsAtomic) {
switch (Size) {
case 1:
MemSetElements(reinterpret_cast<std::atomic<uint8_t>*>(MemData), Value, Length);
break;
case 2:
MemSetElements(reinterpret_cast<std::atomic<uint16_t>*>(MemData), Value, Length);
break;
case 4:
MemSetElements(reinterpret_cast<std::atomic<uint32_t>*>(MemData), Value, Length);
break;
case 8:
MemSetElements(reinterpret_cast<std::atomic<uint64_t>*>(MemData), Value, Length);
break;
default:
LOGMAN_MSG_A_FMT("Unhandled {} size: {}", __func__, Size);
break;
}
}
else {
switch (Size) {
case 1:
MemSetElements(reinterpret_cast<uint8_t*>(MemData), Value, Length);
break;
case 2:
MemSetElements(reinterpret_cast<uint16_t*>(MemData), Value, Length);
break;
case 4:
MemSetElements(reinterpret_cast<uint32_t*>(MemData), Value, Length);
break;
case 8:
MemSetElements(reinterpret_cast<uint64_t*>(MemData), Value, Length);
break;
default:
LOGMAN_MSG_A_FMT("Unhandled {} size: {}", __func__, Size);
break;
}
}
GD = reinterpret_cast<uint64_t>(MemData + (Length * Size));
}
else { // Backward
if (Op->IsAtomic) {
switch (Size) {
case 1:
MemSetElementsInverse(reinterpret_cast<std::atomic<uint8_t>*>(MemData), Value, Length);
break;
case 2:
MemSetElementsInverse(reinterpret_cast<std::atomic<uint16_t>*>(MemData), Value, Length);
break;
case 4:
MemSetElementsInverse(reinterpret_cast<std::atomic<uint32_t>*>(MemData), Value, Length);
break;
case 8:
MemSetElementsInverse(reinterpret_cast<std::atomic<uint64_t>*>(MemData), Value, Length);
break;
default:
LOGMAN_MSG_A_FMT("Unhandled {} size: {}", __func__, Size);
break;
}
}
else {
switch (Size) {
case 1:
MemSetElementsInverse(reinterpret_cast<uint8_t*>(MemData), Value, Length);
break;
case 2:
MemSetElementsInverse(reinterpret_cast<uint16_t*>(MemData), Value, Length);
break;
case 4:
MemSetElementsInverse(reinterpret_cast<uint32_t*>(MemData), Value, Length);
break;
case 8:
MemSetElementsInverse(reinterpret_cast<uint64_t*>(MemData), Value, Length);
break;
default:
LOGMAN_MSG_A_FMT("Unhandled {} size: {}", __func__, Size);
break;
}
}
GD = reinterpret_cast<uint64_t>(MemData - (Length * Size));
}
}
DEF_OP(CacheLineClear) {
auto Op = IROp->C<IR::IROp_CacheLineClear>();
@@ -281,6 +402,15 @@ DEF_OP(CacheLineClear) {
CacheLineFlush(MemData);
}
DEF_OP(CacheLineClean) {
auto Op = IROp->C<IR::IROp_CacheLineClean>();
char *MemData = *GetSrc<char **>(Data->SSAData, Op->Addr);
// 64-byte cache line clear
CacheLineClean(MemData);
}
DEF_OP(CacheLineZero) {
auto Op = IROp->C<IR::IROp_CacheLineZero>();
@@ -902,6 +902,67 @@ DEF_OP(VZip) {
memcpy(GDP, Tmp, OpSize);
}
DEF_OP(VTrn) {
const auto Op = IROp->C<IR::IROp_VTrn>();
const uint8_t OpSize = IROp->Size;
void *Src1 = GetSrc<void*>(Data->SSAData, Op->VectorLower);
void *Src2 = GetSrc<void*>(Data->SSAData, Op->VectorUpper);
uint8_t Tmp[Core::CPUState::XMM_AVX_REG_SIZE]{};
const uint8_t ElementSize = Op->Header.ElementSize;
uint8_t Elements = OpSize / ElementSize;
const uint8_t BaseOffset = IROp->Op == IR::OP_VTRN2 ? 1 : 0;
Elements >>= 1;
switch (ElementSize) {
case 1: {
auto *Dst_d = reinterpret_cast<uint8_t*>(Tmp);
auto *Src1_d = reinterpret_cast<uint8_t*>(Src1);
auto *Src2_d = reinterpret_cast<uint8_t*>(Src2);
for (unsigned i = 0; i < Elements; ++i) {
Dst_d[i*2] = Src1_d[i*2 + BaseOffset];
Dst_d[i*2+1] = Src2_d[i*2 + BaseOffset];
}
break;
}
case 2: {
auto *Dst_d = reinterpret_cast<uint16_t*>(Tmp);
auto *Src1_d = reinterpret_cast<uint16_t*>(Src1);
auto *Src2_d = reinterpret_cast<uint16_t*>(Src2);
for (unsigned i = 0; i < Elements; ++i) {
Dst_d[i*2] = Src1_d[i*2 + BaseOffset];
Dst_d[i*2+1] = Src2_d[i*2 + BaseOffset];
}
break;
}
case 4: {
auto *Dst_d = reinterpret_cast<uint32_t*>(Tmp);
auto *Src1_d = reinterpret_cast<uint32_t*>(Src1);
auto *Src2_d = reinterpret_cast<uint32_t*>(Src2);
for (unsigned i = 0; i < Elements; ++i) {
Dst_d[i*2] = Src1_d[i*2 + BaseOffset];
Dst_d[i*2+1] = Src2_d[i*2 + BaseOffset];
}
break;
}
case 8: {
auto *Dst_d = reinterpret_cast<uint64_t*>(Tmp);
auto *Src1_d = reinterpret_cast<uint64_t*>(Src1);
auto *Src2_d = reinterpret_cast<uint64_t*>(Src2);
for (unsigned i = 0; i < Elements; ++i) {
Dst_d[i*2] = Src1_d[i*2 + BaseOffset];
Dst_d[i*2+1] = Src2_d[i*2 + BaseOffset];
}
break;
}
default:
LOGMAN_MSG_A_FMT("Unknown Element Size: {}", ElementSize);
break;
}
memcpy(GDP, Tmp, OpSize);
}
DEF_OP(VUnZip) {
const auto Op = IROp->C<IR::IROp_VUnZip>();
const uint8_t OpSize = IROp->Size;
@@ -964,7 +1025,9 @@ DEF_OP(VUnZip) {
}
DEF_OP(VBSL) {
auto Op = IROp->C<IR::IROp_VBSL>();
const auto Op = IROp->C<IR::IROp_VBSL>();
const auto OpSize = IROp->Size;
const auto Src1 = *GetSrc<InterpVector256*>(Data->SSAData, Op->VectorMask);
const auto Src2 = *GetSrc<InterpVector256*>(Data->SSAData, Op->VectorTrue);
const auto Src3 = *GetSrc<InterpVector256*>(Data->SSAData, Op->VectorFalse);
@@ -974,7 +1037,8 @@ DEF_OP(VBSL) {
.Upper = (Src2.Upper & Src1.Upper) | (Src3.Upper & ~Src1.Upper),
};
memcpy(GDP, &Tmp, sizeof(Tmp));
memset(GDP, 0, sizeof(InterpVector256));
memcpy(GDP, &Tmp, OpSize);
}
DEF_OP(VCMPEQ) {
@@ -1549,6 +1613,12 @@ DEF_OP(VInsElement) {
Dst_d[Op->DestIdx] = Src2_d[Op->SrcIdx];
break;
}
case 16: {
auto *Dst_d = reinterpret_cast<__uint128_t*>(Tmp);
auto *Src2_d = reinterpret_cast<__uint128_t*>(Src2);
Dst_d[Op->DestIdx] = Src2_d[Op->SrcIdx];
break;
}
default:
LOGMAN_MSG_A_FMT("Unknown Element Size: {}", ElementSize);
break;
@@ -1651,8 +1721,9 @@ DEF_OP(VExtr) {
const auto* VectorsPtr = reinterpret_cast<const uint8_t*>(Vectors.data());
const auto* SrcPtr = VectorsPtr + SanitizedByteIndex;
const auto CopyAmount = std::max(0, int(sizeof(Vectors) - SanitizedByteIndex));
memcpy(GDP, SrcPtr, OpSize);
memcpy(GDP, SrcPtr, CopyAmount);
} else {
uint64_t Offset = Index * ElementSize * 8;
File diff suppressed because it is too large. Load diff
@@ -9,7 +9,7 @@ $end_info$
#include "Interface/HLE/Thunks/Thunks.h"
namespace FEXCore::CPU {
uint64_t Arm64JITCore::GetNamedSymbolLiteral(FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op) {
switch (Op) {
case FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol::SYMBOL_LITERAL_EXITFUNCTION_LINKER:
@@ -22,18 +22,18 @@ uint64_t Arm64JITCore::GetNamedSymbolLiteral(FEXCore::CPU::RelocNamedSymbolLiter
return ~0ULL;
}
void Arm64JITCore::InsertNamedThunkRelocation(vixl::aarch64::Register Reg, const IR::SHA256Sum &Sum) {
void Arm64JITCore::InsertNamedThunkRelocation(ARMEmitter::Register Reg, const IR::SHA256Sum &Sum) {
Relocation MoveABI{};
MoveABI.NamedThunkMove.Header.Type = FEXCore::CPU::RelocationTypes::RELOC_NAMED_THUNK_MOVE;
// Offset is the offset from the entrypoint of the block
auto CurrentCursor = GetCursorAddress<uint8_t *>();
MoveABI.NamedThunkMove.Offset = CurrentCursor - GuestEntry;
MoveABI.NamedThunkMove.Offset = CurrentCursor - CodeData.BlockBegin;
MoveABI.NamedThunkMove.Symbol = Sum;
MoveABI.NamedThunkMove.RegisterIndex = Reg.GetCode();
MoveABI.NamedThunkMove.RegisterIndex = Reg.Idx();
uint64_t Pointer = reinterpret_cast<uint64_t>(EmitterCTX->ThunkHandler->LookupThunk(Sum));
LoadConstant(Reg, Pointer, EmitterCTX->Config.CacheObjectCodeCompilation());
LoadConstant(ARMEmitter::Size::i64Bit, Reg, Pointer, EmitterCTX->Config.CacheObjectCodeCompilation());
Relocations.emplace_back(MoveABI);
}
@@ -41,7 +41,7 @@ Arm64JITCore::NamedSymbolLiteralPair Arm64JITCore::InsertNamedSymbolLiteral(FEXC
uint64_t Pointer = GetNamedSymbolLiteral(Op);
Arm64JITCore::NamedSymbolLiteralPair Lit {
.Lit = Literal(Pointer),
.Lit = Pointer,
.MoveABI = {
.NamedSymbolLiteral = {
.Header = {
@@ -58,22 +58,23 @@ Arm64JITCore::NamedSymbolLiteralPair Arm64JITCore::InsertNamedSymbolLiteral(FEXC
void Arm64JITCore::PlaceNamedSymbolLiteral(NamedSymbolLiteralPair &Lit) {
// Offset is the offset from the entrypoint of the block
auto CurrentCursor = GetCursorAddress<uint8_t *>();
Lit.MoveABI.NamedSymbolLiteral.Offset = CurrentCursor - GuestEntry;
Lit.MoveABI.NamedSymbolLiteral.Offset = CurrentCursor - CodeData.BlockBegin;
place(&Lit.Lit);
Bind(&Lit.Loc);
dc64(Lit.Lit);
Relocations.emplace_back(Lit.MoveABI);
}
void Arm64JITCore::InsertGuestRIPMove(vixl::aarch64::Register Reg, uint64_t Constant) {
void Arm64JITCore::InsertGuestRIPMove(ARMEmitter::Register Reg, uint64_t Constant) {
Relocation MoveABI{};
MoveABI.GuestRIPMove.Header.Type = FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_MOVE;
// Offset is the offset from the entrypoint of the block
auto CurrentCursor = GetCursorAddress<uint8_t *>();
MoveABI.GuestRIPMove.Offset = CurrentCursor - GuestEntry;
MoveABI.GuestRIPMove.Offset = CurrentCursor - CodeData.BlockBegin;
MoveABI.GuestRIPMove.GuestRIP = Constant;
MoveABI.GuestRIPMove.RegisterIndex = Reg.GetCode();
MoveABI.GuestRIPMove.RegisterIndex = Reg.Idx();
LoadConstant(Reg, Constant, EmitterCTX->Config.CacheObjectCodeCompilation());
LoadConstant(ARMEmitter::Size::i64Bit, Reg, Constant, EmitterCTX->Config.CacheObjectCodeCompilation());
Relocations.emplace_back(MoveABI);
}
@@ -87,11 +88,10 @@ bool Arm64JITCore::ApplyRelocations(uint64_t GuestEntry, uint64_t CodeEntry, uin
case FEXCore::CPU::RelocationTypes::RELOC_NAMED_SYMBOL_LITERAL: {
uint64_t Pointer = GetNamedSymbolLiteral(Reloc->NamedSymbolLiteral.Symbol);
// Relocation occurs at the cursorEntry + offset relative to that cursor
GetBuffer()->SetCursorOffset(CursorEntry + Reloc->NamedSymbolLiteral.Offset);
SetCursorOffset(CursorEntry + Reloc->NamedSymbolLiteral.Offset);
// Generate a literal so we can place it
Literal<uint64_t> Lit(Pointer);
place(&Lit);
dc64(Pointer);
DataIndex += sizeof(Reloc->NamedSymbolLiteral);
break;
@@ -103,8 +103,8 @@ bool Arm64JITCore::ApplyRelocations(uint64_t GuestEntry, uint64_t CodeEntry, uin
}
// Relocation occurs at the cursorEntry + offset relative to that cursor.
GetBuffer()->SetCursorOffset(CursorEntry + Reloc->NamedThunkMove.Offset);
LoadConstant(vixl::aarch64::XRegister(Reloc->NamedThunkMove.RegisterIndex), Pointer, true);
SetCursorOffset(CursorEntry + Reloc->NamedThunkMove.Offset);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Register(Reloc->NamedThunkMove.RegisterIndex), Pointer, true);
DataIndex += sizeof(Reloc->NamedThunkMove);
break;
}
@@ -117,8 +117,8 @@ bool Arm64JITCore::ApplyRelocations(uint64_t GuestEntry, uint64_t CodeEntry, uin
}
// Relocation occurs at the cursorEntry + offset relative to that cursor.
GetBuffer()->SetCursorOffset(CursorEntry + Reloc->GuestRIPMove.Offset);
LoadConstant(vixl::aarch64::XRegister(Reloc->GuestRIPMove.RegisterIndex), Pointer, true);
SetCursorOffset(CursorEntry + Reloc->GuestRIPMove.Offset);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Register(Reloc->GuestRIPMove.RegisterIndex), Pointer, true);
DataIndex += sizeof(Reloc->GuestRIPMove);
break;
}
File diff suppressed because it is too large. Load diff
+165 -228
View File
@@ -6,6 +6,7 @@ $end_info$
#include "Interface/Context/Context.h"
#include "FEXCore/IR/IR.h"
#include "Interface/Core/ArchHelpers/CodeEmitter/Emitter.h"
#include "Interface/Core/LookupCache.h"
#include "Interface/Core/JIT/Arm64/JITClass.h"
@@ -17,22 +18,9 @@ $end_info$
#include <Interface/HLE/Thunks/Thunks.h>
namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const *IROp, IR::NodeID Node)
DEF_OP(SignalReturn) {
// First we must reset the stack
ResetStack();
// Now branch to our signal return helper
// This can't be a direct branch since the code needs to live at a constant location
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.SignalReturnHandler)));
br(x0);
}
DEF_OP(CallbackReturn) {
// spill back to CTX
SpillStaticRegs();
@@ -41,14 +29,14 @@ DEF_OP(CallbackReturn) {
// We can now lower the ref counter again
ldr(w2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, SignalHandlerRefCounter)));
sub(w2, w2, 1);
str(w2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, SignalHandlerRefCounter)));
ldr(ARMEmitter::WReg::w2, STATE, offsetof(FEXCore::Core::CpuStateFrame, SignalHandlerRefCounter));
sub(ARMEmitter::Size::i32Bit, ARMEmitter::Reg::r2, ARMEmitter::Reg::r2, 1);
str(ARMEmitter::WReg::w2, STATE, offsetof(FEXCore::Core::CpuStateFrame, SignalHandlerRefCounter));
// We need to adjust an additional 8 bytes to get back to the original "misaligned" RSP state
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RSP])));
add(x2, x2, 8);
str(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RSP])));
ldr(ARMEmitter::XReg::x2, STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RSP]));
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r2, ARMEmitter::Reg::r2, 8);
str(ARMEmitter::XReg::x2, STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RSP]));
PopCalleeSavedRegisters();
@@ -59,39 +47,41 @@ DEF_OP(CallbackReturn) {
DEF_OP(ExitFunction) {
auto Op = IROp->C<IR::IROp_ExitFunction>();
Label FullLookup;
ResetStack();
aarch64::Register RipReg;
uint64_t NewRIP;
if (IsInlineConstant(Op->NewRIP, &NewRIP) || IsInlineEntrypointOffset(Op->NewRIP, &NewRIP)) {
Literal l_BranchHost{ThreadState->CurrentFrame->Pointers.Common.ExitFunctionLinker};
Literal l_BranchGuest{NewRIP};
ARMEmitter::ForwardLabel l_BranchHost;
ARMEmitter::ForwardLabel l_BranchGuest;
ldr(x0, &l_BranchHost);
blr(x0);
ldr(ARMEmitter::XReg::x0, &l_BranchHost);
blr(ARMEmitter::Reg::r0);
Bind(&l_BranchHost);
dc64(ThreadState->CurrentFrame->Pointers.Common.ExitFunctionLinker);
Bind(&l_BranchGuest);
dc64(NewRIP);
place(&l_BranchHost);
place(&l_BranchGuest);
} else {
RipReg = GetReg<RA_64>(Op->NewRIP.ID());
ARMEmitter::ForwardLabel FullLookup;
auto RipReg = GetReg(Op->NewRIP.ID());
// L1 Cache
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.L1Pointer)));
ldr(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.L1Pointer));
and_(x3, RipReg, LookupCache::L1_ENTRIES_MASK);
add(x0, x0, Operand(x3, Shift::LSL, 4));
and_(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r3, RipReg, LookupCache::L1_ENTRIES_MASK);
add(ARMEmitter::XReg::x0, ARMEmitter::XReg::x0, ARMEmitter::XReg::x3, ARMEmitter::ShiftType::LSL, 4);
ldp(x1, x0, MemOperand(x0));
cmp(x0, RipReg);
b(&FullLookup, Condition::ne);
br(x1);
ldp<ARMEmitter::IndexType::OFFSET>(ARMEmitter::XReg::x1, ARMEmitter::XReg::x0, ARMEmitter::Reg::r0, 0);
cmp(ARMEmitter::XReg::x0, RipReg.X());
b(ARMEmitter::Condition::CC_NE, &FullLookup);
br(ARMEmitter::Reg::r1);
bind(&FullLookup);
ldr(TMP1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.DispatcherLoopTop)));
str(RipReg, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.rip)));
Bind(&FullLookup);
ldr(TMP1, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.DispatcherLoopTop));
str(RipReg.X(), STATE, offsetof(FEXCore::Core::CpuStateFrame, State.rip));
br(TMP1);
}
}
@@ -103,66 +93,65 @@ DEF_OP(Jump) {
PendingTargetLabel = &JumpTargets.try_emplace(Target).first->second;
}
#define GRCMP(Node) (Op->CompareSize == 4 ? GetReg<RA_32>(Node) : GetReg<RA_64>(Node))
#define GRFCMP(Node) (Op->CompareSize == 4 ? GetDst(Node).S() : GetDst(Node).D())
static Condition MapBranchCC(IR::CondClassType Cond) {
static ARMEmitter::Condition MapBranchCC(IR::CondClassType Cond) {
switch (Cond.Val) {
case FEXCore::IR::COND_EQ: return Condition::eq;
case FEXCore::IR::COND_NEQ: return Condition::ne;
case FEXCore::IR::COND_SGE: return Condition::ge;
case FEXCore::IR::COND_SLT: return Condition::lt;
case FEXCore::IR::COND_SGT: return Condition::gt;
case FEXCore::IR::COND_SLE: return Condition::le;
case FEXCore::IR::COND_UGE: return Condition::cs;
case FEXCore::IR::COND_ULT: return Condition::cc;
case FEXCore::IR::COND_UGT: return Condition::hi;
case FEXCore::IR::COND_ULE: return Condition::ls;
case FEXCore::IR::COND_FLU: return Condition::lt;
case FEXCore::IR::COND_FGE: return Condition::ge;
case FEXCore::IR::COND_FLEU:return Condition::le;
case FEXCore::IR::COND_FGT: return Condition::gt;
case FEXCore::IR::COND_FU: return Condition::vs;
case FEXCore::IR::COND_FNU: return Condition::vc;
case FEXCore::IR::COND_EQ: return ARMEmitter::Condition::CC_EQ;
case FEXCore::IR::COND_NEQ: return ARMEmitter::Condition::CC_NE;
case FEXCore::IR::COND_SGE: return ARMEmitter::Condition::CC_GE;
case FEXCore::IR::COND_SLT: return ARMEmitter::Condition::CC_LT;
case FEXCore::IR::COND_SGT: return ARMEmitter::Condition::CC_GT;
case FEXCore::IR::COND_SLE: return ARMEmitter::Condition::CC_LE;
case FEXCore::IR::COND_UGE: return ARMEmitter::Condition::CC_CS;
case FEXCore::IR::COND_ULT: return ARMEmitter::Condition::CC_CC;
case FEXCore::IR::COND_UGT: return ARMEmitter::Condition::CC_HI;
case FEXCore::IR::COND_ULE: return ARMEmitter::Condition::CC_LS;
case FEXCore::IR::COND_FLU: return ARMEmitter::Condition::CC_LT;
case FEXCore::IR::COND_FGE: return ARMEmitter::Condition::CC_GE;
case FEXCore::IR::COND_FLEU:return ARMEmitter::Condition::CC_LE;
case FEXCore::IR::COND_FGT: return ARMEmitter::Condition::CC_GT;
case FEXCore::IR::COND_FU: return ARMEmitter::Condition::CC_VS;
case FEXCore::IR::COND_FNU: return ARMEmitter::Condition::CC_VC;
case FEXCore::IR::COND_VS:
case FEXCore::IR::COND_VC:
case FEXCore::IR::COND_MI:
case FEXCore::IR::COND_PL:
default:
LOGMAN_MSG_A_FMT("Unsupported compare type");
return Condition::nv;
return ARMEmitter::Condition::CC_NV;
}
}
DEF_OP(CondJump) {
auto Op = IROp->C<IR::IROp_CondJump>();
Label *TrueTargetLabel = &JumpTargets.try_emplace(Op->TrueBlock.ID()).first->second;
auto TrueTargetLabel = &JumpTargets.try_emplace(Op->TrueBlock.ID()).first->second;
uint64_t Const;
const bool isConst = IsInlineConstant(Op->Cmp2, &Const);
const auto Size = Op->CompareSize == 4 ? ARMEmitter::Size::i32Bit : ARMEmitter::Size::i64Bit;
const auto SubSize = ARMEmitter::ToVectorSizePair(Op->CompareSize == 4 ? ARMEmitter::SubRegSize::i32Bit : ARMEmitter::SubRegSize::i64Bit);
if (isConst && Const == 0 && Op->Cond.Val == FEXCore::IR::COND_EQ) {
LOGMAN_THROW_A_FMT(IsGPR(Op->Cmp1.ID()), "CondJump: Expected GPR");
cbz(GRCMP(Op->Cmp1.ID()), TrueTargetLabel);
cbz(Size, GetReg(Op->Cmp1.ID()), TrueTargetLabel);
} else if (isConst && Const == 0 && Op->Cond.Val == FEXCore::IR::COND_NEQ) {
LOGMAN_THROW_A_FMT(IsGPR(Op->Cmp1.ID()), "CondJump: Expected GPR");
cbnz(GRCMP(Op->Cmp1.ID()), TrueTargetLabel);
cbnz(Size, GetReg(Op->Cmp1.ID()), TrueTargetLabel);
} else {
if (IsGPR(Op->Cmp1.ID())) {
if (isConst) {
cmp(GRCMP(Op->Cmp1.ID()), Const);
cmp(Size, GetReg(Op->Cmp1.ID()), Const);
} else {
cmp(GRCMP(Op->Cmp1.ID()), GRCMP(Op->Cmp2.ID()));
cmp(Size, GetReg(Op->Cmp1.ID()), GetReg(Op->Cmp2.ID()));
}
} else if (IsFPR(Op->Cmp1.ID())) {
fcmp(GRFCMP(Op->Cmp1.ID()), GRFCMP(Op->Cmp2.ID()));
fcmp(SubSize.Scalar, GetVReg(Op->Cmp1.ID()), GetVReg(Op->Cmp2.ID()));
} else {
LOGMAN_MSG_A_FMT("CondJump: Expected GPR or FPR");
}
b(TrueTargetLabel, MapBranchCC(Op->Cond));
b(MapBranchCC(Op->Cond), TrueTargetLabel);
}
PendingTargetLabel = &JumpTargets.try_emplace(Op->FalseBlock.ID()).first->second;
@@ -176,50 +165,59 @@ DEF_OP(Syscall) {
// X2: Pointer to SyscallArguments
FEXCore::IR::SyscallFlags Flags = Op->Flags;
PushDynamicRegsAndLR();
PushDynamicRegsAndLR(TMP1);
if ((Flags & FEXCore::IR::SyscallFlags::NOSYNCSTATEONENTRY) != FEXCore::IR::SyscallFlags::NOSYNCSTATEONENTRY) {
SpillStaticRegs();
}
else {
uint32_t GPRSpillMask = ~0U;
uint32_t FPRSpillMask = ~0U;
if ((Flags & FEXCore::IR::SyscallFlags::NOSYNCSTATEONENTRY) == FEXCore::IR::SyscallFlags::NOSYNCSTATEONENTRY) {
// Need to spill all caller saved registers still
SpillStaticRegs(true, CALLER_GPR_MASK, CALLER_FPR_MASK);
GPRSpillMask = CALLER_GPR_MASK;
FPRSpillMask = CALLER_FPR_MASK;
}
SpillStaticRegs(true, GPRSpillMask, FPRSpillMask);
// Now that we are spilled, store in the state that we are in a syscall
// Still without overwriting registers that matter
// 16bit LoadConstant to be a single instruction
// This gives the signal handler a value to check to see if we are in a syscall at all
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, GPRSpillMask & 0xFFFF);
str(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CpuStateFrame, InSyscallInfo));
uint64_t SPOffset = AlignUp(FEXCore::HLE::SyscallArguments::MAX_ARGS * 8, 16);
sub(sp, sp, SPOffset);
sub(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, ARMEmitter::Reg::rsp, SPOffset);
for (uint32_t i = 0; i < FEXCore::HLE::SyscallArguments::MAX_ARGS; ++i) {
if (Op->Header.Args[i].IsInvalid()) continue;
str(GetReg<RA_64>(Op->Header.Args[i].ID()), MemOperand(sp, i * 8));
str(GetReg(Op->Header.Args[i].ID()).X(), ARMEmitter::Reg::rsp, i * 8);
}
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.SyscallHandlerObj)));
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.SyscallHandlerFunc)));
mov(x1, STATE);
mov(x2, sp);
ldr(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.SyscallHandlerObj));
ldr(ARMEmitter::XReg::x3, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.SyscallHandlerFunc));
mov(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r1, STATE.R());
// SP supporting move
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r2, ARMEmitter::Reg::rsp, 0);
#ifdef VIXL_SIMULATOR
GenerateIndirectRuntimeCall<uint64_t, void*, void*, void*>(x3);
GenerateIndirectRuntimeCall<uint64_t, void*, void*, void*>(ARMEmitter::Reg::r3);
#else
blr(x3);
blr(ARMEmitter::Reg::r3);
#endif
add(sp, sp, SPOffset);
if ((Flags & FEXCore::IR::SyscallFlags::NOSYNCSTATEONENTRY) != FEXCore::IR::SyscallFlags::NOSYNCSTATEONENTRY &&
(Flags & FEXCore::IR::SyscallFlags::NORETURN) != FEXCore::IR::SyscallFlags::NORETURN) {
FillStaticRegs();
}
else {
// Result is now in x0
// Fix the stack and any values that were stepped on
FillStaticRegs(true, CALLER_GPR_MASK, CALLER_FPR_MASK);
}
PopDynamicRegsAndLR();
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, ARMEmitter::Reg::rsp, SPOffset);
if ((Flags & FEXCore::IR::SyscallFlags::NORETURN) != FEXCore::IR::SyscallFlags::NORETURN) {
// Result is now in x0
// Fix the stack and any values that were stepped on
FillStaticRegs(true, GPRSpillMask, FPRSpillMask);
// Now the registers we've spilled are back in their original host registers
// We can safely claim we are no longer in a syscall
str(ARMEmitter::XReg::zr, STATE, offsetof(FEXCore::Core::CpuStateFrame, InSyscallInfo));
PopDynamicRegsAndLR();
// Move result to its destination register
mov(GetReg<RA_64>(Node), x0);
mov(ARMEmitter::Size::i64Bit, GetReg(Node), ARMEmitter::Reg::r0);
}
}
@@ -236,8 +234,8 @@ DEF_OP(InlineSyscall) {
// X6: Arg6 - Doesn't exist in x86-64 land. RA INTERSECT
// One argument is removed from the SyscallArguments::MAX_ARGS since the first argument was syscall number
const static std::array<vixl::aarch64::Register, FEXCore::HLE::SyscallArguments::MAX_ARGS-1> RegArgs = {{
x0, x1, x2, x3, x4, x5
const static std::array<ARMEmitter::XRegister, FEXCore::HLE::SyscallArguments::MAX_ARGS-1> RegArgs = {{
ARMEmitter::XReg::x0, ARMEmitter::XReg::x1, ARMEmitter::XReg::x2, ARMEmitter::XReg::x3, ARMEmitter::XReg::x4, ARMEmitter::XReg::x5
}};
bool Intersects{};
@@ -246,19 +244,15 @@ DEF_OP(InlineSyscall) {
for (uint32_t i = 0; i < FEXCore::HLE::SyscallArguments::MAX_ARGS-1; ++i) {
if (Op->Header.Args[i].IsInvalid()) break;
auto Reg = GetReg<RA_64>(Op->Header.Args[i].ID());
if (Reg.GetCode() == x8.GetCode() ||
Reg.GetCode() == x4.GetCode() ||
Reg.GetCode() == x5.GetCode()) {
auto Reg = GetReg(Op->Header.Args[i].ID());
if (Reg == ARMEmitter::Reg::r8 ||
Reg == ARMEmitter::Reg::r4 ||
Reg == ARMEmitter::Reg::r5) {
SpillMask |= (1U << Reg.GetCode());
SpillMask |= (1U << Reg.Idx());
Intersects = true;
}
}
// XXX: For some reason spilling only the x4, x5, and x8 registers was causing issues
// Come back to this once investigation reveals why it fails the gvisor ioctl test
// For now override to all GPRs
SpillMask = ~0U;
// Ordering is incredibly important here
// We must spill any overlapping registers first THEN claim we are in a syscall without invalidating state at all
@@ -270,66 +264,31 @@ DEF_OP(InlineSyscall) {
// 16bit LoadConstant to be a single instruction
// We must always spill at least one register (x8) so this value always has a bit set
// This gives the signal handler a value to check to see if we are in a syscall at all
LoadConstant(x0, SpillMask & 0xFFFF);
str(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, InSyscallInfo)));
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, SpillMask & 0xFFFF);
str(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CpuStateFrame, InSyscallInfo));
// Now that we have claimed to be a syscall we can set up the arguments
const auto EmitSize = CTX->Config.Is64BitMode() ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
const auto EmitSubSize = CTX->Config.Is64BitMode() ? ARMEmitter::SubRegSize::i64Bit : ARMEmitter::SubRegSize::i32Bit;
if (Intersects) {
for (uint32_t i = 0; i < FEXCore::HLE::SyscallArguments::MAX_ARGS-1; ++i) {
if (Op->Header.Args[i].IsInvalid()) break;
if (CTX->Config.Is64BitMode()) {
auto Reg = GetReg<RA_64>(Op->Header.Args[i].ID());
// In the case of intersection with x4, x5, or x8 then these are currently SRA
// for registers RAX, RBX, and RSI. Which have just been spilled
// Just load back from the context. Could be slightly smarter but this is fairly uncommon
if (Reg.GetCode() == x8.GetCode()) {
ldr(RegArgs[i], MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RSI])));
}
else if (Reg.GetCode() == x4.GetCode()) {
ldr(RegArgs[i], MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RAX])));
}
else if (Reg.GetCode() == x5.GetCode()) {
ldr(RegArgs[i], MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RBX])));
}
auto Reg = GetReg(Op->Header.Args[i].ID());
// In the case of intersection with x4, x5, or x8 then these are currently SRA
// for registers RAX, RBX, and RSI. Which have just been spilled
// Just load back from the context. Could be slightly smarter but this is fairly uncommon
if (Reg == ARMEmitter::Reg::r8) {
ldr(EmitSubSize, RegArgs[i].R(), STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RSI]));
}
}
for (uint32_t i = 0; i < FEXCore::HLE::SyscallArguments::MAX_ARGS-1; ++i) {
if (Op->Header.Args[i].IsInvalid()) break;
if (CTX->Config.Is64BitMode()) {
auto Reg = GetReg<RA_64>(Op->Header.Args[i].ID());
// In the case of intersection with x4, x5, or x8 then these are currently SRA
// for registers RAX, RBX, and RSI. Which have just been spilled
// Just load back from the context. Could be slightly smarter but this is fairly uncommon
if (Reg.GetCode() == x8.GetCode()) {
ldr(RegArgs[i], MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RSI])));
}
else if (Reg.GetCode() == x4.GetCode()) {
ldr(RegArgs[i], MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RAX])));
}
else if (Reg.GetCode() == x5.GetCode()) {
ldr(RegArgs[i], MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RBX])));
}
else {
mov(RegArgs[i], Reg);
}
else if (Reg == ARMEmitter::Reg::r4) {
ldr(EmitSubSize, RegArgs[i].R(), STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RAX]));
}
else if (Reg == ARMEmitter::Reg::r5) {
ldr(EmitSubSize, RegArgs[i].R(), STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RBX]));
}
else {
auto Reg = GetReg<RA_32>(Op->Header.Args[i].ID());
if (Reg.GetCode() == w8.GetCode()) {
ldr(RegArgs[i].W(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RSI])));
}
else if (Reg.GetCode() == w4.GetCode()) {
ldr(RegArgs[i].W(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RAX])));
}
else if (Reg.GetCode() == w5.GetCode()) {
ldr(RegArgs[i].W(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RBX])));
}
else {
uxtw(RegArgs[i].W(), Reg);
}
mov(EmitSize, RegArgs[i].R(), Reg);
}
}
}
@@ -337,16 +296,11 @@ DEF_OP(InlineSyscall) {
for (uint32_t i = 0; i < FEXCore::HLE::SyscallArguments::MAX_ARGS-1; ++i) {
if (Op->Header.Args[i].IsInvalid()) break;
if (CTX->Config.Is64BitMode()) {
mov(RegArgs[i], GetReg<RA_64>(Op->Header.Args[i].ID()));
}
else {
uxtw(RegArgs[i], GetReg<RA_64>(Op->Header.Args[i].ID()));
}
mov(EmitSize, RegArgs[i].R(), GetReg(Op->Header.Args[i].ID()));
}
}
LoadConstant(x8, Op->HostSyscallNumber);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r8, Op->HostSyscallNumber);
svc(0);
// On updated signal mask we can receive a signal RIGHT HERE
@@ -357,16 +311,11 @@ DEF_OP(InlineSyscall) {
// Now the registers we've spilled are back in their original host registers
// We can safely claim we are no longer in a syscall
str(xzr, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, InSyscallInfo)));
str(ARMEmitter::XReg::zr, STATE, offsetof(FEXCore::Core::CpuStateFrame, InSyscallInfo));
// Result is now in x0
// Move result to its destination register
if (CTX->Config.Is64BitMode()) {
mov(GetReg<RA_64>(Node), x0);
}
else {
uxtw(GetReg<RA_64>(Node), x0);
}
mov(EmitSize, GetReg(Node), ARMEmitter::Reg::r0);
}
}
@@ -378,16 +327,16 @@ DEF_OP(Thunk) {
SpillStaticRegs(); // spill to ctx before ra64 spill
PushDynamicRegsAndLR();
PushDynamicRegsAndLR(TMP1);
mov(x0, GetReg<RA_64>(Op->ArgPtr.ID()));
mov(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, GetReg(Op->ArgPtr.ID()));
auto thunkFn = ThreadState->CTX->ThunkHandler->LookupThunk(Op->ThunkNameHash);
LoadConstant(x2, (uintptr_t)thunkFn);
auto thunkFn = static_cast<Context::ContextImpl*>(ThreadState->CTX)->ThunkHandler->LookupThunk(Op->ThunkNameHash);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r2, (uintptr_t)thunkFn);
#ifdef VIXL_SIMULATOR
GenerateIndirectRuntimeCall<void, void*, void*>(x2);
GenerateIndirectRuntimeCall<void, void*, void*>(ARMEmitter::Reg::r2);
#else
blr(x2);
blr(ARMEmitter::Reg::r2);
#endif
PopDynamicRegsAndLR();
@@ -401,43 +350,45 @@ DEF_OP(ValidateCode) {
int len = Op->CodeLength;
int idx = 0;
LoadConstant(GetReg<RA_64>(Node), 0);
LoadConstant(x0, Entry + Op->Offset);
LoadConstant(x1, 1);
LoadConstant(ARMEmitter::Size::i64Bit, GetReg(Node), 0);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, Entry + Op->Offset);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r1, 1);
const auto Dst = GetReg(Node);
while (len >= 8)
{
ldr(x2, MemOperand(x0, idx));
LoadConstant(x3, *(const uint32_t *)(OldCode + idx));
cmp(x2, x3);
csel(GetReg<RA_64>(Node), GetReg<RA_64>(Node), x1, Condition::eq);
ldr(ARMEmitter::XReg::x2, ARMEmitter::Reg::r0, idx);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r3, *(const uint32_t *)(OldCode + idx));
cmp(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r2, ARMEmitter::Reg::r3);
csel(ARMEmitter::Size::i64Bit, Dst, Dst, ARMEmitter::Reg::r1, ARMEmitter::Condition::CC_EQ);
len -= 8;
idx += 8;
}
while (len >= 4)
{
ldr(w2, MemOperand(x0, idx));
LoadConstant(w3, *(const uint32_t *)(OldCode + idx));
cmp(w2, w3);
csel(GetReg<RA_64>(Node), GetReg<RA_64>(Node), x1, Condition::eq);
ldr(ARMEmitter::WReg::w2, ARMEmitter::Reg::r0, idx);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r3, *(const uint32_t *)(OldCode + idx));
cmp(ARMEmitter::Size::i32Bit, ARMEmitter::Reg::r2, ARMEmitter::Reg::r3);
csel(ARMEmitter::Size::i64Bit, Dst, Dst, ARMEmitter::Reg::r1, ARMEmitter::Condition::CC_EQ);
len -= 4;
idx += 4;
}
while (len >= 2)
{
ldrh(w2, MemOperand(x0, idx));
LoadConstant(w3, *(const uint16_t *)(OldCode + idx));
cmp(w2, w3);
csel(GetReg<RA_64>(Node), GetReg<RA_64>(Node), x1, Condition::eq);
ldrh(ARMEmitter::Reg::r2, ARMEmitter::Reg::r0, idx);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r3, *(const uint16_t *)(OldCode + idx));
cmp(ARMEmitter::Size::i32Bit, ARMEmitter::Reg::r2, ARMEmitter::Reg::r3);
csel(ARMEmitter::Size::i64Bit, Dst, Dst, ARMEmitter::Reg::r1, ARMEmitter::Condition::CC_EQ);
len -= 2;
idx += 2;
}
while (len >= 1)
{
ldrb(w2, MemOperand(x0, idx));
LoadConstant(w3, *(const uint8_t *)(OldCode + idx));
cmp(w2, w3);
csel(GetReg<RA_64>(Node), GetReg<RA_64>(Node), x1, Condition::eq);
ldrb(ARMEmitter::Reg::r2, ARMEmitter::Reg::r0, idx);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r3, *(const uint8_t *)(OldCode + idx));
cmp(ARMEmitter::Size::i32Bit, ARMEmitter::Reg::r2, ARMEmitter::Reg::r3);
csel(ARMEmitter::Size::i64Bit, Dst, Dst, ARMEmitter::Reg::r1, ARMEmitter::Condition::CC_EQ);
len -= 1;
idx += 1;
}
@@ -448,17 +399,17 @@ DEF_OP(ThreadRemoveCodeEntry) {
// X0: Thread
// X1: RIP
PushDynamicRegsAndLR();
PushDynamicRegsAndLR(TMP1);
mov(x0, STATE);
LoadConstant(x1, Entry);
mov(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, STATE.R());
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r1, Entry);
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.ThreadRemoveCodeEntryFromJIT)));
ldr(ARMEmitter::XReg::x2, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.ThreadRemoveCodeEntryFromJIT));
SpillStaticRegs();
#ifdef VIXL_SIMULATOR
GenerateIndirectRuntimeCall<void, void*, void*>(x2);
GenerateIndirectRuntimeCall<void, void*, void*>(ARMEmitter::Reg::r2);
#else
blr(x2);
blr(ARMEmitter::Reg::r2);
#endif
FillStaticRegs();
@@ -469,47 +420,33 @@ DEF_OP(ThreadRemoveCodeEntry) {
DEF_OP(CPUID) {
auto Op = IROp->C<IR::IROp_CPUID>();
PushDynamicRegsAndLR();
PushDynamicRegsAndLR(TMP1);
SpillStaticRegs();
// x0 = CPUID Handler
// x1 = CPUID Function
// x2 = CPUID Leaf
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.CPUIDObj)));
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.CPUIDFunction)));
mov(x1, GetReg<RA_64>(Op->Function.ID()));
mov(x2, GetReg<RA_64>(Op->Leaf.ID()));
ldr(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.CPUIDObj));
ldr(ARMEmitter::XReg::x3, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.CPUIDFunction));
mov(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r1, GetReg(Op->Function.ID()));
mov(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r2, GetReg(Op->Leaf.ID()));
#ifdef VIXL_SIMULATOR
GenerateIndirectRuntimeCall<__uint128_t, void*, uint64_t, uint64_t>(x3);
GenerateIndirectRuntimeCall<__uint128_t, void*, uint64_t, uint64_t>(ARMEmitter::Reg::r3);
#else
blr(x3);
blr(ARMEmitter::Reg::r3);
#endif
FillStaticRegs();
PopDynamicRegsAndLR();
// Results are in x0, x1
// Results want to be in a i64v2 vector
auto Dst = GetSrcPair<RA_64>(Node);
mov(Dst.first, x0);
mov(Dst.second, x1);
auto Dst = GetRegPair(Node);
mov(ARMEmitter::Size::i64Bit, Dst.first, ARMEmitter::Reg::r0);
mov(ARMEmitter::Size::i64Bit, Dst.second, ARMEmitter::Reg::r1);
}
#undef DEF_OP
void Arm64JITCore::RegisterBranchHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &Arm64JITCore::Op_##x
REGISTER_OP(SIGNALRETURN, SignalReturn);
REGISTER_OP(CALLBACKRETURN, CallbackReturn);
REGISTER_OP(EXITFUNCTION, ExitFunction);
REGISTER_OP(JUMP, Jump);
REGISTER_OP(CONDJUMP, CondJump);
REGISTER_OP(SYSCALL, Syscall);
REGISTER_OP(INLINESYSCALL, InlineSyscall);
REGISTER_OP(THUNK, Thunk);
REGISTER_OP(VALIDATECODE, ValidateCode);
REGISTER_OP(THREADREMOVECODEENTRY, ThreadRemoveCodeEntry);
REGISTER_OP(CPUID, CPUID);
#undef REGISTER_OP
}
}
@@ -4,12 +4,10 @@ tags: backend|arm64
$end_info$
*/
#include "Interface/Core/ArchHelpers/CodeEmitter/Emitter.h"
#include "Interface/Core/JIT/Arm64/JITClass.h"
namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const *IROp, IR::NodeID Node)
DEF_OP(VInsGPR) {
const auto Op = IROp->C<IR::IROp_VInsGPR>();
@@ -19,8 +17,16 @@ DEF_OP(VInsGPR) {
const auto ElementSize = Op->Header.ElementSize;
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto Dst = GetDst(Node);
const auto DestVector = GetSrc(Op->DestVector.ID());
LOGMAN_THROW_AA_FMT(ElementSize == 8 || ElementSize == 4 || ElementSize == 2 || ElementSize == 1, "Unexpected {} size", __func__);
const auto SubEmitSize = ElementSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
ElementSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
ElementSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
ElementSize == 1 ? ARMEmitter::SubRegSize::i8Bit : ARMEmitter::SubRegSize::i8Bit;
const auto ElementsPer128Bit = 16 / ElementSize;
const auto Dst = GetVReg(Node);
const auto DestVector = GetVReg(Op->DestVector.ID());
const auto Src = GetReg(Op->Src.ID());
if (HostSupportsSVE && Is256Bit) {
const auto ElementSizeBits = ElementSize * 8;
@@ -47,125 +53,111 @@ DEF_OP(VInsGPR) {
if (InUpperLane) {
// Move the upper lane down for the insertion.
const auto CompactPred = p0;
not_(CompactPred.VnB(), PRED_TMP_32B.Zeroing(), PRED_TMP_16B.VnB());
compact(VTMP1.Z().VnD(), CompactPred, DestVector.Z().VnD());
const auto CompactPred = ARMEmitter::PReg::p0;
not_(CompactPred, PRED_TMP_32B.Zeroing(), PRED_TMP_16B);
compact(ARMEmitter::SubRegSize::i64Bit, VTMP1.Z(), CompactPred, DestVector.Z());
}
// Put data in place for destructive SPLICE below.
mov(Dst.Z().VnD(), DestVector.Z().VnD());
mov(Dst.Z(), DestVector.Z());
// Inserts the GPR value into the given V register.
// Also automatically adjusts the index in the case of using the
// moved upper lane.
const auto Insert = [&](const aarch64::VRegister& reg, int index) {
switch (ElementSize) {
case 1:
if (InUpperLane) {
index -= 16;
}
ins(reg.V16B(), index, GetReg<RA_32>(Op->Src.ID()));
break;
case 2:
if (InUpperLane) {
index -= 8;
}
ins(reg.V8H(), index, GetReg<RA_32>(Op->Src.ID()));
break;
case 4:
if (InUpperLane) {
index -= 4;
}
ins(reg.V4S(), index, GetReg<RA_32>(Op->Src.ID()));
break;
case 8:
if (InUpperLane) {
index -= 2;
}
ins(reg.V2D(), index, GetReg<RA_64>(Op->Src.ID()));
break;
default:
LOGMAN_MSG_A_FMT("Unknown Element Size: {}", ElementSize);
break;
const auto Insert = [&](const FEXCore::ARMEmitter::VRegister& reg, int index) {
if (InUpperLane) {
index -= ElementsPer128Bit;
}
ins(SubEmitSize, reg, index, Src);
};
if (InUpperLane) {
Insert(VTMP1, DestIdx);
splice(Dst.Z().VnD(), PRED_TMP_16B, Dst.Z().VnD(), VTMP1.Z().VnD());
splice<ARMEmitter::OpType::Destructive>(ARMEmitter::SubRegSize::i64Bit, Dst.Z(), PRED_TMP_16B, Dst.Z(), VTMP1.Z());
} else {
Insert(Dst, DestIdx);
splice(Dst.Z().VnD(), PRED_TMP_16B, Dst.Z().VnD(), DestVector.Z().VnD());
splice<ARMEmitter::OpType::Destructive>(ARMEmitter::SubRegSize::i64Bit, Dst.Z(), PRED_TMP_16B, Dst.Z(), DestVector.Z());
}
} else {
mov(Dst, DestVector);
switch (ElementSize) {
case 1: {
ins(Dst.V16B(), DestIdx, GetReg<RA_32>(Op->Src.ID()));
break;
}
case 2: {
ins(Dst.V8H(), DestIdx, GetReg<RA_32>(Op->Src.ID()));
break;
}
case 4: {
ins(Dst.V4S(), DestIdx, GetReg<RA_32>(Op->Src.ID()));
break;
}
case 8: {
ins(Dst.V2D(), DestIdx, GetReg<RA_64>(Op->Src.ID()));
break;
}
default:
LOGMAN_MSG_A_FMT("Unknown Element Size: {}", ElementSize);
break;
}
mov(Dst.Q(), DestVector.Q());
ins(SubEmitSize, Dst, DestIdx, Src);
}
}
DEF_OP(VCastFromGPR) {
auto Op = IROp->C<IR::IROp_VCastFromGPR>();
auto Dst = GetVReg(Node);
auto Src = GetReg(Op->Src.ID());
switch (Op->Header.ElementSize) {
case 1:
uxtb(TMP1.W(), GetReg<RA_32>(Op->Src.ID()));
fmov(GetDst(Node).S(), TMP1.W());
uxtb(ARMEmitter::Size::i32Bit, TMP1, Src);
fmov(ARMEmitter::Size::i32Bit, Dst.S(), TMP1);
break;
case 2:
uxth(TMP1.W(), GetReg<RA_32>(Op->Src.ID()));
fmov(GetDst(Node).S(), TMP1.W());
uxth(ARMEmitter::Size::i32Bit, TMP1, Src);
fmov(ARMEmitter::Size::i32Bit, Dst.S(), TMP1);
break;
case 4:
fmov(GetDst(Node).S(), GetReg<RA_32>(Op->Src.ID()).W());
fmov(ARMEmitter::Size::i32Bit, Dst.S(), Src);
break;
case 8:
fmov(GetDst(Node).D(), GetReg<RA_64>(Op->Src.ID()).X());
fmov(ARMEmitter::Size::i64Bit, Dst.D(), Src);
break;
default: LOGMAN_MSG_A_FMT("Unknown castGPR element size: {}", Op->Header.ElementSize);
}
}
DEF_OP(VDupFromGPR) {
const auto Op = IROp->C<IR::IROp_VDupFromGPR>();
const auto OpSize = IROp->Size;
const auto Dst = GetVReg(Node);
const auto Src = GetReg(Op->Src.ID());
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto ElementSize = IROp->ElementSize;
LOGMAN_THROW_AA_FMT(ElementSize == 8 || ElementSize == 4 || ElementSize == 2 || ElementSize == 1,
"Unexpected {} element size: {}", __func__, ElementSize);
const auto SubEmitSize =
ElementSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
ElementSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
ElementSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
ElementSize == 1 ? ARMEmitter::SubRegSize::i8Bit : ARMEmitter::SubRegSize::i8Bit;
if (HostSupportsSVE && Is256Bit) {
dup(SubEmitSize, Dst.Z(), Src);
} else {
dup(SubEmitSize, Dst.Q(), Src);
}
}
DEF_OP(Float_FromGPR_S) {
const auto Op = IROp->C<IR::IROp_Float_FromGPR_S>();
const uint16_t ElementSize = Op->Header.ElementSize;
const uint16_t Conv = (ElementSize << 8) | Op->SrcElementSize;
auto Dst = GetVReg(Node);
auto Src = GetReg(Op->Src.ID());
switch (Conv) {
case 0x0404: { // Float <- int32_t
scvtf(GetDst(Node).S(), GetReg<RA_32>(Op->Src.ID()));
scvtf(ARMEmitter::Size::i32Bit, Dst.S(), Src);
break;
}
case 0x0408: { // Float <- int64_t
scvtf(GetDst(Node).S(), GetReg<RA_64>(Op->Src.ID()));
scvtf(ARMEmitter::Size::i64Bit, Dst.S(), Src);
break;
}
case 0x0804: { // Double <- int32_t
scvtf(GetDst(Node).D(), GetReg<RA_32>(Op->Src.ID()));
scvtf(ARMEmitter::Size::i32Bit, Dst.D(), Src);
break;
}
case 0x0808: { // Double <- int64_t
scvtf(GetDst(Node).D(), GetReg<RA_64>(Op->Src.ID()));
scvtf(ARMEmitter::Size::i64Bit, Dst.D(), Src);
break;
}
default:
@@ -178,13 +170,17 @@ DEF_OP(Float_FromGPR_S) {
DEF_OP(Float_FToF) {
auto Op = IROp->C<IR::IROp_Float_FToF>();
const uint16_t Conv = (Op->Header.ElementSize << 8) | Op->SrcElementSize;
auto Dst = GetVReg(Node);
auto Src = GetVReg(Op->Scalar.ID());
switch (Conv) {
case 0x0804: { // Double <- Float
fcvt(GetDst(Node).D(), GetSrc(Op->Scalar.ID()).S());
fcvt(Dst.D(), Src.S());
break;
}
case 0x0408: { // Float <- Double
fcvt(GetDst(Node).S(), GetSrc(Op->Scalar.ID()).D());
fcvt(Dst.S(), Src.D());
break;
}
default: LOGMAN_MSG_A_FMT("Unknown FCVT sizes: 0x{:x}", Conv);
@@ -198,41 +194,18 @@ DEF_OP(Vector_SToF) {
const auto ElementSize = Op->Header.ElementSize;
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto Dst = GetDst(Node);
const auto Vector = GetSrc(Op->Vector.ID());
LOGMAN_THROW_AA_FMT(ElementSize == 8 || ElementSize == 4 || ElementSize == 2, "Unexpected {} size", __func__);
const auto SubEmitSize = ElementSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
ElementSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
ElementSize == 2 ? ARMEmitter::SubRegSize::i16Bit : ARMEmitter::SubRegSize::i16Bit;
const auto Dst = GetVReg(Node);
const auto Vector = GetVReg(Op->Vector.ID());
if (HostSupportsSVE && Is256Bit) {
const auto Mask = PRED_TMP_32B.Merging();
switch (ElementSize) {
case 2:
scvtf(Dst.Z().VnH(), Mask, Vector.Z().VnH());
break;
case 4:
scvtf(Dst.Z().VnS(), Mask, Vector.Z().VnS());
break;
case 8:
scvtf(Dst.Z().VnD(), Mask, Vector.Z().VnD());
break;
default:
LOGMAN_MSG_A_FMT("Unknown Vector_SToF element size: {}", ElementSize);
break;
}
const auto Mask = PRED_TMP_32B;
scvtf(Dst.Z(), SubEmitSize, Mask.Merging(), Vector.Z(), SubEmitSize);
} else {
switch (ElementSize) {
case 2:
scvtf(Dst.V8H(), Vector.V8H());
break;
case 4:
scvtf(Dst.V4S(), Vector.V4S());
break;
case 8:
scvtf(Dst.V2D(), Vector.V2D());
break;
default:
LOGMAN_MSG_A_FMT("Unknown Vector_SToF element size: {}", ElementSize);
break;
}
scvtf(SubEmitSize, Dst.Q(), Vector.Q());
}
}
@@ -243,41 +216,18 @@ DEF_OP(Vector_FToZS) {
const auto ElementSize = Op->Header.ElementSize;
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto Dst = GetDst(Node);
const auto Vector = GetSrc(Op->Vector.ID());
LOGMAN_THROW_AA_FMT(ElementSize == 8 || ElementSize == 4 || ElementSize == 2, "Unexpected {} size", __func__);
const auto SubEmitSize = ElementSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
ElementSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
ElementSize == 2 ? ARMEmitter::SubRegSize::i16Bit : ARMEmitter::SubRegSize::i16Bit;
const auto Dst = GetVReg(Node);
const auto Vector = GetVReg(Op->Vector.ID());
if (HostSupportsSVE && Is256Bit) {
const auto Mask = PRED_TMP_32B.Merging();
switch (ElementSize) {
case 2:
fcvtzs(Dst.Z().VnH(), Mask, Vector.Z().VnH());
break;
case 4:
fcvtzs(Dst.Z().VnS(), Mask, Vector.Z().VnS());
break;
case 8:
fcvtzs(Dst.Z().VnD(), Mask, Vector.Z().VnD());
break;
default:
LOGMAN_MSG_A_FMT("Unknown Vector_FToZS element size: {}", ElementSize);
break;
}
const auto Mask = PRED_TMP_32B;
fcvtzs(Dst.Z(), SubEmitSize, Mask.Merging(), Vector.Z(), SubEmitSize);
} else {
switch (ElementSize) {
case 2:
fcvtzs(Dst.V8H(), Vector.V8H());
break;
case 4:
fcvtzs(Dst.V4S(), Vector.V4S());
break;
case 8:
fcvtzs(Dst.V2D(), Vector.V2D());
break;
default:
LOGMAN_MSG_A_FMT("Unknown Vector_FToZS element size: {}", ElementSize);
break;
}
fcvtzs(SubEmitSize, Dst.Q(), Vector.Q());
}
}
@@ -288,47 +238,23 @@ DEF_OP(Vector_FToS) {
const auto ElementSize = Op->Header.ElementSize;
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto Dst = GetDst(Node);
const auto Vector = GetSrc(Op->Vector.ID());
LOGMAN_THROW_AA_FMT(ElementSize == 8 || ElementSize == 4 || ElementSize == 2, "Unexpected {} size", __func__);
const auto SubEmitSize = ElementSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
ElementSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
ElementSize == 2 ? ARMEmitter::SubRegSize::i16Bit : ARMEmitter::SubRegSize::i16Bit;
const auto Dst = GetVReg(Node);
const auto Vector = GetVReg(Op->Vector.ID());
if (HostSupportsSVE && Is256Bit) {
const auto Mask = PRED_TMP_32B.Merging();
switch (ElementSize) {
case 2:
frinti(Dst.Z().VnH(), Mask, Vector.Z().VnH());
fcvtzs(Dst.Z().VnH(), Mask, Dst.Z().VnH());
break;
case 4:
frinti(Dst.Z().VnS(), Mask, Vector.Z().VnS());
fcvtzs(Dst.Z().VnS(), Mask, Dst.Z().VnS());
break;
case 8:
frinti(Dst.Z().VnD(), Mask, Vector.Z().VnD());
fcvtzs(Dst.Z().VnD(), Mask, Dst.Z().VnD());
break;
default:
LOGMAN_MSG_A_FMT("Unknown Vector_FToS element size: {}", ElementSize);
break;
}
const auto Mask = PRED_TMP_32B;
frinti(SubEmitSize, Dst.Z(), Mask.Merging(), Vector.Z());
fcvtzs(Dst.Z(), SubEmitSize, Mask.Merging(), Dst.Z(), SubEmitSize);
} else {
switch (ElementSize) {
case 2:
frinti(Dst.V8H(), Vector.V8H());
fcvtzs(Dst.V8H(), Dst.V8H());
break;
case 4:
frinti(Dst.V4S(), Vector.V4S());
fcvtzs(Dst.V4S(), Dst.V4S());
break;
case 8:
frinti(Dst.V2D(), Vector.V2D());
fcvtzs(Dst.V2D(), Dst.V2D());
break;
default:
LOGMAN_MSG_A_FMT("Unknown Vector_FToS element size: {}", ElementSize);
break;
}
const auto Dst = GetVReg(Node);
const auto Vector = GetVReg(Op->Vector.ID());
frinti(SubEmitSize, Dst.Q(), Vector.Q());
fcvtzs(SubEmitSize, Dst.Q(), Dst.Q());
}
}
@@ -340,8 +266,13 @@ DEF_OP(Vector_FToF) {
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto Conv = (ElementSize << 8) | Op->SrcElementSize;
const auto Dst = GetDst(Node);
const auto Vector = GetSrc(Op->Vector.ID());
LOGMAN_THROW_AA_FMT(ElementSize == 8 || ElementSize == 4 || ElementSize == 2, "Unexpected {} size", __func__);
const auto SubEmitSize = ElementSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
ElementSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
ElementSize == 2 ? ARMEmitter::SubRegSize::i16Bit : ARMEmitter::SubRegSize::i16Bit;
const auto Dst = GetVReg(Node);
const auto Vector = GetVReg(Op->Vector.ID());
if (HostSupportsSVE && Is256Bit) {
// Curiously, FCVTLT and FCVTNT have no bottom variants,
@@ -361,23 +292,23 @@ DEF_OP(Vector_FToF) {
switch (Conv) {
case 0x0402: { // Float <- Half
zip1(Dst.Z().VnH(), Vector.Z().VnH(), Vector.Z().VnH());
fcvtlt(Dst.Z().VnS(), Mask, Dst.Z().VnH());
zip1(FEXCore::ARMEmitter::SubRegSize::i16Bit, Dst.Z(), Vector.Z(), Vector.Z());
fcvtlt(FEXCore::ARMEmitter::SubRegSize::i32Bit, Dst.Z(), Mask, Dst.Z());
break;
}
case 0x0804: { // Double <- Float
zip1(Dst.Z().VnS(), Vector.Z().VnS(), Vector.Z().VnS());
fcvtlt(Dst.Z().VnD(), Mask, Dst.Z().VnS());
zip1(FEXCore::ARMEmitter::SubRegSize::i32Bit, Dst.Z(), Vector.Z(), Vector.Z());
fcvtlt(FEXCore::ARMEmitter::SubRegSize::i64Bit, Dst.Z(), Mask, Dst.Z());
break;
}
case 0x0204: { // Half <- Float
fcvtnt(Dst.Z().VnH(), Mask, Vector.Z().VnS());
uzp2(Dst.Z().VnH(), Dst.Z().VnH(), Dst.Z().VnH());
fcvtnt(FEXCore::ARMEmitter::SubRegSize::i16Bit, Dst.Z(), Mask, Vector.Z());
uzp2(FEXCore::ARMEmitter::SubRegSize::i16Bit, Dst.Z(), Dst.Z(), Dst.Z());
break;
}
case 0x0408: { // Float <- Double
fcvtnt(Dst.Z().VnS(), Mask, Vector.Z().VnD());
uzp2(Dst.Z().VnS(), Dst.Z().VnS(), Dst.Z().VnS());
fcvtnt(FEXCore::ARMEmitter::SubRegSize::i32Bit, Dst.Z(), Mask, Vector.Z());
uzp2(FEXCore::ARMEmitter::SubRegSize::i32Bit, Dst.Z(), Dst.Z(), Dst.Z());
break;
}
default:
@@ -386,20 +317,14 @@ DEF_OP(Vector_FToF) {
}
} else {
switch (Conv) {
case 0x0402: { // Float <- Half
fcvtl(Dst.V4S(), Vector.V4H());
break;
}
case 0x0402: // Float <- Half
case 0x0804: { // Double <- Float
fcvtl(Dst.V2D(), Vector.V2S());
break;
}
case 0x0204: { // Half <- Float
fcvtn(Dst.V4H(), Vector.V4S());
fcvtl(SubEmitSize, Dst.D(), Vector.D());
break;
}
case 0x0204: // Half <- Float
case 0x0408: { // Float <- Double
fcvtn(Dst.V2S(), Vector.V2D());
fcvtn(SubEmitSize, Dst.D(), Vector.D());
break;
}
default:
@@ -415,172 +340,56 @@ DEF_OP(Vector_FToI) {
const auto ElementSize = Op->Header.ElementSize;
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
LOGMAN_THROW_AA_FMT(ElementSize == 8 || ElementSize == 4 || ElementSize == 2, "Unexpected {} size", __func__);
const auto Dst = GetDst(Node);
const auto Vector = GetSrc(Op->Vector.ID());
const auto SubEmitSize = ElementSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
ElementSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
ElementSize == 2 ? ARMEmitter::SubRegSize::i16Bit : ARMEmitter::SubRegSize::i16Bit;
const auto Dst = GetVReg(Node);
const auto Vector = GetVReg(Op->Vector.ID());
if (HostSupportsSVE && Is256Bit) {
const auto Mask = PRED_TMP_32B.Merging();
switch (Op->Round) {
case FEXCore::IR::Round_Nearest.Val:
switch (ElementSize) {
case 2:
frintn(Dst.Z().VnH(), Mask, Vector.Z().VnH());
break;
case 4:
frintn(Dst.Z().VnS(), Mask, Vector.Z().VnS());
break;
case 8:
frintn(Dst.Z().VnD(), Mask, Vector.Z().VnD());
break;
}
frintn(SubEmitSize, Dst.Z(), Mask, Vector.Z());
break;
case FEXCore::IR::Round_Negative_Infinity.Val:
switch (ElementSize) {
case 2:
frintm(Dst.Z().VnH(), Mask, Vector.Z().VnH());
break;
case 4:
frintm(Dst.Z().VnS(), Mask, Vector.Z().VnS());
break;
case 8:
frintm(Dst.Z().VnD(), Mask, Vector.Z().VnD());
break;
}
frintm(SubEmitSize, Dst.Z(), Mask, Vector.Z());
break;
case FEXCore::IR::Round_Positive_Infinity.Val:
switch (ElementSize) {
case 2:
frintp(Dst.Z().VnH(), Mask, Vector.Z().VnH());
break;
case 4:
frintp(Dst.Z().VnS(), Mask, Vector.Z().VnS());
break;
case 8:
frintp(Dst.Z().VnD(), Mask, Vector.Z().VnD());
break;
}
frintp(SubEmitSize, Dst.Z(), Mask, Vector.Z());
break;
case FEXCore::IR::Round_Towards_Zero.Val:
switch (ElementSize) {
case 2:
frintz(Dst.Z().VnH(), Mask, Vector.Z().VnH());
break;
case 4:
frintz(Dst.Z().VnS(), Mask, Vector.Z().VnS());
break;
case 8:
frintz(Dst.Z().VnD(), Mask, Vector.Z().VnD());
break;
}
frintz(SubEmitSize, Dst.Z(), Mask, Vector.Z());
break;
case FEXCore::IR::Round_Host.Val:
switch (ElementSize) {
case 2:
frinti(Dst.Z().VnH(), Mask, Vector.Z().VnH());
break;
case 4:
frinti(Dst.Z().VnS(), Mask, Vector.Z().VnS());
break;
case 8:
frinti(Dst.Z().VnD(), Mask, Vector.Z().VnD());
break;
}
frinti(SubEmitSize, Dst.Z(), Mask, Vector.Z());
break;
}
} else {
switch (Op->Round) {
case FEXCore::IR::Round_Nearest.Val:
switch (ElementSize) {
case 2:
frintn(Dst.V8H(), Vector.V8H());
break;
case 4:
frintn(Dst.V4S(), Vector.V4S());
break;
case 8:
frintn(Dst.V2D(), Vector.V2D());
break;
}
frinti(SubEmitSize, Dst.Q(), Vector.Q());
break;
case FEXCore::IR::Round_Negative_Infinity.Val:
switch (ElementSize) {
case 2:
frintm(Dst.V8H(), Vector.V8H());
break;
case 4:
frintm(Dst.V4S(), Vector.V4S());
break;
case 8:
frintm(Dst.V2D(), Vector.V2D());
break;
}
frintm(SubEmitSize, Dst.Q(), Vector.Q());
break;
case FEXCore::IR::Round_Positive_Infinity.Val:
switch (ElementSize) {
case 2:
frintp(Dst.V8H(), Vector.V8H());
break;
case 4:
frintp(Dst.V4S(), Vector.V4S());
break;
case 8:
frintp(Dst.V2D(), Vector.V2D());
break;
}
frintp(SubEmitSize, Dst.Q(), Vector.Q());
break;
case FEXCore::IR::Round_Towards_Zero.Val:
switch (ElementSize) {
case 2:
frintz(Dst.V8H(), Vector.V8H());
break;
case 4:
frintz(Dst.V4S(), Vector.V4S());
break;
case 8:
frintz(Dst.V2D(), Vector.V2D());
break;
}
frintz(SubEmitSize, Dst.Q(), Vector.Q());
break;
case FEXCore::IR::Round_Host.Val:
switch (ElementSize) {
case 2:
frinti(Dst.V8H(), Vector.V8H());
break;
case 4:
frinti(Dst.V4S(), Vector.V4S());
break;
case 8:
frinti(Dst.V2D(), Vector.V2D());
break;
}
frinti(SubEmitSize, Dst.Q(), Vector.Q());
break;
}
}
}
#undef DEF_OP
void Arm64JITCore::RegisterConversionHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &Arm64JITCore::Op_##x
REGISTER_OP(VINSGPR, VInsGPR);
REGISTER_OP(VCASTFROMGPR, VCastFromGPR);
REGISTER_OP(FLOAT_FROMGPR_S, Float_FromGPR_S);
REGISTER_OP(FLOAT_FTOF, Float_FToF);
REGISTER_OP(VECTOR_STOF, Vector_SToF);
REGISTER_OP(VECTOR_FTOZS, Vector_FToZS);
REGISTER_OP(VECTOR_FTOS, Vector_FToS);
REGISTER_OP(VECTOR_FTOF, Vector_FToF);
REGISTER_OP(VECTOR_FTOI, Vector_FToI);
#undef REGISTER_OP
}
}
@@ -4,124 +4,170 @@ tags: backend|arm64
$end_info$
*/
#include "Interface/Core/ArchHelpers/CodeEmitter/Emitter.h"
#include "Interface/Core/JIT/Arm64/JITClass.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const *IROp, IR::NodeID Node)
DEF_OP(AESImc) {
auto Op = IROp->C<IR::IROp_VAESImc>();
aesimc(GetDst(Node).V16B(), GetSrc(Op->Vector.ID()).V16B());
aesimc(GetVReg(Node), GetVReg(Op->Vector.ID()));
}
DEF_OP(AESEnc) {
auto Op = IROp->C<IR::IROp_VAESEnc>();
eor(VTMP2.V16B(), VTMP2.V16B(), VTMP2.V16B());
mov(VTMP1.V16B(), GetSrc(Op->State.ID()).V16B());
aese(VTMP1.V16B(), VTMP2.V16B());
aesmc(VTMP1.V16B(), VTMP1.V16B());
eor(GetDst(Node).V16B(), VTMP1.V16B(), GetSrc(Op->Key.ID()).V16B());
const auto Op = IROp->C<IR::IROp_VAESEnc>();
const auto OpSize = IROp->Size;
const auto Dst = GetVReg(Node);
const auto Key = GetVReg(Op->Key.ID());
const auto State = GetVReg(Op->State.ID());
LOGMAN_THROW_AA_FMT(OpSize == Core::CPUState::XMM_SSE_REG_SIZE,
"Currently only supports 128-bit operations.");
eor(VTMP2.Q(), VTMP2.Q(), VTMP2.Q());
mov(VTMP1.Q(), State.Q());
aese(VTMP1, VTMP2);
aesmc(VTMP1, VTMP1);
eor(Dst.Q(), VTMP1.Q(), Key.Q());
}
DEF_OP(AESEncLast) {
auto Op = IROp->C<IR::IROp_VAESEncLast>();
eor(VTMP2.V16B(), VTMP2.V16B(), VTMP2.V16B());
mov(VTMP1.V16B(), GetSrc(Op->State.ID()).V16B());
aese(VTMP1.V16B(), VTMP2.V16B());
eor(GetDst(Node).V16B(), VTMP1.V16B(), GetSrc(Op->Key.ID()).V16B());
const auto Op = IROp->C<IR::IROp_VAESEncLast>();
const auto OpSize = IROp->Size;
const auto Dst = GetVReg(Node);
const auto Key = GetVReg(Op->Key.ID());
const auto State = GetVReg(Op->State.ID());
LOGMAN_THROW_AA_FMT(OpSize == Core::CPUState::XMM_SSE_REG_SIZE,
"Currently only supports 128-bit operations.");
eor(VTMP2.Q(), VTMP2.Q(), VTMP2.Q());
mov(VTMP1.Q(), State.Q());
aese(VTMP1, VTMP2);
eor(Dst.Q(), VTMP1.Q(), Key.Q());
}
DEF_OP(AESDec) {
auto Op = IROp->C<IR::IROp_VAESDec>();
eor(VTMP2.V16B(), VTMP2.V16B(), VTMP2.V16B());
mov(VTMP1.V16B(), GetSrc(Op->State.ID()).V16B());
aesd(VTMP1.V16B(), VTMP2.V16B());
aesimc(VTMP1.V16B(), VTMP1.V16B());
eor(GetDst(Node).V16B(), VTMP1.V16B(), GetSrc(Op->Key.ID()).V16B());
const auto Op = IROp->C<IR::IROp_VAESDec>();
const auto OpSize = IROp->Size;
const auto Dst = GetVReg(Node);
const auto Key = GetVReg(Op->Key.ID());
const auto State = GetVReg(Op->State.ID());
LOGMAN_THROW_AA_FMT(OpSize == Core::CPUState::XMM_SSE_REG_SIZE,
"Currently only supports 128-bit operations.");
eor(VTMP2.Q(), VTMP2.Q(), VTMP2.Q());
mov(VTMP1.Q(), State.Q());
aesd(VTMP1, VTMP2);
aesimc(VTMP1, VTMP1);
eor(Dst.Q(), VTMP1.Q(), Key.Q());
}
DEF_OP(AESDecLast) {
auto Op = IROp->C<IR::IROp_VAESDecLast>();
eor(VTMP2.V16B(), VTMP2.V16B(), VTMP2.V16B());
mov(VTMP1.V16B(), GetSrc(Op->State.ID()).V16B());
aesd(VTMP1.V16B(), VTMP2.V16B());
eor(GetDst(Node).V16B(), VTMP1.V16B(), GetSrc(Op->Key.ID()).V16B());
const auto Op = IROp->C<IR::IROp_VAESDecLast>();
const auto OpSize = IROp->Size;
const auto Dst = GetVReg(Node);
const auto Key = GetVReg(Op->Key.ID());
const auto State = GetVReg(Op->State.ID());
LOGMAN_THROW_AA_FMT(OpSize == Core::CPUState::XMM_SSE_REG_SIZE,
"Currently only supports 128-bit operations.");
eor(VTMP2.Q(), VTMP2.Q(), VTMP2.Q());
mov(VTMP1.Q(), State.Q());
aesd(VTMP1, VTMP2);
eor(Dst.Q(), VTMP1.Q(), Key.Q());
}
DEF_OP(AESKeyGenAssist) {
auto Op = IROp->C<IR::IROp_VAESKeyGenAssist>();
aarch64::Literal ConstantLiteral (0x0C030609'0306090CULL, 0x040B0E01'0B0E0104ULL);
aarch64::Label PastConstant;
ARMEmitter::ForwardLabel Constant;
ARMEmitter::ForwardLabel PastConstant;
// Do a "regular" AESE step
eor(VTMP2.V16B(), VTMP2.V16B(), VTMP2.V16B());
mov(VTMP1.V16B(), GetSrc(Op->Src.ID()).V16B());
aese(VTMP1.V16B(), VTMP2.V16B());
eor(VTMP2.Q(), VTMP2.Q(), VTMP2.Q());
mov(VTMP1.Q(), GetVReg(Op->Src.ID()).Q());
aese(VTMP1, VTMP2);
// Do a table shuffle to undo ShiftRows
ldr(VTMP3, &ConstantLiteral);
ldr(VTMP3.Q(), &Constant);
// Now EOR in the RCON
if (Op->RCON) {
tbl(VTMP1.V16B(), VTMP1.V16B(), VTMP3.V16B());
tbl(VTMP1.Q(), VTMP1.Q(), VTMP3.Q());
LoadConstant(TMP1, static_cast<uint64_t>(Op->RCON) << 32);
dup(VTMP2.V2D(), TMP1);
eor(GetDst(Node).V16B(), VTMP1.V16B(), VTMP2.V16B());
LoadConstant(ARMEmitter::Size::i64Bit, TMP1, static_cast<uint64_t>(Op->RCON) << 32);
dup(ARMEmitter::SubRegSize::i64Bit, VTMP2.Q(), TMP1);
eor(GetVReg(Node).Q(), VTMP1.Q(), VTMP2.Q());
}
else {
tbl(GetDst(Node).V16B(), VTMP1.V16B(), VTMP3.V16B());
tbl(GetVReg(Node).Q(), VTMP1.Q(), VTMP3.Q());
}
b(&PastConstant);
place(&ConstantLiteral);
bind(&PastConstant);
Bind(&Constant);
dc64(0x040B0E01'0B0E0104ULL);
dc64(0x0C030609'0306090CULL);
Bind(&PastConstant);
}
DEF_OP(CRC32) {
auto Op = IROp->C<IR::IROp_CRC32>();
const auto Dst = GetReg(Node);
const auto Src1 = GetReg(Op->Src1.ID());
const auto Src2 = GetReg(Op->Src2.ID());
switch (Op->SrcSize) {
case 1:
crc32cb(GetReg<RA_32>(Node), GetReg<RA_32>(Op->Src1.ID()), GetReg<RA_32>(Op->Src2.ID()));
crc32cb(Dst.W(), Src1.W(), Src2.W());
break;
case 2:
crc32ch(GetReg<RA_32>(Node), GetReg<RA_32>(Op->Src1.ID()), GetReg<RA_32>(Op->Src2.ID()));
crc32ch(Dst.W(), Src1.W(), Src2.W());
break;
case 4:
crc32cw(GetReg<RA_32>(Node), GetReg<RA_32>(Op->Src1.ID()), GetReg<RA_32>(Op->Src2.ID()));
crc32cw(Dst.W(), Src1.W(), Src2.W());
break;
case 8:
crc32cx(GetReg<RA_32>(Node), GetReg<RA_32>(Op->Src1.ID()), GetReg<RA_64>(Op->Src2.ID()));
crc32cx(Dst.X(), Src1.X(), Src2.X());
break;
default: LOGMAN_MSG_A_FMT("Unknown CRC32 size: {}", Op->SrcSize);
}
}
DEF_OP(PCLMUL) {
auto Op = IROp->C<IR::IROp_PCLMUL>();
const auto Op = IROp->C<IR::IROp_PCLMUL>();
const auto OpSize = IROp->Size;
auto Dst = GetDst(Node).Q();
auto Src1 = GetSrc(Op->Src1.ID()).V2D();
auto Src2 = GetSrc(Op->Src2.ID()).V2D();
const auto Dst = GetVReg(Node);
const auto Src1 = GetVReg(Op->Src1.ID());
const auto Src2 = GetVReg(Op->Src2.ID());
LOGMAN_THROW_AA_FMT(OpSize == Core::CPUState::XMM_SSE_REG_SIZE,
"Currently only supports 128-bit operations.");
switch (Op->Selector) {
case 0b00000000:
pmull(Dst, Src1, Src2);
pmull(ARMEmitter::SubRegSize::i128Bit, Dst.D(), Src1.D(), Src2.D());
break;
case 0b00000001:
mov(VTMP1.V1D(), Src1, 1);
pmull(Dst, VTMP1.V2D(), Src2);
dup(ARMEmitter::SubRegSize::i64Bit, VTMP1.Q(), Src1.Q(), 1);
pmull(ARMEmitter::SubRegSize::i128Bit, Dst.D(), VTMP1.D(), Src2.D());
break;
case 0b00010000:
mov(VTMP1.V1D(), Src2, 1);
pmull(Dst, VTMP1.V2D(), Src1);
dup(ARMEmitter::SubRegSize::i64Bit, VTMP1.Q(), Src2.Q(), 1);
pmull(ARMEmitter::SubRegSize::i128Bit, Dst.D(), VTMP1.D(), Src1.D());
break;
case 0b00010001:
pmull2(Dst, Src1, Src2);
pmull2(ARMEmitter::SubRegSize::i128Bit, Dst.Q(), Src1.Q(), Src2.Q());
break;
default:
LOGMAN_MSG_A_FMT("Unknown PCLMUL selector: {}", Op->Selector);
@@ -130,16 +176,4 @@ DEF_OP(PCLMUL) {
}
#undef DEF_OP
void Arm64JITCore::RegisterEncryptionHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &Arm64JITCore::Op_##x
REGISTER_OP(VAESIMC, AESImc);
REGISTER_OP(VAESENC, AESEnc);
REGISTER_OP(VAESENCLAST, AESEncLast);
REGISTER_OP(VAESDEC, AESDec);
REGISTER_OP(VAESDECLAST, AESDecLast);
REGISTER_OP(VAESKEYGENASSIST, AESKeyGenAssist);
REGISTER_OP(CRC32, CRC32);
REGISTER_OP(PCLMUL, PCLMUL);
#undef REGISTER_OP
}
}
@@ -7,20 +7,12 @@ $end_info$
#include "Interface/Core/JIT/Arm64/JITClass.h"
namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const *IROp, IR::NodeID Node)
DEF_OP(GetHostFlag) {
auto Op = IROp->C<IR::IROp_GetHostFlag>();
ubfx(GetReg<RA_64>(Node), GetReg<RA_64>(Op->Value.ID()), Op->Flag, 1);
ubfx(ARMEmitter::Size::i64Bit, GetReg(Node), GetReg(Op->Value.ID()), Op->Flag, 1);
}
#undef DEF_OP
void Arm64JITCore::RegisterFlagHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &Arm64JITCore::Op_##x
REGISTER_OP(GETHOSTFLAG, GetHostFlag);
#undef REGISTER_OP
}
}
File diff suppressed because it is too large. Load diff
+59 -49
View File
@@ -8,6 +8,7 @@ $end_info$
#include <FEXCore/IR/RegisterAllocationData.h>
#include "Interface/Core/ArchHelpers/Arm64Emitter.h"
#include "Interface/Core/ArchHelpers/CodeEmitter/Emitter.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include <aarch64/assembler-aarch64.h>
@@ -28,18 +29,15 @@ namespace FEXCore::Core {
}
namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
class Arm64JITCore final : public CPUBackend, public Arm64Emitter {
public:
explicit Arm64JITCore(FEXCore::Context::Context *ctx,
explicit Arm64JITCore(FEXCore::Context::ContextImpl *ctx,
FEXCore::Core::InternalThreadState *Thread);
~Arm64JITCore() override;
[[nodiscard]] std::string GetName() override { return "JIT"; }
[[nodiscard]] void *CompileCode(uint64_t Entry,
[[nodiscard]] CPUBackend::CompiledCode CompileCode(uint64_t Entry,
FEXCore::IR::IRListView const *IR,
FEXCore::Core::DebugData *DebugData,
FEXCore::IR::RegisterAllocationData *RAData, bool GDBEnabled) override;
@@ -50,7 +48,7 @@ public:
void ClearCache() override;
static void InitializeSignalHandlers(FEXCore::Context::Context *CTX);
static void InitializeSignalHandlers(FEXCore::Context::ContextImpl *CTX);
void ClearRelocations() override { Relocations.clear(); }
@@ -58,12 +56,13 @@ private:
FEX_CONFIG_OPT(ParanoidTSO, PARANOIDTSO);
const bool HostSupportsSVE{};
Label *PendingTargetLabel;
FEXCore::Context::Context *CTX;
ARMEmitter::BiDirectionalLabel *PendingTargetLabel;
FEXCore::Context::ContextImpl *CTX;
FEXCore::IR::IRListView const *IR;
uint64_t Entry;
CPUBackend::CompiledCode CodeData{};
std::map<IR::NodeID, aarch64::Label> JumpTargets;
std::map<IR::NodeID, ARMEmitter::BiDirectionalLabel> JumpTargets;
/**
* @name Register Allocation
@@ -86,34 +85,57 @@ private:
constexpr static uint8_t RA_64 = 1;
constexpr static uint8_t RA_FPR = 2;
template<uint8_t RAType>
[[nodiscard]] aarch64::Register GetReg(IR::NodeID Node) const;
[[nodiscard]] FEXCore::ARMEmitter::Register GetReg(IR::NodeID Node) const {
const auto Reg = GetPhys(Node);
template<>
[[nodiscard]] aarch64::Register GetReg<RA_32>(IR::NodeID Node) const;
template<>
[[nodiscard]] aarch64::Register GetReg<RA_64>(IR::NodeID Node) const;
LOGMAN_THROW_AA_FMT(Reg.Class == IR::GPRFixedClass.Val || Reg.Class == IR::GPRClass.Val, "Unexpected Class: {}", Reg.Class);
template<uint8_t RAType>
[[nodiscard]] std::pair<aarch64::Register, aarch64::Register> GetSrcPair(IR::NodeID Node) const;
if (Reg.Class == IR::GPRFixedClass.Val) {
return SRA64[Reg.Reg];
} else if (Reg.Class == IR::GPRClass.Val) {
return RA64[Reg.Reg];
}
template<>
[[nodiscard]] std::pair<aarch64::Register, aarch64::Register> GetSrcPair<RA_32>(IR::NodeID Node) const;
template<>
[[nodiscard]] std::pair<aarch64::Register, aarch64::Register> GetSrcPair<RA_64>(IR::NodeID Node) const;
FEX_UNREACHABLE;
}
[[nodiscard]] aarch64::VRegister GetSrc(IR::NodeID Node) const;
[[nodiscard]] aarch64::VRegister GetDst(IR::NodeID Node) const;
[[nodiscard]] FEXCore::ARMEmitter::VRegister GetVReg(IR::NodeID Node) const {
const auto Reg = GetPhys(Node);
LOGMAN_THROW_AA_FMT(Reg.Class == IR::FPRFixedClass.Val || Reg.Class == IR::FPRClass.Val, "Unexpected Class: {}", Reg.Class);
if (Reg.Class == IR::FPRFixedClass.Val) {
return SRAFPR[Reg.Reg];
} else if (Reg.Class == IR::FPRClass.Val) {
return RAFPR[Reg.Reg];
}
FEX_UNREACHABLE;
}
[[nodiscard]] std::pair<FEXCore::ARMEmitter::Register, FEXCore::ARMEmitter::Register> GetRegPair(IR::NodeID Node) const {
const auto Reg = GetPhys(Node);
LOGMAN_THROW_AA_FMT(Reg.Class == IR::GPRPairClass.Val, "Unexpected Class: {}", Reg.Class);
return RA64Pair[Reg.Reg];
}
[[nodiscard]] FEXCore::IR::RegisterClassType GetRegClass(IR::NodeID Node) const;
[[nodiscard]] IR::PhysicalRegister GetPhys(IR::NodeID Node) const;
[[nodiscard]] IR::PhysicalRegister GetPhys(IR::NodeID Node) const {
auto PhyReg = RAData->GetNodeRegister(Node);
LOGMAN_THROW_A_FMT(!PhyReg.IsInvalid(), "Couldn't Allocate register for node: ssa{}. Class: {}", Node, PhyReg.Class);
return PhyReg;
}
[[nodiscard]] bool IsFPR(IR::NodeID Node) const;
[[nodiscard]] bool IsGPR(IR::NodeID Node) const;
[[nodiscard]] MemOperand GenerateMemOperand(uint8_t AccessSize,
aarch64::Register Base,
[[nodiscard]] FEXCore::ARMEmitter::ExtendedMemOperand GenerateMemOperand(uint8_t AccessSize,
FEXCore::ARMEmitter::Register Base,
IR::OrderedNodeWrapper Offset,
IR::MemOffsetType OffsetType,
uint8_t OffsetScale);
@@ -123,8 +145,8 @@ private:
//
// TMP1 is safe to use again once this memory operand is used with its
// equivalent loads or stores that this was called for.
[[nodiscard]] SVEMemOperand GenerateSVEMemOperand(uint8_t AccessSize,
aarch64::Register Base,
[[nodiscard]] FEXCore::ARMEmitter::SVEMemOperand GenerateSVEMemOperand(uint8_t AccessSize,
FEXCore::ARMEmitter::Register Base,
IR::OrderedNodeWrapper Offset,
IR::MemOffsetType OffsetType,
uint8_t OffsetScale);
@@ -162,7 +184,8 @@ private:
* @brief A literal pair relocation object for named symbol literals
*/
struct NamedSymbolLiteralPair {
Literal<uint64_t> Lit;
ARMEmitter::ForwardLabel Loc;
uint64_t Lit;
Relocation MoveABI{};
};
@@ -172,7 +195,7 @@ private:
* @param Reg - The GPR to move the thunk handler in to
* @param Sum - The hash of the thunk
*/
void InsertNamedThunkRelocation(vixl::aarch64::Register Reg, const IR::SHA256Sum &Sum);
void InsertNamedThunkRelocation(ARMEmitter::Register Reg, const IR::SHA256Sum &Sum);
/**
* @brief Inserts a guest GPR move relocation
@@ -180,7 +203,7 @@ private:
* @param Reg - The GPR to move the guest RIP in to
* @param Constant - The guest RIP that will be relocated
*/
void InsertGuestRIPMove(vixl::aarch64::Register Reg, uint64_t Constant);
void InsertGuestRIPMove(ARMEmitter::Register Reg, uint64_t Constant);
/**
* @brief Inserts a named symbol as a literal in memory
@@ -208,23 +231,6 @@ private:
/** @} */
uint32_t SpillSlots{};
/**
* @brief Current guest RIP entrypoint
*/
uint8_t *GuestEntry{};
using OpHandler = void (Arm64JITCore::*)(IR::IROp_Header const *IROp, IR::NodeID Node);
std::array<OpHandler, IR::IROps::OP_LAST + 1> OpHandlers {};
void RegisterALUHandlers();
void RegisterAtomicHandlers();
void RegisterBranchHandlers();
void RegisterConversionHandlers();
void RegisterFlagHandlers();
void RegisterMemoryHandlers();
void RegisterMiscHandlers();
void RegisterMoveHandlers();
void RegisterVectorHandlers();
void RegisterEncryptionHandlers();
#define DEF_OP(x) void Op_##x(IR::IROp_Header const *IROp, IR::NodeID Node)
///< Unhandled handler
@@ -302,7 +308,6 @@ private:
DEF_OP(AtomicFetchNeg);
///< Branch ops
DEF_OP(SignalReturn);
DEF_OP(CallbackReturn);
DEF_OP(ExitFunction);
DEF_OP(Jump);
@@ -317,6 +322,7 @@ private:
///< Conversion ops
DEF_OP(VInsGPR);
DEF_OP(VCastFromGPR);
DEF_OP(VDupFromGPR);
DEF_OP(Float_FromGPR_S);
DEF_OP(Float_FToF);
DEF_OP(Vector_SToF);
@@ -343,9 +349,11 @@ private:
DEF_OP(StoreMem);
DEF_OP(LoadMemTSO);
DEF_OP(StoreMemTSO);
DEF_OP(MemSet);
DEF_OP(ParanoidLoadMemTSO);
DEF_OP(ParanoidStoreMemTSO);
DEF_OP(CacheLineClear);
DEF_OP(CacheLineClean);
DEF_OP(CacheLineZero);
///< Misc ops
@@ -406,6 +414,8 @@ private:
DEF_OP(VZip2);
DEF_OP(VUnZip);
DEF_OP(VUnZip2);
DEF_OP(VTrn);
DEF_OP(VTrn2);
DEF_OP(VBSL);
DEF_OP(VCMPEQ);
DEF_OP(VCMPEQZ);
File diff suppressed because it is too large. Load diff
+64 -89
View File
@@ -5,31 +5,30 @@ $end_info$
*/
#include <syscall.h>
#include "Interface/Core/ArchHelpers/CodeEmitter/Emitter.h"
#include "Interface/Core/JIT/Arm64/JITClass.h"
#include "FEXCore/Debug/InternalThreadState.h"
namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const *IROp, IR::NodeID Node)
DEF_OP(GuestOpcode) {
auto Op = IROp->C<IR::IROp_GuestOpcode>();
// metadata
DebugData->GuestOpcodes.push_back({Op->GuestEntryOffset, GetCursorAddress<uint8_t*>() - GuestEntry});
DebugData->GuestOpcodes.push_back({Op->GuestEntryOffset, GetCursorAddress<uint8_t*>() - CodeData.BlockBegin});
}
DEF_OP(Fence) {
auto Op = IROp->C<IR::IROp_Fence>();
switch (Op->Fence) {
case IR::Fence_Load.Val:
dmb(FullSystem, BarrierReads);
dmb(FEXCore::ARMEmitter::BarrierScope::LD);
break;
case IR::Fence_LoadStore.Val:
dmb(FullSystem, BarrierAll);
dmb(FEXCore::ARMEmitter::BarrierScope::SY);
break;
case IR::Fence_Store.Val:
dmb(FullSystem, BarrierWrites);
dmb(FEXCore::ARMEmitter::BarrierScope::ST);
break;
default: LOGMAN_MSG_A_FMT("Unknown Fence: {}", Op->Fence); break;
}
@@ -52,108 +51,107 @@ DEF_OP(Break) {
uint64_t Constant{};
memcpy(&Constant, &State, sizeof(State));
LoadConstant(x1, Constant);
str(x1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, SynchronousFaultData)));
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r1, Constant);
str(ARMEmitter::XReg::x1, STATE, offsetof(FEXCore::Core::CpuStateFrame, SynchronousFaultData));
switch (Op->Reason.Signal) {
case SIGILL:
ldr(TMP1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.GuestSignal_SIGILL)));
ldr(TMP1, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.GuestSignal_SIGILL));
br(TMP1);
break;
case SIGTRAP:
ldr(TMP1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.GuestSignal_SIGTRAP)));
ldr(TMP1, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.GuestSignal_SIGTRAP));
br(TMP1);
break;
case SIGSEGV:
ldr(TMP1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.GuestSignal_SIGSEGV)));
ldr(TMP1, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.GuestSignal_SIGSEGV));
br(TMP1);
break;
default:
ldr(TMP1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.GuestSignal_SIGTRAP)));
ldr(TMP1, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.GuestSignal_SIGTRAP));
br(TMP1);
break;
}
}
DEF_OP(GetRoundingMode) {
auto Dst = GetReg<RA_64>(Node);
mrs(Dst, FPCR);
lsr(Dst, Dst, 22);
auto Dst = GetReg(Node);
mrs(Dst, ARMEmitter::SystemRegister::FPCR);
lsr(ARMEmitter::Size::i64Bit, Dst, Dst, 22);
// FTZ is already in the correct location
// Rounding mode is different
and_(TMP1, Dst, 0b11);
and_(ARMEmitter::Size::i64Bit, TMP1, Dst, 0b11);
cmp(TMP1, 1);
LoadConstant(TMP3, IR::ROUND_MODE_POSITIVE_INFINITY);
csel(TMP2, TMP3, xzr, vixl::aarch64::Condition::eq);
cmp(ARMEmitter::Size::i64Bit, TMP1, 1);
LoadConstant(ARMEmitter::Size::i64Bit, TMP3, IR::ROUND_MODE_POSITIVE_INFINITY);
csel(ARMEmitter::Size::i64Bit, TMP2, TMP3, ARMEmitter::Reg::zr, ARMEmitter::Condition::CC_EQ);
cmp(TMP1, 2);
LoadConstant(TMP3, IR::ROUND_MODE_NEGATIVE_INFINITY);
csel(TMP2, TMP3, TMP2, vixl::aarch64::Condition::eq);
cmp(ARMEmitter::Size::i64Bit, TMP1, 2);
LoadConstant(ARMEmitter::Size::i64Bit, TMP3, IR::ROUND_MODE_NEGATIVE_INFINITY);
csel(ARMEmitter::Size::i64Bit, TMP2, TMP3, TMP2, ARMEmitter::Condition::CC_EQ);
cmp(TMP1, 3);
LoadConstant(TMP3, IR::ROUND_MODE_TOWARDS_ZERO);
csel(TMP2, TMP3, TMP2, vixl::aarch64::Condition::eq);
cmp(ARMEmitter::Size::i64Bit, TMP1, 3);
LoadConstant(ARMEmitter::Size::i64Bit, TMP3, IR::ROUND_MODE_TOWARDS_ZERO);
csel(ARMEmitter::Size::i64Bit, TMP2, TMP3, TMP2, ARMEmitter::Condition::CC_EQ);
orr(Dst, Dst, TMP2);
orr(ARMEmitter::Size::i64Bit, Dst, Dst, TMP2.R());
bfi(Dst, TMP2, 0, 2);
bfi(ARMEmitter::Size::i64Bit, Dst, TMP2, 0, 2);
}
DEF_OP(SetRoundingMode) {
auto Op = IROp->C<IR::IROp_SetRoundingMode>();
auto Src = GetReg<RA_64>(Op->RoundMode.ID());
auto Src = GetReg(Op->RoundMode.ID());
// Setup the rounding flags correctly
and_(TMP1, Src, 0b11);
and_(ARMEmitter::Size::i64Bit, TMP1, Src, 0b11);
cmp(TMP1, IR::ROUND_MODE_POSITIVE_INFINITY);
LoadConstant(TMP3, 1);
csel(TMP2, TMP3, xzr, vixl::aarch64::Condition::eq);
cmp(ARMEmitter::Size::i64Bit, TMP1, IR::ROUND_MODE_POSITIVE_INFINITY);
LoadConstant(ARMEmitter::Size::i64Bit, TMP3, 1);
csel(ARMEmitter::Size::i64Bit, TMP2, TMP3, ARMEmitter::Reg::zr, ARMEmitter::Condition::CC_EQ);
cmp(TMP1, IR::ROUND_MODE_NEGATIVE_INFINITY);
LoadConstant(TMP3, 2);
csel(TMP2, TMP3, TMP2, vixl::aarch64::Condition::eq);
cmp(ARMEmitter::Size::i64Bit, TMP1, IR::ROUND_MODE_NEGATIVE_INFINITY);
LoadConstant(ARMEmitter::Size::i64Bit, TMP3, 2);
csel(ARMEmitter::Size::i64Bit, TMP2, TMP3, TMP2, ARMEmitter::Condition::CC_EQ);
cmp(TMP1, IR::ROUND_MODE_TOWARDS_ZERO);
LoadConstant(TMP3, 3);
csel(TMP2, TMP3, TMP2, vixl::aarch64::Condition::eq);
cmp(ARMEmitter::Size::i64Bit, TMP1, IR::ROUND_MODE_TOWARDS_ZERO);
LoadConstant(ARMEmitter::Size::i64Bit, TMP3, 3);
csel(ARMEmitter::Size::i64Bit, TMP2, TMP3, TMP2, ARMEmitter::Condition::CC_EQ);
mrs(TMP1, FPCR);
mrs(TMP1, ARMEmitter::SystemRegister::FPCR);
// vixl simulator doesn't support anything beyond ties-to-even rounding
#ifndef VIXL_SIMULATOR
// Insert the rounding flags
bfi(TMP1, TMP2, 22, 2);
bfi(ARMEmitter::Size::i64Bit, TMP1, TMP2, 22, 2);
#endif
// Insert the FTZ flag
lsr(TMP2, Src, 2);
bfi(TMP1, TMP2, 24, 1);
lsr(ARMEmitter::Size::i64Bit, TMP2, Src, 2);
bfi(ARMEmitter::Size::i64Bit, TMP1, TMP2, 24, 1);
// Now save the new FPCR
msr(FPCR, TMP1);
msr(ARMEmitter::SystemRegister::FPCR, TMP1);
}
DEF_OP(Print) {
auto Op = IROp->C<IR::IROp_Print>();
PushDynamicRegsAndLR();
PushDynamicRegsAndLR(TMP1);
SpillStaticRegs();
if (IsGPR(Op->Value.ID())) {
mov(x0, GetReg<RA_64>(Op->Value.ID()));
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.PrintValue)));
mov(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, GetReg(Op->Value.ID()));
ldr(ARMEmitter::XReg::x3, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.PrintValue));
}
else {
fmov(x0, GetSrc(Op->Value.ID()).V1D());
// Bug in vixl that source vector needs to b V1D rather than V2D?
fmov(x1, GetSrc(Op->Value.ID()).V1D(), 1);
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.PrintVectorValue)));
fmov(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, GetVReg(Op->Value.ID()), false);
fmov(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r1, GetVReg(Op->Value.ID()), true);
ldr(ARMEmitter::XReg::x3, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.PrintVectorValue));
}
blr(x3);
blr(ARMEmitter::Reg::r3);
FillStaticRegs();
PopDynamicRegsAndLR();
@@ -173,27 +171,27 @@ DEF_OP(ProcessorID) {
// 16bit LoadConstant to be a single instruction
// We must always spill at least one register (x8) so this value always has a bit set
// This gives the signal handler a value to check to see if we are in a syscall at all
LoadConstant(x0, SpillMask & 0xFFFF);
str(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, InSyscallInfo)));
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, SpillMask & 0xFFFF);
str(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CpuStateFrame, InSyscallInfo));
// Allocate some temporary space for storing the uint32_t CPU and Node IDs
sub(sp, sp, 16);
sub(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, ARMEmitter::Reg::rsp, 16);
// Load the getcpu syscall number
LoadConstant(x8, SYS_getcpu);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r8, SYS_getcpu);
// CPU pointer in x0
add(x0, sp, 0);
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, ARMEmitter::Reg::rsp, 0);
// Node in x1
add(x1, sp, 4);
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r1, ARMEmitter::Reg::rsp, 4);
svc(0);
// On updated signal mask we can receive a signal RIGHT HERE
// Load the values returned by the kernel
ldp(w0, w1, MemOperand(sp));
ldp<ARMEmitter::IndexType::OFFSET>(ARMEmitter::WReg::w0, ARMEmitter::WReg::w1, ARMEmitter::Reg::rsp);
// Deallocate stack space
sub(sp, sp, 16);
sub(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, ARMEmitter::Reg::rsp, 16);
// Now that we are done in the syscall we need to carefully peel back the state
// First unspill the registers from before
@@ -201,14 +199,13 @@ DEF_OP(ProcessorID) {
// Now the registers we've spilled are back in their original host registers
// We can safely claim we are no longer in a syscall
str(xzr, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, InSyscallInfo)));
str(ARMEmitter::XReg::zr, STATE, offsetof(FEXCore::Core::CpuStateFrame, InSyscallInfo));
// Now store the result in the destination in the expected format
// uint32_t Res = (node << 12) | cpu;
// CPU is in w0
// Node is in w1
orr(GetReg<RA_64>(Node), x0, Operand(x1, LSL, 12));
orr(ARMEmitter::Size::i64Bit, GetReg(Node), ARMEmitter::Reg::r0, ARMEmitter::Reg::r1, ARMEmitter::ShiftType::LSL, 12);
}
DEF_OP(RDRAND) {
@@ -216,45 +213,23 @@ DEF_OP(RDRAND) {
// Results are in x0, x1
// Results want to be in a i64v2 vector
auto Dst = GetSrcPair<RA_64>(Node);
auto Dst = GetRegPair(Node);
if (Op->GetReseeded) {
mrs(Dst.first, RNDRRS);
mrs(Dst.first, ARMEmitter::SystemRegister::RNDRRS);
}
else {
mrs(Dst.first, RNDR);
mrs(Dst.first, ARMEmitter::SystemRegister::RNDR);
}
// If the rng number is valid then NZCV is 0b0000, otherwise NZCV is 0b0100
cset(Dst.second, Condition::ne);
cset(ARMEmitter::Size::i64Bit, Dst.second, ARMEmitter::Condition::CC_NE);
}
DEF_OP(Yield) {
hint(SystemHint::YIELD);
yield();
}
#undef DEF_OP
void Arm64JITCore::RegisterMiscHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &Arm64JITCore::Op_##x
REGISTER_OP(DUMMY, NoOp);
REGISTER_OP(IRHEADER, NoOp);
REGISTER_OP(CODEBLOCK, NoOp);
REGISTER_OP(BEGINBLOCK, NoOp);
REGISTER_OP(ENDBLOCK, NoOp);
REGISTER_OP(GUESTOPCODE, GuestOpcode);
REGISTER_OP(FENCE, Fence);
REGISTER_OP(BREAK, Break);
REGISTER_OP(PHI, NoOp);
REGISTER_OP(PHIVALUE, NoOp);
REGISTER_OP(PRINT, Print);
REGISTER_OP(GETROUNDINGMODE, GetRoundingMode);
REGISTER_OP(SETROUNDINGMODE, SetRoundingMode);
REGISTER_OP(INVALIDATEFLAGS, NoOp);
REGISTER_OP(PROCESSORID, ProcessorID);
REGISTER_OP(RDRAND, RDRAND);
REGISTER_OP(YIELD, Yield);
#undef REGISTER_OP
}
}
+21 -54
View File
@@ -7,73 +7,40 @@ $end_info$
#include "Interface/Core/JIT/Arm64/JITClass.h"
namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const *IROp, IR::NodeID Node)
DEF_OP(ExtractElementPair) {
auto Op = IROp->C<IR::IROp_ExtractElementPair>();
switch (Op->Header.Size) {
case 4: {
auto Src = GetSrcPair<RA_32>(Op->Pair.ID());
std::array<aarch64::Register, 2> Regs = {Src.first, Src.second};
mov (GetReg<RA_32>(Node), Regs[Op->Element]);
break;
}
case 8: {
auto Src = GetSrcPair<RA_64>(Op->Pair.ID());
std::array<aarch64::Register, 2> Regs = {Src.first, Src.second};
mov (GetReg<RA_64>(Node), Regs[Op->Element]);
break;
}
default: LOGMAN_MSG_A_FMT("Unknown Size"); break;
}
LOGMAN_THROW_AA_FMT(Op->Header.Size == 4 || Op->Header.Size == 8, "Invalid size");
const auto EmitSize = Op->Header.Size == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
const auto Src = GetRegPair(Op->Pair.ID());
const std::array<ARMEmitter::Register, 2> Regs = {Src.first, Src.second};
mov(EmitSize, GetReg(Node), Regs[Op->Element]);
}
DEF_OP(CreateElementPair) {
auto Op = IROp->C<IR::IROp_CreateElementPair>();
std::pair<aarch64::Register, aarch64::Register> Dst;
aarch64::Register RegFirst;
aarch64::Register RegSecond;
aarch64::Register RegTmp;
LOGMAN_THROW_AA_FMT(IROp->ElementSize == 4 || IROp->ElementSize == 8, "Invalid size");
std::pair<ARMEmitter::Register, ARMEmitter::Register> Dst = GetRegPair(Node);
ARMEmitter::Register RegFirst = GetReg(Op->Lower.ID());
ARMEmitter::Register RegSecond = GetReg(Op->Upper.ID());
ARMEmitter::Register RegTmp = TMP1.R();
switch (IROp->ElementSize) {
case 4: {
Dst = GetSrcPair<RA_32>(Node);
RegFirst = GetReg<RA_32>(Op->Lower.ID());
RegSecond = GetReg<RA_32>(Op->Upper.ID());
RegTmp = w0;
break;
}
case 8: {
Dst = GetSrcPair<RA_64>(Node);
RegFirst = GetReg<RA_64>(Op->Lower.ID());
RegSecond = GetReg<RA_64>(Op->Upper.ID());
RegTmp = x0;
break;
}
default: LOGMAN_MSG_A_FMT("Unknown Size"); break;
}
const auto EmitSize = IROp->ElementSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
if (Dst.first.GetCode() != RegSecond.GetCode()) {
mov(Dst.first, RegFirst);
mov(Dst.second, RegSecond);
} else if (Dst.second.GetCode() != RegFirst.GetCode()) {
mov(Dst.second, RegSecond);
mov(Dst.first, RegFirst);
if (Dst.first.Idx() != RegSecond.Idx()) {
mov(EmitSize, Dst.first, RegFirst);
mov(EmitSize, Dst.second, RegSecond);
} else if (Dst.second.Idx() != RegFirst.Idx()) {
mov(EmitSize, Dst.second, RegSecond);
mov(EmitSize, Dst.first, RegFirst);
} else {
mov(RegTmp, RegFirst);
mov(Dst.second, RegSecond);
mov(Dst.first, RegTmp);
mov(EmitSize, RegTmp, RegFirst);
mov(EmitSize, Dst.second, RegSecond);
mov(EmitSize, Dst.first, RegTmp);
}
}
#undef DEF_OP
void Arm64JITCore::RegisterMoveHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &Arm64JITCore::Op_##x
REGISTER_OP(EXTRACTELEMENTPAIR, ExtractElementPair);
REGISTER_OP(CREATEELEMENTPAIR, CreateElementPair);
#undef REGISTER_OP
}
}
File diff suppressed because it is too large. Load diff
+5 -5
View File
@@ -3,7 +3,7 @@
#include <memory>
namespace FEXCore::Context {
struct Context;
class ContextImpl;
}
namespace FEXCore::Core {
@@ -13,14 +13,14 @@ struct InternalThreadState;
namespace FEXCore::CPU {
class CPUBackend;
[[nodiscard]] std::unique_ptr<CPUBackend> CreateX86JITCore(FEXCore::Context::Context *ctx,
[[nodiscard]] std::unique_ptr<CPUBackend> CreateX86JITCore(FEXCore::Context::ContextImpl *ctx,
FEXCore::Core::InternalThreadState *Thread);
void InitializeX86JITSignalHandlers(FEXCore::Context::Context *CTX);
void InitializeX86JITSignalHandlers(FEXCore::Context::ContextImpl *CTX);
CPUBackendFeatures GetX86JITBackendFeatures();
[[nodiscard]] std::unique_ptr<CPUBackend> CreateArm64JITCore(FEXCore::Context::Context *ctx,
[[nodiscard]] std::unique_ptr<CPUBackend> CreateArm64JITCore(FEXCore::Context::ContextImpl *ctx,
FEXCore::Core::InternalThreadState *Thread);
void InitializeArm64JITSignalHandlers(FEXCore::Context::Context *CTX);
void InitializeArm64JITSignalHandlers(FEXCore::Context::ContextImpl *CTX);
CPUBackendFeatures GetArm64JITBackendFeatures();
} // namespace FEXCore::CPU
@@ -30,15 +30,6 @@ $end_info$
namespace FEXCore::CPU {
#define DEF_OP(x) void X86JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(SignalReturn) {
// Adjust the stack first for a regular return
if (SpillSlots) {
add(rsp, SpillSlots * MaxSpillSlotSize); // + 8 to consume return address
}
jmp(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.SignalReturnHandler)]);
}
DEF_OP(CallbackReturn) {
// Adjust the stack first for a regular return
if (SpillSlots) {
@@ -211,7 +202,7 @@ DEF_OP(Thunk) {
mov(rdi, GetSrc<RA_64>(Op->ArgPtr.ID()));
auto thunkFn = ThreadState->CTX->ThunkHandler->LookupThunk(Op->ThunkNameHash);
auto thunkFn = static_cast<Context::ContextImpl*>(ThreadState->CTX)->ThunkHandler->LookupThunk(Op->ThunkNameHash);
mov(rax, reinterpret_cast<uintptr_t>(thunkFn));
call(rax);
@@ -314,7 +305,6 @@ DEF_OP(CPUID) {
#undef DEF_OP
void X86JITCore::RegisterBranchHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &X86JITCore::Op_##x
REGISTER_OP(SIGNALRETURN, SignalReturn);
REGISTER_OP(CALLBACKRETURN, CallbackReturn);
REGISTER_OP(EXITFUNCTION, ExitFunction);
REGISTER_OP(JUMP, Jump);
@@ -111,6 +111,53 @@ DEF_OP(VCastFromGPR) {
}
}
DEF_OP(VDupFromGPR) {
const auto Op = IROp->C<IR::IROp_VDupFromGPR>();
const auto OpSize = IROp->Size;
const auto ElementSize = IROp->ElementSize;
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto Dst = GetDst(Node);
const auto Src = GetSrc<RA_64>(Op->Src.ID()).cvt64();
vmovq(Dst, Src);
switch (ElementSize) {
case 1:
if (Is256Bit) {
vpbroadcastb(ToYMM(Dst), Dst);
} else {
vpbroadcastb(Dst, Dst);
}
break;
case 2:
if (Is256Bit) {
vpbroadcastw(ToYMM(Dst), Dst);
} else {
vpbroadcastw(Dst, Dst);
}
break;
case 4:
if (Is256Bit) {
vpbroadcastd(ToYMM(Dst), Dst);
} else {
vpbroadcastd(Dst, Dst);
}
break;
case 8:
if (Is256Bit) {
vpbroadcastq(ToYMM(Dst), Dst);
} else {
vpbroadcastq(Dst, Dst);
}
break;
default:
LOGMAN_MSG_A_FMT("Unhandled element size: {}", ElementSize);
return;
}
}
DEF_OP(Float_FromGPR_S) {
const auto Op = IROp->C<IR::IROp_Float_FromGPR_S>();
@@ -186,8 +233,8 @@ DEF_OP(Vector_SToF) {
pextrq(rcx, Vector, 0);
cvtsi2sd(Dst, rcx);
cvtsi2sd(xmm15, rax);
vmovlhps(Dst, Dst, xmm15);
if (Is256Bit) {
movlhps(Dst, xmm15);
vextracti128(xmm15, ToYMM(Vector), 1);
pextrq(rax, xmm15, 1);
@@ -197,6 +244,8 @@ DEF_OP(Vector_SToF) {
movlhps(xmm15, xmm14);
vinserti128(ToYMM(Dst), ToYMM(Dst), xmm15, 1);
} else {
vmovlhps(Dst, Dst, xmm15);
}
break;
default:
@@ -355,6 +404,7 @@ void X86JITCore::RegisterConversionHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &X86JITCore::Op_##x
REGISTER_OP(VINSGPR, VInsGPR);
REGISTER_OP(VCASTFROMGPR, VCastFromGPR);
REGISTER_OP(VDUPFROMGPR, VDupFromGPR);
REGISTER_OP(FLOAT_FROMGPR_S, Float_FromGPR_S);
REGISTER_OP(FLOAT_FTOF, Float_FToF);
REGISTER_OP(VECTOR_STOF, Vector_SToF);
@@ -21,23 +21,67 @@ DEF_OP(AESImc) {
}
DEF_OP(AESEnc) {
auto Op = IROp->C<IR::IROp_VAESEnc>();
vaesenc(GetDst(Node), GetSrc(Op->State.ID()), GetSrc(Op->Key.ID()));
const auto Op = IROp->C<IR::IROp_VAESEnc>();
const auto OpSize = IROp->Size;
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto Dst = GetDst(Node);
const auto Key = GetSrc(Op->Key.ID());
const auto State = GetSrc(Op->State.ID());
if (Is256Bit) {
vaesenc(ToYMM(Dst), ToYMM(State), ToYMM(Key));
} else {
vaesenc(Dst, State, Key);
}
}
DEF_OP(AESEncLast) {
auto Op = IROp->C<IR::IROp_VAESEncLast>();
vaesenclast(GetDst(Node), GetSrc(Op->State.ID()), GetSrc(Op->Key.ID()));
const auto Op = IROp->C<IR::IROp_VAESEncLast>();
const auto OpSize = IROp->Size;
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto Dst = GetDst(Node);
const auto Key = GetSrc(Op->Key.ID());
const auto State = GetSrc(Op->State.ID());
if (Is256Bit) {
vaesenclast(ToYMM(Dst), ToYMM(State), ToYMM(Key));
} else {
vaesenclast(Dst, State, Key);
}
}
DEF_OP(AESDec) {
auto Op = IROp->C<IR::IROp_VAESDec>();
vaesdec(GetDst(Node), GetSrc(Op->State.ID()), GetSrc(Op->Key.ID()));
const auto Op = IROp->C<IR::IROp_VAESDec>();
const auto OpSize = IROp->Size;
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto Dst = GetDst(Node);
const auto Key = GetSrc(Op->Key.ID());
const auto State = GetSrc(Op->State.ID());
if (Is256Bit) {
vaesdec(ToYMM(Dst), ToYMM(State), ToYMM(Key));
} else {
vaesdec(Dst, State, Key);
}
}
DEF_OP(AESDecLast) {
auto Op = IROp->C<IR::IROp_VAESDecLast>();
vaesdeclast(GetDst(Node), GetSrc(Op->State.ID()), GetSrc(Op->Key.ID()));
const auto Op = IROp->C<IR::IROp_VAESDecLast>();
const auto OpSize = IROp->Size;
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto Dst = GetDst(Node);
const auto Key = GetSrc(Op->Key.ID());
const auto State = GetSrc(Op->State.ID());
if (Is256Bit) {
vaesdeclast(ToYMM(Dst), ToYMM(State), ToYMM(Key));
} else {
vaesdeclast(Dst, State, Key);
}
}
DEF_OP(AESKeyGenAssist) {
@@ -76,18 +120,24 @@ DEF_OP(CRC32) {
}
DEF_OP(PCLMUL) {
auto Op = IROp->C<IR::IROp_PCLMUL>();
const auto Op = IROp->C<IR::IROp_PCLMUL>();
const auto OpSize = IROp->Size;
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
auto Dst = GetDst(Node);
auto Src1 = GetSrc(Op->Src1.ID());
auto Src2 = GetSrc(Op->Src2.ID());
const auto Dst = GetDst(Node);
const auto Src1 = GetSrc(Op->Src1.ID());
const auto Src2 = GetSrc(Op->Src2.ID());
switch (Op->Selector) {
case 0b00000000:
case 0b00000001:
case 0b00010000:
case 0b00010001:
vpclmulqdq(Dst, Src1, Src2, Op->Selector);
if (Is256Bit) {
vpclmulqdq(ToYMM(Dst), ToYMM(Src1), ToYMM(Src2), Op->Selector);
} else {
vpclmulqdq(Dst, Src1, Src2, Op->Selector);
}
break;
default:
LOGMAN_MSG_A_FMT("Unknown PCLMUL selector: {}", Op->Selector);
+53 -17
View File
@@ -147,7 +147,12 @@ void X86JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
case FABI_F80_I32: {
PushRegs();
mov(edi, GetSrc<RA_32>(IROp->Args[0].ID()));
if (Info.ABI == FABI_F80_I16) {
movsx(rdi, GetSrc<RA_32>(IROp->Args[0].ID()).cvt16());
}
else {
mov(edi, GetSrc<RA_32>(IROp->Args[0].ID()));
}
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])]);
PopRegs();
@@ -223,7 +228,7 @@ void X86JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
PopRegs();
movzx(GetDst<RA_64>(Node), ax);
movsx(GetDst<RA_64>(Node), ax);
}
break;
case FABI_I32_F80:{
@@ -325,7 +330,7 @@ static uint64_t X86JITCore_ExitFunctionLink(FEXCore::Core::CpuStateFrame *Frame,
}
auto LinkerAddress = Frame->Pointers.Common.ExitFunctionLinker;
Context::Context::ThreadAddBlockLink(Thread, GuestRip, (uintptr_t)record, [record, LinkerAddress]{
Context::ContextImpl::ThreadAddBlockLink(Thread, GuestRip, (uintptr_t)record, [record, LinkerAddress]{
// undo the link
record[0] = LinkerAddress;
});
@@ -337,7 +342,7 @@ static uint64_t X86JITCore_ExitFunctionLink(FEXCore::Core::CpuStateFrame *Frame,
void X86JITCore::Op_NoOp(IR::IROp_Header *IROp, IR::NodeID Node) {
}
X86JITCore::X86JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread)
X86JITCore::X86JITCore(FEXCore::Context::ContextImpl *ctx, FEXCore::Core::InternalThreadState *Thread)
: CPUBackend(Thread, INITIAL_CODE_SIZE, MAX_CODE_SIZE)
, CodeGenerator(0, this, nullptr) // this is not used here
, CTX {ctx} {
@@ -374,7 +379,7 @@ X86JITCore::X86JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalTh
Common.PrintValue = reinterpret_cast<uint64_t>(PrintValue);
Common.PrintVectorValue = reinterpret_cast<uint64_t>(PrintVectorValue);
Common.ThreadRemoveCodeEntryFromJIT = reinterpret_cast<uintptr_t>(&Context::Context::ThreadRemoveCodeEntryFromJit);
Common.ThreadRemoveCodeEntryFromJIT = reinterpret_cast<uintptr_t>(&Context::ContextImpl::ThreadRemoveCodeEntryFromJit);
Common.CPUIDObj = reinterpret_cast<uint64_t>(&CTX->CPUID);
{
@@ -384,7 +389,7 @@ X86JITCore::X86JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalTh
Common.SyscallHandlerObj = reinterpret_cast<uint64_t>(CTX->SyscallHandler);
Common.SyscallHandlerFunc = reinterpret_cast<uint64_t>(FEXCore::Context::HandleSyscall);
Common.ExitFunctionLink = reinterpret_cast<uintptr_t>(&Context::Context::ThreadExitFunctionLink<X86JITCore_ExitFunctionLink>);
Common.ExitFunctionLink = reinterpret_cast<uintptr_t>(&Context::ContextImpl::ThreadExitFunctionLink<X86JITCore_ExitFunctionLink>);
// Fill in the fallback handlers
InterpreterOps::FillFallbackIndexPointers(Common.FallbackHandlerPointers);
@@ -394,9 +399,9 @@ X86JITCore::X86JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalTh
ClearCache();
}
void X86JITCore::InitializeSignalHandlers(FEXCore::Context::Context *CTX) {
void X86JITCore::InitializeSignalHandlers(FEXCore::Context::ContextImpl *CTX) {
CTX->SignalDelegation->RegisterHostSignalHandler(SIGILL, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
return Thread->CTX->Dispatcher->HandleSIGILL(Thread, Signal, info, ucontext);
return static_cast<Context::ContextImpl*>(Thread->CTX)->Dispatcher->HandleSIGILL(Thread, Signal, info, ucontext);
}, true);
}
@@ -582,7 +587,7 @@ std::tuple<X86JITCore::SetCC, X86JITCore::CMovCC, X86JITCore::JCC> X86JITCore::G
return { &CodeGenerator::sete , &CodeGenerator::cmove , &CodeGenerator::je };
}
void *X86JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IRListView const *IR, [[maybe_unused]] FEXCore::Core::DebugData *DebugData, FEXCore::IR::RegisterAllocationData *RAData, bool GDBEnabled) {
CPUBackend::CompiledCode X86JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IRListView const *IR, [[maybe_unused]] FEXCore::Core::DebugData *DebugData, FEXCore::IR::RegisterAllocationData *RAData, bool GDBEnabled) {
FEXCORE_PROFILE_SCOPED("x86::CompileCode");
JumpTargets.clear();
@@ -598,12 +603,27 @@ void *X86JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IRLi
CTX->ClearCodeCache(ThreadState);
}
GuestEntry = getCurr<uint8_t*>();
CodeData.BlockBegin = getCurr<uint8_t*>();
// Put the code header at the start of the data block.
Label JITCodeHeaderLabel{};
L(JITCodeHeaderLabel);
JITCodeHeader *CodeHeader = getCurr<JITCodeHeader *>();
setSize(getSize() + sizeof(JITCodeHeader));
CodeData.BlockEntry = getCurr<uint8_t*>();
// Get the address of the JITCodeHeader and store in to the core state.
// Only two instructions, so very low overhead.
lea(TMP1, ptr [rip + JITCodeHeaderLabel]);
mov(qword [STATE + offsetof(FEXCore::Core::CPUState, InlineJITBlockHeader)], TMP1);
CursorEntry = getSize();
this->IR = IR;
if (GDBEnabled) {
auto GDBSize = CTX->Dispatcher->GenerateGDBPauseCheck(GuestEntry, Entry);
auto GDBSize = CTX->Dispatcher->GenerateGDBPauseCheck(CodeData.BlockBegin, Entry);
setSize(getSize() + GDBSize);
}
@@ -726,7 +746,7 @@ void *X86JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IRLi
if (DebugData) {
DebugData->Subblocks.push_back({
static_cast<uint32_t>(BlockStartHostCode - GuestEntry),
static_cast<uint32_t>(BlockStartHostCode - CodeData.BlockBegin),
static_cast<uint32_t>(getCurr<uint8_t *>() - BlockStartHostCode)
});
}
@@ -739,20 +759,36 @@ void *X86JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IRLi
}
PendingTargetLabel = nullptr;
void *GuestExit = getCurr<void*>();
// Add the JitCodeTail
auto JITBlockTailLocation = getCurr<uint8_t *>();
auto JITBlockTail = getCurr<JITCodeTail*>();
setSize(getSize() + sizeof(JITCodeTail));
// Put the block's RIP entry in the tail.
// This will be used for RIP reconstruction in the future.
// TODO: This needs to be a data RIP relocation once code caching works.
// Current relocation code doesn't support this feature yet.
JITBlockTail->RIP = Entry;
CodeHeader->OffsetToBlockTail = JITBlockTailLocation - CodeData.BlockBegin;
CodeData.Size = getCurr<uint8_t*>() - CodeData.BlockBegin;
JITBlockTail->Size = CodeData.Size;
this->IR = nullptr;
ready();
if (DebugData) {
DebugData->HostCodeSize = reinterpret_cast<uintptr_t>(GuestExit) - reinterpret_cast<uintptr_t>(GuestEntry);
DebugData->HostCodeSize = CodeData.Size;
DebugData->Relocations = &Relocations;
}
return GuestEntry;
return CodeData;
}
std::unique_ptr<CPUBackend> CreateX86JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread) {
std::unique_ptr<CPUBackend> CreateX86JITCore(FEXCore::Context::ContextImpl *ctx, FEXCore::Core::InternalThreadState *Thread) {
return std::make_unique<X86JITCore>(ctx, Thread);
}
@@ -760,7 +796,7 @@ CPUBackendFeatures GetX86JITBackendFeatures() {
return CPUBackendFeatures { };
}
void InitializeX86JITSignalHandlers(FEXCore::Context::Context *CTX) {
void InitializeX86JITSignalHandlers(FEXCore::Context::ContextImpl *CTX) {
X86JITCore::InitializeSignalHandlers(CTX);
}
@@ -51,13 +51,13 @@ const std::array<Xbyak::Xmm, 11> RAXMM_x = { xmm1, xmm2, xmm3, xmm4, xmm5, xmm6
class X86JITCore final : public CPUBackend, public Xbyak::CodeGenerator {
public:
explicit X86JITCore(FEXCore::Context::Context *ctx,
explicit X86JITCore(FEXCore::Context::ContextImpl *ctx,
FEXCore::Core::InternalThreadState *Thread);
~X86JITCore() override;
[[nodiscard]] std::string GetName() override { return "JIT"; }
[[nodiscard]] void *CompileCode(uint64_t Entry,
[[nodiscard]] CPUBackend::CompiledCode CompileCode(uint64_t Entry,
FEXCore::IR::IRListView const *IR,
FEXCore::Core::DebugData *DebugData,
FEXCore::IR::RegisterAllocationData *RAData, bool GDBEnabled) override;
@@ -68,7 +68,7 @@ public:
void ClearCache() override;
static void InitializeSignalHandlers(FEXCore::Context::Context *CTX);
static void InitializeSignalHandlers(FEXCore::Context::ContextImpl *CTX);
void ClearRelocations() override { Relocations.clear(); }
@@ -135,9 +135,10 @@ private:
/** @} */
Label* PendingTargetLabel{};
FEXCore::Context::Context *CTX;
FEXCore::Context::ContextImpl *CTX;
FEXCore::IR::IRListView const *IR;
uint64_t Entry;
CPUBackend::CompiledCode CodeData{};
std::unordered_map<IR::NodeID, Label> JumpTargets;
Xbyak::util::Cpu Features{};
@@ -205,10 +206,6 @@ private:
void EmitDetectionString();
uint32_t SpillSlots{};
/**
* @brief Current guest RIP entrypoint
*/
uint8_t *GuestEntry{};
using SetCC = void (X86JITCore::*)(const Operand& op);
using CMovCC = void (X86JITCore::*)(const Reg& reg, const Operand& op);
@@ -308,7 +305,6 @@ private:
DEF_OP(AtomicFetchNeg);
///< Branch ops
DEF_OP(SignalReturn);
DEF_OP(CallbackReturn);
DEF_OP(ExitFunction);
DEF_OP(Jump);
@@ -322,6 +318,7 @@ private:
///< Conversion ops
DEF_OP(VInsGPR);
DEF_OP(VCastFromGPR);
DEF_OP(VDupFromGPR);
DEF_OP(Float_FromGPR_S);
DEF_OP(Float_FToF);
DEF_OP(Vector_UToF);
@@ -347,7 +344,9 @@ private:
DEF_OP(StoreFlag);
DEF_OP(LoadMem);
DEF_OP(StoreMem);
DEF_OP(MemSet);
DEF_OP(CacheLineClear);
DEF_OP(CacheLineClean);
DEF_OP(CacheLineZero);
///< Misc ops
@@ -408,6 +407,8 @@ private:
DEF_OP(VZip2);
DEF_OP(VUnZip);
DEF_OP(VUnZip2);
DEF_OP(VTrn);
DEF_OP(VTrn2);
DEF_OP(VBSL);
DEF_OP(VCMPEQ);
DEF_OP(VCMPEQZ);
@@ -766,12 +766,112 @@ DEF_OP(StoreMem) {
}
}
DEF_OP(MemSet) {
const auto Op = IROp->C<IR::IROp_MemSet>();
const int32_t Size = Op->Size;
const auto MemReg = GetSrc<RA_64>(Op->Addr.ID());
const auto Value = GetSrc<RA_64>(Op->Value.ID());
const auto Length = GetSrc<RA_64>(Op->Length.ID());
const auto Direction = GetSrc<RA_64>(Op->Direction.ID());
const auto Dst = GetSrc<RA_64>(Node);
// If Direction == 0 then:
// MemReg is incremented (by size)
// else:
// MemReg is decremented (by size)
//
// Counter is decremented regardless.
// TMP1 = rax
// TMP2 = rcx
// TMP4 = rdi
// That leaves us with TMP3 and TMP5
mov(rax, Value);
mov(rcx, Length);
mov(rdi, MemReg);
{
mov(TMP3, Length);
auto CalculateDest = [&]() {
mov(Dst, MemReg);
switch (Size) {
case 1:
break;
case 2:
shl(TMP3, 1);
break;
case 4:
shl(TMP3, 2);
break;
case 8:
shl(TMP3, 3);
break;
default:
LOGMAN_MSG_A_FMT("Unhandled {} size: {}", __func__, Size);
break;
}
};
Label AfterDir;
Label BackwardDir;
cmp(Direction, 0);
jne(BackwardDir);
// Incrementing DF flag.
cld();
CalculateDest();
add(Dst, TMP3);
jmp(AfterDir);
L(BackwardDir);
// Decrementing DF flag.
std();
CalculateDest();
sub(Dst, TMP3);
L(AfterDir);
}
switch (Size) {
case 1:
rep(); stosb();
break;
case 2:
rep(); stosw();
break;
case 4:
rep(); stosd();
break;
case 8:
rep(); stosq();
break;
default:
LOGMAN_MSG_A_FMT("Unhandled {} size: {}", __func__, Size);
break;
}
// Ensure we set DF back to zero. Required by the ABI.
cld();
}
DEF_OP(CacheLineClear) {
auto Op = IROp->C<IR::IROp_CacheLineClear>();
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Addr.ID());
clflush(ptr [MemReg]);
if (Op->Serialize) {
clflush(ptr [MemReg]);
}
else {
clflushopt(ptr [MemReg]);
}
}
DEF_OP(CacheLineClean) {
auto Op = IROp->C<IR::IROp_CacheLineClean>();
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Addr.ID());
clwb(ptr [MemReg]);
}
DEF_OP(CacheLineZero) {
@@ -808,7 +908,9 @@ void X86JITCore::RegisterMemoryHandlers() {
REGISTER_OP(STOREMEM, StoreMem);
REGISTER_OP(LOADMEMTSO, LoadMem);
REGISTER_OP(STOREMEMTSO, StoreMem);
REGISTER_OP(MEMSET, MemSet);
REGISTER_OP(CACHELINECLEAR, CacheLineClear);
REGISTER_OP(CACHELINECLEAN, CacheLineClean);
REGISTER_OP(CACHELINEZERO, CacheLineZero);
#undef REGISTER_OP
}
@@ -24,7 +24,7 @@ namespace FEXCore::CPU {
DEF_OP(GuestOpcode) {
auto Op = IROp->C<IR::IROp_GuestOpcode>();
// metadata
DebugData->GuestOpcodes.push_back({Op->GuestEntryOffset, getCurr<uint8_t*>() - GuestEntry});
DebugData->GuestOpcodes.push_back({Op->GuestEntryOffset, getCurr<uint8_t*>() - CodeData.BlockBegin});
}
DEF_OP(Fence) {
@@ -433,6 +433,7 @@ DEF_OP(VAddP) {
const auto Op = IROp->C<IR::IROp_VAddP>();
const auto OpSize = IROp->Size;
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto ElementSize = Op->Header.ElementSize;
const auto Dst = GetDst(Node);
@@ -478,30 +479,64 @@ DEF_OP(VAddP) {
const auto VectorLowerYMM = ToYMM(VectorLower);
const auto VectorUpperYMM = ToYMM(VectorUpper);
// To behave like ADDP, we need to swap the second and third elements around
// in the 256-bit case. ADDP operates as if both vectors are concatenated
// together and runs down the length of it adding pairs as it goes, whereas
// VPHADDW/D operates on both individual halves of the entire register.
switch (ElementSize) {
case 1:
vmovdqu(ymm15, VectorLowerYMM);
vmovdqu(ymm14, VectorUpperYMM);
if (Is256Bit) {
vmovdqu(ymm15, VectorLowerYMM);
vmovdqu(ymm14, VectorUpperYMM);
vpunpcklbw(ymm0, ymm15, ymm14);
vpunpckhbw(ymm12, ymm15, ymm14);
vpunpcklbw(ymm0, ymm15, ymm14);
vpunpckhbw(ymm12, ymm15, ymm14);
vpunpcklbw(ymm15, ymm0, ymm12);
vpunpckhbw(ymm14, ymm0, ymm12);
vpunpcklbw(ymm15, ymm0, ymm12);
vpunpckhbw(ymm14, ymm0, ymm12);
vpunpcklbw(ymm0, ymm15, ymm14);
vpunpckhbw(ymm12, ymm15, ymm14);
vpunpcklbw(ymm0, ymm15, ymm14);
vpunpckhbw(ymm12, ymm15, ymm14);
vpunpcklbw(ymm15, ymm0, ymm12);
vpunpckhbw(ymm14, ymm0, ymm12);
vpunpcklbw(ymm15, ymm0, ymm12);
vpunpckhbw(ymm14, ymm0, ymm12);
vpaddb(DstYMM, ymm15, ymm14);
vpaddb(DstYMM, ymm15, ymm14);
vpermq(DstYMM, DstYMM, 0b11'01'10'00);
} else {
vmovdqu(xmm15, VectorLower);
vmovdqu(xmm14, VectorUpper);
vpunpcklbw(xmm0, xmm15, xmm14);
vpunpckhbw(xmm12, xmm15, xmm14);
vpunpcklbw(xmm15, xmm0, xmm12);
vpunpckhbw(xmm14, xmm0, xmm12);
vpunpcklbw(xmm0, xmm15, xmm14);
vpunpckhbw(xmm12, xmm15, xmm14);
vpunpcklbw(xmm15, xmm0, xmm12);
vpunpckhbw(xmm14, xmm0, xmm12);
vpaddb(Dst, xmm15, xmm14);
}
break;
case 2:
vphaddw(DstYMM, VectorLowerYMM, VectorUpperYMM);
if (Is256Bit) {
vphaddw(DstYMM, VectorLowerYMM, VectorUpperYMM);
vpermq(DstYMM, DstYMM, 0b11'01'10'00);
} else {
vphaddw(Dst, VectorLower, VectorUpper);
}
break;
case 4:
vphaddd(DstYMM, VectorLowerYMM, VectorUpperYMM);
if (Is256Bit) {
vphaddd(DstYMM, VectorLowerYMM, VectorUpperYMM);
vpermq(DstYMM, DstYMM, 0b11'01'10'00);
} else {
vphaddd(Dst, VectorLower, VectorUpper);
}
break;
default:
LOGMAN_MSG_A_FMT("Unknown Element Size: {}", ElementSize);
@@ -811,10 +846,15 @@ DEF_OP(VFAddP) {
const auto VectorLower = GetSrc(Op->VectorLower.ID());
const auto VectorUpper = GetSrc(Op->VectorUpper.ID());
// To behave like FADDP, we need to swap the second and third elements around
// in the 256-bit case. FADDP operates as if both vectors are concatenated
// together and runs down the length of it adding pairs as it goes, whereas
// VHADDPS operates on both individual halves of the entire register.
switch (ElementSize) {
case 4:
if (Is256Bit) {
vhaddpd(ToYMM(Dst), ToYMM(VectorLower), ToYMM(VectorUpper));
vhaddps(ToYMM(Dst), ToYMM(VectorLower), ToYMM(VectorUpper));
vpermq(ToYMM(Dst), ToYMM(Dst), 0b11'01'10'00);
} else {
vhaddps(Dst, VectorLower, VectorUpper);
}
@@ -822,6 +862,7 @@ DEF_OP(VFAddP) {
case 8:
if (Is256Bit) {
vhaddpd(ToYMM(Dst), ToYMM(VectorLower), ToYMM(VectorUpper));
vpermq(ToYMM(Dst), ToYMM(Dst), 0b11'01'10'00);
} else {
vhaddpd(Dst, VectorLower, VectorUpper);
}
@@ -1709,7 +1750,36 @@ DEF_OP(VUnZip) {
const auto VectorUpper = GetSrc(Op->VectorUpper.ID());
if (OpSize == 8) {
LOGMAN_MSG_A_FMT("Unsupported register size on VUnZip");
switch (ElementSize) {
case 1: {
mov(rax, 0x80'80'80'80'06'04'02'00); // Lower
mov(rcx, 0x80'80'80'80'80'80'80'80); // Upper
vmovq(xmm15, rax);
pinsrq(xmm15, rcx, 1);
vpshufb(xmm14, VectorLower, xmm15);
vpshufb(xmm13, VectorUpper, xmm15);
vpunpckldq(Dst, xmm14, xmm13);
break;
}
case 2: {
mov(rax, 0x80'80'80'80'05'04'01'00); // Lower
mov(rcx, 0x80'80'80'80'80'80'80'80); // Upper
vmovq(xmm15, rax);
pinsrq(xmm15, rcx, 1);
vpshufb(xmm14, VectorLower, xmm15);
vpshufb(xmm13, VectorUpper, xmm15);
vpunpckldq(Dst, xmm14, xmm13);
break;
}
case 4: {
vshufps(Dst, VectorLower, VectorUpper, 0b10'00'10'00);
vmovq(Dst, Dst);
break;
}
default:
LOGMAN_MSG_A_FMT("Unknown Element Size: {}", ElementSize);
break;
}
}
else {
switch (ElementSize) {
@@ -1724,6 +1794,7 @@ DEF_OP(VUnZip) {
vpshufb(ymm14, ToYMM(VectorLower), ymm15);
vpshufb(ymm13, ToYMM(VectorUpper), ymm15);
vpunpcklqdq(ToYMM(Dst), ymm14, ymm13);
vpermq(ToYMM(Dst), ToYMM(Dst), 0b11'01'10'00);
} else {
vpshufb(xmm14, VectorLower, xmm15);
vpshufb(xmm13, VectorUpper, xmm15);
@@ -1743,6 +1814,7 @@ DEF_OP(VUnZip) {
vpshufb(ymm14, ToYMM(VectorLower), ymm15);
vpshufb(ymm13, ToYMM(VectorUpper), ymm15);
vpunpcklqdq(ToYMM(Dst), ymm14, ymm13);
vpermq(ToYMM(Dst), ToYMM(Dst), 0b11'01'10'00);
} else {
vpshufb(xmm14, VectorLower, xmm15);
vpshufb(xmm13, VectorUpper, xmm15);
@@ -1754,6 +1826,7 @@ DEF_OP(VUnZip) {
case 4: {
if (Is256Bit) {
vshufps(ToYMM(Dst), ToYMM(VectorLower), ToYMM(VectorUpper), 0b10'00'10'00);
vpermq(ToYMM(Dst), ToYMM(Dst), 0b11'01'10'00);
} else {
vshufps(Dst, VectorLower, VectorUpper, 0b10'00'10'00);
}
@@ -1762,6 +1835,7 @@ DEF_OP(VUnZip) {
case 8: {
if (Is256Bit) {
vshufpd(ToYMM(Dst), ToYMM(VectorLower), ToYMM(VectorUpper), 0b0'0);
vpermq(ToYMM(Dst), ToYMM(Dst), 0b11'01'10'00);
} else {
vshufpd(Dst, VectorLower, VectorUpper, 0b0'0);
}
@@ -1786,7 +1860,36 @@ DEF_OP(VUnZip2) {
const auto VectorUpper = GetSrc(Op->VectorUpper.ID());
if (OpSize == 8) {
LOGMAN_MSG_A_FMT("Unsupported register size on VUnZip2");
switch (ElementSize) {
case 1: {
mov(rax, 0x80'80'80'80'07'05'03'01); // Lower
mov(rcx, 0x80'80'80'80'80'80'80'80); // Upper
vmovq(xmm15, rax);
pinsrq(xmm15, rcx, 1);
vpshufb(xmm14, VectorLower, xmm15);
vpshufb(xmm13, VectorUpper, xmm15);
vpunpckldq(Dst, xmm14, xmm13);
break;
}
case 2: {
mov(rax, 0x80'80'80'80'07'06'03'02); // Lower
mov(rcx, 0x80'80'80'80'80'80'80'80); // Upper
vmovq(xmm15, rax);
pinsrq(xmm15, rcx, 1);
vpshufb(xmm14, VectorLower, xmm15);
vpshufb(xmm13, VectorUpper, xmm15);
vpunpckldq(Dst, xmm14, xmm13);
break;
}
case 4: {
vshufps(Dst, VectorLower, VectorUpper, 0b11'01'11'01);
vmovq(Dst, Dst);
break;
}
default:
LOGMAN_MSG_A_FMT("Unknown Element Size: {}", ElementSize);
break;
}
}
else {
switch (ElementSize) {
@@ -1801,6 +1904,7 @@ DEF_OP(VUnZip2) {
vpshufb(ymm14, ToYMM(VectorLower), ymm15);
vpshufb(ymm13, ToYMM(VectorUpper), ymm15);
vpunpcklqdq(ToYMM(Dst), ymm14, ymm13);
vpermq(ToYMM(Dst), ToYMM(Dst), 0b11'01'10'00);
} else {
vpshufb(xmm14, VectorLower, xmm15);
vpshufb(xmm13, VectorUpper, xmm15);
@@ -1820,6 +1924,7 @@ DEF_OP(VUnZip2) {
vpshufb(ymm14, ToYMM(VectorLower), ymm15);
vpshufb(ymm13, ToYMM(VectorUpper), ymm15);
vpunpcklqdq(ToYMM(Dst), ymm14, ymm13);
vpermq(ToYMM(Dst), ToYMM(Dst), 0b11'01'10'00);
} else {
vpshufb(xmm14, VectorLower, xmm15);
vpshufb(xmm13, VectorUpper, xmm15);
@@ -1831,6 +1936,7 @@ DEF_OP(VUnZip2) {
case 4: {
if (Is256Bit) {
vshufps(ToYMM(Dst), ToYMM(VectorLower), ToYMM(VectorUpper), 0b11'01'11'01);
vpermq(ToYMM(Dst), ToYMM(Dst), 0b11'01'10'00);
} else {
vshufps(Dst, VectorLower, VectorUpper, 0b11'01'11'01);
}
@@ -1838,7 +1944,8 @@ DEF_OP(VUnZip2) {
}
case 8: {
if (Is256Bit) {
vshufpd(ToYMM(Dst), ToYMM(VectorLower), ToYMM(VectorUpper), 0b1'1);
vshufpd(ToYMM(Dst), ToYMM(VectorLower), ToYMM(VectorUpper), 0b11'11);
vpermq(ToYMM(Dst), ToYMM(Dst), 0b11'01'10'00);
} else {
vshufpd(Dst, VectorLower, VectorUpper, 0b1'1);
}
@@ -1851,6 +1958,191 @@ DEF_OP(VUnZip2) {
}
}
DEF_OP(VTrn) {
const auto Op = IROp->C<IR::IROp_VTrn>();
const auto OpSize = IROp->Size;
const auto ElementSize = Op->Header.ElementSize;
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto Dst = GetDst(Node);
const auto VectorLower = GetSrc(Op->VectorLower.ID());
const auto VectorUpper = GetSrc(Op->VectorUpper.ID());
const auto LoadPshufbReg = [&](Xbyak::Xmm reg, uint64_t lower) {
mov(rax, lower);
mov(rcx, 0x80'80'80'80'80'80'80'80);
vmovq(reg, rax);
pinsrq(reg, rcx, 1);
};
switch (ElementSize) {
case 1: {
LoadPshufbReg(xmm15, 0x0E'0C'0A'08'06'04'02'00);
if (Is256Bit) {
vinserti128(ymm15, ymm15, xmm15, 1);
vpshufb(ymm14, ToYMM(VectorLower), ymm15);
vpshufb(ymm13, ToYMM(VectorUpper), ymm15);
vpunpcklbw(ToYMM(Dst), ymm14, ymm13);
} else {
vpshufb(xmm14, VectorLower, xmm15);
vpshufb(xmm13, VectorUpper, xmm15);
vpunpcklbw(Dst, xmm14, xmm13);
}
break;
}
case 2: {
LoadPshufbReg(xmm15, 0x0D'0C'09'08'05'04'01'00);
if (Is256Bit) {
vinserti128(ymm15, ymm15, xmm15, 1);
vpshufb(ymm14, ToYMM(VectorLower), ymm15);
vpshufb(ymm13, ToYMM(VectorUpper), ymm15);
vpunpcklwd(ToYMM(Dst), ymm14, ymm13);
} else {
vpshufb(xmm14, VectorLower, xmm15);
vpshufb(xmm13, VectorUpper, xmm15);
vpunpcklwd(Dst, xmm14, xmm13);
}
break;
}
case 4: {
LoadPshufbReg(xmm15, 0x0B'0A'09'08'03'02'01'00);
if (Is256Bit) {
vinserti128(ymm15, ymm15, xmm15, 1);
vpshufb(ymm14, ToYMM(VectorLower), ymm15);
vpshufb(ymm13, ToYMM(VectorUpper), ymm15);
vpunpckldq(ToYMM(Dst), ymm14, ymm13);
} else {
vpshufb(xmm14, VectorLower, xmm15);
vpshufb(xmm13, VectorUpper, xmm15);
vpunpckldq(Dst, xmm14, xmm13);
}
break;
}
case 8: {
LoadPshufbReg(xmm15, 0x07'06'05'04'03'02'01'00);
if (Is256Bit) {
vinserti128(ymm15, ymm15, xmm15, 1);
vpshufb(ymm14, ToYMM(VectorLower), ymm15);
vpshufb(ymm13, ToYMM(VectorUpper), ymm15);
vpunpcklqdq(ToYMM(Dst), ymm14, ymm13);
} else {
vpshufb(xmm14, VectorLower, xmm15);
vpshufb(xmm13, VectorUpper, xmm15);
vpunpcklqdq(Dst, xmm14, xmm13);
}
break;
}
default:
LOGMAN_MSG_A_FMT("Unknown Element Size: {}", ElementSize);
return;
}
}
DEF_OP(VTrn2) {
const auto Op = IROp->C<IR::IROp_VTrn2>();
const auto OpSize = IROp->Size;
const auto ElementSize = Op->Header.ElementSize;
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto Dst = GetDst(Node);
const auto VectorLower = GetSrc(Op->VectorLower.ID());
const auto VectorUpper = GetSrc(Op->VectorUpper.ID());
const auto LoadPshufbReg = [&](Xbyak::Xmm reg, uint64_t lower) {
mov(rax, lower);
mov(rcx, 0x80'80'80'80'80'80'80'80);
vmovq(reg, rax);
pinsrq(reg, rcx, 1);
};
switch (ElementSize) {
case 1: {
LoadPshufbReg(xmm15, 0x0F'0D'0B'09'07'05'03'01);
if (Is256Bit) {
vinserti128(ymm15, ymm15, xmm15, 1);
vpshufb(ymm14, ToYMM(VectorLower), ymm15);
vpshufb(ymm13, ToYMM(VectorUpper), ymm15);
vpunpcklbw(ToYMM(Dst), ymm14, ymm13);
} else {
vpshufb(xmm14, VectorLower, xmm15);
vpshufb(xmm13, VectorUpper, xmm15);
vpunpcklbw(Dst, xmm14, xmm13);
}
break;
}
case 2: {
LoadPshufbReg(xmm15, 0x0F'0E'0B'0A'07'06'03'02);
if (Is256Bit) {
vinserti128(ymm15, ymm15, xmm15, 1);
vpshufb(ymm14, ToYMM(VectorLower), ymm15);
vpshufb(ymm13, ToYMM(VectorUpper), ymm15);
vpunpcklwd(ToYMM(Dst), ymm14, ymm13);
} else {
vpshufb(xmm14, VectorLower, xmm15);
vpshufb(xmm13, VectorUpper, xmm15);
vpunpcklwd(Dst, xmm14, xmm13);
}
break;
}
case 4: {
LoadPshufbReg(xmm15, 0x0F'0E'0D'0C'07'06'05'04);
if (Is256Bit) {
vinserti128(ymm15, ymm15, xmm15, 1);
vpshufb(ymm14, ToYMM(VectorLower), ymm15);
vpshufb(ymm13, ToYMM(VectorUpper), ymm15);
vpunpckldq(ToYMM(Dst), ymm14, ymm13);
} else {
vpshufb(xmm14, VectorLower, xmm15);
vpshufb(xmm13, VectorUpper, xmm15);
vpunpckldq(Dst, xmm14, xmm13);
}
break;
}
case 8: {
LoadPshufbReg(xmm15, 0x0F'0E'0D'0C'0B'0A'09'08);
if (Is256Bit) {
vinserti128(ymm15, ymm15, xmm15, 1);
vpshufb(ymm14, ToYMM(VectorLower), ymm15);
vpshufb(ymm13, ToYMM(VectorUpper), ymm15);
vpunpcklqdq(ToYMM(Dst), ymm14, ymm13);
} else {
vpshufb(xmm14, VectorLower, xmm15);
vpshufb(xmm13, VectorUpper, xmm15);
vpunpcklqdq(Dst, xmm14, xmm13);
}
break;
}
default:
LOGMAN_MSG_A_FMT("Unknown Element Size: {}", ElementSize);
return;
}
}
DEF_OP(VBSL) {
const auto Op = IROp->C<IR::IROp_VBSL>();
@@ -2442,7 +2734,22 @@ DEF_OP(VUShr) {
}
DEF_OP(VSShr) {
LOGMAN_MSG_A_FMT("Unimplemented");
const auto Op = IROp->C<IR::IROp_VSShr>();
const auto OpSize = IROp->Size;
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto ElementSize = IROp->ElementSize;
LOGMAN_THROW_AA_FMT(ElementSize == 4, "VSShr only supports 32-bit elements");
const auto Dst = GetDst(Node);
const auto ShiftVector = GetSrc(Op->ShiftVector.ID());
const auto Vector = GetSrc(Op->Vector.ID());
if (Is256Bit) {
vpsravd(ToYMM(Dst), ToYMM(Vector), ToYMM(ShiftVector));
} else {
vpsravd(Dst, Vector, ShiftVector);
}
}
DEF_OP(VUShlS) {
@@ -2679,19 +2986,26 @@ DEF_OP(VInsElement) {
}
};
const auto SrcReg = GetSrcVector(xmm14);
const auto DstReg = GetDstVector(xmm15);
const auto SanitizedDstIdx = SanitizeIndex(DestIdx, DstIsUpper);
const auto SanitizedSrcIdx = SanitizeIndex(SrcIdx, SrcIsUpper);
PerformInsertion(SrcReg, SanitizedSrcIdx, DstReg, SanitizedDstIdx);
vmovapd(ToYMM(Dst), ToYMM(DestVector));
if (DstIsUpper) {
vinserti128(ToYMM(Dst), ToYMM(Dst), DstReg, 1);
const auto Is128BitElement = ElementSize == Core::CPUState::XMM_SSE_REG_SIZE;
if (Is128BitElement) {
vextracti128(xmm14, ToYMM(SrcVector), SrcIdx);
vmovapd(ToYMM(Dst), ToYMM(DestVector));
vinserti128(ToYMM(Dst), ToYMM(Dst), xmm14, DestIdx);
} else {
vinserti128(ToYMM(Dst), ToYMM(Dst), DstReg, 0);
const auto SrcReg = GetSrcVector(xmm14);
const auto DstReg = GetDstVector(xmm15);
const auto SanitizedDstIdx = SanitizeIndex(DestIdx, DstIsUpper);
const auto SanitizedSrcIdx = SanitizeIndex(SrcIdx, SrcIsUpper);
PerformInsertion(SrcReg, SanitizedSrcIdx, DstReg, SanitizedDstIdx);
vmovapd(ToYMM(Dst), ToYMM(DestVector));
if (DstIsUpper) {
vinserti128(ToYMM(Dst), ToYMM(Dst), DstReg, 1);
} else {
vinserti128(ToYMM(Dst), ToYMM(Dst), DstReg, 0);
}
}
} else {
vmovapd(xmm15, DestVector);
@@ -3018,6 +3332,23 @@ DEF_OP(VShlI) {
const auto Vector = GetSrc(Op->Vector.ID());
switch (ElementSize) {
case 1: {
const auto Mask = 0xFFU >> BitShift;
mov(rax, Mask);
vmovq(xmm15, rax);
if (Is256Bit) {
vpsllw(ToYMM(Dst), ToYMM(Vector), BitShift);
vpbroadcastb(ymm15, xmm15);
vpand(ToYMM(Dst), ToYMM(Dst), ymm15);
} else {
vpsllw(Dst, Vector, BitShift);
vpbroadcastb(xmm15, xmm15);
vpand(Dst, Dst, ymm15);
}
break;
}
case 2: {
if (Is256Bit) {
vpsllw(ToYMM(Dst), ToYMM(Vector), BitShift);
@@ -4175,6 +4506,8 @@ void X86JITCore::RegisterVectorHandlers() {
REGISTER_OP(VZIP2, VZip2);
REGISTER_OP(VUNZIP, VUnZip);
REGISTER_OP(VUNZIP2, VUnZip2);
REGISTER_OP(VTRN, VTrn);
REGISTER_OP(VTRN2, VTrn2);
REGISTER_OP(VBSL, VBSL);
REGISTER_OP(VCMPEQ, VCMPEQ);
REGISTER_OP(VCMPEQZ, VCMPEQZ);
+21 -15
View File
@@ -14,9 +14,13 @@ $end_info$
#include <sys/mman.h>
namespace FEXCore {
LookupCache::LookupCache(FEXCore::Context::Context *CTX)
LookupCache::LookupCache(FEXCore::Context::ContextImpl *CTX)
: ctx {CTX} {
TotalCacheSize = ctx->Config.VirtualMemSize / 4096 * 8 + CODE_SIZE + L1_SIZE;
// Setup our PMR map.
BlockLinks = BlockLinks_pma.new_object<BlockLinksMapType>();
// Block cache ends up looking like this
// PageMemoryMap[VirtualMemoryRegion >> 12]
// |
@@ -29,46 +33,48 @@ LookupCache::LookupCache(FEXCore::Context::Context *CTX)
// Allocate a region of memory that we can use to back our block pointers
// We need one pointer per page of virtual memory
// At 64GB of virtual memory this will allocate 128MB of virtual memory space
PagePointer = reinterpret_cast<uintptr_t>(FEXCore::Allocator::mmap(nullptr, ctx->Config.VirtualMemSize / 4096 * 8, PROT_READ | PROT_WRITE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0));
PagePointer = reinterpret_cast<uintptr_t>(FEXCore::Allocator::mmap(nullptr, TotalCacheSize, PROT_READ | PROT_WRITE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0));
// Allocate our memory backing our pages
// We need 32KB per guest page (One pointer per byte)
// XXX: We can drop down to 16KB if we store 4byte offsets from the code base
// We currently limit to 128MB of real memory for caching for the total cache size.
// Can end up being inefficient if we compile a small number of blocks per page
PageMemory = reinterpret_cast<uintptr_t>(FEXCore::Allocator::mmap(nullptr, CODE_SIZE, PROT_READ | PROT_WRITE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0));
PageMemory = PagePointer + ctx->Config.VirtualMemSize / 4096 * 8;
LOGMAN_THROW_AA_FMT(PageMemory != -1ULL, "Failed to allocate page memory");
// L1 Cache
L1Pointer = reinterpret_cast<uintptr_t>(FEXCore::Allocator::mmap(nullptr, L1_SIZE, PROT_READ | PROT_WRITE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0));
L1Pointer = PageMemory + CODE_SIZE;
LOGMAN_THROW_AA_FMT(L1Pointer != -1ULL, "Failed to allocate L1Pointer");
VirtualMemSize = ctx->Config.VirtualMemSize;
}
LookupCache::~LookupCache() {
FEXCore::Allocator::munmap(reinterpret_cast<void*>(PagePointer), ctx->Config.VirtualMemSize / 4096 * 8);
FEXCore::Allocator::munmap(reinterpret_cast<void*>(PageMemory), CODE_SIZE);
FEXCore::Allocator::munmap(reinterpret_cast<void*>(L1Pointer), L1_SIZE);
const size_t TotalCacheSize = ctx->Config.VirtualMemSize / 4096 * 8 + CODE_SIZE + L1_SIZE;
FEXCore::Allocator::munmap(reinterpret_cast<void*>(PagePointer), TotalCacheSize);
// No need to free BlockLinks map.
// These will get freed when their memory allocators are deallocated.
}
void LookupCache::ClearL2Cache() {
std::lock_guard<std::recursive_mutex> lk(WriteLock);
// Clear out the page memory
madvise(reinterpret_cast<void*>(PagePointer), ctx->Config.VirtualMemSize / 4096 * 8, MADV_DONTNEED);
madvise(reinterpret_cast<void*>(PageMemory), CODE_SIZE, MADV_DONTNEED);
// PagePointer and PageMemory are sequential with each other. Clear both at once.
madvise(reinterpret_cast<void*>(PagePointer), ctx->Config.VirtualMemSize / 4096 * 8 + CODE_SIZE, MADV_DONTNEED);
AllocateOffset = 0;
}
void LookupCache::ClearCache() {
std::lock_guard<std::recursive_mutex> lk(WriteLock);
// Clear L1
madvise(reinterpret_cast<void*>(L1Pointer), L1_SIZE, MADV_DONTNEED);
// Clear L2
ClearL2Cache();
// All code is gone, remove links
BlockLinks.clear();
// Clear L1 and L2 by clearing the full cache.
madvise(reinterpret_cast<void*>(PagePointer), TotalCacheSize, MADV_DONTNEED);
// Clear the BlockLinks allocator which frees the BlockLinks map implicitly.
BlockLinks_mbr.release();
// Allocate a new pointer from the BlockLinks pma again.
BlockLinks = BlockLinks_pma.new_object<BlockLinksMapType>();
// All code is gone, clear the block list
BlockList.clear();
}
+20 -10
View File
@@ -1,9 +1,11 @@
#pragma once
#include "Interface/Context/Context.h"
#include <FEXCore/Utils/LogManager.h>
#include <cstdint>
#include <functional>
#include <map>
#include <memory_resource>
#include <stddef.h>
#include <utility>
#include <vector>
@@ -11,9 +13,6 @@
#include <tsl/robin_map.h>
namespace FEXCore {
namespace Context {
struct Context;
}
class LookupCache {
public:
@@ -23,7 +22,7 @@ public:
uintptr_t GuestCode;
};
LookupCache(FEXCore::Context::Context *CTX);
LookupCache(FEXCore::Context::ContextImpl *CTX);
~LookupCache();
uintptr_t FindBlock(uint64_t Address) {
@@ -105,9 +104,9 @@ public:
std::lock_guard<std::recursive_mutex> lk(WriteLock);
// Sever any links to this block
auto lower = BlockLinks.lower_bound({Address, 0});
auto upper = BlockLinks.upper_bound({Address, UINTPTR_MAX});
for (auto it = lower; it != upper; it = BlockLinks.erase(it)) {
auto lower = BlockLinks->lower_bound({Address, 0});
auto upper = BlockLinks->upper_bound({Address, UINTPTR_MAX});
for (auto it = lower; it != upper; it = BlockLinks->erase(it)) {
it->second();
}
@@ -145,7 +144,7 @@ public:
void AddBlockLink(uint64_t GuestDestination, uintptr_t HostLink, const std::function<void()> &delinker) {
std::lock_guard<std::recursive_mutex> lk(WriteLock);
BlockLinks.insert({{GuestDestination, HostLink}, delinker});
BlockLinks->insert({{GuestDestination, HostLink}, delinker});
}
void ClearCache();
@@ -238,17 +237,28 @@ private:
}
};
// Use a monotonic buffer resource to allocate both the std::pmr::map and its members.
// This allows us to quickly clear the block link map by clearing the monotonic allocator.
// If we had allocated the block link map without the MBR, then clearing the map would require slowly
// walking each block member and destructing objects.
//
// This makes `BlockLinks` look like a raw pointer that could memory leak, but since it is backed by the MBR, it won't.
std::pmr::monotonic_buffer_resource BlockLinks_mbr;
using BlockLinksMapType = std::pmr::map<BlockLinkTag, std::function<void()>>;
std::pmr::polymorphic_allocator<std::byte> BlockLinks_pma {&BlockLinks_mbr};
BlockLinksMapType *BlockLinks;
std::map<BlockLinkTag, std::function<void()>> BlockLinks;
tsl::robin_map<uint64_t, uint64_t> BlockList;
size_t TotalCacheSize;
constexpr static size_t CODE_SIZE = 128 * 1024 * 1024;
constexpr static size_t SIZE_PER_PAGE = 4096 * sizeof(LookupCacheEntry);
constexpr static size_t L1_SIZE = L1_ENTRIES * sizeof(LookupCacheEntry);
size_t AllocateOffset {};
FEXCore::Context::Context *ctx;
FEXCore::Context::ContextImpl *ctx;
uint64_t VirtualMemSize{};
};
}
@@ -4,7 +4,7 @@
#include <FEXCore/Config/Config.h>
namespace FEXCore::CodeSerialize {
NamedRegionObjectHandler::NamedRegionObjectHandler(FEXCore::Context::Context *ctx) {
NamedRegionObjectHandler::NamedRegionObjectHandler(FEXCore::Context::ContextImpl *ctx) {
DefaultSerializationConfig.Cookie = CODE_COOKIE;
// Initialize the Arch from CPUID
@@ -13,7 +13,7 @@ namespace {
}
namespace FEXCore::CodeSerialize {
CodeObjectSerializeService::CodeObjectSerializeService(FEXCore::Context::Context *ctx)
CodeObjectSerializeService::CodeObjectSerializeService(FEXCore::Context::ContextImpl *ctx)
: CTX {ctx}
, AsyncHandler { &NamedRegionHandler , this }
, NamedRegionHandler { ctx } {
@@ -253,7 +253,7 @@ namespace FEXCore::CodeSerialize {
class NamedRegionObjectHandler final {
public:
NamedRegionObjectHandler(FEXCore::Context::Context *ctx);
NamedRegionObjectHandler(FEXCore::Context::ContextImpl *ctx);
void HandleNamedRegionObjectJobs();
@@ -338,7 +338,7 @@ namespace FEXCore::CodeSerialize {
*/
class CodeObjectSerializeService final {
public:
CodeObjectSerializeService(FEXCore::Context::Context *ctx);
CodeObjectSerializeService(FEXCore::Context::ContextImpl *ctx);
/**
* @brief Initialize the internal interface
@@ -440,7 +440,7 @@ namespace FEXCore::CodeSerialize {
void NotifyWork() { WorkAvailable.NotifyOne(); }
private:
FEXCore::Context::Context *CTX;
FEXCore::Context::ContextImpl *CTX;
Event WorkAvailable{};
std::unique_ptr<FEXCore::Threads::Thread> WorkerThread;
File diff suppressed because it is too large. Load diff
+243 -26
View File
@@ -75,7 +75,7 @@ public:
OrderedNode* flagsOpDestSigned{};
OrderedNode* flagsOpSrcSigned{};
FEXCore::Context::Context *CTX{};
FEXCore::Context::ContextImpl *CTX{};
// Used during new op bringup
bool ShouldDump {false};
@@ -149,12 +149,13 @@ public:
return false;
}
OpDispatchBuilder(FEXCore::Context::Context *ctx);
OpDispatchBuilder(FEXCore::Context::ContextImpl *ctx);
OpDispatchBuilder(FEXCore::Utils::IntrusivePooledAllocator &Allocator);
void ResetWorkingList();
void ResetDecodeFailure() { DecodeFailure = false; }
void ResetDecodeFailure() { NeedsBlockEnd = DecodeFailure = false; }
bool HadDecodeFailure() const { return DecodeFailure; }
bool NeedsBlockEnder() const { return NeedsBlockEnd; }
void BeginFunction(uint64_t RIP, std::vector<FEXCore::Frontend::Decoder::DecodedBlocks> const *Blocks);
void Finalize();
@@ -167,6 +168,7 @@ public:
void MOVGPRNTOp(OpcodeArgs);
void MOVVectorOp(OpcodeArgs);
void MOVVectorNTOp(OpcodeArgs);
template<FEXCore::IR::IROps ALUIROp, FEXCore::IR::IROps AtomicFetchOp, bool RequiresMask>
void ALUOp(OpcodeArgs);
void INTOp(OpcodeArgs);
void SyscallOp(OpcodeArgs);
@@ -175,7 +177,6 @@ public:
void NOPOp(OpcodeArgs);
void RETOp(OpcodeArgs);
void IRETOp(OpcodeArgs);
void SIGRETOp(OpcodeArgs);
void CallbackReturnOp(OpcodeArgs);
void SecondaryALUOp(OpcodeArgs);
template<uint32_t SrcIndex>
@@ -303,7 +304,6 @@ public:
// SSE
void MOVAPSOp(OpcodeArgs);
void MOVUPSOp(OpcodeArgs);
void MOVLHPSOp(OpcodeArgs);
void MOVLPOp(OpcodeArgs);
void MOVSHDUPOp(OpcodeArgs);
void MOVSLDUPOp(OpcodeArgs);
@@ -322,10 +322,6 @@ public:
void MOVQOp(OpcodeArgs);
template<size_t ElementSize>
void PADDQOp(OpcodeArgs);
template<size_t ElementSize>
void PSUBQOp(OpcodeArgs);
template<size_t ElementSize>
void MOVMSKOp(OpcodeArgs);
void MOVMSKOpOne(OpcodeArgs);
template<size_t ElementSize>
@@ -350,8 +346,6 @@ public:
void PSLLDQ(OpcodeArgs);
template<size_t ElementSize>
void PSRAIOp(OpcodeArgs);
template<size_t ElementSize>
void PAVGOp(OpcodeArgs);
void MOVDDUPOp(OpcodeArgs);
template<size_t DstElementSize>
void CVTGPR_To_FPR(OpcodeArgs);
@@ -378,15 +372,16 @@ public:
void VFCMPOp(OpcodeArgs);
template<size_t ElementSize>
void SHUFOp(OpcodeArgs);
void ANDNOp(OpcodeArgs);
template<size_t ElementSize>
void PINSROp(OpcodeArgs);
void InsertPSOp(OpcodeArgs);
template<size_t ElementSize>
void PExtrOp(OpcodeArgs);
template<size_t ElementSize>
template <size_t ElementSize>
void PSIGN(OpcodeArgs);
template <size_t ElementSize>
void VPSIGN(OpcodeArgs);
// BMI1 Ops
void ANDNBMIOp(OpcodeArgs);
@@ -409,9 +404,58 @@ public:
// AVX Ops
template <IROps IROp, size_t ElementSize>
void AVXVectorALUOp(OpcodeArgs);
template <IROps IROp, size_t ElementSize>
void AVXVectorScalarALUOp(OpcodeArgs);
template <IROps IROp, size_t ElementSize, bool Scalar>
void AVXVectorUnaryOp(OpcodeArgs);
template <size_t ElementSize, size_t DstElementSize, bool Signed>
void AVXExtendVectorElements(OpcodeArgs);
template <size_t ElementSize, bool Scalar>
void AVXVectorRound(OpcodeArgs);
template <size_t SrcElementSize, bool Narrow, bool HostRoundingMode>
void AVXVector_CVT_Float_To_Int(OpcodeArgs);
template <size_t SrcElementSize, bool Widen>
void AVXVector_CVT_Int_To_Float(OpcodeArgs);
template <size_t ElementSize, bool Scalar>
void AVXVFCMPOp(OpcodeArgs);
template <size_t ElementSize>
void VADDSUBPOp(OpcodeArgs);
void VAESDecOp(OpcodeArgs);
void VAESDecLastOp(OpcodeArgs);
void VAESEncOp(OpcodeArgs);
void VAESEncLastOp(OpcodeArgs);
void VAESIMCOp(OpcodeArgs);
void VAESKeyGenAssistOp(OpcodeArgs);
void VANDNOp(OpcodeArgs);
void VBLENDPDOp(OpcodeArgs);
void VPBLENDDOp(OpcodeArgs);
void VPBLENDWOp(OpcodeArgs);
template <size_t ElementSize>
void VBROADCASTOp(OpcodeArgs);
template <size_t ElementSize>
void VDPPOp(OpcodeArgs);
void VEXTRACT128Op(OpcodeArgs);
template <IROps IROp, size_t ElementSize>
void VHADDPOp(OpcodeArgs);
template <size_t ElementSize>
void VHSUBPOp(OpcodeArgs);
void VINSERTOp(OpcodeArgs);
void VINSERTPSOp(OpcodeArgs);
void VMOVAPS_VMOVAPD_Op(OpcodeArgs);
void VMOVUPS_VMOVUPD_Op(OpcodeArgs);
@@ -422,8 +466,81 @@ public:
void VMOVSHDUPOp(OpcodeArgs);
void VMOVSLDUPOp(OpcodeArgs);
void VMOVSDOp(OpcodeArgs);
void VMOVSSOp(OpcodeArgs);
void VMOVVectorNTOp(OpcodeArgs);
template <size_t ElementSize>
void VPACKSSOp(OpcodeArgs);
template <size_t ElementSize>
void VPACKUSOp(OpcodeArgs);
void VPALIGNROp(OpcodeArgs);
void VPERM2Op(OpcodeArgs);
void VPERMDOp(OpcodeArgs);
void VPERMQOp(OpcodeArgs);
template <size_t ElementSize>
void VPERMILImmOp(OpcodeArgs);
template <size_t ElementSize>
void VPERMILRegOp(OpcodeArgs);
void VPHADDSWOp(OpcodeArgs);
void VPHMINPOSUWOp(OpcodeArgs);
template <size_t ElementSize>
void VPHSUBOp(OpcodeArgs);
void VPHSUBSWOp(OpcodeArgs);
void VPMADDWDOp(OpcodeArgs);
void VPMULHRSWOp(OpcodeArgs);
template <bool Signed>
void VPMULHWOp(OpcodeArgs);
template <size_t ElementSize, bool Signed>
void VPMULLOp(OpcodeArgs);
void VPSHUFBOp(OpcodeArgs);
template <size_t ElementSize, bool Low>
void VPSHUFWOp(OpcodeArgs);
template <size_t ElementSize>
void VPSLLOp(OpcodeArgs);
void VPSLLDQOp(OpcodeArgs);
template <size_t ElementSize>
void VPSLLIOp(OpcodeArgs);
template <size_t ElementSize>
void VPSRAOp(OpcodeArgs);
template <size_t ElementSize>
void VPSRAIOp(OpcodeArgs);
void VPSRAVDOp(OpcodeArgs);
template <size_t ElementSize>
void VPSRLDOp(OpcodeArgs);
void VPSRLDQOp(OpcodeArgs);
template <size_t ElementSize>
void VPUNPCKHOp(OpcodeArgs);
template <size_t ElementSize>
void VPUNPCKLOp(OpcodeArgs);
template <size_t ElementSize>
void VPSRLIOp(OpcodeArgs);
template <size_t ElementSize>
void VSHUFOp(OpcodeArgs);
void VZEROOp(OpcodeArgs);
// X87 Ops
@@ -568,12 +685,6 @@ public:
template<bool ToXMM>
void MOVQ2DQ(OpcodeArgs);
template<size_t ElementSize, bool Signed>
void PADDSOp(OpcodeArgs);
template<size_t ElementSize, bool Signed>
void PSUBSOp(OpcodeArgs);
template<size_t ElementSize>
void ADDSUBPOp(OpcodeArgs);
@@ -598,12 +709,7 @@ public:
void MOVBEOp(OpcodeArgs);
template<size_t ElementSize>
void HADDP(OpcodeArgs);
template<size_t ElementSize>
void HSUBP(OpcodeArgs);
template<size_t ElementSize>
void PHADD(OpcodeArgs);
template<size_t ElementSize>
void PHSUB(OpcodeArgs);
@@ -613,6 +719,9 @@ public:
template<uint8_t FenceType>
void FenceOp(OpcodeArgs);
void CLWB(OpcodeArgs);
void CLFLUSHOPT(OpcodeArgs);
void MemFenceOrXSAVEOPT(OpcodeArgs);
void StoreFenceOrCLFlush(OpcodeArgs);
void CLZeroOp(OpcodeArgs);
void RDTSCPOp(OpcodeArgs);
@@ -660,8 +769,6 @@ public:
void InvalidOp(OpcodeArgs);
#undef OpcodeArgs
void SetPackedRFLAG(bool Lower8, OrderedNode *Src);
OrderedNode *GetPackedRFLAG(bool Lower8);
@@ -670,10 +777,120 @@ public:
bool HandledLock = false;
private:
bool DecodeFailure{false};
bool NeedsBlockEnd{false};
FEXCore::IR::IROp_IRHeader *Current_Header{};
OrderedNode *Current_HeaderNode{};
void ALUOpImpl(OpcodeArgs, FEXCore::IR::IROps ALUIROp, FEXCore::IR::IROps AtomicFetchOp, bool RequiresMask);
// Opcode helpers for generalizing behavior across VEX and non-VEX variants.
OrderedNode* ADDSUBPOpImpl(OpcodeArgs, size_t ElementSize,
OrderedNode *Src1, OrderedNode *Src2);
void AVXVectorALUOpImpl(OpcodeArgs, IROps IROp, size_t ElementSize);
void AVXVectorScalarALUOpImpl(OpcodeArgs, IROps IROp, size_t ElementSize);
void AVXVectorUnaryOpImpl(OpcodeArgs, IROps IROp, size_t ElementSize, bool Scalar);
template <size_t ElementSize>
void AVXVectorVariableBlend(OpcodeArgs);
OrderedNode* AESKeyGenAssistImpl(OpcodeArgs);
OrderedNode* AESIMCImpl(OpcodeArgs);
OrderedNode* DPPOpImpl(OpcodeArgs, const X86Tables::DecodedOperand& Src1,
const X86Tables::DecodedOperand& Src2,
const X86Tables::DecodedOperand& Imm, size_t ElementSize);
OrderedNode* ExtendVectorElementsImpl(OpcodeArgs, size_t ElementSize,
size_t DstElementSize, bool Signed);
OrderedNode* HSUBPOpImpl(OpcodeArgs, size_t ElementSize,
const X86Tables::DecodedOperand& Src1Op,
const X86Tables::DecodedOperand& Src2Op);
OrderedNode* InsertPSOpImpl(OpcodeArgs, const X86Tables::DecodedOperand& Src1,
const X86Tables::DecodedOperand& Src2,
const X86Tables::DecodedOperand& Imm);
OrderedNode* PACKSSOpImpl(OpcodeArgs, size_t ElementSize,
OrderedNode *Src1, OrderedNode *Src2);
OrderedNode* PACKUSOpImpl(OpcodeArgs, size_t ElementSize,
OrderedNode *Src1, OrderedNode *Src2);
OrderedNode* PALIGNROpImpl(OpcodeArgs, const X86Tables::DecodedOperand& Src1,
const X86Tables::DecodedOperand& Src2,
const X86Tables::DecodedOperand& Imm);
OrderedNode* PHADDSOpImpl(OpcodeArgs, const X86Tables::DecodedOperand& Src1,
const X86Tables::DecodedOperand& Src2);
OrderedNode* PHMINPOSUWOpImpl(OpcodeArgs);
OrderedNode* PHSUBOpImpl(OpcodeArgs, const X86Tables::DecodedOperand& Src1,
const X86Tables::DecodedOperand& Src2, size_t ElementSize);
OrderedNode* PHSUBSOpImpl(OpcodeArgs, const X86Tables::DecodedOperand& Src1Op,
const X86Tables::DecodedOperand& Src2Op);
OrderedNode* PMADDWDOpImpl(OpcodeArgs, const X86Tables::DecodedOperand& Src1,
const X86Tables::DecodedOperand& Src2);
OrderedNode* PMULHRSWOpImpl(OpcodeArgs, OrderedNode *Src1, OrderedNode *Src2);
OrderedNode* PMULHWOpImpl(OpcodeArgs, bool Signed,
OrderedNode *Src1, OrderedNode *Src2);
OrderedNode* PMULLOpImpl(OpcodeArgs, size_t ElementSize, bool Signed,
OrderedNode *Src1, OrderedNode *Src2);
OrderedNode* PSHUFBOpImpl(OpcodeArgs, const X86Tables::DecodedOperand& Src1,
const X86Tables::DecodedOperand& Src2);
OrderedNode* PSIGNImpl(OpcodeArgs, size_t ElementSize,
OrderedNode *Src1, OrderedNode *Src2);
OrderedNode* PSLLIImpl(OpcodeArgs, size_t ElementSize,
OrderedNode *Src, uint64_t Shift);
OrderedNode* PSLLImpl(OpcodeArgs, size_t ElementSize,
OrderedNode *Src, OrderedNode *ShiftVec);
OrderedNode* PSRAOpImpl(OpcodeArgs, size_t ElementSize,
OrderedNode *Src, OrderedNode *ShiftVec);
OrderedNode* PSRLDOpImpl(OpcodeArgs, size_t ElementSize,
OrderedNode *Src, OrderedNode *ShiftVec);
OrderedNode* SHUFOpImpl(OpcodeArgs, size_t ElementSize,
const X86Tables::DecodedOperand& Src1,
const X86Tables::DecodedOperand& Src2,
const X86Tables::DecodedOperand& Imm);
void VMOVScalarOpImpl(OpcodeArgs, size_t ElementSize);
OrderedNode* VFCMPOpImpl(OpcodeArgs, size_t ElementSize, bool Scalar,
OrderedNode *Src1, OrderedNode *Src2, uint8_t CompType);
void VectorALUOpImpl(OpcodeArgs, IROps IROp, size_t ElementSize);
void VectorALUROpImpl(OpcodeArgs, IROps IROp, size_t ElementSize);
void VectorScalarALUOpImpl(OpcodeArgs, IROps IROp, size_t ElementSize);
void VectorUnaryOpImpl(OpcodeArgs, IROps IROp, size_t ElementSize, bool Scalar);
void VectorUnaryDuplicateOpImpl(OpcodeArgs, IROps IROp, size_t ElementSize);
OrderedNode* VectorRoundImpl(OpcodeArgs, size_t ElementSize,
OrderedNode *Src, uint64_t Mode);
OrderedNode* Vector_CVT_Float_To_IntImpl(OpcodeArgs, size_t SrcElementSize, bool Narrow, bool HostRoundingMode);
OrderedNode* Vector_CVT_Int_To_FloatImpl(OpcodeArgs, size_t SrcElementSize, bool Widen);
#undef OpcodeArgs
OrderedNode *AppendSegmentOffset(OrderedNode *Value, uint32_t Flags, uint32_t DefaultPrefix = 0, bool Override = false);
OrderedNode *GetSegment(uint32_t Flags, uint32_t DefaultPrefix = 0, bool Override = false);
void UpdatePrefixFromSegment(OrderedNode *Segment, uint32_t SegmentReg);
enum class MemoryAccessType {
@@ -35,19 +35,16 @@ void OpDispatchBuilder::SHA1MSG1Op(OpcodeArgs) {
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
auto W0 = _VExtractToGPR(16, 4, Dest, 3);
auto W1 = _VExtractToGPR(16, 4, Dest, 2);
auto W2 = _VExtractToGPR(16, 4, Dest, 1);
auto W3 = _VExtractToGPR(16, 4, Dest, 0);
auto W4 = _VExtractToGPR(16, 4, Src, 3);
auto W5 = _VExtractToGPR(16, 4, Src, 2);
OrderedNode *NewVec{};
NewVec = _VInsElement(16, 4, 3, 1, Dest, Dest);
NewVec = _VInsElement(16, 4, 2, 0, NewVec, Dest);
NewVec = _VInsElement(16, 4, 1, 3, NewVec, Src);
NewVec = _VInsElement(16, 4, 0, 2, NewVec, Src);
auto D3 = _VInsGPR(16, 4, 3, Dest, _Xor(W2, W0));
auto D2 = _VInsGPR(16, 4, 2, D3, _Xor(W3, W1));
auto D1 = _VInsGPR(16, 4, 1, D2, _Xor(W4, W2));
auto D0 = _VInsGPR(16, 4, 0, D1, _Xor(W5, W3));
// [W0, W1, W2, W3] ^ [W2, W3, W4, W5]
OrderedNode *Result = _VXor(16, 1, Dest, NewVec);
StoreResult(FPRClass, Op, D0, -1);
StoreResult(FPRClass, Op, Result, -1);
}
void OpDispatchBuilder::SHA1MSG2Op(OpcodeArgs) {
@@ -264,47 +261,135 @@ void OpDispatchBuilder::SHA256RNDS2Op(OpcodeArgs) {
StoreResult(FPRClass, Op, Res0, -1);
}
void OpDispatchBuilder::AESImcOp(OpcodeArgs) {
OrderedNode* OpDispatchBuilder::AESIMCImpl(OpcodeArgs) {
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
auto Res = _VAESImc(Src);
StoreResult(FPRClass, Op, Res, -1);
return _VAESImc(Src);
}
void OpDispatchBuilder::AESImcOp(OpcodeArgs) {
OrderedNode *Result = AESIMCImpl(Op);
StoreResult(FPRClass, Op, Result, -1);
}
void OpDispatchBuilder::VAESIMCOp(OpcodeArgs) {
OrderedNode *Mixed = AESIMCImpl(Op);
OrderedNode *Result = _VMov(16, Mixed);
StoreResult(FPRClass, Op, Result, -1);
}
void OpDispatchBuilder::AESEncOp(OpcodeArgs) {
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
auto Res = _VAESEnc(Dest, Src);
StoreResult(FPRClass, Op, Res, -1);
OrderedNode *Result = _VAESEnc(16, Dest, Src);
StoreResult(FPRClass, Op, Result, -1);
}
void OpDispatchBuilder::VAESEncOp(OpcodeArgs) {
const auto DstSize = GetDstSize(Op);
const auto Is128Bit = DstSize == Core::CPUState::XMM_SSE_REG_SIZE;
// TODO: Handle 256-bit VAESENC.
LOGMAN_THROW_A_FMT(Is128Bit, "256-bit VAESENC unimplemented");
OrderedNode *State = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *Key = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags, -1);
OrderedNode *Result = _VAESEnc(DstSize, State, Key);
if (Is128Bit) {
Result = _VMov(16, Result);
}
StoreResult(FPRClass, Op, Result, -1);
}
void OpDispatchBuilder::AESEncLastOp(OpcodeArgs) {
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
auto Res = _VAESEncLast(Dest, Src);
StoreResult(FPRClass, Op, Res, -1);
OrderedNode *Result = _VAESEncLast(16, Dest, Src);
StoreResult(FPRClass, Op, Result, -1);
}
void OpDispatchBuilder::VAESEncLastOp(OpcodeArgs) {
const auto DstSize = GetDstSize(Op);
const auto Is128Bit = DstSize == Core::CPUState::XMM_SSE_REG_SIZE;
// TODO: Handle 256-bit VAESENCLAST.
LOGMAN_THROW_A_FMT(Is128Bit, "256-bit VAESENCLAST unimplemented");
OrderedNode *State = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *Key = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags, -1);
OrderedNode *Result = _VAESEncLast(DstSize, State, Key);
if (Is128Bit) {
Result = _VMov(16, Result);
}
StoreResult(FPRClass, Op, Result, -1);
}
void OpDispatchBuilder::AESDecOp(OpcodeArgs) {
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
auto Res = _VAESDec(Dest, Src);
StoreResult(FPRClass, Op, Res, -1);
OrderedNode *Result = _VAESDec(16, Dest, Src);
StoreResult(FPRClass, Op, Result, -1);
}
void OpDispatchBuilder::VAESDecOp(OpcodeArgs) {
const auto DstSize = GetDstSize(Op);
const auto Is128Bit = DstSize == Core::CPUState::XMM_SSE_REG_SIZE;
// TODO: Handle 256-bit VAESDEC.
LOGMAN_THROW_A_FMT(Is128Bit, "256-bit VAESDEC unimplemented");
OrderedNode *State = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *Key = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags, -1);
OrderedNode *Result = _VAESDec(DstSize, State, Key);
if (Is128Bit) {
Result = _VMov(16, Result);
}
StoreResult(FPRClass, Op, Result, -1);
}
void OpDispatchBuilder::AESDecLastOp(OpcodeArgs) {
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
auto Res = _VAESDecLast(Dest, Src);
StoreResult(FPRClass, Op, Res, -1);
OrderedNode *Result = _VAESDecLast(16, Dest, Src);
StoreResult(FPRClass, Op, Result, -1);
}
void OpDispatchBuilder::VAESDecLastOp(OpcodeArgs) {
const auto DstSize = GetDstSize(Op);
const auto Is128Bit = DstSize == Core::CPUState::XMM_SSE_REG_SIZE;
// TODO: Handle 256-bit VAESDECLAST.
LOGMAN_THROW_A_FMT(Is128Bit, "256-bit VAESDECLAST unimplemented");
OrderedNode *State = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *Key = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags, -1);
OrderedNode *Result = _VAESDecLast(DstSize, State, Key);
if (Is128Bit) {
Result = _VMov(16, Result);
}
StoreResult(FPRClass, Op, Result, -1);
}
OrderedNode* OpDispatchBuilder::AESKeyGenAssistImpl(OpcodeArgs) {
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
LOGMAN_THROW_A_FMT(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
const uint64_t RCON = Op->Src[1].Data.Literal.Value;
return _VAESKeyGenAssist(Src, RCON);
}
void OpDispatchBuilder::AESKeyGenAssist(OpcodeArgs) {
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
LOGMAN_THROW_A_FMT(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
uint64_t RCON = Op->Src[1].Data.Literal.Value;
OrderedNode *Result = AESKeyGenAssistImpl(Op);
StoreResult(FPRClass, Op, Result, -1);
}
auto Res = _VAESKeyGenAssist(Src, RCON);
StoreResult(FPRClass, Op, Res, -1);
void OpDispatchBuilder::VAESKeyGenAssistOp(OpcodeArgs) {
OrderedNode *Assist = AESKeyGenAssistImpl(Op);
OrderedNode *Result = _VMov(16, Assist);
StoreResult(FPRClass, Op, Result, -1);
}
void OpDispatchBuilder::PCLMULQDQOp(OpcodeArgs) {
@@ -314,18 +399,25 @@ void OpDispatchBuilder::PCLMULQDQOp(OpcodeArgs) {
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
const auto Selector = static_cast<uint8_t>(Op->Src[1].Data.Literal.Value);
auto Res = _PCLMUL(Dest, Src, Selector);
auto Res = _PCLMUL(16, Dest, Src, Selector);
StoreResult(FPRClass, Op, Res, -1);
}
void OpDispatchBuilder::VPCLMULQDQOp(OpcodeArgs) {
LOGMAN_THROW_A_FMT(Op->Src[2].IsLiteral(), "Selector needs to be literal here");
const auto DstSize = GetDstSize(Op);
const auto Is128Bit = DstSize == Core::CPUState::XMM_SSE_REG_SIZE;
OrderedNode *Src1 = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *Src2 = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags, -1);
const auto Selector = static_cast<uint8_t>(Op->Src[2].Data.Literal.Value);
auto Res = _PCLMUL(Src1, Src2, Selector);
OrderedNode *Res = _PCLMUL(DstSize, Src1, Src2, Selector);
if (Is128Bit) {
Res = _VMov(16, Res);
}
StoreResult(FPRClass, Op, Res, -1);
}
@@ -271,10 +271,8 @@ void OpDispatchBuilder::CalculcateFlags_ADC(uint8_t SrcSize, OrderedNode *Res, O
// SF
{
auto SignBitConst = _Constant(Size - 1);
auto LshrOp = _Lshr(Res, SignBitConst);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(LshrOp);
auto SignOp = _Bfe(1, SrcSize * 8 - 1, Res);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(SignOp);
}
// PF
@@ -342,10 +340,8 @@ void OpDispatchBuilder::CalculcateFlags_SBB(uint8_t SrcSize, OrderedNode *Res, O
// SF
{
auto SignBitConst = _Constant(SrcSize * 8 - 1);
auto LshrOp = _Lshr(Res, SignBitConst);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(LshrOp);
auto SignOp = _Bfe(1, SrcSize * 8 - 1, Res);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(SignOp);
}
// PF
@@ -412,10 +408,8 @@ void OpDispatchBuilder::CalculcateFlags_SUB(uint8_t SrcSize, OrderedNode *Res, O
// SF
{
auto SignBitConst = _Constant(SrcSize * 8 - 1);
auto LshrOp = _Lshr(Res, SignBitConst);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(LshrOp);
auto SignOp = _Bfe(1, SrcSize * 8 - 1, Res);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(SignOp);
}
// PF
@@ -469,10 +463,8 @@ void OpDispatchBuilder::CalculcateFlags_ADD(uint8_t SrcSize, OrderedNode *Res, O
// SF
{
auto SignBitConst = _Constant(SrcSize * 8 - 1);
auto LshrOp = _Lshr(Res, SignBitConst);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(LshrOp);
auto SignOp = _Bfe(1, SrcSize * 8 - 1, Res);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(SignOp);
}
// PF
@@ -583,10 +575,8 @@ void OpDispatchBuilder::CalculcateFlags_Logical(uint8_t SrcSize, OrderedNode *Re
// SF
{
auto SignBitConst = _Constant(SrcSize * 8 - 1);
auto LshrOp = _Lshr(Res, SignBitConst);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(LshrOp);
auto SignOp = _Bfe(1, SrcSize * 8 - 1, Res);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(SignOp);
}
// PF
@@ -750,10 +740,8 @@ void OpDispatchBuilder::CalculcateFlags_SignShiftRight(uint8_t SrcSize, OrderedN
// SF
{
auto SignBitConst = _Constant(SrcSize * 8 - 1);
auto LshrOp = _Lshr(Res, SignBitConst);
COND_FLAG_SET(Src2, RFLAG_SF_LOC, LshrOp);
auto SignBitOp = _Bfe(1, SrcSize * 8 - 1, Res);
COND_FLAG_SET(Src2, RFLAG_SF_LOC, SignBitOp);
}
// OF
@@ -802,15 +790,14 @@ void OpDispatchBuilder::CalculcateFlags_ShiftLeftImmediate(uint8_t SrcSize, Orde
// SF
{
auto LshrOp = _Bfe(1, SrcSize * 8 - 1, Res);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(LshrOp);
auto SignOp = _Bfe(1, SrcSize * 8 - 1, Res);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(SignOp);
// OF
// In the case of left shift. OF is only set from the result of <Top Source Bit> XOR <Top Result Bit>
if (Shift == 1) {
auto SourceBit = _Bfe(1, SrcSize * 8 - 1, Src1);
SetRFLAG<FEXCore::X86State::RFLAG_OF_LOC>(_Xor(SourceBit, LshrOp));
SetRFLAG<FEXCore::X86State::RFLAG_OF_LOC>(_Xor(SourceBit, SignOp));
}
}
}
@@ -851,10 +838,8 @@ void OpDispatchBuilder::CalculcateFlags_SignShiftRightImmediate(uint8_t SrcSize,
// SF
{
auto SignBitConst = _Constant(SrcSize * 8 - 1);
auto LshrOp = _Lshr(Res, SignBitConst);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(LshrOp);
auto SignBitOp = _Bfe(1, SrcSize * 8 - 1, Res);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(SignBitOp);
// OF
// Only defined when Shift is 1 else undefined
@@ -902,10 +887,8 @@ void OpDispatchBuilder::CalculcateFlags_ShiftRightImmediate(uint8_t SrcSize, Ord
// SF
{
auto SignBitConst = _Constant(SrcSize * 8 - 1);
auto LshrOp = _Lshr(Res, SignBitConst);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(LshrOp);
auto SignBitOp = _Bfe(1, SrcSize * 8 - 1, Res);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(SignBitOp);
}
// OF
@@ -1115,10 +1098,8 @@ void OpDispatchBuilder::CalculcateFlags_BLSI(uint8_t SrcSize, OrderedNode *Src)
// SF
{
auto SignBit = _Constant(SrcSize * 8 - 1);
auto SFOp = _Lshr(Src, SignBit);
SetRFLAG<X86State::RFLAG_SF_LOC>(SFOp);
auto SignOp = _Bfe(1, SrcSize * 8 - 1, Src);
SetRFLAG<X86State::RFLAG_SF_LOC>(SignOp);
}
}
@@ -1174,10 +1155,8 @@ void OpDispatchBuilder::CalculcateFlags_BLSR(uint8_t SrcSize, OrderedNode *Resul
// SF
{
auto SignBit = _Constant(SrcSize * 8 - 1);
auto SFOp = _Lshr(Result, SignBit);
SetRFLAG<X86State::RFLAG_SF_LOC>(SFOp);
auto SignOp = _Bfe(1, SrcSize * 8 - 1, Result);
SetRFLAG<X86State::RFLAG_SF_LOC>(SignOp);
}
}
@@ -1230,9 +1209,8 @@ void OpDispatchBuilder::CalculcateFlags_BZHI(uint8_t SrcSize, OrderedNode *Resul
// SF
{
auto SFOp = _Lshr(Result, Bounds);
SetRFLAG<X86State::RFLAG_SF_LOC>(SFOp);
auto SignOp = _Bfe(1, SrcSize * 8 - 1, Result);
SetRFLAG<X86State::RFLAG_SF_LOC>(SignOp);
}
}
File diff suppressed because it is too large. Load diff
@@ -1388,7 +1388,7 @@ void OpDispatchBuilder::X87FCMOV(OpcodeArgs) {
auto a = _LoadContextIndexed(top, 16, MMBaseOffset(), 16, FPRClass);
auto b = _LoadContextIndexed(arg, 16, MMBaseOffset(), 16, FPRClass);
auto Result = _VBSL(VecCond, b, a);
auto Result = _VBSL(16, VecCond, b, a);
// Write to ST[TOP]
_StoreContextIndexed(Result, top, 16, MMBaseOffset(), 16, FPRClass);
@@ -39,7 +39,7 @@ class OrderedNode;
//FST(register to register)
// State loading duplicated from X87.cpp, setting host rounding mode
// See issue
// See issue
void OpDispatchBuilder::FNINITF64(OpcodeArgs) {
// Init FCW to 0x037F
auto NewFCW = _Constant(16, 0x037F);
@@ -76,7 +76,7 @@ void OpDispatchBuilder::X87LDENVF64(OpcodeArgs) {
roundingMode = _And(roundingMode, roundMask);
_SetRoundingMode(roundingMode);
_F80LoadFCW(NewFCW);
_StoreContext(2, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
OrderedNode *MemLocation = _Add(Mem, _Constant(Size * 1));
@@ -184,7 +184,7 @@ void OpDispatchBuilder::FBLDF64(OpcodeArgs) {
void OpDispatchBuilder::FBSTPF64(OpcodeArgs) {
auto orig_top = GetX87Top();
auto data = _LoadContextIndexed(orig_top, 8, MMBaseOffset(), 16, FPRClass);
OrderedNode *converted = _F80CVTTo(data, 8);
converted = _F80BCDStore(converted);
@@ -256,7 +256,7 @@ void OpDispatchBuilder::FSTF64(OpcodeArgs) {
//Convert to 80-bit float
auto result = _F80CVTTo(data, 8);
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, result, 10, 1);
}
}
if ((Op->TableInfo->Flags & X86Tables::InstFlags::FLAGS_POP) != 0) {
// if we are popping then we must first mark this location as empty
@@ -315,7 +315,10 @@ void OpDispatchBuilder::FADDF64(OpcodeArgs) {
// Memory arg
if constexpr (Integer) {
arg = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags, -1);
b = _Float_FromGPR_S(8, 8, arg);
if(width == 16) {
arg = _Sext(16, arg);
}
b = _Float_FromGPR_S(8, width == 64 ? 8 : 4, arg);
} else if constexpr (width == 32) {
arg = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
b = _Float_FToF(8, 4, arg);
@@ -373,7 +376,10 @@ void OpDispatchBuilder::FMULF64(OpcodeArgs) {
// Memory arg
if constexpr (Integer) {
arg = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags, -1);
b = _Float_FromGPR_S(8, 8, arg);
if(width == 16) {
arg = _Sext(16, arg);
}
b = _Float_FromGPR_S(8, width == 64 ? 8 : 4, arg);
} else if constexpr (width == 32) {
arg = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
b = _Float_FToF(8, 4, arg);
@@ -434,7 +440,10 @@ void OpDispatchBuilder::FDIVF64(OpcodeArgs) {
// Memory arg
if constexpr (Integer) {
arg = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags, -1);
b = _Float_FromGPR_S(8, 8, arg);
if(width == 16) {
arg = _Sext(16, arg);
}
b = _Float_FromGPR_S(8, width == 64 ? 8 : 4, arg);
} else if constexpr (width == 32) {
arg = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
b = _Float_FToF(8, 4, arg);
@@ -517,7 +526,10 @@ void OpDispatchBuilder::FSUBF64(OpcodeArgs) {
// Memory arg
if constexpr (Integer) {
arg = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags, -1);
b = _Float_FromGPR_S(8, 8, arg);
if(width == 16) {
arg = _Sext(16, arg);
}
b = _Float_FromGPR_S(8, width == 64 ? 8 : 4, arg);
} else if constexpr (width == 32) {
arg = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
b = _Float_FToF(8, 4, arg);
@@ -676,7 +688,10 @@ void OpDispatchBuilder::FCOMIF64(OpcodeArgs) {
// Memory arg
if constexpr (Integer) {
arg = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags, -1);
b = _Float_FromGPR_S(8, 8, arg);
if(width == 16) {
arg = _Sext(16, arg);
}
b = _Float_FromGPR_S(8, width == 64 ? 8 : 4, arg);
} else if constexpr (width == 32) {
arg = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
b = _Float_FToF(8, 4, arg);
@@ -700,7 +715,7 @@ void OpDispatchBuilder::FCOMIF64(OpcodeArgs) {
OrderedNode *HostFlag_CF = _GetHostFlag(Res, FCMP_FLAG_LT);
OrderedNode *HostFlag_ZF = _GetHostFlag(Res, FCMP_FLAG_EQ);
OrderedNode *HostFlag_Unordered = _GetHostFlag(Res, FCMP_FLAG_UNORDERED);
HostFlag_CF = _Or(HostFlag_CF, HostFlag_Unordered);
HostFlag_ZF = _Or(HostFlag_ZF, HostFlag_Unordered);
@@ -810,8 +825,8 @@ void OpDispatchBuilder::X87BinaryOpF64(OpcodeArgs) {
// Overwrite the op
result.first->Header.Op = IROp;
if constexpr (IROp == IR::OP_F80FPREM ||
IROp == IR::OP_F80FPREM1) {
if constexpr (IROp == IR::OP_F64FPREM ||
IROp == IR::OP_F64FPREM1) {
//TODO: Set C0 to Q2, C3 to Q1, C1 to Q0
SetRFLAG<FEXCore::X86State::X87FLAG_C2_LOC>(_Constant(0));
}
+29 -6
View File
@@ -23,15 +23,38 @@ X86GeneratedCode::X86GeneratedCode() {
// Allocate a page for our emulated guest
CodePtr = AllocateGuestCodeSpace(CODE_SIZE);
SignalReturn = reinterpret_cast<uint64_t>(CodePtr);
CallbackReturn = reinterpret_cast<uint64_t>(CodePtr) + 2;
const std::vector<uint8_t> SignalReturnCode = {
0x0F, 0x36, // SIGRET FEX instruction
constexpr std::array<uint8_t, 2> SignalReturnCode = {
0x0F, 0x37, // CALLBACKRET FEX Instruction
};
memcpy(CodePtr, &SignalReturnCode.at(0), SignalReturnCode.size());
// Signal return handlers need to be bit-exact to what the Linux kernel provides in VDSO.
// GDB and unwinding libraries key off of these instructions to understand if the stack frame is a signal frame or not.
// This two code sections match exactly what libSegFault expects.
//
// Typically this handlers are provided by the 32-bit VDSO thunk library, but that isn't available in all cases.
// Falling back to this generated code segment still allows a backtrace to work, just might not show
// the symbol as VDSO since there is no ELF to parse.
constexpr std::array<uint8_t, 9> sigreturn_32_code = {
0x58, // pop eax
0xb8, 0x77, 0x00, 0x00, 0x00, // mov eax, 0x77
0xcd, 0x80, // int 0x80
0x90, // nop
};
constexpr std::array<uint8_t, 7> rt_sigreturn_32_code = {
0xb8, 0xad, 0x00, 0x00, 0x00, // mov eax, 0xad
0xcd, 0x80, // int 0x80
};
CallbackReturn = reinterpret_cast<uint64_t>(CodePtr);
sigreturn_32 = CallbackReturn + SignalReturnCode.size();
rt_sigreturn_32 = sigreturn_32 + sigreturn_32_code.size();
memcpy(reinterpret_cast<void*>(CallbackReturn), &SignalReturnCode.at(0), SignalReturnCode.size());
memcpy(reinterpret_cast<void*>(sigreturn_32), &sigreturn_32_code.at(0), sigreturn_32_code.size());
memcpy(reinterpret_cast<void*>(rt_sigreturn_32), &rt_sigreturn_32_code.at(0), rt_sigreturn_32_code.size());
mprotect(CodePtr, CODE_SIZE, PROT_READ);
}
X86GeneratedCode::~X86GeneratedCode() {
+2 -1
View File
@@ -15,8 +15,9 @@ public:
X86GeneratedCode();
~X86GeneratedCode();
uint64_t SignalReturn{};
uint64_t CallbackReturn{};
uint64_t sigreturn_32{};
uint64_t rt_sigreturn_32{};
private:
void *CodePtr{};
@@ -24,8 +24,8 @@ void InitializeH0F3ATables(Context::OperatingMode Mode) {
{OPD(0, PF_3A_NONE, 0x0F), 1, X86InstInfo{"PALIGNR", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 1, nullptr}},
{OPD(0, PF_3A_66, 0x08), 1, X86InstInfo{"ROUNDPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x09), 1, X86InstInfo{"ROUNDPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x0A), 1, X86InstInfo{"ROUNDSS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x0B), 1, X86InstInfo{"ROUNDSD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x0A), 1, X86InstInfo{"ROUNDSS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x0B), 1, X86InstInfo{"ROUNDSD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x0C), 1, X86InstInfo{"BLENDPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x0D), 1, X86InstInfo{"BLENDPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x0E), 1, X86InstInfo{"PBLENDW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
@@ -338,7 +338,7 @@ void InitializeSecondaryGroupTables() {
{OPD(TYPE_GROUP_15, PF_NONE, 3), 1, X86InstInfo{"STMXCSR", TYPE_INST, GenFlagsSameSize(SIZE_32BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_NONE, 4), 1, X86InstInfo{"XSAVE", TYPE_PRIV, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_NONE, 5), 1, X86InstInfo{"LFENCE/XRSTOR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_NONE, 6), 1, X86InstInfo{"MFENCE/XSAVEOPT", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_NONE, 6), 1, X86InstInfo{"MFENCE/XSAVEOPT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_NONE, 7), 1, X86InstInfo{"SFENCE/CLFLUSH", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_F3, 0), 1, X86InstInfo{"RDFSBASE", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY, 0, nullptr}},
@@ -356,8 +356,8 @@ void InitializeSecondaryGroupTables() {
{OPD(TYPE_GROUP_15, PF_66, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_66, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_66, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_66, 6), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_66, 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_66, 6), 1, X86InstInfo{"CLWB", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_66, 7), 1, X86InstInfo{"CLFLUSHOPT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_F2, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_F2, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
@@ -42,7 +42,7 @@ void InitializeSecondaryTables(Context::OperatingMode Mode) {
{0x14, 1, X86InstInfo{"UNPCKLPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x15, 1, X86InstInfo{"UNPCKHPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x16, 1, X86InstInfo{"MOVLHPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x17, 1, X86InstInfo{"MOVHPS", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_SF_HIGH_XMM_REG | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x17, 1, X86InstInfo{"MOVHPS", TYPE_INST, GenFlagsSizes(SIZE_64BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x18, 1, X86InstInfo{"", TYPE_GROUP_16, FLAGS_NO_OVERLAY, 0, nullptr}},
{0x19, 7, X86InstInfo{"NOP", TYPE_INST, FLAGS_DEBUG | FLAGS_MODRM | FLAGS_NO_OVERLAY, 0, nullptr}},
@@ -64,6 +64,7 @@ void InitializeSecondaryTables(Context::OperatingMode Mode) {
{0x33, 1, X86InstInfo{"RDPMC", TYPE_PRIV, FLAGS_NO_OVERLAY, 0, nullptr}},
{0x34, 1, X86InstInfo{"SYSENTER", TYPE_PRIV, FLAGS_NO_OVERLAY, 0, nullptr}},
{0x35, 1, X86InstInfo{"SYSEXIT", TYPE_PRIV, FLAGS_NO_OVERLAY, 0, nullptr}},
{0x36, 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NO_OVERLAY, 0, nullptr}},
{0x38, 1, X86InstInfo{"", TYPE_0F38_TABLE, FLAGS_NO_OVERLAY, 0, nullptr}},
{0x39, 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NO_OVERLAY, 0, nullptr}},
{0x3A, 1, X86InstInfo{"", TYPE_0F3A_TABLE, FLAGS_NO_OVERLAY, 0, nullptr}},
@@ -257,8 +258,6 @@ void InitializeSecondaryTables(Context::OperatingMode Mode) {
// FEX reserved instructions
// Unused x86 encoding instruction.
// Used by FEX to know when to do a signal return
{0x36, 1, X86InstInfo{"SIGRET", TYPE_INST, FLAGS_BLOCK_END | FLAGS_NO_OVERLAY | FLAGS_SETS_RIP, 0, nullptr}},
{0x37, 1, X86InstInfo{"CALLBACKRET", TYPE_INST, FLAGS_BLOCK_END | FLAGS_NO_OVERLAY | FLAGS_SETS_RIP, 0, nullptr}},
@@ -353,7 +352,7 @@ void InitializeSecondaryTables(Context::OperatingMode Mode) {
{0xD8, 8, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0xE0, 6, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0xE6, 1, X86InstInfo{"CVTDQ2PD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0xE6, 1, X86InstInfo{"CVTDQ2PD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0xE7, 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0xE8, 8, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
@@ -375,7 +374,7 @@ void InitializeSecondaryTables(Context::OperatingMode Mode) {
{0x24, 6, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0x2A, 1, X86InstInfo{"CVTSI2SD", TYPE_INST, GenFlagsDstSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_SRC_GPR, 0, nullptr}},
{0x2B, 1, X86InstInfo{"MOVNTSD", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x2C, 1, X86InstInfo{"CVTTSD2SI", TYPE_INST, GenFlagsSrcSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_DST_GPR, 0, nullptr}},
{0x2C, 1, X86InstInfo{"CVTTSD2SI", TYPE_INST, GenFlagsSrcSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_DST_GPR, 0, nullptr}},
{0x2D, 1, X86InstInfo{"CVTSD2SI", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_DST_GPR, 0, nullptr}},
{0x2E, 2, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
+222 -225
View File
@@ -19,13 +19,13 @@ void InitializeVEXTables() {
// VEX Map 1
{OPD(1, 0b00, 0x10), 1, X86InstInfo{"VMOVUPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x10), 1, X86InstInfo{"VMOVUPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x10), 1, X86InstInfo{"VMOVSS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b11, 0x10), 1, X86InstInfo{"VMOVSD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b10, 0x10), 1, X86InstInfo{"VMOVSS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b11, 0x10), 1, X86InstInfo{"VMOVSD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b00, 0x11), 1, X86InstInfo{"VMOVUPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x11), 1, X86InstInfo{"VMOVUPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x11), 1, X86InstInfo{"VMOVSS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b11, 0x11), 1, X86InstInfo{"VMOVSD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b10, 0x11), 1, X86InstInfo{"VMOVSS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b11, 0x11), 1, X86InstInfo{"VMOVSD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b00, 0x12), 1, X86InstInfo{"VMOVLPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_XMM_FLAGS | FLAGS_VEX_1ST_SRC, 0, nullptr}},
{OPD(1, 0b01, 0x12), 1, X86InstInfo{"VMOVLPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_XMM_FLAGS | FLAGS_VEX_1ST_SRC, 0, nullptr}},
@@ -35,11 +35,11 @@ void InitializeVEXTables() {
{OPD(1, 0b00, 0x13), 1, X86InstInfo{"VMOVLPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x13), 1, X86InstInfo{"VMOVLPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b00, 0x14), 1, X86InstInfo{"VUNPCKLPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x14), 1, X86InstInfo{"VUNPCKLPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x14), 1, X86InstInfo{"VUNPCKLPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x14), 1, X86InstInfo{"VUNPCKLPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b00, 0x15), 1, X86InstInfo{"VUNPCKHPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x15), 1, X86InstInfo{"VUNPCKHPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x15), 1, X86InstInfo{"VUNPCKHPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x15), 1, X86InstInfo{"VUNPCKHPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b00, 0x16), 1, X86InstInfo{"VMOVHPS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_XMM_FLAGS | FLAGS_VEX_1ST_SRC, 0, nullptr}},
{OPD(1, 0b01, 0x16), 1, X86InstInfo{"VMOVHPD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_XMM_FLAGS | FLAGS_VEX_1ST_SRC, 0, nullptr}},
@@ -48,19 +48,19 @@ void InitializeVEXTables() {
{OPD(1, 0b00, 0x17), 1, X86InstInfo{"VMOVHPS", TYPE_INST, GenFlagsSizes(SIZE_64BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x17), 1, X86InstInfo{"VMOVHPD", TYPE_INST, GenFlagsSizes(SIZE_64BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b00, 0x50), 1, X86InstInfo{"VMOVMSKPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x50), 1, X86InstInfo{"VMOVMSKPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x50), 1, X86InstInfo{"VMOVMSKPS", TYPE_INST, GenFlagsSizes(SIZE_32BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_REG_ONLY | FLAGS_XMM_FLAGS | FLAGS_SF_DST_GPR, 0, nullptr}},
{OPD(1, 0b01, 0x50), 1, X86InstInfo{"VMOVMSKPD", TYPE_INST, GenFlagsSizes(SIZE_32BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_REG_ONLY | FLAGS_XMM_FLAGS | FLAGS_SF_DST_GPR, 0, nullptr}},
{OPD(1, 0b00, 0x51), 1, X86InstInfo{"VSQRTPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x51), 1, X86InstInfo{"VSQRTPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b10, 0x51), 1, X86InstInfo{"VSQRTSS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b11, 0x51), 1, X86InstInfo{"VSQRTSD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x51), 1, X86InstInfo{"VSQRTPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x51), 1, X86InstInfo{"VSQRTPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x51), 1, X86InstInfo{"VSQRTSS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b11, 0x51), 1, X86InstInfo{"VSQRTSD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b00, 0x52), 1, X86InstInfo{"VRSQRTPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b10, 0x52), 1, X86InstInfo{"VRSQRTSS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x52), 1, X86InstInfo{"VRSQRTPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x52), 1, X86InstInfo{"VRSQRTSS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b00, 0x53), 1, X86InstInfo{"VRCPPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b10, 0x53), 1, X86InstInfo{"VRCPSS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x53), 1, X86InstInfo{"VRCPPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x53), 1, X86InstInfo{"VRCPSS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b00, 0x54), 1, X86InstInfo{"VANDPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x54), 1, X86InstInfo{"VANDPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
@@ -74,39 +74,39 @@ void InitializeVEXTables() {
{OPD(1, 0b00, 0x57), 1, X86InstInfo{"VXORPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x57), 1, X86InstInfo{"VXORPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x60), 1, X86InstInfo{"VPUNPCKLBW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x61), 1, X86InstInfo{"VPUNPCKLWD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x62), 1, X86InstInfo{"VPUNPCKLDQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x63), 1, X86InstInfo{"VPACKSSWB", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x64), 1, X86InstInfo{"VPCMPGTB", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x65), 1, X86InstInfo{"VPVMPGTW", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x66), 1, X86InstInfo{"VPVMPGTD", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x67), 1, X86InstInfo{"VPACKUSWB", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x60), 1, X86InstInfo{"VPUNPCKLBW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x61), 1, X86InstInfo{"VPUNPCKLWD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x62), 1, X86InstInfo{"VPUNPCKLDQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x63), 1, X86InstInfo{"VPACKSSWB", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x64), 1, X86InstInfo{"VPCMPGTB", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x65), 1, X86InstInfo{"VPCMPGTW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x66), 1, X86InstInfo{"VPCMPGTD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x67), 1, X86InstInfo{"VPACKUSWB", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x70), 1, X86InstInfo{"VPSHUFD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b10, 0x70), 1, X86InstInfo{"VPSHUFHW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b11, 0x70), 1, X86InstInfo{"VPSHUFLW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x70), 1, X86InstInfo{"VPSHUFD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(1, 0b10, 0x70), 1, X86InstInfo{"VPSHUFHW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(1, 0b11, 0x70), 1, X86InstInfo{"VPSHUFLW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(1, 0b01, 0x71), 1, X86InstInfo{"", TYPE_VEX_GROUP_12, FLAGS_NONE, 0, nullptr}}, // VEX Group 12
{OPD(1, 0b01, 0x72), 1, X86InstInfo{"", TYPE_VEX_GROUP_13, FLAGS_NONE, 0, nullptr}}, // VEX Group 13
{OPD(1, 0b01, 0x73), 1, X86InstInfo{"", TYPE_VEX_GROUP_14, FLAGS_NONE, 0, nullptr}}, // VEX Group 14
{OPD(1, 0b01, 0x74), 1, X86InstInfo{"VPCMPEQB", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x75), 1, X86InstInfo{"VPCMPEQW", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x76), 1, X86InstInfo{"VPCMPEQD", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x74), 1, X86InstInfo{"VPCMPEQB", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x75), 1, X86InstInfo{"VPCMPEQW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x76), 1, X86InstInfo{"VPCMPEQD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b00, 0x77), 1, X86InstInfo{"VZERO*", TYPE_INST, GenFlagsDstSize(SIZE_128BIT), 0, nullptr}},
{OPD(1, 0b00, 0xC2), 1, X86InstInfo{"VCMPccPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xC2), 1, X86InstInfo{"VCMPccPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b10, 0xC2), 1, X86InstInfo{"VCMPccSS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b11, 0xC2), 1, X86InstInfo{"VCMPccSD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0xC2), 1, X86InstInfo{"VCMPccPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(1, 0b01, 0xC2), 1, X86InstInfo{"VCMPccPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(1, 0b10, 0xC2), 1, X86InstInfo{"VCMPccSS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(1, 0b11, 0xC2), 1, X86InstInfo{"VCMPccSD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(1, 0b01, 0xC4), 1, X86InstInfo{"VPINSRW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xC5), 1, X86InstInfo{"VPEXTRW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xC5), 1, X86InstInfo{"VPEXTRW", TYPE_INST, GenFlagsSizes(SIZE_32BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_REG_ONLY | FLAGS_SF_DST_GPR | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(1, 0b00, 0xC6), 1, X86InstInfo{"VSHUFPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xC6), 1, X86InstInfo{"VSHUFPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0xC6), 1, X86InstInfo{"VSHUFPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(1, 0b01, 0xC6), 1, X86InstInfo{"VSHUFPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
// The above ops are defined from `Table A-17. VEX Opcode Map 1, Low Nibble = [0h:7h]` of AMD Architecture programmer's manual Volume 3
// This table doesn't state which VEX.pp is for which instruction
@@ -124,69 +124,68 @@ void InitializeVEXTables() {
{OPD(1, 0b00, 0x2B), 1, X86InstInfo{"VMOVNTPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x2B), 1, X86InstInfo{"VMOVNTPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x2C), 1, X86InstInfo{"VCVTTSS2SI", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b11, 0x2C), 1, X86InstInfo{"VCVTTSD2SI", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b10, 0x2C), 1, X86InstInfo{"VCVTTSS2SI", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_DST_GPR, 0, nullptr}},
{OPD(1, 0b11, 0x2C), 1, X86InstInfo{"VCVTTSD2SI", TYPE_INST, GenFlagsSrcSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_DST_GPR, 0, nullptr}},
{OPD(1, 0b10, 0x2D), 1, X86InstInfo{"VCVTSS2SI", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b11, 0x2D), 1, X86InstInfo{"VCVTSD2SI", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b10, 0x2D), 1, X86InstInfo{"VCVTSS2SI", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_DST_GPR, 0, nullptr}},
{OPD(1, 0b11, 0x2D), 1, X86InstInfo{"VCVTSD2SI", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_DST_GPR, 0, nullptr}},
{OPD(1, 0b00, 0x2E), 1, X86InstInfo{"VUCOMISS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x2E), 1, X86InstInfo{"VUCOMISD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x2E), 1, X86InstInfo{"VUCOMISS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x2E), 1, X86InstInfo{"VUCOMISD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b00, 0x2F), 1, X86InstInfo{"VUCOMISS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x2F), 1, X86InstInfo{"VUCOMISD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x2F), 1, X86InstInfo{"VCOMISS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x2F), 1, X86InstInfo{"VCOMISD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b00, 0x58), 1, X86InstInfo{"VADDPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x58), 1, X86InstInfo{"VADDPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x58), 1, X86InstInfo{"VADDSS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b11, 0x58), 1, X86InstInfo{"VADDSD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b10, 0x58), 1, X86InstInfo{"VADDSS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b11, 0x58), 1, X86InstInfo{"VADDSD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b00, 0x59), 1, X86InstInfo{"VMULPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x59), 1, X86InstInfo{"VMULPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b10, 0x59), 1, X86InstInfo{"VMULSS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b11, 0x59), 1, X86InstInfo{"VMULSD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x59), 1, X86InstInfo{"VMULPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x59), 1, X86InstInfo{"VMULPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x59), 1, X86InstInfo{"VMULSS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b11, 0x59), 1, X86InstInfo{"VMULSD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b00, 0x5B), 1, X86InstInfo{"VCVTDQ2PS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x5B), 1, X86InstInfo{"VCVTPS2DQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b10, 0x5B), 1, X86InstInfo{"VCVTPS2DQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x5B), 1, X86InstInfo{"VCVTDQ2PS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x5B), 1, X86InstInfo{"VCVTPS2DQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x5B), 1, X86InstInfo{"VCVTTPS2DQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b00, 0x5C), 1, X86InstInfo{"VSUBPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x5C), 1, X86InstInfo{"VSUBPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b10, 0x5C), 1, X86InstInfo{"VSUBSS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b11, 0x5C), 1, X86InstInfo{"VSUBSD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x5C), 1, X86InstInfo{"VSUBPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x5C), 1, X86InstInfo{"VSUBPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x5C), 1, X86InstInfo{"VSUBSS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b11, 0x5C), 1, X86InstInfo{"VSUBSD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b00, 0x5D), 1, X86InstInfo{"VMINPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x5D), 1, X86InstInfo{"VMINPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b10, 0x5D), 1, X86InstInfo{"VMINSS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b11, 0x5D), 1, X86InstInfo{"VMINSD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x5D), 1, X86InstInfo{"VMINPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x5D), 1, X86InstInfo{"VMINPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x5D), 1, X86InstInfo{"VMINSS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b11, 0x5D), 1, X86InstInfo{"VMINSD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b00, 0x5E), 1, X86InstInfo{"VDIVPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x5E), 1, X86InstInfo{"VDIVPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b10, 0x5E), 1, X86InstInfo{"VDIVSS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b11, 0x5E), 1, X86InstInfo{"VDIVSD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x5E), 1, X86InstInfo{"VDIVPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x5E), 1, X86InstInfo{"VDIVPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x5E), 1, X86InstInfo{"VDIVSS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b11, 0x5E), 1, X86InstInfo{"VDIVSD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b00, 0x5F), 1, X86InstInfo{"VMAXPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x5F), 1, X86InstInfo{"VMAXPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b10, 0x5F), 1, X86InstInfo{"VMAXSS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b11, 0x5F), 1, X86InstInfo{"VMAXSD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x5F), 1, X86InstInfo{"VMAXPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x5F), 1, X86InstInfo{"VMAXPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x5F), 1, X86InstInfo{"VMAXSS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b11, 0x5F), 1, X86InstInfo{"VMAXSD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x68), 1, X86InstInfo{"VPUNPCKHBW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x69), 1, X86InstInfo{"VPUNPCKHWD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x6A), 1, X86InstInfo{"VPUNPCKHDQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x6B), 1, X86InstInfo{"VPACKSSDW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x6C), 1, X86InstInfo{"VPUNPCKLQDQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x6D), 1, X86InstInfo{"VPUNPCKHQDQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x68), 1, X86InstInfo{"VPUNPCKHBW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x69), 1, X86InstInfo{"VPUNPCKHWD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x6A), 1, X86InstInfo{"VPUNPCKHDQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x6B), 1, X86InstInfo{"VPACKSSDW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x6C), 1, X86InstInfo{"VPUNPCKLQDQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x6D), 1, X86InstInfo{"VPUNPCKHQDQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x6E), 1, X86InstInfo{"VMOV*", TYPE_INST, GenFlagsDstSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_SRC_GPR, 0, nullptr}},
{OPD(1, 0b01, 0x6F), 1, X86InstInfo{"VMOVDQA", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x6F), 1, X86InstInfo{"VMOVDQU", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x7C), 1, X86InstInfo{"VHADDPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b11, 0x7C), 1, X86InstInfo{"VHADDPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x7C), 1, X86InstInfo{"VHADDPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b11, 0x7C), 1, X86InstInfo{"VHADDPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x7D), 1, X86InstInfo{"VHSUBPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b11, 0x7D), 1, X86InstInfo{"VHSUBPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x7D), 1, X86InstInfo{"VHSUBPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b11, 0x7D), 1, X86InstInfo{"VHSUBPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x7E), 1, X86InstInfo{"VMOV*", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_DST_GPR | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x7E), 1, X86InstInfo{"VMOVQ", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
@@ -199,142 +198,142 @@ void InitializeVEXTables() {
{OPD(1, 0b10, 0xAE), 1, X86InstInfo{"", TYPE_VEX_GROUP_15, FLAGS_NONE, 0, nullptr}}, // VEX Group 15
{OPD(1, 0b11, 0xAE), 1, X86InstInfo{"", TYPE_VEX_GROUP_15, FLAGS_NONE, 0, nullptr}}, // VEX Group 15
{OPD(1, 0b01, 0xD0), 1, X86InstInfo{"VADDSUBPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b11, 0xD0), 1, X86InstInfo{"VADDSUBPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xD0), 1, X86InstInfo{"VADDSUBPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b11, 0xD0), 1, X86InstInfo{"VADDSUBPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xD1), 1, X86InstInfo{"VPSRLW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xD2), 1, X86InstInfo{"VPSRLD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xD3), 1, X86InstInfo{"VPSRLQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xD1), 1, X86InstInfo{"VPSRLW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xD2), 1, X86InstInfo{"VPSRLD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xD3), 1, X86InstInfo{"VPSRLQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xD4), 1, X86InstInfo{"VPADDQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xD5), 1, X86InstInfo{"VPMULLW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xD5), 1, X86InstInfo{"VPMULLW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xD6), 1, X86InstInfo{"VMOVQ", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xD7), 1, X86InstInfo{"VPMOVMSKB", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_DST_GPR | FLAGS_SF_MOD_REG_ONLY, 0, nullptr}},
{OPD(1, 0b01, 0xD7), 1, X86InstInfo{"VPMOVMSKB", TYPE_UNDEC, FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_DST_GPR | FLAGS_SF_MOD_REG_ONLY, 0, nullptr}},
{OPD(1, 0b01, 0xD8), 1, X86InstInfo{"VPSUBUSB", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xD9), 1, X86InstInfo{"VPSUBUSW", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xDA), 1, X86InstInfo{"VPMINUB", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xD8), 1, X86InstInfo{"VPSUBUSB", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xD9), 1, X86InstInfo{"VPSUBUSW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xDA), 1, X86InstInfo{"VPMINUB", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xDB), 1, X86InstInfo{"VPAND", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xDC), 1, X86InstInfo{"VPADDUSB", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xDD), 1, X86InstInfo{"VPADDUSW", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xDE), 1, X86InstInfo{"VPMAXUB", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xDC), 1, X86InstInfo{"VPADDUSB", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xDD), 1, X86InstInfo{"VPADDUSW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xDE), 1, X86InstInfo{"VPMAXUB", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xDF), 1, X86InstInfo{"VPANDN", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xE0), 1, X86InstInfo{"VPAVGB", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xE1), 1, X86InstInfo{"VPSRAW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xE2), 1, X86InstInfo{"VPSRAD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xE3), 1, X86InstInfo{"VPAVGW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xE4), 1, X86InstInfo{"VPMULHUW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xE5), 1, X86InstInfo{"VPMULHW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xE0), 1, X86InstInfo{"VPAVGB", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xE1), 1, X86InstInfo{"VPSRAW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xE2), 1, X86InstInfo{"VPSRAD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xE3), 1, X86InstInfo{"VPAVGW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xE4), 1, X86InstInfo{"VPMULHUW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xE5), 1, X86InstInfo{"VPMULHW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xE6), 1, X86InstInfo{"VCVTTPD2DQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b10, 0xE6), 1, X86InstInfo{"VCVTDQ2PD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b11, 0xE6), 1, X86InstInfo{"VCVTPD2DQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xE6), 1, X86InstInfo{"VCVTTPD2DQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0xE6), 1, X86InstInfo{"VCVTDQ2PD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b11, 0xE6), 1, X86InstInfo{"VCVTPD2DQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xE7), 1, X86InstInfo{"VMOVNTDQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xE8), 1, X86InstInfo{"VPSUBSB", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xE9), 1, X86InstInfo{"VPSUBSW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xEA), 1, X86InstInfo{"VPMINSW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xE8), 1, X86InstInfo{"VPSUBSB", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xE9), 1, X86InstInfo{"VPSUBSW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xEA), 1, X86InstInfo{"VPMINSW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xEB), 1, X86InstInfo{"VPOR", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xEC), 1, X86InstInfo{"VPADDSB", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xED), 1, X86InstInfo{"VPADDSW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xEE), 1, X86InstInfo{"VPMAXSW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xEC), 1, X86InstInfo{"VPADDSB", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xED), 1, X86InstInfo{"VPADDSW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xEE), 1, X86InstInfo{"VPMAXSW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xEF), 1, X86InstInfo{"VPXOR", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b11, 0xF0), 1, X86InstInfo{"VLDDQU", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b11, 0xF0), 1, X86InstInfo{"VLDDQU", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xF1), 1, X86InstInfo{"VPSLLW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xF2), 1, X86InstInfo{"VPSLLD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xF3), 1, X86InstInfo{"VPSLLQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xF4), 1, X86InstInfo{"VPMULUDQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xF5), 1, X86InstInfo{"VPMADDWD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xF1), 1, X86InstInfo{"VPSLLW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xF2), 1, X86InstInfo{"VPSLLD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xF3), 1, X86InstInfo{"VPSLLQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xF4), 1, X86InstInfo{"VPMULUDQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xF5), 1, X86InstInfo{"VPMADDWD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xF6), 1, X86InstInfo{"VPSADBW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xF7), 1, X86InstInfo{"VMASKMOVDQU", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xF7), 1, X86InstInfo{"VMASKMOVDQU", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_REG_ONLY | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xF8), 1, X86InstInfo{"VPSUBB", TYPE_INST, GenFlagsSameSize(SIZE_256BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xF9), 1, X86InstInfo{"VPSUBW", TYPE_INST, GenFlagsSameSize(SIZE_256BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xFA), 1, X86InstInfo{"VPSUBD", TYPE_INST, GenFlagsSameSize(SIZE_256BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xFB), 1, X86InstInfo{"VPSUBQ", TYPE_INST, GenFlagsSameSize(SIZE_256BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xF8), 1, X86InstInfo{"VPSUBB", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xF9), 1, X86InstInfo{"VPSUBW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xFA), 1, X86InstInfo{"VPSUBD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xFB), 1, X86InstInfo{"VPSUBQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xFC), 1, X86InstInfo{"VPADDB", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xFD), 1, X86InstInfo{"VPADDW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xFE), 1, X86InstInfo{"VPADDD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
// VEX Map 2
{OPD(2, 0b01, 0x00), 1, X86InstInfo{"VPSHUFB", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x01), 1, X86InstInfo{"VPADDW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x02), 1, X86InstInfo{"VPHADDD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x03), 1, X86InstInfo{"VPHADDSW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x00), 1, X86InstInfo{"VPSHUFB", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x01), 1, X86InstInfo{"VPHADDW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x02), 1, X86InstInfo{"VPHADDD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x03), 1, X86InstInfo{"VPHADDSW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x04), 1, X86InstInfo{"VPMADDUBSW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x05), 1, X86InstInfo{"VPHSUBW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x06), 1, X86InstInfo{"VPHSUBD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x07), 1, X86InstInfo{"VPHSUBSW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x05), 1, X86InstInfo{"VPHSUBW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x06), 1, X86InstInfo{"VPHSUBD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x07), 1, X86InstInfo{"VPHSUBSW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x08), 1, X86InstInfo{"VPSIGNB", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x09), 1, X86InstInfo{"VPSIGNW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x0A), 1, X86InstInfo{"VPSIGND", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x0B), 1, X86InstInfo{"VPMULHRSW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x0C), 1, X86InstInfo{"VPERMILPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x0D), 1, X86InstInfo{"VPERMILPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x08), 1, X86InstInfo{"VPSIGNB", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x09), 1, X86InstInfo{"VPSIGNW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x0A), 1, X86InstInfo{"VPSIGND", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x0B), 1, X86InstInfo{"VPMULHRSW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x0C), 1, X86InstInfo{"VPERMILPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x0D), 1, X86InstInfo{"VPERMILPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x0E), 1, X86InstInfo{"VTESTPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x0F), 1, X86InstInfo{"VTESTPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x13), 1, X86InstInfo{"VCVTPH2PS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x16), 1, X86InstInfo{"VPERMPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x17), 1, X86InstInfo{"VPTEST", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x16), 1, X86InstInfo{"VPERMPS", TYPE_INST, GenFlagsSameSize(SIZE_256BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x17), 1, X86InstInfo{"VPTEST", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x18), 1, X86InstInfo{"VBROADCASTSS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x19), 1, X86InstInfo{"VBROADCASTSD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x1A), 1, X86InstInfo{"VBROADCASTF128", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x1C), 1, X86InstInfo{"VPABSB", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x1D), 1, X86InstInfo{"VPABSW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x1E), 1, X86InstInfo{"VPABSD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x18), 1, X86InstInfo{"VBROADCASTSS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x19), 1, X86InstInfo{"VBROADCASTSD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x1A), 1, X86InstInfo{"VBROADCASTF128", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x1C), 1, X86InstInfo{"VPABSB", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x1D), 1, X86InstInfo{"VPABSW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x1E), 1, X86InstInfo{"VPABSD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x20), 1, X86InstInfo{"VPMOVSXBW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x21), 1, X86InstInfo{"VPMOVSXBD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x22), 1, X86InstInfo{"VPMOVSXBQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x23), 1, X86InstInfo{"VPMOVSXWD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x24), 1, X86InstInfo{"VPMOVSXWQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x25), 1, X86InstInfo{"VPMOVSXDQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x20), 1, X86InstInfo{"VPMOVSXBW", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x21), 1, X86InstInfo{"VPMOVSXBD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x22), 1, X86InstInfo{"VPMOVSXBQ", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_16BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x23), 1, X86InstInfo{"VPMOVSXWD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x24), 1, X86InstInfo{"VPMOVSXWQ", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x25), 1, X86InstInfo{"VPMOVSXDQ", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x28), 1, X86InstInfo{"VPMULDQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x29), 1, X86InstInfo{"VPCMPEQQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x28), 1, X86InstInfo{"VPMULDQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x29), 1, X86InstInfo{"VPCMPEQQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x2A), 1, X86InstInfo{"VMOVNTDQA", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x2B), 1, X86InstInfo{"VPACKUSDW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x2B), 1, X86InstInfo{"VPACKUSDW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x2C), 1, X86InstInfo{"VMASKMOVPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x2D), 1, X86InstInfo{"VMASKMOVPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x2E), 1, X86InstInfo{"VMASKMOVPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x2F), 1, X86InstInfo{"VMASKMOVPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x30), 1, X86InstInfo{"VPMOVZXBW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x31), 1, X86InstInfo{"VPMOVZXBD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x32), 1, X86InstInfo{"VPMOVZXBQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x33), 1, X86InstInfo{"VPMOVZXWD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x34), 1, X86InstInfo{"VPMOVZXWQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x35), 1, X86InstInfo{"VPMOVZXDQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x36), 1, X86InstInfo{"VPERMD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x37), 1, X86InstInfo{"VPVMPGTQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x30), 1, X86InstInfo{"VPMOVZXBW", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x31), 1, X86InstInfo{"VPMOVZXBD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x32), 1, X86InstInfo{"VPMOVZXBQ", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_16BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x33), 1, X86InstInfo{"VPMOVZXWD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x34), 1, X86InstInfo{"VPMOVZXWQ", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x35), 1, X86InstInfo{"VPMOVZXDQ", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x36), 1, X86InstInfo{"VPERMD", TYPE_INST, GenFlagsSameSize(SIZE_256BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x37), 1, X86InstInfo{"VPCMPGTQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x38), 1, X86InstInfo{"VPMINSB", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x39), 1, X86InstInfo{"VPMINSD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x3A), 1, X86InstInfo{"VPMINUW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x3B), 1, X86InstInfo{"VPMINUD", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x3C), 1, X86InstInfo{"VPMAXSB", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x3D), 1, X86InstInfo{"VPMAXSD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x3E), 1, X86InstInfo{"VPMAXUW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x3F), 1, X86InstInfo{"VPMAXUD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x38), 1, X86InstInfo{"VPMINSB", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x39), 1, X86InstInfo{"VPMINSD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x3A), 1, X86InstInfo{"VPMINUW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x3B), 1, X86InstInfo{"VPMINUD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x3C), 1, X86InstInfo{"VPMAXSB", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x3D), 1, X86InstInfo{"VPMAXSD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x3E), 1, X86InstInfo{"VPMAXUW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x3F), 1, X86InstInfo{"VPMAXUD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x40), 1, X86InstInfo{"VPMULLD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x41), 1, X86InstInfo{"VPHMINPOSUW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x40), 1, X86InstInfo{"VPMULLD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x41), 1, X86InstInfo{"VPHMINPOSUW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x45), 1, X86InstInfo{"VPSRLV", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x46), 1, X86InstInfo{"VPSRAVD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x46), 1, X86InstInfo{"VPSRAVD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x47), 1, X86InstInfo{"VPSLLV", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x58), 1, X86InstInfo{"VPBROADCASTD", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{OPD(2, 0b01, 0x59), 1, X86InstInfo{"VPBROADCASTQ", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{OPD(2, 0b01, 0x5A), 1, X86InstInfo{"VBBROADCASTI128", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{OPD(2, 0b01, 0x58), 1, X86InstInfo{"VPBROADCASTD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x59), 1, X86InstInfo{"VPBROADCASTQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x5A), 1, X86InstInfo{"VBROADCASTI128", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x78), 1, X86InstInfo{"VPBROADCASTB", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{OPD(2, 0b01, 0x79), 1, X86InstInfo{"VPBROADCASTW", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{OPD(2, 0b01, 0x78), 1, X86InstInfo{"VPBROADCASTB", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x79), 1, X86InstInfo{"VPBROADCASTW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x8C), 1, X86InstInfo{"VPMASKMOV", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x8E), 1, X86InstInfo{"VPMASKMOV", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
@@ -380,11 +379,11 @@ void InitializeVEXTables() {
{OPD(2, 0b01, 0xB6), 1, X86InstInfo{"VFMADDSUB231", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0xB7), 1, X86InstInfo{"VFMSUBADD231", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0xDB), 1, X86InstInfo{"VAESIMC", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0xDC), 1, X86InstInfo{"VAESENC", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0xDD), 1, X86InstInfo{"VAESENCLAST", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0xDE), 1, X86InstInfo{"VAESDEC", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0xDF), 1, X86InstInfo{"VAESDECLAST", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0xDB), 1, X86InstInfo{"VAESIMC", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0xDC), 1, X86InstInfo{"VAESENC", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0xDD), 1, X86InstInfo{"VAESENCLAST", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0xDE), 1, X86InstInfo{"VAESDEC", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0xDF), 1, X86InstInfo{"VAESDECLAST", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b00, 0xF2), 1, X86InstInfo{"ANDN", TYPE_INST, FLAGS_MODRM | FLAGS_VEX_1ST_SRC, 0, nullptr}},
@@ -406,49 +405,47 @@ void InitializeVEXTables() {
{OPD(2, 0b11, 0xF7), 1, X86InstInfo{"SHRX", TYPE_INST, FLAGS_MODRM | FLAGS_VEX_2ND_SRC, 0, nullptr}},
// VEX Map 3
{OPD(3, 0b01, 0x00), 1, X86InstInfo{"VPERMQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(3, 0b01, 0x01), 1, X86InstInfo{"VPERMPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(3, 0b01, 0x02), 1, X86InstInfo{"VPBLENDD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(3, 0b01, 0x04), 1, X86InstInfo{"VPERMILPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(3, 0b01, 0x05), 1, X86InstInfo{"VPERMILPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(3, 0b01, 0x06), 1, X86InstInfo{"VPERM2F128", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(3, 0b01, 0x00), 1, X86InstInfo{"VPERMQ", TYPE_INST, GenFlagsSameSize(SIZE_256BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x01), 1, X86InstInfo{"VPERMPD", TYPE_INST, GenFlagsSameSize(SIZE_256BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x02), 1, X86InstInfo{"VPBLENDD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x04), 1, X86InstInfo{"VPERMILPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x05), 1, X86InstInfo{"VPERMILPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x06), 1, X86InstInfo{"VPERM2F128", TYPE_INST, GenFlagsSameSize(SIZE_256BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x08), 1, X86InstInfo{"VROUNDPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(3, 0b01, 0x09), 1, X86InstInfo{"VROUNDPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(3, 0b01, 0x0A), 1, X86InstInfo{"VROUNDSS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(3, 0b01, 0x0B), 1, X86InstInfo{"VROUNDSD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(3, 0b01, 0x0C), 1, X86InstInfo{"VBLENDPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(3, 0b01, 0x0D), 1, X86InstInfo{"VBLENDPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(3, 0b01, 0x0E), 1, X86InstInfo{"VBLENDW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(3, 0b01, 0x0F), 1, X86InstInfo{"VPALIGNR", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x08), 1, X86InstInfo{"VROUNDPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x09), 1, X86InstInfo{"VROUNDPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x0A), 1, X86InstInfo{"VROUNDSS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x0B), 1, X86InstInfo{"VROUNDSD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x0C), 1, X86InstInfo{"VBLENDPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x0D), 1, X86InstInfo{"VBLENDPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x0E), 1, X86InstInfo{"VPBLENDW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x0F), 1, X86InstInfo{"VPALIGNR", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x14), 1, X86InstInfo{"VPEXTRB", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(3, 0b01, 0x15), 1, X86InstInfo{"VPEXTRW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(3, 0b01, 0x16), 1, X86InstInfo{"VPEXTRD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(3, 0b01, 0x17), 1, X86InstInfo{"VEXTRACTPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(3, 0b01, 0x14), 1, X86InstInfo{"VPEXTRB", TYPE_INST, GenFlagsSizes(SIZE_32BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_DST_GPR | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x15), 1, X86InstInfo{"VPEXTRW", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_DST_GPR | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x16), 1, X86InstInfo{"VPEXTRD", TYPE_INST, GenFlagsSizes(SIZE_32BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_DST_GPR | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x17), 1, X86InstInfo{"VEXTRACTPS", TYPE_INST, GenFlagsSizes(SIZE_32BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_DST_GPR | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x18), 1, X86InstInfo{"VINSERTF128", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(3, 0b01, 0x19), 1, X86InstInfo{"VEXTRACTF128", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(3, 0b01, 0x18), 1, X86InstInfo{"VINSERTF128", TYPE_INST, GenFlagsSameSize(SIZE_256BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x19), 1, X86InstInfo{"VEXTRACTF128", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_256BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x1D), 1, X86InstInfo{"VCVTPS2PH", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(3, 0b01, 0x20), 1, X86InstInfo{"VPINSRB", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(3, 0b01, 0x21), 1, X86InstInfo{"VINSERTPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(3, 0b01, 0x21), 1, X86InstInfo{"VINSERTPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x22), 1, X86InstInfo{"VPINSRD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(3, 0b01, 0x38), 1, X86InstInfo{"VINSERTI128", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(3, 0b01, 0x39), 1, X86InstInfo{"VEXTRACTI128", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(3, 0b01, 0x38), 1, X86InstInfo{"VINSERTI128", TYPE_INST, GenFlagsSameSize(SIZE_256BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x39), 1, X86InstInfo{"VEXTRACTI128", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_256BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x40), 1, X86InstInfo{"VDPPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(3, 0b01, 0x41), 1, X86InstInfo{"VDPPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(3, 0b01, 0x40), 1, X86InstInfo{"VDPPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x41), 1, X86InstInfo{"VDPPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x42), 1, X86InstInfo{"VMPSADBW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(3, 0b01, 0x44), 1, X86InstInfo{"VPCLMULQDQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x46), 1, X86InstInfo{"VPERM2I128", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(3, 0b01, 0x46), 1, X86InstInfo{"VPERM2I128", TYPE_INST, GenFlagsSameSize(SIZE_256BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x48), 1, X86InstInfo{"VPERMILzz2PS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(3, 0b01, 0x49), 1, X86InstInfo{"VPERMILzz2PD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(3, 0b01, 0x4A), 1, X86InstInfo{"VBLENDVPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(3, 0b01, 0x4B), 1, X86InstInfo{"VBLENDVPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(3, 0b01, 0x4C), 1, X86InstInfo{"VBLENDVB", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(3, 0b01, 0x4A), 1, X86InstInfo{"VBLENDVPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x4B), 1, X86InstInfo{"VBLENDVPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x4C), 1, X86InstInfo{"VPBLENDVB", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x5C), 1, X86InstInfo{"VFMADDSUBPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(3, 0b01, 0x5D), 1, X86InstInfo{"VFMADDSUBPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
@@ -478,7 +475,7 @@ void InitializeVEXTables() {
{OPD(3, 0b01, 0x7E), 1, X86InstInfo{"VFNMSUBSS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(3, 0b01, 0x7F), 1, X86InstInfo{"VFNMSUBSD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(3, 0b01, 0xDF), 1, X86InstInfo{"VAESKEYGENASSIST", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(3, 0b01, 0xDF), 1, X86InstInfo{"VAESKEYGENASSIST", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b11, 0xF0), 1, X86InstInfo{"RORX", TYPE_INST, FLAGS_MODRM, 1, nullptr}},
@@ -488,21 +485,21 @@ void InitializeVEXTables() {
#define OPD(group, pp, opcode) (((group - TYPE_VEX_GROUP_12) << 4) | (pp << 3) | (opcode))
static constexpr U8U8InfoStruct VEXGroupTable[] = {
{OPD(TYPE_VEX_GROUP_12, 1, 0b010), 1, X86InstInfo{"VPSRLW", TYPE_UNDEC, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_VEX_GROUP_12, 1, 0b100), 1, X86InstInfo{"VPSRAW", TYPE_UNDEC, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_VEX_GROUP_12, 1, 0b110), 1, X86InstInfo{"VPSLLW", TYPE_UNDEC, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_VEX_GROUP_12, 1, 0b010), 1, X86InstInfo{"VPSRLW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_DST | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(TYPE_VEX_GROUP_12, 1, 0b100), 1, X86InstInfo{"VPSRAW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_DST | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(TYPE_VEX_GROUP_12, 1, 0b110), 1, X86InstInfo{"VPSLLW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_DST | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(TYPE_VEX_GROUP_13, 1, 0b010), 1, X86InstInfo{"VPSRLD", TYPE_UNDEC, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_VEX_GROUP_13, 1, 0b100), 1, X86InstInfo{"VPSRAD", TYPE_UNDEC, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_VEX_GROUP_13, 1, 0b110), 1, X86InstInfo{"VPSLLD", TYPE_UNDEC, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_VEX_GROUP_13, 1, 0b010), 1, X86InstInfo{"VPSRLD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_DST | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(TYPE_VEX_GROUP_13, 1, 0b100), 1, X86InstInfo{"VPSRAD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_DST | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(TYPE_VEX_GROUP_13, 1, 0b110), 1, X86InstInfo{"VPSLLD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_DST | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(TYPE_VEX_GROUP_14, 1, 0b010), 1, X86InstInfo{"VPSRLQ", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(TYPE_VEX_GROUP_14, 1, 0b011), 1, X86InstInfo{"VPSRLDQ", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(TYPE_VEX_GROUP_14, 1, 0b110), 1, X86InstInfo{"VPSLLQ", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(TYPE_VEX_GROUP_14, 1, 0b111), 1, X86InstInfo{"VPSLLDQ", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(TYPE_VEX_GROUP_14, 1, 0b010), 1, X86InstInfo{"VPSRLQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_DST | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(TYPE_VEX_GROUP_14, 1, 0b011), 1, X86InstInfo{"VPSRLDQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_DST | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(TYPE_VEX_GROUP_14, 1, 0b110), 1, X86InstInfo{"VPSLLQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_DST | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(TYPE_VEX_GROUP_14, 1, 0b111), 1, X86InstInfo{"VPSLLDQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_DST | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(TYPE_VEX_GROUP_15, 1, 0b010), 1, X86InstInfo{"VLDMXCSR", TYPE_UNDEC, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_VEX_GROUP_15, 1, 0b011), 1, X86InstInfo{"VSTMXCSR", TYPE_UNDEC, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_VEX_GROUP_15, 0, 0b010), 1, X86InstInfo{"VLDMXCSR", TYPE_INST, GenFlagsSameSize(SIZE_32BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_VEX_GROUP_15, 0, 0b011), 1, X86InstInfo{"VSTMXCSR", TYPE_INST, GenFlagsSameSize(SIZE_32BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_VEX_GROUP_17, 0, 0b001), 1, X86InstInfo{"BLSR", TYPE_INST, FLAGS_MODRM | FLAGS_VEX_DST, 0, nullptr}},
{OPD(TYPE_VEX_GROUP_17, 0, 0b010), 1, X86InstInfo{"BLSMSK", TYPE_INST, FLAGS_MODRM | FLAGS_VEX_DST, 0, nullptr}},
+16 -6
View File
@@ -167,6 +167,8 @@ namespace FEXCore {
* address to another function. The original callee address is passed
* to the target function through an implicit argument stored in r11.
*
* For 32-bit the implicit argument is stored in the lower 32-bits of mm0.
*
* The primary use case of this is ensuring that host function pointers
* returned from thunked APIs can safely be called by the guest.
*/
@@ -177,7 +179,7 @@ namespace FEXCore {
};
auto args = reinterpret_cast<args_t*>(argsv);
auto CTX = Thread->CTX;
auto CTX = static_cast<Context::ContextImpl*>(Thread->CTX);
LOGMAN_THROW_AA_FMT(args->original_callee, "Tried to link null pointer address to guest function");
LOGMAN_THROW_AA_FMT(args->target_addr, "Tried to link address to null pointer guest function");
@@ -199,7 +201,12 @@ namespace FEXCore {
const uint8_t GPRSize = CTX->GetGPRSize();
emit->_StoreRegister(emit->_Constant(Entrypoint), false, offsetof(Core::CPUState, gregs[X86State::REG_R11]), IR::GPRClass, IR::GPRFixedClass, GPRSize);
if (GPRSize == 8) {
emit->_StoreRegister(emit->_Constant(Entrypoint), false, offsetof(Core::CPUState, gregs[X86State::REG_R11]), IR::GPRClass, IR::GPRFixedClass, GPRSize);
}
else {
emit->_StoreRegister(emit->_Constant(Entrypoint), false, offsetof(Core::CPUState, mm[0][0]), IR::GPRClass, IR::GPRFixedClass, GPRSize);
}
emit->_ExitFunction(emit->_Constant(GuestThunkEntrypoint));
}, CTX->ThunkHandler.get(), (void*)args->target_addr);
@@ -257,13 +264,16 @@ namespace FEXCore {
}
static void LoadLib(void *ArgsV) {
auto CTX = Thread->CTX;
auto CTX = static_cast<Context::ContextImpl*>(Thread->CTX);
auto Args = reinterpret_cast<LoadlibArgs*>(ArgsV);
auto Name = Args->Name;
auto SOName = CTX->Config.ThunkHostLibsPath() + "/" + (const char*)Name + "-host.so";
auto SOName = (CTX->Config.Is64BitMode() ?
CTX->Config.ThunkHostLibsPath() :
CTX->Config.ThunkHostLibsPath32())
+ "/" + (const char*)Name + "-host.so";
LogMan::Msg::DFmt("LoadLib: {} -> {}", Name, SOName);
@@ -311,7 +321,7 @@ namespace FEXCore {
auto &[Name, rv] = *reinterpret_cast<ArgsRV_t*>(ArgsRV);
auto CTX = Thread->CTX;
auto CTX = static_cast<Context::ContextImpl*>(Thread->CTX);
auto That = reinterpret_cast<ThunkHandler_impl*>(CTX->ThunkHandler.get());
{
@@ -375,7 +385,7 @@ namespace FEXCore {
HostToGuestTrampolinePtr* MakeHostTrampolineForGuestFunction(void* HostPacker, uintptr_t GuestTarget, uintptr_t GuestUnpacker) {
LOGMAN_THROW_AA_FMT(GuestTarget, "Tried to create host-trampoline to null pointer guest function");
const auto CTX = Thread->CTX;
const auto CTX = static_cast<Context::ContextImpl*>(Thread->CTX);
const auto ThunkHandler = reinterpret_cast<ThunkHandler_impl *>(CTX->ThunkHandler.get());
const GuestcallInfo gci = { GuestUnpacker, GuestTarget };
+1 -1
View File
@@ -11,7 +11,7 @@ $end_info$
#include <vector>
namespace FEXCore::Context {
struct Context;
class ContextImpl;
}
namespace FEXCore::Core {
+5 -2
View File
@@ -17,6 +17,9 @@
namespace FEXCore::Core {
struct DebugData;
}
namespace FEXCore::Context {
class ContextImpl;
}
namespace FEXCore::IR {
class RegisterAllocationData;
@@ -87,7 +90,7 @@ namespace FEXCore::IR {
class AOTIRCaptureCache final {
public:
AOTIRCaptureCache(FEXCore::Context::Context *ctx) : CTX {ctx} {}
AOTIRCaptureCache(FEXCore::Context::ContextImpl *ctx) : CTX {ctx} {}
void FinalizeAOTIRCache();
void AOTIRCaptureCacheWriteoutQueue_Flush();
@@ -131,7 +134,7 @@ namespace FEXCore::IR {
}
private:
FEXCore::Context::Context *CTX;
FEXCore::Context::ContextImpl *CTX;
std::shared_mutex AOTIRCacheLock;
std::shared_mutex AOTIRCaptureCacheWriteoutLock;
+45 -18
View File
@@ -22,7 +22,7 @@
"",
"Eg:",
"IR op with no result and no arguments",
" SignalReturn",
" CallbackReturn",
"",
"IR op with result and no arguments",
" GPR = ProcessorID",
@@ -264,9 +264,6 @@
"Break BreakDefinition:$Reason": {
"HasSideEffects": true
},
"SignalReturn": {
"HasSideEffects": true
},
"CallbackReturn": {
"HasSideEffects": true
},
@@ -479,9 +476,25 @@
]
},
"CacheLineClear GPR:$Addr": {
"GPR = MemSet i1:$IsAtomic, u8:$Size, GPR:$Prefix, GPR:$Addr, GPR:$Value, GPR:$Length, GPR:$Direction": {
"Desc": ["Duplicates behaviour of x86 STOS repeat",
"Returns the final address that gets generated without the prefix appended."
],
"HasSideEffects": true,
"DestSize": "8"
},
"CacheLineClear GPR:$Addr, i1:$Serialize": {
"Desc": ["Does a 64 byte cacheline clear at the address specified",
"Only clears the data cachelines. Doesn't do any zeroing"
"Only clears the data cachelines. Doesn't do any zeroing",
"Can skip serialization if requested."
],
"HasSideEffects": true
},
"CacheLineClean GPR:$Addr": {
"Desc": ["Does a 64 byte cacheline cleanat the address specified",
"Only cleans the data cachelines. Doesn't do any zeroing",
"Skips the invalidation step of the CacheLineClear operation"
],
"HasSideEffects": true
},
@@ -1183,6 +1196,14 @@
"DestSize": "RegisterSize",
"NumElements": "RegisterSize / ElementSize"
},
"FPR = VTrn u8:#RegisterSize, u8:#ElementSize, FPR:$VectorLower, FPR:$VectorUpper": {
"DestSize": "RegisterSize",
"NumElements": "RegisterSize / ElementSize"
},
"FPR = VTrn2 u8:#RegisterSize, u8:#ElementSize, FPR:$VectorLower, FPR:$VectorUpper": {
"DestSize": "RegisterSize",
"NumElements": "RegisterSize / ElementSize"
},
"FPR = VFAdd u8:#RegisterSize, u8:#ElementSize, FPR:$Vector1, FPR:$Vector2": {
"DestSize": "RegisterSize",
@@ -1345,12 +1366,12 @@
"DestSize": "RegisterSize"
},
"FPR = VBSL FPR:$VectorMask, FPR:$VectorTrue, FPR:$VectorFalse": {
"FPR = VBSL u8:#RegisterSize, FPR:$VectorMask, FPR:$VectorTrue, FPR:$VectorFalse": {
"Desc": ["Does a vector bitwise select.",
"If the bit in the field is 1 then the corresponding bit is pulled from VectorTrue",
"If the bit in the field is 0 then the corresponding bit is pulled from VectorFalse"
],
"DestSize": "16"
"DestSize": "RegisterSize"
}
},
"Conv": {
@@ -1362,6 +1383,12 @@
"NumElements": "RegisterSize / ElementSize"
},
"FPR = VDupFromGPR u8:#RegisterSize, u8:#ElementSize, GPR:$Src": {
"Desc": ["Broadcasts a value in a GPR into each ElementSize-sized element in a vector"],
"DestSize": "RegisterSize",
"NumElements": "RegisterSize / ElementSize"
},
"FPR = Float_FromGPR_S u8:#DstElementSize, u8:$SrcElementSize, GPR:$Src": {
"Desc": ["Scalar op: Converts signed GPR to Scalar float",
"Zeroes the upper bits of the vector register"
@@ -1410,21 +1437,21 @@
"Desc": "Does a stage of the inverse mix column transformation",
"DestSize": "16"
},
"FPR = VAESEnc FPR:$State, FPR:$Key": {
"FPR = VAESEnc u8:#RegisterSize, FPR:$State, FPR:$Key": {
"Desc": "Does a step of AES encryption",
"DestSize": "16"
"DestSize": "RegisterSize"
},
"FPR = VAESEncLast FPR:$State, FPR:$Key": {
"FPR = VAESEncLast u8:#RegisterSize, FPR:$State, FPR:$Key": {
"Desc": "Does the last step of AES encryption",
"DestSize": "16"
"DestSize": "RegisterSize"
},
"FPR = VAESDec FPR:$State, FPR:$Key": {
"FPR = VAESDec u8:#RegisterSize, FPR:$State, FPR:$Key": {
"Desc": "Does a step of AES decryption",
"DestSize": "16"
"DestSize": "RegisterSize"
},
"FPR = VAESDecLast FPR:$State, FPR:$Key": {
"FPR = VAESDecLast u8:#RegisterSize, FPR:$State, FPR:$Key": {
"Desc": "Does the last step of AES decryption",
"DestSize": "16"
"DestSize": "RegisterSize"
},
"FPR = VAESKeyGenAssist FPR:$Src, u8:$RCON": {
"Desc": "Assists in key generation",
@@ -1435,7 +1462,7 @@
],
"DestSize": "std::max<uint8_t>(4, GetOpSize(_Src1))"
},
"FPR = PCLMUL FPR:$Src1, FPR:$Src2, u8:$Selector": {
"FPR = PCLMUL u8:#RegisterSize, FPR:$Src1, FPR:$Src2, u8:$Selector": {
"Desc": [
"Performs carryless multiplication of 64-bit elements depending on the selector.",
"Selector = 0b00000000: Uses low 64-bit elements from both input vectors",
@@ -1443,7 +1470,7 @@
"Selector = 0b00010000: Uses low 64-bit element from Src1 and high 64-bit element from Src2",
"Selector = 0b00010001: Uses high 64-bit elements from both input vectors"
],
"DestSize": "16"
"DestSize": "RegisterSize"
}
},
"F64": {
Loaded 100 of 498 files, more files were not shown because too many files have changed in this diff. Show more