Compare commits

...
454 Commits
Author SHA1 Message Date
Ryan Houdek cae4f2f873 Docs: Update for release FEX-2206 2022-06-04 12:55:42 -07:00
Stefanos Kornilios Mitsis Poiitidis 0fc6d6b6b5 Merge pull request #1749 from Sonicadvance1/atomic_tests
unittests: Reenable atomic tests on ARMv8.0
2022-06-04 14:02:25 +03:00
Stefanos Kornilios Mitsis Poiitidis 7227ee9b2e Merge pull request #1748 from Sonicadvance1/gvisor_investigations
unittests: Investigate failing CI changes
2022-06-04 13:49:48 +03:00
Stefanos Kornilios Mitsis Poiitidis c6153d6a52 Merge pull request #1747 from Sonicadvance1/struct_verifier_fixes
Struct verifier fixes and reenable
2022-06-04 13:46:30 +03:00
Ryan Houdek d82d2944a9 Merge pull request #1745 from FEX-Emu/skmp/mtrack-fixes
mtrack: Fixes 32-bit shmat, shmdt tracking, guaranteed invalidation atomicity
2022-06-04 00:36:58 -07:00
Ryan Houdek 58ad400519 unittests: Reenable atomic tests on ARMv8.0 2022-06-04 00:30:02 -07:00
Ryan Houdek d23c76d0a9 Arm64: Work with more unaligned atomic operations
The latest ARMv8.0 toolchain is implementing fetch_add with a bic rather
than an and. Not sure why they started doing this but support the
remaining logical operations in our unaligned atomics handler.

Fixes the Interpreter ARMv8.0 atomic ops.

Fixes #1742
2022-06-04 00:30:02 -07:00
Ryan Houdek 6a5b9e2a93 unittests: Investigate failing CI changes
One gvisor test didn't expect a file header to change layout.
Another one was testing behaviour that was removed from upstream Linux

Fixes #1741
2022-06-03 23:03:59 -07:00
Ryan Houdek f33a93b0a1 github: Reenable struct verifier tests 2022-06-03 19:15:32 -07:00
Ryan Houdek 015200f511 StructVerifier: Reenable DRM testing 2022-06-03 19:12:33 -07:00
Ryan Houdek 92f48819b6 IoctlEmulation: Fix DRM includes
These were being overridden by system includes.
2022-06-03 19:12:33 -07:00
Ryan Houdek 0c6483cad5 External: Update drm-headers 2022-06-03 18:50:33 -07:00
Stefanos Kornilios Misis Poiitidis ee02b1ca51 Review feedback 2022-06-03 14:15:22 +03:00
Stefanos Kornilios Misis Poiitidis 33845a3112 Mtrack: Remove race conditions around concurrent invalidation and compilation 2022-06-03 12:41:38 +03:00
Stefanos Kornilios Misis Poiitidis 096ed29b5e Mtrack/x86: Track shmat, shmdt via ipc syscall as well 2022-06-03 12:41:32 +03:00
Ryan Houdek 3bbff8a948 Merge pull request #1744 from lioncash/sha256
OpcodeDispatcher: Implement SHA256 instructions
2022-06-02 13:53:52 -07:00
lioncash 726918b82c CPUID: Enable SHA extension bit
Now that all SHA instructions have an implementation, we can enable the
CPUID bit for it.
2022-06-02 15:59:35 -04:00
lioncash 0f59a18223 unittests: Disable SHA256 tests
Currently our x86 CI doesn't have SHA instruction extensions.
2022-06-02 15:58:17 -04:00
lioncash 3402cde334 OpcodeDispatcher: Implement SHA256RNDS2 2022-06-02 15:56:22 -04:00
lioncash 8f53c6bb96 OpcodeDispatcher: Implement SHA256MSG2 2022-06-02 15:12:46 -04:00
lioncash 0d6e4631a3 OpcodeDispatcher: Implement SHA256MSG1 2022-06-02 15:03:33 -04:00
Ryan Houdek 8dd9a5bd38 Merge pull request #1739 from lioncash/sha1
OpcodeDispatcher: Handle SHA-1 instructions
2022-06-02 11:28:54 -07:00
lioncash 903cf84874 unittests: Disable SHA-1 tests for now
Currently the x86 CI doesn't support the SHA instruction extension set
2022-06-02 14:02:30 -04:00
lioncash e997da48c7 OpcodeDispatcher: Implement SHA1RNDS4 2022-06-02 13:37:42 -04:00
lioncash 5ff89fd171 OpcodeDispatcher: Implement SHA1MSG2 2022-06-02 13:37:42 -04:00
lioncash fad4254c0e OpcodeDispatcher: Implement SHA1MSG1 2022-06-02 13:37:42 -04:00
lioncash 2fb3c4f11c OpcodeDispatcher: Implement SHA1NEXTE 2022-06-02 13:37:42 -04:00
Stefanos Kornilios Mitsis Poiitidis ce5297b75f Merge pull request #1738 from Sonicadvance1/workaround_tests
unittests: Workaround runner issues
2022-06-02 15:48:51 +03:00
Ryan Houdek b2b0c277f6 github: Disable struct verifier
Needs to be validated again. Xavier is really hating it.
2022-06-02 05:00:10 -07:00
Ryan Houdek 46919979ce StructVerifier: Ensure drm include is in place
drm testing disabled while investigations occur
2022-06-02 04:46:00 -07:00
Ryan Houdek 8b716c6a22 unittests ASM: Disable failing ARMv8 tests 2022-06-02 04:40:13 -07:00
Ryan Houdek 75090f8f6c GVisor: Disable failing unit tests 2022-06-02 04:40:13 -07:00
Stefanos Kornilios Mitsis Poiitidis c14c0c2e3b Merge pull request #1736 from Sonicadvance1/argument_injector
AppConfig: Inject --no-sandbox in to steamwebhelper
2022-06-01 10:19:07 +03:00
Ryan Houdek 95efd18b73 AppConfig: Inject --no-sandbox in to steamwebhelper
Steam's webhelper has started enabling its sandbox which completely
breaks under FEX since we don't support seccomp.

Curiously the Chromium code is actually supposed to support a fallback
namespace only mode, which is used in glibc 2.34 environments.

The startup script for this will try to use this namespace only mode,
but Chromium developers never tested this in an environment that doesn't
support seccomp.

Due to an early check in their sandbox code, it checks for bpf support
before checking for which sandbox mode it is entering. This returns
early with a false statement which brings the entire browser instance
down with an assert.

Inject the --no-sandbox argument so we get around this and the sandbox
is disabled.
2022-05-31 17:12:08 -07:00
Ryan Houdek c9319a768f Config: Add the ability to inject command line arguments
Simple enough since we control the full emulation
2022-05-31 17:11:42 -07:00
Mai M da48020882 Merge pull request #1730 from Sonicadvance1/support_pause
OpcodeDispatcher: Implements support for PAUSE
2022-05-26 19:20:04 -04:00
Ryan Houdek ce4380e136 unittests: Adds basic PAUSE test 2022-05-26 08:51:00 -07:00
Ryan Houdek c633661121 OpcodeDispatcher: Implements support for PAUSE
The pause instruction is architecturally defined to be the `REP NOP`
instruction.

This allows people to use this instruction as a backwards compatible
pause without checking for CPUID support. In fact there is no way to
check if the hardware implements this as a `REP NOP` or a `PAUSE`.

If you have new enough hardware then this just ends up being a PAUSE.

Pass this PAUSE over to our host to help out applications that are
writing spin loops with a PAUSE in it, which we were deleting
previously.
2022-05-26 08:47:56 -07:00
Ryan Houdek 4bfd1dde1f IR: Implements support for Yield IR op
This is an IR op that produces nor consumes any SSA values, but has side
effects.

Turns in to the pause instruction on x86 and yield instruction on
AArch64.
2022-05-26 08:46:14 -07:00
Mai M fe11bd2242 Merge pull request #1726 from Sonicadvance1/fix_pextrb
OpcodeDispatcher: Fixes pextrb with high registers
2022-05-24 23:05:51 -04:00
Ryan Houdek 814f0c3c93 unittests: Adds unit test for the high pextrb encoding
Previous unit tests didn't cover this edge case.
2022-05-24 16:31:09 -07:00
Ryan Houdek 7e904056d3 OpcodeDispatcher: Fixes pextrb with high registers
Previously if the instruction was encoded to use rsp, rbp, rsi, or rdi
then due to how these were encoded in modrm this would hit the frontend
path for writing to the high 8 bits of a 16bit register.

This is because it's instruction specific if an 8-bit modrm instruction
chooses to use the high 8-bit region or the upper 4 registers.
See the ModRM.reg section of `ModRM.reg and .r/m Field Encodings`
specifically to see what each encoding stands for. Has four different
meanings per encoding depending on instruction.

This instruction doesn't actually write to registers at 8-bit size, it
extracts an element at 8-bit size and then zero extends it to the full
GPR.

When storing to memory it always stores to memory at the size of the
element extracted.

I grepped around the instruction tables to see if there were any other
instances of this mistake. This was the only one.

Fixes #1472 and also gets Psychonauts 2 running.
2022-05-24 16:23:22 -07:00
Mai M 969d8f866c Merge pull request #1724 from Sonicadvance1/v5.18
v5.18 support
2022-05-23 14:43:11 -04:00
Ryan Houdek a2d0b7d7c4 Linux: Expose v5.18 host to guest 2022-05-23 11:00:16 -07:00
Ryan Houdek 97a8fa77bc Ioctl: Update drm msm for v5.18
Fixes #1602
2022-05-23 10:59:29 -07:00
Ryan Houdek c1296cc64d Update external drm-headers 2022-05-23 10:58:50 -07:00
Ryan Houdek 1dee54a9d8 Merge pull request #1722 from Hypnotron/main
Fix dangling curl hyphen
2022-05-21 11:24:16 -07:00
The Hypnotron 17e5d73e64 Fix dangling curl hyphen
Fixes a regression in fa87c73b9ee60a334eace2cdc3097725cbaf5b88; curl complains "curl: option -: is unknown" when trying to fetch a RootFS without this.
2022-05-21 14:12:48 -04:00
Mai M d523b7a6c7 Merge pull request #1720 from Sonicadvance1/workaround_libstdcxx_bug
FEXLogServer: Stop improper use of std::erase_if
2022-05-20 09:02:38 -04:00
Ryan Houdek 24ad208778 FEXLogServer: Stop improper use of std::erase_if
std::erase_if shouldn't allow you to modify the object passed in to the
predicate.
libstdc++ hasn't always enforced this but now it does with libstdc++12
2022-05-20 03:47:24 -07:00
Ryan Houdek fa87c73b9e Merge pull request #1719 from Sonicadvance1/fex_rootfs_fetcher_no_rety
FEXRootFSFetcher: Don't continue download
2022-05-20 03:13:38 -07:00
Ryan Houdek 0ed96544e1 Merge pull request #1721 from Sonicadvance1/fix_clone3_stack
Syscalls: Fixes clone3 stack pointer
2022-05-20 02:46:05 -07:00
Ryan Houdek d28ccc59ac Syscalls: Fixes clone3 stack pointer
clone2 stack pointer passed in points to the highest address for the
stack.

clone3 switches this around and gives us a base pointer and a size.

glibc started using clone3 for its thread cloning which finally caught
this bug. Necessary to run any application under the Ubuntu 22.04 rootfs
since that uses a new enough glibc to encounter this.
2022-05-19 22:51:17 -07:00
Ryan Houdek eaa75c1ed2 FEXRootFSFetcher: Don't continue download
While our CDN supports download continue, the backblaze storage backing
does not.
2022-05-19 22:44:22 -07:00
Ryan Houdek 4f4263263b Merge pull request #1718 from FEX-Emu/skmp/fix-x86tables-leave
X86Tables: Leave shouldn't end block
2022-05-19 00:22:27 -07:00
Stefanos Kornilios Misis Poiitidis ae00654694 X86Tables: Leave shouldn't end block 2022-05-19 08:51:02 +03:00
Stefanos Kornilios Mitsis Poiitidis a7156276e9 Merge pull request #1716 from FEX-Emu/skmp/jitsymbols-file-offsets
JitSymbols: Print file+offset if possible
2022-05-17 16:06:18 +03:00
Stefanos Kornilios Misis Poiitidis 29859d2491 JitSymbols: Print file offsets if possible 2022-05-17 15:05:47 +03:00
Stefanos Kornilios Mitsis Poiitidis 5460a24ea9 Merge pull request #1558 from FEX-Emu/skmp/smc-memtrack
SMC detection via segfaults
2022-05-16 17:37:07 +03:00
Stefanos Kornilios Misis Poiitidis a284adcd19 SMC: Add mprotect based tracking, --smc=mtrack, make default 2022-05-16 15:51:09 +03:00
Stefanos Kornilios Mitsis Poiitidis 73d43c1d55 Merge pull request #1700 from FEX-Emu/skmp/standarized-todo
Standarized TODO markers: FEX_TODO, FEX_TODO_ISSUE
2022-05-16 14:17:16 +03:00
Stefanos Kornilios Misis Poiitidis 256df76674 FEX_TODO: Convert some XXX to FEX_TODO 2022-05-16 12:22:42 +03:00
Ryan Houdek c8dc663b0b Merge pull request #1709 from Sonicadvance1/remove_debug_statement
OpcodeDispatcher: Remove debugging dump statement
2022-05-14 19:31:01 -07:00
Ryan Houdek ba78dff1f8 Merge pull request #1707 from Sonicadvance1/non_temporal
OpcodeDispatcher: Adds support for non-temporal loadstores
2022-05-14 19:30:53 -07:00
Ryan Houdek 1e597bfbed Merge pull request #1706 from Sonicadvance1/ref_count_shared_mutex
FEXCore: Adds refcount_shared_mutex class
2022-05-14 19:30:43 -07:00
Ryan Houdek b78af2fdaf Merge pull request #1684 from Sonicadvance1/testharness_named_regions
TestHarnessRunner: Use guest mapper for test harness files
2022-05-14 19:18:13 -07:00
Ryan Houdek 5379f0a9c7 Merge pull request #1677 from neobrain/refactor_scopedsignalmask
Clean up and document ScopedSignalMask
2022-05-14 19:13:37 -07:00
Ryan Houdek f1f523e525 OpcodeDispatcher: Remove debugging dump statement 2022-05-14 19:09:10 -07:00
Ryan Houdek c79d79e08b TestHarnessRunner: Use guest mapper for test harness files
This will allow it to get picked up for named region handling. Thus
ending up in the code caching for testing.
2022-05-14 19:08:32 -07:00
Ryan Houdek ee2d417d21 Merge pull request #1691 from Sonicadvance1/object_cache_named_region
Object cache named region no-op implementation
2022-05-14 18:20:05 -07:00
Ryan Houdek cebdde599a ObjectCache: Adds no-op named region object loading
This does the setup for handling the named region object loading and
closing using the async interface.

This exercises the async interface while the async thread itself only
does the minimum no-op steps required to fake loading and saving.
2022-05-14 18:09:38 -07:00
Ryan Houdek c3ac72a01e Core: Do named region async code object cache usage 2022-05-14 18:02:43 -07:00
Ryan Houdek 13f3c6e75a Merge pull request #1690 from Sonicadvance1/job_handler
Core: Adds Code Object Cache service
2022-05-14 18:00:40 -07:00
Ryan Houdek b3cd4edb3b Core: Clear relocations after the cache service had a chance to copy them
Can't clear the relocations vector until after the code object cache
service has consumed them.
2022-05-14 17:49:18 -07:00
Ryan Houdek afe10c1666 Core: Adds Code Object Cache service
The no-op interface is hooked up to the point of exercising it in the
most minimal of sense.

If the configuration is set to enable read-only or read/write object
code then it will spin up the async worker thread as well, but it
doesn't do anything yet.
2022-05-14 17:38:20 -07:00
Ryan Houdek d9d30916ba ObjectCache: Adds no-op object cache
Currently unused.

Showcases the main interface in to the service implemented as no-ops
currently.
2022-05-14 17:38:20 -07:00
Ryan Houdek 3f6c1c0e68 JITs: Return pointer to internal relocation vectors
Currently unused.

This will be used by the Code Object Serialization service soon.
2022-05-14 17:36:15 -07:00
Ryan Houdek 5b2cc77109 CPUBackend: Adds RelocateJITObjectCode virtual function
When a backend supports relocations it will override this function.
returning nullptr meaning no relocation done.

Currently unused but will be soon.
2022-05-14 17:36:14 -07:00
Ryan Houdek b824023ec6 InternalThreadState: Adds Object Cache job ref counter mutex
Currently unused but will be used soon.

Removes old `IsCompileService` bool as well.
2022-05-14 17:36:14 -07:00
Ryan Houdek 6ce1be0880 InternalThreadState: Adds Relocations pointer to DebugData
This will be used soon to pass relocation data to the JIT object cache.
2022-05-14 17:36:14 -07:00
Ryan Houdek c5dacab2ee Merge pull request #1688 from Sonicadvance1/jit_relocations
JIT relocation handling support
2022-05-14 17:25:34 -07:00
Ryan Houdek 4d24b85d57 OpcodeDispatcher: Adds support for non-temporal loadstores
x86 has eight instructions that are non-temporal.
Only one of which is a load-NT.

Adds a memory access type classification to our LoadSource/StoreResult
helpers.

This lets us explicitly choose Default, TSO, NonTSO, and Stream.

Stream currently just behaves like NonTSO so at some point in the future
we can add non-temporal loadstores to the IR.

Main thing is to move these NT accesses to non-TSO.
2022-05-14 02:53:44 -07:00
Ryan Houdek e967b447e6 FEXCore: Adds refcount_shared_mutex class
This class is similar to std::shared_mutex except it is safe for the
same thread to increment or decrement the ref counter multiple times.

This can be passed to regular std locks.

This will be required with the code object cache service soon.
2022-05-14 00:51:15 -07:00
Ryan Houdek 9bc631a427 Merge pull request #1705 from Sonicadvance1/fix_fsgsbase
32-bit FSGS instruction fixes.
2022-05-14 00:04:42 -07:00
Ryan Houdek c480ef137d unittests: Only disable fsgs tests on host
Since the x86 CI machine doesn't have a new enough kernel for this.
2022-05-13 02:43:28 -07:00
Ryan Houdek 15629e790e unittests: Adds fsgsbase 32-bit tests
Ensures that the upper 32-bits are zero'd rather than inserted.
2022-05-13 02:42:28 -07:00
Ryan Houdek 3f08d8b691 OpcodeDispatcher: Fixes 32-bit fs/gs write instructions
Documentation claims that these insert the lower 32-bits leaving the
upper bits unaffected.
Hardware testing proves that the upper 32-bits of the base registers are
zero'd.

Additional documentation also concurs that this is the case.
2022-05-13 02:42:28 -07:00
Ryan Houdek 9d9d171aad OpcodeDispatcher: Only expose fsgs instructions in 64-bit
These aren't supported in 32-bit
2022-05-13 02:42:28 -07:00
Ryan Houdek 65218c8285 unittests: Update tests to use canonical addresses 2022-05-13 02:34:02 -07:00
Ryan Houdek 27f2e0b06d Merge pull request #1704 from Sonicadvance1/fix_instruction_rerun
Arm64: Fix LDAPUR/STLUR DMB backpatch
2022-05-13 00:49:21 -07:00
Ryan Houdek ae75983b54 Arm64: Fix LDAPUR/STLUR DMB backpatch
This ended up in the wrong commit. We need to rerun the DMB that we
patched in.
2022-05-12 08:03:18 -07:00
Ryan Houdek f8ba373e18 Merge pull request #1702 from Sonicadvance1/support_rcpc2
Arm64: Adds support for RCPC2 extension
2022-05-12 07:33:02 -07:00
Ryan Houdek 2feae06209 Arm64: Adds support for RCPC2 extension
This allows us to have RCPC loadstore operations with a 9-bit signed
offset.

This gives us a small range of [-256,256) of immediate encoding range on
our TSO loadstore operations.
Updates the inline constant pass in ConstProp to support this range on
TSO IR ops if the host supports RCPC2.

Apple M1 supports this extension, didn't test with Cortex-X2/A710.
2022-05-12 07:13:47 -07:00
Stefanos Kornilios Mitsis Poiitidis 2e0534924a Merge pull request #1699 from FEX-Emu/skmp/add-fwrapv
CMake: C/C++ flags for defined singed overflow warping
2022-05-11 11:10:46 +03:00
Stefanos Kornilios Misis Poiitidis ad1fd7f54b FexHeaderUtils: Add TodoDefines 2022-05-11 11:08:28 +03:00
Tony Wasserka 9aaace51e1 ScopedSignalMask: Add usage guidelines 2022-05-10 17:18:36 +02:00
Tony Wasserka 58841142ee Merge pull request #1693 from neobrain/feature_linker
CMake: Add option to use the mold linker
2022-05-10 16:29:39 +02:00
Stefanos Kornilios Misis Poiitidis a6a816fb38 CMake: C/C++ flags for defined singed overflow warping 2022-05-10 17:14:17 +03:00
Ryan Houdek 70988ccfee Merge pull request #1694 from Sonicadvance1/fix_RCPC
Arm64: Fixes AtomicSwap
2022-05-09 23:30:34 -07:00
Ryan Houdek a8d9caf0c0 x64Jit: Adds relocation handling support
The JIT currently doesn't use this. This is just the handling code
itself.

One line disabled handling Guest RIP move relocations until the JIT
object cache is enabled.
2022-05-09 19:56:33 -07:00
Ryan Houdek bc22186093 Arm64: Adds relocation handling support
The JIT currently doesn't use this. This is just the handling code
itself.

One line disabled handling Guest RIP move relocations until the JIT
Object cache is enabled.
2022-05-09 19:56:33 -07:00
Ryan Houdek 317416b2e0 Context: Adds Cache object code config option 2022-05-09 19:56:32 -07:00
Ryan Houdek b5ae9e4c97 Merge pull request #1686 from Sonicadvance1/add_relocation_definitions
ArchHelpers: Adds relocation struct defines
2022-05-09 19:49:02 -07:00
Ryan Houdek 099737ca05 ArchHelpers: Adds relocation struct defines
Pulled from #1548 with one of the unused relocation types removed.

Unused for now.
2022-05-09 19:30:09 -07:00
Ryan Houdek e90164b519 Arm64: Fixes AtomicSwap
It wasn't using acquire semantics, only release semantics.
This was causing the swap to load data from a stale cacheline, causing
the futex system in glibc to break.

This break only occured if you tried going down the RCPC codepath
because of edge case memory ordering problems.

This then enables the RCPC code path now since it works.
2022-05-09 16:50:43 -07:00
Tony Wasserka 933c1af7e8 CMake: Add option to use the mold linker 2022-05-09 17:00:10 +02:00
Stefanos Kornilios Mitsis Poiitidis b9d878b1f4 Merge pull request #1672 from FEX-Emu/skmp/add-guest-mmap-munmap
Syscalls/Linux: Add guest[Mmap/Munmap]
2022-05-09 12:55:40 +03:00
Parallels 45a9a83c79 Loaders: Use bind_front instead of lambdas to bind GuestM(un)map 2022-05-09 12:39:48 +03:00
Parallels 560cfc757c HarnessHelper: Allocations need to be MAP_FIXED 2022-05-09 11:57:19 +03:00
Stefanos Kornilios Misis Poiitidis e92f51e415 Syscalls/Linux: Add GuestMmap & GuestMunmap, update code to use it 2022-05-09 11:57:08 +03:00
Ryan Houdek 278ca52d97 Merge pull request #1683 from Sonicadvance1/code_cache_config
Config: Adds code cache config option
2022-05-08 18:44:32 -07:00
Ryan Houdek 912dbfe5bd Merge pull request #1689 from Sonicadvance1/AArch64_MoveConstant_ADR
Arm64Emitter: Optimize constants with ADRP and ADR
2022-05-08 18:43:36 -07:00
Ryan Houdek 687f46fc71 Arm64Emitter: Optimize constants with ADRP and ADR
In a large number of cases we are moving pointers within a 4GB region
and some marginal pointers that are within 1MB.

This is only used in the case that MOVZ can't be used.

NOP padding still occurs after these instructions to ensure that if they
are being used with relocations it will still get padded to a full 4
instruction length.

Not all hardware fuses these and LLVM claims that Cortex beyond A72 even
doesn't, but it'll still be faster.
2022-05-08 18:33:43 -07:00
Ryan Houdek 6e9e5b3bd6 Config: Adds code object cache config option
This will be used soon
2022-05-06 10:26:44 -07:00
Stefanos Kornilios Mitsis Poiitidis 4fbc266b18 Merge pull request #1685 from Sonicadvance1/fix_tmp_file_flags
EmulatedFiles: Fixes temporary file flags
2022-05-06 10:38:42 +03:00
Ryan Houdek d1ac406895 EmulatedFiles: Fixes temporary file flags
mode and flags were being combined incorrectly.
2022-05-05 21:50:34 -07:00
Stefanos Kornilios Mitsis Poiitidis ce0f5db6f7 Merge pull request #1671 from FEX-Emu/skmp/refactor-guest-mman-tracking
Syscalls/Linux: Refactor guest mman tracking
2022-05-03 15:37:55 +03:00
Stefanos Kornilios Misis Poiitidis efb42c1ad1 Syscalls/Linux: Refactor guest mman tracking 2022-05-03 15:26:32 +03:00
Stefanos Kornilios Mitsis Poiitidis d8109880f4 Merge pull request #1670 from FEX-Emu/skmp/processwide-code-invalidations
Core: context-wide guest code invalidations
2022-05-03 15:23:10 +03:00
Stefanos Kornilios Misis Poiitidis 09be28a443 LookupCache: Cleanups 2022-05-02 16:59:50 +03:00
Tony Wasserka fb0bb8dd2c ScopedSignalMask: Unify implementation 2022-05-02 11:27:26 +02:00
Ryan Houdek 8e36f5331f Merge pull request #1669 from FEX-Emu/skmp/movable-lock-guards
ScopedSignalMask: Add shared mutex support, move constructors
2022-05-01 16:27:47 -07:00
Stefanos Kornilios Misis Poiitidis a365a70275 Review feedback 2022-05-02 01:51:17 +03:00
Stefanos Kornilios Misis Poiitidis b5a4e5920d Core: Rename FlushCodeRange to InvalidateGuestCodeRange 2022-05-02 01:49:54 +03:00
Stefanos Kornilios Misis Poiitidis 94d2ed85a7 Core: Add support for process-wide code invalidation, rename IR invalidate op to do thread specific invalidation 2022-05-02 01:49:49 +03:00
Ryan Houdek 90f338d7db Merge pull request #1674 from FEX-Emu/skmp/fexloader-fix-aotir-create_directories
FEXLoader: Fix create_directories check for aotir .path file writting
2022-05-01 15:36:28 -07:00
Stefanos Kornilios Misis Poiitidis df78f5d50e FEXLoader: Fix create_directories check for aotir .path file writting 2022-05-02 01:25:02 +03:00
Ryan Houdek b2b4c2bdcf Merge pull request #1673 from FEX-Emu/skmp/shmdt-fixes
Linux/MemAllocator32Bit: Add missing lock to shmdt, fix error returns
2022-05-01 15:22:08 -07:00
Stefanos Kornilios Misis Poiitidis a888da436b Linux/MemAllocator32Bit: Add missing lock to shmdt, fix error returns 2022-05-02 01:00:53 +03:00
Stefanos Kornilios Misis Poiitidis 3cb8ae9a9c ScopedSignalMask: Add shared mutex support, move constructors 2022-04-30 16:00:00 +03:00
Ryan Houdek db3854e391 Merge pull request #1664 from CallumDev/f64-fldcw-impl
F64: Implement FCW using host rounding mode
2022-04-28 15:11:43 -07:00
CallumDev 05b4b095fe F64: Set host RoundingMode for all FCW loads 2022-04-29 05:07:44 +09:30
CallumDev 93926641d9 F64: Implement FLDCW using host rounding mode 2022-04-29 04:39:57 +09:30
Ryan Houdek 89d6752d3d Merge pull request #1662 from CallumDev/f64-int-fixes
F64: Fix FILD and FIST for Size < 8
2022-04-28 08:28:11 -07:00
CallumDev e4f95fec79 F64: Fix FILD and FIST for Size < 8 2022-04-29 00:42:56 +09:30
Ryan Houdek 8a7f39559c Merge pull request #1627 from Sonicadvance1/wip_reclaimable_pool_allocator
FEXCore: Reclaimable thread pool allocator
2022-04-26 10:21:36 -07:00
Ryan Houdek 753d0ede6c FEXCore: Reclaimable thread pool allocator
Creates a pool allocator for OpcodeDispatcher and IRCompaction that
shares memory allocations between threads in a pool and supports
reclaiming stale allocations from participating threads.

A thread will use a heuristic to keep its claimed memory allocation
around if it is allocating a lot of code. If it slows down then it will
start putting the memory allocation back in to the thread pool.

Additionally if the allocation has been "disowned" and gone to sleep
while still retaining the allocation, then another thread can inspect
 these stale allocations and reclaim it from the idling thread. Saving
further memory.

This needs some more work and cleanup but this is an interesting concept
that saves a decent amount of memory even in a basic test.

Causes teeworlds' title screen to go from 754MB to 599MB in my simple
test. 79.4% the memory usage is a good start.
2022-04-26 10:01:56 -07:00
Ryan Houdek da2e44d024 Merge pull request #1658 from wannacu/main
AOTIR: copy RAData and IRList, make sure data is accessible
2022-04-25 19:04:23 -07:00
wannacu 7b379fc3cf AOTIR: copy RAData and IRList, make sure data is accessible 2022-04-26 09:27:18 +08:00
Ryan Houdek ec38d58b37 Merge pull request #1659 from Sonicadvance1/fexbash_ps1
FEXBash: Set PS1 to make it more obvious when running under FEX
2022-04-25 12:30:17 -07:00
Ryan Houdek f6a74a710d FEXBash: Set PS1 to make it more obvious when running under FEX
This requires us to pass in --no-rc to bash since otherwise PS1 gets
overwritten by shell variables and nothing happens.

Which this is fine for the common use case of just wanting to run a
basic bash script under emulation.
2022-04-25 11:59:04 -07:00
Ryan Houdek 3fd136b0da Merge pull request #1657 from Sonicadvance1/fix_32bit_mmap
Linux: Fixes 32-bit mmap
2022-04-24 11:17:09 -07:00
Ryan Houdek 72e82d0304 Linux: Fixes 32-bit mmap
This went unnoticed for so long since most applications are new enough
to use mmap2 instead of mmap.
This was just completely broken.

Fixes #1630
2022-04-24 10:56:56 -07:00
Ryan Houdek 2f7dcb8d93 Merge pull request #1656 from Sonicadvance1/v5.17_support
V5.17 support
2022-04-24 10:55:39 -07:00
Ryan Houdek 128a24d699 Linux: Updates supported guest Linux version to v5.17 2022-04-23 11:59:18 -07:00
Ryan Houdek cd94a8f0ac Linux: Adds support for new v5.17 virtio IOCTL 2022-04-23 11:58:56 -07:00
Ryan Houdek df5e0e5df9 Linux: Adds support for new v5.17 syscall 2022-04-23 11:58:38 -07:00
Ryan Houdek 458bbf4ef7 Linux: Updates syscalls for v5.17 2022-04-23 11:57:37 -07:00
Ryan Houdek 82319c9deb Scripts: Updates generate syscall numbers to support renaming
Instead of manually renaming the three syscalls each time, let the
script do it automatically.
2022-04-23 11:56:33 -07:00
Ryan Houdek 3bc4df7295 Updates drm headers to v5.17 2022-04-23 11:56:06 -07:00
Ryan Houdek 42a6320935 Merge pull request #1585 from CallumDev/x87f64
Emulate reduced-precision X87 with 64-bit host FPU ops
2022-04-21 09:22:03 -07:00
CallumDev 843fe378db Document that X87ReducedPrecision reduces accuracy 2022-04-22 01:38:18 +09:30
CallumDev 3679673d5b Implement FNSAVE and FRSTOR in F64 2022-04-22 01:36:56 +09:30
CallumDev 3a269f04d2 Add remaining possible X87F64 tests. Tweak FPREM 2022-04-22 01:36:56 +09:30
CallumDev 0021723b50 Fix F64 FSCALE 2022-04-22 01:36:56 +09:30
CallumDev 1c4b0272e8 X87F64: Implement FXTRACT using bit ops, add test 2022-04-22 01:36:56 +09:30
CallumDev 9640216124 X87F64: Add working BCD test 2022-04-22 01:36:56 +09:30
CallumDev 0f8f2bf2b4 X87F64 fix integer load, add tests 2022-04-22 01:36:56 +09:30
CallumDev ef899b7b1a F64: Working FABS and FCHS 2022-04-22 01:36:56 +09:30
CallumDev b85abf725d Implement FCOM in 64-bit ops 2022-04-22 01:36:56 +09:30
CallumDev 3f89e46d66 Add X87ReducedPrecision to FEXConfig 2022-04-22 01:36:56 +09:30
CallumDev d02712ddc6 X87F64: Basic implementation of FRNDINT 2022-04-22 01:36:56 +09:30
CallumDev 91f48c63ff Implement F64 ops in Interpreter 2022-04-22 01:36:56 +09:30
CallumDev cd16769e57 F64: Fix Arm64 JIT compile error 2022-04-22 01:36:56 +09:30
CallumDev 0a778e802f Unit Tests for x87F64 2022-04-22 01:36:56 +09:30
CallumDev d4d5f4d1dd Introduce F64 codegen for reduced precision X87 2022-04-22 01:36:44 +09:30
Mai M b1033ed7c6 Merge pull request #1652 from Sonicadvance1/remove_compile_service
CompileService: Removes no longer necessary service thread
2022-04-19 22:34:15 -04:00
Ryan Houdek 253333a4cf CompileService: Removes no longer necessary service thread
Since we are masking signals before compiling code, we no longer will
receive a signal in the middle of compiling code.

This makes the compile service never be invoked so we can just remove
it.

We still have some locations in the syscall handling that isn't signal
safe, but compileservice wouldn't have fixed those anyway.
2022-04-19 18:52:01 -07:00
Ryan Houdek 50595ac3a9 X86Dispatcher: Disable signals when compiling just like on AArch64 2022-04-19 18:29:21 -07:00
Ryan Houdek 37f1e55ed5 Docs: Update for release FEX-2204 2022-04-19 01:19:00 -07:00
Ryan Houdek 8ad14728f6 Merge pull request #1644 from Sonicadvance1/ldiv_minor_opt
JITArm64: Get long divide out of the hot path
2022-04-01 18:23:30 -07:00
Ryan Houdek 6b3cd3d31d Merge pull request #1645 from Sonicadvance1/update_aarch64_fit
Scripts: Updates AArch64 fit for Clang 14
2022-04-01 18:23:13 -07:00
Ryan Houdek fba698cb74 Scripts: Updates AArch64 fit for Clang 14
Clang now supports these latest ARMv9 CPUs
2022-04-01 18:08:02 -07:00
Ryan Houdek 0946b123bb JITArm64: Get long divide out of the hot path
For 128-bit divides, we can very quickly check at runtime if we can
avoid the long divide and just do a 64-bit divide.

For unsigned just check if the top bits are all zero.
For signed just check if the top bits match bit 63 of the lower bits.

Additionally, keep the long divide handlers inside of the dispatcher.
This keeps the majority of the code bloat out of the code block itself,
significantly reducing block size for something doing these divides.
Also a fairly large icache improvement from this.

Hard performance number improvements here are hard to get since it
heavily depends on the application, also only occurs on x86-64.

Seems to have helped FTL and Dead Cells performance quite a bit though.
2022-03-31 09:33:19 -07:00
Ryan Houdek b43937a7a1 Merge pull request #1643 from Sonicadvance1/fix_termux
SignalDelegator: Adds missing include
2022-03-29 21:05:28 -07:00
Ryan Houdek 4564eba20d SignalDelegator: Adds missing include
Fixes Termux building.
Fixes #1642
2022-03-29 20:47:22 -07:00
Ryan Houdek 5cc0c0a3da Merge pull request #1641 from philpax/docs-remove-stale-text
docs: Remove stale text
2022-03-29 02:56:20 -07:00
Philpax f8e7c75f86 docs: Remove stale text 2022-03-29 11:27:43 +02:00
Ryan Houdek 042cd354dc Merge pull request #1633 from Sonicadvance1/disable_instructions_on_host_missing
OpcodeDispatcher: Fixes SIGILL on unsupported host instructions
2022-03-23 13:48:04 -07:00
Ryan Houdek 977bda97b2 Merge pull request #1635 from Sonicadvance1/4000_0001h
CPUID: Adds 4000_0001h function
2022-03-23 13:42:45 -07:00
Ryan Houdek 4cf48ca9bb CPUID: Adds 4000_0001h function
Exposes the host architecture through this CPUID function. Only exposes
the architectures we support. Not burning 16-bits on using ELF machine
definitions here.

Uses 4 bits still for future expansion.
2022-03-22 16:53:44 -07:00
Mai M a247df50ea Merge pull request #1624 from Sonicadvance1/cleanup_ir_after_use
FEXCore: Delete IR after it is used
2022-03-22 13:23:05 -04:00
Mai M 1f1c214944 Merge pull request #1634 from Sonicadvance1/cpuid_documentation
Documentation: Adds hypervisor CPUID information
2022-03-22 12:57:15 -04:00
Ryan Houdek d16db4ebde OpcodeDispatcher: Fixes SIGILL on unsupported host instructions
If the host doesn't support the instructions required for implementing
an instruction then don't even add them to the opcodedispatcher.

This means that we will never try emitting instructions that the host
doesn't support (For these instructions anyway) and successfully passes
the guest SIGILL for these particular instructions.

Fixes #1631
2022-03-21 23:03:04 -07:00
Ryan Houdek ae1c563082 Documentation: Adds hypervisor CPUID information
Currently we only implement function 4000_0000h. This will expand in the
future but this is all we have right now.
2022-03-21 22:46:48 -07:00
Ryan Houdek ebd0edbab7 Merge pull request #1632 from FEX-Emu/skmp/flush-test-harness
TestHarnessRunner: Flush log on asserts
2022-03-21 12:50:43 -07:00
Stefanos Kornilios Misis Poiitidis e87e9d269a TestHarnessRunner: Flush log on asserts 2022-03-21 21:32:35 +02:00
Ryan Houdek 187c64182b Merge pull request #1628 from Sonicadvance1/fix_finit
OpcodeDispatcher: Fixes FNINIT
2022-03-17 20:37:57 -07:00
Ryan Houdek 60c7ea6e5f Merge pull request #1620 from Sonicadvance1/fix_1618
FEXCore: Fixes #1618
2022-03-17 20:36:06 -07:00
Ryan Houdek 6f1b4b0eee OpcodeDispatcher: Fixes FNINIT
Was incorrectly setting the FCW to 037h when it was supposed to be
037Fh.

Fixes a bug in a visual novel where its CPUID state wouldn't initialize
if this was set incorrectly.
2022-03-17 20:27:22 -07:00
Ryan Houdek fb69300397 FEXCore: Delete IR after it is used
For the JIT cores we don't need to keep IR around, it's only necessary
for the Interpreter. So once the AOT IR service is done dealing with the
IR, check to see if we can delete it.

This causes teeworld's title screen memory usage to go from 730MB to
566MB. 77.5% the memory usage there.

This is effectively an infinite memory leak if the codespace wasn't ever
overwritten or invalidated. So larger memory usage programs would end up
having a larger impact.
2022-03-13 19:01:40 -07:00
Ryan Houdek 5677924525 Merge pull request #1621 from Sonicadvance1/fix_1584
Softfloat: Fixes FSCALE
2022-03-13 18:57:06 -07:00
Ryan Houdek 8422fc632d Merge pull request #1623 from Sonicadvance1/remove_unused_debug_data
FEXCore: Removes unused debug data
2022-03-13 18:56:50 -07:00
Ryan Houdek d33cd744fb FEXCore: Removes unused debug data
This isn't used anywhere. Just remove these.
If we get the imgui debugger running again then we can add even more
stats to sort block costs by.
2022-03-13 18:40:49 -07:00
Ryan Houdek 3b0fb27ae9 Softfloat: Fixes FSCALE
I misread the implementation details of this instruction when
implementing.

The pseudocode says `ST(0) = ST(0) ∗ 2^rndint(ST(1))` so I understood
the instruction to use the current rounding mode of the host to extract
the integer portion of `ST(1)`.

The actual implementation is in the details of the statement `the
integer portion of the floating- point value in ST(1).`

This behaves like round towards zero/truncate, additional hardware
testing and documentation reading confirms this.

Fixes #1584
2022-03-13 14:11:31 -07:00
Ryan Houdek 4603e09a04 FEXCore: Fixes #1618 2022-03-13 13:42:37 -07:00
Ryan Houdek 7b0265ffe2 Merge pull request #1617 from Sonicadvance1/gdbstub_improvements4
GDBServer improvements: Three's a crowd
2022-03-13 13:24:41 -07:00
Ryan Houdek ec54560a38 GDBServer: Fixes memory reading
memory-map is not something we want to use. Adds a comment about it and
disables it.

Also changes core events to wait for an event from GDBStub for waking up
which fixes a hang.
2022-03-13 13:05:00 -07:00
Ryan Houdek 6a5abd3672 Merge pull request #1616 from Sonicadvance1/gdbstub_improvements3
Gdbstub improvements: The sequel
2022-03-13 13:03:25 -07:00
Ryan Houdek 3e6af39c42 GDBServer: Zero initialize some variables to fix connection stability
Otherwise you always had to attempt connecting twice in a row
2022-03-13 12:50:01 -07:00
Ryan Houdek cf82ffc052 GDBServer: Let gdb know when the library map has updated
We need to fetch the full list of map files from the memory map and hand
it over to gdb.
It will then fetch all the libraries from the remote host and give us
backtraces
2022-03-13 12:50:01 -07:00
Ryan Houdek b190150281 FEXCore: Merges redundant string trimming implementations 2022-03-13 12:50:01 -07:00
Ryan Houdek 53ffe5df43 Merge pull request #1613 from Sonicadvance1/gdbstub_improvements2
GDBServer improvements
2022-03-13 12:49:17 -07:00
Ryan Houdek d39df8d3ed Merge pull request #1614 from Sonicadvance1/add_comment
JIT: Adds comment to EmitDetectionString
2022-03-10 14:53:41 -08:00
Ryan Houdek eeb2b928b9 JIT: Adds comment to EmitDetectionString 2022-03-10 14:28:45 -08:00
Ryan Houdek 6cf24a748f GdbServer: Document what PassSignals is for 2022-03-10 14:23:21 -08:00
Ryan Houdek 3b4fd180de GDBServer: Pass auxv better
Fixes 32-bit auxv as well.
2022-03-10 14:21:46 -08:00
Ryan Houdek 7300c7a853 SignalDelegator: Remove anti-pattern usage 2022-03-10 14:00:51 -08:00
Ryan Houdek 376f6db3ac GDBServer: Reformat code to two space tabs
No functional change
2022-03-10 14:00:51 -08:00
Ryan Houdek ad3a960717 GDBServer: Support sending gdb the correct signal
Instead of just sending SIGSEGV, pass the real signal
2022-03-10 13:57:17 -08:00
Ryan Houdek 0a1ef867ee SignalDelegator: Support multiple backend host handlers
This will be necessary for gdbserver to handle signals indepedentally of
the FEX handling.
2022-03-10 13:57:17 -08:00
Ryan Houdek a7fe69deea CPU: Stop trying to initialize signal handlers per thread
These are static per process and only need to be initialized once.
We are going to support multiple signal handlers from the backend after
this, so can only install once.
2022-03-10 13:57:17 -08:00
Ryan Houdek a8c3b6d46f GDBServer: Capture signal capture numbers
This will allow us to wire this to a future signal handler for gdbserver
2022-03-10 13:57:17 -08:00
Ryan Houdek 60db2655bb GDBServer: Encode the return to pread correctly
This encodes the resulting data as raw binary rather than any special
escaped encoding
2022-03-10 13:57:10 -08:00
Ryan Houdek fd717b6995 GDBServer: Expose program offsets better
We were incorrectly returning programing offsets
Get the program offset from the frontend so we can know what to give gdb
2022-03-09 19:07:45 -08:00
Ryan Houdek ba37388fe3 GDBServer: Expose auxv values
We already expose these in the code loader, pump it through gdb
2022-03-09 19:07:45 -08:00
Ryan Houdek 68f32d85c9 CodeLoader: Expose base ELF loaded offset
Useful for gdb
2022-03-09 19:07:45 -08:00
Ryan Houdek 2aa77e85de LinuxSyscalls: Expose CodeLoader through syscall interface 2022-03-09 19:07:45 -08:00
Ryan Houdek 91665fdf0e Merge pull request #1610 from wannacu/main
FileManager: Fix realpath failed on debian buster
2022-03-09 18:58:52 -08:00
Ryan Houdek 23a1c64bf7 Merge pull request #1612 from Sonicadvance1/tag_memory_allocations
JITs: Emit identification string in the code buffers
2022-03-09 18:52:46 -08:00
Ryan Houdek c2dcf06632 JITs: Emit identification string in the code buffers
At the start of each code buffer, emit a small string for letting memory
inspection know if a code region is for the JITs.
2022-03-09 18:29:37 -08:00
Ryan Houdek fad91bb818 Merge pull request #1609 from Sonicadvance1/fix_map_32bit
Linux: Fixes MAP_32BIT supported range
2022-03-08 17:33:34 -08:00
wannacu 0e769ece26 docs: Update Readme_CN.md 2022-03-08 18:10:08 +08:00
wannacu 9a780b40a2 Docs: Add Chinese README 2022-03-08 18:04:09 +08:00
wannacu 898873e9e3 FileManager: Fix realpath failed on debian buster
This happend on debian buster when run realpath(i386) on arm64 host.
2022-03-08 16:42:05 +08:00
Ryan Houdek 52292e5f7e Linux: Fixes MAP_32BIT supported range
I accidentally committed a 32-bit range that was significantly smaller
than what it should be.
While the minimal range worked for simple cases, it didn't work for
anything complex.
Give it the full range it needs.

Fixes #1600
2022-03-06 17:46:17 -08:00
Ryan Houdek 5de6c866b7 Merge pull request #1608 from Sonicadvance1/termux_build_option
Adds a cmake option for forcing a termux build
2022-03-06 13:54:20 -08:00
Ryan Houdek ec0cd3aec4 Adds a cmake option for forcing a termux build
This is necessary when cross-compiling rather than building on-device
2022-03-06 12:56:25 -08:00
Mai M f5f9512d9a Merge pull request #1606 from Sonicadvance1/fhu_page_size
Change page define usages over to self-defined
2022-03-06 15:53:15 -05:00
Mai M fb27cb4356 Merge pull request #1607 from Sonicadvance1/disable_guis_termux
Disables GUI applications in a Termux build
2022-03-06 15:52:37 -05:00
Mai M 94664580c8 Merge pull request #1605 from Sonicadvance1/update_docs_termux
Update ReleaseProcess docs for Termux
2022-03-06 15:52:10 -05:00
Ryan Houdek 99a93fa9ea Disables GUI applications in a Termux build 2022-03-06 08:09:55 -08:00
Ryan Houdek 4cb6918506 Change page define usages over to self-defined
In the case of an AArch64 builder is using 16kb or 64kb pages like is
common on servers then it would fail to compile, even if the resulting
application would only ever run on 4k page hosts.

Resolve this by removing the build check and hardcoding 4kb pages for
each of our uses. We still require 4kb pages to run, so this mostly just
removes the weirdness where it is 16kb builder + 4k runner. Would have
broken some of our assumptions when running.
2022-03-06 07:33:10 -08:00
Ryan Houdek 9cc743bf84 Update ReleaseProcess docs for Termux
FEX hardly works on Termux as-is, but we should make sure to document
how to update the packages otherwise we will quickly become outdated on
their package management.
2022-03-06 05:59:33 -08:00
Ryan Houdek a408749eef Docs: Update for release FEX-2203 2022-03-06 04:49:31 -08:00
Ryan Houdek d8a3687ac3 Merge pull request #1604 from Sonicadvance1/fix_cmpxchg_66h
OpcodeDispatcher: Fixes CMPXCHG8B/16B with 66h/72h/73h prefix
2022-03-06 04:10:58 -08:00
Ryan Houdek a32c7f6ce2 unittests: Adds cmpxchg unit tests for prefixes 2022-03-06 03:54:45 -08:00
Ryan Houdek 540feb857b OpcodeDispatcher: Fixes CMPXCHG8B/16B with 66h/72h/73h prefix
The documentation is incorrect about this instruction. It claims that
you use 66h prefix to choose between operating at 8B or 16B.
This is incorrect, real hardware only responds to REX.W for choosing the
operating size. These other prefixes are ignored but is still accepted as
an instruction decoding.
2022-03-06 03:54:45 -08:00
Ryan Houdek a3902a0d2d FEXCore: Fixes usage of GPRPair in operations
These were working around the previous quirks by accident
2022-03-06 03:54:45 -08:00
Ryan Houdek 5fbd01536f IR: Fixes some GPRPair IR op definitions
These were always wrong but how it the operations were handled meant
that it happened to work even though the IR representation was broken
2022-03-06 03:54:45 -08:00
Ryan Houdek bebcab0277 Merge pull request #1603 from Sonicadvance1/rng_support
FEXCore: Adds support for RDRAND/RDSEED
2022-03-06 03:54:20 -08:00
Ryan Houdek d0f17d400e unittests: Adds RDRAND/RDSEED unit tests 2022-03-06 03:40:25 -08:00
Ryan Houdek 11a07eb3f2 FEXCore: Adds support for RDRAND/RDSEED
This matches the AArch64 implementation fairly well.
Bundles RDRAND and RDSEED together for simplification, both instructions
are a single flag on AArch64.
2022-03-06 03:40:25 -08:00
Ryan Houdek cb13e1bdb9 X86Tables: Fixes secondary group decoding
If we're hitting these group tables then it needs the ignore overlay extension
since the prefixes are used to select ops inside this table.
2022-03-06 03:40:25 -08:00
Ryan Houdek 3e8c6d0be0 Merge pull request #1601 from Sonicadvance1/new_ir_json
IR: New IR JSON format
2022-03-06 03:40:05 -08:00
Ryan Houdek 847b6e5026 LinuxSyscalls: Fixes struct verifier on Ubuntu 20.04
The `linux/types.h` header needs to be included before the rest
otherwise we are missing types.
2022-03-04 17:44:26 -08:00
Ryan Houdek 7f6e9d3dae unittests/IR: Resolves fallout from recent JSON changes 2022-03-04 16:39:11 -08:00
Ryan Houdek e244e8f66b x86Jit: Resolves fallout from recent JSON changes 2022-03-04 16:39:11 -08:00
Ryan Houdek 0adaa85d97 ArmJit: Resolves fallout from recent JSON changes 2022-03-04 16:39:11 -08:00
Ryan Houdek f1d34c7407 Interpreter: Resolves fallout from recent JSON changes 2022-03-04 16:39:11 -08:00
Ryan Houdek 7a9492ceca IRPasses: Resolves fallout from recent JSON changes 2022-03-04 16:39:11 -08:00
Ryan Houdek 39192cd092 OpcodeDispatcher: Resolves fallout from recent JSON changes
Also fixes a bug in FCVTIntTo where it was getting passed an FPR when it
wants a GPR. Would cause RA problems on the JITs
2022-03-04 05:10:45 -08:00
Ryan Houdek 73abf9b9dd IRParser: Resolves fallout from recent JSON changes 2022-03-04 04:07:25 -08:00
Ryan Houdek 0387c24e21 IREmitter: Resolves the fallout from recent JSON changes 2022-03-04 04:07:25 -08:00
Ryan Houdek 9d14bc7846 IR: Updates generators and json to new format
This greatly simplifies the IR format by using string parsing for
gathering the information.
Tons of redundant information removed.
Significantly more difficult to mess up adding a new IR op.

Significantly improves the generator functions in IREmitter
2022-03-04 03:51:44 -08:00
Ryan Houdek ac32ecadbe Merge pull request #1597 from Sonicadvance1/3dnow_and_back_again
OpcodeDispatcher: Implements all the 3DNow! instructions
2022-03-04 03:36:40 -08:00
Mai M 08938ecb11 Merge pull request #1596 from Sonicadvance1/fix_old_kernel_bug
LinuxAllocator: Fixes bug with old kernels and hint allocation
2022-03-02 00:26:51 -05:00
Ryan Houdek 4a3cbf1b1f Merge pull request #1598 from Sonicadvance1/hypervisor_cpuid
CPUID: Implements leaf 4000_0000
2022-03-01 21:02:32 -08:00
Ryan Houdek 0a0bc21c88 CPUID: Implements leaf 4000_0000
This region is reserved for hypervisor uses. Let's follow other examples
and return a hypervisor vendor id signature as another way for software
to find if it is running under FEX-Emu.
2022-03-01 20:47:03 -08:00
Mai M a7ad7f4456 Merge pull request #1593 from Sonicadvance1/add_robin_map
Adds tsl::robin_map
2022-03-01 23:24:34 -05:00
Mai M 46ac05e3cf Merge pull request #1599 from Sonicadvance1/deprecated_distutils
Scripts: Stop using deprecated Distutils
2022-03-01 23:23:55 -05:00
Ryan Houdek 9911fe68d4 Scripts: Stop using deprecated Distutils
According to PEP 386: https://www.python.org/dev/peps/pep-0386/

distutils is deprecated and will be removed in an upcoming python
version.

Switch over to pkg_resources for version parsing and comparison
2022-03-01 07:10:41 -08:00
Ryan Houdek 57a56545b2 Merge pull request #1589 from Sonicadvance1/fix_vixl_assert
Update vixl to fix assert
2022-03-01 05:35:31 -08:00
Ryan Houdek b4e0565907 Merge pull request #1590 from Sonicadvance1/termux_fixes
Termux fixes
2022-03-01 04:08:48 -08:00
Ryan Houdek 5137af5bae Fixes epoxy include in FEXConfig and FEXLogServer 2022-03-01 03:55:19 -08:00
Ryan Houdek 385ed2c2ec Update imgui to fix autodetect 2022-03-01 03:55:16 -08:00
Ryan Houdek 3418cc8054 LinuxSyscalls: More type fixes 2022-03-01 03:55:15 -08:00
Ryan Houdek 1f835ea1f4 LinuxSyscalls: Fixes semid and ipc types
Newer headers redefine semid_ds and ipc_perm as semid64_ds and
ipc64_perm silently.
Use the new types directly since we are a 64-bit only application.
2022-03-01 03:55:12 -08:00
Ryan Houdek 0a40753624 More missing include fixes 2022-03-01 03:55:10 -08:00
Ryan Houdek 609538e758 Utils/Allocator: Use a namespace alias for pmr
This still lives under experimental in Termux environment
2022-03-01 03:55:08 -08:00
Ryan Houdek 09010e1292 LinuxSyscalls: Remove unused headers now
These don't even exist on termux.
2022-03-01 03:54:01 -08:00
Ryan Houdek bddf3871ba LinuxSyscalls/x32/Types: Fixes stat type definitions
The time argument definitions are defines in termux.
Rename our definition name of these so we don't get caught by define.
2022-03-01 03:50:18 -08:00
Ryan Houdek 1ca9e56502 LinuxSyscalls/x64/Types: Fix guest_stat definition
kernel types don't exist in termux. Use uint64_t and int64_t directly.

Also using reserved `__` causes compile failure.
2022-03-01 03:50:17 -08:00
Ryan Houdek fe58a9ae2c LinuxSyscalls/Thread: Don't use set_robust_list on Termux
Would get caught by seccomp and crash FEX
2022-03-01 03:50:15 -08:00
Ryan Houdek bc7c4dfa74 LinuxSyscalls/x32/Types: Don't redefine SIGEV defines 2022-03-01 03:50:12 -08:00
Ryan Houdek 7af7b1309d Msg: Switch msqd_t to FEX defined type 2022-03-01 03:50:10 -08:00
Ryan Houdek dd4630750f LinuxSyscalls/Types: Adds missing types for Termux 2022-03-01 03:50:07 -08:00
Ryan Houdek 302da029eb Work around Termux not supporting hardlinks
The Android filesystem they are on just doesn't support them
Instead of hardlinking FEXLoader to FEXInterpter, just build the
executable twice and eat the filesystem cost.
2022-03-01 03:50:05 -08:00
Ryan Houdek 3cc59bc68e LinuxSyscalls: Switches to a bunch of raw syscalls
For older and Termux build environments these helper libc functions
don't exist.
2022-03-01 03:49:31 -08:00
Ryan Houdek 30c27851ef FEXRootFSFetcher: Termux build environments 2022-03-01 03:30:47 -08:00
Ryan Houdek 9b29ff61e9 Adds some missing headers 2022-03-01 03:30:45 -08:00
Ryan Houdek 91f780223b Stop self-defining PAGE_SIZE
We only work on targets with 4096 byte page sizes.
Adds a cmake compile test to ensure this is adhered to.
2022-03-01 03:30:43 -08:00
Ryan Houdek 2933a00b12 Merge pull request #1579 from Sonicadvance1/optimize_syscalls_with_flags
Allow classifying syscalls with flags
2022-03-01 03:23:43 -08:00
Ryan Houdek a0efd2b01f Resolve comments. 2022-02-28 21:04:03 -08:00
Ryan Houdek 3e9dbda146 Classify syscalls 2022-02-28 21:04:03 -08:00
Ryan Houdek 6501715a3c Allow classifying syscalls with flags
In some cases we can generate more optimal code if we have more
information about a syscall which number gets const-propagated.

In particular optimizing through syscalls, not synchronizing state, and
never returning.

- Noreturn is used by a syscall that never returns, like exit.

This means that it never needs to try and synchronize state coming back

- Not synchronizing state and optimizing through syscalls

Useful for syscalls that don't read the state past arguments and only
returns a value.
2022-02-28 21:03:54 -08:00
Ryan Houdek 2af23d9bec Merge pull request #1591 from Sonicadvance1/new_cpus_in_native_fit
Scripts: Updates CPU fitting script for latest CPUs
2022-02-28 07:05:45 -08:00
Ryan Houdek 3ba2d6cfb4 Merge pull request #1586 from Sonicadvance1/testharness_env
TestHarnessRunner: Wire up environment variable option setting
2022-02-28 07:05:28 -08:00
Ryan Houdek abd266441c unittests: Implements 3DNow! unit tests
Covers the full space, of which there aren't many.

3DNow! unit tests are disabled on the CI runner since the x86 CPU in CI
doesn't support it.
2022-02-28 04:06:14 -08:00
Ryan Houdek b1b6078518 CPUID: Enables 3DNow! + Extensions
Now that we support these
2022-02-28 04:05:18 -08:00
Ryan Houdek a9a89bea9f OpcodeDispatcher: Implements all the 3DNow! instructions
This picks up all the instruction implementations, including 3DNow!
Extended and the Geode specific instructions that were added.

Most of these match preexisting SSE instructions except that they
operate at 64-bit and in the MMX registers.
2022-02-28 04:03:55 -08:00
Ryan Houdek 9b1b2e6496 X86Tables: Fills out 3DNow tables
Fully decoded the same way and adds the Geode specific instruction
decodings as well.
2022-02-28 04:02:57 -08:00
Ryan Houdek 638da92f45 Frontend: Fixes minor bug decoding 3DNow!
We already decoded the modrm `rm` bits, check while decoding modrm to
ensure we don't try decoding it again.
Was causing double decoding of SIB and displacement bytes, breaking
things
2022-02-28 04:01:32 -08:00
Ryan Houdek 0f3a169cc4 Opdispatcher: Minor bug fix with unimplemented op
If multiblock isn't enabled then on Unimplemented op we shouldn't create
a new block.

Was causing IR validation to get angry
2022-02-28 04:00:41 -08:00
Ryan Houdek c7dd176799 IR: Adds a VRev64 op
This directly matches the AArch64 instruction and will be used shortly
2022-02-28 04:00:13 -08:00
Ryan Houdek 23daf4cb72 LinuxAllocator: Fixes bug with old kernels and hint allocation
In the face of an application using MAP_FIXED_NOREPLACE AND the host
linux kernel doesn't understand this flag. Then we were falling down the
hint allocation path which would allocate a pointer in 64-bit space,
returning this pointer to a 32-bit userspace and breaking things.

Now when the hint fails with this flag, we know that it intersecting a
range and can early exit.
2022-02-26 21:12:44 -08:00
Ryan Houdek bb7fa84fb8 Scripts: Updates CPU fitting script for latest CPUs
Clang-13 doesn't yet understand the latest ARM CPUs so just document them
and set to the closest thing.
2022-02-26 02:14:52 -08:00
Ryan Houdek 3325ba52b9 Adds tsl::robin_map
This will be used with the code serialization service soon
2022-02-26 00:43:43 -08:00
Ryan Houdek 7f47fe6d73 Update vixl to fix assert
Any hardware using MTE will assert without this
2022-02-24 13:46:56 -08:00
Ryan Houdek ee165379c5 TestHarnessRunner: Wire up environment variable option setting
Wire up the environment variable option setting so asm files can set
these and it works
2022-02-24 13:37:39 -08:00
Ryan Houdek fa554d3096 Merge pull request #1588 from Azkali/main
Improve compatibility with older uapi kernel headers
2022-02-24 01:34:24 -08:00
The Great Wizard Azkali d60710d3f1 Define proper statx syscall depending on CPU architecture 2022-02-24 10:20:28 +01:00
Azkali 75988b2ae5 Improve compatibility with older uapi kernel headers
Following up the work previously done in 2079f6b3c7.
Adding more defines for older Linux uapi headers missing some defines.
2022-02-24 09:49:56 +01:00
Mai M 30803c66f7 Merge pull request #1587 from Sonicadvance1/fix_missing_telemetry_names
Telemetry: Fix missing telemetry names
2022-02-22 21:46:00 -05:00
Ryan Houdek a9d838fa27 Telemetry: Fix missing telemetry names
Didn't have names for tearing
2022-02-22 18:30:40 -08:00
Ryan Houdek 2b8f60c108 TestHarness: Support for asm files having the option to set config options
Allows some something like the following:
"Env": {
  "FEX_MAXINST": "500"
}

Not that I would recommend overriding MAXINST in the asm tests, as
command line overrides that
2022-02-21 14:53:27 -08:00
Mai M 5ec6ee5b69 Merge pull request #1578 from Sonicadvance1/update_vixl
Updates vixl for new cursor updating methods
2022-02-17 16:18:26 -05:00
Mai M 0db7205f61 Merge pull request #1582 from Sonicadvance1/ccache_option
Adds option to disable ccache
2022-02-16 22:49:48 -05:00
Ryan Houdek 6027494d69 Adds option to disable ccache
Can be useful when running static analysis tools
2022-02-16 19:24:24 -08:00
Mai M 252dcfe26f Merge pull request #1581 from Sonicadvance1/add_required_growsdown
FEXLoader: Adds back required MAP_GROWSDOWN
2022-02-16 19:07:40 -05:00
Ryan Houdek 944c93c10c FEXLoader: Adds back required MAP_GROWSDOWN
I was overzealous with my removal of MAP_GROWSDOWN.
We still require the primary thread to have this flag set.
2022-02-16 15:41:44 -08:00
Ryan Houdek a5fb7e7313 Merge pull request #1580 from Sonicadvance1/fix_hostthunks_install
Fixes Host and guest thunks install path
2022-02-15 15:46:10 -08:00
Ryan Houdek 364b3380fc Fixes Host and guest thunks install path
Hosts were using the cmake install path with $DESTDIR which duplicates
paths.

GuestThunks were doing some magic that wasn't actually necessary
2022-02-15 15:36:31 -08:00
Ryan Houdek a0edab8040 Updates vixl for new cursor updating methods
These will be required for code cache
2022-02-14 16:33:51 -08:00
Ryan Houdek b65194f433 Merge pull request #1576 from Sonicadvance1/move_x87_constant_helpers
JIT: Implements x87 fallback helpers as lookups in to state
2022-02-14 14:13:33 -08:00
Ryan Houdek 0d1c9cd7df JIT: Implements x87 fallback helpers as lookups in to state
This allows x87 fallbacks to be loaded from the upcoming code cache
without relocations.

Only 40 pointers necessary to store and means x87 code won't hit
relocations heavily.

Probably improves performance slightly on the x86 host side but should
be neglible.

Needs #1574 and #1575 merged first.
2022-02-14 14:00:10 -08:00
Ryan Houdek 99dcda7f8c Merge pull request #1575 from Sonicadvance1/move_constant_functions_x86
JITx86: Switches over to loading pointers from state
2022-02-14 13:59:07 -08:00
Ryan Houdek 2aa4d33de5 JITx86: Switches over to loading pointers from state
Just like the previous AArch64 JIT.
These pointers are process or thread specific depending on the pointer
and should be loaded from the State object.

Performance here might slightly increase.

This is required for code cache on x86

Needs #1574 merged first.
2022-02-14 13:49:54 -08:00
Ryan Houdek 1ef78d0a00 Merge pull request #1574 from Sonicadvance1/move_constant_functions
ARMJIT: Switches over to loading pointers from state
2022-02-14 13:39:28 -08:00
Stefanos Kornilios Mitsis Poiitidis 64d1840f1e Merge pull request #1571 from Sonicadvance1/fix_syscall_strace
Linux: Fix missing types for syscall strace
2022-02-14 19:05:48 +02:00
Stefanos Kornilios Mitsis Poiitidis 34530236f9 Merge pull request #1573 from Sonicadvance1/disable_int_tests_with_no_int
unittests: Disables Interpreter tests when its disabled
2022-02-14 19:00:13 +02:00
Ryan Houdek f5a9da082f ARMJIT: Switches over to loading pointers from state
These pointers are process or thread specific depending on which pointer
it is.
All of these pointers end up getting used inside of the JIT blocks
themselves and with code caching would result in a ton of relocations
occuring inside the code.

The pointers used within the dispatcher don't currently matter since I'm
not expecting to cache the dispatcher itself. It's only a page per
thread after all. This does move us significantly closer towards using a
single dispatcher for all threads though.

The performance impact of this change is unlikely to be felt at all,
some locations have less code generation which could improve perf
slightly. Some locations move from a 1-3 cycle constant calculation to a
4 cycle load, hard to be felt since it gets hidden by other
instructions.

x86-64 JIT will be added soon after this
2022-02-13 23:06:26 -08:00
Ryan Houdek 9b49e8cb59 FEXCore: Adds utility class for class member function casting
Adds validation for our class member casting to ensure we don't try
casting a virtual member.
2022-02-13 23:06:25 -08:00
Ryan Houdek 1b6d20b731 unittests: Disables Interpreter tests when its disabled
Would result in failures if you weren't expecting it.
2022-02-11 12:55:10 -08:00
Mai M e9c7c76174 Merge pull request #1572 from Sonicadvance1/LoadConstant_no_opt
Arm64Emitter: Allow non-optimizing LoadConstant
2022-02-11 00:50:56 -05:00
Ryan Houdek cbce06d012 Arm64Emitter: Allow non-optimizing LoadConstant
This is pulled from the code cache PR. Will be necessary for supporting
relocations.

Not currently being used but will be once we have code caching in place.
2022-02-10 20:04:02 -08:00
Ryan Houdek d95326b23b Linux: Fix missing types for syscall strace 2022-02-10 19:46:43 -08:00
Mai M 43fada7555 Merge pull request #1568 from Sonicadvance1/fix_musl_load
ELFCodeLoader: Fixes typo in AT_BASE calculation
2022-02-10 21:25:48 -05:00
Mai M afa7172cb1 Merge pull request #1570 from Sonicadvance1/remove_growsdown
Removes MAP_GROWSDOWN usage
2022-02-10 21:25:23 -05:00
Ryan Houdek b88d8a7cc4 Removes MAP_GROWSDOWN usage
This is just a memory leak waiting to happen.
Only the primary thread in an application really should have this set
since the kernel cleans it up.

We only ever allocate the primary thread of the guest application then
every host thread's stack on top of that. It's up to the guest when it
is cloning to set up new stack pointers, we don't manage that.

We are already allocating the first thread's size at the soft stack
limit with RLIMIT_STACK anyway.

Fixes #1556
2022-02-10 18:03:18 -08:00
Ryan Houdek 6add09b78b ELFCodeLoader: Fixes typo in AT_BASE calculation
Fixes executing musl applications with the dynamic linker.

It was using the main executable's p_offset instead of the
interpreter's.
Wasn't a problem with glibc since it uses a different symbol to find the
base (Don't ask me why it does this).

musl dynamic linker on the other hand just uses AT_BASE directly and
since it was calculated incorrectly it was crashing.

Testing application was `ls` which had a p_offset of 0x40, so it would
try and read some values from AT_BASE, starting at an offset below where
it was mapped.
2022-02-10 17:43:58 -08:00
Mai M 4bb3a54ccf Merge pull request #1566 from neobrain/refactor_thunk_misc
Miscellaneous thunk cleanups
2022-02-10 15:33:34 -05:00
Mai M 76f86e51a6 Merge pull request #1567 from Sonicadvance1/fix_fexgetconfig_rootfs
FEXGetConfig: Fix --current-rootfs option
2022-02-10 15:15:26 -05:00
Ryan Houdek 2a1b27df58 FEXGetConfig: Fix --current-rootfs option
If the configured rootfs wasn't a squashfs then it was failing to return
the directory.

Now it works for both squashfs and directory rootfs again.
2022-02-10 11:45:29 -08:00
Tony Wasserka 68426735a5 Thunks: Clean up ASTMatcher-based testing helpers
The run_thunkgen* helpers now parse generated source code and return its AST
representation, so HasASTMatching helper calls don't each need to redundantly
compile it themselves. This also ensures the generator output actually compiles
in tests where we didn't explicitly check that before.

This also allows printing the full AST of the generator output on test
failures. This must be enabled manually by changing a variable in the ostream
output operator for SourceWithAST.
2022-02-10 12:11:52 +01:00
Tony Wasserka ee6b558000 Thunks: Fix warning about unused field 2022-02-10 12:11:52 +01:00
Tony Wasserka 3c7872c6d5 Thunks: Rename FrontendAction to GenerateThunkLibsAction 2022-02-10 12:11:52 +01:00
Tony Wasserka 4aed6fc56c Thunks/gen: Remove now unneeded code 2022-02-10 12:11:52 +01:00
Tony Wasserka 5d555a10a2 Thunks: Explicitly put thunks into the text library section
Previously, defining zero-initialized variables right before LOAD_LIB
could cause the compiler to put thunk definitions into bss, hence triggering
errors during assembly ("attempt to store non-zero value in section `.bss'").
2022-02-10 12:11:52 +01:00
Ryan Houdek 5854d4ad1c Merge pull request #1565 from neobrain/refactor_thunk_ide_integration
Enable proper IDE integration of thunk libraries
2022-02-10 02:44:29 -08:00
Tony Wasserka bd6999eb87 CMake: Clean up build architecture for ThunkLibs
Host thunk libraries are always built as part of the main project now.
Guest thunk libraries are still cross-compiled in a CMake ExternalProject,
but *additionally* there are CMake targets in the main project to make
sure IDE engines can properly handle guest source files.
2022-02-10 11:23:43 +01:00
Tony Wasserka 3ebb2eaf0d Thunks: Fix guest libs build on clang 2022-02-10 11:23:41 +01:00
Mai M defd3be30c Merge pull request #1563 from Sonicadvance1/fix_auto_script
Updates Readme to fix install script
2022-02-09 18:18:39 -05:00
Mai M a8e5a0a68e Merge pull request #1562 from Sonicadvance1/remove_debug_memory_mapping
FEXLoader: Removes memory mapping check on startup
2022-02-09 18:18:18 -05:00
Ryan Houdek f27c43ff8a Updates Readme to fix install script
Fixes an issue where the FEXRootFSFetcher wouldn't get a any user input
and just fail out.
Save it to the tmp folder and execute from there instead.

Fixes #1557
2022-02-09 14:01:48 -08:00
Ryan Houdek 8840fc818c FEXLoader: Removes memory mapping check on startup
FEX always builds with PIE and we don't hit this issue anymore anyway.
If some application wants to inject a page in to the lower 32-bits then
we have no reason to complain about it anymore. Just let it go and
hopefully they know what they are doing.

Fixes #1559
2022-02-09 13:52:05 -08:00
Mai M ffcaf294e3 Merge pull request #1555 from Sonicadvance1/weirdo_edge_case
OpcodeDispatcher: Fixes weirdo edge case in segment moving
2022-02-08 00:26:55 -05:00
Ryan Houdek 6b7a84bef2 OpcodeDispatcher: Fixes weirdo edge case in segment moving
Just noticed this while casually reading the x86 architecture manuals.
The move segment registers instructions ignore the REX.R prefix on the
segment register.

Previously this was expected to create an invalid register selection.
A little bit silly but sure, support it.
2022-02-07 21:12:53 -08:00
Mai M 287c65dc64 Merge pull request #1554 from Sonicadvance1/fix_tricky_stat
Linux: x32: Fixes tricky stat64 defines
2022-02-06 19:52:31 -05:00
Mai M 62397bb19c Merge pull request #1553 from Sonicadvance1/fix_sigevent
Linux: Make sure to use correct accessors for sigevent
2022-02-06 19:52:14 -05:00
Ryan Houdek 0216bcf27f Linux: x32: Fixes tricky stat64 defines
Some build environments use a define to change stat64 and statfs64 to be
the same definition as stat and statfs.

Check if the define exists and if it does then remove the 64bit
constructors.
2022-02-06 16:05:27 -08:00
Ryan Houdek b981fcfe42 Linux: Make sure to use correct accessors for sigevent
Some of these are defined differently depending on environment
2022-02-06 15:33:07 -08:00
Mai M cb491a8acb Merge pull request #1552 from Sonicadvance1/fix_ucontext_copy
UContext: Fixes 32-bit siginfo_t copying definition
2022-02-06 18:23:04 -05:00
Mai M 1b99495b4b Merge pull request #1551 from Sonicadvance1/fix_older_env
Some fixes for older environments
2022-02-06 18:22:26 -05:00
Ryan Houdek 3c5a2cec90 UContext: Fixes 32-bit siginfo_t copying definition
The host provided siginfo_t definition can vary depending on the build
environment.
What doesn't change however is how the data is laid out.
It's always 128bytes, The first three 32-bit words are always known.
The 64-bit host side always has an additional 32-bit pad member.
Then the remaining bytes is the sifields.
2022-02-06 15:03:56 -08:00
Ryan Houdek 9a64e7f100 Linux: Renamed some 64-bit syscall names
Some build environments use defines to rename these. Which breaks our
naming
2022-02-06 14:19:30 -08:00
Ryan Houdek 8e8baec47a Linux: Use raw syscalls for pkey syscalls
For older libc environments
2022-02-06 14:19:30 -08:00
Ryan Houdek 59ca60e39f Fixes a bunch of header includes
Necessary for older build environments
2022-02-06 14:19:30 -08:00
Ryan Houdek 8c956e6ce1 Docs: Update for release FEX-2202 2022-02-05 22:48:35 -08:00
Ryan Houdek 832d013c92 Merge pull request #1513 from Sonicadvance1/reduce_flags_memory_usage
FEXCore: Defer a significant number of ALU flag calculation
2022-02-04 16:22:24 -08:00
Mai M 1c24206117 Merge pull request #1550 from Sonicadvance1/fix_weirdo_crc32
OpcodeDispatcher: Fixes CRC32 decoding in 0F_38 table
2022-02-04 01:01:50 -05:00
Ryan Houdek f1979c15a2 unittests: Adds new CRC32 unittests
The instruction decode tables for crc32 introduced some dumb.
`F2h` and `F2h && 66h` prefixes both work for crc32.
This is a failure on Intel's part for sticking crc32 in to the vector
table.

MOVBE without any prefixes also does the same garbage where prefix `66h`
acts as an operand prefix size ONLY.
2022-02-03 21:11:50 -08:00
Ryan Houdek 556a1dab24 OpcodeDispatcher: Fixes CRC32 decoding in 0F_38 table
This table is particularly terrible. CRC32 is the first instruction in
this table that needs either prefix `72h` OR `66h && F2h`

For 8bit CRC32, this ignores the 66h operand size override prefix.
  - But our table decoding didn't handle this
For 16bit/32bit/64bit CRC32 this behaviour changes depending on 66h
prefix AND REX.W
  - 66h prefix is ignored when REX.W is set, always 64bit but it falls
    down the other table path

This is an absolutely weird edge case that nobody should hit, but here
we are.
2022-02-03 21:11:50 -08:00
Mai M caffad8562 Merge pull request #1549 from Sonicadvance1/implement_pcmpgtq
OpcodeDispatcher: Implements PCMPGTQ
2022-02-03 21:46:13 -05:00
Mai M 5978143141 Merge pull request #1547 from Sonicadvance1/remove_system_xxhash
CMake: Always use local xxhash to statically link
2022-02-03 21:45:58 -05:00
Ryan Houdek 594c70b5e0 OpcodeDispatcher: Implements PCMPGTQ
I thought we already had this implemented but I guess it was missed.

Required for SSE 4.2
2022-02-03 18:36:29 -08:00
Ryan Houdek 655e6989ca FEXCore: Defer a significant number of ALU flag calculation
This was mainly an optimization around memory usage. ALU ops tend to
bloat the IR quite heavily, but I also noticed a 2-4% uplift in
performance of some applications. So a nice side effect.

Should let us more aggressively target reducing our IR intrusive
allocator size since this is quite reduced.

In a pedantic heavy ALU op code block this reduces the number of IR ops
from 14,756 IR ops to 2,016 prior to optimization.
After optimization both had reduced down to 50 IR ops, proving the
output IR was the same.
2022-02-03 01:37:59 -08:00
Ryan Houdek afeb228a89 CMake: Always use local xxhash to statically link
Dynamically linking xxhash is causing problems with pressure-vessel.

With this in place we only have the typical C++ dependencies
```
$ ldd ./Bin/FEXLoader
        linux-vdso.so.1 (0x00007fff44d9d000)
        libstdc++.so.6 => /lib/x86_64-linux-gnu/libstdc++.so.6 (0x00007f4c4d884000)
        libm.so.6 => /lib/x86_64-linux-gnu/libm.so.6 (0x00007f4c4d7a0000)
        libgcc_s.so.1 => /lib/x86_64-linux-gnu/libgcc_s.so.1 (0x00007f4c4d786000)
        libc.so.6 => /lib/x86_64-linux-gnu/libc.so.6 (0x00007f4c4d55e000)
        /lib64/ld-linux-x86-64.so.2 (0x00007f4c4e0fa000)
```
2022-02-03 01:31:43 -08:00
Mai M d308a438ea Merge pull request #1546 from Sonicadvance1/fix_fexconfig
Fixes FEXConfig build
2022-02-01 21:13:08 -05:00
Ryan Houdek 260fc8ba52 Fixes FEXConfig build
Oops. This was added late and didn't test it.
2022-02-01 15:44:06 -08:00
Mai M ade0d0f241 Merge pull request #1543 from Sonicadvance1/fixes_for_1423
Linux: Fixes for older build environments
2022-02-01 16:52:40 -05:00
Mai M 11a5105547 Merge pull request #1544 from Sonicadvance1/allow_disable_interpreter
Adds an option to disable the IR interpreter
2022-02-01 16:52:22 -05:00
Ryan Houdek 10ad5db686 Adds an option to disable the IR interpreter
By default we won't build with the interpeter to reduce user confusion.
The interpreter isn't really useful to end users so remove it.

Completely removes it from building except for the fallback operations.

This also removes the selection from FEXConfig to remove selection
confusion there.

File Stats:
FEXLoader Size with Interpreter:    3422768 bytes
FEXLoader Size without Interpreter: 3301944 bytes
Size difference:                    96.4699915%
Bytes removed:                      120824 bytes
4k pages removed:                   29.498046875 -> 30 rounded up

VM Stats (Reported from bloaty):
Memory Size with Interpreter:    6.50Mi
Memory Size without Interpreter: 6.38Mi
Size difference:                 98.1538462%
2022-02-01 13:00:29 -08:00
Ryan Houdek 68c441575d Linux: Fixes for older build environments
Should resolve the new building issues from #1423
2022-02-01 12:17:09 -08:00
Ryan Houdek 334a8ef87c Merge pull request #1542 from Sonicadvance1/fix_pressure_vessel_hangs
Fix pressure vessel hangs
2022-01-31 08:57:20 -08:00
Ryan Houdek b7a76af72f Merge pull request #1541 from Sonicadvance1/implement_crc
OpcodeDispatcher: Implements CRC32 instruction
2022-01-31 08:57:01 -08:00
Ryan Houdek 4c92b562b8 Merge pull request #1540 from Sonicadvance1/remove_extract
OpcodeDispatcher: Removes extraneous extract in VFCMP
2022-01-31 08:56:47 -08:00
Stefanos Kornilios Mitsis Poiitidis 9d08451903 Merge pull request #1536 from Sonicadvance1/fix_orbitals
Softfloat: Stop doing special handling for FREM
2022-01-31 16:51:18 +02:00
Stefanos Kornilios Mitsis Poiitidis c252f8bfc5 Merge pull request #1539 from Sonicadvance1/fix_wrong_offsets
IR: Fixes some wrong offsets in passes
2022-01-31 15:40:33 +02:00
Ryan Houdek dc7ec6377b Linux: Safely handle Filemanagement mutex on fork
If an application is forking heavily with threaded file accesses
happening then the mutex can end up in an unknown state.

On fork make sure to lock the mutex then immediately unlock after fork
occurs.

This final step resolves hanging that pressure-vessel hits on startup.
Since it is doing a ton of file opening and forking during
initialization.
2022-01-30 18:15:57 -08:00
Ryan Houdek ce6f4edaaa FileManagement: Use ScopedSignalMaskWithMutex
When using mutexes in syscall helpers we need to be extra careful around
signals.
2022-01-30 18:15:57 -08:00
Ryan Houdek 983c35ea3b Allocator: Use ScopedSignalMaskWithMutex
Instead of just a basic mutex, also mask the signals.
This fixes the problem where we can end up receiving a signal in the
middle of memory allocation. Thus leaving the locked mutex in a broken
state.

This more closely matches the Linux kernel behaviour.
Since if you're in the middle of a memory allocating syscall, you won't
get signaled.
2022-01-30 18:15:57 -08:00
Ryan Houdek 70aaa1117a FEXHeaderUtils: Adds ScopedSignalMaskWithMutex
This class allows a scoped region lock a mutex and mask signals.

This is necessary for thread and signal safety coming up
2022-01-30 18:15:57 -08:00
Ryan Houdek 59e9859087 unittests: Implements CRC32 unit tests 2022-01-30 15:38:26 -08:00
Ryan Houdek d9453ff639 OpcodeDispatcher: Implements CRC32 instruction
Now that the rest of the code matches behaviour, we just need to pass
this through.

Easy enough and get Horizon Zero Dawn running.
2022-01-30 15:38:26 -08:00
Ryan Houdek 70754991d1 CPUID: Fill out CPUID for SSE4.2 feature
Currently force disabled until the rest of SSE 4.2 is enabled
This is to remind us in the future that SSE4.2 can only be enabled in
CPUID with CRC32 instruction support.
2022-01-30 15:38:26 -08:00
Ryan Houdek 57ebfceb48 HostFeatures: Check for CRC32 op support
Available with CRC32 bit on Arm64 or SSE4.2 on x86-64
2022-01-30 15:38:26 -08:00
Ryan Houdek 9e224d2bb0 x86 JIT: Implements CRC32 op 2022-01-30 15:38:26 -08:00
Ryan Houdek e43bd04901 JITArm64: Implements CRC32 op 2022-01-30 15:38:26 -08:00
Ryan Houdek 48762e03a6 Interpreter: Implements CRC32 op 2022-01-30 15:38:26 -08:00
Ryan Houdek 858924309e IR: Implements CRC32 op 2022-01-29 23:33:53 -08:00
Ryan Houdek cab02d1e65 OpcodeDispatcher: Removes extraneous extract in VFCMP
We don't need to extract the element to compare it.
2022-01-28 22:22:34 -08:00
Ryan Houdek 2a64f80567 IR: Fixes some wrong offsets in passes
GPR and FPR ending offsets were off by one here. Just a quick fix.
2022-01-28 22:19:51 -08:00
Ryan Houdek 174ddea99d Softfloat: Stop doing special handling for FREM
This isn't correct and breaks games.
This makes the FREM and REM1 implementation the same.
While not 100% correct, it is still better than before.
New issues will be created to handle the differences in the future.

Fixes #1374.
Also fixes most of the HL2 issues, just not the seam issue.
2022-01-28 19:42:57 -08:00
Ryan Houdek 6fb0e3c0cf Softfloat: Allow x87 fallback for all ops 2022-01-28 19:42:25 -08:00
Ryan Houdek 2b044bbdf4 Disable fprem unittests
These are about to be broken
2022-01-28 19:39:03 -08:00
Mai M ea76de0fd2 Merge pull request #1533 from Sonicadvance1/revise_posix_tests
unittests: Revise POSIX tests known failures and disabled
2022-01-25 15:19:24 -05:00
Ryan Houdek 9c8642e0dc unittests: Revise POSIX tests known failures and disabled
Some of these behaviours have changed now, particularly around signal
handling.

Some things still fail now of course. But most everything is now
documented as to why it is failing or disabled.

Fixes #955
2022-01-25 11:41:45 -08:00
Ryan Houdek 13f35f7b79 Merge pull request #1530 from Sonicadvance1/rootfs_fetcher_fixes
FEXRootFSFetcher: Fixes some edge case behaviours
2022-01-25 10:29:55 -08:00
Ryan Houdek e2798e370e Merge pull request #1518 from Sonicadvance1/fix_signed_branch
JIT: Fixes signed displacement wraparound on 32-bit
2022-01-25 10:29:46 -08:00
Ryan Houdek d9149548b5 Merge pull request #1531 from Sonicadvance1/fix_sockopt
Linux: Fixes 32-bit getsockopt and setsockopt
2022-01-25 09:09:16 -08:00
Ryan Houdek 4823933f79 Linux: Fixes 32-bit getsockopt and setsockopt
On Set, we have four options that need to be converted.
On Get, we have two options that need to be converted.

This fixes a crash that Tomb Raider 2013 was having on launch.
2022-01-24 17:14:50 -08:00
Ryan Houdek ee04067424 FEXRootFSFetcher: Fixes some edge case behaviours
Makes curl do its continue feature to give the users the best chance of
downloading a rootfs. We don't need to restart the full file transfer on
failure. Helps people with slower connections.

On failure to download, asks the user if they want to retry the download
rather than just exiting with a weird error about hash failure.

Once the image is downloaded, now changes options depending on if
squashfuse or unsquashfs works.

Prevents the user from selecting a bad option and getting unexpected
behaviour. Ideally we would do a squashfs mount test as well for
platforms that don't have working FUSE, like termux. This is harder to
get right and its for an unsupported platform, so I'm not going to
invest more time with it.

Fixes #1525
Fixes #1526
Fixes #1527
2022-01-23 22:54:05 -08:00
Ryan Houdek e4aef26ef5 FEXRootFSFetcher: Adds helper namespace for tool checking
Location to check if curl, squashfuse, and unsquashfs are working.

unsquashfs is a bit more complex where it needs to parse the help output
to see if zstd is supported
2022-01-23 22:44:31 -08:00
Ryan Houdek a41dc8eafa FEXRootFSFetcher: Fix pipe redirecting
In the case of launching without stdout/stderr then redirection could
have these constants be a redirected FD that sits in the same fd number.

Use -2 to indicate no redirection.
Use -1 to indicate closing traditional stderr/stdout
The rest will indicate if stdout and stderr should be replaced as
normal.
Making sure not to close the incoming fds if they matched the
stdout/stderr FD numbers.
2022-01-23 22:41:39 -08:00
Ryan Houdek f41cd8deff OpcodeDispatcher: Renamed GetDynamicPC to GetRelocatedPC
For clarity.
2022-01-23 18:52:53 -08:00
Ryan Houdek 2a0c3cce30 Core: Have GetDynamicPC mask based on operating size
This ensures on 32-bit we overflow correctly under relocation.
2022-01-23 18:50:16 -08:00
Ryan Houdek e817f5d98c unittests: Adds 32-bit tests for signed displacement wraparound
A bit meta since it needs to JIT some minor code but easy enough.
Ensures something like #1517 won't happen again.
2022-01-23 18:38:44 -08:00
Ryan Houdek 8b8cda9b80 JIT: Fixes signed displacement wraparound on 32-bit
This cropped up mostly with multiblock and `jmp <signed displacement>`
This also happened with non multiblock `jcc <signed displacement>`

Due to how IR relocations occur, this needs to happen fairly late but
isn't a big deal.

Fixes #1517
2022-01-23 18:38:43 -08:00
Ryan Houdek 8e3893df07 Merge pull request #1523 from lioncash/vixl-update
Externals: Update vixl
2022-01-20 15:13:36 -08:00
lioncash 51b335914c github: Synchronize submodules before checking them out
Ensures that we don't get stale remotes.
2022-01-20 17:57:45 -05:00
lioncash 8835d57ae3 Arm64Emitter: Adjust XRegister to Register
With the updated API, we need to make use of Register as opposed to
XRegister in our arrays.
2022-01-20 16:42:04 -05:00
lioncash eba1b65fb0 Externals: Update vixl to updated branch
Now we have access to some SVE goodies.
2022-01-20 16:42:02 -05:00
Ryan Houdek a3a138ef7e Merge pull request #1520 from lioncash/vixl
External: Point vixl submodule towards FEX's fork
2022-01-14 13:40:18 -08:00
lioncash 62d9a494cd External: Point vixl submodule towards FEX's fork
This allows it to be managed by all organization members
2022-01-14 12:21:19 -05:00
Stefanos Kornilios Mitsis Poiitidis 6744a06a53 Merge pull request #1519 from Sonicadvance1/aarch64_single_instruction_opt
AArch64: Single instruction optimization for AESKeyGenAssist
2022-01-14 15:41:57 +02:00
Ryan Houdek 140e9824b7 Merge pull request #1516 from lioncash/fmt
externals: Update fmt to 8.1.1
2022-01-14 01:55:50 -08:00
Ryan Houdek bad84f61fa AArch64: Single instruction optimization for AESKeyGenAssist
No need to do adr when loads can do a 1MB offset loadstore
2022-01-14 01:49:21 -08:00
lioncash 2296126af3 externals: Update fmt to 8.1.1
Brings along a bunch of enhancements and ensures we always build against
the latest version.

Also fixes up a few issues that arose due to changes in fmt
2022-01-13 14:48:35 -05:00
Stefanos Kornilios Mitsis Poiitidis 0a8717d9a8 Merge pull request #1515 from Sonicadvance1/fix_ptest
OpcodeDispatcher: Fixes ptest flags calculation.
2022-01-13 10:38:38 +02:00
Ryan Houdek 4e2220c27f unittests: Adds ptest unit test to ensure correct flag setting
ptest wasn't correctly setting OF, SF, AF, and PF to zero until now.
Do a unit test to ensure correct behaviour here
2022-01-11 16:45:56 -08:00
Ryan Houdek c87e11cee9 OpcodeDispatcher: Fixes ptest flags calculation.
We were missing four flags that require setting zero.
2022-01-11 16:45:11 -08:00
Ryan Houdek 7768f6965a Merge pull request #1501 from Sonicadvance1/finish_siginfo_32bit
Linux: Handles the remaining 32-bit siginfo_t usage
2022-01-11 00:02:55 -08:00
Ryan Houdek 023aaaae0c Merge pull request #1499 from Sonicadvance1/resolve_rootfs_path_in_interpreter
FEXLoader: Resolve the absolute path to rootfs if possible
2022-01-11 00:02:27 -08:00
Stefanos Kornilios Mitsis Poiitidis a2aa9f3fc1 Merge pull request #1512 from Sonicadvance1/fix_ssa_id_print
IR: Fixes SSA ID printing
2022-01-11 09:09:06 +02:00
Ryan Houdek e46ec9a0ce IR: Fixes SSA ID printing
These should print as decimal. They were ending up as hex
2022-01-10 18:09:03 -08:00
Ryan Houdek 784cbdd973 Merge pull request #1500 from Sonicadvance1/rootfsfetch_check_curl
FEXRootFSFetcher: Check if curl is installed and fail before running
2022-01-10 16:23:05 -08:00
Ryan Houdek 6022715a9b Merge pull request #1510 from Sonicadvance1/fix_asan_cpuid
CPUID: Fixes ASAN problem with reading midr
2022-01-10 02:12:10 -08:00
Ryan Houdek 9eb5ba5ad1 Merge pull request #1509 from Sonicadvance1/fix_logserver_sync
SocketLogging: Fixes MsgHandler not syncing with Assert level
2022-01-10 02:12:01 -08:00
Ryan Houdek 82e5977709 Merge pull request #1506 from Sonicadvance1/fix_apitest_syscalls
APITests: Fixes InterruptableConditionVariable test to use the syscal…
2022-01-10 02:11:44 -08:00
Ryan Houdek 7c08b67dff Merge pull request #1504 from Sonicadvance1/fix_unittest_rootfs_define
unittests: Fixes ROOTFS needing to be defined prior to cmake
2022-01-10 02:11:35 -08:00
Ryan Houdek 609587f9ee Merge pull request #1503 from Sonicadvance1/implement_bcd_tests
unittests: Adds a BCD unit test
2022-01-10 02:11:07 -08:00
Ryan Houdek 5dda3a1599 Merge pull request #1497 from Sonicadvance1/fix_alternative_links
Linux: Fixes emulatedpath with symlink following
2022-01-10 02:10:54 -08:00
Ryan Houdek 73aaa4c3a6 CPUID: Fixes ASAN problem with reading midr
Needs to be a string_view for the MIDR for the StrConv helper to work in
this instance.
There is no null terminator character when reading from the file is why.
2022-01-10 01:21:56 -08:00
Ryan Houdek 3ba5371d36 SocketLogging: Fixes MsgHandler not syncing with Assert level
AssertHandler by default synchronizes but MsgHandler with Assert level
should also synchronize.

Fixes an issue where LogMan::Msg::AFmt wasn't syncing so the
FEXLogServer would never see the messages.
2022-01-10 01:16:53 -08:00
Ryan Houdek bf581decde Merge pull request #1507 from Sonicadvance1/fix_warnings
Fixes some of the warnings that cropped up
2022-01-10 01:12:39 -08:00
Ryan Houdek 250504502a Fixes some of the warnings that cropped up 2022-01-10 00:46:10 -08:00
Ryan Houdek a0e826feaf APITests: Fixes InterruptableConditionVariable test to use the syscall wrappers.
Fixes a build error on old Ubuntu
2022-01-09 22:00:57 -08:00
Ryan Houdek eb17edec05 Merge pull request #1502 from Sonicadvance1/fix_fexlog_server_message
FEXLogServer: Stop duplicating and dropping messages
2022-01-09 03:22:24 -08:00
Ryan Houdek 228aed98c7 unittests: Fixes ROOTFS needing to be defined prior to cmake
cmake will bake in the environment variable in to the build scripts.
Instead have the guest_test_runner fetch it at runtime.

This means if you forget to set ROOTFS prior to running cmake, you can
now set it afterwards and rerun with just ctest instead of a cmake
dance.

Fixes #315
2022-01-09 01:56:13 -08:00
Ryan Houdek 6ba4aec88e unittests: Adds a BCD unit test
Nothing really fantastical found here. Just that sub-precision results
weren't rounded correctly on store

Fixes #770
2022-01-09 01:36:26 -08:00
Ryan Houdek b8d6b2cd4a F80: Ensures BCDStore rounds to the current rounding mode
BCD storing will round any subprecision results depending on the current
rounding mode.
2022-01-09 01:35:28 -08:00
Ryan Houdek d4e2f42f90 FEXLogServer: Stop duplicating and dropping messages
In the case that multiple messages appearing in a single packet then we
were repeating the first message and dropping any subsequent messages.

Fixes #1496
2022-01-09 00:22:12 -08:00
Ryan Houdek 3055c23365 Linux: Handles the remaining 32-bit siginfo_t usage
Just need to translate them between 32-bit and 64-bit versions.

Fixes #1254
2022-01-09 00:03:17 -08:00
Ryan Houdek cb7feaefbb Types: Allows passing 64-bit host siginfo_t to 32-bit siginfo_t
Needed for waitid
2022-01-09 00:00:13 -08:00
Ryan Houdek 5a5a498ed6 FEXRootFSFetcher: Check if curl is installed and fail before running
Before doing anything that requires curl, actually check if it is
installed.
Then instruct the user to install curl before using.

Doesn't try installing curl itself since we don't have a clean way to
execute sudo from potentially GUI.

Fixes #1498
2022-01-08 21:35:29 -08:00
Ryan Houdek 285ed8f1e0 FEXRootFSFetcher: Adds new Exec function with stdout,stderr redirection
Just so we can test for applications without spamming terminal
2022-01-08 21:34:59 -08:00
Ryan Houdek a68da468a5 FEXRootFSFetcher: ExecAndWaitForResponse sign extend program result
Only the lower 8bits of the execve result is the program result.
Makes sure to sign extend it so -1 is a true -1 instead of 255
2022-01-08 21:33:24 -08:00
Ryan Houdek d59aa6874e FEXLoader: Resolve the absolute path to rootfs if possible
If the user passes in an absolute path then check to see if it exists in
the rootfs before executing.

Useful for launching applications directly out of the rootfs with
FEXInterpreter.

In the case that the absolute path doesn't exist in the rootfs then
fallback to the host system as usual
2022-01-07 03:42:41 -08:00
Ryan Houdek 19fd89d2bf Linux: Fixes emulatedpath with symlink following
Some syscalls support `AT_SYMLINK_NOFOLLOW` In these instances we need
to follow the symlink on a couple of syscalls.

Fixes executing wine using the basic wine path
eg:
FEXBash "wine dxcapsviewer.exe"
2022-01-07 02:26:26 -08:00
Mai M e36beb8dbe Merge pull request #1495 from Sonicadvance1/add_tune_arch
CMake: Adds TUNE_ARCH option
2022-01-05 18:09:26 -05:00
Ryan Houdek 70f447b265 CMake: Adds TUNE_ARCH option
I forgot about this option working for tuning arch on AArch64. This will
be used in PPA releases in the future. Will leave the previous option
since it can be used in testing.
2022-01-05 13:48:35 -08:00
Ryan Houdek a0026c92a8 Merge pull request #1494 from Seas0/main
ThunkLibs: Add meta data to libvulkan_device
2022-01-04 23:37:54 -08:00
Seas0 8fc5f66a5b ThunkLibs: Add meta data to libvulkan_device 2022-01-05 14:28:26 +08:00
498 changed files with 24311 additions and 9687 deletions

No files matched your search

+4 -2
View File
@@ -31,7 +31,9 @@ jobs:
- name : submodule checkout
# Need to update submodules
run: git submodule update --init --depth 1
run: |
git submodule sync --recursive
git submodule update --init --depth 1
- name: Clean Build Environment
run: rm -Rf ${{runner.workspace}}/build
@@ -49,7 +51,7 @@ jobs:
# Note the current convention is to use the -S and -B options here to specify source
# and build directories, but this is only available with CMake 3.13 and higher.
# The CMake binaries on the Github Actions machines are (as of this writing) 3.12
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True -DENABLE_INTERPRETER=True
- name: Build
working-directory: ${{runner.workspace}}/build
+5 -1
View File
@@ -1,7 +1,7 @@
[submodule "External/vixl"]
shallow = true
path = External/vixl
url = https://github.com/Sonicadvance1/vixl.git
url = https://github.com/FEX-Emu/vixl.git
[submodule "External/cpp-optparse"]
path = External/cpp-optparse
url = https://github.com/Sonicadvance1/cpp-optparse
@@ -45,3 +45,7 @@
[submodule "External/Catch2"]
path = External/Catch2
url = https://github.com/catchorg/Catch2.git
[submodule "External/robin-map"]
shallow = true
path = External/robin-map
url = https://github.com/Tessil/robin-map.git
+53 -38
View File
@@ -7,7 +7,8 @@ option(ENABLE_CLANG_FORMAT "Run clang format over the source" FALSE)
option(ENABLE_IWYU "Enables include what you use program" FALSE)
option(ENABLE_LTO "Enable LTO with compilation" TRUE)
option(ENABLE_XRAY "Enable building with LLVM X-Ray" FALSE)
option(ENABLE_LLD "Enable linking with LLD" FALSE)
option(ENABLE_LLD "Enable linking with lld" FALSE)
option(ENABLE_MOLD "Enable linking with mold" FALSE)
option(ENABLE_ASAN "Enables Clang ASAN" FALSE)
option(ENABLE_TSAN "Enables Clang TSAN" FALSE)
option(ENABLE_ASSERTIONS "Enables assertions in build" FALSE)
@@ -19,6 +20,9 @@ option(ENABLE_JEMALLOC "Enables jemalloc allocator" TRUE)
option(ENABLE_OFFLINE_TELEMETRY "Enables FEX offline telemetry" TRUE)
option(ENABLE_COMPILE_TIME_TRACE "Enables time trace compile option" FALSE)
option(ENABLE_LIBCXX "Enables LLVM libc++" FALSE)
option(ENABLE_INTERPRETER "Enables FEX's Interpreter" FALSE)
option(ENABLE_CCACHE "Enables ccache for compile caching" TRUE)
option(ENABLE_TERMUX_BUILD "Forces building for Termux on a non-Termux build machine" FALSE)
set (X86_C_COMPILER "x86_64-linux-gnu-gcc" CACHE STRING "c compiler for compiling x86 guest libs")
set (X86_CXX_COMPILER "x86_64-linux-gnu-g++" CACHE STRING "c++ compiler for compiling x86 guest libs")
@@ -26,6 +30,7 @@ set (DATA_DIRECTORY "${CMAKE_INSTALL_PREFIX}/share/fex-emu" CACHE PATH "global d
# These options are meant for package management
set (TUNE_CPU "native" CACHE STRING "Override the CPU the build is tuned for")
set (TUNE_ARCH "generic" CACHE STRING "Override the Arch the build is tuned for")
set (OVERRIDE_VERSION "detect" CACHE STRING "Override the FEX version in the format of <MMYY>{.<REV>}")
string(TOUPPER "${CMAKE_BUILD_TYPE}" CMAKE_BUILD_TYPE)
@@ -38,6 +43,11 @@ if (ENABLE_ASSERTIONS)
add_definitions(-DASSERTIONS_ENABLED=1)
endif()
if (ENABLE_INTERPRETER)
message(STATUS "Interpreter enabled")
add_definitions(-DINTERPRETER_ENABLED=1)
endif()
set(CMAKE_CXX_STANDARD 20)
set(CMAKE_EXPORT_COMPILE_COMMANDS ON)
set(CMAKE_RUNTIME_OUTPUT_DIRECTORY ${CMAKE_BINARY_DIR}/Bin)
@@ -72,10 +82,12 @@ if (CMAKE_SYSTEM_PROCESSOR MATCHES "aarch64")
add_definitions(-D_M_ARM_64=1)
endif()
find_program(CCACHE_PROGRAM ccache)
if(CCACHE_PROGRAM)
message(STATUS "CCache enabled")
set_property(GLOBAL PROPERTY RULE_LAUNCH_COMPILE "${CCACHE_PROGRAM}")
if (ENABLE_CCACHE)
find_program(CCACHE_PROGRAM ccache)
if(CCACHE_PROGRAM)
message(STATUS "CCache enabled")
set_property(GLOBAL PROPERTY RULE_LAUNCH_COMPILE "${CCACHE_PROGRAM}")
endif()
endif()
if (ENABLE_XRAY)
@@ -89,9 +101,14 @@ if (ENABLE_COMPILE_TIME_TRACE)
endif()
set (PTHREAD_LIB pthread)
if (ENABLE_LLD)
if (ENABLE_LLD AND ENABLE_MOLD)
message (FATAL_ERROR "Cannot enable both lld and mold")
elseif (ENABLE_LLD)
set (LD_OVERRIDE "-fuse-ld=lld")
link_libraries(${LD_OVERRIDE})
add_link_options(${LD_OVERRIDE})
elseif (ENABLE_MOLD)
add_link_options("-fuse-ld=mold")
endif()
if (ENABLE_LIBCXX)
@@ -105,9 +122,16 @@ if (NOT ENABLE_OFFLINE_TELEMETRY)
add_definitions(-DFEX_DISABLE_TELEMETRY=1)
endif()
if(DEFINED ENV{TERMUX_VERSION} OR ENABLE_TERMUX_BUILD)
add_definitions(-DTERMUX_BUILD=1)
set(TERMUX_BUILD 1)
# Termux doesn't support Jemalloc due to bad interactions between emutls, jemalloc, and scudo
set(ENABLE_JEMALLOC FALSE)
endif()
if (ENABLE_STATIC_PIE)
if (_M_ARM_64 AND ENABLE_LLD)
message (FATAL_ERROR "Static linking does not currently work with AArch64+LLD. Use GNU ld for now.")
message (FATAL_ERROR "Static linking does not currently work with AArch64+lld. Use GNU ld for now.")
endif()
file(WRITE ${PROJECT_BINARY_DIR}/CMakeFiles/CMakeTmp/Determine_iplt.c
@@ -258,6 +282,8 @@ set (CMAKE_LINKER_FLAGS_RELWITHDEBINFO "${CMAKE_LINKER_FLAGS_RELWITHDEBINFO} -fn
set (CMAKE_CXX_FLAGS_RELEASE "${CMAKE_CXX_FLAGS_RELEASE} -fomit-frame-pointer")
set (CMAKE_LINKER_FLAGS_RELEASE "${CMAKE_LINKER_FLAGS_RELEASE} -fomit-frame-pointer")
include_directories(External/robin-map/include/)
add_subdirectory(External/vixl/)
include_directories(External/vixl/src/)
@@ -268,13 +294,9 @@ endif()
find_package(PkgConfig REQUIRED)
find_package(Python 3.0 REQUIRED COMPONENTS Interpreter)
pkg_check_modules(XXHASH libxxhash>=0.8.0 QUIET)
if (NOT XXHASH_FOUND)
message(STATUS "xxHash not found. Using Externals")
add_subdirectory(External/xxhash/)
include_directories(External/xxhash/)
endif()
add_subdirectory(External/xxhash/)
include_directories(External/xxhash/)
add_definitions(-Wno-trigraphs)
add_definitions(-DGLOBAL_DATA_DIRECTORY="${DATA_DIRECTORY}/")
@@ -335,6 +357,15 @@ if(ENABLE_WERROR OR ENABLE_STRICT_WERROR)
endif()
endif()
if (NOT TUNE_ARCH STREQUAL "generic")
check_cxx_compiler_flag("-march=${TUNE_ARCH}" COMPILER_SUPPORTS_ARCH_TYPE)
if(COMPILER_SUPPORTS_ARCH_TYPE)
add_compile_options("-march=${TUNE_ARCH}")
else()
message(FATAL_ERROR "Trying to compile arch type '${TUNE_ARCH}' but the compiler doesn't support this")
endif()
endif()
if (TUNE_CPU STREQUAL "native")
if(_M_ARM_64)
if (CMAKE_CXX_COMPILER_VERSION VERSION_GREATER_EQUAL 999999.0)
@@ -468,30 +499,14 @@ endif()
if (BUILD_THUNKS)
add_subdirectory(ThunkLibs/Generator)
# Thunk targets for both host libraries and IDE integration
add_subdirectory(ThunkLibs/HostLibs)
# Thunk targets for IDE integration of guest code, only
add_subdirectory(ThunkLibs/GuestLibs)
# Thunk targets for guest libraries
include(ExternalProject)
ExternalProject_Add(host-libs
PREFIX host-libs
SOURCE_DIR "${CMAKE_CURRENT_SOURCE_DIR}/ThunkLibs/HostLibs"
BINARY_DIR "Host"
CMAKE_ARGS
"-DCMAKE_BUILD_TYPE=${CMAKE_BUILD_TYPE}"
"-DCMAKE_INSTALL_PREFIX=${CMAKE_INSTALL_PREFIX}"
"-DGENERATOR_EXE=$<TARGET_FILE:thunkgen>"
INSTALL_COMMAND ""
BUILD_ALWAYS ON
DEPENDS thunkgen
)
install(
CODE "MESSAGE(\"-- Installing: host-libs\")"
CODE "
EXECUTE_PROCESS(COMMAND ${CMAKE_COMMAND} --build . --target ThunkHostsInstall
WORKING_DIRECTORY ${CMAKE_BINARY_DIR}/Host
)"
DEPENDS host-libs
)
ExternalProject_Add(guest-libs
PREFIX guest-libs
SOURCE_DIR "${CMAKE_CURRENT_SOURCE_DIR}/ThunkLibs/GuestLibs"
@@ -511,7 +526,7 @@ if (BUILD_THUNKS)
install(
CODE "MESSAGE(\"-- Installing: guest-libs\")"
CODE "
EXECUTE_PROCESS(COMMAND ${CMAKE_COMMAND} --build . --target ThunkGuestsInstall
EXECUTE_PROCESS(COMMAND ${CMAKE_COMMAND} --build . --target install
WORKING_DIRECTORY ${CMAKE_BINARY_DIR}/Guest
)"
DEPENDS guest-libs
+5
View File
@@ -0,0 +1,5 @@
{
"Config": {
"AdditionalArguments": "--no-sandbox"
}
}
+2 -12
View File
@@ -18,14 +18,8 @@ This project aims to provide a fast and functional x86-64 emulation library that
* Portable library implementation in order to support easy integration in to applications
### Target Host Architecture
The target host architecture for this library is AArch64. Specifically the ARMv8.1 version or newer.
The CPU IR is designed with AArch64 in mind but there is a desire to run the recompiled code on other architectures as well.
Multiple architecture support is desired for easier bringup and debugging, performance isn't as much of a priority there (ex. x86-64(guest) translated to x86-64(host))
### Not currently goals but will be in the future
* 32bit x86 support
* This will be a desire in the future, but to lower the amount of work required, decided to push this off for now.
* Integration in to WINE
* Later generation of x86-64 instruction sets
* Including AVX, F16C, XOP, FMA, AVX2, etc
The CPU IR is designed with AArch64 in mind but should allow for other architectures as well.
x86-64 host support is available for ease of development, but is not a priority.
### Not desired
* Kernel space emulation
* CPL0-2 emulation
@@ -33,7 +27,3 @@ Multiple architecture support is desired for easier bringup and debugging, perfo
* IRQs
* SVM
* "Cycle Accurate" emulation
### Dependencies
* clang-tidy if you want to ensure the code stays tidy
* cmake
* A C++17 compliant compiler (There are assumptions made about using Clang and LTO)
+6 -43
View File
@@ -7,17 +7,12 @@ OpClasses = collections.OrderedDict()
def get_ir_classes(ops, defines):
global OpClasses
for op_key, op_vals in ops.items():
if not ("Last" in op_vals):
OpClass = "#Unknown"
for op_class, opslist in ops.items():
if not (op_class in OpClasses):
OpClasses[op_class] = []
if ("OpClass" in op_vals):
OpClass = op_vals["OpClass"]
if not (OpClass in OpClasses):
OpClasses[OpClass] = []
OpClasses[OpClass].append([op_key, op_vals])
for op, op_val in opslist.items():
OpClasses[op_class].append([op, op_val])
# Sort the dictionary after we are done parsing it
OpClasses = collections.OrderedDict(sorted(OpClasses.items()))
@@ -38,41 +33,9 @@ def print_ir_ops():
op_key = op[0]
op_vals = op[1]
output_file.write("## %s\n" % (op_key))
HasDest = ("HasDest" in op_vals and op_vals["HasDest"] == True)
HasSSAArgs = ("SSAArgs" in op_vals and len(op_vals["SSAArgs"]) > 0)
HasSSAArgNames = "SSANames" in op_vals
HasArgs = "Args" in op_vals
SSAArgsCount = 0
ArgCount = 0
if (HasSSAArgs):
SSAArgsCount = int(op_vals["SSAArgs"])
if (HasArgs):
ArgCount = len(op_vals["Args"])
TotalArgsCount = SSAArgsCount + (ArgCount / 2)
output_file.write(">")
if (HasDest):
output_file.write("%dest = ")
output_file.write("%s " % op_key)
ArgComma = (", ", "")
if (HasSSAArgs):
for i in range(0, SSAArgsCount):
FinalArg = (i + 1) == TotalArgsCount
if (HasSSAArgNames):
output_file.write("%%%s%s" % (op_vals["SSANames"][i], ArgComma[FinalArg]))
else:
output_file.write("%%ssa%d%s" % (i, ArgComma[FinalArg]))
if (HasArgs):
Args = op_vals["Args"]
for i in range(0, ArgCount, 2):
FinalArg = ((i / 2) + SSAArgsCount + 1) == TotalArgsCount
data_type = Args[i]
data_name = Args[i + 1]
output_file.write("\<%s %s\>%s" % (data_type, data_name, ArgComma[FinalArg]))
output_file.write(op_key)
output_file.write("\n\n")
Vendored Regular → Executable
+469 -382
View File
File diff suppressed because it is too large. Load diff
+26 -15
View File
@@ -80,16 +80,19 @@ set (SRCS
Interface/Context/Context.cpp
Interface/Core/LookupCache.cpp
Interface/Core/BlockSamplingData.cpp
Interface/Core/CompileService.cpp
Interface/Core/Core.cpp
Interface/Core/CPUID.cpp
Interface/Core/Frontend.cpp
Interface/Core/GdbServer.cpp
Interface/Core/HostFeatures.cpp
Interface/Core/ObjectCache/JobHandling.cpp
Interface/Core/ObjectCache/NamedRegionObjectHandler.cpp
Interface/Core/ObjectCache/ObjectCacheService.cpp
Interface/Core/OpcodeDispatcher/Crypto.cpp
Interface/Core/OpcodeDispatcher/Flags.cpp
Interface/Core/OpcodeDispatcher/Vector.cpp
Interface/Core/OpcodeDispatcher/X87.cpp
Interface/Core/OpcodeDispatcher/X87F64.cpp
Interface/Core/OpcodeDispatcher.cpp
Interface/Core/SignalDelegator.cpp
Interface/Core/X86Tables.cpp
@@ -100,19 +103,7 @@ set (SRCS
Interface/Core/Dispatcher/Dispatcher.cpp
Interface/Core/Dispatcher/X86Dispatcher.cpp
Interface/Core/Dispatcher/Arm64Dispatcher.cpp
Interface/Core/Interpreter/InterpreterCore.cpp
Interface/Core/Interpreter/InterpreterOps.cpp
Interface/Core/Interpreter/ALUOps.cpp
Interface/Core/Interpreter/AtomicOps.cpp
Interface/Core/Interpreter/BranchOps.cpp
Interface/Core/Interpreter/ConversionOps.cpp
Interface/Core/Interpreter/EncryptionOps.cpp
Interface/Core/Interpreter/F80Ops.cpp
Interface/Core/Interpreter/FlagOps.cpp
Interface/Core/Interpreter/MemoryOps.cpp
Interface/Core/Interpreter/MiscOps.cpp
Interface/Core/Interpreter/MoveOps.cpp
Interface/Core/Interpreter/VectorOps.cpp
Interface/Core/Interpreter/InterpreterFallbacks.cpp
Interface/Core/X86Tables/BaseTables.cpp
Interface/Core/X86Tables/DDDTables.cpp
Interface/Core/X86Tables/EVEXTables.cpp
@@ -152,6 +143,23 @@ set (SRCS
Utils/Threads.cpp
)
if (ENABLE_INTERPRETER)
list(APPEND SRCS
Interface/Core/Interpreter/InterpreterCore.cpp
Interface/Core/Interpreter/InterpreterOps.cpp
Interface/Core/Interpreter/ALUOps.cpp
Interface/Core/Interpreter/AtomicOps.cpp
Interface/Core/Interpreter/BranchOps.cpp
Interface/Core/Interpreter/ConversionOps.cpp
Interface/Core/Interpreter/EncryptionOps.cpp
Interface/Core/Interpreter/F80Ops.cpp
Interface/Core/Interpreter/FlagOps.cpp
Interface/Core/Interpreter/MemoryOps.cpp
Interface/Core/Interpreter/MiscOps.cpp
Interface/Core/Interpreter/MoveOps.cpp
Interface/Core/Interpreter/VectorOps.cpp)
endif()
if(_M_ARM_64)
list(APPEND SRCS
Interface/Core/ArchHelpers/Arm64.cpp)
@@ -179,7 +187,9 @@ if (ENABLE_JIT_X86_64)
Interface/Core/JIT/x86_64/MemoryOps.cpp
Interface/Core/JIT/x86_64/MiscOps.cpp
Interface/Core/JIT/x86_64/MoveOps.cpp
Interface/Core/JIT/x86_64/VectorOps.cpp)
Interface/Core/JIT/x86_64/VectorOps.cpp
Interface/Core/JIT/x86_64/x64Relocations.cpp
)
list(APPEND DEFINES -DJIT_X86_64)
endif()
@@ -324,6 +334,7 @@ function(AddDefaultOptionsToTarget Name)
-Wno-trigraphs
-ffunction-sections
-fwrapv
)
if (GCC_COLOR)
+12 -4
View File
@@ -23,7 +23,7 @@ namespace FEXCore {
// Linux perf format is very straightforward
// `<HostPtr> <Size> <Name>\n`
fmt::print(fp.get(), "{:x} {:x} JIT_0x{:x}_{:x}\n", HostAddr, CodeSize, GuestAddr, HostAddr);
fmt::print(fp.get(), "{} {:x} JIT_0x{:x}_{}\n", HostAddr, CodeSize, GuestAddr, HostAddr);
}
void JITSymbols::Register(const void *HostAddr, uint32_t CodeSize, std::string_view Name) {
@@ -31,7 +31,15 @@ namespace FEXCore {
// Linux perf format is very straightforward
// `<HostPtr> <Size> <Name>\n`
fmt::print(fp.get(), "{:x} {:x} {}_{:x}\n", HostAddr, CodeSize, Name, HostAddr);
fmt::print(fp.get(), "{} {:x} {}_{}\n", HostAddr, CodeSize, Name, HostAddr);
}
void JITSymbols::Register(const void *HostAddr, uint32_t CodeSize, std::string_view Name, uintptr_t Offset) {
if (!fp) return;
// Linux perf format is very straightforward
// `<HostPtr> <Size> <Name>\n`
fmt::print(fp.get(), "{} {:x} {}+0x{:x} ({})\n", HostAddr, CodeSize, Name, Offset, HostAddr);
}
void JITSymbols::RegisterNamedRegion(const void *HostAddr, uint32_t CodeSize, std::string_view Name) {
@@ -39,7 +47,7 @@ namespace FEXCore {
// Linux perf format is very straightforward
// `<HostPtr> <Size> <Name>\n`
fmt::print(fp.get(), "{:x} {:x} {}\n", HostAddr, CodeSize, Name);
fmt::print(fp.get(), "{} {:x} {}\n", HostAddr, CodeSize, Name);
}
void JITSymbols::RegisterJITSpace(const void *HostAddr, uint32_t CodeSize) {
@@ -47,7 +55,7 @@ namespace FEXCore {
// Linux perf format is very straightforward
// `<HostPtr> <Size> <Name>\n`
fmt::print(fp.get(), "{:x} {:x} FEXJIT\n", HostAddr, CodeSize);
fmt::print(fp.get(), "{} {:x} FEXJIT\n", HostAddr, CodeSize);
}
} // namespace FEXCore
+1
View File
@@ -13,6 +13,7 @@ public:
void Register(const void *HostAddr, uint64_t GuestAddr, uint32_t CodeSize);
void Register(const void *HostAddr, uint32_t CodeSize, std::string_view Name);
void Register(const void *HostAddr, uint32_t CodeSize, std::string_view Name, uintptr_t Offset);
void RegisterNamedRegion(const void *HostAddr, uint32_t CodeSize, std::string_view Name);
void RegisterJITSpace(const void *HostAddr, uint32_t CodeSize);
+269 -9
View File
@@ -57,51 +57,183 @@ struct X80SoftFloat {
// Ops
static X80SoftFloat FADD(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm (R"(
fninit;
fldt %[rhs]; # st1
fldt %[lhs]; # st0
faddp;
fstpt %[result];
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
, [rhs] "m" (rhs)
: "st", "st(1)");
return Result;
#else
return extF80_add(lhs, rhs);
#endif
}
static X80SoftFloat FSUB(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm (R"(
fninit;
fldt %[rhs]; # st1
fldt %[lhs]; # st0
fsubp;
fstpt %[result];
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
, [rhs] "m" (rhs)
: "st", "st(1)");
return Result;
#else
return extF80_sub(lhs, rhs);
#endif
}
static X80SoftFloat FMUL(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm (R"(
fninit;
fldt %[rhs]; # st1
fldt %[lhs]; # st0
fmulp;
fstpt %[result];
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
, [rhs] "m" (rhs)
: "st", "st(1)");
return Result;
#else
return extF80_mul(lhs, rhs);
#endif
}
static X80SoftFloat FDIV(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm (R"(
fninit;
fldt %[rhs]; # st1
fldt %[lhs]; # st0
fdivp;
fstpt %[result];
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
, [rhs] "m" (rhs)
: "st", "st(1)");
return Result;
#else
return extF80_div(lhs, rhs);
#endif
}
static X80SoftFloat FREM(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
X80SoftFloat Rem = extF80_rem(lhs, rhs);
if (SignBit(Rem)) {
Rem = extF80_add(Rem, rhs);
}
else {
Rem.Sign = SignBit(lhs);
}
#if defined(DEBUG_X86_FLOAT)
BIGFLOAT Result;
asm (R"(
fninit;
fldt %[rhs]; # st1
fldt %[lhs]; # st0
fprem;
fstpt %[result];
ffreep %%st(0);
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
, [rhs] "m" (rhs)
: "st", "st(1)");
return Rem;
return Result;
#else
return extF80_rem(lhs, rhs);
#endif
}
static X80SoftFloat FREM1(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
#if defined(DEBUG_X86_FLOAT)
BIGFLOAT Result;
asm (R"(
fninit;
fldt %[rhs]; # st1
fldt %[lhs]; # st0
fprem1;
fstpt %[result];
ffreep %%st(0);
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
, [rhs] "m" (rhs)
: "st", "st(1)");
return Result;
#else
return extF80_rem(lhs, rhs);
#endif
}
static X80SoftFloat FRNDINT(X80SoftFloat const &lhs) {
return extF80_roundToInt(lhs, softfloat_roundingMode, false);
}
static X80SoftFloat FRNDINT(X80SoftFloat const &lhs, uint_fast8_t RoundMode) {
return extF80_roundToInt(lhs, RoundMode, false);
}
static X80SoftFloat FXTRACT_SIG(X80SoftFloat const &lhs) {
#if defined(DEBUG_X86_FLOAT)
BIGFLOAT Result;
asm (R"(
fninit;
fldt %[lhs]; # st0
fxtract;
fstpt %[result];
ffreep %%st(0);
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
: "st", "st(1)");
return Result;
#else
X80SoftFloat Tmp = lhs;
Tmp.Exponent = 0x3FFF;
Tmp.Sign = lhs.Sign;
return Tmp;
#endif
}
static X80SoftFloat FXTRACT_EXP(X80SoftFloat const &lhs) {
#if defined(DEBUG_X86_FLOAT)
BIGFLOAT Result;
asm (R"(
fninit;
fldt %[lhs]; # st0
fxtract;
ffreep %%st(0);
fstpt %[result];
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
: "st", "st(1)");
return Result;
#else
int32_t TrueExp = lhs.Exponent - ExponentBias;
return i32_to_extF80(TrueExp);
#endif
}
static void FCMP(X80SoftFloat const &lhs, X80SoftFloat const &rhs, bool *eq, bool *lt, bool *nan) {
@@ -112,61 +244,189 @@ struct X80SoftFloat {
static X80SoftFloat FSCALE(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
WARN_ONCE_FMT("x87: Application used FSCALE which may have accuracy problems");
X80SoftFloat Int = FRNDINT(rhs);
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm (R"(
fninit;
fldt %[rhs]; # st1
fldt %[lhs]; # st0
fscale; # st0 = st0 * 2^(rdint(st1))
fstpt %[result];
ffreep %%st(0);
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
, [rhs] "m" (rhs)
: "st", "st(1)");
return Result;
#else
X80SoftFloat Int = FRNDINT(rhs, softfloat_round_minMag);
BIGFLOAT Src2_d = Int;
Src2_d = exp2l(Src2_d);
X80SoftFloat Src2_X80 = Src2_d;
X80SoftFloat Result = extF80_mul(lhs, Src2_X80);
return Result;
#endif
}
static X80SoftFloat F2XM1(X80SoftFloat const &lhs) {
WARN_ONCE_FMT("x87: Application used F2XM1 which may have accuracy problems");
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm (R"(
fninit;
fldt %[lhs]; # st0
f2xm1; # st0 = 2^st(0) - 1
fstpt %[result];
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
: "st");
return Result;
#else
BIGFLOAT Src1_d = lhs;
BIGFLOAT Result = exp2l(Src1_d);
Result -= 1.0;
return Result;
#endif
}
static X80SoftFloat FYL2X(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
WARN_ONCE_FMT("x87: Application used FYL2X which may have accuracy problems");
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm (R"(
fninit;
fldt %[rhs]; # st(1)
fldt %[lhs]; # st(0)
fyl2x; # st(1) * log2l(st(0))
fstpt %[result];
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
, [rhs] "m" (rhs)
: "st", "st(1)");
return Result;
#else
BIGFLOAT Src1_d = lhs;
BIGFLOAT Src2_d = rhs;
BIGFLOAT Tmp = Src2_d * log2l(Src1_d);
return Tmp;
#endif
}
static X80SoftFloat FATAN(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
WARN_ONCE_FMT("x87: Application used FATAN which may have accuracy problems");
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm (R"(
fninit;
fldt %[lhs];
fldt %[rhs];
fpatan;
fstpt %[result];
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
, [rhs] "m" (rhs)
: "st", "st(1)");
return Result;
#else
BIGFLOAT Src1_d = lhs;
BIGFLOAT Src2_d = rhs;
BIGFLOAT Tmp = atan2l(Src1_d, Src2_d);
return Tmp;
#endif
}
static X80SoftFloat FTAN(X80SoftFloat const &lhs) {
WARN_ONCE_FMT("x87: Application used FTAN which may have accuracy problems");
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm (R"(
fninit;
fldt %[lhs]; # st0
fptan;
ffreep %%st(0);
fstpt %[result];
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
: "st");
return Result;
#else
BIGFLOAT Src_d = lhs;
Src_d = tanl(Src_d);
return Src_d;
#endif
}
static X80SoftFloat FSIN(X80SoftFloat const &lhs) {
WARN_ONCE_FMT("x87: Application used FSIN which may have accuracy problems");
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm (R"(
fninit;
fldt %[lhs]; # st0
fsin;
fstpt %[result];
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
: "st");
return Result;
#else
BIGFLOAT Src_d = lhs;
Src_d = sinl(Src_d);
return Src_d;
#endif
}
static X80SoftFloat FCOS(X80SoftFloat const &lhs) {
WARN_ONCE_FMT("x87: Application used FCOS which may have accuracy problems");
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm (R"(
fninit;
fldt %[lhs]; # st0
fcos;
fstpt %[result];
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
: "st");
return Result;
#else
BIGFLOAT Src_d = lhs;
Src_d = cosl(Src_d);
return Src_d;
#endif
}
static X80SoftFloat FSQRT(X80SoftFloat const &lhs) {
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm (R"(
fninit;
fldt %[lhs]; # st0
fsqrt;
fstpt %[result];
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
: "st");
return Result;
#else
return extF80_sqrt(lhs);
#endif
}
operator float() const {
+29
View File
@@ -0,0 +1,29 @@
#pragma once
#include <string>
namespace FEXCore::StringUtils {
// Trim the left side of the string of whitespace and new lines
[[maybe_unused]] static std::string LeftTrim(std::string String, std::string TrimTokens = " \t\n\r") {
size_t pos = std::string::npos;
if ((pos = String.find_first_not_of(TrimTokens)) != std::string::npos) {
String.erase(0, pos);
}
return String;
}
// Trim the right side of the string of whitespace and new lines
[[maybe_unused]] static std::string RightTrim(std::string String, std::string TrimTokens = " \t\n\r") {
size_t pos = std::string::npos;
if ((pos = String.find_last_not_of(TrimTokens)) != std::string::npos) {
String.erase(String.begin() + pos + 1, String.end());
}
return String;
}
// Trim both the left and right of the string of whitespace and new lines
[[maybe_unused]] static std::string Trim(std::string String, std::string TrimTokens = " \t\n\r") {
return RightTrim(LeftTrim(String, TrimTokens), TrimTokens);
}
}
+18 -25
View File
@@ -1,4 +1,5 @@
#include "Common/StringConv.h"
#include "Common/StringUtils.h"
#include "Common/Paths.h"
#include "Utils/FileLoading.h"
@@ -370,29 +371,6 @@ namespace JSON {
return {};
}
std::string ltrim(std::string String) {
size_t pos = std::string::npos;
if ((pos = String.find_first_not_of(" \t\n\r")) != std::string::npos) {
String.erase(0, pos);
}
return String;
}
std::string rtrim(std::string String) {
size_t pos = std::string::npos;
if ((pos = String.find_last_not_of(" \t\n\r")) != std::string::npos) {
String.erase(String.begin() + pos + 1, String.end());
}
return String;
}
std::string trim(std::string String) {
return rtrim(ltrim(String));
}
std::string FindContainerPrefix() {
// We only support pressure-vessel at the moment
const static std::string ContainerManager = "/run/host/container-manager";
@@ -401,7 +379,7 @@ namespace JSON {
if (FEXCore::FileLoading::LoadFile(Manager, ContainerManager)) {
// Trim the whitespace, may contain a newline
std::string ManagerStr = Manager.data();
ManagerStr = trim(ManagerStr);
ManagerStr = FEXCore::StringUtils::Trim(ManagerStr);
if (strncmp(ManagerStr.data(), "pressure-vessel", Manager.size()) == 0) {
// We are running inside of pressure vessel
// Our $CMAKE_INSTALL_PREFIX paths are now inside of /run/host/$CMAKE_INSTALL_PREFIX
@@ -434,12 +412,27 @@ namespace JSON {
#else
constexpr uint32_t MaxCoreNumber = 1;
#endif
if (Core > MaxCoreNumber) {
#ifdef INTERPRETER_ENABLED
constexpr uint32_t MinCoreNumber = 0;
#else
constexpr uint32_t MinCoreNumber = 1;
#endif
if (Core > MaxCoreNumber || Core < MinCoreNumber) {
// Sanitize the core option by setting the core to the JIT if invalid
FEXCore::Config::EraseSet(FEXCore::Config::CONFIG_CORE, std::to_string(FEXCore::Config::CONFIG_IRJIT));
}
}
if (FEXCore::Config::Exists(FEXCore::Config::CONFIG_CACHEOBJECTCODECOMPILATION)) {
FEX_CONFIG_OPT(CacheObjectCodeCompilation, CACHEOBJECTCODECOMPILATION);
FEX_CONFIG_OPT(Core, CORE);
if (CacheObjectCodeCompilation() && Core() == FEXCore::Config::CONFIG_INTERPRETER) {
// If running the interpreter then disable cache code compilation
FEXCore::Config::Erase(FEXCore::Config::CONFIG_CACHEOBJECTCODECOMPILATION);
}
}
std::string ContainerPrefix { FindContainerPrefix() };
auto ExpandPathIfExists = [&ContainerPrefix](FEXCore::Config::ConfigOption Config, std::string PathName) {
auto NewPath = ExpandPath(ContainerPrefix, PathName);
+30 -4
View File
@@ -38,6 +38,17 @@
"Number of physical hardware threads to tell the process we have.",
"0 will auto detect."
]
},
"CacheObjectCodeCompilation": {
"Type": "uint32",
"Default": "FEXCore::Config::ConfigObjectCodeHandler::CONFIG_NONE",
"TextDefault": "none",
"Choices": [ "none", "read", "readwrite" ],
"ArgumentHandler": "CacheObjectCodeHandler",
"Desc": [
"Cache JIT object code to drive.",
"Allows JIT code to be shared between applications"
]
}
},
"Emulation": {
@@ -104,6 +115,13 @@
"This can be useful for setting environment variables that thunks can pick up.",
"Typically isn't necessary since the guest libc isn't thunked. But is possible."
]
},
"AdditionalArguments": {
"Type": "strarray",
"Default": "",
"Desc": [
"Allows the user to pass additional arguments to the application"
]
}
},
"Debug": {
@@ -223,14 +241,15 @@
"Hacks": {
"SMCChecks": {
"Type": "uint8",
"Default": "FEXCore::Config::CONFIG_SMC_MMAN",
"TextDefault": "mman",
"Default": "FEXCore::Config::CONFIG_SMC_MTRACK",
"TextDefault": "mtrack",
"ArgumentHandler": "SMCCheckHandler",
"Desc": [
"Checks code for modification before execution.",
"\tnone: No checks",
"\tmman: Invalidate on mmap, mprotect, munmap",
"\tfull: Validate code before every run (slow)"
"\tmtrack: Page tracking based invalidation",
"\tfull: Validate code before every run (slow)",
"\tmman: Invalidate on mmap, mprotect, munmap (deprecated, use mtrack)"
]
},
"TSOEnabled": {
@@ -241,6 +260,13 @@
"Highly likely to break any multithreaded application if disabled."
]
},
"X87ReducedPrecision": {
"Type": "bool",
"Default": "false",
"Desc": [
"Emulates X87 floating point using 64-bit precision. This reduces emulation accuracy and may result in rendering bugs."
]
},
"ABILocalFlags": {
"Type": "bool",
"Default": "false",
+5 -5
View File
@@ -149,7 +149,7 @@ namespace FEXCore::Context {
void CleanupAfterFork(FEXCore::Context::Context *CTX, FEXCore::Core::InternalThreadState *Thread) {
CTX->CleanupAfterFork(Thread);
}
void SetSignalDelegator(FEXCore::Context::Context *CTX, FEXCore::SignalDelegator *SignalDelegation) {
CTX->SignalDelegation = SignalDelegation;
}
@@ -186,11 +186,11 @@ namespace FEXCore::Context {
CTX->WriteFilesWithCode(Writer);
}
void AddNamedRegion(FEXCore::Context::Context *CTX, uintptr_t Base, uintptr_t Length, uintptr_t Offset, const std::string& Name) {
return CTX->AddNamedRegion(Base, Length, Offset, Name);
IR::AOTIRCacheEntry *LoadAOTIRCacheEntry(FEXCore::Context::Context *CTX, const std::string &Name) {
return CTX->LoadAOTIRCacheEntry(Name);
}
void RemoveNamedRegion(FEXCore::Context::Context *CTX, uintptr_t Base, uintptr_t Length) {
return CTX->RemoveNamedRegion(Base, Length);
void UnloadAOTIRCacheEntry(FEXCore::Context::Context *CTX, IR::AOTIRCacheEntry *Entry) {
return CTX->UnloadAOTIRCacheEntry(Entry);
}
namespace Debug {
+27 -17
View File
@@ -4,6 +4,7 @@
#include "Interface/Core/CPUID.h"
#include "Interface/Core/HostFeatures.h"
#include "Interface/Core/X86HelperGen.h"
#include "Interface/Core/ObjectCache/ObjectCacheService.h"
#include "Interface/IR/AOTIR.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/Context.h>
@@ -34,6 +35,10 @@ class CodeLoader;
class ThunkHandler;
class GdbServer;
namespace CodeSerialize {
class CodeObjectSerializeService;
}
namespace CPU {
class Arm64JITCore;
class X86JITCore;
@@ -98,6 +103,8 @@ namespace FEXCore::Context {
FEX_CONFIG_OPT(LibraryJITNaming, LIBRARYJITNAMING);
FEX_CONFIG_OPT(BlockJITNaming, BLOCKJITNAMING);
FEX_CONFIG_OPT(ParanoidTSO, PARANOIDTSO);
FEX_CONFIG_OPT(CacheObjectCodeCompilation, CACHEOBJECTCODECOMPILATION);
FEX_CONFIG_OPT(x87ReducedPrecision, X87REDUCEDPRECISION);
} Config;
using IntCallbackReturn = FEX_NAKED void(*)(FEXCore::Core::InternalThreadState *Thread, volatile void *Host_RSP);
@@ -155,11 +162,13 @@ namespace FEXCore::Context {
void RegisterHostSignalHandler(int Signal, HostSignalDelegatorFunction Func, bool Required);
void RegisterFrontendHostSignalHandler(int Signal, HostSignalDelegatorFunction Func, bool Required);
static void RemoveCodeEntry(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP);
// Must be called from owning thread
static void RemoveThreadCodeEntry(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP);
// Wrapper which takes CpuStateFrame instead of InternalThreadState
static void RemoveCodeEntryFromJit(FEXCore::Core::CpuStateFrame *Frame, uint64_t GuestRIP) {
RemoveCodeEntry(Frame->Thread, GuestRIP);
// Must be called from owning thread
static void RemoveThreadCodeEntryFromJit(FEXCore::Core::CpuStateFrame *Frame, uint64_t GuestRIP) {
RemoveThreadCodeEntry(Frame->Thread, GuestRIP);
}
// Debugger interface
@@ -196,18 +205,6 @@ namespace FEXCore::Context {
// same as CompileBlock, but aborts on failure
void CompileBlockJit(FEXCore::Core::CpuStateFrame *Frame, uint64_t GuestRIP);
/**
* @brief Initializes the JIT compilers for the thread
*
* @param State The internal FEX thread state object
* @param CompileThread Is this for the compile service or not?
*
* InitializeCompiler is called inside of CreateThread, so you likely don't need this
* This is exposed because the CompileService needs to initialize compilers while copying data from
* the paired InternalThreadState that it is compiling code for
*/
void InitializeCompiler(FEXCore::Core::InternalThreadState* State, bool CompileThread);
// Used for thread creation from syscalls
/**
* @brief Used to create FEX thread objects in preparation for creating a true OS thread
@@ -269,8 +266,8 @@ namespace FEXCore::Context {
uint8_t GetGPRSize() const { return Config.Is64BitMode ? 8 : 4; }
void AddNamedRegion(uintptr_t Base, uintptr_t Size, uintptr_t Offset, const std::string &filename);
void RemoveNamedRegion(uintptr_t Base, uintptr_t Size);
IR::AOTIRCacheEntry *LoadAOTIRCacheEntry(const std::string &filename);
void UnloadAOTIRCacheEntry(IR::AOTIRCacheEntry *Entry);
FEXCore::JITSymbols Symbols;
@@ -297,6 +294,9 @@ namespace FEXCore::Context {
IRCaptureCache.SetAOTIRRenamer(CacheRenamer);
}
FEXCore::Utils::PooledAllocatorMMap OpDispatcherAllocator;
FEXCore::Utils::PooledAllocatorMMap FrontendAllocator;
protected:
void ClearCodeCache(FEXCore::Core::InternalThreadState *Thread, bool AlsoClearIRCache);
@@ -310,6 +310,15 @@ namespace FEXCore::Context {
*/
void InitializeThreadData(FEXCore::Core::InternalThreadState *Thread);
/**
* @brief Initializes the JIT compilers for the thread
*
* @param State The internal FEX thread state object
*
* InitializeCompiler is called inside of CreateThread, so you likely don't need this
*/
void InitializeCompiler(FEXCore::Core::InternalThreadState* State);
void WaitForIdleWithTimeout();
void NotifyPause();
@@ -323,6 +332,7 @@ namespace FEXCore::Context {
std::unique_ptr<GdbServer> DebugServer;
IR::AOTIRCaptureCache IRCaptureCache;
std::unique_ptr<FEXCore::CodeSerialize::CodeObjectSerializeService> CodeObjectCacheService;
bool StartPaused = false;
FEX_CONFIG_OPT(AppFilename, APP_FILENAME);
+155 -109
View File
@@ -513,7 +513,8 @@ uint64_t HandleCASPAL_ARMv8(void *_ucontext, void *_info, uint32_t Instr) {
//Only 32-bit pairs
for(int i = 1; i < 10; i++) {
uint32_t NextInstr = PC[i];
if ((NextInstr & FEXCore::ArchHelpers::Arm64::ALU_OP_MASK) == FEXCore::ArchHelpers::Arm64::CMP_INST) {
if ((NextInstr & FEXCore::ArchHelpers::Arm64::ALU_OP_MASK) == FEXCore::ArchHelpers::Arm64::CMP_INST ||
(NextInstr & FEXCore::ArchHelpers::Arm64::ALU_OP_MASK) == FEXCore::ArchHelpers::Arm64::CMP_SHIFT_INST) {
ExpectedReg1 = GetRmReg(NextInstr);
} else if ((NextInstr & FEXCore::ArchHelpers::Arm64::CCMP_MASK) == FEXCore::ArchHelpers::Arm64::CCMP_INST) {
ExpectedReg2 = GetRmReg(NextInstr);
@@ -535,7 +536,6 @@ uint64_t HandleCASPAL_ARMv8(void *_ucontext, void *_info, uint32_t Instr) {
}
bool HandleAtomicVectorStore(void *_ucontext, void *_info, uint32_t Instr) {
mcontext_t* mcontext = &reinterpret_cast<ucontext_t*>(_ucontext)->uc_mcontext;
siginfo_t* info = reinterpret_cast<siginfo_t*>(_info);
if (info->si_code != BUS_ADRALN) {
@@ -1580,7 +1580,7 @@ bool HandleAtomicMemOp(void *_ucontext, void *_info, uint32_t Instr) {
return false;
}
bool HandleAtomicLoad(void *_ucontext, void *_info, uint32_t Instr) {
bool HandleAtomicLoad(void *_ucontext, void *_info, uint32_t Instr, int64_t Offset) {
mcontext_t* mcontext = &reinterpret_cast<ucontext_t*>(_ucontext)->uc_mcontext;
siginfo_t* info = reinterpret_cast<siginfo_t*>(_info);
@@ -1593,7 +1593,7 @@ bool HandleAtomicLoad(void *_ucontext, void *_info, uint32_t Instr) {
uint32_t ResultReg = Instr & 0b11111;
uint32_t AddressReg = (Instr >> 5) & 0b11111;
uint64_t Addr = mcontext->regs[AddressReg];
uint64_t Addr = mcontext->regs[AddressReg] + Offset;
if (Size == 2) {
auto Res = DoLoad16(Addr);
@@ -1623,7 +1623,7 @@ bool HandleAtomicLoad(void *_ucontext, void *_info, uint32_t Instr) {
return false;
}
bool HandleAtomicStore(void *_ucontext, void *_info, uint32_t Instr) {
bool HandleAtomicStore(void *_ucontext, void *_info, uint32_t Instr, int64_t Offset) {
mcontext_t* mcontext = &reinterpret_cast<ucontext_t*>(_ucontext)->uc_mcontext;
siginfo_t* info = reinterpret_cast<siginfo_t*>(_info);
@@ -1636,7 +1636,7 @@ bool HandleAtomicStore(void *_ucontext, void *_info, uint32_t Instr) {
uint32_t DataReg = Instr & 0x1F;
uint32_t AddressReg = (Instr >> 5) & 0b11111;
uint64_t Addr = mcontext->regs[AddressReg];
uint64_t Addr = mcontext->regs[AddressReg] + Offset;
constexpr bool DoRetry = false;
if (Size == 2) {
@@ -1746,7 +1746,8 @@ static uint64_t HandleCAS_NoAtomics(void *_ucontext, void *_info)
#endif
DesiredReg = GetRdReg(NextInstr);
}
else if ((NextInstr & FEXCore::ArchHelpers::Arm64::ALU_OP_MASK) == FEXCore::ArchHelpers::Arm64::CMP_INST) {
else if ((NextInstr & FEXCore::ArchHelpers::Arm64::ALU_OP_MASK) == FEXCore::ArchHelpers::Arm64::CMP_INST ||
(NextInstr & FEXCore::ArchHelpers::Arm64::ALU_OP_MASK) == FEXCore::ArchHelpers::Arm64::CMP_SHIFT_INST) {
ExpectedReg = GetRmReg(NextInstr);
}
}
@@ -1829,11 +1830,13 @@ uint64_t HandleAtomicLoadstoreExclusive(void *_ucontext, void *_info) {
// Scan forward at most five instructions to find our instructions
for (size_t i = 1; i < 6; ++i) {
uint32_t NextInstr = PC[i];
if ((NextInstr & FEXCore::ArchHelpers::Arm64::ALU_OP_MASK) == FEXCore::ArchHelpers::Arm64::ADD_INST) {
if ((NextInstr & FEXCore::ArchHelpers::Arm64::ALU_OP_MASK) == FEXCore::ArchHelpers::Arm64::ADD_INST ||
(NextInstr & FEXCore::ArchHelpers::Arm64::ALU_OP_MASK) == FEXCore::ArchHelpers::Arm64::ADD_SHIFT_INST) {
AtomicOp = ExclusiveAtomicPairType::TYPE_ADD;
DataSourceReg = GetRmReg(NextInstr);
}
else if ((NextInstr & FEXCore::ArchHelpers::Arm64::ALU_OP_MASK) == FEXCore::ArchHelpers::Arm64::SUB_INST) {
else if ((NextInstr & FEXCore::ArchHelpers::Arm64::ALU_OP_MASK) == FEXCore::ArchHelpers::Arm64::SUB_INST ||
(NextInstr & FEXCore::ArchHelpers::Arm64::ALU_OP_MASK) == FEXCore::ArchHelpers::Arm64::SUB_SHIFT_INST) {
uint32_t RnReg = GetRnReg(NextInstr);
if (RnReg == REGISTER_MASK) {
// Zero reg means neg
@@ -1844,21 +1847,34 @@ uint64_t HandleAtomicLoadstoreExclusive(void *_ucontext, void *_info) {
}
DataSourceReg = GetRmReg(NextInstr);
}
else if ((NextInstr & FEXCore::ArchHelpers::Arm64::ALU_OP_MASK) == FEXCore::ArchHelpers::Arm64::CMP_INST) {
return HandleCAS_NoAtomics(_ucontext, _info); //ARMv8.0 CAS
else if ((NextInstr & FEXCore::ArchHelpers::Arm64::ALU_OP_MASK) == FEXCore::ArchHelpers::Arm64::CMP_INST ||
(NextInstr & FEXCore::ArchHelpers::Arm64::ALU_OP_MASK) == FEXCore::ArchHelpers::Arm64::CMP_SHIFT_INST ) {
return HandleCAS_NoAtomics(_ucontext, _info); //ARMv8.0 CAS
}
else if ((NextInstr & FEXCore::ArchHelpers::Arm64::ALU_OP_MASK) == FEXCore::ArchHelpers::Arm64::AND_INST) {
AtomicOp = ExclusiveAtomicPairType::TYPE_AND;
DataSourceReg = GetRmReg(NextInstr);
}
else if ((NextInstr & FEXCore::ArchHelpers::Arm64::ALU_OP_MASK) == FEXCore::ArchHelpers::Arm64::BIC_INST) {
AtomicOp = ExclusiveAtomicPairType::TYPE_BIC;
DataSourceReg = GetRmReg(NextInstr);
}
else if ((NextInstr & FEXCore::ArchHelpers::Arm64::ALU_OP_MASK) == FEXCore::ArchHelpers::Arm64::OR_INST) {
AtomicOp = ExclusiveAtomicPairType::TYPE_OR;
DataSourceReg = GetRmReg(NextInstr);
}
else if ((NextInstr & FEXCore::ArchHelpers::Arm64::ALU_OP_MASK) == FEXCore::ArchHelpers::Arm64::ORN_INST) {
AtomicOp = ExclusiveAtomicPairType::TYPE_ORN;
DataSourceReg = GetRmReg(NextInstr);
}
else if ((NextInstr & FEXCore::ArchHelpers::Arm64::ALU_OP_MASK) == FEXCore::ArchHelpers::Arm64::EOR_INST) {
AtomicOp = ExclusiveAtomicPairType::TYPE_EOR;
DataSourceReg = GetRmReg(NextInstr);
}
else if ((NextInstr & FEXCore::ArchHelpers::Arm64::ALU_OP_MASK) == FEXCore::ArchHelpers::Arm64::EON_INST) {
AtomicOp = ExclusiveAtomicPairType::TYPE_EON;
DataSourceReg = GetRmReg(NextInstr);
}
else if ((NextInstr & FEXCore::ArchHelpers::Arm64::STLXR_MASK) == FEXCore::ArchHelpers::Arm64::STLXR_INST) {
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
// Just double check that the memory destination matches
@@ -1889,40 +1905,53 @@ uint64_t HandleAtomicLoadstoreExclusive(void *_ucontext, void *_info) {
uint32_t Size = 1 << (Instr >> 30);
constexpr bool DoRetry = true;
auto NOPExpected = []<typename AtomicType>(AtomicType SrcVal, AtomicType) -> AtomicType {
return SrcVal;
};
auto ADDDesired = []<typename AtomicType>(AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return SrcVal + Desired;
};
auto SUBDesired = []<typename AtomicType>(AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return SrcVal - Desired;
};
auto ANDDesired = []<typename AtomicType>(AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return SrcVal & Desired;
};
auto BICDesired = []<typename AtomicType>(AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return SrcVal & ~Desired;
};
auto ORDesired = []<typename AtomicType>(AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return SrcVal | Desired;
};
auto ORNDesired = []<typename AtomicType>(AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return SrcVal | ~Desired;
};
auto EORDesired = []<typename AtomicType>(AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return SrcVal ^ Desired;
};
auto EONDesired = []<typename AtomicType>(AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return SrcVal ^ ~Desired;
};
auto NEGDesired = []<typename AtomicType>(AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return -SrcVal;
};
auto SWAPDesired = []<typename AtomicType>(AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return Desired;
};
if (Size == 2) {
using AtomicType = uint16_t;
auto NOPExpected = [](AtomicType SrcVal, AtomicType) -> AtomicType {
return SrcVal;
};
auto ADDDesired = [](AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return SrcVal + Desired;
};
auto SUBDesired = [](AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return SrcVal - Desired;
};
auto ANDDesired = [](AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return SrcVal & Desired;
};
auto ORDesired = [](AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return SrcVal | Desired;
};
auto EORDesired = [](AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return SrcVal ^ Desired;
};
auto NEGDesired = [](AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return -SrcVal;
};
auto SWAPDesired = [](AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return Desired;
};
CASDesiredFn<AtomicType> DesiredFunction{};
switch (AtomicOp) {
@@ -1938,12 +1967,21 @@ uint64_t HandleAtomicLoadstoreExclusive(void *_ucontext, void *_info) {
case ExclusiveAtomicPairType::TYPE_AND:
DesiredFunction = ANDDesired;
break;
case ExclusiveAtomicPairType::TYPE_BIC:
DesiredFunction = BICDesired;
break;
case ExclusiveAtomicPairType::TYPE_OR:
DesiredFunction = ORDesired;
break;
case ExclusiveAtomicPairType::TYPE_ORN:
DesiredFunction = ORNDesired;
break;
case ExclusiveAtomicPairType::TYPE_EOR:
DesiredFunction = EORDesired;
break;
case ExclusiveAtomicPairType::TYPE_EON:
DesiredFunction = EONDesired;
break;
case ExclusiveAtomicPairType::TYPE_NEG:
DesiredFunction = NEGDesired;
break;
@@ -1967,38 +2005,6 @@ uint64_t HandleAtomicLoadstoreExclusive(void *_ucontext, void *_info) {
}
else if (Size == 4) {
using AtomicType = uint32_t;
auto NOPExpected = [](AtomicType SrcVal, AtomicType) -> AtomicType {
return SrcVal;
};
auto ADDDesired = [](AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return SrcVal + Desired;
};
auto SUBDesired = [](AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return SrcVal - Desired;
};
auto ANDDesired = [](AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return SrcVal & Desired;
};
auto ORDesired = [](AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return SrcVal | Desired;
};
auto EORDesired = [](AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return SrcVal ^ Desired;
};
auto NEGDesired = [](AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return -SrcVal;
};
auto SWAPDesired = [](AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return Desired;
};
CASDesiredFn<AtomicType> DesiredFunction{};
switch (AtomicOp) {
@@ -2014,12 +2020,21 @@ uint64_t HandleAtomicLoadstoreExclusive(void *_ucontext, void *_info) {
case ExclusiveAtomicPairType::TYPE_AND:
DesiredFunction = ANDDesired;
break;
case ExclusiveAtomicPairType::TYPE_BIC:
DesiredFunction = BICDesired;
break;
case ExclusiveAtomicPairType::TYPE_OR:
DesiredFunction = ORDesired;
break;
case ExclusiveAtomicPairType::TYPE_ORN:
DesiredFunction = ORNDesired;
break;
case ExclusiveAtomicPairType::TYPE_EOR:
DesiredFunction = EORDesired;
break;
case ExclusiveAtomicPairType::TYPE_EON:
DesiredFunction = EONDesired;
break;
case ExclusiveAtomicPairType::TYPE_NEG:
DesiredFunction = NEGDesired;
break;
@@ -2043,38 +2058,6 @@ uint64_t HandleAtomicLoadstoreExclusive(void *_ucontext, void *_info) {
}
else if (Size == 8) {
using AtomicType = uint64_t;
auto NOPExpected = [](AtomicType SrcVal, AtomicType) -> AtomicType {
return SrcVal;
};
auto ADDDesired = [](AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return SrcVal + Desired;
};
auto SUBDesired = [](AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return SrcVal - Desired;
};
auto ANDDesired = [](AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return SrcVal & Desired;
};
auto ORDesired = [](AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return SrcVal | Desired;
};
auto EORDesired = [](AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return SrcVal ^ Desired;
};
auto NEGDesired = [](AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return -SrcVal;
};
auto SWAPDesired = [](AtomicType SrcVal, AtomicType Desired) -> AtomicType {
return Desired;
};
CASDesiredFn<AtomicType> DesiredFunction{};
switch (AtomicOp) {
@@ -2090,12 +2073,21 @@ uint64_t HandleAtomicLoadstoreExclusive(void *_ucontext, void *_info) {
case ExclusiveAtomicPairType::TYPE_AND:
DesiredFunction = ANDDesired;
break;
case ExclusiveAtomicPairType::TYPE_BIC:
DesiredFunction = BICDesired;
break;
case ExclusiveAtomicPairType::TYPE_OR:
DesiredFunction = ORDesired;
break;
case ExclusiveAtomicPairType::TYPE_ORN:
DesiredFunction = ORNDesired;
break;
case ExclusiveAtomicPairType::TYPE_EOR:
DesiredFunction = EORDesired;
break;
case ExclusiveAtomicPairType::TYPE_EON:
DesiredFunction = EONDesired;
break;
case ExclusiveAtomicPairType::TYPE_NEG:
DesiredFunction = NEGDesired;
break;
@@ -2141,7 +2133,7 @@ bool HandleSIGBUS(bool ParanoidTSO, int Signal, void *info, void *ucontext) {
if ((Instr & 0x3F'FF'FC'00) == 0x08'DF'FC'00 || // LDAR*
(Instr & 0x3F'FF'FC'00) == 0x38'BF'C0'00) { // LDAPR*
if (ParanoidTSO) {
if (FEXCore::ArchHelpers::Arm64::HandleAtomicLoad(ucontext, info, Instr)) {
if (FEXCore::ArchHelpers::Arm64::HandleAtomicLoad(ucontext, info, Instr, 0)) {
// Skip this instruction now
ArchHelpers::Context::SetPc(ucontext, ArchHelpers::Context::GetPc(ucontext) + 4);
return true;
@@ -2165,7 +2157,7 @@ bool HandleSIGBUS(bool ParanoidTSO, int Signal, void *info, void *ucontext) {
}
else if ( (Instr & 0x3F'FF'FC'00) == 0x08'9F'FC'00) { // STLR*
if (ParanoidTSO) {
if (FEXCore::ArchHelpers::Arm64::HandleAtomicStore(ucontext, info, Instr)) {
if (FEXCore::ArchHelpers::Arm64::HandleAtomicStore(ucontext, info, Instr, 0)) {
// Skip this instruction now
ArchHelpers::Context::SetPc(ucontext, ArchHelpers::Context::GetPc(ucontext) + 4);
return true;
@@ -2187,6 +2179,60 @@ bool HandleSIGBUS(bool ParanoidTSO, int Signal, void *info, void *ucontext) {
ArchHelpers::Context::SetPc(ucontext, ArchHelpers::Context::GetPc(ucontext) - 4);
}
}
else if ((Instr & RCPC2_MASK) == LDAPUR_INST) { // LDAPUR*
// Extract the 9-bit offset from the instruction
int32_t Offset = static_cast<int32_t>(Instr) << 11 >> 23;
if (ParanoidTSO) {
if (FEXCore::ArchHelpers::Arm64::HandleAtomicLoad(ucontext, info, Instr, Offset)) {
// Skip this instruction now
ArchHelpers::Context::SetPc(ucontext, ArchHelpers::Context::GetPc(ucontext) + 4);
return true;
}
else {
LogMan::Msg::EFmt("Unhandled JIT SIGBUS LDAPUR*: PC: {} Instruction: 0x{:08x}\n", fmt::ptr(PC), PC[0]);
return false;
}
}
else {
uint32_t LDUR = 0b0011'1000'0100'0000'0000'0000'0000'0000;
LDUR |= Size << 30;
LDUR |= AddrReg << 5;
LDUR |= DataReg;
LDUR |= Instr & (0b1'1111'1111 << 9);
PC[-1] = DMB;
PC[0] = LDUR;
PC[1] = DMB;
// Back up one instruction and have another go
ArchHelpers::Context::SetPc(ucontext, ArchHelpers::Context::GetPc(ucontext) - 4);
}
}
else if ((Instr & RCPC2_MASK) == STLUR_INST) { // STLUR*
// Extract the 9-bit offset from the instruction
int32_t Offset = static_cast<int32_t>(Instr) << 11 >> 23;
if (ParanoidTSO) {
if (FEXCore::ArchHelpers::Arm64::HandleAtomicStore(ucontext, info, Instr, Offset)) {
// Skip this instruction now
ArchHelpers::Context::SetPc(ucontext, ArchHelpers::Context::GetPc(ucontext) + 4);
return true;
}
else {
LogMan::Msg::EFmt("Unhandled JIT SIGBUS LDLUR*: PC: {} Instruction: 0x{:08x}\n", fmt::ptr(PC), PC[0]);
return false;
}
}
else {
uint32_t STUR = 0b0011'1000'0000'0000'0000'0000'0000'0000;
STUR |= Size << 30;
STUR |= AddrReg << 5;
STUR |= DataReg;
STUR |= Instr & (0b1'1111'1111 << 9);
PC[-1] = DMB;
PC[0] = STUR;
PC[1] = DMB;
// Back up one instruction and have another go
ArchHelpers::Context::SetPc(ucontext, ArchHelpers::Context::GetPc(ucontext) - 4);
}
}
else if ((Instr & FEXCore::ArchHelpers::Arm64::LDAXP_MASK) == FEXCore::ArchHelpers::Arm64::LDAXP_INST) { // LDAXP
//Should be compare and swap pair only. LDAXP not used elsewhere
uint64_t BytesToSkip = FEXCore::ArchHelpers::Arm64::HandleCASPAL_ARMv8(ucontext, info, Instr);
+22 -9
View File
@@ -12,6 +12,10 @@ namespace FEXCore::ArchHelpers::Arm64 {
constexpr uint32_t ATOMIC_MEM_MASK = 0x3B200C00;
constexpr uint32_t ATOMIC_MEM_INST = 0x38200000;
constexpr uint32_t RCPC2_MASK = 0x3F'E0'0C'00;
constexpr uint32_t LDAPUR_INST = 0x19'40'00'00;
constexpr uint32_t STLUR_INST = 0x19'00'00'00;
constexpr uint32_t LDAXP_MASK = 0xBF'FF'80'00;
constexpr uint32_t LDAXP_INST = 0x88'7F'80'00;
@@ -27,13 +31,19 @@ namespace FEXCore::ArchHelpers::Arm64 {
constexpr uint32_t CBNZ_MASK = 0x7F'00'00'00;
constexpr uint32_t CBNZ_INST = 0x35'00'00'00;
constexpr uint32_t ALU_OP_MASK = 0x7F'00'00'00;
constexpr uint32_t ADD_INST = 0x0B'00'00'00;
constexpr uint32_t SUB_INST = 0x4B'00'00'00;
constexpr uint32_t CMP_INST = 0x6B'00'00'00;
constexpr uint32_t AND_INST = 0x0A'00'00'00;
constexpr uint32_t OR_INST = 0x2A'00'00'00;
constexpr uint32_t EOR_INST = 0x4A'00'00'00;
constexpr uint32_t ALU_OP_MASK = 0x7F'20'00'00;
constexpr uint32_t ADD_INST = 0x0B'00'00'00;
constexpr uint32_t SUB_INST = 0x4B'00'00'00;
constexpr uint32_t ADD_SHIFT_INST = 0x0B'20'00'00;
constexpr uint32_t SUB_SHIFT_INST = 0x4B'20'00'00;
constexpr uint32_t CMP_INST = 0x6B'00'00'00;
constexpr uint32_t CMP_SHIFT_INST = 0x6B'20'00'00;
constexpr uint32_t AND_INST = 0x0A'00'00'00;
constexpr uint32_t BIC_INST = 0x0A'20'00'00;
constexpr uint32_t OR_INST = 0x2A'00'00'00;
constexpr uint32_t ORN_INST = 0x2A'20'00'00;
constexpr uint32_t EOR_INST = 0x4A'00'00'00;
constexpr uint32_t EON_INST = 0x4A'20'00'00;
constexpr uint32_t CCMP_MASK = 0x7F'E0'0C'10;
constexpr uint32_t CCMP_INST = 0x7A'40'00'00;
@@ -46,8 +56,11 @@ namespace FEXCore::ArchHelpers::Arm64 {
TYPE_ADD,
TYPE_SUB,
TYPE_AND,
TYPE_BIC,
TYPE_OR,
TYPE_ORN,
TYPE_EOR,
TYPE_EON,
TYPE_NEG, // This is just a sub with zero. Need to know the differences
};
@@ -83,8 +96,8 @@ namespace FEXCore::ArchHelpers::Arm64 {
return (Instr >> RM_OFFSET) & REGISTER_MASK;
}
bool HandleAtomicLoad(void *_ucontext, void *_info, uint32_t Instr);
bool HandleAtomicStore(void *_ucontext, void *_info, uint32_t Instr);
bool HandleAtomicLoad(void *_ucontext, void *_info, uint32_t Instr, int64_t Offset);
bool HandleAtomicStore(void *_ucontext, void *_info, uint32_t Instr, int64_t Offset);
bool HandleAtomicLoad128(void *_ucontext, void *_info, uint32_t Instr);
uint64_t HandleAtomicLoadstoreExclusive(void *_ucontext, void *_info);
bool HandleCASPAL(void *_ucontext, void *_info, uint32_t Instr);
@@ -1,5 +1,7 @@
#include "Interface/Core/ArchHelpers/Arm64Emitter.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Context/Context.h"
#include "Interface/HLE/Thunks/Thunks.h"
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Utils/LogManager.h>
@@ -16,7 +18,9 @@ namespace FEXCore::CPU {
#define STATE x28
// We want vixl to not allocate a default buffer. Jit and dispatcher will manually create one.
Arm64Emitter::Arm64Emitter(FEXCore::Context::Context *ctx, size_t size) : vixl::aarch64::Assembler(size, vixl::aarch64::PositionDependentCode) {
Arm64Emitter::Arm64Emitter(FEXCore::Context::Context *ctx, size_t size)
: vixl::aarch64::Assembler(size, vixl::aarch64::PositionDependentCode)
, EmitterCTX {ctx} {
CPU.SetUp();
auto Features = vixl::CPUFeatures::InferFromOS();
@@ -28,20 +32,77 @@ Arm64Emitter::Arm64Emitter(FEXCore::Context::Context *ctx, size_t size) : vixl::
SetCPUFeatures(Features);
}
void Arm64Emitter::LoadConstant(vixl::aarch64::Register Reg, uint64_t Constant) {
void Arm64Emitter::LoadConstant(vixl::aarch64::Register Reg, uint64_t Constant, bool NOPPad) {
bool Is64Bit = Reg.IsX();
int Segments = Is64Bit ? 4 : 2;
if (Is64Bit && ((~Constant)>> 16) == 0) {
movn(Reg, (~Constant) & 0xFFFF);
if (NOPPad) {
nop(); nop(); nop();
}
return;
}
movz(Reg, (Constant) & 0xFFFF, 0);
for (int i = 1; i < Segments; ++i) {
int NumMoves = 1;
int RequiredMoveSegments{};
// Count the number of move segments
// We only want to use ADRP+ADD if we have more than 1 segment
for (size_t i = 0; i < Segments; ++i) {
uint16_t Part = (Constant >> (i * 16)) & 0xFFFF;
if (Part) {
movk(Reg, Part, i * 16);
if (Part != 0) {
++RequiredMoveSegments;
}
}
// ADRP+ADD is specifically optimized in hardware
// Check if we can use this
auto PC = GetCursorAddress<uint64_t>();
// PC aligned to page
uint64_t AlignedPC = PC & ~0xFFFULL;
// Offset from aligned PC
int64_t AlignedOffset = static_cast<int64_t>(Constant) - static_cast<int64_t>(AlignedPC);
// If the aligned offset is within the 4GB window then we can use ADRP+ADD
// and the number of move segments more than 1
if (RequiredMoveSegments > 1 && vixl::IsInt32(AlignedOffset)) {
// If this is 4k page aligned then we only need ADRP
if ((AlignedOffset & 0xFFF) == 0) {
adrp(Reg, AlignedOffset >> 12);
}
else {
// If the constant is within 1MB of PC then we can still use ADR to load in a single instruction
// 21-bit signed integer here
int64_t SmallOffset = static_cast<int64_t>(Constant) - static_cast<int64_t>(PC);
if (vixl::IsInt21(SmallOffset)) {
adr(Reg, SmallOffset);
}
else {
// Need to use ADRP + ADD
adrp(Reg, AlignedOffset >> 12);
add(Reg, Reg, Constant & 0xFFF);
NumMoves = 2;
}
}
}
else {
movz(Reg, (Constant) & 0xFFFF, 0);
for (int i = 1; i < Segments; ++i) {
uint16_t Part = (Constant >> (i * 16)) & 0xFFFF;
if (Part) {
movk(Reg, Part, i * 16);
++NumMoves;
}
}
}
if (NOPPad) {
for (int i = NumMoves; i < Segments; ++i) {
nop();
}
}
}
@@ -50,7 +111,7 @@ void Arm64Emitter::PushCalleeSavedRegisters() {
// We need to save pairs of registers
// We save r19-r30
MemOperand PairOffset(sp, -16, PreIndex);
const std::array<std::pair<vixl::aarch64::XRegister, vixl::aarch64::XRegister>, 6> CalleeSaved = {{
const std::array<std::pair<vixl::aarch64::Register, vixl::aarch64::Register>, 6> CalleeSaved = {{
{x19, x20},
{x21, x22},
{x23, x24},
@@ -113,7 +174,7 @@ void Arm64Emitter::PopCalleeSavedRegisters() {
}
MemOperand PairOffset(sp, 16, PostIndex);
const std::array<std::pair<vixl::aarch64::XRegister, vixl::aarch64::XRegister>, 6> CalleeSaved = {{
const std::array<std::pair<vixl::aarch64::Register, vixl::aarch64::Register>, 6> CalleeSaved = {{
{x29, x30},
{x27, x28},
{x25, x26},
@@ -128,51 +189,75 @@ void Arm64Emitter::PopCalleeSavedRegisters() {
}
void Arm64Emitter::SpillStaticRegs(bool FPRs, uint32_t SpillMask) {
void Arm64Emitter::SpillStaticRegs(bool FPRs, uint32_t GPRSpillMask, uint32_t FPRSpillMask) {
if (StaticRegisterAllocation()) {
for (size_t i = 0; i < SRA64.size(); i+=2) {
auto Reg1 = SRA64[i];
auto Reg2 = SRA64[i+1];
if (((1U << Reg1.GetCode()) & SpillMask) &&
((1U << Reg2.GetCode()) & SpillMask)) {
if (((1U << Reg1.GetCode()) & GPRSpillMask) &&
((1U << Reg2.GetCode()) & GPRSpillMask)) {
stp(Reg1, Reg2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i])));
}
else if (((1U << Reg1.GetCode()) & SpillMask)) {
else if (((1U << Reg1.GetCode()) & GPRSpillMask)) {
str(Reg1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i])));
}
else if (((1U << Reg2.GetCode()) & SpillMask)) {
else if (((1U << Reg2.GetCode()) & GPRSpillMask)) {
str(Reg2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i+1])));
}
}
if (FPRs) {
for (size_t i = 0; i < SRAFPR.size(); i+=2) {
stp(SRAFPR[i].Q(), SRAFPR[i+1].Q(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.xmm[i][0])));
auto Reg1 = SRAFPR[i];
auto Reg2 = SRAFPR[i+1];
if (((1U << Reg1.GetCode()) & FPRSpillMask) &&
((1U << Reg2.GetCode()) & FPRSpillMask)) {
stp(Reg1.Q(), Reg2.Q(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.xmm[i][0])));
}
else if (((1U << Reg1.GetCode()) & FPRSpillMask)) {
str(Reg1.Q(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.xmm[i][0])));
}
else if (((1U << Reg2.GetCode()) & FPRSpillMask)) {
str(Reg2.Q(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.xmm[i+1][0])));
}
}
}
}
}
void Arm64Emitter::FillStaticRegs(bool FPRs, uint32_t FillMask) {
void Arm64Emitter::FillStaticRegs(bool FPRs, uint32_t GPRFillMask, uint32_t FPRFillMask) {
if (StaticRegisterAllocation()) {
for (size_t i = 0; i < SRA64.size(); i+=2) {
auto Reg1 = SRA64[i];
auto Reg2 = SRA64[i+1];
if (((1U << Reg1.GetCode()) & FillMask) &&
((1U << Reg2.GetCode()) & FillMask)) {
if (((1U << Reg1.GetCode()) & GPRFillMask) &&
((1U << Reg2.GetCode()) & GPRFillMask)) {
ldp(Reg1, Reg2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i])));
}
else if (((1U << Reg1.GetCode()) & FillMask)) {
else if (((1U << Reg1.GetCode()) & GPRFillMask)) {
ldr(Reg1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i])));
}
else if (((1U << Reg2.GetCode()) & FillMask)) {
else if (((1U << Reg2.GetCode()) & GPRFillMask)) {
ldr(Reg2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i+1])));
}
}
if (FPRs) {
for (size_t i = 0; i < SRAFPR.size(); i+=2) {
ldp(SRAFPR[i].Q(), SRAFPR[i+1].Q(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.xmm[i][0])));
auto Reg1 = SRAFPR[i];
auto Reg2 = SRAFPR[i+1];
if (((1U << Reg1.GetCode()) & FPRFillMask) &&
((1U << Reg2.GetCode()) & FPRFillMask)) {
ldp(Reg1.Q(), Reg2.Q(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.xmm[i][0])));
}
else if (((1U << Reg1.GetCode()) & FPRFillMask)) {
ldr(Reg1.Q(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.xmm[i][0])));
}
else if (((1U << Reg2.GetCode()) & FPRFillMask)) {
ldr(Reg2.Q(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.xmm[i+1][0])));
}
}
}
}
@@ -240,8 +325,126 @@ void Arm64Emitter::ResetStack() {
void Arm64Emitter::Align16B() {
uint64_t CurrentOffset = GetCursorAddress<uint64_t>();
for (uint64_t i = (16 - (CurrentOffset & 0xF)); i != 0; i -= 4) {
nop();
nop();
}
}
uint64_t Arm64Emitter::GetNamedSymbolLiteral(FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op) {
switch (Op) {
case FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol::SYMBOL_LITERAL_EXITFUNCTION_LINKER:
return Dispatcher->ExitFunctionLinkerAddress;
break;
default:
ERROR_AND_DIE_FMT("Unknown named symbol literal: {}", static_cast<uint32_t>(Op));
break;
}
return ~0ULL;
}
void Arm64Emitter::InsertNamedThunkRelocation(vixl::aarch64::Register Reg, const IR::SHA256Sum &Sum) {
Relocation MoveABI{};
MoveABI.NamedThunkMove.Header.Type = FEXCore::CPU::RelocationTypes::RELOC_NAMED_THUNK_MOVE;
// Offset is the offset from the entrypoint of the block
auto CurrentCursor = GetCursorAddress<uint64_t>();
MoveABI.NamedThunkMove.Offset = CurrentCursor - GuestEntry;
MoveABI.NamedThunkMove.Symbol = Sum;
MoveABI.NamedThunkMove.RegisterIndex = Reg.GetCode();
uint64_t Pointer = reinterpret_cast<uint64_t>(EmitterCTX->ThunkHandler->LookupThunk(Sum));
LoadConstant(Reg, Pointer, EmitterCTX->Config.CacheObjectCodeCompilation());
Relocations.emplace_back(MoveABI);
}
Arm64Emitter::NamedSymbolLiteralPair Arm64Emitter::InsertNamedSymbolLiteral(FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op) {
uint64_t Pointer = GetNamedSymbolLiteral(Op);
Arm64Emitter::NamedSymbolLiteralPair Lit {
.Lit = Literal(Pointer),
.MoveABI = {
.NamedSymbolLiteral = {
.Header = {
.Type = FEXCore::CPU::RelocationTypes::RELOC_NAMED_SYMBOL_LITERAL,
},
.Symbol = Op,
.Offset = 0,
},
},
};
return Lit;
}
void Arm64Emitter::PlaceNamedSymbolLiteral(NamedSymbolLiteralPair &Lit) {
// Offset is the offset from the entrypoint of the block
auto CurrentCursor = GetCursorAddress<uint64_t>();
Lit.MoveABI.NamedSymbolLiteral.Offset = CurrentCursor - GuestEntry;
place(&Lit.Lit);
Relocations.emplace_back(Lit.MoveABI);
}
void Arm64Emitter::InsertGuestRIPMove(vixl::aarch64::Register Reg, uint64_t Constant) {
Relocation MoveABI{};
MoveABI.GuestRIPMove.Header.Type = FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_MOVE;
// Offset is the offset from the entrypoint of the block
auto CurrentCursor = GetCursorAddress<uint64_t>();
MoveABI.GuestRIPMove.Offset = CurrentCursor - GuestEntry;
MoveABI.GuestRIPMove.GuestRIP = Constant;
MoveABI.GuestRIPMove.RegisterIndex = Reg.GetCode();
LoadConstant(Reg, Constant, EmitterCTX->Config.CacheObjectCodeCompilation());
Relocations.emplace_back(MoveABI);
}
bool Arm64Emitter::ApplyRelocations(uint64_t GuestEntry, uint64_t CodeEntry, uint64_t CursorEntry, size_t NumRelocations, const char* EntryRelocations) {
size_t DataIndex{};
for (size_t j = 0; j < NumRelocations; ++j) {
const FEXCore::CPU::Relocation *Reloc = reinterpret_cast<const FEXCore::CPU::Relocation *>(&EntryRelocations[DataIndex]);
LOGMAN_THROW_A_FMT((DataIndex % alignof(Relocation)) == 0, "Alignment of relocation wasn't adhered to");
switch (Reloc->Header.Type) {
case FEXCore::CPU::RelocationTypes::RELOC_NAMED_SYMBOL_LITERAL: {
uint64_t Pointer = GetNamedSymbolLiteral(Reloc->NamedSymbolLiteral.Symbol);
// Relocation occurs at the cursorEntry + offset relative to that cursor
GetBuffer()->SetCursorOffset(CursorEntry + Reloc->NamedSymbolLiteral.Offset);
// Generate a literal so we can place it
Literal<uint64_t> Lit(Pointer);
place(&Lit);
DataIndex += sizeof(Reloc->NamedSymbolLiteral);
break;
}
case FEXCore::CPU::RelocationTypes::RELOC_NAMED_THUNK_MOVE: {
uint64_t Pointer = reinterpret_cast<uint64_t>(EmitterCTX->ThunkHandler->LookupThunk(Reloc->NamedThunkMove.Symbol));
if (Pointer == ~0ULL) {
return false;
}
// Relocation occurs at the cursorEntry + offset relative to that cursor.
GetBuffer()->SetCursorOffset(CursorEntry + Reloc->NamedThunkMove.Offset);
LoadConstant(vixl::aarch64::XRegister(Reloc->NamedThunkMove.RegisterIndex), Pointer, true);
DataIndex += sizeof(Reloc->NamedThunkMove);
break;
}
case FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_MOVE: {
// XXX: Reenable once the JIT Object Cache is upstream
// XXX: Should spin the relocation list, create a list of guest RIP moves, and ask for them all once, reduces lock contention.
uint64_t Pointer = ~0ULL; // EmitterCTX->JITObjectCache->FindRelocatedRIP(Reloc->GuestRIPMove.GuestRIP);
if (Pointer == ~0ULL) {
return false;
}
// Relocation occurs at the cursorEntry + offset relative to that cursor.
GetBuffer()->SetCursorOffset(CursorEntry + Reloc->GuestRIPMove.Offset);
LoadConstant(vixl::aarch64::XRegister(Reloc->GuestRIPMove.RegisterIndex), Pointer, true);
DataIndex += sizeof(Reloc->GuestRIPMove);
break;
}
}
}
return true;
}
}
@@ -1,5 +1,8 @@
#pragma once
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Core/ObjectCache/Relocations.h"
#include "aarch64/assembler-aarch64.h"
#include "aarch64/constants-aarch64.h"
#include "aarch64/cpu-aarch64.h"
@@ -60,10 +63,19 @@ class Arm64Emitter : public vixl::aarch64::Assembler {
protected:
Arm64Emitter(FEXCore::Context::Context *ctx, size_t size);
std::unique_ptr<FEXCore::CPU::Dispatcher> Dispatcher;
FEXCore::Context::Context *EmitterCTX;
vixl::aarch64::CPU CPU;
void LoadConstant(vixl::aarch64::Register Reg, uint64_t Constant);
void SpillStaticRegs(bool FPRs = true, uint32_t SpillMask = ~0U);
void FillStaticRegs(bool FPRs = true, uint32_t FillMask = ~0U);
void LoadConstant(vixl::aarch64::Register Reg, uint64_t Constant, bool NOPPad = false);
void SpillStaticRegs(bool FPRs = true, uint32_t GPRSpillMask = ~0U, uint32_t FPRSpillMask = ~0U);
void FillStaticRegs(bool FPRs = true, uint32_t GPRFillMask = ~0U, uint32_t FPRFillMask = ~0U);
static constexpr uint32_t CALLER_GPR_MASK = 0b0011'1111'1111'1111'1111;
// This isn't technically true because the lower 64-bits of v8..v15 are callee saved
// We can't guarantee only the lower 64bits are used so flush everything
static constexpr uint32_t CALLER_FPR_MASK = ~0U;
void PushDynamicRegsAndLR();
void PopDynamicRegsAndLR();
@@ -73,8 +85,66 @@ protected:
void ResetStack();
void Align16B();
/**
* @name Relocations
* @{ */
uint64_t GetNamedSymbolLiteral(FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op);
/**
* @brief A literal pair relocation object for named symbol literals
*/
struct NamedSymbolLiteralPair {
Literal<uint64_t> Lit;
Relocation MoveABI{};
};
/**
* @brief Inserts a thunk relocation
*
* @param Reg - The GPR to move the thunk handler in to
* @param Sum - The hash of the thunk
*/
void InsertNamedThunkRelocation(vixl::aarch64::Register Reg, const IR::SHA256Sum &Sum);
/**
* @brief Inserts a guest GPR move relocation
*
* @param Reg - The GPR to move the guest RIP in to
* @param Constant - The guest RIP that will be relocated
*/
void InsertGuestRIPMove(vixl::aarch64::Register Reg, uint64_t Constant);
/**
* @brief Inserts a named symbol as a literal in memory
*
* Need to use `PlaceNamedSymbolLiteral` with the return value to place the literal in the desired location
*
* @param Op The named symbol to place
*
* @return A temporary `NamedSymbolLiteralPair`
*/
NamedSymbolLiteralPair InsertNamedSymbolLiteral(FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op);
/**
* @brief Place the named symbol literal relocation in memory
*
* @param Lit - Which literal to place
*/
void PlaceNamedSymbolLiteral(NamedSymbolLiteralPair &Lit);
std::vector<FEXCore::CPU::Relocation> Relocations;
///< Relocation code loading
bool ApplyRelocations(uint64_t GuestEntry, uint64_t CodeEntry, uint64_t CursorEntry, size_t NumRelocations, const char* EntryRelocations);
/** @} */
uint32_t SpillSlots{};
/**
* @brief Current guest RIP entrypoint
*/
uint64_t GuestEntry{};
FEX_CONFIG_OPT(StaticRegisterAllocation, SRA);
};
+55 -7
View File
@@ -125,7 +125,8 @@ void CPUIDEmu::SetupHostHybridFlag() {
// Only read 18 bytes for a 64bit value prefixed with 0x
if (FEXCore::FileLoading::LoadFile(Data, MIDRPath, 18)) {
uint64_t NewMIDR{};
if (FEXCore::StrConv::Conv(&Data.at(0), &NewMIDR)) {
std::string_view MIDRView(&Data.at(0), 18);
if (FEXCore::StrConv::Conv(MIDRView, &NewMIDR)) {
if (MIDR != 0 && MIDR != NewMIDR) {
// CPU mismatch, claim hybrid
Hybrid = true;
@@ -403,6 +404,8 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_0h(uint32_t Leaf) {
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_01h(uint32_t Leaf) {
FEXCore::CPUID::FunctionResults Res{};
uint32_t CoreCount = Cores();
// XXX: Enable once the rest of the SSE4.2 instructions are emulated
uint32_t SupportsSSE42 = CTX->HostFeatures.SupportsCRC && false ? 1 : 0;
Res.eax = FAMILY_IDENTIFIER;
@@ -432,7 +435,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_01h(uint32_t Leaf) {
(0 << 17) | // Process-context identifiers
(0 << 18) | // Prefetching from memory mapped device
(1 << 19) | // SSE4.1
(0 << 20) | // SSE4.2
(SupportsSSE42 << 20) | // SSE4.2
(0 << 21) | // X2APIC
(1 << 22) | // MOVBE
(1 << 23) | // POPCNT
@@ -442,7 +445,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_01h(uint32_t Leaf) {
(0 << 27) | // OSXSAVE
(SUPPORTS_AVX << 28) | // AVX
(0 << 29) | // F16C
(0 << 30) | // RDRAND
(CTX->HostFeatures.SupportsRAND << 30) | // RDRAND
(0 << 31); // Hypervisor always returns zero
Res.edx =
@@ -644,7 +647,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_07h(uint32_t Leaf) {
(0 << 15) | // Intel Resource Directory Technology Allocation
(0 << 16) | // Reserved
(0 << 17) | // Reserved
(0 << 18) | // RDSEED
(CTX->HostFeatures.SupportsRAND << 18) | // RDSEED
(1 << 19) | // ADCX and ADOX instructions
(0 << 20) | // SMAP Supervisor mode access prevention and CLAC/STAC instructions
(0 << 21) | // Reserved
@@ -655,7 +658,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_07h(uint32_t Leaf) {
(0 << 26) | // Reserved
(0 << 27) | // Reserved
(0 << 28) | // Reserved
(0 << 29) | // SHA instructions
(1 << 29) | // SHA instructions
(0 << 30) | // Reserved
(0 << 31); // Reserved
@@ -809,6 +812,47 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_1Ah(uint32_t Leaf) {
return Res;
}
// Hypervisor CPUID information leaf
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_4000_0000h(uint32_t Leaf) {
FEXCore::CPUID::FunctionResults Res{};
// Maximum supported hypervisor leafs
// We only expose the information leaf
//
// Common courtesy to follow VMWare's "Hypervisor CPUID Interface proposal"
// 4000_0000h - Information leaf. Advertising to the software which hypervisor this is
// 4000_0001h - 4000_000Fh - Hypervisor specific leafs. FEX can use these for anything
// 4000_0010h - 4000_00FFh - "Generic Leafs" - Try not to overwrite, other hypervisors might expect information in these
//
// CPUID documentation information:
// 4000_0000h - 4FFF_FFFFh - No existing or future CPU will return information in this range
// Reserved entirely for VMs to do whatever they want.
Res.eax = 0x40000001;
// EBX, EDX, ECX become the hypervisor ID signature
constexpr static char HypervisorID[12] = "FEXIFEXIEMU";
memcpy(&Res.ebx, HypervisorID, sizeof(HypervisorID));
return Res;
}
// Hypervisor CPUID information leaf
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_4000_0001h(uint32_t Leaf) {
FEXCore::CPUID::FunctionResults Res{};
if (Leaf == 0) {
// EAX[3:0] Is the host architecture that FEX is running under
#ifdef _M_X86_64
// EAX[3:0] = 1 = x86_64 host architecture
Res.eax |= 0b0001;
#elif defined(_M_ARM_64)
// EAX[3:0] = 2 = AArch64 host architecture
Res.eax |= 0b0010;
#else
// EAX[3:0] = 0 = Unknown architecture
#endif
}
return Res;
}
// Highest extended function implemented
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0000h(uint32_t Leaf) {
FEXCore::CPUID::FunctionResults Res{};
@@ -899,8 +943,8 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0001h(uint32_t Leaf) {
(1 << 27) | // RDTSCP
(0 << 28) | // Reserved
(1 << 29) | // Long Mode
(0 << 30) | // 3DNow! Extensions
(0 << 31); // 3DNow!
(1 << 30) | // 3DNow! Extensions
(1 << 31); // 3DNow!
return Res;
}
@@ -1201,6 +1245,10 @@ void CPUIDEmu::Init(FEXCore::Context::Context *ctx) {
#ifndef CPUID_AMD
RegisterFunction(0x1A, &CPUIDEmu::Function_1Ah);
#endif
// Hypervisor CPUID information leaf
RegisterFunction(0x4000'0000, &CPUIDEmu::Function_4000_0000h);
RegisterFunction(0x4000'0001, &CPUIDEmu::Function_4000_0001h);
// Largest extended function number
RegisterFunction(0x8000'0000, &CPUIDEmu::Function_8000_0000h);
// Processor vendor
+3
View File
@@ -3,6 +3,7 @@
#include <cstdint>
#include <unordered_map>
#include <utility>
#include <vector>
#include <FEXCore/Core/CPUID.h>
#include <FEXCore/Config/Config.h>
@@ -78,6 +79,8 @@ private:
FEXCore::CPUID::FunctionResults Function_0Dh(uint32_t Leaf);
FEXCore::CPUID::FunctionResults Function_15h(uint32_t Leaf);
FEXCore::CPUID::FunctionResults Function_1Ah(uint32_t Leaf);
FEXCore::CPUID::FunctionResults Function_4000_0000h(uint32_t Leaf);
FEXCore::CPUID::FunctionResults Function_4000_0001h(uint32_t Leaf);
FEXCore::CPUID::FunctionResults Function_8000_0000h(uint32_t Leaf);
FEXCore::CPUID::FunctionResults Function_8000_0001h(uint32_t Leaf);
FEXCore::CPUID::FunctionResults Function_8000_0002h(uint32_t Leaf);
@@ -1,165 +0,0 @@
#include "Interface/Context/Context.h"
#include "Interface/Core/LookupCache.h"
#include "Interface/Core/CompileService.h"
#include "Interface/Core/OpcodeDispatcher.h"
#include "FEXCore/Debug/InternalThreadState.h"
#include "FEXCore/HLE/Linux/ThreadManagement.h"
#include "Interface/IR/PassManager.h"
#include <FEXCore/Core/CPUBackend.h>
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Core/SignalDelegator.h>
#include <FEXCore/Utils/Event.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/Threads.h>
#include <memory>
#include <pthread.h>
#include <stdio.h>
namespace FEXCore {
static void* ThreadHandler(void *Arg) {
FEXCore::CompileService *This = reinterpret_cast<FEXCore::CompileService*>(Arg);
This->ExecutionThread();
return nullptr;
}
CompileService::CompileService(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread)
: CTX {ctx}
, ParentThread {Thread} {
CompileThreadData = std::make_unique<FEXCore::Core::InternalThreadState>();
CompileThreadData->IsCompileService = true;
// We need a compiler for this work thread
CTX->InitializeCompiler(CompileThreadData.get(), true);
CompileThreadData->CPUBackend->CopyNecessaryDataForCompileThread(ParentThread->CPUBackend.get());
uint64_t OldMask = FEXCore::Threads::SetSignalMask(~0ULL);
WorkerThread = FEXCore::Threads::Thread::Create(ThreadHandler, this);
FEXCore::Threads::SetSignalMask(OldMask);
}
void CompileService::Initialize() {
// Share CompileService which = this
CompileThreadData->CompileService = ParentThread->CompileService;
}
void CompileService::Shutdown() {
ShuttingDown = true;
// Kick the working thread
StartWork.NotifyAll();
WorkerThread->join(nullptr);
}
void CompileService::ClearCache(FEXCore::Core::InternalThreadState *Thread) {
// On cache clear we need to spin down the execution thread to ensure it isn't trying to give us more work items
if (CompileMutex.try_lock()) {
// We can only clear these things if we pulled the compile mutex
// Grab the work queue and clear it
// We don't need to grab the queue mutex since this thread will no longer receive any work events
// Threads are bounded 1:1
while (!WorkQueue.empty()) {
WorkQueue.pop();
}
// Go through the garbage collection array and clear it
// It's safe to clear things that aren't marked safe since we are clearing cache
GCArray.clear();
LOGMAN_THROW_A_FMT(CompileThreadData->LocalIRCache.empty(), "Compile service must never have LocalIRCache");
CompileMutex.unlock();
}
// Clear the inverse cache of what is calling us from the Context ClearCache routine
auto SelectedThread = Thread->IsCompileService ? ParentThread : Thread;
SelectedThread->LookupCache->ClearCache();
SelectedThread->CPUBackend->ClearCache();
}
CompileService::WorkItem *CompileService::CompileCode(uint64_t RIP) {
WorkItem* ResultItem = nullptr;
{
// Tell the worker thread to compile code for us
auto Item = std::make_unique<WorkItem>();
Item->RIP = RIP;
// Fill the threads work queue
std::scoped_lock lk(QueueMutex);
ResultItem = WorkQueue.emplace(std::move(Item)).get();
}
// Notify the thread that it has more work
StartWork.NotifyAll();
return ResultItem;
}
void CompileService::ExecutionThread() {
// Set our thread name so we can see its relation
char ThreadName[16]{};
snprintf(ThreadName, 16, "%ld-CS", ParentThread->ThreadManager.TID.load());
pthread_setname_np(pthread_self(), ThreadName);
while (true) {
// Wait for work
StartWork.Wait();
if (ShuttingDown.load()) {
break;
}
std::scoped_lock lk(CompileMutex);
size_t WorkItems{};
do {
// Grab a work item
std::unique_ptr<WorkItem> Item{};
{
std::scoped_lock lk(QueueMutex);
WorkItems = WorkQueue.size();
if (WorkItems != 0) {
Item = std::move(WorkQueue.front());
WorkQueue.pop();
}
}
// If we had a work item then work on it
if (Item) {
// Make sure it's not in lookup cache by accident
LOGMAN_THROW_A_FMT(CompileThreadData->LookupCache->FindBlock(Item->RIP) == 0, "Compile Service must never have entries in the LookupCache");
// Code isn't in cache, compile now
// Set our thread state's RIP
CompileThreadData->CurrentFrame->State.rip = Item->RIP;
auto [CodePtr, IRList, DebugData, RAData, Generated, StartAddr, Length] = CTX->CompileCode(CompileThreadData.get(), Item->RIP);
LOGMAN_THROW_A_FMT(Generated == true, "Compile Service doesn't have IR Cache");
if (!CodePtr) {
// XXX: We currently have the expectation that compile service code will be significantly smaller than regular thread's code
ERROR_AND_DIE_FMT("Couldn't compile code for thread at RIP: 0x{:x}", Item->RIP);
}
Item->CodePtr = CodePtr;
Item->IRList = IRList;
Item->DebugData = DebugData;
Item->RAData = RAData;
Item->StartAddr = StartAddr;
Item->Length = Length;
auto& GCItem = GCArray.emplace_back(std::move(Item));
GCItem->ServiceWorkDone.NotifyAll();
}
} while (WorkItems != 0);
// Clean up any safe entries in our GC array if we have any.
std::erase_if(GCArray, [](const auto& Entry) {
return Entry->SafeToClear.load(std::memory_order_relaxed);
});
}
}
}
-68
View File
@@ -1,68 +0,0 @@
#pragma once
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXCore/Utils/Event.h>
#include <FEXCore/Utils/Threads.h>
#include <atomic>
#include <memory>
#include <mutex>
#include <queue>
#include <stdint.h>
#include <vector>
namespace FEXCore {
namespace Context {
struct Context;
}
namespace IR {
class IRListView;
class RegisterAllocationData;
};
class CompileService final {
public:
CompileService(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread);
void Initialize();
void Shutdown();
struct WorkItem {
// Incoming
uint64_t RIP{};
// Outgoing
void *CodePtr{};
FEXCore::IR::IRListView *IRList{};
FEXCore::IR::RegisterAllocationData *RAData{};
FEXCore::Core::DebugData *DebugData{};
uint64_t StartAddr;
uint64_t Length;
// Communication
Event ServiceWorkDone{};
std::atomic_bool SafeToClear{};
};
WorkItem *CompileCode(uint64_t RIP);
void ClearCache(FEXCore::Core::InternalThreadState *Thread);
// Public for threading
void ExecutionThread();
bool IsAddressInJITCode(uint64_t Address) const {
return CompileThreadData->CPUBackend->IsAddressInJITCode(Address, false, false);
}
private:
FEXCore::Context::Context *CTX;
FEXCore::Core::InternalThreadState *ParentThread;
std::unique_ptr<FEXCore::Threads::Thread> WorkerThread;
std::unique_ptr<FEXCore::Core::InternalThreadState> CompileThreadData;
std::mutex QueueMutex{};
std::mutex CompileMutex{};
std::queue<std::unique_ptr<WorkItem>> WorkQueue{};
std::vector<std::unique_ptr<WorkItem>> GCArray{};
Event StartWork{};
std::atomic_bool ShuttingDown{false};
};
}
+172 -87
View File
@@ -9,11 +9,11 @@ $end_info$
#include "Interface/Context/Context.h"
#include "Interface/Core/LookupCache.h"
#include "Interface/Core/CompileService.h"
#include "Interface/Core/Core.h"
#include "Interface/Core/CPUID.h"
#include "Interface/Core/Frontend.h"
#include "Interface/Core/GdbServer.h"
#include "Interface/Core/ObjectCache/ObjectCacheService.h"
#include "Interface/Core/OpcodeDispatcher.h"
#include "Interface/Core/Interpreter/InterpreterCore.h"
#include "Interface/Core/JIT/JITCore.h"
@@ -42,6 +42,7 @@ $end_info$
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/Threads.h>
#include <FEXHeaderUtils/Syscalls.h>
#include <FEXHeaderUtils/TodoDefines.h>
#include <algorithm>
#include <array>
@@ -147,10 +148,17 @@ namespace FEXCore::Context {
#ifdef BLOCKSTATS
BlockData = std::make_unique<FEXCore::BlockSamplingData>();
#endif
if (Config.CacheObjectCodeCompilation() != FEXCore::Config::ConfigObjectCodeHandler::CONFIG_NONE) {
CodeObjectCacheService = std::make_unique<FEXCore::CodeSerialize::CodeObjectSerializeService>(this);
}
}
Context::~Context() {
{
if (CodeObjectCacheService) {
CodeObjectCacheService->Shutdown();
}
for (auto &Thread : Threads) {
if (Thread->ExecutionThread->joinable()) {
Thread->ExecutionThread->join(nullptr);
@@ -158,10 +166,6 @@ namespace FEXCore::Context {
}
for (auto &Thread : Threads) {
if (Thread->CompileService) {
Thread->CompileService->Shutdown();
}
delete Thread;
}
Threads.clear();
@@ -190,6 +194,32 @@ namespace FEXCore::Context {
}
FEXCore::Core::InternalThreadState* Context::InitCore(FEXCore::CodeLoader *Loader) {
// Initialize the CPU core signal handlers
switch (Config.Core) {
#ifdef INTERPRETER_ENABLED
case FEXCore::Config::CONFIG_INTERPRETER:
FEXCore::CPU::InitializeInterpreterSignalHandlers(this);
break;
#endif
case FEXCore::Config::CONFIG_IRJIT:
#if (_M_X86_64 && JIT_X86_64)
FEXCore::CPU::InitializeX86JITSignalHandlers(this);
#elif (_M_ARM_64 && JIT_ARM64)
FEXCore::CPU::InitializeArm64JITSignalHandlers(this);
#else
ERROR_AND_DIE_FMT("FEXCore has been compiled without a viable JIT core");
#endif
break;
case FEXCore::Config::CONFIG_CUSTOM:
// Do nothing
break;
default:
ERROR_AND_DIE_FMT("Unknown core configuration");
break;
}
// Initialize GDBServer after the signal handlers are installed
// It may install its own handlers that need to be executed AFTER the CPU cores
if (Config.GdbServer) {
StartGdbServer();
}
@@ -322,6 +352,9 @@ namespace FEXCore::Context {
// Walk the threads and tell them to clear their caches
// Useful when our block size is set to a large number and we need to step a single instruction
for (auto &Thread : Threads) {
// Wait for thread to be fully constructed
// XXX: Look into thread partial construction issues
while(Thread->RunningEvents.WaitingToStart.load()) ;
ClearCodeCache(Thread, true);
}
}
@@ -426,6 +459,7 @@ namespace FEXCore::Context {
: nullptr),
decltype(Entry.DebugData)(new Core::DebugData())
};
std::lock_guard<std::recursive_mutex> lk(Thread->LookupCache->WriteLock);
Thread->LocalIRCache.insert({Addr, std::move(Entry)});
};
@@ -468,7 +502,7 @@ namespace FEXCore::Context {
Thread->StartRunning.NotifyAll();
}
void Context::InitializeCompiler(FEXCore::Core::InternalThreadState* State, bool CompileThread) {
void Context::InitializeCompiler(FEXCore::Core::InternalThreadState* State) {
State->OpDispatcher = std::make_unique<FEXCore::IR::OpDispatchBuilder>(this);
State->OpDispatcher->SetMultiblock(Config.Multiblock);
State->LookupCache = std::make_unique<FEXCore::LookupCache>(this);
@@ -486,23 +520,25 @@ namespace FEXCore::Context {
bool DoSRA = false;
#endif
State->PassManager->AddDefaultPasses(Config.Core == FEXCore::Config::CONFIG_IRJIT, DoSRA);
State->PassManager->AddDefaultPasses(this, Config.Core == FEXCore::Config::CONFIG_IRJIT, DoSRA);
State->PassManager->AddDefaultValidationPasses();
State->PassManager->RegisterSyscallHandler(SyscallHandler);
// Create CPU backend
switch (Config.Core) {
#ifdef INTERPRETER_ENABLED
case FEXCore::Config::CONFIG_INTERPRETER:
State->CPUBackend = FEXCore::CPU::CreateInterpreterCore(this, State, CompileThread);
State->CPUBackend = FEXCore::CPU::CreateInterpreterCore(this, State);
break;
#endif
case FEXCore::Config::CONFIG_IRJIT:
State->PassManager->InsertRegisterAllocationPass(DoSRA);
#if (_M_X86_64 && JIT_X86_64)
State->CPUBackend = FEXCore::CPU::CreateX86JITCore(this, State, CompileThread);
State->CPUBackend = FEXCore::CPU::CreateX86JITCore(this, State);
#elif (_M_ARM_64 && JIT_ARM64)
State->CPUBackend = FEXCore::CPU::CreateArm64JITCore(this, State, CompileThread);
State->CPUBackend = FEXCore::CPU::CreateArm64JITCore(this, State);
#else
ERROR_AND_DIE_FMT("FEXCore has been compiled without a viable JIT core");
#endif
@@ -534,7 +570,7 @@ namespace FEXCore::Context {
// Set up the thread manager state
Thread->ThreadManager.parent_tid = ParentTID;
InitializeCompiler(Thread, false);
InitializeCompiler(Thread);
InitializeThreadData(Thread);
return Thread;
@@ -596,24 +632,25 @@ namespace FEXCore::Context {
// Clean up dead stacks
FEXCore::Threads::Thread::CleanupAfterFork();
if (LiveThread->CompileService) {
// If this live thread had a compile service then it no longer exists
// Erase the shared_ptr
LiveThread->CompileService.reset();
}
}
void Context::AddBlockMapping(FEXCore::Core::InternalThreadState *Thread, uint64_t Address, void *Ptr, uint64_t Start, uint64_t Length) {
Thread->LookupCache->AddBlockMapping(Address, Ptr, Start, Length);
// Only call MarkGuestExecutableRange if new pages are marked as containing code
if (Thread->LookupCache->AddBlockMapping(Address, Ptr, Start, Length)) {
Thread->CTX->SyscallHandler->MarkGuestExecutableRange(Start, Length);
}
}
void Context::ClearCodeCache(FEXCore::Core::InternalThreadState *Thread, bool AlsoClearIRCache) {
{
// Ensure the Code Object Serialization service has fully serialized this thread's data before clearing the cache
// Use the thread's object cache ref counter for this
CodeSerialize::CodeObjectSerializeService::WaitForEmptyJobQueue(&Thread->ObjectCacheRefCounter);
}
std::lock_guard<std::recursive_mutex> lk(Thread->LookupCache->WriteLock);
Thread->LookupCache->ClearCache();
Thread->CPUBackend->ClearCache();
if (Thread->CompileService) {
Thread->CompileService->ClearCache(Thread);
}
if (AlsoClearIRCache) {
Thread->LocalIRCache.clear();
@@ -649,15 +686,16 @@ namespace FEXCore::Context {
}
};
static void ValidateIR(FEXCore::Core::InternalThreadState *Thread) {
static void ValidateIR(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread) {
// Convert to text, Parse, Convert to text again and make sure the texts match
std::stringstream out;
static auto compaction = IR::CreateIRCompaction();
static auto compaction = IR::CreateIRCompaction(ctx->OpDispatcherAllocator);
compaction->Run(Thread->OpDispatcher.get());
auto NewIR = Thread->OpDispatcher->ViewIR();
Dump(&out, &NewIR, nullptr);
out.seekg(0);
auto reparsed = IR::Parse(&out);
FEXCore::Utils::PooledAllocatorMalloc Allocator;
auto reparsed = IR::Parse(Allocator, &out);
if (reparsed == nullptr) {
LOGMAN_MSG_A_FMT("Failed to parse IR\n");
} else {
@@ -685,6 +723,8 @@ namespace FEXCore::Context {
auto CodeBlocks = Thread->FrontendDecoder->GetDecodedBlocks();
Thread->OpDispatcher->ReownOrClaimBuffer();
Thread->OpDispatcher->ResetWorkingList();
Thread->OpDispatcher->BeginFunction(GuestRIP, CodeBlocks);
const uint8_t GPRSize = GetGPRSize();
@@ -721,7 +761,7 @@ namespace FEXCore::Context {
Thread->OpDispatcher->SetTrueJumpTarget(InvalidateCodeCond, CodeWasChangedBlock);
Thread->OpDispatcher->SetCurrentCodeBlock(CodeWasChangedBlock);
Thread->OpDispatcher->_RemoveCodeEntry();
Thread->OpDispatcher->_RemoveThreadCodeEntry();
Thread->OpDispatcher->_ExitFunction(Thread->OpDispatcher->_EntrypointOffset(Block.Entry + BlockInstructionsLength - GuestRIP, GPRSize));
auto NextOpBlock = Thread->OpDispatcher->CreateNewCodeBlockAfter(CurrentBlock);
@@ -785,7 +825,7 @@ namespace FEXCore::Context {
}
if (Thread->CTX->Config.ValidateIRarser) {
ValidateIR(Thread);
ValidateIR(this, Thread);
}
}
@@ -809,7 +849,8 @@ namespace FEXCore::Context {
auto RAData = Thread->PassManager->HasPass("RA") ? Thread->PassManager->GetPass<IR::RegisterAllocationPass>("RA")->PullAllocationData() : nullptr;
auto IRList = Thread->OpDispatcher->CreateIRCopy();
Thread->OpDispatcher->ResetWorkingList();
Thread->OpDispatcher->DelayedDisownBuffer();
Thread->FrontendDecoder->DelayedDisownBuffer();
return {
.IRList = IRList,
@@ -829,6 +870,7 @@ namespace FEXCore::Context {
uint64_t StartAddr {};
uint64_t Length {};
std::lock_guard<std::recursive_mutex> lk(Thread->LookupCache->WriteLock);
// Do we already have this in the IR cache?
auto LocalEntry = Thread->LocalIRCache.find(GuestRIP);
@@ -844,6 +886,25 @@ namespace FEXCore::Context {
GeneratedIR = false;
}
// JIT Code object cache lookup
if (CodeObjectCacheService) {
auto CodeCacheEntry = CodeObjectCacheService->FetchCodeObjectFromCache(GuestRIP);
if (CodeCacheEntry) {
auto CompiledCode = Thread->CPUBackend->RelocateJITObjectCode(GuestRIP, CodeCacheEntry);
if (CompiledCode) {
return {
.CompiledCode = CompiledCode,
.IRData = nullptr, // No IR data generated
.DebugData = nullptr, // nullptr here ensures that code serialization doesn't occur on from cache read
.RAData = nullptr, // No RA data generated
.GeneratedIR = false, // nullptr here ensures IR cache mechanisms won't run
.StartAddr = 0, // Unused
.Length = 0, // Unused
};
}
}
}
// AOT IR bookkeeping and cache
{
auto [IRCopy, RACopy, DebugDataCopy, _StartAddr, _Length, _GeneratedIR] = IRCaptureCache.PreGenerateIRFetch(GuestRIP, IRList);
@@ -869,10 +930,6 @@ namespace FEXCore::Context {
StartAddr = _StartAddr;
Length = _Length;
// Initialize metadata
DebugData->GuestCodeSize = TotalInstructionsLength;
DebugData->GuestInstructionCount = TotalInstructions;
// Increment stats
Thread->Stats.BlocksCompiled.fetch_add(1);
@@ -909,6 +966,9 @@ namespace FEXCore::Context {
uintptr_t Context::CompileBlock(FEXCore::Core::CpuStateFrame *Frame, uint64_t GuestRIP) {
auto Thread = Frame->Thread;
// Needs to be held for SMC interactions around concurrent compile and invalidation hazards
auto InvalidationLk = Thread->CTX->SyscallHandler->CompileCodeLock(GuestRIP);
// Is the code in the cache?
// The backends only check L1 and L2, not L3
if (auto HostCode = Thread->LookupCache->FindBlock(GuestRIP)) {
@@ -920,63 +980,66 @@ namespace FEXCore::Context {
FEXCore::Core::DebugData *DebugData {};
FEXCore::IR::RegisterAllocationData *RAData {};
bool DecrementRefCount = false;
bool GeneratedIR {};
uint64_t StartAddr {}, Length {};
if (Thread->CompileBlockReentrantRefCount != 0) {
if (!Thread->CompileService) {
Thread->CompileService = std::make_shared<FEXCore::CompileService>(this, Thread);
Thread->CompileService->Initialize();
}
auto* WorkItem = Thread->CompileService->CompileCode(GuestRIP);
WorkItem->ServiceWorkDone.Wait();
// Return here with the data in place
CodePtr = WorkItem->CodePtr;
IRList = WorkItem->IRList;
DebugData = WorkItem->DebugData;
RAData = WorkItem->RAData;
StartAddr = WorkItem->StartAddr;
Length = WorkItem->Length;
WorkItem->SafeToClear = true;
// The compile service will always generate IR + DebugData + RAData
// Remove the entries here to make sure we don't fail to insert later on
RemoveCodeEntry(Thread, GuestRIP);
GeneratedIR = true;
} else {
++Thread->CompileBlockReentrantRefCount;
DecrementRefCount = true;
auto [Code, IR, Data, RA, Generated, _StartAddr, _Length] = CompileCode(Thread, GuestRIP);
CodePtr = Code;
IRList = IR;
DebugData = Data;
RAData = RA;
GeneratedIR = Generated;
StartAddr = _StartAddr;
Length = _Length;
}
auto [Code, IR, Data, RA, Generated, _StartAddr, _Length] = CompileCode(Thread, GuestRIP);
CodePtr = Code;
IRList = IR;
DebugData = Data;
RAData = RA;
GeneratedIR = Generated;
StartAddr = _StartAddr;
Length = _Length;
if (CodePtr == nullptr) {
if (DecrementRefCount)
--Thread->CompileBlockReentrantRefCount;
return 0;
}
// The core managed to compile the code.
if (Config.BlockJITNaming()) {
if (DebugData) {
auto GuestRIPLookup = this->SyscallHandler->LookupAOTIRCacheEntry(GuestRIP);
if (DebugData->Subblocks.size()) {
for (auto& Subblock: DebugData->Subblocks) {
if (GuestRIPLookup.Entry) {
Symbols.Register(CodePtr, DebugData->HostCodeSize, GuestRIPLookup.Entry->Filename, GuestRIP - GuestRIPLookup.Offset);
} else {
Symbols.Register((void*)Subblock.HostCodeStart, GuestRIP, Subblock.HostCodeSize);
}
}
} else {
if (GuestRIPLookup.Entry) {
Symbols.Register(CodePtr, DebugData->HostCodeSize, GuestRIPLookup.Entry->Filename, GuestRIP - GuestRIPLookup.Offset);
} else {
Symbols.Register(CodePtr, GuestRIP, DebugData->HostCodeSize);
}
}
}
}
// Tell the object cache service to serialize the code if enabled
if (CodeObjectCacheService &&
Config.CacheObjectCodeCompilation == FEXCore::Config::ConfigObjectCodeHandler::CONFIG_READWRITE &&
DebugData) {
CodeObjectCacheService->AsyncAddSerializationJob(std::make_unique<CodeSerialize::AsyncJobHandler::SerializationJobData>(
CodeSerialize::AsyncJobHandler::SerializationJobData {
.GuestRIP = GuestRIP,
.GuestCodeLength = Length,
.GuestCodeHash = 0,
.HostCodeBegin = CodePtr,
.HostCodeLength = DebugData->HostCodeSize,
.HostCodeHash = 0,
.ThreadJobRefCount = &Thread->ObjectCacheRefCounter,
.Relocations = std::move(*DebugData->Relocations),
}
));
}
// Clear any relocations that might have been generated
Thread->CPUBackend->ClearRelocations();
if (IRCaptureCache.PostCompileCode(
Thread,
CodePtr,
@@ -986,15 +1049,11 @@ namespace FEXCore::Context {
RAData,
IRList,
DebugData,
GeneratedIR,
DecrementRefCount)) {
GeneratedIR)) {
// Early exit
return (uintptr_t)CodePtr;
}
if (DecrementRefCount)
--Thread->CompileBlockReentrantRefCount;
// Insert to lookup cache
AddBlockMapping(Thread, GuestRIP, CodePtr, StartAddr, Length);
@@ -1029,8 +1088,15 @@ namespace FEXCore::Context {
Thread->RunningEvents.Running = false;
}
{
// Ensure the Code Object Serialization service has fully serialized this thread's data before clearing the cache
// Use the thread's object cache ref counter for this
CodeSerialize::CodeObjectSerializeService::WaitForEmptyJobQueue(&Thread->ObjectCacheRefCounter);
}
// If it is the parent thread that died then just leave
// XXX: This doesn't make sense when the parent thread doesn't outlive its children
FEX_TODO("This doesn't make sense when the parent thread doesn't outlive its children");
if (Thread->ThreadManager.parent_tid == 0) {
CoreShuttingDown.store(true);
Thread->ExitReason = FEXCore::Context::ExitReason::EXIT_SHUTDOWN;
@@ -1051,21 +1117,32 @@ namespace FEXCore::Context {
}
}
void FlushCodeRange(FEXCore::Core::InternalThreadState *Thread, uint64_t Start, uint64_t Length) {
static void InvalidateGuestThreadCodeRange(FEXCore::Core::InternalThreadState *Thread, uint64_t Start, uint64_t Length) {
std::lock_guard<std::recursive_mutex> lk(Thread->LookupCache->WriteLock);
if (Thread->CTX->Config.SMCChecks == FEXCore::Config::CONFIG_SMC_MMAN) {
auto lower = Thread->LookupCache->CodePages.lower_bound(Start >> 12);
auto upper = Thread->LookupCache->CodePages.upper_bound((Start + Length) >> 12);
auto lower = Thread->LookupCache->CodePages.lower_bound(Start >> 12);
auto upper = Thread->LookupCache->CodePages.upper_bound((Start + Length - 1) >> 12);
for (auto it = lower; it != upper; it++) {
for (auto Address: it->second)
Context::RemoveCodeEntry(Thread, Address);
it->second.clear();
for (auto it = lower; it != upper; it++) {
for (auto Address: it->second) {
Context::RemoveThreadCodeEntry(Thread, Address);
}
it->second.clear();
}
}
void InvalidateGuestCodeRange(FEXCore::Context::Context *CTX, uint64_t Start, uint64_t Length) {
std::lock_guard<std::mutex> lk(CTX->ThreadCreationMutex);
for (auto &Thread : CTX->Threads) {
if (Thread->RunningEvents.Running.load()) {
InvalidateGuestThreadCodeRange(Thread, Start, Length);
}
}
}
void Context::RemoveCodeEntry(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP) {
void Context::RemoveThreadCodeEntry(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP) {
std::lock_guard<std::recursive_mutex> lk(Thread->LookupCache->WriteLock);
Thread->LocalIRCache.erase(GuestRIP);
Thread->LookupCache->Erase(GuestRIP);
}
@@ -1076,7 +1153,7 @@ namespace FEXCore::Context {
Thread->CurrentFrame->State.rip = RIP;
// Erase the RIP from all the storage backings if it exists
RemoveCodeEntry(Thread, RIP);
RemoveThreadCodeEntry(Thread, RIP);
// We don't care if compilation passes or not
CompileBlock(Thread->CurrentFrame, RIP);
@@ -1093,6 +1170,7 @@ namespace FEXCore::Context {
}
bool Context::GetDebugDataForRIP(uint64_t RIP, FEXCore::Core::DebugData *Data) {
std::lock_guard<std::recursive_mutex> lk(ParentThread->LookupCache->WriteLock);
auto it = ParentThread->LocalIRCache.find(RIP);
if (it == ParentThread->LocalIRCache.end()) {
return false;
@@ -1118,12 +1196,19 @@ namespace FEXCore::Context {
return Result;
}
void Context::AddNamedRegion(uintptr_t Base, uintptr_t Size, uintptr_t Offset, const std::string &filename) {
IRCaptureCache.AddNamedRegion(Base, Size, Offset, filename);
IR::AOTIRCacheEntry *Context::LoadAOTIRCacheEntry(const std::string &filename) {
auto rv = IRCaptureCache.LoadAOTIRCacheEntry(filename);
if (DebugServer) {
DebugServer->AlertLibrariesChanged();
}
return rv;
}
void Context::RemoveNamedRegion(uintptr_t Base, uintptr_t Size) {
IRCaptureCache.RemoveNamedRegion(Base, Size);
void Context::UnloadAOTIRCacheEntry(IR::AOTIRCacheEntry *Entry) {
IRCaptureCache.UnloadAOTIRCacheEntry(Entry);
if (DebugServer) {
DebugServer->AlertLibrariesChanged();
}
}
void ConfigureAOTGen(FEXCore::Core::InternalThreadState *Thread, std::set<uint64_t> *ExternalBranches, uint64_t SectionMaxAddress) {
@@ -39,7 +39,7 @@ static constexpr size_t MAX_DISPATCHER_CODE_SIZE = 4096;
#define STATE x28
Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, DispatcherConfig &config)
: Dispatcher(ctx, Thread), Arm64Emitter(ctx, MAX_DISPATCHER_CODE_SIZE) {
: FEXCore::CPU::Dispatcher(ctx, Thread), Arm64Emitter(ctx, MAX_DISPATCHER_CODE_SIZE) {
SRAEnabled = config.StaticRegisterAssignment;
SetAllowAssembler(true);
@@ -53,10 +53,7 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
// Ptr();
// }
uint64_t VirtualMemorySize = Thread->LookupCache->GetVirtualMemorySize();
Literal l_VirtualMemory {VirtualMemorySize};
Literal l_PagePtr {Thread->LookupCache->GetPagePointer()};
Literal l_L1Ptr {Thread->LookupCache->GetL1Pointer()};
Literal l_CTX {reinterpret_cast<uintptr_t>(CTX)};
Literal l_Sleep {reinterpret_cast<uint64_t>(SleepThread)};
Literal l_CompileBlock {GetCompileBlockPtr()};
@@ -99,7 +96,7 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
auto RipReg = x2;
// L1 Cache
ldr(x0, &l_L1Ptr);
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.L1Pointer)));
and_(x3, RipReg, LookupCache::L1_ENTRIES_MASK);
add(x0, x0, Operand(x3, Shift::LSL, 4));
@@ -121,11 +118,12 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
ldr(x0, &l_PagePtr);
// Mask the address by the virtual address size so we can check for aliases
uint64_t VirtualMemorySize = Thread->LookupCache->GetVirtualMemorySize();
if (std::popcount(VirtualMemorySize) == 1) {
and_(x3, RipReg, Thread->LookupCache->GetVirtualMemorySize() - 1);
and_(x3, RipReg, VirtualMemorySize - 1);
}
else {
ldr(x3, &l_VirtualMemory);
LoadConstant(x3, VirtualMemorySize);
and_(x3, RipReg, x3);
}
@@ -159,7 +157,7 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
// If we've made it here then we have a real compiled block
{
// update L1 cache
ldr(x0, &l_L1Ptr);
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.L1Pointer)));
and_(x1, RipReg, LookupCache::L1_ENTRIES_MASK);
add(x0, x0, Operand(x1, Shift::LSL, 4));
@@ -438,9 +436,89 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
b(&LoopTop);
}
place(&l_VirtualMemory);
// Long division helpers
uint64_t LUDIVHandler{};
uint64_t LDIVHandler{};
uint64_t LUREMHandler{};
uint64_t LREMHandler{};
{
LUDIVHandler = GetCursorAddress<uint64_t>();
PushDynamicRegsAndLR();
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.LUDIV)));
SpillStaticRegs();
blr(x3);
FillStaticRegs();
// Result is now in x0
// Fix the stack and any values that were stepped on
PopDynamicRegsAndLR();
// Go back to our code block
ret();
}
{
LDIVHandler = GetCursorAddress<uint64_t>();
PushDynamicRegsAndLR();
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.LDIV)));
SpillStaticRegs();
blr(x3);
FillStaticRegs();
// Result is now in x0
// Fix the stack and any values that were stepped on
PopDynamicRegsAndLR();
// Go back to our code block
ret();
}
{
LUREMHandler = GetCursorAddress<uint64_t>();
PushDynamicRegsAndLR();
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.LUREM)));
SpillStaticRegs();
blr(x3);
FillStaticRegs();
// Result is now in x0
// Fix the stack and any values that were stepped on
PopDynamicRegsAndLR();
// Go back to our code block
ret();
}
{
LREMHandler = GetCursorAddress<uint64_t>();
PushDynamicRegsAndLR();
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.LREM)));
SpillStaticRegs();
blr(x3);
FillStaticRegs();
// Result is now in x0
// Fix the stack and any values that were stepped on
PopDynamicRegsAndLR();
// Go back to our code block
ret();
}
place(&l_PagePtr);
place(&l_L1Ptr);
place(&l_CTX);
place(&l_Sleep);
place(&l_CompileBlock);
@@ -461,6 +539,24 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
if (CTX->Config.GlobalJITNaming()) {
CTX->Symbols.RegisterJITSpace(reinterpret_cast<void*>(DispatchPtr), End - reinterpret_cast<uint64_t>(DispatchPtr));
}
// Setup dispatcher specific pointers that need to be accessed from JIT code
{
auto &Pointers = ThreadState->CurrentFrame->Pointers.AArch64;
Pointers.DispatcherLoopTop = AbsoluteLoopTopAddress;
Pointers.DispatcherLoopTopFillSRA = AbsoluteLoopTopAddressFillSRA;
Pointers.ThreadStopHandlerSpillSRA = ThreadStopHandlerAddressSpillSRA;
Pointers.ThreadPauseHandlerSpillSRA = ThreadPauseHandlerAddressSpillSRA;
Pointers.UnimplementedInstructionHandler = UnimplementedInstructionAddress;
Pointers.OverflowExceptionHandler = OverflowExceptionInstructionAddress;
Pointers.SignalReturnHandler = SignalHandlerReturnAddress;
Pointers.L1Pointer = Thread->LookupCache->GetL1Pointer();
Pointers.LUDIVHandler = LUDIVHandler;
Pointers.LDIVHandler = LDIVHandler;
Pointers.LUREMHandler = LUREMHandler;
Pointers.LREMHandler = LREMHandler;
}
}
void Arm64Dispatcher::SpillSRA(void *ucontext, uint32_t IgnoreMask) {
@@ -2,7 +2,6 @@
#include "Interface/Core/ArchHelpers/MContext.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Core/X86HelperGen.h"
#include "Interface/Core/CompileService.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/CoreState.h>
@@ -16,9 +15,9 @@
#include <atomic>
#include <condition_variable>
#include <bits/types/siginfo_t.h>
#include <csignal>
#include <cstring>
#include <signal.h>
namespace FEXCore::CPU {
@@ -759,7 +758,7 @@ void Dispatcher::RemoveCodeBuffer(uint8_t* start_to_remove) {
}
}
bool Dispatcher::IsAddressInJITCode(uint64_t Address, bool IncludeDispatcher, bool IncludeCompileService) const {
bool Dispatcher::IsAddressInJITCode(uint64_t Address, bool IncludeDispatcher) const {
for (auto [start, end] : CodeBuffers) {
if (Address >= start && Address < end) {
return true;
@@ -770,9 +769,6 @@ bool Dispatcher::IsAddressInJITCode(uint64_t Address, bool IncludeDispatcher, bo
return true;
}
if (IncludeCompileService && ThreadState->CompileService && ThreadState->CompileService->IsAddressInJITCode(Address)) {
return true;
}
return false;
}
@@ -5,8 +5,8 @@
#include "Interface/Context/Context.h"
#include "Interface/Core/ArchHelpers/MContext.h"
#include <bits/types/stack_t.h>
#include <cstdint>
#include <signal.h>
#include <stddef.h>
#include <stack>
#include <tuple>
@@ -77,7 +77,7 @@ public:
void RemoveCodeBuffer(uint8_t* start);
bool IsAddressInJITCode(uint64_t Address, bool IncludeDispatcher = true, bool IncludeCompileService = true) const;
bool IsAddressInJITCode(uint64_t Address, bool IncludeDispatcher = true) const;
bool IsAddressInDispatcher(uint64_t Address) const {
return Address >= Start && Address < End;
}
@@ -95,7 +95,7 @@ X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::Inte
mov(rdx, qword [STATE + offsetof(FEXCore::Core::CPUState, rip)]);
// L1 Cache
mov(r13, Thread->LookupCache->GetL1Pointer());
mov(r13, qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.L1Pointer)]);
mov(rax, rdx);
and_(rax, LookupCache::L1_ENTRIES_MASK);
@@ -114,8 +114,9 @@ X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::Inte
mov(r13, Thread->LookupCache->GetPagePointer());
// Full lookup
uint64_t VirtualMemorySize = Thread->LookupCache->GetVirtualMemorySize();
mov(rax, rdx);
mov(rbx, Thread->LookupCache->GetVirtualMemorySize() - 1);
mov(rbx, VirtualMemorySize - 1);
and_(rax, rbx);
shr(rax, 12);
@@ -142,8 +143,7 @@ X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::Inte
je(NoBlock);
// Update L1
mov(r13, Thread->LookupCache->GetL1Pointer());
mov(r13, qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.L1Pointer)]);
mov(rcx, rdx);
and_(rcx, LookupCache::L1_ENTRIES_MASK);
shl(rcx, 1);
@@ -193,10 +193,37 @@ X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::Inte
ret();
}
constexpr bool SignalSafeCompile = true;
// Block creation
{
L(NoBlock);
if (SignalSafeCompile) {
// When compiling code, mask all signals to reduce the chance of reentrant allocations
// RDI: SETMASK
// RSI: Pointer to mask value (uint64_t)
// RDX: Pointer to old mask value (uint64_t)
// R10: Size of mask, sizeof(uint64_t)
// RAX: Syscall
// Backup rdx
mov(r9, rdx);
mov(rdi, ~0ULL);
sub(rsp, 16);
mov(qword [rsp], rdi);
mov(qword [rsp + 8], rdi);
mov(rdi, SIG_SETMASK);
mov(rsi, rsp);
mov(rdx, rsp);
mov(r10, 8);
mov(rax, SYS_rt_sigprocmask);
syscall();
mov(rdx, r9);
}
// {rdi, rsi, rdx}
mov(rdi, reinterpret_cast<uint64_t>(CTX));
mov(rsi, STATE);
@@ -204,12 +231,57 @@ X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::Inte
call(rax);
if (SignalSafeCompile) {
// Now restore the signal mask
// Living in the same location
// Backup rdx
mov(r9, rdx);
mov(rdi, SIG_SETMASK);
mov(rsi, rsp);
mov(rdx, 0); // Don't care about result
mov(r10, 8);
mov(rax, SYS_rt_sigprocmask);
syscall();
// Bring stack back
add(rsp, 16);
mov(rdx, r9);
}
// rdx already contains RIP here
jmp(LoopTop);
}
{
ExitFunctionLinkerAddress = getCurr<uint64_t>();
if (SignalSafeCompile) {
// When compiling code, mask all signals to reduce the chance of reentrant allocations
// RDI: SETMASK
// RSI: Pointer to mask value (uint64_t)
// RDX: Pointer to old mask value (uint64_t)
// R10: Size of mask, sizeof(uint64_t)
// RAX: Syscall
// Backup rax
mov(r9, rax);
mov(rdi, ~0ULL);
sub(rsp, 16);
mov(qword [rsp], rdi);
mov(qword [rsp + 8], rdi);
mov(rdi, SIG_SETMASK);
mov(rsi, rsp);
mov(rdx, rsp);
mov(r10, 8);
mov(rax, SYS_rt_sigprocmask);
syscall();
mov(rax, r9);
}
// {rdi, rsi, rdx}
mov(rdi, config.ExitFunctionLinkThis);
mov(rsi, STATE);
@@ -217,7 +289,28 @@ X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::Inte
mov(rax, config.ExitFunctionLink);
call(rax);
jmp(rax);
if (SignalSafeCompile) {
// Now restore the signal mask
// Living in the same location
// Backup rax
mov(r9, rax);
mov(rdi, SIG_SETMASK);
mov(rsi, rsp);
mov(rdx, 0); // Don't care about result
mov(r10, 8);
mov(rax, SYS_rt_sigprocmask);
syscall();
// Bring stack back
add(rsp, 16);
jmp(r9);
}
else {
jmp(rax);
}
}
{
@@ -252,8 +345,7 @@ X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::Inte
// XXX: XMM?
// Make sure to adjust the refcounter so we don't clear the cache now
mov(rax, reinterpret_cast<uint64_t>(&SignalHandlerRefCounter));
add(dword [rax], 1);
add(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.SignalHandlerRefCountPointer)], 1);
// Now push the callback return trampoline to the guest stack
// Guest will be misaligned because calling a thunk won't correct the guest's stack once we call the callback from the host
@@ -337,6 +429,20 @@ X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::Inte
if (CTX->Config.GlobalJITNaming()) {
CTX->Symbols.RegisterJITSpace(reinterpret_cast<void*>(Start), End-Start);
}
// Setup dispatcher specific pointers that need to be accessed from JIT code
{
auto &Pointers = ThreadState->CurrentFrame->Pointers.X86;
Pointers.DispatcherLoopTop = AbsoluteLoopTopAddress;
Pointers.DispatcherLoopTopFillSRA = AbsoluteLoopTopAddressFillSRA;
Pointers.ThreadStopHandler = ThreadStopHandlerAddress;
Pointers.ThreadPauseHandler = ThreadPauseHandlerAddress;
Pointers.UnimplementedInstructionHandler = UnimplementedInstructionAddress;
Pointers.OverflowExceptionHandler = OverflowExceptionInstructionAddress;
Pointers.SignalReturnHandler = SignalHandlerReturnAddress;
Pointers.L1Pointer = Thread->LookupCache->GetL1Pointer();
}
}
X86Dispatcher::~X86Dispatcher() {
+27 -25
View File
@@ -177,17 +177,12 @@ static uint32_t MapVEXToReg(uint8_t vvvv, bool HasXMM) {
Decoder::Decoder(FEXCore::Context::Context *ctx)
: CTX {ctx}
, OSABI { ctx->SyscallHandler ? ctx->SyscallHandler->GetOSABI() : FEXCore::HLE::SyscallOSABI::OS_UNKNOWN } {
// Using mmap is a start-up time optimization
// Take advantage of page faulting to reduce startup time for minimal runtime cost
DecodedBuffer =
reinterpret_cast<FEXCore::X86Tables::DecodedInst *>(
FEXCore::Allocator::mmap(0, sizeof(FEXCore::X86Tables::DecodedInst) * DefaultDecodedBufferSize,
PROT_READ | PROT_WRITE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0));
, OSABI { ctx->SyscallHandler ? ctx->SyscallHandler->GetOSABI() : FEXCore::HLE::SyscallOSABI::OS_UNKNOWN }
, PoolObject {ctx->FrontendAllocator, sizeof(FEXCore::X86Tables::DecodedInst) * DefaultDecodedBufferSize} {
}
Decoder::~Decoder() {
FEXCore::Allocator::munmap(DecodedBuffer, sizeof(FEXCore::X86Tables::DecodedInst) * DefaultDecodedBufferSize);
PoolObject.UnclaimBuffer();
}
uint8_t Decoder::ReadByte() {
@@ -583,8 +578,11 @@ bool Decoder::NormalOp(FEXCore::X86Tables::X86InstInfo const *Info, uint16_t Op,
return false;
}
else {
auto Disp = DecodeModRMs_Disp[Has16BitAddressing];
(this->*Disp)(&NonGPR, ModRM);
// Only decode if we haven't pre-decoded
if (NonGPR.IsNone()) {
auto Disp = DecodeModRMs_Disp[Has16BitAddressing];
(this->*Disp)(&NonGPR, ModRM);
}
}
return true;
@@ -857,7 +855,7 @@ bool Decoder::DecodeInstruction(uint64_t PC) {
case 0x0F: {// Escape Op
uint8_t EscapeOp = ReadByte();
switch (EscapeOp) {
case 0x0F: { // 3DNow!
case 0x0F: [[unlikely]] { // 3DNow!
// 3DNow! Instruction Encoding: 0F 0F [ModRM] [SIB] [Displacement] [Opcode]
// Decode ModRM
uint8_t ModRMByte = ReadByte();
@@ -870,8 +868,12 @@ bool Decoder::DecodeInstruction(uint64_t PC) {
const bool Has16BitAddressing = !CTX->Config.Is64BitMode &&
DecodeInst->Flags & DecodeFlags::FLAG_ADDRESS_SIZE;
auto Disp = DecodeModRMs_Disp[Has16BitAddressing];
(this->*Disp)(&DecodeInst->Src[0], ModRM);
// All 3DNow! instructions have the second argument as the rm handler
// We need to decode it upfront to get the displacement out of the way
if (ModRM.mod != 0b11) {
auto Disp = DecodeModRMs_Disp[Has16BitAddressing];
(this->*Disp)(&DecodeInst->Src[0], ModRM);
}
// Take a peek at the op just past the displacement
uint8_t LocalOp = ReadByte();
@@ -880,20 +882,19 @@ bool Decoder::DecodeInstruction(uint64_t PC) {
}
case 0x38: { // F38 Table!
constexpr uint16_t PF_38_NONE = 0;
constexpr uint16_t PF_38_66 = 1;
constexpr uint16_t PF_38_F2 = 2;
constexpr uint16_t PF_38_F3 = 3;
constexpr uint16_t PF_38_66 = (1U << 0);
constexpr uint16_t PF_38_F2 = (1U << 1);
constexpr uint16_t PF_38_F3 = (1U << 2);
uint16_t Prefix = PF_38_NONE;
if (DecodeInst->LastEscapePrefix == 0xF2) {
// Repeat prefix or instruction-specific
Prefix = PF_38_F2;
} else if (DecodeInst->LastEscapePrefix == 0xF3) {
// Repeat prefix or instruction-specific
Prefix = PF_38_F3;
} else if (DecodeInst->LastEscapePrefix == 0x66) {
// Operand size
Prefix = PF_38_66;
if (DecodeInst->Flags & DecodeFlags::FLAG_OPERAND_SIZE) {
Prefix |= PF_38_66;
}
if (DecodeInst->Flags & DecodeFlags::FLAG_REPNE_PREFIX) {
Prefix |= PF_38_F2;
}
if (DecodeInst->Flags & DecodeFlags::FLAG_REP_PREFIX) {
Prefix |= PF_38_F3;
}
uint16_t LocalOp = (Prefix << 8) | ReadByte();
@@ -1142,6 +1143,7 @@ void Decoder::DecodeInstructionsAtEntry(uint8_t const* _InstStream, uint64_t PC)
DecodedSize = 0;
MaxCondBranchForward = 0;
MaxCondBranchBackwards = ~0ULL;
DecodedBuffer = PoolObject.ReownOrClaimBuffer();
// XXX: Load symbol data
SymbolAvailable = false;
+7 -1
View File
@@ -35,9 +35,14 @@ public:
uint64_t DecodedMinAddress {};
uint64_t DecodedMaxAddress {~0ULL};
void SetSectionMaxAddress(uint64_t v) { SectionMaxAddress = v; }
void SetExternalBranches(std::set<uint64_t> *v) { ExternalBranches = v; }
void DelayedDisownBuffer() {
PoolObject.DelayedDisownBuffer();
}
private:
// To pass any information from instruction prefixes
// down into the actual instruction handling machinery.
@@ -63,6 +68,7 @@ private:
static constexpr size_t DefaultDecodedBufferSize = 0x10000;
FEXCore::X86Tables::DecodedInst *DecodedBuffer{};
Utils::FixedSizePooledAllocation<FEXCore::X86Tables::DecodedInst*, 5000, 500> PoolObject;
size_t DecodedSize {};
uint8_t const *InstStream;
File diff suppressed because it is too large. Load diff
+14 -1
View File
@@ -6,8 +6,10 @@ $end_info$
#pragma once
#include <FEXCore/Config/Config.h>
#include <FEXCore/Utils/Event.h>
#include <FEXCore/Utils/Threads.h>
#include <atomic>
#include <istream>
#include <memory>
#include <mutex>
@@ -27,6 +29,10 @@ public:
// Public for threading
void GdbServerLoop();
void AlertLibrariesChanged() {
LibraryMapChanged = true;
}
private:
void Break(int signal);
@@ -38,6 +44,9 @@ private:
void SendACK(std::ostream &stream, bool NACK);
Event ThreadBreakEvent{};
void WaitForThreadWakeup();
struct HandledPacketType {
std::string Response{};
enum ResponseType {
@@ -74,9 +83,13 @@ private:
bool NoAckMode{false};
bool NonStopMode{false};
std::string ThreadString{};
std::string MemoryMapString{};
std::string OSDataString{};
void buildLibraryMap();
std::atomic<bool> LibraryMapChanged = true;
std::string LibraryMapString{};
// Used to keep track of which signals to pass to the guest
std::array<bool, SignalDelegator::MAX_SIGNALS + 1> PassSignals{};
uint32_t CurrentDebuggingThread{};
int ListenSocket{};
FEX_CONFIG_OPT(Filename, APP_FILENAME);
+20 -5
View File
@@ -53,14 +53,14 @@ HostFeatures::HostFeatures() {
#ifdef _M_ARM_64
auto Features = vixl::CPUFeatures::InferFromOS();
SupportsAES = Features.Has(vixl::CPUFeatures::Feature::kAES);
SupportsCRC = Features.Has(vixl::CPUFeatures::Feature::kCRC32);
SupportsAtomics = Features.Has(vixl::CPUFeatures::Feature::kAtomics);
SupportsRAND = Features.Has(vixl::CPUFeatures::Feature::kRNG);
// Only supported when FEAT_AFP is supported
SupportsFlushInputsToZero = Features.Has(vixl::CPUFeatures::Feature::kAFP);
// RCPC is bugged on Snapdragon 865
// Causes glibc cond16 test to immediately throw assert
// __pthread_mutex_cond_lock: Assertion `mutex->__data.__owner == 0'
SupportsRCPC = false; //Features.Has(vixl::CPUFeatures::Feature::kRCpc);
SupportsRCPC = Features.Has(vixl::CPUFeatures::Feature::kRCpc);
SupportsTSOImm9 = Features.Has(vixl::CPUFeatures::Feature::kRCpcImm);
// We need to get the CPU's cache line size
// We expect sane targets that have correct cacheline sizes across clusters
@@ -78,6 +78,21 @@ HostFeatures::HostFeatures() {
#ifdef _M_X86_64
Xbyak::util::Cpu Features{};
SupportsAES = Features.has(Xbyak::util::Cpu::tAESNI);
SupportsCRC = Features.has(Xbyak::util::Cpu::tSSE42);
SupportsRAND = Features.has(Xbyak::util::Cpu::tRDRAND) && Features.has(Xbyak::util::Cpu::tRDSEED);
SupportsRCPC = true;
SupportsTSOImm9 = true;
// xbyak doesn't know how to check for CLZero
uint32_t eax, ebx, ecx, edx;
// First ensure we support a new enough extended CPUID function range
__cpuid(0x8000'0000, eax, ebx, ecx, edx);
if (eax >= 0x8000'0008U) {
// CLZero defined in 8000_00008_EBX[bit 0]
__cpuid(0x8000'0008, eax, ebx, ecx, edx);
SupportsCLZERO = ebx & 1;
}
SupportsFlushInputsToZero = true;
SupportsFloatExceptions = true;
#else
+3
View File
@@ -15,9 +15,12 @@ class HostFeatures final {
uint32_t DCacheLineSize{};
uint32_t ICacheLineSize{};
bool SupportsAES{};
bool SupportsCRC{};
bool SupportsCLZERO{};
bool SupportsAtomics{};
bool SupportsRCPC{};
bool SupportsTSOImm9{};
bool SupportsRAND{};
// Float exception behaviour
bool SupportsFlushInputsToZero{};
+14 -11
View File
@@ -18,7 +18,7 @@ namespace FEXCore::CPU {
DEF_OP(TruncElementPair) {
auto Op = IROp->C<IR::IROp_TruncElementPair>();
switch (Op->Size) {
switch (IROp->Size) {
case 4: {
uint64_t *Src = GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[0]);
uint64_t Result{};
@@ -27,7 +27,7 @@ DEF_OP(TruncElementPair) {
GD = Result;
break;
}
default: LOGMAN_MSG_A_FMT("Unhandled Truncation size: {}", Op->Size); break;
default: LOGMAN_MSG_A_FMT("Unhandled Truncation size: {}", IROp->Size); break;
}
}
@@ -38,7 +38,13 @@ DEF_OP(Constant) {
DEF_OP(EntrypointOffset) {
auto Op = IROp->C<IR::IROp_EntrypointOffset>();
GD = Data->CurrentEntry + Op->Offset;
uint64_t Mask = ~0ULL;
uint8_t OpSize = IROp->Size;
if (OpSize == 4) {
Mask = 0xFFFF'FFFFULL;
}
GD = (Data->CurrentEntry + Op->Offset) & Mask;
}
DEF_OP(InlineConstant) {
@@ -835,9 +841,8 @@ DEF_OP(Bfi) {
DEF_OP(Bfe) {
auto Op = IROp->C<IR::IROp_Bfe>();
uint8_t OpSize = IROp->Size;
LOGMAN_THROW_A_FMT(OpSize <= 8, "OpSize is too large for BFE: {}", OpSize);
LOGMAN_THROW_A_FMT(IROp->Size <= 8, "OpSize is too large for BFE: {}", IROp->Size);
uint64_t SourceMask = (1ULL << Op->Width) - 1;
if (Op->Width == 64)
SourceMask = ~0ULL;
@@ -848,9 +853,8 @@ DEF_OP(Bfe) {
DEF_OP(Sbfe) {
auto Op = IROp->C<IR::IROp_Sbfe>();
uint8_t OpSize = IROp->Size;
LOGMAN_THROW_A_FMT(OpSize <= 8, "OpSize is too large for SBFE: {}", OpSize);
LOGMAN_THROW_A_FMT(IROp->Size <= 8, "OpSize is too large for SBFE: {}", IROp->Size);
int64_t Src = *GetSrc<int64_t*>(Data->SSAData, Op->Header.Args[0]);
uint64_t ShiftLeftAmount = (64 - (Op->Width + Op->lsb));
uint64_t ShiftRightAmount = ShiftLeftAmount + Op->lsb;
@@ -889,15 +893,14 @@ DEF_OP(Select) {
DEF_OP(VExtractToGPR) {
auto Op = IROp->C<IR::IROp_VExtractToGPR>();
uint8_t OpSize = IROp->Size;
uint32_t SourceSize = GetOpSize(Data->CurrentIR, Op->Header.Args[0]);
LOGMAN_THROW_A_FMT(OpSize <= 16, "OpSize is too large for VExtractToGPR: {}", OpSize);
LOGMAN_THROW_A_FMT(IROp->Size <= 16, "OpSize is too large for VExtractToGPR: {}", IROp->Size);
if (SourceSize == 16) {
__uint128_t SourceMask = (1ULL << (Op->Header.ElementSize * 8)) - 1;
uint64_t Shift = Op->Header.ElementSize * Op->Idx * 8;
uint64_t Shift = Op->Header.ElementSize * Op->Index * 8;
if (Op->Header.ElementSize == 8)
SourceMask = ~0ULL;
@@ -908,7 +911,7 @@ DEF_OP(VExtractToGPR) {
}
else {
uint64_t SourceMask = (1ULL << (Op->Header.ElementSize * 8)) - 1;
uint64_t Shift = Op->Header.ElementSize * Op->Idx * 8;
uint64_t Shift = Op->Header.ElementSize * Op->Index * 8;
if (Op->Header.ElementSize == 8)
SourceMask = ~0ULL;
+112 -113
View File
@@ -313,30 +313,29 @@ uint64_t AtomicCompareAndSwap(uint64_t expected, uint64_t desired, uint64_t *add
#define DEF_OP(x) void InterpreterOps::Op_##x(IR::IROp_Header *IROp, IROpData *Data, IR::NodeID Node)
DEF_OP(CASPair) {
auto Op = IROp->C<IR::IROp_CASPair>();
uint8_t OpSize = IROp->Size;
// Size is the size of each pair element
switch (OpSize) {
switch (IROp->ElementSize) {
case 4: {
GD = AtomicCompareAndSwap(
*GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[0]),
*GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]),
*GetSrc<uint64_t**>(Data->SSAData, Op->Header.Args[2])
*GetSrc<uint64_t*>(Data->SSAData, Op->Expected),
*GetSrc<uint64_t*>(Data->SSAData, Op->Desired),
*GetSrc<uint64_t**>(Data->SSAData, Op->Addr)
);
break;
}
case 8: {
std::atomic<__uint128_t> *MemData = *GetSrc<std::atomic<__uint128_t> **>(Data->SSAData, Op->Header.Args[2]);
std::atomic<__uint128_t> *MemData = *GetSrc<std::atomic<__uint128_t> **>(Data->SSAData, Op->Addr);
__uint128_t Src1 = *GetSrc<__uint128_t*>(Data->SSAData, Op->Header.Args[0]);
__uint128_t Src2 = *GetSrc<__uint128_t*>(Data->SSAData, Op->Header.Args[1]);
__uint128_t Src1 = *GetSrc<__uint128_t*>(Data->SSAData, Op->Expected);
__uint128_t Src2 = *GetSrc<__uint128_t*>(Data->SSAData, Op->Desired);
__uint128_t Expected = Src1;
bool Result = MemData->compare_exchange_strong(Expected, Src2);
memcpy(GDP, Result ? &Src1 : &Expected, 16);
break;
}
default: LOGMAN_MSG_A_FMT("Unknown CAS size: {}", OpSize); break;
default: LOGMAN_MSG_A_FMT("Unknown CAS size: {}", IROp->ElementSize); break;
}
}
@@ -347,33 +346,33 @@ DEF_OP(CAS) {
switch (OpSize) {
case 1: {
GD = AtomicCompareAndSwap(
*GetSrc<uint8_t*>(Data->SSAData, Op->Header.Args[0]),
*GetSrc<uint8_t*>(Data->SSAData, Op->Header.Args[1]),
*GetSrc<uint8_t**>(Data->SSAData, Op->Header.Args[2])
*GetSrc<uint8_t*>(Data->SSAData, Op->Expected),
*GetSrc<uint8_t*>(Data->SSAData, Op->Desired),
*GetSrc<uint8_t**>(Data->SSAData, Op->Addr)
);
break;
}
case 2: {
GD = AtomicCompareAndSwap(
*GetSrc<uint16_t*>(Data->SSAData, Op->Header.Args[0]),
*GetSrc<uint16_t*>(Data->SSAData, Op->Header.Args[1]),
*GetSrc<uint16_t**>(Data->SSAData, Op->Header.Args[2])
*GetSrc<uint16_t*>(Data->SSAData, Op->Expected),
*GetSrc<uint16_t*>(Data->SSAData, Op->Desired),
*GetSrc<uint16_t**>(Data->SSAData, Op->Addr)
);
break;
}
case 4: {
GD = AtomicCompareAndSwap(
*GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[0]),
*GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[1]),
*GetSrc<uint32_t**>(Data->SSAData, Op->Header.Args[2])
*GetSrc<uint32_t*>(Data->SSAData, Op->Expected),
*GetSrc<uint32_t*>(Data->SSAData, Op->Desired),
*GetSrc<uint32_t**>(Data->SSAData, Op->Addr)
);
break;
}
case 8: {
GD = AtomicCompareAndSwap(
*GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[0]),
*GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]),
*GetSrc<uint64_t**>(Data->SSAData, Op->Header.Args[2])
*GetSrc<uint64_t*>(Data->SSAData, Op->Expected),
*GetSrc<uint64_t*>(Data->SSAData, Op->Desired),
*GetSrc<uint64_t**>(Data->SSAData, Op->Addr)
);
break;
}
@@ -385,26 +384,26 @@ DEF_OP(AtomicAdd) {
auto Op = IROp->C<IR::IROp_AtomicAdd>();
switch (IROp->Size) {
case 1: {
std::atomic<uint8_t> *MemData = *GetSrc<std::atomic<uint8_t> **>(Data->SSAData, Op->Header.Args[0]);
uint8_t Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint8_t> *MemData = *GetSrc<std::atomic<uint8_t> **>(Data->SSAData, Op->Addr);
uint8_t Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Value);
*MemData += Src;
break;
}
case 2: {
std::atomic<uint16_t> *MemData = *GetSrc<std::atomic<uint16_t> **>(Data->SSAData, Op->Header.Args[0]);
uint16_t Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint16_t> *MemData = *GetSrc<std::atomic<uint16_t> **>(Data->SSAData, Op->Addr);
uint16_t Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Value);
*MemData += Src;
break;
}
case 4: {
std::atomic<uint32_t> *MemData = *GetSrc<std::atomic<uint32_t> **>(Data->SSAData, Op->Header.Args[0]);
uint32_t Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint32_t> *MemData = *GetSrc<std::atomic<uint32_t> **>(Data->SSAData, Op->Addr);
uint32_t Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Value);
*MemData += Src;
break;
}
case 8: {
std::atomic<uint64_t> *MemData = *GetSrc<std::atomic<uint64_t> **>(Data->SSAData, Op->Header.Args[0]);
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint64_t> *MemData = *GetSrc<std::atomic<uint64_t> **>(Data->SSAData, Op->Addr);
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Value);
*MemData += Src;
break;
}
@@ -416,26 +415,26 @@ DEF_OP(AtomicSub) {
auto Op = IROp->C<IR::IROp_AtomicSub>();
switch (IROp->Size) {
case 1: {
std::atomic<uint8_t> *MemData = *GetSrc<std::atomic<uint8_t> **>(Data->SSAData, Op->Header.Args[0]);
uint8_t Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint8_t> *MemData = *GetSrc<std::atomic<uint8_t> **>(Data->SSAData, Op->Addr);
uint8_t Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Value);
*MemData -= Src;
break;
}
case 2: {
std::atomic<uint16_t> *MemData = *GetSrc<std::atomic<uint16_t> **>(Data->SSAData, Op->Header.Args[0]);
uint16_t Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint16_t> *MemData = *GetSrc<std::atomic<uint16_t> **>(Data->SSAData, Op->Addr);
uint16_t Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Value);
*MemData -= Src;
break;
}
case 4: {
std::atomic<uint32_t> *MemData = *GetSrc<std::atomic<uint32_t> **>(Data->SSAData, Op->Header.Args[0]);
uint32_t Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint32_t> *MemData = *GetSrc<std::atomic<uint32_t> **>(Data->SSAData, Op->Addr);
uint32_t Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Value);
*MemData -= Src;
break;
}
case 8: {
std::atomic<uint64_t> *MemData = *GetSrc<std::atomic<uint64_t> **>(Data->SSAData, Op->Header.Args[0]);
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint64_t> *MemData = *GetSrc<std::atomic<uint64_t> **>(Data->SSAData, Op->Addr);
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Value);
*MemData -= Src;
break;
}
@@ -447,26 +446,26 @@ DEF_OP(AtomicAnd) {
auto Op = IROp->C<IR::IROp_AtomicAnd>();
switch (IROp->Size) {
case 1: {
std::atomic<uint8_t> *MemData = *GetSrc<std::atomic<uint8_t> **>(Data->SSAData, Op->Header.Args[0]);
uint8_t Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint8_t> *MemData = *GetSrc<std::atomic<uint8_t> **>(Data->SSAData, Op->Addr);
uint8_t Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Value);
*MemData &= Src;
break;
}
case 2: {
std::atomic<uint16_t> *MemData = *GetSrc<std::atomic<uint16_t> **>(Data->SSAData, Op->Header.Args[0]);
uint16_t Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint16_t> *MemData = *GetSrc<std::atomic<uint16_t> **>(Data->SSAData, Op->Addr);
uint16_t Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Value);
*MemData &= Src;
break;
}
case 4: {
std::atomic<uint32_t> *MemData = *GetSrc<std::atomic<uint32_t> **>(Data->SSAData, Op->Header.Args[0]);
uint32_t Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint32_t> *MemData = *GetSrc<std::atomic<uint32_t> **>(Data->SSAData, Op->Addr);
uint32_t Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Value);
*MemData &= Src;
break;
}
case 8: {
std::atomic<uint64_t> *MemData = *GetSrc<std::atomic<uint64_t> **>(Data->SSAData, Op->Header.Args[0]);
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint64_t> *MemData = *GetSrc<std::atomic<uint64_t> **>(Data->SSAData, Op->Addr);
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Value);
*MemData &= Src;
break;
}
@@ -478,26 +477,26 @@ DEF_OP(AtomicOr) {
auto Op = IROp->C<IR::IROp_AtomicOr>();
switch (IROp->Size) {
case 1: {
std::atomic<uint8_t> *MemData = *GetSrc<std::atomic<uint8_t> **>(Data->SSAData, Op->Header.Args[0]);
uint8_t Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint8_t> *MemData = *GetSrc<std::atomic<uint8_t> **>(Data->SSAData, Op->Addr);
uint8_t Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Value);
*MemData |= Src;
break;
}
case 2: {
std::atomic<uint16_t> *MemData = *GetSrc<std::atomic<uint16_t> **>(Data->SSAData, Op->Header.Args[0]);
uint16_t Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint16_t> *MemData = *GetSrc<std::atomic<uint16_t> **>(Data->SSAData, Op->Addr);
uint16_t Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Value);
*MemData |= Src;
break;
}
case 4: {
std::atomic<uint32_t> *MemData = *GetSrc<std::atomic<uint32_t> **>(Data->SSAData, Op->Header.Args[0]);
uint32_t Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint32_t> *MemData = *GetSrc<std::atomic<uint32_t> **>(Data->SSAData, Op->Addr);
uint32_t Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Value);
*MemData |= Src;
break;
}
case 8: {
std::atomic<uint64_t> *MemData = *GetSrc<std::atomic<uint64_t> **>(Data->SSAData, Op->Header.Args[0]);
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint64_t> *MemData = *GetSrc<std::atomic<uint64_t> **>(Data->SSAData, Op->Addr);
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Value);
*MemData |= Src;
break;
}
@@ -509,26 +508,26 @@ DEF_OP(AtomicXor) {
auto Op = IROp->C<IR::IROp_AtomicXor>();
switch (IROp->Size) {
case 1: {
std::atomic<uint8_t> *MemData = *GetSrc<std::atomic<uint8_t> **>(Data->SSAData, Op->Header.Args[0]);
uint8_t Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint8_t> *MemData = *GetSrc<std::atomic<uint8_t> **>(Data->SSAData, Op->Addr);
uint8_t Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Value);
*MemData ^= Src;
break;
}
case 2: {
std::atomic<uint16_t> *MemData = *GetSrc<std::atomic<uint16_t> **>(Data->SSAData, Op->Header.Args[0]);
uint16_t Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint16_t> *MemData = *GetSrc<std::atomic<uint16_t> **>(Data->SSAData, Op->Addr);
uint16_t Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Value);
*MemData ^= Src;
break;
}
case 4: {
std::atomic<uint32_t> *MemData = *GetSrc<std::atomic<uint32_t> **>(Data->SSAData, Op->Header.Args[0]);
uint32_t Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint32_t> *MemData = *GetSrc<std::atomic<uint32_t> **>(Data->SSAData, Op->Addr);
uint32_t Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Value);
*MemData ^= Src;
break;
}
case 8: {
std::atomic<uint64_t> *MemData = *GetSrc<std::atomic<uint64_t> **>(Data->SSAData, Op->Header.Args[0]);
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint64_t> *MemData = *GetSrc<std::atomic<uint64_t> **>(Data->SSAData, Op->Addr);
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Value);
*MemData ^= Src;
break;
}
@@ -540,29 +539,29 @@ DEF_OP(AtomicSwap) {
auto Op = IROp->C<IR::IROp_AtomicSwap>();
switch (IROp->Size) {
case 1: {
std::atomic<uint8_t> *MemData = *GetSrc<std::atomic<uint8_t> **>(Data->SSAData, Op->Header.Args[0]);
uint8_t Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint8_t> *MemData = *GetSrc<std::atomic<uint8_t> **>(Data->SSAData, Op->Addr);
uint8_t Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Value);
uint8_t Previous = MemData->exchange(Src);
GD = Previous;
break;
}
case 2: {
std::atomic<uint16_t> *MemData = *GetSrc<std::atomic<uint16_t> **>(Data->SSAData, Op->Header.Args[0]);
uint16_t Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint16_t> *MemData = *GetSrc<std::atomic<uint16_t> **>(Data->SSAData, Op->Addr);
uint16_t Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Value);
uint16_t Previous = MemData->exchange(Src);
GD = Previous;
break;
}
case 4: {
std::atomic<uint32_t> *MemData = *GetSrc<std::atomic<uint32_t> **>(Data->SSAData, Op->Header.Args[0]);
uint32_t Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint32_t> *MemData = *GetSrc<std::atomic<uint32_t> **>(Data->SSAData, Op->Addr);
uint32_t Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Value);
uint32_t Previous = MemData->exchange(Src);
GD = Previous;
break;
}
case 8: {
std::atomic<uint64_t> *MemData = *GetSrc<std::atomic<uint64_t> **>(Data->SSAData, Op->Header.Args[0]);
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint64_t> *MemData = *GetSrc<std::atomic<uint64_t> **>(Data->SSAData, Op->Addr);
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Value);
uint64_t Previous = MemData->exchange(Src);
GD = Previous;
break;
@@ -575,29 +574,29 @@ DEF_OP(AtomicFetchAdd) {
auto Op = IROp->C<IR::IROp_AtomicFetchAdd>();
switch (IROp->Size) {
case 1: {
std::atomic<uint8_t> *MemData = *GetSrc<std::atomic<uint8_t> **>(Data->SSAData, Op->Header.Args[0]);
uint8_t Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint8_t> *MemData = *GetSrc<std::atomic<uint8_t> **>(Data->SSAData, Op->Addr);
uint8_t Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Value);
uint8_t Previous = MemData->fetch_add(Src);
GD = Previous;
break;
}
case 2: {
std::atomic<uint16_t> *MemData = *GetSrc<std::atomic<uint16_t> **>(Data->SSAData, Op->Header.Args[0]);
uint16_t Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint16_t> *MemData = *GetSrc<std::atomic<uint16_t> **>(Data->SSAData, Op->Addr);
uint16_t Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Value);
uint16_t Previous = MemData->fetch_add(Src);
GD = Previous;
break;
}
case 4: {
std::atomic<uint32_t> *MemData = *GetSrc<std::atomic<uint32_t> **>(Data->SSAData, Op->Header.Args[0]);
uint32_t Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint32_t> *MemData = *GetSrc<std::atomic<uint32_t> **>(Data->SSAData, Op->Addr);
uint32_t Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Value);
uint32_t Previous = MemData->fetch_add(Src);
GD = Previous;
break;
}
case 8: {
std::atomic<uint64_t> *MemData = *GetSrc<std::atomic<uint64_t> **>(Data->SSAData, Op->Header.Args[0]);
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint64_t> *MemData = *GetSrc<std::atomic<uint64_t> **>(Data->SSAData, Op->Addr);
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Value);
uint64_t Previous = MemData->fetch_add(Src);
GD = Previous;
break;
@@ -610,29 +609,29 @@ DEF_OP(AtomicFetchSub) {
auto Op = IROp->C<IR::IROp_AtomicFetchSub>();
switch (IROp->Size) {
case 1: {
std::atomic<uint8_t> *MemData = *GetSrc<std::atomic<uint8_t> **>(Data->SSAData, Op->Header.Args[0]);
uint8_t Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint8_t> *MemData = *GetSrc<std::atomic<uint8_t> **>(Data->SSAData, Op->Addr);
uint8_t Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Value);
uint8_t Previous = MemData->fetch_sub(Src);
GD = Previous;
break;
}
case 2: {
std::atomic<uint16_t> *MemData = *GetSrc<std::atomic<uint16_t> **>(Data->SSAData, Op->Header.Args[0]);
uint16_t Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint16_t> *MemData = *GetSrc<std::atomic<uint16_t> **>(Data->SSAData, Op->Addr);
uint16_t Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Value);
uint16_t Previous = MemData->fetch_sub(Src);
GD = Previous;
break;
}
case 4: {
std::atomic<uint32_t> *MemData = *GetSrc<std::atomic<uint32_t> **>(Data->SSAData, Op->Header.Args[0]);
uint32_t Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint32_t> *MemData = *GetSrc<std::atomic<uint32_t> **>(Data->SSAData, Op->Addr);
uint32_t Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Value);
uint32_t Previous = MemData->fetch_sub(Src);
GD = Previous;
break;
}
case 8: {
std::atomic<uint64_t> *MemData = *GetSrc<std::atomic<uint64_t> **>(Data->SSAData, Op->Header.Args[0]);
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint64_t> *MemData = *GetSrc<std::atomic<uint64_t> **>(Data->SSAData, Op->Addr);
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Value);
uint64_t Previous = MemData->fetch_sub(Src);
GD = Previous;
break;
@@ -645,29 +644,29 @@ DEF_OP(AtomicFetchAnd) {
auto Op = IROp->C<IR::IROp_AtomicFetchAnd>();
switch (IROp->Size) {
case 1: {
std::atomic<uint8_t> *MemData = *GetSrc<std::atomic<uint8_t> **>(Data->SSAData, Op->Header.Args[0]);
uint8_t Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint8_t> *MemData = *GetSrc<std::atomic<uint8_t> **>(Data->SSAData, Op->Addr);
uint8_t Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Value);
uint8_t Previous = MemData->fetch_and(Src);
GD = Previous;
break;
}
case 2: {
std::atomic<uint16_t> *MemData = *GetSrc<std::atomic<uint16_t> **>(Data->SSAData, Op->Header.Args[0]);
uint16_t Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint16_t> *MemData = *GetSrc<std::atomic<uint16_t> **>(Data->SSAData, Op->Addr);
uint16_t Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Value);
uint16_t Previous = MemData->fetch_and(Src);
GD = Previous;
break;
}
case 4: {
std::atomic<uint32_t> *MemData = *GetSrc<std::atomic<uint32_t> **>(Data->SSAData, Op->Header.Args[0]);
uint32_t Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint32_t> *MemData = *GetSrc<std::atomic<uint32_t> **>(Data->SSAData, Op->Addr);
uint32_t Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Value);
uint32_t Previous = MemData->fetch_and(Src);
GD = Previous;
break;
}
case 8: {
std::atomic<uint64_t> *MemData = *GetSrc<std::atomic<uint64_t> **>(Data->SSAData, Op->Header.Args[0]);
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint64_t> *MemData = *GetSrc<std::atomic<uint64_t> **>(Data->SSAData, Op->Addr);
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Value);
uint64_t Previous = MemData->fetch_and(Src);
GD = Previous;
break;
@@ -680,29 +679,29 @@ DEF_OP(AtomicFetchOr) {
auto Op = IROp->C<IR::IROp_AtomicFetchOr>();
switch (IROp->Size) {
case 1: {
std::atomic<uint8_t> *MemData = *GetSrc<std::atomic<uint8_t> **>(Data->SSAData, Op->Header.Args[0]);
uint8_t Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint8_t> *MemData = *GetSrc<std::atomic<uint8_t> **>(Data->SSAData, Op->Addr);
uint8_t Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Value);
uint8_t Previous = MemData->fetch_or(Src);
GD = Previous;
break;
}
case 2: {
std::atomic<uint16_t> *MemData = *GetSrc<std::atomic<uint16_t> **>(Data->SSAData, Op->Header.Args[0]);
uint16_t Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint16_t> *MemData = *GetSrc<std::atomic<uint16_t> **>(Data->SSAData, Op->Addr);
uint16_t Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Value);
uint16_t Previous = MemData->fetch_or(Src);
GD = Previous;
break;
}
case 4: {
std::atomic<uint32_t> *MemData = *GetSrc<std::atomic<uint32_t> **>(Data->SSAData, Op->Header.Args[0]);
uint32_t Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint32_t> *MemData = *GetSrc<std::atomic<uint32_t> **>(Data->SSAData, Op->Addr);
uint32_t Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Value);
uint32_t Previous = MemData->fetch_or(Src);
GD = Previous;
break;
}
case 8: {
std::atomic<uint64_t> *MemData = *GetSrc<std::atomic<uint64_t> **>(Data->SSAData, Op->Header.Args[0]);
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint64_t> *MemData = *GetSrc<std::atomic<uint64_t> **>(Data->SSAData, Op->Addr);
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Value);
uint64_t Previous = MemData->fetch_or(Src);
GD = Previous;
break;
@@ -715,29 +714,29 @@ DEF_OP(AtomicFetchXor) {
auto Op = IROp->C<IR::IROp_AtomicFetchXor>();
switch (IROp->Size) {
case 1: {
std::atomic<uint8_t> *MemData = *GetSrc<std::atomic<uint8_t> **>(Data->SSAData, Op->Header.Args[0]);
uint8_t Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint8_t> *MemData = *GetSrc<std::atomic<uint8_t> **>(Data->SSAData, Op->Addr);
uint8_t Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Value);
uint8_t Previous = MemData->fetch_xor(Src);
GD = Previous;
break;
}
case 2: {
std::atomic<uint16_t> *MemData = *GetSrc<std::atomic<uint16_t> **>(Data->SSAData, Op->Header.Args[0]);
uint16_t Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint16_t> *MemData = *GetSrc<std::atomic<uint16_t> **>(Data->SSAData, Op->Addr);
uint16_t Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Value);
uint16_t Previous = MemData->fetch_xor(Src);
GD = Previous;
break;
}
case 4: {
std::atomic<uint32_t> *MemData = *GetSrc<std::atomic<uint32_t> **>(Data->SSAData, Op->Header.Args[0]);
uint32_t Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint32_t> *MemData = *GetSrc<std::atomic<uint32_t> **>(Data->SSAData, Op->Addr);
uint32_t Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Value);
uint32_t Previous = MemData->fetch_xor(Src);
GD = Previous;
break;
}
case 8: {
std::atomic<uint64_t> *MemData = *GetSrc<std::atomic<uint64_t> **>(Data->SSAData, Op->Header.Args[0]);
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint64_t> *MemData = *GetSrc<std::atomic<uint64_t> **>(Data->SSAData, Op->Addr);
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Value);
uint64_t Previous = MemData->fetch_xor(Src);
GD = Previous;
break;
@@ -751,22 +750,22 @@ DEF_OP(AtomicFetchNeg) {
switch (IROp->Size) {
case 1: {
using Type = uint8_t;
GD = AtomicFetchNeg(*GetSrc<Type**>(Data->SSAData, Op->Header.Args[0]));
GD = AtomicFetchNeg(*GetSrc<Type**>(Data->SSAData, Op->Addr));
break;
}
case 2: {
using Type = uint16_t;
GD = AtomicFetchNeg(*GetSrc<Type**>(Data->SSAData, Op->Header.Args[0]));
GD = AtomicFetchNeg(*GetSrc<Type**>(Data->SSAData, Op->Addr));
break;
}
case 4: {
using Type = uint32_t;
GD = AtomicFetchNeg(*GetSrc<Type**>(Data->SSAData, Op->Header.Args[0]));
GD = AtomicFetchNeg(*GetSrc<Type**>(Data->SSAData, Op->Addr));
break;
}
case 8: {
using Type = uint64_t;
GD = AtomicFetchNeg(*GetSrc<Type**>(Data->SSAData, Op->Header.Args[0]));
GD = AtomicFetchNeg(*GetSrc<Type**>(Data->SSAData, Op->Addr));
break;
}
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", IROp->Size);
@@ -33,10 +33,6 @@ DEF_OP(GuestCallIndirect) {
LogMan::Msg::DFmt("Unimplemented");
}
DEF_OP(GuestReturn) {
LogMan::Msg::DFmt("Unimplemented");
}
DEF_OP(SignalReturn) {
SignalReturn(Data->State);
}
@@ -151,8 +147,8 @@ DEF_OP(ValidateCode) {
}
}
DEF_OP(RemoveCodeEntry) {
Data->State->CTX->RemoveCodeEntry(Data->State, Data->CurrentEntry);
DEF_OP(RemoveThreadCodeEntry) {
Data->State->CTX->RemoveThreadCodeEntry(Data->State, Data->CurrentEntry);
}
DEF_OP(CPUID) {
@@ -19,7 +19,7 @@ DEF_OP(VInsGPR) {
__uint128_t Src1 = *GetSrc<__uint128_t*>(Data->SSAData, Op->Header.Args[0]);
__uint128_t Src2 = *GetSrc<__uint128_t*>(Data->SSAData, Op->Header.Args[1]);
uint64_t Offset = Op->Index * Op->Header.ElementSize * 8;
uint64_t Offset = Op->DestIdx * Op->Header.ElementSize * 8;
__uint128_t Mask = (1ULL << (Op->Header.ElementSize * 8)) - 1;
if (Op->Header.ElementSize == 8) {
Mask = ~0ULL;
@@ -298,6 +298,63 @@ namespace AES {
}
}
namespace CRC32 {
// CRC32 per byte lookup table.
constexpr std::array<uint32_t, 256> CRC32CTable = []() consteval {
std::array<uint32_t, 256> Table{};
// Clang 11.x doesn't support bitreverse as a consteval
// constexpr uint32_t Polynomial = 0x1EDC6F41;
constexpr uint32_t PolynomialRev = 0x82F63B78; //__builtin_bitreverse32(Polynomial);
for (size_t Char = 0; Char < std::size(Table); ++Char) {
uint32_t CurrentChar = Char;
for (size_t i = 0; i < 8; ++i) {
if (CurrentChar & 1) {
CurrentChar = (CurrentChar >> 1) ^ PolynomialRev;
}
else {
CurrentChar >>= 1;
}
}
Table[Char] = CurrentChar;
}
return Table;
}();
uint32_t crc32cb(uint32_t Accumulator, uint8_t data) {
Accumulator = CRC32CTable[(uint8_t)Accumulator ^ data] ^ Accumulator >> 8;
return Accumulator;
}
uint32_t crc32ch(uint32_t Accumulator, uint16_t data) {
Accumulator = CRC32CTable[(uint8_t)Accumulator ^ ((data >> 0) & 0xFF)] ^ Accumulator >> 8;
Accumulator = CRC32CTable[(uint8_t)Accumulator ^ ((data >> 8) & 0xFF)] ^ Accumulator >> 8;
return Accumulator;
}
uint32_t crc32cw(uint32_t Accumulator, uint32_t data) {
Accumulator = CRC32CTable[(uint8_t)Accumulator ^ ((data >> 0) & 0xFF)] ^ Accumulator >> 8;
Accumulator = CRC32CTable[(uint8_t)Accumulator ^ ((data >> 8) & 0xFF)] ^ Accumulator >> 8;
Accumulator = CRC32CTable[(uint8_t)Accumulator ^ ((data >> 16) & 0xFF)] ^ Accumulator >> 8;
Accumulator = CRC32CTable[(uint8_t)Accumulator ^ ((data >> 24) & 0xFF)] ^ Accumulator >> 8;
return Accumulator;
}
uint32_t crc32cx(uint32_t Accumulator, uint64_t data) {
Accumulator = CRC32CTable[(uint8_t)Accumulator ^ ((data >> 0) & 0xFF)] ^ Accumulator >> 8;
Accumulator = CRC32CTable[(uint8_t)Accumulator ^ ((data >> 8) & 0xFF)] ^ Accumulator >> 8;
Accumulator = CRC32CTable[(uint8_t)Accumulator ^ ((data >> 16) & 0xFF)] ^ Accumulator >> 8;
Accumulator = CRC32CTable[(uint8_t)Accumulator ^ ((data >> 24) & 0xFF)] ^ Accumulator >> 8;
Accumulator = CRC32CTable[(uint8_t)Accumulator ^ ((data >> 32) & 0xFF)] ^ Accumulator >> 8;
Accumulator = CRC32CTable[(uint8_t)Accumulator ^ ((data >> 40) & 0xFF)] ^ Accumulator >> 8;
Accumulator = CRC32CTable[(uint8_t)Accumulator ^ ((data >> 48) & 0xFF)] ^ Accumulator >> 8;
Accumulator = CRC32CTable[(uint8_t)Accumulator ^ ((data >> 56) & 0xFF)] ^ Accumulator >> 8;
return Accumulator;
}
}
namespace FEXCore::CPU {
#define DEF_OP(x) void InterpreterOps::Op_##x(IR::IROp_Header *IROp, IROpData *Data, IR::NodeID Node)
@@ -429,6 +486,33 @@ DEF_OP(AESKeyGenAssist) {
memcpy(GDP, &Tmp, sizeof(Tmp));
}
DEF_OP(CRC32) {
auto Op = IROp->C<IR::IROp_CRC32>();
uint32_t Src1 = *GetSrc<uint32_t*>(Data->SSAData, Op->Src1);
uint8_t *Src2 = GetSrc<uint8_t*>(Data->SSAData, Op->Src2);
uint32_t Tmp{};
switch (Op->SrcSize) {
case 1:
Tmp = CRC32::crc32cb(Src1, *(uint8_t*)Src2);
break;
case 2:
Tmp = CRC32::crc32ch(Src1, *(uint16_t*)Src2);
break;
case 4:
Tmp = CRC32::crc32cw(Src1, *(uint32_t*)Src2);
break;
case 8:
Tmp = CRC32::crc32cx(Src1, *(uint64_t*)Src2);
break;
default:
LOGMAN_MSG_A_FMT("Unknown CRC32C size: {}", Op->SrcSize);
break;
}
memcpy(GDP, &Tmp, sizeof(Tmp));
}
#undef DEF_OP
} // namespace FEXCore::CPU
@@ -158,7 +158,7 @@ DEF_OP(F80CVTINT) {
DEF_OP(F80CVTTO) {
auto Op = IROp->C<IR::IROp_F80CVTTo>();
switch (Op->Size) {
switch (Op->SrcSize) {
case 4: {
float Src = *GetSrc<float *>(Data->SSAData, Op->Header.Args[0]);
X80SoftFloat Tmp = Src;
@@ -171,14 +171,14 @@ DEF_OP(F80CVTTO) {
memcpy(GDP, &Tmp, sizeof(X80SoftFloat));
break;
}
default: LogMan::Msg::DFmt("Unhandled size: {}", Op->Size);
default: LogMan::Msg::DFmt("Unhandled size: {}", Op->SrcSize);
}
}
DEF_OP(F80CVTTOINT) {
auto Op = IROp->C<IR::IROp_F80CVTToInt>();
switch (Op->Size) {
switch (Op->SrcSize) {
case 2: {
int16_t Src = *GetSrc<int16_t*>(Data->SSAData, Op->Header.Args[0]);
X80SoftFloat Tmp = Src;
@@ -191,7 +191,7 @@ DEF_OP(F80CVTTOINT) {
memcpy(GDP, &Tmp, sizeof(X80SoftFloat));
break;
}
default: LogMan::Msg::DFmt("Unhandled size: {}", Op->Size);
default: LogMan::Msg::DFmt("Unhandled size: {}", Op->SrcSize);
}
}
@@ -323,7 +323,7 @@ DEF_OP(F80BCDLOAD) {
DEF_OP(F80BCDSTORE) {
auto Op = IROp->C<IR::IROp_F80BCDStore>();
X80SoftFloat Src1 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[0]);
X80SoftFloat Src1 = X80SoftFloat::FRNDINT(*GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[0]));
bool Negative = Src1.Sign;
// Clear the Sign bit
@@ -356,6 +356,81 @@ DEF_OP(F80BCDSTORE) {
memcpy(GDP, BCD, 10);
}
DEF_OP(F64SIN) {
auto Op = IROp->C<IR::IROp_F64SIN>();
double Src = *GetSrc<double*>(Data->SSAData, Op->Header.Args[0]);
double Tmp = sin(Src);
memcpy(GDP, &Tmp, sizeof(double));
}
DEF_OP(F64COS) {
auto Op = IROp->C<IR::IROp_F64COS>();
double Src = *GetSrc<double*>(Data->SSAData, Op->Header.Args[0]);
double Tmp = cos(Src);
memcpy(GDP, &Tmp, sizeof(double));
}
DEF_OP(F64TAN) {
auto Op = IROp->C<IR::IROp_F64TAN>();
double Src = *GetSrc<double*>(Data->SSAData, Op->Header.Args[0]);
double Tmp = tan(Src);
memcpy(GDP, &Tmp, sizeof(double));
}
DEF_OP(F64F2XM1) {
auto Op = IROp->C<IR::IROp_F64F2XM1>();
double Src = *GetSrc<double*>(Data->SSAData, Op->Header.Args[0]);
double Tmp = exp2(Src) - 1.0;
memcpy(GDP, &Tmp, sizeof(double));
}
DEF_OP(F64ATAN) {
auto Op = IROp->C<IR::IROp_F64ATAN>();
double Src1 = *GetSrc<double*>(Data->SSAData, Op->Header.Args[0]);
double Src2 = *GetSrc<double*>(Data->SSAData, Op->Header.Args[1]);
double Tmp = atan2(Src1, Src2);
memcpy(GDP, &Tmp, sizeof(double));
}
DEF_OP(F64FPREM) {
auto Op = IROp->C<IR::IROp_F64FPREM>();
double Src1 = *GetSrc<double*>(Data->SSAData, Op->Header.Args[0]);
double Src2 = *GetSrc<double*>(Data->SSAData, Op->Header.Args[1]);
double Tmp = fmod(Src1, Src2);
memcpy(GDP, &Tmp, sizeof(double));
}
DEF_OP(F64FPREM1) {
auto Op = IROp->C<IR::IROp_F64FPREM1>();
double Src1 = *GetSrc<double*>(Data->SSAData, Op->Header.Args[0]);
double Src2 = *GetSrc<double*>(Data->SSAData, Op->Header.Args[1]);
double Tmp = remainder(Src1, Src2);
memcpy(GDP, &Tmp, sizeof(double));
}
DEF_OP(F64FYL2X) {
auto Op = IROp->C<IR::IROp_F64FYL2X>();
double Src1 = *GetSrc<double*>(Data->SSAData, Op->Header.Args[0]);
double Src2 = *GetSrc<double*>(Data->SSAData, Op->Header.Args[1]);
double Tmp = Src2 * log2(Src1);
memcpy(GDP, &Tmp, sizeof(double));
}
DEF_OP(F64SCALE) {
auto Op = IROp->C<IR::IROp_F64SCALE>();
double Src1 = *GetSrc<double*>(Data->SSAData, Op->Header.Args[0]);
double Src2 = *GetSrc<double*>(Data->SSAData, Op->Header.Args[1]);
double trunc = (double)(int64_t)(Src2); //truncate
double Tmp = Src1 * exp2(trunc);
memcpy(GDP, &Tmp, sizeof(double));
}
#undef DEF_OP
} // namespace FEXCore::CPU
@@ -222,11 +222,80 @@ struct OpHandlers<IR::OP_F80SCALE> {
}
};
template<>
struct OpHandlers<IR::OP_F64SIN> {
static double handle(double src) {
return sin(src);
}
};
template<>
struct OpHandlers<IR::OP_F64COS> {
static double handle(double src) {
return cos(src);
}
};
template<>
struct OpHandlers<IR::OP_F64TAN> {
static double handle(double src) {
return tan(src);
}
};
template<>
struct OpHandlers<IR::OP_F64F2XM1> {
static double handle(double src) {
return exp2(src) - 1.0;
}
};
template<>
struct OpHandlers<IR::OP_F64ATAN> {
static double handle(double src1, double src2) {
return atan2(src1, src2);
}
};
template<>
struct OpHandlers<IR::OP_F64FPREM> {
static double handle(double src1, double src2) {
return fmod(src1, src2);
}
};
template<>
struct OpHandlers<IR::OP_F64FPREM1> {
static double handle(double src1, double src2) {
return remainder(src1, src2);
}
};
template<>
struct OpHandlers<IR::OP_F64FYL2X> {
static double handle(double src1, double src2) {
return src2 * log2(src1);
}
};
template<>
struct OpHandlers<IR::OP_F64SCALE> {
static double handle(double src1, double src2) {
double trunc = (double)(int64_t)(src2); //truncate
return src1 * exp2(trunc);
}
};
template<>
struct OpHandlers<IR::OP_F80BCDSTORE> {
static X80SoftFloat handle(X80SoftFloat Src1) {
bool Negative = Src1.Sign;
Src1 = X80SoftFloat::FRNDINT(Src1);
// Clear the Sign bit
Src1.Sign = 0;
@@ -21,8 +21,7 @@ using DestMapType = std::vector<uint32_t>;
class InterpreterCore final : public CPUBackend {
public:
explicit InterpreterCore(FEXCore::Context::Context *ctx,
FEXCore::Core::InternalThreadState *Thread,
bool CompileThread);
FEXCore::Core::InternalThreadState *Thread);
[[nodiscard]] std::string GetName() override { return "Interpreter"; }
@@ -37,6 +36,10 @@ public:
void CreateAsmDispatch(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread);
static void InitializeSignalHandlers(FEXCore::Context::Context *CTX);
bool NeedsRetainedIRCopy() const override { return true; }
private:
FEXCore::Context::Context *CTX;
FEXCore::Core::InternalThreadState *State;
@@ -11,7 +11,6 @@
#include <FEXCore/Utils/LogManager.h>
#include <memory>
#include <bits/types/stack_t.h>
#include <signal.h>
#include <stdint.h>
#include <unordered_map>
@@ -35,32 +34,32 @@ static void InterpreterExecution(FEXCore::Core::CpuStateFrame *Frame) {
InterpreterOps::InterpretIR(Thread, Thread->CurrentFrame->State.rip, LocalEntry->second.IR.get(), LocalEntry->second.DebugData.get());
}
InterpreterCore::InterpreterCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, bool CompileThread)
InterpreterCore::InterpreterCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread)
: CTX {ctx}
, State {Thread} {
if (!CompileThread &&
CTX->Config.Core == FEXCore::Config::CONFIG_INTERPRETER) {
CreateAsmDispatch(ctx, Thread);
CTX->SignalDelegation->RegisterHostSignalHandler(SignalDelegator::SIGNAL_FOR_PAUSE, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
InterpreterCore *Core = reinterpret_cast<InterpreterCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleSignalPause(Signal, info, ucontext);
}, true);
CreateAsmDispatch(ctx, Thread);
}
void InterpreterCore::InitializeSignalHandlers(FEXCore::Context::Context *CTX) {
CTX->SignalDelegation->RegisterHostSignalHandler(SignalDelegator::SIGNAL_FOR_PAUSE, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
InterpreterCore *Core = reinterpret_cast<InterpreterCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleSignalPause(Signal, info, ucontext);
}, true);
#ifdef _M_ARM_64
CTX->SignalDelegation->RegisterHostSignalHandler(SIGBUS, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
return FEXCore::ArchHelpers::Arm64::HandleSIGBUS(true, Signal, info, ucontext);
}, true);
CTX->SignalDelegation->RegisterHostSignalHandler(SIGBUS, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
return FEXCore::ArchHelpers::Arm64::HandleSIGBUS(true, Signal, info, ucontext);
}, true);
#endif
auto GuestSignalHandler = [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack) -> bool {
InterpreterCore *Core = reinterpret_cast<InterpreterCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleGuestSignal(Signal, info, ucontext, GuestAction, GuestStack);
};
auto GuestSignalHandler = [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack) -> bool {
InterpreterCore *Core = reinterpret_cast<InterpreterCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleGuestSignal(Signal, info, ucontext, GuestAction, GuestStack);
};
for (uint32_t Signal = 0; Signal <= SignalDelegator::MAX_SIGNALS; ++Signal) {
CTX->SignalDelegation->RegisterHostSignalHandlerForGuest(Signal, GuestSignalHandler);
}
for (uint32_t Signal = 0; Signal <= SignalDelegator::MAX_SIGNALS; ++Signal) {
CTX->SignalDelegation->RegisterHostSignalHandlerForGuest(Signal, GuestSignalHandler);
}
}
@@ -68,8 +67,12 @@ void *InterpreterCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR:
return reinterpret_cast<void*>(InterpreterExecution);
}
std::unique_ptr<CPUBackend> CreateInterpreterCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, bool CompileThread) {
return std::make_unique<InterpreterCore>(ctx, Thread, CompileThread);
std::unique_ptr<CPUBackend> CreateInterpreterCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread) {
return std::make_unique<InterpreterCore>(ctx, Thread);
}
void InitializeInterpreterSignalHandlers(FEXCore::Context::Context *CTX) {
InterpreterCore::InitializeSignalHandlers(CTX);
}
}
@@ -14,7 +14,8 @@ namespace FEXCore::CPU {
class CPUBackend;
[[nodiscard]] std::unique_ptr<CPUBackend> CreateInterpreterCore(FEXCore::Context::Context *ctx,
FEXCore::Core::InternalThreadState *Thread,
bool CompileThread);
FEXCore::Core::InternalThreadState *Thread);
void InitializeInterpreterSignalHandlers(FEXCore::Context::Context *CTX);
} // namespace FEXCore::CPU
@@ -0,0 +1,313 @@
#include "FEXCore/Core/CoreState.h"
#include "Interface/Core/Interpreter/InterpreterOps.h"
#include "Interface/Core/Interpreter/F80Ops.h"
#include <cstddef>
#include <cstdint>
namespace FEXCore::CPU {
template<typename R, typename... Args>
static FallbackInfo GetFallbackInfo(R(*fn)(Args...), FEXCore::Core::FallbackHandlerIndex HandlerIndex) {
return {FABI_UNKNOWN, (void*)fn, HandlerIndex};
}
template<>
FallbackInfo GetFallbackInfo(X80SoftFloat(*fn)(float), FEXCore::Core::FallbackHandlerIndex HandlerIndex) {
return {FABI_F80_F32, (void*)fn, HandlerIndex};
}
template<>
FallbackInfo GetFallbackInfo(X80SoftFloat(*fn)(double), FEXCore::Core::FallbackHandlerIndex HandlerIndex) {
return {FABI_F80_F64, (void*)fn, HandlerIndex};
}
template<>
FallbackInfo GetFallbackInfo(X80SoftFloat(*fn)(int16_t), FEXCore::Core::FallbackHandlerIndex HandlerIndex) {
return {FABI_F80_I16, (void*)fn, HandlerIndex};
}
template<>
FallbackInfo GetFallbackInfo(void(*fn)(uint16_t), FEXCore::Core::FallbackHandlerIndex HandlerIndex) {
return {FABI_VOID_U16, (void*)fn, HandlerIndex};
}
template<>
FallbackInfo GetFallbackInfo(X80SoftFloat(*fn)(int32_t), FEXCore::Core::FallbackHandlerIndex HandlerIndex) {
return {FABI_F80_I32, (void*)fn, HandlerIndex};
}
template<>
FallbackInfo GetFallbackInfo(float(*fn)(X80SoftFloat), FEXCore::Core::FallbackHandlerIndex HandlerIndex) {
return {FABI_F32_F80, (void*)fn, HandlerIndex};
}
template<>
FallbackInfo GetFallbackInfo(double(*fn)(X80SoftFloat), FEXCore::Core::FallbackHandlerIndex HandlerIndex) {
return {FABI_F64_F80, (void*)fn, HandlerIndex};
}
template<>
FallbackInfo GetFallbackInfo(double(*fn)(double), FEXCore::Core::FallbackHandlerIndex HandlerIndex) {
return {FABI_F64_F64, (void*)fn, HandlerIndex};
}
template<>
FallbackInfo GetFallbackInfo(double(*fn)(double,double), FEXCore::Core::FallbackHandlerIndex HandlerIndex) {
return {FABI_F64_F64_F64, (void*)fn, HandlerIndex};
}
template<>
FallbackInfo GetFallbackInfo(int16_t(*fn)(X80SoftFloat), FEXCore::Core::FallbackHandlerIndex HandlerIndex) {
return {FABI_I16_F80, (void*)fn, HandlerIndex};
}
template<>
FallbackInfo GetFallbackInfo(int32_t(*fn)(X80SoftFloat), FEXCore::Core::FallbackHandlerIndex HandlerIndex) {
return {FABI_I32_F80, (void*)fn, HandlerIndex};
}
template<>
FallbackInfo GetFallbackInfo(int64_t(*fn)(X80SoftFloat), FEXCore::Core::FallbackHandlerIndex HandlerIndex) {
return {FABI_I64_F80, (void*)fn, HandlerIndex};
}
template<>
FallbackInfo GetFallbackInfo(uint64_t(*fn)(X80SoftFloat, X80SoftFloat), FEXCore::Core::FallbackHandlerIndex HandlerIndex) {
return {FABI_I64_F80_F80, (void*)fn, HandlerIndex};
}
template<>
FallbackInfo GetFallbackInfo(X80SoftFloat(*fn)(X80SoftFloat), FEXCore::Core::FallbackHandlerIndex HandlerIndex) {
return {FABI_F80_F80, (void*)fn, HandlerIndex};
}
template<>
FallbackInfo GetFallbackInfo(X80SoftFloat(*fn)(X80SoftFloat, X80SoftFloat), FEXCore::Core::FallbackHandlerIndex HandlerIndex) {
return {FABI_F80_F80_F80, (void*)fn, HandlerIndex};
}
void InterpreterOps::FillFallbackIndexPointers(uint64_t *Info) {
Info[Core::OPINDEX_F80LOADFCW] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80LOADFCW>::handle, Core::OPINDEX_F80LOADFCW).fn);
Info[Core::OPINDEX_F80CVTTO_4] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTO>::handle4, Core::OPINDEX_F80CVTTO_4).fn);
Info[Core::OPINDEX_F80CVTTO_8] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTO>::handle8, Core::OPINDEX_F80CVTTO_8).fn);
Info[Core::OPINDEX_F80CVT_4] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVT>::handle4, Core::OPINDEX_F80CVT_4).fn);
Info[Core::OPINDEX_F80CVT_8] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVT>::handle8, Core::OPINDEX_F80CVT_8).fn);
Info[Core::OPINDEX_F80CVTINT_2] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle2, Core::OPINDEX_F80CVTINT_2).fn);
Info[Core::OPINDEX_F80CVTINT_4] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle4, Core::OPINDEX_F80CVTINT_4).fn);
Info[Core::OPINDEX_F80CVTINT_8] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle8, Core::OPINDEX_F80CVTINT_8).fn);
Info[Core::OPINDEX_F80CVTINT_TRUNC2] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle2t, Core::OPINDEX_F80CVTINT_TRUNC2).fn);
Info[Core::OPINDEX_F80CVTINT_TRUNC4] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle4t, Core::OPINDEX_F80CVTINT_TRUNC4).fn);
Info[Core::OPINDEX_F80CVTINT_TRUNC8] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle8t, Core::OPINDEX_F80CVTINT_TRUNC8).fn);
Info[Core::OPINDEX_F80CMP_0] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<0>, Core::OPINDEX_F80CMP_0).fn);
Info[Core::OPINDEX_F80CMP_1] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<1>, Core::OPINDEX_F80CMP_1).fn);
Info[Core::OPINDEX_F80CMP_2] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<2>, Core::OPINDEX_F80CMP_2).fn);
Info[Core::OPINDEX_F80CMP_3] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<3>, Core::OPINDEX_F80CMP_3).fn);
Info[Core::OPINDEX_F80CMP_4] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<4>, Core::OPINDEX_F80CMP_4).fn);
Info[Core::OPINDEX_F80CMP_5] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<5>, Core::OPINDEX_F80CMP_5).fn);
Info[Core::OPINDEX_F80CMP_6] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<6>, Core::OPINDEX_F80CMP_6).fn);
Info[Core::OPINDEX_F80CMP_7] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<7>, Core::OPINDEX_F80CMP_7).fn);
Info[Core::OPINDEX_F80CVTTOINT_2] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTOINT>::handle2, Core::OPINDEX_F80CVTTOINT_2).fn);
Info[Core::OPINDEX_F80CVTTOINT_4] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTOINT>::handle4, Core::OPINDEX_F80CVTTOINT_4).fn);
// Unary
Info[Core::OPINDEX_F80ROUND] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80ROUND>::handle, Core::OPINDEX_F80ROUND).fn);
Info[Core::OPINDEX_F80F2XM1] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80F2XM1>::handle, Core::OPINDEX_F80F2XM1).fn);
Info[Core::OPINDEX_F80TAN] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80TAN>::handle, Core::OPINDEX_F80TAN).fn);
Info[Core::OPINDEX_F80SQRT] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80SQRT>::handle, Core::OPINDEX_F80SQRT).fn);
Info[Core::OPINDEX_F80SIN] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80SIN>::handle, Core::OPINDEX_F80SIN).fn);
Info[Core::OPINDEX_F80COS] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80COS>::handle, Core::OPINDEX_F80COS).fn);
Info[Core::OPINDEX_F80XTRACT_EXP] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80XTRACT_EXP>::handle, Core::OPINDEX_F80XTRACT_EXP).fn);
Info[Core::OPINDEX_F80XTRACT_SIG] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80XTRACT_SIG>::handle, Core::OPINDEX_F80XTRACT_SIG).fn);
Info[Core::OPINDEX_F80BCDSTORE] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80BCDSTORE>::handle, Core::OPINDEX_F80BCDSTORE).fn);
Info[Core::OPINDEX_F80BCDLOAD] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80BCDLOAD>::handle, Core::OPINDEX_F80BCDLOAD).fn);
// Binary
Info[Core::OPINDEX_F80ADD] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80ADD>::handle, Core::OPINDEX_F80ADD).fn);
Info[Core::OPINDEX_F80SUB] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80SUB>::handle, Core::OPINDEX_F80SUB).fn);
Info[Core::OPINDEX_F80MUL] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80MUL>::handle, Core::OPINDEX_F80MUL).fn);
Info[Core::OPINDEX_F80DIV] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80DIV>::handle, Core::OPINDEX_F80DIV).fn);
Info[Core::OPINDEX_F80FYL2X] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80FYL2X>::handle, Core::OPINDEX_F80FYL2X).fn);
Info[Core::OPINDEX_F80ATAN] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80ATAN>::handle, Core::OPINDEX_F80ATAN).fn);
Info[Core::OPINDEX_F80FPREM1] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80FPREM1>::handle, Core::OPINDEX_F80FPREM1).fn);
Info[Core::OPINDEX_F80FPREM] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80FPREM>::handle, Core::OPINDEX_F80FPREM).fn);
Info[Core::OPINDEX_F80SCALE] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80SCALE>::handle, Core::OPINDEX_F80SCALE).fn);
// Double Precision
Info[Core::OPINDEX_F64SIN] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F64SIN>::handle, Core::OPINDEX_F64SIN).fn);
Info[Core::OPINDEX_F64COS] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F64COS>::handle, Core::OPINDEX_F64COS).fn);
Info[Core::OPINDEX_F64TAN] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F64TAN>::handle, Core::OPINDEX_F64TAN).fn);
Info[Core::OPINDEX_F64ATAN] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F64ATAN>::handle, Core::OPINDEX_F64ATAN).fn);
Info[Core::OPINDEX_F64F2XM1] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F64F2XM1>::handle, Core::OPINDEX_F64F2XM1).fn);
Info[Core::OPINDEX_F64FYL2X] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F64FYL2X>::handle, Core::OPINDEX_F64FYL2X).fn);
Info[Core::OPINDEX_F64FPREM] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F64FPREM>::handle, Core::OPINDEX_F64FPREM).fn);
Info[Core::OPINDEX_F64FPREM1] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F64FPREM1>::handle, Core::OPINDEX_F64FPREM1).fn);
Info[Core::OPINDEX_F64SCALE] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F64SCALE>::handle, Core::OPINDEX_F64SCALE).fn);
}
bool InterpreterOps::GetFallbackHandler(IR::IROp_Header *IROp, FallbackInfo *Info) {
uint8_t OpSize = IROp->Size;
switch(IROp->Op) {
case IR::OP_F80LOADFCW: {
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80LOADFCW>::handle, Core::OPINDEX_F80LOADFCW);
return true;
}
case IR::OP_F80CVTTO: {
auto Op = IROp->C<IR::IROp_F80CVTTo>();
switch (Op->SrcSize) {
case 4: {
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTO>::handle4, Core::OPINDEX_F80CVTTO_4);
return true;
}
case 8: {
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTO>::handle8, Core::OPINDEX_F80CVTTO_8);
return true;
}
default: LogMan::Msg::DFmt("Unhandled size: {}", OpSize);
}
break;
}
case IR::OP_F80CVT: {
switch (OpSize) {
case 4: {
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVT>::handle4, Core::OPINDEX_F80CVT_4);
return true;
}
case 8: {
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVT>::handle8, Core::OPINDEX_F80CVT_8);
return true;
}
default: LogMan::Msg::DFmt("Unhandled size: {}", OpSize);
}
break;
}
case IR::OP_F80CVTINT: {
auto Op = IROp->C<IR::IROp_F80CVTInt>();
switch (OpSize) {
case 2: {
if (Op->Truncate) {
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle2t, Core::OPINDEX_F80CVTINT_TRUNC2);
}
else {
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle2, Core::OPINDEX_F80CVTINT_2);
}
return true;
}
case 4: {
if (Op->Truncate) {
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle4t, Core::OPINDEX_F80CVTINT_TRUNC4);
}
else {
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle4, Core::OPINDEX_F80CVTINT_4);
}
return true;
}
case 8: {
if (Op->Truncate) {
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle8t, Core::OPINDEX_F80CVTINT_TRUNC8);
}
else {
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle8, Core::OPINDEX_F80CVTINT_8);
}
return true;
}
default: LogMan::Msg::DFmt("Unhandled size: {}", OpSize);
}
break;
}
case IR::OP_F80CMP: {
auto Op = IROp->C<IR::IROp_F80Cmp>();
static constexpr std::array handlers{
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<0>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<1>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<2>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<3>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<4>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<5>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<6>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<7>,
};
*Info = GetFallbackInfo(handlers[Op->Flags], (Core::FallbackHandlerIndex)(Core::OPINDEX_F80CMP_0 + Op->Flags));
return true;
}
case IR::OP_F80CVTTOINT: {
auto Op = IROp->C<IR::IROp_F80CVTToInt>();
switch (Op->SrcSize) {
case 2: {
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTOINT>::handle2, Core::OPINDEX_F80CVTTOINT_2);
return true;
}
case 4: {
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTOINT>::handle4, Core::OPINDEX_F80CVTTOINT_4);
return true;
}
default: LogMan::Msg::DFmt("Unhandled size: {}", OpSize);
}
break;
}
#define COMMON_X87_OP(OP) \
case IR::OP_F80##OP: { \
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80##OP>::handle, Core::OPINDEX_F80##OP); \
return true; \
}
#define COMMON_F64_OP(OP) \
case IR::OP_F64##OP: { \
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F64##OP>::handle, Core::OPINDEX_F64##OP); \
return true; \
}
// Unary
COMMON_X87_OP(ROUND)
COMMON_X87_OP(F2XM1)
COMMON_X87_OP(TAN)
COMMON_X87_OP(SQRT)
COMMON_X87_OP(SIN)
COMMON_X87_OP(COS)
COMMON_X87_OP(XTRACT_EXP)
COMMON_X87_OP(XTRACT_SIG)
COMMON_X87_OP(BCDSTORE)
COMMON_X87_OP(BCDLOAD)
// Binary
COMMON_X87_OP(ADD)
COMMON_X87_OP(SUB)
COMMON_X87_OP(MUL)
COMMON_X87_OP(DIV)
COMMON_X87_OP(FYL2X)
COMMON_X87_OP(ATAN)
COMMON_X87_OP(FPREM1)
COMMON_X87_OP(FPREM)
COMMON_X87_OP(SCALE)
// Double Precision Unary
COMMON_F64_OP(F2XM1)
COMMON_F64_OP(TAN)
COMMON_F64_OP(SIN)
COMMON_F64_OP(COS)
// Double Precision Binary
COMMON_F64_OP(FYL2X)
COMMON_F64_OP(ATAN)
COMMON_F64_OP(FPREM1)
COMMON_F64_OP(FPREM)
COMMON_F64_OP(SCALE)
default:
break;
}
return false;
}
}
@@ -114,7 +114,6 @@ constexpr OpHandlerArray InterpreterOpHandlers = [] {
// Branch ops
REGISTER_OP(GUESTCALLDIRECT, GuestCallDirect);
REGISTER_OP(GUESTCALLINDIRECT, GuestCallIndirect);
REGISTER_OP(GUESTRETURN, GuestReturn);
REGISTER_OP(SIGNALRETURN, SignalReturn);
REGISTER_OP(CALLBACKRETURN, CallbackReturn);
REGISTER_OP(EXITFUNCTION, ExitFunction);
@@ -124,7 +123,7 @@ constexpr OpHandlerArray InterpreterOpHandlers = [] {
REGISTER_OP(INLINESYSCALL, InlineSyscall);
REGISTER_OP(THUNK, Thunk);
REGISTER_OP(VALIDATECODE, ValidateCode);
REGISTER_OP(REMOVECODEENTRY, RemoveCodeEntry);
REGISTER_OP(REMOVETHREADCODEENTRY, RemoveThreadCodeEntry);
REGISTER_OP(CPUID, CPUID);
// Conversion ops
@@ -176,6 +175,8 @@ constexpr OpHandlerArray InterpreterOpHandlers = [] {
REGISTER_OP(SETROUNDINGMODE, SetRoundingMode);
REGISTER_OP(INVALIDATEFLAGS, NoOp);
REGISTER_OP(PROCESSORID, ProcessorID);
REGISTER_OP(RDRAND, RDRAND);
REGISTER_OP(YIELD, Yield);
// Move ops
REGISTER_OP(EXTRACTELEMENTPAIR, ExtractElementPair);
@@ -185,8 +186,6 @@ constexpr OpHandlerArray InterpreterOpHandlers = [] {
// Vector ops
REGISTER_OP(VECTORZERO, VectorZero);
REGISTER_OP(VECTORIMM, VectorImm);
REGISTER_OP(CREATEVECTOR2, CreateVector2);
REGISTER_OP(CREATEVECTOR4, CreateVector4);
REGISTER_OP(SPLATVECTOR2, SplatVector);
REGISTER_OP(SPLATVECTOR4, SplatVector);
REGISTER_OP(VMOV, VMov);
@@ -275,6 +274,7 @@ constexpr OpHandlerArray InterpreterOpHandlers = [] {
REGISTER_OP(VSMULL2, VSMull2);
REGISTER_OP(VUABDL, VUABDL);
REGISTER_OP(VTBL1, VTBL1);
REGISTER_OP(VREV64, VRev64);
// Encryption ops
REGISTER_OP(VAESIMC, AESImc);
@@ -283,6 +283,7 @@ constexpr OpHandlerArray InterpreterOpHandlers = [] {
REGISTER_OP(VAESDEC, AESDec);
REGISTER_OP(VAESDECLAST, AESDecLast);
REGISTER_OP(VAESKEYGENASSIST, AESKeyGenAssist);
REGISTER_OP(CRC32, CRC32);
// F80 ops
REGISTER_OP(F80LOADFCW, F80LOADFCW);
@@ -311,6 +312,17 @@ constexpr OpHandlerArray InterpreterOpHandlers = [] {
REGISTER_OP(F80BCDLOAD, F80BCDLOAD);
REGISTER_OP(F80BCDSTORE, F80BCDSTORE);
// F64 ops
REGISTER_OP(F64SIN, F64SIN);
REGISTER_OP(F64COS, F64COS);
REGISTER_OP(F64TAN, F64TAN);
REGISTER_OP(F64F2XM1, F64F2XM1);
REGISTER_OP(F64ATAN, F64ATAN);
REGISTER_OP(F64FPREM, F64FPREM);
REGISTER_OP(F64FPREM1, F64FPREM1);
REGISTER_OP(F64FYL2X, F64FYL2X);
REGISTER_OP(F64SCALE, F64SCALE);
return Handlers;
}();
@@ -321,214 +333,9 @@ void InterpreterOps::Op_Unhandled(FEXCore::IR::IROp_Header *IROp, IROpData *Data
void InterpreterOps::Op_NoOp(FEXCore::IR::IROp_Header *IROp, IROpData *Data, IR::NodeID Node) {
}
template<typename R, typename... Args>
static FallbackInfo GetFallbackInfo(R(*fn)(Args...)) {
return {FABI_UNKNOWN, (void*)fn};
}
template<>
FallbackInfo GetFallbackInfo(X80SoftFloat(*fn)(float)) {
return {FABI_F80_F32, (void*)fn};
}
template<>
FallbackInfo GetFallbackInfo(X80SoftFloat(*fn)(double)) {
return {FABI_F80_F64, (void*)fn};
}
template<>
FallbackInfo GetFallbackInfo(X80SoftFloat(*fn)(int16_t)) {
return {FABI_F80_I16, (void*)fn};
}
template<>
FallbackInfo GetFallbackInfo(void(*fn)(uint16_t)) {
return {FABI_VOID_U16, (void*)fn};
}
template<>
FallbackInfo GetFallbackInfo(X80SoftFloat(*fn)(int32_t)) {
return {FABI_F80_I32, (void*)fn};
}
template<>
FallbackInfo GetFallbackInfo(float(*fn)(X80SoftFloat)) {
return {FABI_F32_F80, (void*)fn};
}
template<>
FallbackInfo GetFallbackInfo(double(*fn)(X80SoftFloat)) {
return {FABI_F64_F80, (void*)fn};
}
template<>
FallbackInfo GetFallbackInfo(int16_t(*fn)(X80SoftFloat)) {
return {FABI_I16_F80, (void*)fn};
}
template<>
FallbackInfo GetFallbackInfo(int32_t(*fn)(X80SoftFloat)) {
return {FABI_I32_F80, (void*)fn};
}
template<>
FallbackInfo GetFallbackInfo(int64_t(*fn)(X80SoftFloat)) {
return {FABI_I64_F80, (void*)fn};
}
template<>
FallbackInfo GetFallbackInfo(uint64_t(*fn)(X80SoftFloat, X80SoftFloat)) {
return {FABI_I64_F80_F80, (void*)fn};
}
template<>
FallbackInfo GetFallbackInfo(X80SoftFloat(*fn)(X80SoftFloat)) {
return {FABI_F80_F80, (void*)fn};
}
template<>
FallbackInfo GetFallbackInfo(X80SoftFloat(*fn)(X80SoftFloat, X80SoftFloat)) {
return {FABI_F80_F80_F80, (void*)fn};
}
bool InterpreterOps::GetFallbackHandler(IR::IROp_Header *IROp, FallbackInfo *Info) {
uint8_t OpSize = IROp->Size;
switch(IROp->Op) {
case IR::OP_F80LOADFCW: {
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80LOADFCW>::handle);
return true;
}
case IR::OP_F80CVTTO: {
auto Op = IROp->C<IR::IROp_F80CVTTo>();
switch (Op->Size) {
case 4: {
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTO>::handle4);
return true;
}
case 8: {
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTO>::handle8);
return true;
}
default: LogMan::Msg::DFmt("Unhandled size: {}", OpSize);
}
break;
}
case IR::OP_F80CVT: {
switch (OpSize) {
case 4: {
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVT>::handle4);
return true;
}
case 8: {
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVT>::handle8);
return true;
}
default: LogMan::Msg::DFmt("Unhandled size: {}", OpSize);
}
break;
}
case IR::OP_F80CVTINT: {
auto Op = IROp->C<IR::IROp_F80CVTInt>();
switch (OpSize) {
case 2: {
*Info = GetFallbackInfo(Op->Truncate ? &FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle2t : &FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle2);
return true;
}
case 4: {
*Info = GetFallbackInfo(Op->Truncate ? &FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle4t : &FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle4);
return true;
}
case 8: {
*Info = GetFallbackInfo(Op->Truncate ? &FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle8t : &FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle8);
return true;
}
default: LogMan::Msg::DFmt("Unhandled size: {}", OpSize);
}
break;
}
case IR::OP_F80CMP: {
auto Op = IROp->C<IR::IROp_F80Cmp>();
static constexpr std::array handlers{
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<0>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<1>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<2>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<3>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<4>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<5>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<6>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<7>,
};
*Info = GetFallbackInfo(handlers[Op->Flags]);
return true;
}
case IR::OP_F80CVTTOINT: {
auto Op = IROp->C<IR::IROp_F80CVTToInt>();
switch (Op->Size) {
case 2: {
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTOINT>::handle2);
return true;
}
case 4: {
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTOINT>::handle4);
return true;
}
default: LogMan::Msg::DFmt("Unhandled size: {}", OpSize);
}
break;
}
#define COMMON_X87_OP(OP) \
case IR::OP_F80##OP: { \
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80##OP>::handle); \
return true; \
}
// Unary
COMMON_X87_OP(ROUND)
COMMON_X87_OP(F2XM1)
COMMON_X87_OP(TAN)
COMMON_X87_OP(SQRT)
COMMON_X87_OP(SIN)
COMMON_X87_OP(COS)
COMMON_X87_OP(XTRACT_EXP)
COMMON_X87_OP(XTRACT_SIG)
COMMON_X87_OP(BCDSTORE)
COMMON_X87_OP(BCDLOAD)
// Binary
COMMON_X87_OP(ADD)
COMMON_X87_OP(SUB)
COMMON_X87_OP(MUL)
COMMON_X87_OP(DIV)
COMMON_X87_OP(FYL2X)
COMMON_X87_OP(ATAN)
COMMON_X87_OP(FPREM1)
COMMON_X87_OP(FPREM)
COMMON_X87_OP(SCALE)
default:
break;
}
return false;
}
void InterpreterOps::InterpretIR(FEXCore::Core::InternalThreadState *Thread, uint64_t Entry, FEXCore::IR::IRListView *CurrentIR, FEXCore::Core::DebugData *DebugData) {
volatile void *StackEntry = alloca(0);
// Debug data is only passed in debug builds
#ifndef NDEBUG
// TODO: should be moved to an IR Op
Thread->Stats.InstructionsExecuted.fetch_add(DebugData->GuestInstructionCount);
#endif
uintptr_t ListSize = CurrentIR->GetSSACount();
static_assert(sizeof(FEXCore::IR::IROp_Header) == 4);
@@ -1,6 +1,7 @@
#pragma once
#include <stdint.h>
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IntrusiveIRList.h>
@@ -27,6 +28,8 @@ namespace FEXCore::CPU {
FABI_F80_I32,
FABI_F32_F80,
FABI_F64_F80,
FABI_F64_F64,
FABI_F64_F64_F64,
FABI_I16_F80,
FABI_I32_F80,
FABI_I64_F80,
@@ -38,12 +41,14 @@ namespace FEXCore::CPU {
struct FallbackInfo {
FallbackABI ABI;
void *fn;
FEXCore::Core::FallbackHandlerIndex HandlerIndex;
};
class InterpreterOps {
public:
static void InterpretIR(FEXCore::Core::InternalThreadState *Thread, uint64_t Entry, FEXCore::IR::IRListView *CurrentIR, FEXCore::Core::DebugData *DebugData);
static void FillFallbackIndexPointers(uint64_t *Info);
static bool GetFallbackHandler(IR::IROp_Header *IROp, FallbackInfo *Info);
struct IROpData {
@@ -139,7 +144,6 @@ namespace FEXCore::CPU {
///< Branch ops
DEF_OP(GuestCallDirect);
DEF_OP(GuestCallIndirect);
DEF_OP(GuestReturn);
DEF_OP(SignalReturn);
DEF_OP(CallbackReturn);
DEF_OP(ExitFunction);
@@ -149,7 +153,7 @@ namespace FEXCore::CPU {
DEF_OP(InlineSyscall);
DEF_OP(Thunk);
DEF_OP(ValidateCode);
DEF_OP(RemoveCodeEntry);
DEF_OP(RemoveThreadCodeEntry);
DEF_OP(CPUID);
///< Conversion ops
@@ -194,6 +198,8 @@ namespace FEXCore::CPU {
DEF_OP(GetRoundingMode);
DEF_OP(SetRoundingMode);
DEF_OP(ProcessorID);
DEF_OP(RDRAND);
DEF_OP(Yield);
///< Move ops
DEF_OP(ExtractElementPair);
@@ -203,8 +209,6 @@ namespace FEXCore::CPU {
///< Vector ops
DEF_OP(VectorZero);
DEF_OP(VectorImm);
DEF_OP(CreateVector2);
DEF_OP(CreateVector4);
DEF_OP(SplatVector);
DEF_OP(VMov);
DEF_OP(VAnd);
@@ -290,6 +294,7 @@ namespace FEXCore::CPU {
DEF_OP(VSMull2);
DEF_OP(VUABDL);
DEF_OP(VTBL1);
DEF_OP(VRev64);
///< Encryption ops
DEF_OP(AESImc);
@@ -298,6 +303,7 @@ namespace FEXCore::CPU {
DEF_OP(AESDec);
DEF_OP(AESDecLast);
DEF_OP(AESKeyGenAssist);
DEF_OP(CRC32);
///< F80 ops
DEF_OP(F80LOADFCW);
@@ -325,6 +331,17 @@ namespace FEXCore::CPU {
DEF_OP(F80CMP);
DEF_OP(F80BCDLOAD);
DEF_OP(F80BCDSTORE);
//< F64 ops
DEF_OP(F64SIN);
DEF_OP(F64COS);
DEF_OP(F64TAN);
DEF_OP(F64F2XM1);
DEF_OP(F64ATAN);
DEF_OP(F64FPREM);
DEF_OP(F64FPREM1);
DEF_OP(F64FYL2X);
DEF_OP(F64SCALE);
#undef DEF_OP
template<typename unsigned_type, typename signed_type, typename float_type>
[[nodiscard]] static bool IsConditionTrue(uint8_t Cond, uint64_t Src1, uint64_t Src2) {
@@ -14,6 +14,7 @@ $end_info$
#ifdef _M_X86_64
#include <xmmintrin.h>
#endif
#include <sys/random.h>
namespace FEXCore::CPU {
[[noreturn]]
@@ -148,6 +149,19 @@ DEF_OP(ProcessorID) {
GD = (CPUNode << 12) | CPU;
}
DEF_OP(RDRAND) {
// We are ignoring Op->GetReseeded in the interpreter
uint64_t *DstPtr = GetDest<uint64_t*>(Data->SSAData, Node);
ssize_t Result = ::getrandom(&DstPtr[0], 8, 0);
// Second result is if we managed to read a valid random number or not
DstPtr[1] = Result == 8 ? 1 : 0;
}
DEF_OP(Yield) {
// Nop implementation
}
#undef DEF_OP
} // namespace FEXCore::CPU
@@ -26,8 +26,8 @@ DEF_OP(CreateElementPair) {
uint8_t *Dst = GetDest<uint8_t*>(Data->SSAData, Node);
memcpy(Dst, Src_Lower, Op->Header.Size);
memcpy(Dst + Op->Header.Size, Src_Upper, Op->Header.Size);
memcpy(Dst, Src_Lower, IROp->ElementSize);
memcpy(Dst + IROp->ElementSize, Src_Upper, IROp->ElementSize);
}
DEF_OP(Mov) {
@@ -7,6 +7,7 @@ $end_info$
#include "Interface/Core/Interpreter/InterpreterClass.h"
#include "Interface/Core/Interpreter/InterpreterOps.h"
#include "Interface/Core/Interpreter/InterpreterDefines.h"
#include <FEXCore/Utils/BitUtils.h>
#include <bit>
#include <cstdint>
@@ -38,39 +39,6 @@ DEF_OP(VectorImm) {
memcpy(GDP, Tmp, OpSize);
}
DEF_OP(CreateVector2) {
auto Op = IROp->C<IR::IROp_CreateVector2>();
uint8_t OpSize = IROp->Size;
LOGMAN_THROW_A_FMT(OpSize <= 16, "Can't handle a vector of size: {}", OpSize);
void *Src1 = GetSrc<void*>(Data->SSAData, Op->Header.Args[0]);
void *Src2 = GetSrc<void*>(Data->SSAData, Op->Header.Args[1]);
uint8_t Tmp[16];
uint8_t ElementSize = OpSize / 2;
#define CREATE_VECTOR(elementsize, type) \
case elementsize: { \
auto *Dst_d = reinterpret_cast<type*>(Tmp); \
auto *Src1_d = reinterpret_cast<type*>(Src1); \
auto *Src2_d = reinterpret_cast<type*>(Src2); \
Dst_d[0] = *Src1_d; \
Dst_d[1] = *Src2_d; \
break; \
}
switch (ElementSize) {
CREATE_VECTOR(1, uint8_t)
CREATE_VECTOR(2, uint16_t)
CREATE_VECTOR(4, uint32_t)
CREATE_VECTOR(8, uint64_t)
default: LOGMAN_MSG_A_FMT("Unknown Element Size: {}", ElementSize); break;
}
#undef CREATE_VECTOR
memcpy(GDP, Tmp, OpSize);
}
DEF_OP(CreateVector4) {
LOGMAN_MSG_A_FMT("Unimplemented");
}
DEF_OP(SplatVector) {
auto Op = IROp->C<IR::IROp_SplatVector2>();
uint8_t OpSize = IROp->Size;
@@ -1402,10 +1370,9 @@ DEF_OP(VInsScalarElement) {
DEF_OP(VExtractElement) {
auto Op = IROp->C<IR::IROp_VExtractElement>();
uint8_t OpSize = IROp->Size;
uint32_t SourceSize = GetOpSize(Data->CurrentIR, Op->Header.Args[0]);
LOGMAN_THROW_A_FMT(OpSize <= 16, "OpSize is too large for VExtractElement: {}", OpSize);
LOGMAN_THROW_A_FMT(IROp->Size <= 16, "OpSize is too large for VExtractElement: {}", IROp->Size);
if (SourceSize == 16) {
__uint128_t SourceMask = (1ULL << (Op->Header.ElementSize * 8)) - 1;
uint64_t Shift = Op->Header.ElementSize * Op->Index * 8;
@@ -1931,6 +1898,39 @@ DEF_OP(VTBL1) {
memcpy(GDP, Tmp, OpSize);
}
DEF_OP(VRev64) {
auto Op = IROp->C<IR::IROp_VRev64>();
uint8_t OpSize = IROp->Size;
void *Src = GetSrc<void*>(Data->SSAData, Op->Header.Args[0]);
uint8_t Tmp[16];
uint8_t Elements = OpSize / 8;
// The element working size is always 64-bit
// The defined element size in the op is the operating size of the element swapping
auto Func8 = [](auto a) { return BSwap64(a); };
auto Func16 = [](auto a) {
return (a >> 48) | // Element[3] -> Element[0]
((a >> 16) & 0xFFFF'0000U) | // Element[2] -> Element[1]
((a << 16) & 0xFFFF'0000'0000ULL) | // Element[1] -> Element[2]
(a << 48); // Element[0] -> Element[3]
};
auto Func32 = [](auto a) {
return (a >> 32) | (a << 32);
};
switch (Op->Header.ElementSize) {
DO_VECTOR_1SRC_OP(1, uint64_t, Func8)
DO_VECTOR_1SRC_OP(2, uint64_t, Func16)
DO_VECTOR_1SRC_OP(4, uint64_t, Func32)
default: LOGMAN_MSG_A_FMT("Unknown Element Size: {}", Op->Header.ElementSize); break;
}
memcpy(GDP, Tmp, Op->Header.Size);
}
#undef DEF_OP
} // namespace FEXCore::CPU
+133 -83
View File
@@ -12,37 +12,13 @@ namespace FEXCore::CPU {
#define GRD(Node) (IROp->Size <= 4 ? GetDst<RA_32>(Node) : GetDst<RA_64>(Node))
#define GRS(Node) (IROp->Size <= 4 ? GetReg<RA_32>(Node) : GetReg<RA_64>(Node))
static uint64_t LUDIV(uint64_t SrcHigh, uint64_t SrcLow, uint64_t Divisor) {
__uint128_t Source = (static_cast<__uint128_t>(SrcHigh) << 64) | SrcLow;
__uint128_t Res = Source / Divisor;
return Res;
}
static int64_t LDIV(int64_t SrcHigh, int64_t SrcLow, int64_t Divisor) {
__int128_t Source = (static_cast<__int128_t>(SrcHigh) << 64) | SrcLow;
__int128_t Res = Source / Divisor;
return Res;
}
static uint64_t LUREM(uint64_t SrcHigh, uint64_t SrcLow, uint64_t Divisor) {
__uint128_t Source = (static_cast<__uint128_t>(SrcHigh) << 64) | SrcLow;
__uint128_t Res = Source % Divisor;
return Res;
}
static int64_t LREM(int64_t SrcHigh, int64_t SrcLow, int64_t Divisor) {
__int128_t Source = (static_cast<__int128_t>(SrcHigh) << 64) | SrcLow;
__int128_t Res = Source % Divisor;
return Res;
}
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(TruncElementPair) {
auto Op = IROp->C<IR::IROp_TruncElementPair>();
switch (Op->Size) {
switch (IROp->Size) {
case 4: {
auto Dst = GetSrcPair<RA_32>(Node);
auto Src = GetSrcPair<RA_32>(Op->Header.Args[0].ID());
@@ -50,7 +26,7 @@ DEF_OP(TruncElementPair) {
mov(Dst.second, Src.second);
break;
}
default: LOGMAN_MSG_A_FMT("Unhandled Truncation size: {}", Op->Size); break;
default: LOGMAN_MSG_A_FMT("Unhandled Truncation size: {}", IROp->Size); break;
}
}
@@ -65,7 +41,13 @@ DEF_OP(EntrypointOffset) {
auto Constant = Entry + Op->Offset;
auto Dst = GetReg<RA_64>(Node);
LoadConstant(Dst, Constant);
uint64_t Mask = ~0ULL;
uint8_t OpSize = IROp->Size;
if (OpSize == 4) {
Mask = 0xFFFF'FFFFULL;
}
LoadConstant(Dst, Constant & Mask);
}
DEF_OP(InlineConstant) {
@@ -654,23 +636,42 @@ DEF_OP(LDiv) {
break;
}
case 8: {
PushDynamicRegsAndLR();
auto Upper64Bit = GetReg<RA_64>(Op->Header.Args[1].ID());
auto Lower64Bit = GetReg<RA_64>(Op->Header.Args[0].ID());
auto Divisor = GetReg<RA_64>(Op->Header.Args[2].ID());
Label Only64Bit{};
Label LongDIVRet{};
mov(x0, GetReg<RA_64>(Op->Header.Args[1].ID()));
mov(x1, GetReg<RA_64>(Op->Header.Args[0].ID()));
mov(x2, GetReg<RA_64>(Op->Header.Args[2].ID()));
// Check if the upper bits match the top bit of the lower 64-bits
// Sign extend the top bit of lower bits
sbfx(TMP1, Lower64Bit, 63, 1);
eor(TMP1, TMP1, Upper64Bit);
LoadConstant(x3, reinterpret_cast<uint64_t>(LDIV));
SpillStaticRegs();
blr(x3);
FillStaticRegs();
// If the sign bit matches then the result is zero
cbz(TMP1, &Only64Bit);
// Result is now in x0
// Fix the stack and any values that were stepped on
PopDynamicRegsAndLR();
// Long divide
{
mov(x0, Upper64Bit);
mov(x1, Lower64Bit);
mov(x2, Divisor);
// Move result to its destination register
mov(GetReg<RA_64>(Node), x0);
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.LDIVHandler)));
blr(x3);
// Move result to its destination register
mov(GetReg<RA_64>(Node), x0);
// Skip 64-bit path
b(&LongDIVRet);
}
bind(&Only64Bit);
// 64-Bit only
{
sdiv(GetReg<RA_64>(Node), Lower64Bit, Divisor);
}
bind(&LongDIVRet);
break;
}
default: LOGMAN_MSG_A_FMT("Unknown LDIV Size: {}", Size); break;
@@ -697,23 +698,38 @@ DEF_OP(LUDiv) {
break;
}
case 8: {
PushDynamicRegsAndLR();
auto Upper64Bit = GetReg<RA_64>(Op->Header.Args[1].ID());
auto Lower64Bit = GetReg<RA_64>(Op->Header.Args[0].ID());
auto Divisor = GetReg<RA_64>(Op->Header.Args[2].ID());
Label Only64Bit{};
Label LongDIVRet{};
mov(x0, GetReg<RA_64>(Op->Header.Args[1].ID()));
mov(x1, GetReg<RA_64>(Op->Header.Args[0].ID()));
mov(x2, GetReg<RA_64>(Op->Header.Args[2].ID()));
// Check the upper bits for zero
// If the upper bits are zero then we can do a 64-bit divide
cbz(Upper64Bit, &Only64Bit);
LoadConstant(x3, reinterpret_cast<uint64_t>(LUDIV));
SpillStaticRegs();
blr(x3);
FillStaticRegs();
// Long divide
{
mov(x0, Upper64Bit);
mov(x1, Lower64Bit);
mov(x2, Divisor);
// Result is now in x0
// Fix the stack and any values that were stepped on
PopDynamicRegsAndLR();
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.LUDIVHandler)));
blr(x3);
// Move result to its destination register
mov(GetReg<RA_64>(Node), x0);
// Move result to its destination register
mov(GetReg<RA_64>(Node), x0);
// Skip 64-bit path
b(&LongDIVRet);
}
bind(&Only64Bit);
// 64-Bit only
{
udiv(GetReg<RA_64>(Node), Lower64Bit, Divisor);
}
bind(&LongDIVRet);
break;
}
default: LOGMAN_MSG_A_FMT("Unknown LUDIV Size: {}", Size); break;
@@ -750,23 +766,42 @@ DEF_OP(LRem) {
break;
}
case 8: {
PushDynamicRegsAndLR();
auto Upper64Bit = GetReg<RA_64>(Op->Header.Args[1].ID());
auto Lower64Bit = GetReg<RA_64>(Op->Header.Args[0].ID());
auto Divisor = GetReg<RA_64>(Op->Header.Args[2].ID());
Label Only64Bit{};
Label LongDIVRet{};
mov(x0, GetReg<RA_64>(Op->Header.Args[1].ID()));
mov(x1, GetReg<RA_64>(Op->Header.Args[0].ID()));
mov(x2, GetReg<RA_64>(Op->Header.Args[2].ID()));
// Check if the upper bits match the top bit of the lower 64-bits
// Sign extend the top bit of lower bits
sbfx(TMP1, Lower64Bit, 63, 1);
eor(TMP1, TMP1, Upper64Bit);
LoadConstant(x3, reinterpret_cast<uint64_t>(LREM));
SpillStaticRegs();
blr(x3);
FillStaticRegs();
// If the sign bit matches then the result is zero
cbz(TMP1, &Only64Bit);
// Result is now in x0
// Fix the stack and any values that were stepped on
PopDynamicRegsAndLR();
// Long divide
{
mov(x0, Upper64Bit);
mov(x1, Lower64Bit);
mov(x2, Divisor);
// Move result to its destination register
mov(GetReg<RA_64>(Node), x0);
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.LREMHandler)));
blr(x3);
// Move result to its destination register
mov(GetReg<RA_64>(Node), x0);
// Skip 64-bit path
b(&LongDIVRet);
}
bind(&Only64Bit);
// 64-Bit only
{
sdiv(TMP1, Lower64Bit, Divisor);
msub(GetReg<RA_64>(Node), TMP1, Divisor, Lower64Bit);
}
bind(&LongDIVRet);
break;
}
default: LOGMAN_MSG_A_FMT("Unknown LREM Size: {}", Size); break;
@@ -799,24 +834,39 @@ DEF_OP(LURem) {
break;
}
case 8: {
auto Upper64Bit = GetReg<RA_64>(Op->Header.Args[1].ID());
auto Lower64Bit = GetReg<RA_64>(Op->Header.Args[0].ID());
auto Divisor = GetReg<RA_64>(Op->Header.Args[2].ID());
Label Only64Bit{};
Label LongDIVRet{};
PushDynamicRegsAndLR();
// Check the upper bits for zero
// If the upper bits are zero then we can do a 64-bit divide
cbz(Upper64Bit, &Only64Bit);
mov(x0, GetReg<RA_64>(Op->Header.Args[1].ID()));
mov(x1, GetReg<RA_64>(Op->Header.Args[0].ID()));
mov(x2, GetReg<RA_64>(Op->Header.Args[2].ID()));
// Long divide
{
mov(x0, Upper64Bit);
mov(x1, Lower64Bit);
mov(x2, Divisor);
LoadConstant(x3, reinterpret_cast<uint64_t>(LUREM));
SpillStaticRegs();
blr(x3);
FillStaticRegs();
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.LUREMHandler)));
blr(x3);
// Move result to its destination register
mov(GetReg<RA_64>(Node), x0);
// Fix the stack and any values that were stepped on
PopDynamicRegsAndLR();
// Skip 64-bit path
b(&LongDIVRet);
}
// Result is now in x0
// Move result to its destination register
mov(GetReg<RA_64>(Node), x0);
bind(&Only64Bit);
// 64-Bit only
{
udiv(TMP1, Lower64Bit, Divisor);
msub(GetReg<RA_64>(Node), TMP1, Divisor, Lower64Bit);
}
bind(&LongDIVRet);
break;
}
default: LOGMAN_MSG_A_FMT("Unknown LUREM Size: {}", OpSize); break;
@@ -1087,16 +1137,16 @@ DEF_OP(VExtractToGPR) {
uint8_t OpSize = IROp->Size;
switch (OpSize) {
case 1:
umov(GetReg<RA_32>(Node), GetSrc(Op->Header.Args[0].ID()).V16B(), Op->Idx);
umov(GetReg<RA_32>(Node), GetSrc(Op->Header.Args[0].ID()).V16B(), Op->Index);
break;
case 2:
umov(GetReg<RA_32>(Node), GetSrc(Op->Header.Args[0].ID()).V8H(), Op->Idx);
umov(GetReg<RA_32>(Node), GetSrc(Op->Header.Args[0].ID()).V8H(), Op->Index);
break;
case 4:
umov(GetReg<RA_32>(Node), GetSrc(Op->Header.Args[0].ID()).V4S(), Op->Idx);
umov(GetReg<RA_32>(Node), GetSrc(Op->Header.Args[0].ID()).V4S(), Op->Index);
break;
case 8:
umov(GetReg<RA_64>(Node), GetSrc(Op->Header.Args[0].ID()).V2D(), Op->Idx);
umov(GetReg<RA_64>(Node), GetSrc(Op->Header.Args[0].ID()).V2D(), Op->Index);
break;
default: LOGMAN_MSG_A_FMT("Unhandled ExtractElementSize: {}", OpSize);
}
+99 -103
View File
@@ -12,18 +12,17 @@ using namespace vixl::aarch64;
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(CASPair) {
auto Op = IROp->C<IR::IROp_CASPair>();
uint8_t OpSize = IROp->Size;
// Size is the size of each pair element
auto Dst = GetSrcPair<RA_64>(Node);
auto Expected = GetSrcPair<RA_64>(Op->Header.Args[0].ID());
auto Desired = GetSrcPair<RA_64>(Op->Header.Args[1].ID());
auto MemSrc = GetReg<RA_64>(Op->Header.Args[2].ID());
auto Expected = GetSrcPair<RA_64>(Op->Expected.ID());
auto Desired = GetSrcPair<RA_64>(Op->Desired.ID());
auto MemSrc = GetReg<RA_64>(Op->Addr.ID());
if (CTX->HostFeatures.SupportsAtomics) {
mov(TMP3, Expected.first);
mov(TMP4, Expected.second);
switch (OpSize) {
switch (IROp->ElementSize) {
case 4:
caspal(TMP3.W(), TMP4.W(), Desired.first.W(), Desired.second.W(), MemOperand(MemSrc));
mov(Dst.first.W(), TMP3.W());
@@ -34,11 +33,11 @@ DEF_OP(CASPair) {
mov(Dst.first, TMP3);
mov(Dst.second, TMP4);
break;
default: LOGMAN_MSG_A_FMT("Unsupported: {}", OpSize);
default: LOGMAN_MSG_A_FMT("Unsupported: {}", IROp->ElementSize);
}
}
else {
switch (OpSize) {
switch (IROp->ElementSize) {
case 4: {
aarch64::Label LoopTop;
aarch64::Label LoopNotExpected;
@@ -91,7 +90,7 @@ DEF_OP(CASPair) {
bind(&LoopExpected);
break;
}
default: LOGMAN_MSG_A_FMT("Unsupported: {}", OpSize);
default: LOGMAN_MSG_A_FMT("Unsupported: {}", IROp->ElementSize);
}
}
}
@@ -99,16 +98,13 @@ DEF_OP(CASPair) {
DEF_OP(CAS) {
auto Op = IROp->C<IR::IROp_CAS>();
uint8_t OpSize = IROp->Size;
// Args[0]: Expected
// Args[1]: Desired
// Args[2]: Pointer
// DataSrc = *Src1
// if (DataSrc == Src3) { *Src1 == Src2; } Src2 = DataSrc
// This will write to memory! Careful!
auto Expected = GetReg<RA_64>(Op->Header.Args[0].ID());
auto Desired = GetReg<RA_64>(Op->Header.Args[1].ID());
auto MemSrc = GetReg<RA_64>(Op->Header.Args[2].ID());
auto Expected = GetReg<RA_64>(Op->Expected.ID());
auto Desired = GetReg<RA_64>(Op->Desired.ID());
auto MemSrc = GetReg<RA_64>(Op->Addr.ID());
if (CTX->HostFeatures.SupportsAtomics) {
mov(TMP2, Expected);
@@ -216,14 +212,14 @@ DEF_OP(CAS) {
DEF_OP(AtomicAdd) {
auto Op = IROp->C<IR::IROp_AtomicAdd>();
auto MemSrc = GetReg<RA_64>(Op->Header.Args[0].ID());
auto MemSrc = GetReg<RA_64>(Op->Addr.ID());
if (CTX->HostFeatures.SupportsAtomics) {
switch (IROp->Size) {
case 1: staddlb(GetReg<RA_32>(Op->Header.Args[1].ID()), MemOperand(MemSrc)); break;
case 2: staddlh(GetReg<RA_32>(Op->Header.Args[1].ID()), MemOperand(MemSrc)); break;
case 4: staddl(GetReg<RA_32>(Op->Header.Args[1].ID()), MemOperand(MemSrc)); break;
case 8: staddl(GetReg<RA_64>(Op->Header.Args[1].ID()), MemOperand(MemSrc)); break;
case 1: staddlb(GetReg<RA_32>(Op->Value.ID()), MemOperand(MemSrc)); break;
case 2: staddlh(GetReg<RA_32>(Op->Value.ID()), MemOperand(MemSrc)); break;
case 4: staddl(GetReg<RA_32>(Op->Value.ID()), MemOperand(MemSrc)); break;
case 8: staddl(GetReg<RA_64>(Op->Value.ID()), MemOperand(MemSrc)); break;
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", IROp->Size);
}
}
@@ -234,7 +230,7 @@ DEF_OP(AtomicAdd) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxrb(TMP2.W(), MemOperand(MemSrc));
add(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
add(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxrb(TMP2.W(), TMP2.W(), MemOperand(MemSrc));
cbnz(TMP2.W(), &LoopTop);
break;
@@ -243,7 +239,7 @@ DEF_OP(AtomicAdd) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxrh(TMP2.W(), MemOperand(MemSrc));
add(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
add(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxrh(TMP2.W(), TMP2.W(), MemOperand(MemSrc));
cbnz(TMP2.W(), &LoopTop);
break;
@@ -252,7 +248,7 @@ DEF_OP(AtomicAdd) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxr(TMP2.W(), MemOperand(MemSrc));
add(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
add(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxr(TMP2.W(), TMP2.W(), MemOperand(MemSrc));
cbnz(TMP2.W(), &LoopTop);
break;
@@ -261,7 +257,7 @@ DEF_OP(AtomicAdd) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxr(TMP2, MemOperand(MemSrc));
add(TMP2, TMP2, GetReg<RA_64>(Op->Header.Args[1].ID()));
add(TMP2, TMP2, GetReg<RA_64>(Op->Value.ID()));
stlxr(TMP2, TMP2, MemOperand(MemSrc));
cbnz(TMP2, &LoopTop);
break;
@@ -274,10 +270,10 @@ DEF_OP(AtomicAdd) {
DEF_OP(AtomicSub) {
auto Op = IROp->C<IR::IROp_AtomicSub>();
auto MemSrc = GetReg<RA_64>(Op->Header.Args[0].ID());
auto MemSrc = GetReg<RA_64>(Op->Addr.ID());
if (CTX->HostFeatures.SupportsAtomics) {
neg(TMP2, GetReg<RA_64>(Op->Header.Args[1].ID()));
neg(TMP2, GetReg<RA_64>(Op->Value.ID()));
switch (IROp->Size) {
case 1: staddlb(TMP2.W(), MemOperand(MemSrc)); break;
case 2: staddlh(TMP2.W(), MemOperand(MemSrc)); break;
@@ -293,7 +289,7 @@ DEF_OP(AtomicSub) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxrb(TMP2.W(), MemOperand(MemSrc));
sub(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
sub(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxrb(TMP2.W(), TMP2.W(), MemOperand(MemSrc));
cbnz(TMP2.W(), &LoopTop);
break;
@@ -302,7 +298,7 @@ DEF_OP(AtomicSub) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxrh(TMP2.W(), MemOperand(MemSrc));
sub(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
sub(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxrh(TMP2.W(), TMP2.W(), MemOperand(MemSrc));
cbnz(TMP2.W(), &LoopTop);
break;
@@ -311,7 +307,7 @@ DEF_OP(AtomicSub) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxr(TMP2.W(), MemOperand(MemSrc));
sub(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
sub(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxr(TMP2.W(), TMP2.W(), MemOperand(MemSrc));
cbnz(TMP2.W(), &LoopTop);
break;
@@ -320,7 +316,7 @@ DEF_OP(AtomicSub) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxr(TMP2, MemOperand(MemSrc));
sub(TMP2, TMP2, GetReg<RA_64>(Op->Header.Args[1].ID()));
sub(TMP2, TMP2, GetReg<RA_64>(Op->Value.ID()));
stlxr(TMP2, TMP2, MemOperand(MemSrc));
cbnz(TMP2, &LoopTop);
break;
@@ -333,10 +329,10 @@ DEF_OP(AtomicSub) {
DEF_OP(AtomicAnd) {
auto Op = IROp->C<IR::IROp_AtomicAnd>();
auto MemSrc = GetReg<RA_64>(Op->Header.Args[0].ID());
auto MemSrc = GetReg<RA_64>(Op->Addr.ID());
if (CTX->HostFeatures.SupportsAtomics) {
mvn(TMP2, GetReg<RA_64>(Op->Header.Args[1].ID()));
mvn(TMP2, GetReg<RA_64>(Op->Value.ID()));
switch (IROp->Size) {
case 1: stclrlb(TMP2.W(), MemOperand(MemSrc)); break;
case 2: stclrlh(TMP2.W(), MemOperand(MemSrc)); break;
@@ -352,7 +348,7 @@ DEF_OP(AtomicAnd) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxrb(TMP2.W(), MemOperand(MemSrc));
and_(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
and_(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxrb(TMP2.W(), TMP2.W(), MemOperand(MemSrc));
cbnz(TMP2.W(), &LoopTop);
break;
@@ -361,7 +357,7 @@ DEF_OP(AtomicAnd) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxrh(TMP2.W(), MemOperand(MemSrc));
and_(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
and_(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxrh(TMP2.W(), TMP2.W(), MemOperand(MemSrc));
cbnz(TMP2.W(), &LoopTop);
break;
@@ -370,7 +366,7 @@ DEF_OP(AtomicAnd) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxr(TMP2.W(), MemOperand(MemSrc));
and_(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
and_(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxr(TMP2.W(), TMP2.W(), MemOperand(MemSrc));
cbnz(TMP2.W(), &LoopTop);
break;
@@ -379,7 +375,7 @@ DEF_OP(AtomicAnd) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxr(TMP2, MemOperand(MemSrc));
and_(TMP2, TMP2, GetReg<RA_64>(Op->Header.Args[1].ID()));
and_(TMP2, TMP2, GetReg<RA_64>(Op->Value.ID()));
stlxr(TMP2, TMP2, MemOperand(MemSrc));
cbnz(TMP2, &LoopTop);
break;
@@ -392,14 +388,14 @@ DEF_OP(AtomicAnd) {
DEF_OP(AtomicOr) {
auto Op = IROp->C<IR::IROp_AtomicOr>();
auto MemSrc = GetReg<RA_64>(Op->Header.Args[0].ID());
auto MemSrc = GetReg<RA_64>(Op->Addr.ID());
if (CTX->HostFeatures.SupportsAtomics) {
switch (IROp->Size) {
case 1: stsetlb(GetReg<RA_32>(Op->Header.Args[1].ID()), MemOperand(MemSrc)); break;
case 2: stsetlh(GetReg<RA_32>(Op->Header.Args[1].ID()), MemOperand(MemSrc)); break;
case 4: stsetl(GetReg<RA_32>(Op->Header.Args[1].ID()), MemOperand(MemSrc)); break;
case 8: stsetl(GetReg<RA_64>(Op->Header.Args[1].ID()), MemOperand(MemSrc)); break;
case 1: stsetlb(GetReg<RA_32>(Op->Value.ID()), MemOperand(MemSrc)); break;
case 2: stsetlh(GetReg<RA_32>(Op->Value.ID()), MemOperand(MemSrc)); break;
case 4: stsetl(GetReg<RA_32>(Op->Value.ID()), MemOperand(MemSrc)); break;
case 8: stsetl(GetReg<RA_64>(Op->Value.ID()), MemOperand(MemSrc)); break;
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", IROp->Size);
}
}
@@ -410,7 +406,7 @@ DEF_OP(AtomicOr) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxrb(TMP2.W(), MemOperand(MemSrc));
orr(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
orr(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxrb(TMP2.W(), TMP2.W(), MemOperand(MemSrc));
cbnz(TMP2.W(), &LoopTop);
break;
@@ -419,7 +415,7 @@ DEF_OP(AtomicOr) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxrh(TMP2.W(), MemOperand(MemSrc));
orr(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
orr(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxrh(TMP2.W(), TMP2.W(), MemOperand(MemSrc));
cbnz(TMP2.W(), &LoopTop);
break;
@@ -428,7 +424,7 @@ DEF_OP(AtomicOr) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxr(TMP2.W(), MemOperand(MemSrc));
orr(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
orr(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxr(TMP2.W(), TMP2.W(), MemOperand(MemSrc));
cbnz(TMP2.W(), &LoopTop);
break;
@@ -437,7 +433,7 @@ DEF_OP(AtomicOr) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxr(TMP2, MemOperand(MemSrc));
orr(TMP2, TMP2, GetReg<RA_64>(Op->Header.Args[1].ID()));
orr(TMP2, TMP2, GetReg<RA_64>(Op->Value.ID()));
stlxr(TMP2, TMP2, MemOperand(MemSrc));
cbnz(TMP2, &LoopTop);
break;
@@ -450,14 +446,14 @@ DEF_OP(AtomicOr) {
DEF_OP(AtomicXor) {
auto Op = IROp->C<IR::IROp_AtomicXor>();
auto MemSrc = GetReg<RA_64>(Op->Header.Args[0].ID());
auto MemSrc = GetReg<RA_64>(Op->Addr.ID());
if (CTX->HostFeatures.SupportsAtomics) {
switch (IROp->Size) {
case 1: steorlb(GetReg<RA_32>(Op->Header.Args[1].ID()), MemOperand(MemSrc)); break;
case 2: steorlh(GetReg<RA_32>(Op->Header.Args[1].ID()), MemOperand(MemSrc)); break;
case 4: steorl(GetReg<RA_32>(Op->Header.Args[1].ID()), MemOperand(MemSrc)); break;
case 8: steorl(GetReg<RA_64>(Op->Header.Args[1].ID()), MemOperand(MemSrc)); break;
case 1: steorlb(GetReg<RA_32>(Op->Value.ID()), MemOperand(MemSrc)); break;
case 2: steorlh(GetReg<RA_32>(Op->Value.ID()), MemOperand(MemSrc)); break;
case 4: steorl(GetReg<RA_32>(Op->Value.ID()), MemOperand(MemSrc)); break;
case 8: steorl(GetReg<RA_64>(Op->Value.ID()), MemOperand(MemSrc)); break;
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", IROp->Size);
}
}
@@ -468,7 +464,7 @@ DEF_OP(AtomicXor) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxrb(TMP2.W(), MemOperand(MemSrc));
eor(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
eor(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxrb(TMP2.W(), TMP2.W(), MemOperand(MemSrc));
cbnz(TMP2.W(), &LoopTop);
break;
@@ -477,7 +473,7 @@ DEF_OP(AtomicXor) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxrh(TMP2.W(), MemOperand(MemSrc));
eor(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
eor(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxrh(TMP2.W(), TMP2.W(), MemOperand(MemSrc));
cbnz(TMP2.W(), &LoopTop);
break;
@@ -486,7 +482,7 @@ DEF_OP(AtomicXor) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxr(TMP2.W(), MemOperand(MemSrc));
eor(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
eor(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxr(TMP2.W(), TMP2.W(), MemOperand(MemSrc));
cbnz(TMP2.W(), &LoopTop);
break;
@@ -495,7 +491,7 @@ DEF_OP(AtomicXor) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxr(TMP2, MemOperand(MemSrc));
eor(TMP2, TMP2, GetReg<RA_64>(Op->Header.Args[1].ID()));
eor(TMP2, TMP2, GetReg<RA_64>(Op->Value.ID()));
stlxr(TMP2, TMP2, MemOperand(MemSrc));
cbnz(TMP2, &LoopTop);
break;
@@ -508,15 +504,15 @@ DEF_OP(AtomicXor) {
DEF_OP(AtomicSwap) {
auto Op = IROp->C<IR::IROp_AtomicSwap>();
auto MemSrc = GetReg<RA_64>(Op->Header.Args[0].ID());
auto MemSrc = GetReg<RA_64>(Op->Addr.ID());
if (CTX->HostFeatures.SupportsAtomics) {
mov(TMP2, GetReg<RA_64>(Op->Header.Args[1].ID()));
mov(TMP2, GetReg<RA_64>(Op->Value.ID()));
switch (IROp->Size) {
case 1: swplb(TMP2.W(), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 2: swplh(TMP2.W(), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 4: swpl(TMP2.W(), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 8: swpl(TMP2.X(), GetReg<RA_64>(Node), MemOperand(MemSrc)); break;
case 1: swpalb(TMP2.W(), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 2: swpalh(TMP2.W(), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 4: swpal(TMP2.W(), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 8: swpal(TMP2.X(), GetReg<RA_64>(Node), MemOperand(MemSrc)); break;
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", IROp->Size);
}
}
@@ -527,7 +523,7 @@ DEF_OP(AtomicSwap) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxrb(TMP2.W(), MemOperand(MemSrc));
stlxrb(TMP4.W(), GetReg<RA_32>(Op->Header.Args[1].ID()), MemOperand(MemSrc));
stlxrb(TMP4.W(), GetReg<RA_32>(Op->Value.ID()), MemOperand(MemSrc));
cbnz(TMP4.W(), &LoopTop);
uxtb(GetReg<RA_32>(Node), TMP2.W());
break;
@@ -536,7 +532,7 @@ DEF_OP(AtomicSwap) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxrh(TMP2.W(), MemOperand(MemSrc));
stlxrh(TMP4.W(), GetReg<RA_32>(Op->Header.Args[1].ID()), MemOperand(MemSrc));
stlxrh(TMP4.W(), GetReg<RA_32>(Op->Value.ID()), MemOperand(MemSrc));
cbnz(TMP4.W(), &LoopTop);
uxtw(GetReg<RA_32>(Node), TMP2.W());
break;
@@ -545,7 +541,7 @@ DEF_OP(AtomicSwap) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxr(TMP2.W(), MemOperand(MemSrc));
stlxr(TMP4.W(), GetReg<RA_32>(Op->Header.Args[1].ID()), MemOperand(MemSrc));
stlxr(TMP4.W(), GetReg<RA_32>(Op->Value.ID()), MemOperand(MemSrc));
cbnz(TMP4.W(), &LoopTop);
mov(GetReg<RA_32>(Node), TMP2.W());
break;
@@ -554,7 +550,7 @@ DEF_OP(AtomicSwap) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxr(TMP2, MemOperand(MemSrc));
stlxr(TMP4, GetReg<RA_64>(Op->Header.Args[1].ID()), MemOperand(MemSrc));
stlxr(TMP4, GetReg<RA_64>(Op->Value.ID()), MemOperand(MemSrc));
cbnz(TMP4, &LoopTop);
mov(GetReg<RA_64>(Node), TMP2.X());
break;
@@ -566,14 +562,14 @@ DEF_OP(AtomicSwap) {
DEF_OP(AtomicFetchAdd) {
auto Op = IROp->C<IR::IROp_AtomicFetchAdd>();
auto MemSrc = GetReg<RA_64>(Op->Header.Args[0].ID());
auto MemSrc = GetReg<RA_64>(Op->Addr.ID());
if (CTX->HostFeatures.SupportsAtomics) {
switch (IROp->Size) {
case 1: ldaddalb(GetReg<RA_32>(Op->Header.Args[1].ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 2: ldaddalh(GetReg<RA_32>(Op->Header.Args[1].ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 4: ldaddal(GetReg<RA_32>(Op->Header.Args[1].ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 8: ldaddal(GetReg<RA_64>(Op->Header.Args[1].ID()), GetReg<RA_64>(Node), MemOperand(MemSrc)); break;
case 1: ldaddalb(GetReg<RA_32>(Op->Value.ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 2: ldaddalh(GetReg<RA_32>(Op->Value.ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 4: ldaddal(GetReg<RA_32>(Op->Value.ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 8: ldaddal(GetReg<RA_64>(Op->Value.ID()), GetReg<RA_64>(Node), MemOperand(MemSrc)); break;
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", IROp->Size);
}
}
@@ -584,7 +580,7 @@ DEF_OP(AtomicFetchAdd) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxrb(TMP2.W(), MemOperand(MemSrc));
add(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
add(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxrb(TMP4.W(), TMP3.W(), MemOperand(MemSrc));
cbnz(TMP4.W(), &LoopTop);
mov(GetReg<RA_32>(Node), TMP2.W());
@@ -594,7 +590,7 @@ DEF_OP(AtomicFetchAdd) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxrh(TMP2.W(), MemOperand(MemSrc));
add(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
add(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxrh(TMP4.W(), TMP3.W(), MemOperand(MemSrc));
cbnz(TMP4.W(), &LoopTop);
mov(GetReg<RA_32>(Node), TMP2.W());
@@ -604,7 +600,7 @@ DEF_OP(AtomicFetchAdd) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxr(TMP2.W(), MemOperand(MemSrc));
add(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
add(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxr(TMP4.W(), TMP3.W(), MemOperand(MemSrc));
cbnz(TMP4.W(), &LoopTop);
mov(GetReg<RA_32>(Node), TMP2.W());
@@ -614,7 +610,7 @@ DEF_OP(AtomicFetchAdd) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxr(TMP2, MemOperand(MemSrc));
add(TMP3, TMP2, GetReg<RA_64>(Op->Header.Args[1].ID()));
add(TMP3, TMP2, GetReg<RA_64>(Op->Value.ID()));
stlxr(TMP4, TMP3, MemOperand(MemSrc));
cbnz(TMP4, &LoopTop);
mov(GetReg<RA_64>(Node), TMP2);
@@ -627,10 +623,10 @@ DEF_OP(AtomicFetchAdd) {
DEF_OP(AtomicFetchSub) {
auto Op = IROp->C<IR::IROp_AtomicFetchSub>();
auto MemSrc = GetReg<RA_64>(Op->Header.Args[0].ID());
auto MemSrc = GetReg<RA_64>(Op->Addr.ID());
if (CTX->HostFeatures.SupportsAtomics) {
neg(TMP2, GetReg<RA_64>(Op->Header.Args[1].ID()));
neg(TMP2, GetReg<RA_64>(Op->Value.ID()));
switch (IROp->Size) {
case 1: ldaddalb(TMP2.W(), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 2: ldaddalh(TMP2.W(), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
@@ -646,7 +642,7 @@ DEF_OP(AtomicFetchSub) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxrb(TMP2.W(), MemOperand(MemSrc));
sub(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
sub(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxrb(TMP4.W(), TMP3.W(), MemOperand(MemSrc));
cbnz(TMP4.W(), &LoopTop);
mov(GetReg<RA_32>(Node), TMP2.W());
@@ -656,7 +652,7 @@ DEF_OP(AtomicFetchSub) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxrh(TMP2.W(), MemOperand(MemSrc));
sub(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
sub(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxrh(TMP4.W(), TMP3.W(), MemOperand(MemSrc));
cbnz(TMP4.W(), &LoopTop);
mov(GetReg<RA_32>(Node), TMP2.W());
@@ -666,7 +662,7 @@ DEF_OP(AtomicFetchSub) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxr(TMP2.W(), MemOperand(MemSrc));
sub(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
sub(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxr(TMP4.W(), TMP3.W(), MemOperand(MemSrc));
cbnz(TMP4.W(), &LoopTop);
mov(GetReg<RA_32>(Node), TMP2.W());
@@ -676,7 +672,7 @@ DEF_OP(AtomicFetchSub) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxr(TMP2, MemOperand(MemSrc));
sub(TMP3, TMP2, GetReg<RA_64>(Op->Header.Args[1].ID()));
sub(TMP3, TMP2, GetReg<RA_64>(Op->Value.ID()));
stlxr(TMP4, TMP3, MemOperand(MemSrc));
cbnz(TMP4, &LoopTop);
mov(GetReg<RA_64>(Node), TMP2);
@@ -689,10 +685,10 @@ DEF_OP(AtomicFetchSub) {
DEF_OP(AtomicFetchAnd) {
auto Op = IROp->C<IR::IROp_AtomicFetchAnd>();
auto MemSrc = GetReg<RA_64>(Op->Header.Args[0].ID());
auto MemSrc = GetReg<RA_64>(Op->Addr.ID());
if (CTX->HostFeatures.SupportsAtomics) {
mvn(TMP2, GetReg<RA_64>(Op->Header.Args[1].ID()));
mvn(TMP2, GetReg<RA_64>(Op->Value.ID()));
switch (IROp->Size) {
case 1: ldclralb(TMP2.W(), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 2: ldclralh(TMP2.W(), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
@@ -708,7 +704,7 @@ DEF_OP(AtomicFetchAnd) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxrb(TMP2.W(), MemOperand(MemSrc));
and_(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
and_(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxrb(TMP4.W(), TMP3.W(), MemOperand(MemSrc));
cbnz(TMP4.W(), &LoopTop);
mov(GetReg<RA_32>(Node), TMP2.W());
@@ -718,7 +714,7 @@ DEF_OP(AtomicFetchAnd) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxrh(TMP2.W(), MemOperand(MemSrc));
and_(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
and_(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxrh(TMP4.W(), TMP3.W(), MemOperand(MemSrc));
cbnz(TMP4.W(), &LoopTop);
mov(GetReg<RA_32>(Node), TMP2.W());
@@ -728,7 +724,7 @@ DEF_OP(AtomicFetchAnd) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxr(TMP2.W(), MemOperand(MemSrc));
and_(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
and_(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxr(TMP4.W(), TMP3.W(), MemOperand(MemSrc));
cbnz(TMP4.W(), &LoopTop);
mov(GetReg<RA_32>(Node), TMP2.W());
@@ -738,7 +734,7 @@ DEF_OP(AtomicFetchAnd) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxr(TMP2, MemOperand(MemSrc));
and_(TMP3, TMP2, GetReg<RA_64>(Op->Header.Args[1].ID()));
and_(TMP3, TMP2, GetReg<RA_64>(Op->Value.ID()));
stlxr(TMP4, TMP3, MemOperand(MemSrc));
cbnz(TMP4, &LoopTop);
mov(GetReg<RA_64>(Node), TMP2);
@@ -751,14 +747,14 @@ DEF_OP(AtomicFetchAnd) {
DEF_OP(AtomicFetchOr) {
auto Op = IROp->C<IR::IROp_AtomicFetchOr>();
auto MemSrc = GetReg<RA_64>(Op->Header.Args[0].ID());
auto MemSrc = GetReg<RA_64>(Op->Addr.ID());
if (CTX->HostFeatures.SupportsAtomics) {
switch (IROp->Size) {
case 1: ldsetalb(GetReg<RA_32>(Op->Header.Args[1].ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 2: ldsetalh(GetReg<RA_32>(Op->Header.Args[1].ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 4: ldsetal(GetReg<RA_32>(Op->Header.Args[1].ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 8: ldsetal(GetReg<RA_64>(Op->Header.Args[1].ID()), GetReg<RA_64>(Node), MemOperand(MemSrc)); break;
case 1: ldsetalb(GetReg<RA_32>(Op->Value.ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 2: ldsetalh(GetReg<RA_32>(Op->Value.ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 4: ldsetal(GetReg<RA_32>(Op->Value.ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 8: ldsetal(GetReg<RA_64>(Op->Value.ID()), GetReg<RA_64>(Node), MemOperand(MemSrc)); break;
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", IROp->Size);
}
}
@@ -769,7 +765,7 @@ DEF_OP(AtomicFetchOr) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxrb(TMP2.W(), MemOperand(MemSrc));
orr(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
orr(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxrb(TMP4.W(), TMP3.W(), MemOperand(MemSrc));
cbnz(TMP4.W(), &LoopTop);
mov(GetReg<RA_32>(Node), TMP2.W());
@@ -779,7 +775,7 @@ DEF_OP(AtomicFetchOr) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxrh(TMP2.W(), MemOperand(MemSrc));
orr(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
orr(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxrh(TMP4.W(), TMP3.W(), MemOperand(MemSrc));
cbnz(TMP4.W(), &LoopTop);
mov(GetReg<RA_32>(Node), TMP2.W());
@@ -789,7 +785,7 @@ DEF_OP(AtomicFetchOr) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxr(TMP2.W(), MemOperand(MemSrc));
orr(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
orr(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxr(TMP4.W(), TMP3.W(), MemOperand(MemSrc));
cbnz(TMP4.W(), &LoopTop);
mov(GetReg<RA_32>(Node), TMP2.W());
@@ -799,7 +795,7 @@ DEF_OP(AtomicFetchOr) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxr(TMP2, MemOperand(MemSrc));
orr(TMP3, TMP2, GetReg<RA_64>(Op->Header.Args[1].ID()));
orr(TMP3, TMP2, GetReg<RA_64>(Op->Value.ID()));
stlxr(TMP4, TMP3, MemOperand(MemSrc));
cbnz(TMP4, &LoopTop);
mov(GetReg<RA_64>(Node), TMP2);
@@ -812,14 +808,14 @@ DEF_OP(AtomicFetchOr) {
DEF_OP(AtomicFetchXor) {
auto Op = IROp->C<IR::IROp_AtomicFetchXor>();
auto MemSrc = GetReg<RA_64>(Op->Header.Args[0].ID());
auto MemSrc = GetReg<RA_64>(Op->Addr.ID());
if (CTX->HostFeatures.SupportsAtomics) {
switch (IROp->Size) {
case 1: ldeoralb(GetReg<RA_32>(Op->Header.Args[1].ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 2: ldeoralh(GetReg<RA_32>(Op->Header.Args[1].ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 4: ldeoral(GetReg<RA_32>(Op->Header.Args[1].ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 8: ldeoral(GetReg<RA_64>(Op->Header.Args[1].ID()), GetReg<RA_64>(Node), MemOperand(MemSrc)); break;
case 1: ldeoralb(GetReg<RA_32>(Op->Value.ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 2: ldeoralh(GetReg<RA_32>(Op->Value.ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 4: ldeoral(GetReg<RA_32>(Op->Value.ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 8: ldeoral(GetReg<RA_64>(Op->Value.ID()), GetReg<RA_64>(Node), MemOperand(MemSrc)); break;
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", IROp->Size);
}
}
@@ -830,7 +826,7 @@ DEF_OP(AtomicFetchXor) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxrb(TMP2.W(), MemOperand(MemSrc));
eor(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
eor(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxrb(TMP4.W(), TMP3.W(), MemOperand(MemSrc));
cbnz(TMP4.W(), &LoopTop);
mov(GetReg<RA_32>(Node), TMP2.W());
@@ -840,7 +836,7 @@ DEF_OP(AtomicFetchXor) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxrh(TMP2.W(), MemOperand(MemSrc));
eor(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
eor(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxrh(TMP4.W(), TMP3.W(), MemOperand(MemSrc));
cbnz(TMP4.W(), &LoopTop);
mov(GetReg<RA_32>(Node), TMP2.W());
@@ -850,7 +846,7 @@ DEF_OP(AtomicFetchXor) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxr(TMP2.W(), MemOperand(MemSrc));
eor(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
eor(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxr(TMP4.W(), TMP3.W(), MemOperand(MemSrc));
cbnz(TMP4.W(), &LoopTop);
mov(GetReg<RA_32>(Node), TMP2.W());
@@ -860,7 +856,7 @@ DEF_OP(AtomicFetchXor) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxr(TMP2, MemOperand(MemSrc));
eor(TMP3, TMP2, GetReg<RA_64>(Op->Header.Args[1].ID()));
eor(TMP3, TMP2, GetReg<RA_64>(Op->Value.ID()));
stlxr(TMP4, TMP3, MemOperand(MemSrc));
cbnz(TMP4, &LoopTop);
mov(GetReg<RA_64>(Node), TMP2);
@@ -873,7 +869,7 @@ DEF_OP(AtomicFetchXor) {
DEF_OP(AtomicFetchNeg) {
auto Op = IROp->C<IR::IROp_AtomicFetchNeg>();
auto MemSrc = GetReg<RA_64>(Op->Header.Args[0].ID());
auto MemSrc = GetReg<RA_64>(Op->Addr.ID());
// TMP2-TMP3
switch (IROp->Size) {
@@ -4,6 +4,7 @@ tags: backend|arm64
$end_info$
*/
#include "FEXCore/IR/IR.h"
#include "Interface/Core/LookupCache.h"
#include "Interface/Core/JIT/Arm64/JITClass.h"
@@ -26,17 +27,13 @@ DEF_OP(GuestCallIndirect) {
LogMan::Msg::DFmt("Unimplemented");
}
DEF_OP(GuestReturn) {
LogMan::Msg::DFmt("Unimplemented");
}
DEF_OP(SignalReturn) {
// First we must reset the stack
ResetStack();
// Now branch to our signal return helper
// This can't be a direct branch since the code needs to live at a constant location
LoadConstant(x0, ThreadSharedData.SignalReturnInstruction);
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.SignalReturnHandler)));
br(x0);
}
@@ -49,7 +46,7 @@ DEF_OP(CallbackReturn) {
ResetStack();
// We can now lower the ref counter again
LoadConstant(x0, reinterpret_cast<uint64_t>(ThreadSharedData.SignalHandlerRefCounterPtr));
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.SignalHandlerRefCountPointer)));
ldr(w2, MemOperand(x0));
sub(w2, w2, 1);
str(w2, MemOperand(x0));
@@ -76,7 +73,7 @@ DEF_OP(ExitFunction) {
uint64_t NewRIP;
if (IsInlineConstant(Op->NewRIP, &NewRIP) || IsInlineEntrypointOffset(Op->NewRIP, &NewRIP)) {
Literal l_BranchHost{ThreadSharedData.Dispatcher->ExitFunctionLinkerAddress};
Literal l_BranchHost{Dispatcher->ExitFunctionLinkerAddress};
Literal l_BranchGuest{NewRIP};
ldr(x0, &l_BranchHost);
@@ -88,7 +85,7 @@ DEF_OP(ExitFunction) {
RipReg = GetReg<RA_64>(Op->Header.Args[0].ID());
// L1 Cache
LoadConstant(x0, ThreadState->LookupCache->GetL1Pointer());
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.L1Pointer)));
and_(x3, RipReg, LookupCache::L1_ENTRIES_MASK);
add(x0, x0, Operand(x3, Shift::LSL, 4));
@@ -99,7 +96,7 @@ DEF_OP(ExitFunction) {
br(x1);
bind(&FullLookup);
LoadConstant(TMP1, ThreadSharedData.Dispatcher->AbsoluteLoopTopAddress);
ldr(TMP1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.DispatcherLoopTop)));
str(RipReg, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.rip)));
br(TMP1);
}
@@ -184,8 +181,16 @@ DEF_OP(Syscall) {
// X1: ThreadState
// X2: Pointer to SyscallArguments
FEXCore::IR::SyscallFlags Flags = Op->Flags;
PushDynamicRegsAndLR();
SpillStaticRegs();
if ((Flags & FEXCore::IR::SyscallFlags::NOSYNCSTATEONENTRY) != FEXCore::IR::SyscallFlags::NOSYNCSTATEONENTRY) {
SpillStaticRegs();
}
else {
// Need to spill all caller saved registers still
SpillStaticRegs(true, CALLER_GPR_MASK, CALLER_FPR_MASK);
}
uint64_t SPOffset = AlignUp(FEXCore::HLE::SyscallArguments::MAX_ARGS * 8, 16);
sub(sp, sp, SPOffset);
@@ -194,22 +199,30 @@ DEF_OP(Syscall) {
str(GetReg<RA_64>(Op->Header.Args[i].ID()), MemOperand(sp, i * 8));
}
LoadConstant(x0, reinterpret_cast<uint64_t>(CTX->SyscallHandler));
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.SyscallHandlerObj)));
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.SyscallHandlerFunc)));
mov(x1, STATE);
mov(x2, sp);
LoadConstant(x3, reinterpret_cast<uint64_t>(FEXCore::Context::HandleSyscall));
blr(x3);
add(sp, sp, SPOffset);
// Result is now in x0
// Fix the stack and any values that were stepped on
FillStaticRegs();
if ((Flags & FEXCore::IR::SyscallFlags::NOSYNCSTATEONENTRY) != FEXCore::IR::SyscallFlags::NOSYNCSTATEONENTRY &&
(Flags & FEXCore::IR::SyscallFlags::NORETURN) != FEXCore::IR::SyscallFlags::NORETURN) {
FillStaticRegs();
}
else {
// Result is now in x0
// Fix the stack and any values that were stepped on
FillStaticRegs(true, CALLER_GPR_MASK, CALLER_FPR_MASK);
}
PopDynamicRegsAndLR();
// Move result to its destination register
mov(GetReg<RA_64>(Node), x0);
if ((Flags & FEXCore::IR::SyscallFlags::NORETURN) != FEXCore::IR::SyscallFlags::NORETURN) {
// Move result to its destination register
mov(GetReg<RA_64>(Node), x0);
}
}
DEF_OP(InlineSyscall) {
@@ -340,21 +353,23 @@ DEF_OP(InlineSyscall) {
svc(0);
// On updated signal mask we can receive a signal RIGHT HERE
// Now that we are done in the syscall we need to carefully peel back the state
// First unspill the registers from before
FillStaticRegs(false, SpillMask);
if ((Op->Flags & FEXCore::IR::SyscallFlags::NORETURN) != FEXCore::IR::SyscallFlags::NORETURN) {
// Now that we are done in the syscall we need to carefully peel back the state
// First unspill the registers from before
FillStaticRegs(false, SpillMask);
// Now the registers we've spilled are back in their original host registers
// We can safely claim we are no longer in a syscall
str(xzr, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, InSyscallInfo)));
// Now the registers we've spilled are back in their original host registers
// We can safely claim we are no longer in a syscall
str(xzr, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, InSyscallInfo)));
// Result is now in x0
// Move result to its destination register
if (CTX->Config.Is64BitMode()) {
mov(GetReg<RA_64>(Node), x0);
}
else {
uxtw(GetReg<RA_64>(Node), x0);
// Result is now in x0
// Move result to its destination register
if (CTX->Config.Is64BitMode()) {
mov(GetReg<RA_64>(Node), x0);
}
else {
uxtw(GetReg<RA_64>(Node), x0);
}
}
}
@@ -427,7 +442,7 @@ DEF_OP(ValidateCode) {
}
}
DEF_OP(RemoveCodeEntry) {
DEF_OP(RemoveThreadCodeEntry) {
// Arguments are passed as follows:
// X0: Thread
// X1: RIP
@@ -437,7 +452,7 @@ DEF_OP(RemoveCodeEntry) {
mov(x0, STATE);
LoadConstant(x1, Entry);
LoadConstant(x2, reinterpret_cast<uintptr_t>(&Context::Context::RemoveCodeEntryFromJit));
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.RemoveThreadCodeEntryFromJIT)));
SpillStaticRegs();
blr(x2);
FillStaticRegs();
@@ -454,19 +469,10 @@ DEF_OP(CPUID) {
// x0 = CPUID Handler
// x1 = CPUID Function
// x2 = CPUID Leaf
LoadConstant(x0, reinterpret_cast<uint64_t>(&CTX->CPUID));
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.CPUIDObj)));
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.CPUIDFunction)));
mov(x1, GetReg<RA_64>(Op->Header.Args[0].ID()));
mov(x2, GetReg<RA_64>(Op->Header.Args[1].ID()));
using ClassPtrType = FEXCore::CPUID::FunctionResults (FEXCore::CPUIDEmu::*)(uint32_t, uint32_t);
union PtrCast {
ClassPtrType ClassPtr;
uintptr_t Data;
};
PtrCast Ptr;
Ptr.ClassPtr = &FEXCore::CPUIDEmu::RunFunction;
LoadConstant(x3, Ptr.Data);
SpillStaticRegs();
blr(x3);
FillStaticRegs();
@@ -485,7 +491,6 @@ void Arm64JITCore::RegisterBranchHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &Arm64JITCore::Op_##x
REGISTER_OP(GUESTCALLDIRECT, GuestCallDirect);
REGISTER_OP(GUESTCALLINDIRECT, GuestCallIndirect);
REGISTER_OP(GUESTRETURN, GuestReturn);
REGISTER_OP(SIGNALRETURN, SignalReturn);
REGISTER_OP(CALLBACKRETURN, CallbackReturn);
REGISTER_OP(EXITFUNCTION, ExitFunction);
@@ -495,7 +500,7 @@ void Arm64JITCore::RegisterBranchHandlers() {
REGISTER_OP(INLINESYSCALL, InlineSyscall);
REGISTER_OP(THUNK, Thunk);
REGISTER_OP(VALIDATECODE, ValidateCode);
REGISTER_OP(REMOVECODEENTRY, RemoveCodeEntry);
REGISTER_OP(REMOVETHREADCODEENTRY, RemoveThreadCodeEntry);
REGISTER_OP(CPUID, CPUID);
#undef REGISTER_OP
}
@@ -16,19 +16,19 @@ DEF_OP(VInsGPR) {
mov(GetDst(Node), GetSrc(Op->Header.Args[0].ID()));
switch (Op->Header.ElementSize) {
case 1: {
ins(GetDst(Node).V16B(), Op->Index, GetReg<RA_32>(Op->Header.Args[1].ID()));
ins(GetDst(Node).V16B(), Op->DestIdx, GetReg<RA_32>(Op->Header.Args[1].ID()));
break;
}
case 2: {
ins(GetDst(Node).V8H(), Op->Index, GetReg<RA_32>(Op->Header.Args[1].ID()));
ins(GetDst(Node).V8H(), Op->DestIdx, GetReg<RA_32>(Op->Header.Args[1].ID()));
break;
}
case 4: {
ins(GetDst(Node).V4S(), Op->Index, GetReg<RA_32>(Op->Header.Args[1].ID()));
ins(GetDst(Node).V4S(), Op->DestIdx, GetReg<RA_32>(Op->Header.Args[1].ID()));
break;
}
case 8: {
ins(GetDst(Node).V2D(), Op->Index, GetReg<RA_64>(Op->Header.Args[1].ID()));
ins(GetDst(Node).V2D(), Op->DestIdx, GetReg<RA_64>(Op->Header.Args[1].ID()));
break;
}
default: LOGMAN_MSG_A_FMT("Unknown Element Size: {}", Op->Header.ElementSize); break;
@@ -54,7 +54,7 @@ DEF_OP(AESDecLast) {
DEF_OP(AESKeyGenAssist) {
auto Op = IROp->C<IR::IROp_VAESKeyGenAssist>();
aarch64::Label Constant;
aarch64::Literal ConstantLiteral (0x0C030609'0306090CULL, 0x040B0E01'0B0E0104ULL);
aarch64::Label PastConstant;
// Do a "regular" AESE step
@@ -63,8 +63,7 @@ DEF_OP(AESKeyGenAssist) {
aese(VTMP1.V16B(), VTMP2.V16B());
// Do a table shuffle to undo ShiftRows
adr(TMP1.X(), &Constant);
ldr(VTMP3, MemOperand(TMP1.X()));
ldr(VTMP3, &ConstantLiteral);
// Now EOR in the RCON
if (Op->RCON) {
@@ -80,14 +79,29 @@ DEF_OP(AESKeyGenAssist) {
}
b(&PastConstant);
bind(&Constant);
dc32(0x0B0E0104);
dc32(0x040B0E01);
dc32(0x0306090C);
dc32(0x0C030609);
place(&ConstantLiteral);
bind(&PastConstant);
}
DEF_OP(CRC32) {
auto Op = IROp->C<IR::IROp_CRC32>();
switch (Op->SrcSize) {
case 1:
crc32cb(GetReg<RA_32>(Node), GetReg<RA_32>(Op->Src1.ID()), GetReg<RA_32>(Op->Src2.ID()));
break;
case 2:
crc32ch(GetReg<RA_32>(Node), GetReg<RA_32>(Op->Src1.ID()), GetReg<RA_32>(Op->Src2.ID()));
break;
case 4:
crc32cw(GetReg<RA_32>(Node), GetReg<RA_32>(Op->Src1.ID()), GetReg<RA_32>(Op->Src2.ID()));
break;
case 8:
crc32cx(GetReg<RA_32>(Node), GetReg<RA_32>(Op->Src1.ID()), GetReg<RA_64>(Op->Src2.ID()));
break;
default: LOGMAN_MSG_A_FMT("Unknown CRC32 size: {}", Op->SrcSize);
}
}
#undef DEF_OP
void Arm64JITCore::RegisterEncryptionHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &Arm64JITCore::Op_##x
@@ -97,7 +111,7 @@ void Arm64JITCore::RegisterEncryptionHandlers() {
REGISTER_OP(VAESDEC, AESDec);
REGISTER_OP(VAESDECLAST, AESDecLast);
REGISTER_OP(VAESKEYGENASSIST, AESKeyGenAssist);
REGISTER_OP(CRC32, CRC32);
#undef REGISTER_OP
}
}
+192 -98
View File
@@ -21,10 +21,13 @@ $end_info$
#include "Interface/IR/Passes/RegisterAllocationPass.h"
#include "Utils/MemberFunctionToPointer.h"
#include <FEXCore/Core/X86Enums.h>
#include <FEXCore/Core/UContext.h>
#include <FEXCore/Utils/Allocator.h>
#include <FEXCore/Utils/CompilerDefs.h>
#include "Interface/Core/Interpreter/InterpreterOps.h"
#include <sys/mman.h>
@@ -32,13 +35,42 @@ $end_info$
#include <unistd.h>
#include <string.h>
namespace FEXCore::CPU {
void Arm64JITCore::CopyNecessaryDataForCompileThread(CPUBackend *Original) {
Arm64JITCore *Core = reinterpret_cast<Arm64JITCore*>(Original);
ThreadSharedData = Core->ThreadSharedData;
namespace {
static uint64_t LUDIV(uint64_t SrcHigh, uint64_t SrcLow, uint64_t Divisor) {
__uint128_t Source = (static_cast<__uint128_t>(SrcHigh) << 64) | SrcLow;
__uint128_t Res = Source / Divisor;
return Res;
}
static int64_t LDIV(int64_t SrcHigh, int64_t SrcLow, int64_t Divisor) {
__int128_t Source = (static_cast<__int128_t>(SrcHigh) << 64) | SrcLow;
__int128_t Res = Source / Divisor;
return Res;
}
static uint64_t LUREM(uint64_t SrcHigh, uint64_t SrcLow, uint64_t Divisor) {
__uint128_t Source = (static_cast<__uint128_t>(SrcHigh) << 64) | SrcLow;
__uint128_t Res = Source % Divisor;
return Res;
}
static int64_t LREM(int64_t SrcHigh, int64_t SrcLow, int64_t Divisor) {
__int128_t Source = (static_cast<__int128_t>(SrcHigh) << 64) | SrcLow;
__int128_t Res = Source % Divisor;
return Res;
}
static void PrintValue(uint64_t Value) {
LogMan::Msg::DFmt("Value: 0x{:x}", Value);
}
static void PrintVectorValue(uint64_t Value, uint64_t ValueUpper) {
LogMan::Msg::DFmt("Value: 0x{:016x}'{:016x}", ValueUpper, Value);
}
}
namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
@@ -56,8 +88,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
PushDynamicRegsAndLR();
uxth(w0, GetReg<RA_32>(IROp->Args[0].ID()));
LoadConstant(x1, (uintptr_t)Info.fn);
ldr(x1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.FallbackHandlerPointers[Info.HandlerIndex])));
blr(x1);
PopDynamicRegsAndLR();
@@ -72,8 +103,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
PushDynamicRegsAndLR();
fmov(v0.S(), GetSrc(IROp->Args[0].ID()).S()) ;
LoadConstant(x0, (uintptr_t)Info.fn);
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.FallbackHandlerPointers[Info.HandlerIndex])));
blr(x0);
PopDynamicRegsAndLR();
@@ -92,8 +122,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
PushDynamicRegsAndLR();
mov(v0.D(), GetSrc(IROp->Args[0].ID()).D());
LoadConstant(x0, (uintptr_t)Info.fn);
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.FallbackHandlerPointers[Info.HandlerIndex])));
blr(x0);
PopDynamicRegsAndLR();
@@ -118,8 +147,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
else {
mov(w0, GetReg<RA_32>(IROp->Args[0].ID()));
}
LoadConstant(x1, (uintptr_t)Info.fn);
ldr(x1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.FallbackHandlerPointers[Info.HandlerIndex])));
blr(x1);
PopDynamicRegsAndLR();
@@ -140,8 +168,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
umov(x0, GetSrc(IROp->Args[0].ID()).V2D(), 0);
umov(w1, GetSrc(IROp->Args[0].ID()).V8H(), 4);
LoadConstant(x2, (uintptr_t)Info.fn);
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.FallbackHandlerPointers[Info.HandlerIndex])));
blr(x2);
PopDynamicRegsAndLR();
@@ -160,8 +187,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
umov(x0, GetSrc(IROp->Args[0].ID()).V2D(), 0);
umov(w1, GetSrc(IROp->Args[0].ID()).V8H(), 4);
LoadConstant(x2, (uintptr_t)Info.fn);
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.FallbackHandlerPointers[Info.HandlerIndex])));
blr(x2);
PopDynamicRegsAndLR();
@@ -172,6 +198,43 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
}
break;
case FABI_F64_F64: {
SpillStaticRegs();
PushDynamicRegsAndLR();
mov(v0.D(), GetSrc(IROp->Args[0].ID()).D());
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.FallbackHandlerPointers[Info.HandlerIndex])));
blr(x0);
PopDynamicRegsAndLR();
FillStaticRegs();
mov(GetDst(Node).D(), v0.D());
}
break;
case FABI_F64_F64_F64: {
SpillStaticRegs();
PushDynamicRegsAndLR();
mov(v0.D(), GetSrc(IROp->Args[0].ID()).D());
mov(v1.D(), GetSrc(IROp->Args[1].ID()).D());
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.FallbackHandlerPointers[Info.HandlerIndex])));
blr(x0);
PopDynamicRegsAndLR();
FillStaticRegs();
mov(GetDst(Node).D(), v0.D());
}
break;
case FABI_I16_F80:{
SpillStaticRegs();
@@ -180,8 +243,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
umov(x0, GetSrc(IROp->Args[0].ID()).V2D(), 0);
umov(w1, GetSrc(IROp->Args[0].ID()).V8H(), 4);
LoadConstant(x2, (uintptr_t)Info.fn);
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.FallbackHandlerPointers[Info.HandlerIndex])));
blr(x2);
PopDynamicRegsAndLR();
@@ -199,8 +261,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
umov(x0, GetSrc(IROp->Args[0].ID()).V2D(), 0);
umov(w1, GetSrc(IROp->Args[0].ID()).V8H(), 4);
LoadConstant(x2, (uintptr_t)Info.fn);
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.FallbackHandlerPointers[Info.HandlerIndex])));
blr(x2);
PopDynamicRegsAndLR();
@@ -218,8 +279,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
umov(x0, GetSrc(IROp->Args[0].ID()).V2D(), 0);
umov(w1, GetSrc(IROp->Args[0].ID()).V8H(), 4);
LoadConstant(x2, (uintptr_t)Info.fn);
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.FallbackHandlerPointers[Info.HandlerIndex])));
blr(x2);
PopDynamicRegsAndLR();
@@ -240,8 +300,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
umov(x2, GetSrc(IROp->Args[1].ID()).V2D(), 0);
umov(w3, GetSrc(IROp->Args[1].ID()).V8H(), 4);
LoadConstant(x4, (uintptr_t)Info.fn);
ldr(x4, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.FallbackHandlerPointers[Info.HandlerIndex])));
blr(x4);
PopDynamicRegsAndLR();
@@ -259,8 +318,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
umov(x0, GetSrc(IROp->Args[0].ID()).V2D(), 0);
umov(w1, GetSrc(IROp->Args[0].ID()).V8H(), 4);
LoadConstant(x2, (uintptr_t)Info.fn);
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.FallbackHandlerPointers[Info.HandlerIndex])));
blr(x2);
PopDynamicRegsAndLR();
@@ -283,8 +341,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
umov(x2, GetSrc(IROp->Args[1].ID()).V2D(), 0);
umov(w3, GetSrc(IROp->Args[1].ID()).V8H(), 4);
LoadConstant(x4, (uintptr_t)Info.fn);
ldr(x4, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.FallbackHandlerPointers[Info.HandlerIndex])));
blr(x4);
PopDynamicRegsAndLR();
@@ -332,28 +389,10 @@ void Arm64JITCore::FreeCodeBuffer(CodeBuffer Buffer) {
Dispatcher->RemoveCodeBuffer(Buffer.Ptr);
}
Arm64JITCore::Arm64JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, bool CompileThread)
Arm64JITCore::Arm64JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread)
: Arm64Emitter(ctx, 0)
, CTX {ctx}
, ThreadState {Thread} {
{
DispatcherConfig config;
config.ExitFunctionLink = reinterpret_cast<uintptr_t>(&ExitFunctionLink);
config.ExitFunctionLinkThis = reinterpret_cast<uintptr_t>(this);
config.StaticRegisterAssignment = ctx->Config.StaticRegisterAllocation;
Dispatcher = std::make_unique<Arm64Dispatcher>(CTX, ThreadState, config);
DispatchPtr = Dispatcher->DispatchPtr;
CallbackPtr = Dispatcher->CallbackPtr;
}
// Can't allocate a code buffer until after dispatcher is created
InitialCodeBuffer = AllocateNewCodeBuffer(Arm64JITCore::INITIAL_CODE_SIZE);
*GetBuffer() = vixl::CodeBuffer(InitialCodeBuffer.Ptr, InitialCodeBuffer.Size);
SetAllowAssembler(true);
CurrentCodeBuffer = &InitialCodeBuffer;
RAPass = Thread->PassManager->GetPass<IR::RegisterAllocationPass>("RA");
#if DEBUG
@@ -393,50 +432,97 @@ Arm64JITCore::Arm64JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::Intern
RegisterVectorHandlers();
RegisterEncryptionHandlers();
if (!CompileThread) {
ThreadSharedData.SignalHandlerRefCounterPtr = &Dispatcher->SignalHandlerRefCounter;
ThreadSharedData.SignalReturnInstruction = Dispatcher->SignalHandlerReturnAddress;
ThreadSharedData.UnimplementedInstructionAddress = Dispatcher->UnimplementedInstructionAddress;
{
DispatcherConfig config;
config.ExitFunctionLink = reinterpret_cast<uintptr_t>(&ExitFunctionLink);
config.ExitFunctionLinkThis = reinterpret_cast<uintptr_t>(this);
config.StaticRegisterAssignment = ctx->Config.StaticRegisterAllocation;
ThreadSharedData.Dispatcher = Dispatcher.get();
// This will register the host signal handler per thread, which is fine
CTX->SignalDelegation->RegisterHostSignalHandler(SIGILL, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
Arm64JITCore *Core = reinterpret_cast<Arm64JITCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleSIGILL(Signal, info, ucontext);
}, true);
CTX->SignalDelegation->RegisterHostSignalHandler(SIGBUS, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
Arm64JITCore *Core = reinterpret_cast<Arm64JITCore*>(Thread->CPUBackend.get());
if (!Core->Dispatcher->IsAddressInJITCode(ArchHelpers::Context::GetPc(ucontext))) {
// Wasn't a sigbus in JIT code
return false;
}
return FEXCore::ArchHelpers::Arm64::HandleSIGBUS(Core->CTX->Config.ParanoidTSO(), Signal, info, ucontext);
}, true);
CTX->SignalDelegation->RegisterHostSignalHandler(SignalDelegator::SIGNAL_FOR_PAUSE, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
Arm64JITCore *Core = reinterpret_cast<Arm64JITCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleSignalPause(Signal, info, ucontext);
}, true);
auto GuestSignalHandler = [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack) -> bool {
Arm64JITCore *Core = reinterpret_cast<Arm64JITCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleGuestSignal(Signal, info, ucontext, GuestAction, GuestStack);
};
for (uint32_t Signal = 0; Signal <= SignalDelegator::MAX_SIGNALS; ++Signal) {
CTX->SignalDelegation->RegisterHostSignalHandlerForGuest(Signal, GuestSignalHandler);
}
Dispatcher = std::make_unique<Arm64Dispatcher>(CTX, ThreadState, config);
DispatchPtr = Dispatcher->DispatchPtr;
CallbackPtr = Dispatcher->CallbackPtr;
}
{
// Set up pointers that the JIT needs to load
auto &Pointers = ThreadState->CurrentFrame->Pointers.AArch64;
// Process specific
Pointers.LUDIV = reinterpret_cast<uint64_t>(LUDIV);
Pointers.LDIV = reinterpret_cast<uint64_t>(LDIV);
Pointers.LUREM = reinterpret_cast<uint64_t>(LUREM);
Pointers.LREM = reinterpret_cast<uint64_t>(LREM);
Pointers.PrintValue = reinterpret_cast<uint64_t>(PrintValue);
Pointers.PrintVectorValue = reinterpret_cast<uint64_t>(PrintVectorValue);
Pointers.RemoveThreadCodeEntryFromJIT = reinterpret_cast<uintptr_t>(&Context::Context::RemoveThreadCodeEntryFromJit);
Pointers.CPUIDObj = reinterpret_cast<uint64_t>(&CTX->CPUID);
{
FEXCore::Utils::MemberFunctionToPointerCast PMF(&FEXCore::CPUIDEmu::RunFunction);
Pointers.CPUIDFunction = PMF.GetConvertedPointer();
}
Pointers.SyscallHandlerObj = reinterpret_cast<uint64_t>(CTX->SyscallHandler);
Pointers.SyscallHandlerFunc = reinterpret_cast<uint64_t>(FEXCore::Context::HandleSyscall);
// Fill in the fallback handlers
InterpreterOps::FillFallbackIndexPointers(Pointers.FallbackHandlerPointers);
// Thread Specific
Pointers.SignalHandlerRefCountPointer = reinterpret_cast<uint64_t>(&Dispatcher->SignalHandlerRefCounter);
}
// Can't allocate a code buffer until after dispatcher is created
InitialCodeBuffer = AllocateNewCodeBuffer(Arm64JITCore::INITIAL_CODE_SIZE);
*GetBuffer() = vixl::CodeBuffer(InitialCodeBuffer.Ptr, InitialCodeBuffer.Size);
SetAllowAssembler(true);
EmitDetectionString();
CurrentCodeBuffer = &InitialCodeBuffer;
}
void Arm64JITCore::InitializeSignalHandlers(FEXCore::Context::Context *CTX) {
CTX->SignalDelegation->RegisterHostSignalHandler(SIGILL, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
Arm64JITCore *Core = reinterpret_cast<Arm64JITCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleSIGILL(Signal, info, ucontext);
}, true);
CTX->SignalDelegation->RegisterHostSignalHandler(SIGBUS, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
Arm64JITCore *Core = reinterpret_cast<Arm64JITCore*>(Thread->CPUBackend.get());
if (!Core->Dispatcher->IsAddressInJITCode(ArchHelpers::Context::GetPc(ucontext))) {
// Wasn't a sigbus in JIT code
return false;
}
return FEXCore::ArchHelpers::Arm64::HandleSIGBUS(Core->CTX->Config.ParanoidTSO(), Signal, info, ucontext);
}, true);
CTX->SignalDelegation->RegisterHostSignalHandler(SignalDelegator::SIGNAL_FOR_PAUSE, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
Arm64JITCore *Core = reinterpret_cast<Arm64JITCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleSignalPause(Signal, info, ucontext);
}, true);
auto GuestSignalHandler = [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack) -> bool {
Arm64JITCore *Core = reinterpret_cast<Arm64JITCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleGuestSignal(Signal, info, ucontext, GuestAction, GuestStack);
};
for (uint32_t Signal = 0; Signal <= SignalDelegator::MAX_SIGNALS; ++Signal) {
CTX->SignalDelegation->RegisterHostSignalHandlerForGuest(Signal, GuestSignalHandler);
}
}
void Arm64JITCore::EmitDetectionString() {
const char JITString[] = "FEXJIT::Arm64JITCore::";
auto Buffer = GetBuffer();
Buffer->EmitString(JITString);
Buffer->Align();
}
void Arm64JITCore::ClearCache() {
// Get the backing code buffer
auto Buffer = GetBuffer();
if (*ThreadSharedData.SignalHandlerRefCounterPtr == 0) {
if (Dispatcher->SignalHandlerRefCounter == 0) {
if (!CodeBuffers.empty()) {
// If we have more than one code buffer we are tracking then walk them and delete
// This is a cleanup step
@@ -473,6 +559,7 @@ void Arm64JITCore::ClearCache() {
EmplaceNewCodeBuffer(NewCodeBuffer);
*Buffer = vixl::CodeBuffer(NewCodeBuffer.Ptr, NewCodeBuffer.Size);
}
EmitDetectionString();
}
Arm64JITCore::~Arm64JITCore() {
@@ -582,7 +669,12 @@ bool Arm64JITCore::IsInlineEntrypointOffset(const IR::OrderedNodeWrapper& WNode,
if (OpHeader->Op == IR::IROps::OP_INLINEENTRYPOINTOFFSET) {
auto Op = OpHeader->C<IR::IROp_InlineEntrypointOffset>();
if (Value) {
*Value = Entry + Op->Offset;
uint64_t Mask = ~0ULL;
uint8_t OpSize = OpHeader->Size;
if (OpSize == 4) {
Mask = 0xFFFF'FFFFULL;
}
*Value = (Entry + Op->Offset) & Mask;
}
return true;
} else {
@@ -647,9 +739,9 @@ void *Arm64JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IR
// X1-X3 = Temp
// X4-r18 = RA
auto GuestEntry = GetCursorAddress<uint64_t>();
GuestEntry = GetCursorAddress<uint64_t>();
if (CTX->GetGdbServerStatus()) {
if (CTX->GetGdbServerStatus()) {
aarch64::Label RunBlock;
// If we have a gdb server running then run in a less efficient mode that checks if we need to exit
@@ -668,7 +760,7 @@ void *Arm64JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IR
str(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.rip)));
// Stop the thread
LoadConstant(x0, ThreadSharedData.Dispatcher->ThreadPauseHandlerAddressSpillSRA);
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.ThreadPauseHandlerSpillSRA)));
br(x0);
}
bind(&RunBlock);
@@ -696,6 +788,7 @@ void *Arm64JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IR
LOGMAN_THROW_A_FMT(BlockIROp->Header.Op == IR::OP_CODEBLOCK, "IR type failed to be a code block");
#endif
uintptr_t BlockStartHostCode = GetCursorAddress<uintptr_t>();
{
const auto Node = IR->GetID(BlockNode);
const auto IsTarget = JumpTargets.try_emplace(Node).first;
@@ -710,10 +803,6 @@ void *Arm64JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IR
bind(&IsTarget->second);
}
if (DebugData) {
DebugData->Subblocks.push_back({GetCursorAddress<uintptr_t>(), 0, IR->GetID(BlockNode)});
}
for (auto [CodeNode, IROp] : IR->GetCode(BlockNode)) {
const auto ID = IR->GetID(CodeNode);
@@ -723,7 +812,7 @@ void *Arm64JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IR
}
if (DebugData) {
DebugData->Subblocks.back().HostCodeSize = GetCursorAddress<uintptr_t>() - DebugData->Subblocks.back().HostCodeStart;
DebugData->Subblocks.push_back({BlockStartHostCode, static_cast<uint32_t>(GetCursorAddress<uintptr_t>() - BlockStartHostCode)});
}
}
@@ -741,6 +830,7 @@ void *Arm64JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IR
if (DebugData) {
DebugData->HostCodeSize = reinterpret_cast<uintptr_t>(CodeEnd) - reinterpret_cast<uintptr_t>(GuestEntry);
DebugData->Relocations = &Relocations;
}
this->IR = nullptr;
@@ -757,11 +847,11 @@ uint64_t Arm64JITCore::ExitFunctionLink(Arm64JITCore *core, FEXCore::Core::CpuSt
if (!HostCode) {
//fmt::print("ExitFunctionLink: Aborting, {:X} not in cache\n", GuestRip);
Frame->State.rip = GuestRip;
return core->ThreadSharedData.Dispatcher->AbsoluteLoopTopAddress;
return core->Dispatcher->AbsoluteLoopTopAddress;
}
uintptr_t branch = (uintptr_t)(record) - 8;
auto LinkerAddress = core->ThreadSharedData.Dispatcher->ExitFunctionLinkerAddress;
auto LinkerAddress = core->Dispatcher->ExitFunctionLinkerAddress;
auto offset = HostCode/4 - branch/4;
if (IsInt26(offset)) {
@@ -797,7 +887,11 @@ uint64_t Arm64JITCore::ExitFunctionLink(Arm64JITCore *core, FEXCore::Core::CpuSt
return HostCode;
}
std::unique_ptr<CPUBackend> CreateArm64JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, bool CompileThread) {
return std::make_unique<Arm64JITCore>(ctx, Thread, CompileThread);
std::unique_ptr<CPUBackend> CreateArm64JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread) {
return std::make_unique<Arm64JITCore>(ctx, Thread);
}
void InitializeArm64JITSignalHandlers(FEXCore::Context::Context *CTX) {
Arm64JITCore::InitializeSignalHandlers(CTX);
}
}
+15 -20
View File
@@ -43,8 +43,7 @@ public:
};
explicit Arm64JITCore(FEXCore::Context::Context *ctx,
FEXCore::Core::InternalThreadState *Thread,
bool CompileThread);
FEXCore::Core::InternalThreadState *Thread);
~Arm64JITCore() override;
[[nodiscard]] std::string GetName() override { return "JIT"; }
@@ -63,11 +62,14 @@ public:
static constexpr size_t INITIAL_CODE_SIZE = 1024 * 1024 * 16;
[[nodiscard]] CodeBuffer AllocateNewCodeBuffer(size_t Size);
void CopyNecessaryDataForCompileThread(CPUBackend *Original) override;
bool IsAddressInJITCode(uint64_t Address, bool IncludeDispatcher = true, bool IncludeCompileService = true) const override {
return Dispatcher->IsAddressInJITCode(Address, IncludeDispatcher, IncludeCompileService);
bool IsAddressInJITCode(uint64_t Address, bool IncludeDispatcher = true) const override {
return Dispatcher->IsAddressInJITCode(Address, IncludeDispatcher);
}
static void InitializeSignalHandlers(FEXCore::Context::Context *CTX);
void ClearRelocations() override { Relocations.clear(); }
private:
FEX_CONFIG_OPT(ParanoidTSO, PARANOIDTSO);
@@ -173,16 +175,8 @@ private:
static uint64_t ExitFunctionLink(Arm64JITCore *core, FEXCore::Core::CpuStateFrame *Frame, uint64_t *record);
struct CompilerSharedData {
uint64_t SignalReturnInstruction{};
uint64_t UnimplementedInstructionAddress{};
uint64_t OverflowExceptionInstructionAddress{};
uint32_t *SignalHandlerRefCounterPtr{};
FEXCore::CPU::Dispatcher *Dispatcher{};
};
CompilerSharedData ThreadSharedData;
// This is purely a debugging aid for developers to see if they are in JIT code space when inspecting raw memory
void EmitDetectionString();
IR::RegisterAllocationPass *RAPass;
IR::RegisterAllocationData *RAData;
@@ -277,7 +271,6 @@ private:
///< Branch ops
DEF_OP(GuestCallDirect);
DEF_OP(GuestCallIndirect);
DEF_OP(GuestReturn);
DEF_OP(SignalReturn);
DEF_OP(CallbackReturn);
DEF_OP(ExitFunction);
@@ -287,7 +280,7 @@ private:
DEF_OP(InlineSyscall);
DEF_OP(Thunk);
DEF_OP(ValidateCode);
DEF_OP(RemoveCodeEntry);
DEF_OP(RemoveThreadCodeEntry);
DEF_OP(CPUID);
///< Conversion ops
@@ -305,7 +298,7 @@ private:
DEF_OP(GetHostFlag);
///< Memory ops
DEF_OP(LoadContext);
DEF_OP(LoadContext);
DEF_OP(StoreContext);
DEF_OP(LoadRegister);
DEF_OP(StoreRegister);
@@ -336,6 +329,8 @@ private:
DEF_OP(GetRoundingMode);
DEF_OP(SetRoundingMode);
DEF_OP(ProcessorID);
DEF_OP(RDRAND);
DEF_OP(Yield);
///< Move ops
DEF_OP(ExtractElementPair);
@@ -345,8 +340,6 @@ private:
///< Vector ops
DEF_OP(VectorZero);
DEF_OP(VectorImm);
DEF_OP(CreateVector2);
DEF_OP(CreateVector4);
DEF_OP(SplatVector2);
DEF_OP(SplatVector4);
DEF_OP(VMov);
@@ -434,6 +427,7 @@ private:
DEF_OP(VSMull2);
DEF_OP(VUABDL);
DEF_OP(VTBL1);
DEF_OP(VRev64);
///< Encryption ops
DEF_OP(AESImc);
@@ -442,6 +436,7 @@ private:
DEF_OP(AESDec);
DEF_OP(AESDecLast);
DEF_OP(AESKeyGenAssist);
DEF_OP(CRC32);
#undef DEF_OP
};
@@ -617,13 +617,43 @@ DEF_OP(LoadMem) {
DEF_OP(LoadMemTSO) {
auto Op = IROp->C<IR::IROp_LoadMemTSO>();
auto MemSrc = MemOperand(GetReg<RA_64>(Op->Header.Args[0].ID()));
auto MemReg = GetReg<RA_64>(Op->Addr.ID());
auto MemSrc = GenerateMemOperand(IROp->Size, MemReg, Op->Offset, Op->OffsetType, Op->OffsetScale);
if (!Op->Offset.IsInvalid()) {
LOGMAN_MSG_A_FMT("LoadMemTSO: No offset allowed");
if (CTX->HostFeatures.SupportsTSOImm9) {
// RCPC2 means that the offset must be an inline constant
LOGMAN_THROW_A_FMT(MemSrc.IsRegisterOffset() == false, "RCPC2 doesn't support register offset. Only Immediate offset");
}
else {
LOGMAN_THROW_A_FMT(Op->Offset.IsInvalid(), "LoadMemTSO: No offset allowed");
}
if (CTX->HostFeatures.SupportsRCPC && Op->Class == FEXCore::IR::GPRClass) {
if (CTX->HostFeatures.SupportsTSOImm9 && Op->Class == FEXCore::IR::GPRClass) {
if (IROp->Size == 1) {
// 8bit load is always aligned to natural alignment
auto Dst = GetReg<RA_64>(Node);
ldapurb(Dst, MemSrc);
}
else {
// Aligned
nop();
auto Dst = GetReg<RA_64>(Node);
switch (IROp->Size) {
case 2:
ldapurh(Dst, MemSrc);
break;
case 4:
ldapur(Dst.W(), MemSrc);
break;
case 8:
ldapur(Dst, MemSrc);
break;
default: LOGMAN_MSG_A_FMT("Unhandled LoadMemTSO size: {}", IROp->Size);
}
nop();
}
}
else if (CTX->HostFeatures.SupportsRCPC && Op->Class == FEXCore::IR::GPRClass) {
if (IROp->Size == 1) {
// 8bit load is always aligned to natural alignment
auto Dst = GetReg<RA_64>(Node);
@@ -698,29 +728,29 @@ DEF_OP(LoadMemTSO) {
DEF_OP(StoreMem) {
auto Op = IROp->C<IR::IROp_StoreMem>();
auto MemReg = GetReg<RA_64>(Op->Header.Args[0].ID());
auto MemReg = GetReg<RA_64>(Op->Addr.ID());
auto MemSrc = GenerateMemOperand(IROp->Size, MemReg, Op->Offset, Op->OffsetType, Op->OffsetScale);
if (Op->Class == FEXCore::IR::GPRClass) {
switch (IROp->Size) {
case 1:
strb(GetReg<RA_64>(Op->Header.Args[1].ID()), MemSrc);
strb(GetReg<RA_64>(Op->Value.ID()), MemSrc);
break;
case 2:
strh(GetReg<RA_64>(Op->Header.Args[1].ID()), MemSrc);
strh(GetReg<RA_64>(Op->Value.ID()), MemSrc);
break;
case 4:
str(GetReg<RA_32>(Op->Header.Args[1].ID()), MemSrc);
str(GetReg<RA_32>(Op->Value.ID()), MemSrc);
break;
case 8:
str(GetReg<RA_64>(Op->Header.Args[1].ID()), MemSrc);
str(GetReg<RA_64>(Op->Value.ID()), MemSrc);
break;
default: LOGMAN_MSG_A_FMT("Unhandled StoreMem size: {}", IROp->Size);
}
}
else {
auto Src = GetSrc(Op->Header.Args[1].ID());
auto Src = GetSrc(Op->Value.ID());
switch (IROp->Size) {
case 1:
str(Src.B(), MemSrc);
@@ -744,28 +774,56 @@ DEF_OP(StoreMem) {
DEF_OP(StoreMemTSO) {
auto Op = IROp->C<IR::IROp_StoreMemTSO>();
auto MemSrc = MemOperand(GetReg<RA_64>(Op->Header.Args[0].ID()));
if (!Op->Offset.IsInvalid()) {
LOGMAN_MSG_A_FMT("StoreMemTSO: No offset allowed");
auto MemReg = GetReg<RA_64>(Op->Addr.ID());
auto MemSrc = GenerateMemOperand(IROp->Size, MemReg, Op->Offset, Op->OffsetType, Op->OffsetScale);
if (CTX->HostFeatures.SupportsTSOImm9) {
// RCPC2 means that the offset must be an inline constant
LOGMAN_THROW_A_FMT(MemSrc.IsRegisterOffset() == false, "RCPC2 doesn't support register offset. Only Immediate offset");
}
else {
LOGMAN_THROW_A_FMT(Op->Offset.IsInvalid(), "StoreMemTSO: No offset allowed");
}
if (Op->Class == FEXCore::IR::GPRClass) {
if (CTX->HostFeatures.SupportsTSOImm9 && Op->Class == FEXCore::IR::GPRClass) {
if (IROp->Size == 1) {
// 8bit load is always aligned to natural alignment
stlrb(GetReg<RA_64>(Op->Header.Args[1].ID()), MemSrc);
stlurb(GetReg<RA_64>(Op->Value.ID()), MemSrc);
}
else {
nop();
switch (IROp->Size) {
case 2:
stlrh(GetReg<RA_64>(Op->Header.Args[1].ID()), MemSrc);
stlurh(GetReg<RA_64>(Op->Value.ID()), MemSrc);
break;
case 4:
stlr(GetReg<RA_32>(Op->Header.Args[1].ID()), MemSrc);
stlur(GetReg<RA_32>(Op->Value.ID()), MemSrc);
break;
case 8:
stlr(GetReg<RA_64>(Op->Header.Args[1].ID()), MemSrc);
stlur(GetReg<RA_64>(Op->Value.ID()), MemSrc);
break;
default: LOGMAN_MSG_A_FMT("Unhandled StoreMemTSO size: {}", IROp->Size);
}
nop();
}
}
else if (Op->Class == FEXCore::IR::GPRClass) {
if (IROp->Size == 1) {
// 8bit load is always aligned to natural alignment
stlrb(GetReg<RA_64>(Op->Value.ID()), MemSrc);
}
else {
nop();
switch (IROp->Size) {
case 2:
stlrh(GetReg<RA_64>(Op->Value.ID()), MemSrc);
break;
case 4:
stlr(GetReg<RA_32>(Op->Value.ID()), MemSrc);
break;
case 8:
stlr(GetReg<RA_64>(Op->Value.ID()), MemSrc);
break;
default: LOGMAN_MSG_A_FMT("Unhandled StoreMemTSO size: {}", IROp->Size);
}
@@ -774,7 +832,7 @@ DEF_OP(StoreMemTSO) {
}
else {
dmb(InnerShareable, BarrierAll);
auto Src = GetSrc(Op->Header.Args[1].ID());
auto Src = GetSrc(Op->Value.ID());
switch (IROp->Size) {
case 1:
str(Src.B(), MemSrc);
@@ -800,7 +858,7 @@ DEF_OP(StoreMemTSO) {
DEF_OP(ParanoidLoadMemTSO) {
auto Op = IROp->C<IR::IROp_LoadMemTSO>();
auto MemSrc = MemOperand(GetReg<RA_64>(Op->Header.Args[0].ID()));
auto MemSrc = MemOperand(GetReg<RA_64>(Op->Addr.ID()));
if (!Op->Offset.IsInvalid()) {
LOGMAN_MSG_A_FMT("ParanoidLoadMemTSO: No offset allowed");
@@ -857,7 +915,7 @@ DEF_OP(ParanoidLoadMemTSO) {
DEF_OP(ParanoidStoreMemTSO) {
auto Op = IROp->C<IR::IROp_StoreMemTSO>();
auto MemSrc = MemOperand(GetReg<RA_64>(Op->Header.Args[0].ID()));
auto MemSrc = MemOperand(GetReg<RA_64>(Op->Addr.ID()));
if (!Op->Offset.IsInvalid()) {
LOGMAN_MSG_A_FMT("ParanoidStoreMemTSO: No offset allowed");
@@ -866,25 +924,25 @@ DEF_OP(ParanoidStoreMemTSO) {
if (Op->Class == FEXCore::IR::GPRClass) {
if (IROp->Size == 1) {
// 8bit load is always aligned to natural alignment
stlrb(GetReg<RA_64>(Op->Header.Args[1].ID()), MemSrc);
stlrb(GetReg<RA_64>(Op->Value.ID()), MemSrc);
}
else {
switch (IROp->Size) {
case 2:
stlrh(GetReg<RA_64>(Op->Header.Args[1].ID()), MemSrc);
stlrh(GetReg<RA_64>(Op->Value.ID()), MemSrc);
break;
case 4:
stlr(GetReg<RA_32>(Op->Header.Args[1].ID()), MemSrc);
stlr(GetReg<RA_32>(Op->Value.ID()), MemSrc);
break;
case 8:
stlr(GetReg<RA_64>(Op->Header.Args[1].ID()), MemSrc);
stlr(GetReg<RA_64>(Op->Value.ID()), MemSrc);
break;
default: LOGMAN_MSG_A_FMT("Unhandled ParanoidStoreMemTSO size: {}", IROp->Size);
}
}
}
else {
auto Src = GetSrc(Op->Header.Args[1].ID());
auto Src = GetSrc(Op->Value.ID());
if (IROp->Size == 1) {
// 8bit load is always aligned to natural alignment
mov(TMP1.W(), Src.V16B(), 0);
+31 -15
View File
@@ -7,14 +7,6 @@ $end_info$
#include "Interface/Core/JIT/Arm64/JITClass.h"
namespace FEXCore::CPU {
static void PrintValue(uint64_t Value) {
LogMan::Msg::DFmt("Value: 0x{:x}", Value);
}
static void PrintVectorValue(uint64_t Value, uint64_t ValueUpper) {
LogMan::Msg::DFmt("Value: 0x{:016x}'{:016x}", ValueUpper, Value);
}
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
@@ -44,7 +36,7 @@ DEF_OP(Break) {
break;
case FEXCore::IR::Break_Overflow: // overflow
ResetStack();
LoadConstant(TMP1, ThreadSharedData.Dispatcher->OverflowExceptionInstructionAddress);
ldr(TMP1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.OverflowExceptionHandler)));
br(TMP1);
break;
case FEXCore::IR::Break_Halt: { // HLT
@@ -54,14 +46,13 @@ DEF_OP(Break) {
add(sp, TMP1, 0);
// Now we need to jump to the thread stop handler
LoadConstant(TMP1, ThreadSharedData.Dispatcher->ThreadStopHandlerAddressSpillSRA);
ldr(TMP1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.ThreadStopHandlerSpillSRA)));
br(TMP1);
break;
}
case FEXCore::IR::Break_Interrupt3: { // INT3
ResetStack();
LoadConstant(TMP1, ThreadSharedData.Dispatcher->ThreadPauseHandlerAddressSpillSRA);
ldr(TMP1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.ThreadPauseHandlerSpillSRA)));
br(TMP1);
break;
}
@@ -69,7 +60,7 @@ DEF_OP(Break) {
{
ResetStack();
LoadConstant(TMP1, ThreadSharedData.Dispatcher->UnimplementedInstructionAddress);
ldr(TMP1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.UnimplementedInstructionHandler)));
br(TMP1);
break;
@@ -143,13 +134,13 @@ DEF_OP(Print) {
if (IsGPR(Op->Header.Args[0].ID())) {
mov(x0, GetReg<RA_64>(Op->Header.Args[0].ID()));
LoadConstant(x3, reinterpret_cast<uint64_t>(PrintValue));
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.PrintValue)));
}
else {
fmov(x0, GetSrc(Op->Header.Args[0].ID()).V1D());
// Bug in vixl that source vector needs to b V1D rather than V2D?
fmov(x1, GetSrc(Op->Header.Args[0].ID()).V1D(), 1);
LoadConstant(x3, reinterpret_cast<uint64_t>(PrintVectorValue));
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.PrintVectorValue)));
}
SpillStaticRegs();
blr(x3);
@@ -210,6 +201,28 @@ DEF_OP(ProcessorID) {
orr(GetReg<RA_64>(Node), x0, Operand(x1, LSL, 12));
}
DEF_OP(RDRAND) {
auto Op = IROp->C<IR::IROp_RDRAND>();
// Results are in x0, x1
// Results want to be in a i64v2 vector
auto Dst = GetSrcPair<RA_64>(Node);
if (Op->GetReseeded) {
mrs(Dst.first, RNDRRS);
}
else {
mrs(Dst.first, RNDR);
}
// If the rng number is valid then NZCV is 0b0000, otherwise NZCV is 0b0100
cset(Dst.second, Condition::ne);
}
DEF_OP(Yield) {
hint(SystemHint::YIELD);
}
#undef DEF_OP
void Arm64JITCore::RegisterMiscHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &Arm64JITCore::Op_##x
@@ -227,6 +240,9 @@ void Arm64JITCore::RegisterMiscHandlers() {
REGISTER_OP(SETROUNDINGMODE, SetRoundingMode);
REGISTER_OP(INVALIDATEFLAGS, NoOp);
REGISTER_OP(PROCESSORID, ProcessorID);
REGISTER_OP(RDRAND, RDRAND);
REGISTER_OP(YIELD, Yield);
#undef REGISTER_OP
}
}
@@ -37,7 +37,7 @@ DEF_OP(CreateElementPair) {
aarch64::Register RegSecond;
aarch64::Register RegTmp;
switch (Op->Header.Size) {
switch (IROp->ElementSize) {
case 4: {
Dst = GetSrcPair<RA_32>(Node);
RegFirst = GetReg<RA_32>(Op->Header.Args[0].ID());
@@ -42,14 +42,6 @@ DEF_OP(VectorImm) {
}
}
DEF_OP(CreateVector2) {
LOGMAN_MSG_A_FMT("Unimplemented");
}
DEF_OP(CreateVector4) {
LOGMAN_MSG_A_FMT("Unimplemented");
}
DEF_OP(SplatVector2) {
auto Op = IROp->C<IR::IROp_SplatVector2>();
uint8_t OpSize = IROp->Size;
@@ -1758,8 +1750,7 @@ DEF_OP(VInsScalarElement) {
DEF_OP(VExtractElement) {
auto Op = IROp->C<IR::IROp_VExtractElement>();
uint8_t OpSize = IROp->Size;
switch (OpSize) {
switch (Op->Header.Size) {
case 1:
mov(GetDst(Node).B(), GetSrc(Op->Header.Args[0].ID()).V16B(), Op->Index);
break;
@@ -1772,7 +1763,7 @@ DEF_OP(VExtractElement) {
case 8:
mov(GetDst(Node).D(), GetSrc(Op->Header.Args[0].ID()).V2D(), Op->Index);
break;
default: LOGMAN_MSG_A_FMT("Unhandled VExtractElement element size: {}", OpSize);
default: LOGMAN_MSG_A_FMT("Unhandled VExtractElement element size: {}", Op->Header.Size);
}
}
@@ -2318,13 +2309,27 @@ DEF_OP(VTBL1) {
}
}
DEF_OP(VRev64) {
auto Op = IROp->C<IR::IROp_VRev64>();
uint8_t OpSize = IROp->Size;
uint8_t Elements = OpSize / Op->Header.ElementSize;
// Vector
switch (Op->Header.ElementSize) {
case 1:
case 2:
case 4:
rev64(GetDst(Node).VCast(OpSize * 8, Elements), GetSrc(Op->Header.Args[0].ID()).VCast(OpSize * 8, Elements));
break;
case 8:
default: LOGMAN_MSG_A_FMT("Invalid Element Size: {}", Op->Header.ElementSize); break;
}
}
#undef DEF_OP
void Arm64JITCore::RegisterVectorHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &Arm64JITCore::Op_##x
REGISTER_OP(VECTORZERO, VectorZero);
REGISTER_OP(VECTORIMM, VectorImm);
REGISTER_OP(CREATEVECTOR2, CreateVector2);
REGISTER_OP(CREATEVECTOR4, CreateVector4);
REGISTER_OP(SPLATVECTOR2, SplatVector2);
REGISTER_OP(SPLATVECTOR4, SplatVector4);
REGISTER_OP(VMOV, VMov);
@@ -2413,6 +2418,7 @@ void Arm64JITCore::RegisterVectorHandlers() {
REGISTER_OP(VSMULL2, VSMull2);
REGISTER_OP(VUABDL, VUABDL);
REGISTER_OP(VTBL1, VTBL1);
REGISTER_OP(VREV64, VRev64);
#undef REGISTER_OP
}
}
+5 -4
View File
@@ -14,10 +14,11 @@ namespace FEXCore::CPU {
class CPUBackend;
[[nodiscard]] std::unique_ptr<CPUBackend> CreateX86JITCore(FEXCore::Context::Context *ctx,
FEXCore::Core::InternalThreadState *Thread,
bool CompileThread);
FEXCore::Core::InternalThreadState *Thread);
void InitializeX86JITSignalHandlers(FEXCore::Context::Context *CTX);
[[nodiscard]] std::unique_ptr<CPUBackend> CreateArm64JITCore(FEXCore::Context::Context *ctx,
FEXCore::Core::InternalThreadState *Thread,
bool CompileThread);
FEXCore::Core::InternalThreadState *Thread);
void InitializeArm64JITSignalHandlers(FEXCore::Context::Context *CTX);
} // namespace FEXCore::CPU
@@ -24,7 +24,7 @@ namespace FEXCore::CPU {
DEF_OP(TruncElementPair) {
auto Op = IROp->C<IR::IROp_TruncElementPair>();
switch (Op->Size) {
switch (IROp->Size) {
case 4: {
auto Dst = GetSrcPair<RA_32>(Node);
auto Src = GetSrcPair<RA_32>(Op->Header.Args[0].ID());
@@ -32,7 +32,7 @@ DEF_OP(TruncElementPair) {
mov(Dst.second, Src.second);
break;
}
default: LOGMAN_MSG_A_FMT("Unhandled Truncation size: {}", Op->Size); break;
default: LOGMAN_MSG_A_FMT("Unhandled Truncation size: {}", IROp->Size); break;
}
}
@@ -45,7 +45,13 @@ DEF_OP(EntrypointOffset) {
auto Op = IROp->C<IR::IROp_EntrypointOffset>();
auto Constant = Entry + Op->Offset;
mov(GetDst<RA_64>(Node), Constant);
uint64_t Mask = ~0ULL;
uint8_t OpSize = IROp->Size;
if (OpSize == 4) {
Mask = 0xFFFF'FFFFULL;
}
mov(GetDst<RA_64>(Node), Constant & Mask);
}
DEF_OP(InlineConstant) {
@@ -1144,19 +1150,19 @@ DEF_OP(VExtractToGPR) {
switch (Op->Header.ElementSize) {
case 1: {
pextrb(GetDst<RA_32>(Node), GetSrc(Op->Header.Args[0].ID()), Op->Idx);
pextrb(GetDst<RA_32>(Node), GetSrc(Op->Header.Args[0].ID()), Op->Index);
break;
}
case 2: {
pextrw(GetDst<RA_32>(Node), GetSrc(Op->Header.Args[0].ID()), Op->Idx);
pextrw(GetDst<RA_32>(Node), GetSrc(Op->Header.Args[0].ID()), Op->Index);
break;
}
case 4: {
pextrd(GetDst<RA_32>(Node), GetSrc(Op->Header.Args[0].ID()), Op->Idx);
pextrd(GetDst<RA_32>(Node), GetSrc(Op->Header.Args[0].ID()), Op->Index);
break;
}
case 8: {
pextrq(GetDst<RA_64>(Node), GetSrc(Op->Header.Args[0].ID()), Op->Idx);
pextrq(GetDst<RA_64>(Node), GetSrc(Op->Header.Args[0].ID()), Op->Index);
break;
}
default: LOGMAN_MSG_A_FMT("Unknown Element Size: {}", Op->Header.ElementSize); break;
@@ -18,20 +18,16 @@ namespace FEXCore::CPU {
#define DEF_OP(x) void X86JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(CASPair) {
auto Op = IROp->C<IR::IROp_CAS>();
uint8_t OpSize = IROp->Size;
// Args[0]: Desired
// Args[1]: Expected
// Args[2]: Pointer
// DataSrc = *Src1
// if (DataSrc == Src3) { *Src1 == Src2; } Src2 = DataSrc
// This will write to memory! Careful!
// Third operand must be a calculated guest memory address
//OrderedNode *CASResult = _CAS(Src3, Src2, Src1);
auto Dst = GetSrcPair<RA_64>(Node);
auto Expected = GetSrcPair<RA_64>(Op->Header.Args[0].ID());
auto Desired = GetSrcPair<RA_64>(Op->Header.Args[1].ID());
auto MemSrc = GetSrc<RA_64>(Op->Header.Args[2].ID());
auto Expected = GetSrcPair<RA_64>(Op->Expected.ID());
auto Desired = GetSrcPair<RA_64>(Op->Desired.ID());
auto MemSrc = GetSrc<RA_64>(Op->Addr.ID());
Xbyak::Reg MemReg = MemSrc;
@@ -47,7 +43,7 @@ DEF_OP(CASPair) {
lock();
switch (OpSize) {
switch (IROp->ElementSize) {
case 4: {
cmpxchg8b(dword [MemReg]);
// EDX:EAX now contains the result
@@ -62,7 +58,7 @@ DEF_OP(CASPair) {
mov(Dst.second, rdx);
break;
}
default: LOGMAN_MSG_A_FMT("Unsupported: {}", OpSize);
default: LOGMAN_MSG_A_FMT("Unsupported: {}", IROp->ElementSize);
}
}
@@ -70,18 +66,15 @@ DEF_OP(CAS) {
auto Op = IROp->C<IR::IROp_CAS>();
uint8_t OpSize = IROp->Size;
// Args[0]: Desired
// Args[1]: Expected
// Args[2]: Pointer
// DataSrc = *Src1
// if (DataSrc == Src3) { *Src1 == Src2; } Src2 = DataSrc
// This will write to memory! Careful!
// Third operand must be a calculated guest memory address
//OrderedNode *CASResult = _CAS(Src3, Src2, Src1);
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Header.Args[2].ID());
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Addr.ID());
mov(rax, GetSrc<RA_64>(Op->Header.Args[0].ID()));
mov(rax, GetSrc<RA_64>(Op->Expected.ID()));
// RCX now contains pointer
// RAX contains our expected value
@@ -90,23 +83,23 @@ DEF_OP(CAS) {
lock();
switch (OpSize) {
case 1: {
cmpxchg(byte [MemReg], GetSrc<RA_8>(Op->Header.Args[1].ID()));
cmpxchg(byte [MemReg], GetSrc<RA_8>(Op->Desired.ID()));
movzx(GetDst<RA_64>(Node), al);
break;
}
case 2: {
cmpxchg(word [MemReg], GetSrc<RA_16>(Op->Header.Args[1].ID()));
cmpxchg(word [MemReg], GetSrc<RA_16>(Op->Desired.ID()));
movzx(GetDst<RA_64>(Node), ax);
break;
}
case 4: {
cmpxchg(dword [MemReg], GetSrc<RA_32>(Op->Header.Args[1].ID()));
cmpxchg(dword [MemReg], GetSrc<RA_32>(Op->Desired.ID()));
// RAX now contains the result
mov (GetDst<RA_64>(Node), eax);
break;
}
case 8: {
cmpxchg(qword [MemReg], GetSrc<RA_64>(Op->Header.Args[1].ID()));
cmpxchg(qword [MemReg], GetSrc<RA_64>(Op->Desired.ID()));
// RAX now contains the result
mov (GetDst<RA_64>(Node), rax);
break;
@@ -118,21 +111,21 @@ DEF_OP(CAS) {
DEF_OP(AtomicAdd) {
auto Op = IROp->C<IR::IROp_AtomicAdd>();
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Header.Args[0].ID());
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Addr.ID());
lock();
switch (IROp->Size) {
case 1:
add(byte [MemReg], GetSrc<RA_8>(Op->Header.Args[1].ID()));
add(byte [MemReg], GetSrc<RA_8>(Op->Value.ID()));
break;
case 2:
add(word [MemReg], GetSrc<RA_16>(Op->Header.Args[1].ID()));
add(word [MemReg], GetSrc<RA_16>(Op->Value.ID()));
break;
case 4:
add(dword [MemReg], GetSrc<RA_32>(Op->Header.Args[1].ID()));
add(dword [MemReg], GetSrc<RA_32>(Op->Value.ID()));
break;
case 8:
add(qword [MemReg], GetSrc<RA_64>(Op->Header.Args[1].ID()));
add(qword [MemReg], GetSrc<RA_64>(Op->Value.ID()));
break;
default: LOGMAN_MSG_A_FMT("Unhandled AtomicAdd size: {}", IROp->Size);
}
@@ -141,20 +134,20 @@ DEF_OP(AtomicAdd) {
DEF_OP(AtomicSub) {
auto Op = IROp->C<IR::IROp_AtomicSub>();
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Header.Args[0].ID());
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Addr.ID());
lock();
switch (IROp->Size) {
case 1:
sub(byte [MemReg], GetSrc<RA_8>(Op->Header.Args[1].ID()));
sub(byte [MemReg], GetSrc<RA_8>(Op->Value.ID()));
break;
case 2:
sub(word [MemReg], GetSrc<RA_16>(Op->Header.Args[1].ID()));
sub(word [MemReg], GetSrc<RA_16>(Op->Value.ID()));
break;
case 4:
sub(dword [MemReg], GetSrc<RA_32>(Op->Header.Args[1].ID()));
sub(dword [MemReg], GetSrc<RA_32>(Op->Value.ID()));
break;
case 8:
sub(qword [MemReg], GetSrc<RA_64>(Op->Header.Args[1].ID()));
sub(qword [MemReg], GetSrc<RA_64>(Op->Value.ID()));
break;
default: LOGMAN_MSG_A_FMT("Unhandled AtomicAdd size: {}", IROp->Size);
}
@@ -163,20 +156,20 @@ DEF_OP(AtomicSub) {
DEF_OP(AtomicAnd) {
auto Op = IROp->C<IR::IROp_AtomicAnd>();
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Header.Args[0].ID());
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Addr.ID());
lock();
switch (IROp->Size) {
case 1:
and_(byte [MemReg], GetSrc<RA_8>(Op->Header.Args[1].ID()));
and_(byte [MemReg], GetSrc<RA_8>(Op->Value.ID()));
break;
case 2:
and_(word [MemReg], GetSrc<RA_16>(Op->Header.Args[1].ID()));
and_(word [MemReg], GetSrc<RA_16>(Op->Value.ID()));
break;
case 4:
and_(dword [MemReg], GetSrc<RA_32>(Op->Header.Args[1].ID()));
and_(dword [MemReg], GetSrc<RA_32>(Op->Value.ID()));
break;
case 8:
and_(qword [MemReg], GetSrc<RA_64>(Op->Header.Args[1].ID()));
and_(qword [MemReg], GetSrc<RA_64>(Op->Value.ID()));
break;
default: LOGMAN_MSG_A_FMT("Unhandled AtomicAdd size: {}", IROp->Size);
}
@@ -185,20 +178,20 @@ DEF_OP(AtomicAnd) {
DEF_OP(AtomicOr) {
auto Op = IROp->C<IR::IROp_AtomicOr>();
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Header.Args[0].ID());
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Addr.ID());
lock();
switch (IROp->Size) {
case 1:
or_(byte [MemReg], GetSrc<RA_8>(Op->Header.Args[1].ID()));
or_(byte [MemReg], GetSrc<RA_8>(Op->Value.ID()));
break;
case 2:
or_(word [MemReg], GetSrc<RA_16>(Op->Header.Args[1].ID()));
or_(word [MemReg], GetSrc<RA_16>(Op->Value.ID()));
break;
case 4:
or_(dword [MemReg], GetSrc<RA_32>(Op->Header.Args[1].ID()));
or_(dword [MemReg], GetSrc<RA_32>(Op->Value.ID()));
break;
case 8:
or_(qword [MemReg], GetSrc<RA_64>(Op->Header.Args[1].ID()));
or_(qword [MemReg], GetSrc<RA_64>(Op->Value.ID()));
break;
default: LOGMAN_MSG_A_FMT("Unhandled AtomicAdd size: {}", IROp->Size);
}
@@ -207,20 +200,20 @@ DEF_OP(AtomicOr) {
DEF_OP(AtomicXor) {
auto Op = IROp->C<IR::IROp_AtomicXor>();
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Header.Args[0].ID());
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Addr.ID());
lock();
switch (IROp->Size) {
case 1:
xor_(byte [MemReg], GetSrc<RA_8>(Op->Header.Args[1].ID()));
xor_(byte [MemReg], GetSrc<RA_8>(Op->Value.ID()));
break;
case 2:
xor_(word [MemReg], GetSrc<RA_16>(Op->Header.Args[1].ID()));
xor_(word [MemReg], GetSrc<RA_16>(Op->Value.ID()));
break;
case 4:
xor_(dword [MemReg], GetSrc<RA_32>(Op->Header.Args[1].ID()));
xor_(dword [MemReg], GetSrc<RA_32>(Op->Value.ID()));
break;
case 8:
xor_(qword [MemReg], GetSrc<RA_64>(Op->Header.Args[1].ID()));
xor_(qword [MemReg], GetSrc<RA_64>(Op->Value.ID()));
break;
default: LOGMAN_MSG_A_FMT("Unhandled AtomicAdd size: {}", IROp->Size);
}
@@ -230,26 +223,26 @@ DEF_OP(AtomicSwap) {
auto Op = IROp->C<IR::IROp_AtomicSwap>();
Xbyak::Reg MemReg = rax;
mov(MemReg, GetSrc<RA_64>(Op->Header.Args[0].ID()));
mov(MemReg, GetSrc<RA_64>(Op->Addr.ID()));
switch (IROp->Size) {
case 1:
movzx(GetDst<RA_64>(Node), GetSrc<RA_8>(Op->Header.Args[1].ID()));
movzx(GetDst<RA_64>(Node), GetSrc<RA_8>(Op->Value.ID()));
lock();
xchg(byte [MemReg], GetDst<RA_8>(Node));
break;
case 2:
movzx(GetDst<RA_64>(Node), GetSrc<RA_16>(Op->Header.Args[1].ID()));
movzx(GetDst<RA_64>(Node), GetSrc<RA_16>(Op->Value.ID()));
lock();
xchg(word [MemReg], GetDst<RA_16>(Node));
break;
case 4:
mov(GetDst<RA_64>(Node), GetSrc<RA_32>(Op->Header.Args[1].ID()));
mov(GetDst<RA_64>(Node), GetSrc<RA_32>(Op->Value.ID()));
lock();
xchg(dword [MemReg], GetDst<RA_32>(Node));
break;
case 8:
mov(GetDst<RA_64>(Node), GetSrc<RA_64>(Op->Header.Args[1].ID()));
mov(GetDst<RA_64>(Node), GetSrc<RA_64>(Op->Value.ID()));
lock();
xchg(qword [MemReg], GetDst<RA_64>(Node));
break;
@@ -260,28 +253,28 @@ DEF_OP(AtomicSwap) {
DEF_OP(AtomicFetchAdd) {
auto Op = IROp->C<IR::IROp_AtomicFetchAdd>();
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Header.Args[0].ID());
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Addr.ID());
switch (IROp->Size) {
case 1:
movzx(rcx, GetSrc<RA_8>(Op->Header.Args[1].ID()));
movzx(rcx, GetSrc<RA_8>(Op->Value.ID()));
lock();
xadd(byte [MemReg], cl);
movzx(GetDst<RA_32>(Node), cl);
break;
case 2:
movzx(rcx, GetSrc<RA_16>(Op->Header.Args[1].ID()));
movzx(rcx, GetSrc<RA_16>(Op->Value.ID()));
lock();
xadd(word [MemReg], cx);
movzx(GetDst<RA_32>(Node), cx);
break;
case 4:
mov(ecx, GetSrc<RA_32>(Op->Header.Args[1].ID()));
mov(ecx, GetSrc<RA_32>(Op->Value.ID()));
lock();
xadd(dword [MemReg], ecx);
mov(GetDst<RA_64>(Node), ecx);
break;
case 8:
mov(rcx, GetSrc<RA_64>(Op->Header.Args[1].ID()));
mov(rcx, GetSrc<RA_64>(Op->Value.ID()));
lock();
xadd(qword [MemReg], rcx);
mov(GetDst<RA_64>(Node), rcx);
@@ -293,31 +286,31 @@ DEF_OP(AtomicFetchAdd) {
DEF_OP(AtomicFetchSub) {
auto Op = IROp->C<IR::IROp_AtomicFetchSub>();
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Header.Args[0].ID());
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Addr.ID());
switch (IROp->Size) {
case 1:
mov(cl, GetSrc<RA_8>(Op->Header.Args[1].ID()));
mov(cl, GetSrc<RA_8>(Op->Value.ID()));
neg(cl);
lock();
xadd(byte [MemReg], cl);
movzx(GetDst<RA_32>(Node), cl);
break;
case 2:
mov(cx, GetSrc<RA_16>(Op->Header.Args[1].ID()));
mov(cx, GetSrc<RA_16>(Op->Value.ID()));
neg(cx);
lock();
xadd(word [MemReg], cx);
movzx(GetDst<RA_32>(Node), cx);
break;
case 4:
mov(ecx, GetSrc<RA_32>(Op->Header.Args[1].ID()));
mov(ecx, GetSrc<RA_32>(Op->Value.ID()));
neg(ecx);
lock();
xadd(dword [MemReg], ecx);
mov(GetDst<RA_32>(Node), ecx);
break;
case 8:
mov(rcx, GetSrc<RA_64>(Op->Header.Args[1].ID()));
mov(rcx, GetSrc<RA_64>(Op->Value.ID()));
neg(rcx);
lock();
xadd(qword [MemReg], rcx);
@@ -331,7 +324,7 @@ DEF_OP(AtomicFetchAnd) {
auto Op = IROp->C<IR::IROp_AtomicFetchAnd>();
// TMP1 = rax
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Header.Args[0].ID());
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Addr.ID());
switch (IROp->Size) {
case 1: {
@@ -341,7 +334,7 @@ DEF_OP(AtomicFetchAnd) {
L(Loop);
mov(TMP2.cvt8(), TMP1.cvt8());
mov(TMP3.cvt8(), TMP1.cvt8());
and_(TMP2.cvt8(), GetSrc<RA_8>(Op->Header.Args[1].ID()));
and_(TMP2.cvt8(), GetSrc<RA_8>(Op->Value.ID()));
// Updates RAX with the value from memory
lock(); cmpxchg(byte [MemReg], TMP2.cvt8());
@@ -357,7 +350,7 @@ DEF_OP(AtomicFetchAnd) {
L(Loop);
mov(TMP2.cvt16(), TMP1.cvt16());
mov(TMP3.cvt16(), TMP1.cvt16());
and_(TMP2.cvt16(), GetSrc<RA_16>(Op->Header.Args[1].ID()));
and_(TMP2.cvt16(), GetSrc<RA_16>(Op->Value.ID()));
// Updates RAX with the value from memory
lock(); cmpxchg(word [MemReg], TMP2.cvt16());
@@ -374,7 +367,7 @@ DEF_OP(AtomicFetchAnd) {
L(Loop);
mov(TMP2.cvt32(), TMP1.cvt32());
mov(TMP3.cvt32(), TMP1.cvt32());
and_(TMP2.cvt32(), GetSrc<RA_32>(Op->Header.Args[1].ID()));
and_(TMP2.cvt32(), GetSrc<RA_32>(Op->Value.ID()));
// Updates RAX with the value from memory
lock(); cmpxchg(dword [MemReg], TMP2.cvt32());
@@ -391,7 +384,7 @@ DEF_OP(AtomicFetchAnd) {
L(Loop);
mov(TMP2.cvt64(), TMP1.cvt64());
mov(TMP3.cvt64(), TMP1.cvt64());
and_(TMP2.cvt64(), GetSrc<RA_64>(Op->Header.Args[1].ID()));
and_(TMP2.cvt64(), GetSrc<RA_64>(Op->Value.ID()));
// Updates RAX with the value from memory
lock(); cmpxchg(qword [MemReg], TMP2.cvt64());
@@ -409,7 +402,7 @@ DEF_OP(AtomicFetchOr) {
auto Op = IROp->C<IR::IROp_AtomicFetchOr>();
// TMP1 = rax
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Header.Args[0].ID());
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Addr.ID());
switch (IROp->Size) {
case 1: {
mov(TMP1.cvt8(), byte [MemReg]);
@@ -418,7 +411,7 @@ DEF_OP(AtomicFetchOr) {
L(Loop);
mov(TMP2.cvt8(), TMP1.cvt8());
mov(TMP3.cvt8(), TMP1.cvt8());
or_(TMP2.cvt8(), GetSrc<RA_8>(Op->Header.Args[1].ID()));
or_(TMP2.cvt8(), GetSrc<RA_8>(Op->Value.ID()));
// Updates RAX with the value from memory
lock(); cmpxchg(byte [MemReg], TMP2.cvt8());
@@ -434,7 +427,7 @@ DEF_OP(AtomicFetchOr) {
L(Loop);
mov(TMP2.cvt16(), TMP1.cvt16());
mov(TMP3.cvt16(), TMP1.cvt16());
or_(TMP2.cvt16(), GetSrc<RA_16>(Op->Header.Args[1].ID()));
or_(TMP2.cvt16(), GetSrc<RA_16>(Op->Value.ID()));
// Updates RAX with the value from memory
lock(); cmpxchg(word [MemReg], TMP2.cvt16());
@@ -451,7 +444,7 @@ DEF_OP(AtomicFetchOr) {
L(Loop);
mov(TMP2.cvt32(), TMP1.cvt32());
mov(TMP3.cvt32(), TMP1.cvt32());
or_(TMP2.cvt32(), GetSrc<RA_32>(Op->Header.Args[1].ID()));
or_(TMP2.cvt32(), GetSrc<RA_32>(Op->Value.ID()));
// Updates RAX with the value from memory
lock(); cmpxchg(dword [MemReg], TMP2.cvt32());
@@ -468,7 +461,7 @@ DEF_OP(AtomicFetchOr) {
L(Loop);
mov(TMP2.cvt64(), TMP1.cvt64());
mov(TMP3.cvt64(), TMP1.cvt64());
or_(TMP2.cvt64(), GetSrc<RA_64>(Op->Header.Args[1].ID()));
or_(TMP2.cvt64(), GetSrc<RA_64>(Op->Value.ID()));
// Updates RAX with the value from memory
lock(); cmpxchg(qword [MemReg], TMP2.cvt64());
@@ -486,7 +479,7 @@ DEF_OP(AtomicFetchXor) {
auto Op = IROp->C<IR::IROp_AtomicFetchXor>();
// TMP1 = rax
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Header.Args[0].ID());
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Addr.ID());
switch (IROp->Size) {
case 1: {
mov(TMP1.cvt8(), byte [MemReg]);
@@ -495,7 +488,7 @@ DEF_OP(AtomicFetchXor) {
L(Loop);
mov(TMP2.cvt8(), TMP1.cvt8());
mov(TMP3.cvt8(), TMP1.cvt8());
xor_(TMP2.cvt8(), GetSrc<RA_8>(Op->Header.Args[1].ID()));
xor_(TMP2.cvt8(), GetSrc<RA_8>(Op->Value.ID()));
// Updates RAX with the value from memory
lock(); cmpxchg(byte [MemReg], TMP2.cvt8());
@@ -511,7 +504,7 @@ DEF_OP(AtomicFetchXor) {
L(Loop);
mov(TMP2.cvt16(), TMP1.cvt16());
mov(TMP3.cvt16(), TMP1.cvt16());
xor_(TMP2.cvt16(), GetSrc<RA_16>(Op->Header.Args[1].ID()));
xor_(TMP2.cvt16(), GetSrc<RA_16>(Op->Value.ID()));
// Updates RAX with the value from memory
lock(); cmpxchg(word [MemReg], TMP2.cvt16());
@@ -528,7 +521,7 @@ DEF_OP(AtomicFetchXor) {
L(Loop);
mov(TMP2.cvt32(), TMP1.cvt32());
mov(TMP3.cvt32(), TMP1.cvt32());
xor_(TMP2.cvt32(), GetSrc<RA_32>(Op->Header.Args[1].ID()));
xor_(TMP2.cvt32(), GetSrc<RA_32>(Op->Value.ID()));
// Updates RAX with the value from memory
lock(); cmpxchg(dword [MemReg], TMP2.cvt32());
@@ -545,7 +538,7 @@ DEF_OP(AtomicFetchXor) {
L(Loop);
mov(TMP2.cvt64(), TMP1.cvt64());
mov(TMP3.cvt64(), TMP1.cvt64());
xor_(TMP2.cvt64(), GetSrc<RA_64>(Op->Header.Args[1].ID()));
xor_(TMP2.cvt64(), GetSrc<RA_64>(Op->Value.ID()));
// Updates RAX with the value from memory
lock(); cmpxchg(qword [MemReg], TMP2.cvt64());
@@ -562,7 +555,7 @@ DEF_OP(AtomicFetchXor) {
DEF_OP(AtomicFetchNeg) {
auto Op = IROp->C<IR::IROp_AtomicFetchNeg>();
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Header.Args[0].ID());
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Addr.ID());
switch (IROp->Size) {
case 1: {
mov(TMP1.cvt8(), byte [MemReg]);
@@ -37,18 +37,13 @@ DEF_OP(GuestCallIndirect) {
LogMan::Msg::DFmt("Unimplemented");
}
DEF_OP(GuestReturn) {
LogMan::Msg::DFmt("Unimplemented");
}
DEF_OP(SignalReturn) {
// Adjust the stack first for a regular return
if (SpillSlots) {
add(rsp, SpillSlots * 16); // + 8 to consume return address
}
mov(TMP1, ThreadSharedData.SignalHandlerReturnAddress);
jmp(TMP1);
jmp(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.SignalReturnHandler)]);
}
DEF_OP(CallbackReturn) {
@@ -58,8 +53,7 @@ DEF_OP(CallbackReturn) {
}
// Make sure to adjust the refcounter so we don't clear the cache now
mov(rax, reinterpret_cast<uint64_t>(ThreadSharedData.SignalHandlerRefCounterPtr));
sub(dword [rax], 1);
sub(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.SignalHandlerRefCountPointer)], 1);
// We need to adjust an additional 8 bytes to get back to the original "misaligned" RSP state
add(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RSP])], 8);
@@ -97,14 +91,15 @@ DEF_OP(ExitFunction) {
jmp(qword[rax]);
L(l_BranchHost);
dq(ThreadSharedData.Dispatcher->ExitFunctionLinkerAddress);
dq(Dispatcher->ExitFunctionLinkerAddress);
L(l_BranchGuest);
dq(NewRIP);
} else {
Xbyak::Reg RipReg = GetSrc<RA_64>(Op->NewRIP.ID());
// L1 Cache
mov(rcx, ThreadState->LookupCache->GetL1Pointer());
mov(rcx, qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.L1Pointer)]);
mov(rax, RipReg);
and_(rax, LookupCache::L1_ENTRIES_MASK);
@@ -117,9 +112,8 @@ DEF_OP(ExitFunction) {
jmp(qword[LookupBase + 0]);
L(FullLookup);
mov(rax, ThreadSharedData.Dispatcher->AbsoluteLoopTopAddress);
mov(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, State.rip)], RipReg);
jmp(rax);
jmp(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.DispatcherLoopTop)]);
}
#ifdef BLOCKSTATS
@@ -187,15 +181,13 @@ DEF_OP(Syscall) {
}
mov(rsi, STATE); // Move thread in to rsi
mov(rdi, reinterpret_cast<uint64_t>(CTX->SyscallHandler));
mov(rdi, qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.SyscallHandlerObj)]);
mov(rdx, rsp);
mov(rax, reinterpret_cast<uint64_t>(FEXCore::Context::HandleSyscall));
if (NumPush & 1)
sub(rsp, 8); // Align
// {rdi, rsi, rdx}
call(rax);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.SyscallHandlerFunc)]);
if (NumPush & 1)
add(rsp, 8); // Align
@@ -267,7 +259,7 @@ DEF_OP(ValidateCode) {
}
}
DEF_OP(RemoveCodeEntry) {
DEF_OP(RemoveThreadCodeEntry) {
auto NumPush = RA64.size();
for (auto &Reg : RA64)
@@ -280,9 +272,7 @@ DEF_OP(RemoveCodeEntry) {
mov(rax, Entry); // imm64 move
mov(rsi, rax);
mov(rax, reinterpret_cast<uintptr_t>(&Context::Context::RemoveCodeEntryFromJit));
call(rax);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.RemoveThreadCodeEntryFromJIT)]);
if (NumPush & 1)
add(rsp, 8); // Align
@@ -294,13 +284,6 @@ DEF_OP(RemoveCodeEntry) {
DEF_OP(CPUID) {
auto Op = IROp->C<IR::IROp_CPUID>();
using ClassPtrType = FEXCore::CPUID::FunctionResults (FEXCore::CPUIDEmu::*)(uint32_t Function, uint32_t Leaf);
union {
ClassPtrType ClassPtr;
uint64_t Raw;
} Ptr;
Ptr.ClassPtr = &CPUIDEmu::RunFunction;
for (auto &Reg : RA64)
push(Reg);
@@ -313,18 +296,15 @@ DEF_OP(CPUID) {
// rsi can be in the source registers, so copy argument to edx first
mov (edx, GetSrc<RA_32>(Op->Header.Args[1].ID()));
mov (esi, GetSrc<RA_32>(Op->Header.Args[0].ID()));
mov (rdi, reinterpret_cast<uint64_t>(&CTX->CPUID));
mov (rdi, qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.CPUIDObj)]);
auto NumPush = RA64.size();
if (NumPush & 1)
sub(rsp, 8); // Align
mov(rax, Ptr.Raw);
// {rdi, rsi, rdx}
call(rax);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.CPUIDFunction)]);
if (NumPush & 1)
add(rsp, 8); // Align
@@ -342,7 +322,6 @@ void X86JITCore::RegisterBranchHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &X86JITCore::Op_##x
REGISTER_OP(GUESTCALLDIRECT, GuestCallDirect);
REGISTER_OP(GUESTCALLINDIRECT, GuestCallIndirect);
REGISTER_OP(GUESTRETURN, GuestReturn);
REGISTER_OP(SIGNALRETURN, SignalReturn);
REGISTER_OP(CALLBACKRETURN, CallbackReturn);
REGISTER_OP(EXITFUNCTION, ExitFunction);
@@ -351,7 +330,7 @@ void X86JITCore::RegisterBranchHandlers() {
REGISTER_OP(SYSCALL, Syscall);
REGISTER_OP(THUNK, Thunk);
REGISTER_OP(VALIDATECODE, ValidateCode);
REGISTER_OP(REMOVECODEENTRY, RemoveCodeEntry);
REGISTER_OP(REMOVETHREADCODEENTRY, RemoveThreadCodeEntry);
REGISTER_OP(CPUID, CPUID);
#undef REGISTER_OP
}
@@ -18,23 +18,23 @@ namespace FEXCore::CPU {
#define DEF_OP(x) void X86JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(VInsGPR) {
auto Op = IROp->C<IR::IROp_VInsGPR>();
movapd(GetDst(Node), GetSrc(Op->Header.Args[0].ID()));
movapd(GetDst(Node), GetSrc(Op->DestVector.ID()));
switch (Op->Header.ElementSize) {
case 1: {
pinsrb(GetDst(Node), GetSrc<RA_32>(Op->Header.Args[1].ID()), Op->Index);
pinsrb(GetDst(Node), GetSrc<RA_32>(Op->Src.ID()), Op->DestIdx);
break;
}
case 2: {
pinsrw(GetDst(Node), GetSrc<RA_32>(Op->Header.Args[1].ID()), Op->Index);
pinsrw(GetDst(Node), GetSrc<RA_32>(Op->Src.ID()), Op->DestIdx);
break;
}
case 4: {
pinsrd(GetDst(Node), GetSrc<RA_32>(Op->Header.Args[1].ID()), Op->Index);
pinsrd(GetDst(Node), GetSrc<RA_32>(Op->Src.ID()), Op->DestIdx);
break;
}
case 8: {
pinsrq(GetDst(Node), GetSrc<RA_64>(Op->Header.Args[1].ID()), Op->Index);
pinsrq(GetDst(Node), GetSrc<RA_64>(Op->Src.ID()), Op->DestIdx);
break;
}
default: LOGMAN_MSG_A_FMT("Unknown Element Size: {}", Op->Header.ElementSize); break;
@@ -45,6 +45,36 @@ DEF_OP(AESKeyGenAssist) {
vaeskeygenassist(GetDst(Node), GetSrc(Op->Header.Args[0].ID()), Op->RCON);
}
DEF_OP(CRC32) {
auto Op = IROp->C<IR::IROp_CRC32>();
switch (IROp->Size) {
case 4:
mov(TMP1, GetSrc<RA_32>(Op->Header.Args[1].ID()));
mov(GetDst<RA_32>(Node), GetSrc<RA_32>(Op->Header.Args[0].ID()));
break;
case 8:
mov(TMP1, GetSrc<RA_64>(Op->Header.Args[1].ID()));
mov(GetDst<RA_64>(Node), GetSrc<RA_64>(Op->Header.Args[0].ID()));
break;
default: LOGMAN_MSG_A_FMT("Unknown CRC32 size: {}", IROp->Size);
}
switch (Op->SrcSize) {
case 1:
crc32(GetDst<RA_32>(Node).cvt32(), TMP1.cvt8());
break;
case 2:
crc32(GetDst<RA_32>(Node).cvt32(), TMP1.cvt16());
break;
case 4:
crc32(GetDst<RA_32>(Node).cvt32(), TMP1.cvt32());
break;
case 8:
crc32(GetDst<RA_64>(Node).cvt64(), TMP1.cvt64());
break;
}
}
#undef DEF_OP
void X86JITCore::RegisterEncryptionHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &X86JITCore::Op_##x
@@ -54,7 +84,7 @@ void X86JITCore::RegisterEncryptionHandlers() {
REGISTER_OP(VAESDEC, AESDec);
REGISTER_OP(VAESDECLAST, AESDecLast);
REGISTER_OP(VAESKEYGENASSIST, AESKeyGenAssist);
REGISTER_OP(CRC32, CRC32);
#undef REGISTER_OP
}
}
+129 -80
View File
@@ -15,6 +15,8 @@ $end_info$
#include "Interface/IR/PassManager.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
#include "Utils/MemberFunctionToPointer.h"
#include <FEXCore/Core/CPUBackend.h>
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Core/SignalDelegator.h>
@@ -27,7 +29,6 @@ $end_info$
#include <algorithm>
#include <array>
#include <bits/types/stack_t.h>
#include <memory>
#include <stddef.h>
#include <stdint.h>
@@ -42,6 +43,16 @@ $end_info$
// #define DEBUG_RA 1
// #define DEBUG_CYCLES
namespace {
static void PrintValue(uint64_t Value) {
LogMan::Msg::DFmt("Value: 0x{:x}", Value);
}
static void PrintVectorValue(uint64_t Value, uint64_t ValueUpper) {
LogMan::Msg::DFmt("Value: 0x{:016x}'{:016x}", ValueUpper, Value);
}
}
namespace FEXCore::CPU {
CodeBuffer AllocateNewCodeBuffer(FEXCore::Context::Context *CTX, size_t Size) {
@@ -64,11 +75,6 @@ void FreeCodeBuffer(CodeBuffer Buffer) {
FEXCore::Allocator::munmap(Buffer.Ptr, Buffer.Size);
}
void X86JITCore::CopyNecessaryDataForCompileThread(CPUBackend *Original) {
X86JITCore *Core = reinterpret_cast<X86JITCore*>(Original);
ThreadSharedData = Core->ThreadSharedData;
}
void X86JITCore::PushRegs() {
sub(rsp, 16 * RAXMM_x.size());
for (size_t i = 0; i < RAXMM_x.size(); ++i) {
@@ -109,9 +115,7 @@ void X86JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
case FABI_VOID_U16: {
PushRegs();
mov(edi, GetSrc<RA_32>(IROp->Args[0].ID()));
mov(rax, (uintptr_t)Info.fn);
call(rax);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.FallbackHandlerPointers[Info.HandlerIndex])]);
PopRegs();
break;
@@ -120,9 +124,7 @@ void X86JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
PushRegs();
movss(xmm0, GetSrc(IROp->Args[0].ID()));
mov(rax, (uintptr_t)Info.fn);
call(rax);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.FallbackHandlerPointers[Info.HandlerIndex])]);
PopRegs();
@@ -136,9 +138,7 @@ void X86JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
PushRegs();
movsd(xmm0, GetSrc(IROp->Args[0].ID()));
mov(rax, (uintptr_t)Info.fn);
call(rax);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.FallbackHandlerPointers[Info.HandlerIndex])]);
PopRegs();
@@ -153,9 +153,7 @@ void X86JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
PushRegs();
mov(edi, GetSrc<RA_32>(IROp->Args[0].ID()));
mov(rax, (uintptr_t)Info.fn);
call(rax);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.FallbackHandlerPointers[Info.HandlerIndex])]);
PopRegs();
@@ -171,9 +169,7 @@ void X86JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
movq(rdi, GetSrc(IROp->Args[0].ID()));
pextrq(rsi, GetSrc(IROp->Args[0].ID()), 1);
mov(rax, (uintptr_t)Info.fn);
call(rax);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.FallbackHandlerPointers[Info.HandlerIndex])]);
PopRegs();
@@ -187,9 +183,34 @@ void X86JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
movq(rdi, GetSrc(IROp->Args[0].ID()));
pextrq(rsi, GetSrc(IROp->Args[0].ID()), 1);
mov(rax, (uintptr_t)Info.fn);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.FallbackHandlerPointers[Info.HandlerIndex])]);
call(rax);
PopRegs();
movsd(GetDst(Node), xmm0);
}
break;
case FABI_F64_F64: {
PushRegs();
movsd(xmm0, GetSrc(IROp->Args[0].ID()));
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.FallbackHandlerPointers[Info.HandlerIndex])]);
PopRegs();
movsd(GetDst(Node), xmm0);
}
break;
case FABI_F64_F64_F64: {
PushRegs();
movsd(xmm0, GetSrc(IROp->Args[0].ID()));
movsd(xmm1, GetSrc(IROp->Args[1].ID()));
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.FallbackHandlerPointers[Info.HandlerIndex])]);
PopRegs();
@@ -203,9 +224,7 @@ void X86JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
movq(rdi, GetSrc(IROp->Args[0].ID()));
pextrq(rsi, GetSrc(IROp->Args[0].ID()), 1);
mov(rax, (uintptr_t)Info.fn);
call(rax);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.FallbackHandlerPointers[Info.HandlerIndex])]);
PopRegs();
@@ -218,9 +237,7 @@ void X86JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
movq(rdi, GetSrc(IROp->Args[0].ID()));
pextrq(rsi, GetSrc(IROp->Args[0].ID()), 1);
mov(rax, (uintptr_t)Info.fn);
call(rax);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.FallbackHandlerPointers[Info.HandlerIndex])]);
PopRegs();
@@ -233,9 +250,7 @@ void X86JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
movq(rdi, GetSrc(IROp->Args[0].ID()));
pextrq(rsi, GetSrc(IROp->Args[0].ID()), 1);
mov(rax, (uintptr_t)Info.fn);
call(rax);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.FallbackHandlerPointers[Info.HandlerIndex])]);
PopRegs();
@@ -251,9 +266,7 @@ void X86JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
movq(rdx, GetSrc(IROp->Args[1].ID()));
pextrq(rcx, GetSrc(IROp->Args[1].ID()), 1);
mov(rax, (uintptr_t)Info.fn);
call(rax);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.FallbackHandlerPointers[Info.HandlerIndex])]);
PopRegs();
@@ -266,9 +279,7 @@ void X86JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
movq(rdi, GetSrc(IROp->Args[0].ID()));
pextrq(rsi, GetSrc(IROp->Args[0].ID()), 1);
mov(rax, (uintptr_t)Info.fn);
call(rax);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.FallbackHandlerPointers[Info.HandlerIndex])]);
PopRegs();
@@ -286,9 +297,7 @@ void X86JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
movq(rdx, GetSrc(IROp->Args[1].ID()));
pextrq(rcx, GetSrc(IROp->Args[1].ID()), 1);
mov(rax, (uintptr_t)Info.fn);
call(rax);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.FallbackHandlerPointers[Info.HandlerIndex])]);
PopRegs();
@@ -311,13 +320,14 @@ void X86JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
void X86JITCore::Op_NoOp(IR::IROp_Header *IROp, IR::NodeID Node) {
}
X86JITCore::X86JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, CodeBuffer Buffer, bool CompileThread)
X86JITCore::X86JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, CodeBuffer Buffer)
: CodeGenerator(Buffer.Size, Buffer.Ptr, nullptr)
, CTX {ctx}
, ThreadState {Thread}
, InitialCodeBuffer {Buffer}
{
CurrentCodeBuffer = &InitialCodeBuffer;
EmitDetectionString();
RAPass = Thread->PassManager->GetPass<IR::RegisterAllocationPass>("RA");
@@ -346,41 +356,58 @@ X86JITCore::X86JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalTh
RegisterVectorHandlers();
RegisterEncryptionHandlers();
if (!CompileThread) {
DispatcherConfig config;
config.ExitFunctionLink = reinterpret_cast<uintptr_t>(&ExitFunctionLink);
config.ExitFunctionLinkThis = reinterpret_cast<uintptr_t>(this);
DispatcherConfig config;
config.ExitFunctionLink = reinterpret_cast<uintptr_t>(&ExitFunctionLink);
config.ExitFunctionLinkThis = reinterpret_cast<uintptr_t>(this);
config.StaticRegisterAssignment = ctx->Config.StaticRegisterAllocation;
Dispatcher = std::make_unique<X86Dispatcher>(CTX, ThreadState, config);
DispatchPtr = Dispatcher->DispatchPtr;
CallbackPtr = Dispatcher->CallbackPtr;
Dispatcher = std::make_unique<X86Dispatcher>(CTX, ThreadState, config);
DispatchPtr = Dispatcher->DispatchPtr;
CallbackPtr = Dispatcher->CallbackPtr;
ThreadSharedData.SignalHandlerRefCounterPtr = &Dispatcher->SignalHandlerRefCounter;
ThreadSharedData.SignalHandlerReturnAddress = Dispatcher->SignalHandlerReturnAddress;
ThreadSharedData.UnimplementedInstructionAddress = Dispatcher->UnimplementedInstructionAddress;
ThreadSharedData.OverflowExceptionInstructionAddress = Dispatcher->OverflowExceptionInstructionAddress;
{
// Set up pointers that the JIT needs to load
auto &Pointers = ThreadState->CurrentFrame->Pointers.X86;
// Process specific
Pointers.PrintValue = reinterpret_cast<uint64_t>(PrintValue);
Pointers.PrintVectorValue = reinterpret_cast<uint64_t>(PrintVectorValue);
Pointers.RemoveThreadCodeEntryFromJIT = reinterpret_cast<uintptr_t>(&Context::Context::RemoveThreadCodeEntryFromJit);
Pointers.CPUIDObj = reinterpret_cast<uint64_t>(&CTX->CPUID);
ThreadSharedData.Dispatcher = Dispatcher.get();
// This will register the host signal handler per thread, which is fine
CTX->SignalDelegation->RegisterHostSignalHandler(SIGILL, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
X86JITCore *Core = reinterpret_cast<X86JITCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleSIGILL(Signal, info, ucontext);
}, true);
CTX->SignalDelegation->RegisterHostSignalHandler(SignalDelegator::SIGNAL_FOR_PAUSE, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
X86JITCore *Core = reinterpret_cast<X86JITCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleSignalPause(Signal, info, ucontext);
}, true);
auto GuestSignalHandler = [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack) -> bool {
X86JITCore *Core = reinterpret_cast<X86JITCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleGuestSignal(Signal, info, ucontext, GuestAction, GuestStack);
};
for (uint32_t Signal = 0; Signal <= SignalDelegator::MAX_SIGNALS; ++Signal) {
CTX->SignalDelegation->RegisterHostSignalHandlerForGuest(Signal, GuestSignalHandler);
{
FEXCore::Utils::MemberFunctionToPointerCast PMF(&FEXCore::CPUIDEmu::RunFunction);
Pointers.CPUIDFunction = PMF.GetConvertedPointer();
}
Pointers.SyscallHandlerObj = reinterpret_cast<uint64_t>(CTX->SyscallHandler);
Pointers.SyscallHandlerFunc = reinterpret_cast<uint64_t>(FEXCore::Context::HandleSyscall);
// Fill in the fallback handlers
InterpreterOps::FillFallbackIndexPointers(Pointers.FallbackHandlerPointers);
// Thread Specific
Pointers.SignalHandlerRefCountPointer = reinterpret_cast<uint64_t>(&Dispatcher->SignalHandlerRefCounter);
}
}
void X86JITCore::InitializeSignalHandlers(FEXCore::Context::Context *CTX) {
CTX->SignalDelegation->RegisterHostSignalHandler(SIGILL, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
X86JITCore *Core = reinterpret_cast<X86JITCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleSIGILL(Signal, info, ucontext);
}, true);
CTX->SignalDelegation->RegisterHostSignalHandler(SignalDelegator::SIGNAL_FOR_PAUSE, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
X86JITCore *Core = reinterpret_cast<X86JITCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleSignalPause(Signal, info, ucontext);
}, true);
auto GuestSignalHandler = [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack) -> bool {
X86JITCore *Core = reinterpret_cast<X86JITCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleGuestSignal(Signal, info, ucontext, GuestAction, GuestStack);
};
for (uint32_t Signal = 0; Signal <= SignalDelegator::MAX_SIGNALS; ++Signal) {
CTX->SignalDelegation->RegisterHostSignalHandlerForGuest(Signal, GuestSignalHandler);
}
}
@@ -394,8 +421,15 @@ X86JITCore::~X86JITCore() {
FreeCodeBuffer(InitialCodeBuffer);
}
void X86JITCore::EmitDetectionString() {
const char JITString[] = "FEXJIT::X86JITCore::";
for (char c : JITString) {
db(c);
}
}
void X86JITCore::ClearCache() {
if (*ThreadSharedData.SignalHandlerRefCounterPtr == 0) {
if (Dispatcher->SignalHandlerRefCounter == 0) {
if (!CodeBuffers.empty()) {
// If we have more than one code buffer we are tracking then walk them and delete
// This is a cleanup step
@@ -432,6 +466,8 @@ void X86JITCore::ClearCache() {
EmplaceNewCodeBuffer(NewCodeBuffer);
setNewBuffer(NewCodeBuffer.Ptr, NewCodeBuffer.Size);
}
EmitDetectionString();
}
IR::PhysicalRegister X86JITCore::GetPhys(IR::NodeID Node) const {
@@ -553,7 +589,12 @@ bool X86JITCore::IsInlineEntrypointOffset(const IR::OrderedNodeWrapper& WNode, u
if (OpHeader->Op == IR::IROps::OP_INLINEENTRYPOINTOFFSET) {
auto Op = OpHeader->C<IR::IROp_InlineEntrypointOffset>();
if (Value) {
*Value = Entry + Op->Offset;
uint64_t Mask = ~0ULL;
uint8_t OpSize = OpHeader->Size;
if (OpSize == 4) {
Mask = 0xFFFF'FFFFULL;
}
*Value = (Entry + Op->Offset) & Mask;
}
return true;
} else {
@@ -608,6 +649,7 @@ void *X86JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IRLi
}
void *GuestEntry = getCurr<void*>();
CursorEntry = getSize();
this->IR = IR;
if (CTX->GetGdbServerStatus()) {
@@ -622,7 +664,7 @@ void *X86JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IRLi
cmp(dword [rax + (offsetof(FEXCore::Context::Context, Config.RunningMode))], 0);
je(RunBlock);
// Else we need to pause now
mov(rax, ThreadSharedData.Dispatcher->ThreadPauseHandlerAddress);
mov(rax, Dispatcher->ThreadPauseHandlerAddress);
jmp(rax);
ud2();
@@ -760,7 +802,9 @@ void *X86JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IRLi
if (DebugData) {
DebugData->HostCodeSize = reinterpret_cast<uintptr_t>(GuestExit) - reinterpret_cast<uintptr_t>(GuestEntry);
DebugData->Relocations = &Relocations;
}
return GuestEntry;
}
@@ -772,10 +816,10 @@ uint64_t X86JITCore::ExitFunctionLink(X86JITCore *core, FEXCore::Core::CpuStateF
if (!HostCode) {
Thread->CurrentFrame->State.rip = GuestRip;
return core->ThreadSharedData.Dispatcher->AbsoluteLoopTopAddress;
return core->Dispatcher->AbsoluteLoopTopAddress;
}
auto LinkerAddress = core->ThreadSharedData.Dispatcher->ExitFunctionLinkerAddress;
auto LinkerAddress = core->Dispatcher->ExitFunctionLinkerAddress;
Thread->LookupCache->AddBlockLink(GuestRip, (uintptr_t)record, [record, LinkerAddress]{
// undo the link
record[0] = LinkerAddress;
@@ -785,7 +829,12 @@ uint64_t X86JITCore::ExitFunctionLink(X86JITCore *core, FEXCore::Core::CpuStateF
return HostCode;
}
std::unique_ptr<CPUBackend> CreateX86JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, bool CompileThread) {
return std::make_unique<X86JITCore>(ctx, Thread, AllocateNewCodeBuffer(ctx, CompileThread ? X86JITCore::MAX_CODE_SIZE : X86JITCore::INITIAL_CODE_SIZE), CompileThread);
std::unique_ptr<CPUBackend> CreateX86JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread) {
return std::make_unique<X86JITCore>(ctx, Thread, AllocateNewCodeBuffer(ctx, X86JITCore::INITIAL_CODE_SIZE));
}
void InitializeX86JITSignalHandlers(FEXCore::Context::Context *CTX) {
X86JITCore::InitializeSignalHandlers(CTX);
}
}
+77 -20
View File
@@ -8,6 +8,7 @@ $end_info$
#include "Interface/Core/BlockSamplingData.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Core/ObjectCache/Relocations.h"
#define XBYAK64
#include <xbyak/xbyak.h>
@@ -58,8 +59,7 @@ class X86JITCore final : public CPUBackend, public Xbyak::CodeGenerator {
public:
explicit X86JITCore(FEXCore::Context::Context *ctx,
FEXCore::Core::InternalThreadState *Thread,
CodeBuffer Buffer,
bool CompileThread);
CodeBuffer Buffer);
~X86JITCore() override;
[[nodiscard]] std::string GetName() override { return "JIT"; }
@@ -77,13 +77,77 @@ public:
static constexpr size_t INITIAL_CODE_SIZE = 1024 * 1024 * 16;
static constexpr size_t MAX_CODE_SIZE = 1024 * 1024 * 256;
void CopyNecessaryDataForCompileThread(CPUBackend *Original) override;
bool IsAddressInJITCode(uint64_t Address, bool IncludeDispatcher = true, bool IncludeCompileService = true) const override {
return Dispatcher->IsAddressInJITCode(Address, IncludeDispatcher, IncludeCompileService);
bool IsAddressInJITCode(uint64_t Address, bool IncludeDispatcher = true) const override {
return Dispatcher->IsAddressInJITCode(Address, IncludeDispatcher);
}
static void InitializeSignalHandlers(FEXCore::Context::Context *CTX);
void ClearRelocations() override { Relocations.clear(); }
private:
/**
* @name Relocations
* @{ */
uint64_t GetNamedSymbolLiteral(FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op);
void LoadConstantWithPadding(Xbyak::Reg Reg, uint64_t Constant);
/**
* @brief A literal pair relocation object for named symbol literals
*/
struct NamedSymbolLiteralPair {
Label Offset;
Relocation MoveABI{};
};
/**
* @brief Inserts a thunk relocation
*
* @param Reg - The GPR to move the thunk handler in to
* @param Sum - The hash of the thunk
*/
void InsertNamedThunkRelocation(Xbyak::Reg Reg, const IR::SHA256Sum &Sum);
/**
* @brief Inserts a guest GPR move relocation
*
* @param Reg - The GPR to move the guest RIP in to
* @param Constant - The guest RIP that will be relocated
*/
void InsertGuestRIPMove(Xbyak::Reg Reg, uint64_t Constant);
/**
* @brief Inserts a named symbol as a literal in memory
*
* Need to use `PlaceNamedSymbolLiteral` with the return value to place the literal in the desired location
*
* @param Op The named symbol to place
*
* @return A temporary `NamedSymbolLiteralPair`
*/
NamedSymbolLiteralPair InsertNamedSymbolLiteral(FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op);
/**
* @brief Place the named symbol literal relocation in memory
*
* @param Lit - Which literal to place
*/
void PlaceNamedSymbolLiteral(NamedSymbolLiteralPair &Lit);
std::vector<FEXCore::CPU::Relocation> Relocations;
///< Relocation code loading
bool ApplyRelocations(uint64_t GuestEntry, uint64_t CodeEntry, uint64_t CursorEntry, size_t NumRelocations, const char* EntryRelocations);
/**
* @brief Current guest RIP entrypoint
*/
uint64_t CursorEntry{};
/** @} */
Label* PendingTargetLabel{};
FEXCore::Context::Context *CTX;
FEXCore::Core::InternalThreadState *ThreadState;
@@ -152,6 +216,9 @@ private:
static uint64_t ExitFunctionLink(X86JITCore* code, FEXCore::Core::CpuStateFrame *Frame, uint64_t *record);
// This is purely a debugging aid for developers to see if they are in JIT code space when inspecting raw memory
void EmitDetectionString();
// This is the initial code buffer that we will fall back to
// In a program without signals and code clearing, we will typically
// only have this code buffer
@@ -164,17 +231,6 @@ private:
// This is the current code buffer that we are tracking
CodeBuffer *CurrentCodeBuffer{};
struct CompilerSharedData {
uint64_t SignalHandlerReturnAddress{};
uint64_t UnimplementedInstructionAddress{};
uint64_t OverflowExceptionInstructionAddress{};
uint32_t *SignalHandlerRefCounterPtr{};
FEXCore::CPU::Dispatcher *Dispatcher{};
};
CompilerSharedData ThreadSharedData;
uint32_t SpillSlots{};
using SetCC = void (X86JITCore::*)(const Operand& op);
using CMovCC = void (X86JITCore::*)(const Reg& reg, const Operand& op);
@@ -276,7 +332,6 @@ private:
///< Branch ops
DEF_OP(GuestCallDirect);
DEF_OP(GuestCallIndirect);
DEF_OP(GuestReturn);
DEF_OP(SignalReturn);
DEF_OP(CallbackReturn);
DEF_OP(ExitFunction);
@@ -285,7 +340,7 @@ private:
DEF_OP(Syscall);
DEF_OP(Thunk);
DEF_OP(ValidateCode);
DEF_OP(RemoveCodeEntry);
DEF_OP(RemoveThreadCodeEntry);
DEF_OP(CPUID);
///< Conversion ops
@@ -329,6 +384,8 @@ private:
DEF_OP(GetRoundingMode);
DEF_OP(SetRoundingMode);
DEF_OP(ProcessorID);
DEF_OP(RDRAND);
DEF_OP(Yield);
///< Move ops
DEF_OP(ExtractElementPair);
@@ -338,8 +395,6 @@ private:
///< Vector ops
DEF_OP(VectorZero);
DEF_OP(VectorImm);
DEF_OP(CreateVector2);
DEF_OP(CreateVector4);
DEF_OP(SplatVector);
DEF_OP(VMov);
DEF_OP(VAnd);
@@ -426,6 +481,7 @@ private:
DEF_OP(VSMull2);
DEF_OP(VUABDL);
DEF_OP(VTBL1);
DEF_OP(VRev64);
///< Encryption ops
DEF_OP(AESImc);
@@ -434,6 +490,7 @@ private:
DEF_OP(AESDec);
DEF_OP(AESDecLast);
DEF_OP(AESKeyGenAssist);
DEF_OP(CRC32);
#undef DEF_OP
};
@@ -23,7 +23,7 @@ DEF_OP(LoadContext) {
auto Op = IROp->C<IR::IROp_LoadContext>();
uint8_t OpSize = IROp->Size;
if (Op->Class.Val == 0) {
if (Op->Class == IR::GPRClass) {
switch (OpSize) {
case 1: {
movzx(GetDst<RA_32>(Node), byte [STATE + Op->Offset]);
@@ -84,23 +84,23 @@ DEF_OP(StoreContext) {
auto Op = IROp->C<IR::IROp_StoreContext>();
uint8_t OpSize = IROp->Size;
if (Op->Class.Val == 0) {
if (Op->Class == IR::GPRClass) {
switch (OpSize) {
case 1: {
mov(byte [STATE + Op->Offset], GetSrc<RA_8>(Op->Header.Args[0].ID()));
mov(byte [STATE + Op->Offset], GetSrc<RA_8>(Op->Value.ID()));
}
break;
case 2: {
mov(word [STATE + Op->Offset], GetSrc<RA_16>(Op->Header.Args[0].ID()));
mov(word [STATE + Op->Offset], GetSrc<RA_16>(Op->Value.ID()));
}
break;
case 4: {
mov(dword [STATE + Op->Offset], GetSrc<RA_32>(Op->Header.Args[0].ID()));
mov(dword [STATE + Op->Offset], GetSrc<RA_32>(Op->Value.ID()));
}
break;
case 8: {
mov(qword [STATE + Op->Offset], GetSrc<RA_64>(Op->Header.Args[0].ID()));
mov(qword [STATE + Op->Offset], GetSrc<RA_64>(Op->Value.ID()));
}
break;
case 16:
@@ -112,27 +112,27 @@ DEF_OP(StoreContext) {
else {
switch (OpSize) {
case 1: {
pextrb(byte [STATE + Op->Offset], GetSrc(Op->Header.Args[0].ID()), 0);
pextrb(byte [STATE + Op->Offset], GetSrc(Op->Value.ID()), 0);
}
break;
case 2: {
pextrw(word [STATE + Op->Offset], GetSrc(Op->Header.Args[0].ID()), 0);
pextrw(word [STATE + Op->Offset], GetSrc(Op->Value.ID()), 0);
}
break;
case 4: {
vmovd(dword [STATE + Op->Offset], GetSrc(Op->Header.Args[0].ID()));
vmovd(dword [STATE + Op->Offset], GetSrc(Op->Value.ID()));
}
break;
case 8: {
vmovq(qword [STATE + Op->Offset], GetSrc(Op->Header.Args[0].ID()));
vmovq(qword [STATE + Op->Offset], GetSrc(Op->Value.ID()));
}
break;
case 16: {
if (Op->Offset % 16 == 0)
movaps(xword [STATE + Op->Offset], GetSrc(Op->Header.Args[0].ID()));
movaps(xword [STATE + Op->Offset], GetSrc(Op->Value.ID()));
else
movups(xword [STATE + Op->Offset], GetSrc(Op->Header.Args[0].ID()));
movups(xword [STATE + Op->Offset], GetSrc(Op->Value.ID()));
}
break;
default: LOGMAN_MSG_A_FMT("Unhandled StoreContext size: {}", OpSize);
@@ -143,9 +143,9 @@ DEF_OP(StoreContext) {
DEF_OP(LoadContextIndexed) {
auto Op = IROp->C<IR::IROp_LoadContextIndexed>();
size_t size = IROp->Size;
Reg index = GetSrc<RA_64>(Op->Header.Args[0].ID());
Reg index = GetSrc<RA_64>(Op->Index.ID());
if (Op->Class.Val == 0) {
if (Op->Class == IR::GPRClass) {
switch (Op->Stride) {
case 1:
case 2:
@@ -245,11 +245,11 @@ DEF_OP(LoadContextIndexed) {
DEF_OP(StoreContextIndexed) {
auto Op = IROp->C<IR::IROp_StoreContextIndexed>();
Reg index = GetSrc<RA_64>(Op->Header.Args[1].ID());
Reg index = GetSrc<RA_64>(Op->Index.ID());
size_t size = IROp->Size;
if (Op->Class.Val == 0) {
auto value = GetSrc<RA_64>(Op->Header.Args[0].ID());
if (Op->Class == IR::GPRClass) {
auto value = GetSrc<RA_64>(Op->Value.ID());
lea(rax, dword [STATE + Op->BaseOffset]);
switch (Op->Stride) {
@@ -269,7 +269,7 @@ DEF_OP(StoreContextIndexed) {
}
}
else {
auto value = GetSrc(Op->Header.Args[0].ID());
auto value = GetSrc(Op->Value.ID());
switch (Op->Stride) {
case 1:
case 2:
@@ -339,19 +339,19 @@ DEF_OP(SpillRegister) {
if (Op->Class == FEXCore::IR::GPRClass) {
switch (OpSize) {
case 1: {
mov(byte [rsp + SlotOffset], GetSrc<RA_8>(Op->Header.Args[0].ID()));
mov(byte [rsp + SlotOffset], GetSrc<RA_8>(Op->Value.ID()));
break;
}
case 2: {
mov(word [rsp + SlotOffset], GetSrc<RA_16>(Op->Header.Args[0].ID()));
mov(word [rsp + SlotOffset], GetSrc<RA_16>(Op->Value.ID()));
break;
}
case 4: {
mov(dword [rsp + SlotOffset], GetSrc<RA_32>(Op->Header.Args[0].ID()));
mov(dword [rsp + SlotOffset], GetSrc<RA_32>(Op->Value.ID()));
break;
}
case 8: {
mov(qword [rsp + SlotOffset], GetSrc<RA_64>(Op->Header.Args[0].ID()));
mov(qword [rsp + SlotOffset], GetSrc<RA_64>(Op->Value.ID()));
break;
}
default: LOGMAN_MSG_A_FMT("Unhandled SpillRegister size: {}", OpSize);
@@ -359,15 +359,15 @@ DEF_OP(SpillRegister) {
} else if (Op->Class == FEXCore::IR::FPRClass) {
switch (OpSize) {
case 4: {
movss(dword [rsp + SlotOffset], GetSrc(Op->Header.Args[0].ID()));
movss(dword [rsp + SlotOffset], GetSrc(Op->Value.ID()));
break;
}
case 8: {
movsd(qword [rsp + SlotOffset], GetSrc(Op->Header.Args[0].ID()));
movsd(qword [rsp + SlotOffset], GetSrc(Op->Value.ID()));
break;
}
case 16: {
movaps(xword [rsp + SlotOffset], GetSrc(Op->Header.Args[0].ID()));
movaps(xword [rsp + SlotOffset], GetSrc(Op->Value.ID()));
break;
}
default: LOGMAN_MSG_A_FMT("Unhandled SpillRegister size: {}", OpSize);
@@ -435,7 +435,7 @@ DEF_OP(LoadFlag) {
DEF_OP(StoreFlag) {
auto Op = IROp->C<IR::IROp_StoreFlag>();
mov (rax, GetSrc<RA_64>(Op->Header.Args[0].ID()));
mov (rax, GetSrc<RA_64>(Op->Value.ID()));
mov(byte [STATE + (offsetof(FEXCore::Core::CPUState, flags[0]) + Op->Flag)], al);
}
@@ -469,7 +469,7 @@ DEF_OP(LoadMem) {
auto MemPtr = GenerateModRM(MemReg, Op->Offset, Op->OffsetType, Op->OffsetScale);
if (Op->Class.Val == 0) {
if (Op->Class == IR::GPRClass) {
auto Dst = GetDst<RA_64>(Node);
switch (IROp->Size) {
@@ -537,19 +537,19 @@ DEF_OP(StoreMem) {
auto MemPtr = GenerateModRM(MemReg, Op->Offset, Op->OffsetType, Op->OffsetScale);
if (Op->Class.Val == 0) {
if (Op->Class == IR::GPRClass) {
switch (IROp->Size) {
case 1:
mov(byte [MemPtr], GetSrc<RA_8>(Op->Header.Args[1].ID()));
mov(byte [MemPtr], GetSrc<RA_8>(Op->Value.ID()));
break;
case 2:
mov(word [MemPtr], GetSrc<RA_16>(Op->Header.Args[1].ID()));
mov(word [MemPtr], GetSrc<RA_16>(Op->Value.ID()));
break;
case 4:
mov(dword [MemPtr], GetSrc<RA_32>(Op->Header.Args[1].ID()));
mov(dword [MemPtr], GetSrc<RA_32>(Op->Value.ID()));
break;
case 8:
mov(qword [MemPtr], GetSrc<RA_64>(Op->Header.Args[1].ID()));
mov(qword [MemPtr], GetSrc<RA_64>(Op->Value.ID()));
break;
default: LOGMAN_MSG_A_FMT("Unhandled StoreMem size: {}", IROp->Size);
}
@@ -557,22 +557,22 @@ DEF_OP(StoreMem) {
else {
switch (IROp->Size) {
case 1:
pextrb(byte [MemPtr], GetSrc(Op->Header.Args[1].ID()), 0);
pextrb(byte [MemPtr], GetSrc(Op->Value.ID()), 0);
break;
case 2:
pextrw(word [MemPtr], GetSrc(Op->Header.Args[1].ID()), 0);
pextrw(word [MemPtr], GetSrc(Op->Value.ID()), 0);
break;
case 4:
vmovd(dword [MemPtr], GetSrc(Op->Header.Args[1].ID()));
vmovd(dword [MemPtr], GetSrc(Op->Value.ID()));
break;
case 8:
vmovq(qword [MemPtr], GetSrc(Op->Header.Args[1].ID()));
vmovq(qword [MemPtr], GetSrc(Op->Value.ID()));
break;
case 16:
if (IROp->Size == Op->Align)
movups(xword [MemPtr], GetSrc(Op->Header.Args[1].ID()));
movups(xword [MemPtr], GetSrc(Op->Value.ID()));
else
movups(xword [MemPtr], GetSrc(Op->Header.Args[1].ID()));
movups(xword [MemPtr], GetSrc(Op->Value.ID()));
break;
default: LOGMAN_MSG_A_FMT("Unhandled StoreMem size: {}", IROp->Size);
}
+30 -24
View File
@@ -18,14 +18,6 @@ $end_info$
#include <xbyak/xbyak.h>
namespace FEXCore::CPU {
static void PrintValue(uint64_t Value) {
LogMan::Msg::DFmt("Value: 0x{:x}", Value);
}
static void PrintVectorValue(uint64_t Value, uint64_t ValueUpper) {
LogMan::Msg::DFmt("Value: 0x{:016x}'{:016x}", ValueUpper, Value);
}
#define DEF_OP(x) void X86JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(Fence) {
@@ -53,8 +45,7 @@ DEF_OP(Break) {
break;
case FEXCore::IR::Break_Overflow: // overflow
// Need to be outside of JIT cache space to ensure cache clearing correctness
mov(TMP1, ThreadSharedData.Dispatcher->OverflowExceptionInstructionAddress);
jmp(TMP1);
jmp(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.OverflowExceptionHandler)]);
break;
case FEXCore::IR::Break_Halt: { // HLT
// Time to quit
@@ -62,8 +53,7 @@ DEF_OP(Break) {
mov(rsp, qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, ReturningStackLocation)]);
// Now we need to jump to the thread stop handler
mov(TMP1, ThreadSharedData.Dispatcher->ThreadStopHandlerAddress);
jmp(TMP1);
jmp(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.ThreadStopHandler)]);
break;
}
case FEXCore::IR::Break_Interrupt3: // INT3
@@ -75,8 +65,7 @@ DEF_OP(Break) {
}
// This jump target needs to be a constant offset here
mov(TMP1, ThreadSharedData.Dispatcher->ThreadPauseHandlerAddress);
jmp(TMP1);
jmp(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.ThreadPauseHandler)]);
}
else {
// If we don't have a gdb server attached then....crash?
@@ -84,8 +73,7 @@ DEF_OP(Break) {
mov(rsp, qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, ReturningStackLocation)]);
// Now we need to jump to the thread stop handler
mov(TMP1, ThreadSharedData.Dispatcher->ThreadStopHandlerAddress);
jmp(TMP1);
jmp(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.ThreadStopHandler)]);
}
break;
}
@@ -96,9 +84,7 @@ DEF_OP(Break) {
}
// Need to be outside of JIT cache space to ensure cache clearing correctness
mov(TMP1, ThreadSharedData.Dispatcher->UnimplementedInstructionAddress);
jmp(TMP1);
jmp(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.UnimplementedInstructionHandler)]);
break;
}
default: LOGMAN_MSG_A_FMT("Unknown Break reason: {}", Op->Reason);
@@ -144,18 +130,15 @@ DEF_OP(Print) {
PushRegs();
if (IsGPR(Op->Header.Args[0].ID())) {
mov (rdi, GetSrc<RA_64>(Op->Header.Args[0].ID()));
mov(rax, reinterpret_cast<uintptr_t>(PrintValue));
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.PrintValue)]);
}
else {
pextrq(rdi, GetSrc(Op->Header.Args[0].ID()), 0);
pextrq(rsi, GetSrc(Op->Header.Args[0].ID()), 1);
mov(rax, reinterpret_cast<uintptr_t>(PrintVectorValue));
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.PrintVectorValue)]);
}
call(rax);
PopRegs();
}
@@ -166,6 +149,27 @@ DEF_OP(ProcessorID) {
mov (GetDst<RA_32>(Node), ecx);
}
DEF_OP(RDRAND) {
auto Op = IROp->C<IR::IROp_RDRAND>();
auto Dst = GetSrcPair<RA_64>(Node);
if (Op->GetReseeded) {
rdrand(Dst.first);
}
else {
rdseed(Dst.first);
}
// In the case of RDRAND or RDSEED returning a valid number then CF = 1, else 0
mov (Dst.second, 0);
setc(Dst.second.cvt8());
}
DEF_OP(Yield) {
pause();
}
#undef DEF_OP
void X86JITCore::RegisterMiscHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &X86JITCore::Op_##x
@@ -183,6 +187,8 @@ void X86JITCore::RegisterMiscHandlers() {
REGISTER_OP(SETROUNDINGMODE, SetRoundingMode);
REGISTER_OP(INVALIDATEFLAGS, NoOp);
REGISTER_OP(PROCESSORID, ProcessorID);
REGISTER_OP(RDRAND, RDRAND);
REGISTER_OP(YIELD, Yield);
#undef REGISTER_OP
}
}
@@ -42,7 +42,7 @@ DEF_OP(CreateElementPair) {
Xbyak::Reg RegSecond;
Xbyak::Reg RegTmp;
switch (Op->Header.Size) {
switch (IROp->ElementSize) {
case 4: {
Dst = GetSrcPair<RA_32>(Node);
RegFirst = GetSrc<RA_32>(Op->Header.Args[0].ID());
@@ -67,14 +67,6 @@ DEF_OP(VectorImm) {
}
}
DEF_OP(CreateVector2) {
LOGMAN_MSG_A_FMT("Unimplemented");
}
DEF_OP(CreateVector4) {
LOGMAN_MSG_A_FMT("Unimplemented");
}
DEF_OP(SplatVector) {
auto Op = IROp->C<IR::IROp_SplatVector2>();
uint8_t OpSize = IROp->Size;
@@ -1545,7 +1537,7 @@ DEF_OP(VInsScalarElement) {
DEF_OP(VExtractElement) {
auto Op = IROp->C<IR::IROp_VExtractElement>();
switch (Op->Header.ElementSize) {
switch (Op->Header.Size) {
case 1: {
pextrb(eax, GetSrc(Op->Header.Args[0].ID()), Op->Index);
pinsrb(GetDst(Node), eax, 0);
@@ -1566,7 +1558,7 @@ DEF_OP(VExtractElement) {
pinsrq(GetDst(Node), rax, 0);
break;
}
default: LOGMAN_MSG_A_FMT("Unknown Element Size: {}", Op->Header.ElementSize); break;
default: LOGMAN_MSG_A_FMT("Unknown Element Size: {}", Op->Header.Size); break;
}
}
@@ -2151,13 +2143,79 @@ DEF_OP(VTBL1) {
}
}
DEF_OP(VRev64) {
auto Op = IROp->C<IR::IROp_VDupElement>();
switch (Op->Header.ElementSize) {
case 1: {
mov(rax, 0x00'01'02'03'04'05'06'07); // Lower
vmovq(xmm15, rax);
if (IROp->Size == 16) {
// Full 8bit byteswap in each 64-bit element
mov(rcx, 0x08'09'0A'0B'0C'0D'0E'0F); // Upper
pinsrq(xmm15, rcx, 1);
}
else {
// 8byte, upper bits get zero
// Full 8bit byteswap in each 64-bit element
mov(rcx, 0x80'80'80'80'80'80'80'80); // Upper
pinsrq(xmm15, rcx, 1);
}
vpshufb(GetDst(Node), GetSrc(Op->Header.Args[0].ID()), xmm15);
break;
}
case 2: {
// Full 16-bit byteswap in each 64-bit element
mov(rax, 0x01'00'03'02'05'04'07'06); // Lower
vmovq(xmm15, rax);
if (IROp->Size == 16) {
mov(rcx, 0x09'08'0B'0A'0D'0C'0F'0E); // Upper
pinsrq(xmm15, rcx, 1);
}
else {
// 8byte, upper bits get zero
// Full 8bit byteswap in each 64-bit element
mov(rcx, 0x80'80'80'80'80'80'80'80); // Upper
pinsrq(xmm15, rcx, 1);
}
vpshufb(GetDst(Node), GetSrc(Op->Header.Args[0].ID()), xmm15);
break;
}
case 4: {
if (IROp->Size == 16) {
vpshufd(GetDst(Node),
GetSrc(Op->Header.Args[0].ID()),
(0b11 << 0) |
(0b10 << 2) |
(0b01 << 4) |
(0b00 << 6));
}
else {
vpshufd(GetDst(Node),
GetSrc(Op->Header.Args[0].ID()),
(0b01 << 0) |
(0b00 << 2) |
(0b11 << 4) | // Last two don't matter, will be overwritten with zero
(0b11 << 6));
// Zero upper 64-bits
mov(rcx, 0);
pinsrq(GetDst(Node), rcx, 1);
}
break;
}
default: LOGMAN_MSG_A_FMT("Unknown Element Size: {}", Op->Header.ElementSize); break;
}
}
#undef DEF_OP
void X86JITCore::RegisterVectorHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &X86JITCore::Op_##x
REGISTER_OP(VECTORZERO, VectorZero);
REGISTER_OP(VECTORIMM, VectorImm);
REGISTER_OP(CREATEVECTOR2, CreateVector2);
REGISTER_OP(CREATEVECTOR4, CreateVector4);
REGISTER_OP(SPLATVECTOR2, SplatVector);
REGISTER_OP(SPLATVECTOR4, SplatVector);
REGISTER_OP(VMOV, VMov);
@@ -2246,6 +2304,7 @@ void X86JITCore::RegisterVectorHandlers() {
REGISTER_OP(VSMULL2, VSMull2);
REGISTER_OP(VUABDL, VUABDL);
REGISTER_OP(VTBL1, VTBL1);
REGISTER_OP(VREV64, VRev64);
#undef REGISTER_OP
}
}
@@ -0,0 +1,138 @@
/*
$info$
tags: backend|x86-64
desc: relocation logic of the x86-64 splatter backend
$end_info$
*/
#include "Interface/Core/JIT/x86_64/JITClass.h"
#include "Interface/HLE/Thunks/Thunks.h"
namespace FEXCore::CPU {
uint64_t X86JITCore::GetNamedSymbolLiteral(FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op) {
switch (Op) {
case FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol::SYMBOL_LITERAL_EXITFUNCTION_LINKER:
return Dispatcher->ExitFunctionLinkerAddress;
break;
default:
ERROR_AND_DIE_FMT("Unknown named symbol literal: {}", static_cast<uint32_t>(Op));
break;
}
return ~0ULL;
}
void X86JITCore::LoadConstantWithPadding(Xbyak::Reg Reg, uint64_t Constant) {
// The maximum size a move constant can be in bytes
// Need to NOP pad to this size to ensure backpatching is always the same size
// Calculated as:
// [Rex]
// [Mov op]
// [8 byte constant]
//
// All other move types are smaller than this. xbyak will use a NOP slide which is quite quick
constexpr static size_t MAX_MOVE_SIZE = 10;
auto StartingOffset = getSize();
mov(Reg, Constant);
auto MoveSize = getSize() - StartingOffset;
auto NOPPadSize = MAX_MOVE_SIZE - MoveSize;
nop(NOPPadSize);
}
X86JITCore::NamedSymbolLiteralPair X86JITCore::InsertNamedSymbolLiteral(FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op) {
NamedSymbolLiteralPair Lit {
.MoveABI = {
.NamedSymbolLiteral = {
.Header = {
.Type = FEXCore::CPU::RelocationTypes::RELOC_NAMED_SYMBOL_LITERAL,
},
.Symbol = Op,
.Offset = 0,
},
},
};
return Lit;
}
void X86JITCore::PlaceNamedSymbolLiteral(NamedSymbolLiteralPair &Lit) {
// Offset is the offset from the entrypoint of the block
auto CurrentCursor = getSize();
Lit.MoveABI.NamedSymbolLiteral.Offset = CurrentCursor - CursorEntry;
uint64_t Pointer = GetNamedSymbolLiteral(Lit.MoveABI.NamedSymbolLiteral.Symbol);
L(Lit.Offset);
dq(Pointer);
Relocations.emplace_back(Lit.MoveABI);
}
void X86JITCore::InsertGuestRIPMove(Xbyak::Reg Reg, uint64_t Constant) {
Relocation MoveABI{};
MoveABI.GuestRIPMove.Header.Type = FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_MOVE;
// Offset is the offset from the entrypoint of the block
auto CurrentCursor = getSize();
MoveABI.GuestRIPMove.Offset = CurrentCursor - CursorEntry;
MoveABI.GuestRIPMove.GuestRIP = Constant;
MoveABI.GuestRIPMove.RegisterIndex = Reg.getIdx();
if (CTX->Config.CacheObjectCodeCompilation()) {
LoadConstantWithPadding(Reg, Constant);
}
else {
mov(Reg, Constant);
}
Relocations.emplace_back(MoveABI);
}
bool X86JITCore::ApplyRelocations(uint64_t GuestEntry, uint64_t CodeEntry, uint64_t CursorEntry, size_t NumRelocations, const char* EntryRelocations) {
size_t DataIndex{};
for (size_t j = 0; j < NumRelocations; ++j) {
const FEXCore::CPU::Relocation *Reloc = reinterpret_cast<const FEXCore::CPU::Relocation *>(&EntryRelocations[DataIndex]);
LOGMAN_THROW_A_FMT((DataIndex % alignof(Relocation)) == 0, "Alignment of relocation wasn't adhered to");
switch (Reloc->Header.Type) {
case FEXCore::CPU::RelocationTypes::RELOC_NAMED_SYMBOL_LITERAL: {
uint64_t Pointer = GetNamedSymbolLiteral(Reloc->NamedSymbolLiteral.Symbol);
// Relocation occurs at the cursorEntry + offset relative to that cursor.
setSize(CursorEntry + Reloc->NamedSymbolLiteral.Offset);
// Place the pointer
dq(Pointer);
DataIndex += sizeof(Reloc->NamedSymbolLiteral);
break;
}
case FEXCore::CPU::RelocationTypes::RELOC_NAMED_THUNK_MOVE: {
uint64_t Pointer = reinterpret_cast<uint64_t>(CTX->ThunkHandler->LookupThunk(Reloc->NamedThunkMove.Symbol));
if (Pointer == ~0ULL) {
return false;
}
// Relocation occurs at the cursorEntry + offset relative to that cursor.
setSize(CursorEntry + Reloc->NamedThunkMove.Offset);
LoadConstantWithPadding(Xbyak::Reg64(Reloc->NamedThunkMove.RegisterIndex), Pointer);
DataIndex += sizeof(Reloc->NamedThunkMove);
break;
}
case FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_MOVE:
// XXX: Reenable once the JIT Object Cache is upstream
// XXX: Should spin the relocation list, create a list of guest RIP moves, and ask for them all once, reduces lock contention.
uint64_t Pointer = ~0ULL; // EmitterCTX->JITObjectCache->FindRelocatedRIP(Reloc->GuestRIPMove.GuestRIP);
if (Pointer == ~0ULL) {
return false;
}
// Relocation occurs at the cursorEntry + offset relative to that cursor.
setSize(CursorEntry + Reloc->GuestRIPMove.Offset);
LoadConstantWithPadding(Xbyak::Reg64(Reloc->GuestRIPMove.RegisterIndex), Pointer);
DataIndex += sizeof(Reloc->GuestRIPMove);
break;
}
}
return true;
}
}
@@ -61,6 +61,7 @@ void LookupCache::HintUsedRange(uint64_t Address, uint64_t Size) {
}
void LookupCache::ClearL2Cache() {
std::lock_guard<std::recursive_mutex> lk(WriteLock);
// Clear out the page memory
madvise(reinterpret_cast<void*>(PagePointer), ctx->Config.VirtualMemSize / 4096 * 8, MADV_DONTNEED);
madvise(reinterpret_cast<void*>(PageMemory), CODE_SIZE, MADV_DONTNEED);
@@ -68,6 +69,8 @@ void LookupCache::ClearL2Cache() {
}
void LookupCache::ClearCache() {
std::lock_guard<std::recursive_mutex> lk(WriteLock);
// Clear L1
madvise(reinterpret_cast<void*>(L1Pointer), L1_SIZE, MADV_DONTNEED);
// Clear L2
+65 -44
View File
@@ -7,6 +7,7 @@
#include <stddef.h>
#include <utility>
#include <vector>
#include <mutex>
namespace FEXCore {
namespace Context {
@@ -28,31 +29,63 @@ public:
uintptr_t End() { return 0; }
uintptr_t FindBlock(uint64_t Address) {
auto HostCode = FindCodePointerForAddress(Address);
if (HostCode) {
return HostCode;
} else {
auto HostCode = BlockList.find(Address);
// Try L1, no lock needed
auto &L1Entry = reinterpret_cast<LookupCacheEntry*>(L1Pointer)[Address & L1_ENTRIES_MASK];
if (L1Entry.GuestCode == Address) {
return L1Entry.HostCode;
}
if (HostCode != BlockList.end()) {
CacheBlockMapping(Address, HostCode->second);
return HostCode->second;
} else {
return 0;
// L2 and L3 need to be locked
std::lock_guard<std::recursive_mutex> lk(WriteLock);
// Try L2
const auto PageIndex = (Address & (VirtualMemSize -1)) >> 12;
const auto PageOffset = Address & (0x0FFF);
const auto Pointers = reinterpret_cast<uintptr_t*>(PagePointer);
auto LocalPagePointer = Pointers[PageIndex];
// Do we a page pointer for this address?
if (LocalPagePointer) {
// Find there pointer for the address in the blocks
auto BlockPointers = reinterpret_cast<LookupCacheEntry*>(LocalPagePointer);
if (BlockPointers[PageOffset].GuestCode == Address)
{
L1Entry.GuestCode = Address;
L1Entry.HostCode = BlockPointers[PageOffset].HostCode;
return L1Entry.HostCode;
}
}
// Try L3
auto HostCode = BlockList.find(Address);
if (HostCode != BlockList.end()) {
CacheBlockMapping(Address, HostCode->second);
return HostCode->second;
}
// Failed to find
return 0;
}
std::map<uint64_t, std::vector<uint64_t>> CodePages;
void AddBlockMapping(uint64_t Address, void *HostCode, uint64_t Start, uint64_t Length) {
// Returns true if new pages are marked as containing code
bool AddBlockMapping(uint64_t Address, void *HostCode, uint64_t Start, uint64_t Length) {
std::lock_guard<std::recursive_mutex> lk(WriteLock);
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
auto InsertPoint =
#endif
BlockList.emplace(Address, (uintptr_t)HostCode);
LOGMAN_THROW_A_FMT(InsertPoint.second == true, "Dupplicate block mapping added");
bool rv = false;
for (auto CurrentPage = Start >> 12, EndPage = (Start + Length) >> 12; CurrentPage <= EndPage; CurrentPage++) {
rv |= CodePages[CurrentPage].size() == 0;
CodePages[CurrentPage].push_back(Address);
}
@@ -61,10 +94,14 @@ public:
auto &L1Entry = reinterpret_cast<LookupCacheEntry*>(L1Pointer)[Address & L1_ENTRIES_MASK];
L1Entry.GuestCode = Address;
L1Entry.HostCode = (uintptr_t)HostCode;
return rv;
}
void Erase(uint64_t Address) {
std::lock_guard<std::recursive_mutex> lk(WriteLock);
// Sever any links to this block
auto lower = BlockLinks.lower_bound({Address, 0});
auto upper = BlockLinks.upper_bound({Address, UINTPTR_MAX});
@@ -78,7 +115,10 @@ public:
// Do L1
auto &L1Entry = reinterpret_cast<LookupCacheEntry*>(L1Pointer)[Address & L1_ENTRIES_MASK];
if (L1Entry.GuestCode == Address) {
L1Entry.GuestCode = L1Entry.HostCode = 0;
L1Entry.GuestCode = 0;
// Leave L1Entry.HostCode as is, so that concurrent lookups won't read a null pointer
// This is a soft guarantee for cross thread invalidation, as atomics are not used
// and it hasn't been thoroughly tested
}
// Do full map
@@ -101,6 +141,8 @@ public:
void AddBlockLink(uint64_t GuestDestination, uintptr_t HostLink, const std::function<void()> &delinker) {
std::lock_guard<std::recursive_mutex> lk(WriteLock);
BlockLinks.insert({{GuestDestination, HostLink}, delinker});
}
@@ -116,8 +158,19 @@ public:
constexpr static size_t L1_ENTRIES = 1 * 1024 * 1024; // Must be a power of 2
constexpr static size_t L1_ENTRIES_MASK = L1_ENTRIES - 1;
// This needs to be taken before reads or writes to L2, L3, CodePages, Thread::LocalIRCache,
// and before writes to L1. Concurrent access from a thread that this LookupCache doesn't belong to
// may only happen during cross thread invalidation (::Erase).
// All other operations must be done from the owning thread.
// Some care is taken so that L1 lookups can be done without locks, and even tearing is unlikely to lead to a crash.
// This approach has not been fully vetted yet.
// Also note that L1 lookups might be inlined in the JIT Dispatcher and/or block ends.
std::recursive_mutex WriteLock;
private:
void CacheBlockMapping(uint64_t Address, uintptr_t HostCode) {
std::lock_guard<std::recursive_mutex> lk(WriteLock);
// Do L1
auto &L1Entry = reinterpret_cast<LookupCacheEntry*>(L1Pointer)[Address & L1_ENTRIES_MASK];
L1Entry.GuestCode = Address;
@@ -167,38 +220,6 @@ private:
return PageMemory + NewBase;
}
uintptr_t FindCodePointerForAddress(uint64_t Address) {
// Do L1
auto &L1Entry = reinterpret_cast<LookupCacheEntry*>(L1Pointer)[Address & L1_ENTRIES_MASK];
if (L1Entry.GuestCode == Address) {
return L1Entry.HostCode;
}
auto FullAddress = Address;
Address = Address & (VirtualMemSize -1);
uint64_t PageOffset = Address & (0x0FFF);
Address >>= 12;
uintptr_t *Pointers = reinterpret_cast<uintptr_t*>(PagePointer);
uint64_t LocalPagePointer = Pointers[Address];
if (!LocalPagePointer) {
// We don't have a page pointer for this address
return 0;
}
// Find there pointer for the address in the blocks
auto BlockPointers = reinterpret_cast<LookupCacheEntry*>(LocalPagePointer);
if (BlockPointers[PageOffset].GuestCode == FullAddress)
{
L1Entry.GuestCode = FullAddress;
return L1Entry.HostCode = BlockPointers[PageOffset].HostCode;
}
else
return 0;
}
uintptr_t PagePointer;
uintptr_t PageMemory;
uintptr_t L1Pointer;
@@ -0,0 +1,84 @@
#pragma once
#include <cstdint>
namespace FEXCore::CodeSerialize {
// If any of the config options mismatch on load then the cache won't be used
// Any of these will result in codegen changes
struct CodeObjectSerializationConfig {
// Cookie in the header of the file, isn't part of the config hash
uint64_t Cookie{};
// Instructions per block configuration
int32_t MaxInstPerBlock{};
// Follows CPUID 4000_0001_EAX[3:0]
unsigned Arch : 4;
// Multiblock enabled
bool MultiBlock : 1;
// TSO enabled
bool TSOEnabled : 1;
// ABI local flag unsafe optimization
bool ABILocalFlags : 1;
// ABI no PF unsafe optimization
bool ABINoPF : 1;
// Static register allocation enabled
bool SRA : 1;
// Paranoid TSO mode enabled
bool ParanoidTSO : 1;
// Guest code execution mode (We don't support live mode switch)
bool Is64BitMode : 1;
// SMC checks style
unsigned SMCChecks : 2;
// x87 reduced precision
bool x87ReducedPrecision : 1;
// Padding to remove uninitialized data warning from asan
// Shows remaining amount of bits available for config
unsigned _Pad : 18;
bool operator==(CodeObjectSerializationConfig const &other) const {
return Cookie == other.Cookie &&
MaxInstPerBlock == other.MaxInstPerBlock &&
Arch == other.Arch &&
MultiBlock == other.MultiBlock &&
TSOEnabled == other.TSOEnabled &&
ABILocalFlags == other.ABILocalFlags &&
ABINoPF == other.ABINoPF &&
SRA == other.SRA &&
ParanoidTSO == other.ParanoidTSO &&
Is64BitMode == other.Is64BitMode &&
SMCChecks == other.SMCChecks &&
x87ReducedPrecision == other.x87ReducedPrecision;
}
static uint64_t GetHash(CodeObjectSerializationConfig const &other) {
// For < 64-bits of data just pack directly
// Skip the cookie
uint64_t Hash{};
Hash <<= 32; Hash |= other.MaxInstPerBlock;
Hash <<= 1; Hash |= other.Arch;
Hash <<= 1; Hash |= other.MultiBlock;
Hash <<= 1; Hash |= other.TSOEnabled;
Hash <<= 1; Hash |= other.ABILocalFlags;
Hash <<= 1; Hash |= other.ABINoPF;
Hash <<= 1; Hash |= other.SRA;
Hash <<= 1; Hash |= other.ParanoidTSO;
Hash <<= 1; Hash |= other.Is64BitMode;
Hash <<= 2; Hash |= other.SMCChecks;
Hash <<= 1; Hash |= other.x87ReducedPrecision;
return Hash;
}
};
static_assert(sizeof(CodeObjectSerializationConfig) == 16, "Size changed");
static_assert((sizeof(CodeObjectSerializationConfig) - sizeof(uint64_t)) == 8, "Config size exceeded 64its. Need to change how the hash is generated!");
}
@@ -0,0 +1,127 @@
#include "Interface/Context/Context.h"
#include "Interface/Core/ObjectCache/ObjectCacheService.h"
#include <FEXCore/Config/Config.h>
#include <fcntl.h>
#include <filesystem>
#include <memory>
#include <string>
#include <sys/uio.h>
#include <sys/mman.h>
#include <xxhash.h>
namespace FEXCore::CodeSerialize {
void AsyncJobHandler::AsyncAddNamedRegionJob(uintptr_t Base, uintptr_t Size, uintptr_t Offset, const std::string &filename) {
// This function adds a named region *JOB* to our named region handler
// This needs to be as fast as possible to keep out of the way of the JIT
auto BaseFilename = std::filesystem::path(filename).filename().string();
if (!BaseFilename.empty()) {
// Create a new entry that once set up will be put in to our section object map
auto Entry = std::make_unique<CodeRegionEntry>(
Base,
Size,
Offset,
filename,
NamedRegionHandler->DefaultCodeHeader(Base, Offset)
);
// Lock the job ref counter so we can block anything attempting to use the entry before it is loaded
Entry->NamedJobRefCountMutex.lock();
CodeRegionMapType::iterator EntryIterator;
{
std::unique_lock lk {CodeObjectCacheService->GetEntryMapMutex()};
auto &EntryMap = CodeObjectCacheService->GetEntryMap();
auto it = EntryMap.emplace(Base, std::move(Entry));
if (!it.second) {
// This happens when an application overwrites a previous region without unmapping what was there
// Lock this entry's Named job reference counter.
// Once this passes then we know that this section has been loaded.
it.first->second->NamedJobRefCountMutex.lock();
// Finalize anything the region needs to do first.
CodeObjectCacheService->DoCodeRegionClosure(it.first->second->Base, it.first->second.get());
// munmap the file that was mapped
FEXCore::Allocator::munmap(it.first->second->CodeData, it.first->second->FileSize);
// Remove this entry from the unrelocated map as well
{
std::unique_lock lk2 {CodeObjectCacheService->GetUnrelocatedEntryMapMutex()};
CodeObjectCacheService->GetUnrelocatedEntryMap().erase(it.first->second->EntryHeader.OriginalBase);
}
// Now overwrite the entry in the map
it = EntryMap.insert_or_assign(Base, std::move(Entry));
EntryIterator = it.first;
}
else {
// No overwrite, just insert
EntryIterator = it.first;
}
}
// Now that this entry has been added to the map, we can insert a load job using the entry iterator.
// This allows us to quickly unblock the JIT thread when it is loading multiple regions and have the async thread
// do the loading for us.
//
// Create the async work queue job now so it can load
NamedRegionHandler->AsyncAddNamedRegionWorkItem(BaseFilename, filename, true, EntryIterator);
// Tell the async thread that it has work to do
CodeObjectCacheService->NotifyWork();
}
}
void AsyncJobHandler::AsyncRemoveNamedRegionJob(uintptr_t Base, uintptr_t Size) {
// Removing a named region through the job system
// We need to find the entry that we are deleting first
std::unique_ptr<CodeRegionEntry> EntryPointer;
{
std::unique_lock lk {CodeObjectCacheService->GetEntryMapMutex()};
auto &EntryMap = CodeObjectCacheService->GetEntryMap();
auto it = EntryMap.find(Base);
if (it != EntryMap.end()) {
// Lock the job ref counter since we are erasing it
// Once this passes it will have been loaded
it->second->NamedJobRefCountMutex.lock();
// Take the pointer from the map
EntryPointer = std::move(it->second);
// We can now unmap the file data
FEXCore::Allocator::munmap(EntryPointer->CodeData, EntryPointer->FileSize);
// Remove this from the entry map
EntryMap.erase(it);
// Remove this entry from the unrelocated map as well
{
std::unique_lock lk2 {CodeObjectCacheService->GetUnrelocatedEntryMapMutex()};
CodeObjectCacheService->GetUnrelocatedEntryMap().erase(EntryPointer->EntryHeader.OriginalBase);
}
}
else {
// Tried to remove something that wasn't in our code object tracking
return;
}
// Create the async work queue job now so it can finalize what it needs to do
NamedRegionHandler->AsyncRemoveNamedRegionWorkItem(Base, Size, std::move(EntryPointer));
// Tell the async thread that it has work to do
CodeObjectCacheService->NotifyWork();
}
}
void AsyncJobHandler::AsyncAddSerializationJob(std::unique_ptr<SerializationJobData> Data) {
// XXX: Actually add serialization job
}
}
@@ -0,0 +1,71 @@
#include "Interface/Context/Context.h"
#include "Interface/Core/ObjectCache/ObjectCacheService.h"
#include <FEXCore/Config/Config.h>
namespace FEXCore::CodeSerialize {
NamedRegionObjectHandler::NamedRegionObjectHandler(FEXCore::Context::Context *ctx) {
DefaultSerializationConfig.Cookie = CODE_COOKIE;
// Initialize the Arch from CPUID
uint32_t Arch = ctx->CPUID.RunFunction(0x4000'0001, 0).eax & 0xF;
DefaultSerializationConfig.Arch = Arch;
DefaultSerializationConfig.MaxInstPerBlock = ctx->Config.MaxInstPerBlock;
DefaultSerializationConfig.MultiBlock = ctx->Config.Multiblock;
DefaultSerializationConfig.TSOEnabled = ctx->Config.TSOEnabled;
DefaultSerializationConfig.ABILocalFlags = ctx->Config.ABILocalFlags;
DefaultSerializationConfig.ABINoPF = ctx->Config.ABINoPF;
DefaultSerializationConfig.SRA = ctx->Config.StaticRegisterAllocation;
DefaultSerializationConfig.ParanoidTSO = ctx->Config.ParanoidTSO;
DefaultSerializationConfig.Is64BitMode = ctx->Config.Is64BitMode;
DefaultSerializationConfig.SMCChecks = ctx->Config.SMCChecks;
DefaultSerializationConfig.x87ReducedPrecision = ctx->Config.x87ReducedPrecision;
}
void NamedRegionObjectHandler::AddNamedRegionObject(CodeRegionMapType::iterator Entry, const std::string &base_filename, const std::string &filename, bool Executable) {
// XXX: Add named region objects
// XXX: Until entry loading is complete just claim it is loaded
Entry->second->NamedJobRefCountMutex.unlock();
}
void NamedRegionObjectHandler::RemoveNamedRegionObject(uintptr_t Base, uintptr_t Size, std::unique_ptr<CodeRegionEntry> Entry) {
// XXX: Remove named region objects
// XXX: Until entry loading is complete just claim it is loaded
Entry->NamedJobRefCountMutex.unlock();
}
void NamedRegionObjectHandler::HandleNamedRegionObjectJobs() {
// Walk through all of our jobs sequentially until the work queue is empty
while (NamedWorkQueueJobs.load()) {
std::unique_ptr<AsyncJobHandler::NamedRegionWorkItem> WorkItem;
{
// Lock the work queue mutex for a short moment and grab an item from the list
std::unique_lock lk {NamedWorkQueueMutex};
size_t WorkItems = WorkQueue.size();
if (WorkItems != 0) {
WorkItem = std::move(WorkQueue.front());
WorkQueue.pop();
}
// Atomically update the number of jobs
--NamedWorkQueueJobs;
}
if (WorkItem) {
if (WorkItem->GetType() == AsyncJobHandler::NamedRegionJobType::JOB_ADD_NAMED_REGION) {
auto WorkAdd = static_cast<AsyncJobHandler::WorkItemAddNamedRegion *>(WorkItem.get());
AddNamedRegionObject(WorkAdd->Entry, WorkAdd->BaseFilename, WorkAdd->Filename, WorkAdd->Executable);
}
if (WorkItem->GetType() == AsyncJobHandler::NamedRegionJobType::JOB_REMOVE_NAMED_REGION) {
auto WorkRemove = static_cast<AsyncJobHandler::WorkItemRemoveNamedRegion *>(WorkItem.get());
RemoveNamedRegionObject(WorkRemove->Base, WorkRemove->Size, std::move(WorkRemove->Entry));
}
}
}
}
}
@@ -0,0 +1,85 @@
#include "Interface/Core/ObjectCache/ObjectCacheService.h"
#include <FEXCore/Config/Config.h>
#include <memory>
namespace {
static void* ThreadHandler(void *Arg) {
FEXCore::CodeSerialize::CodeObjectSerializeService *This = reinterpret_cast<FEXCore::CodeSerialize::CodeObjectSerializeService*>(Arg);
This->ExecutionThread();
return nullptr;
}
}
namespace FEXCore::CodeSerialize {
CodeObjectSerializeService::CodeObjectSerializeService(FEXCore::Context::Context *ctx)
: CTX {ctx}
, AsyncHandler { &NamedRegionHandler , this }
, NamedRegionHandler { ctx } {
Initialize();
}
void CodeObjectSerializeService::Shutdown() {
if (CTX->Config.CacheObjectCodeCompilation() == FEXCore::Config::ConfigObjectCodeHandler::CONFIG_NONE) {
return;
}
WorkerThreadShuttingDown = true;
// Kick the working thread
WorkAvailable.NotifyAll();
if (WorkerThread->joinable()) {
// Wait for worker thread to close down
WorkerThread->join(nullptr);
}
}
void CodeObjectSerializeService::Initialize() {
// Add a canary so we don't crash on empty map iterator handling
auto it = AddressToEntryMap.insert_or_assign(~0ULL, std::make_unique<CodeRegionEntry>());
UnrelocatedAddressToEntryMap.insert_or_assign(~0ULL, it.first->second.get());
uint64_t OldMask = FEXCore::Threads::SetSignalMask(~0ULL);
WorkerThread = FEXCore::Threads::Thread::Create(ThreadHandler, this);
FEXCore::Threads::SetSignalMask(OldMask);
}
void CodeObjectSerializeService::DoCodeRegionClosure(uint64_t Base, CodeRegionEntry *it) {
if (Base == ~0ULL) {
// Don't do closure on canary
return;
}
// XXX: Do code region closure
}
CodeObjectFileSection const *CodeObjectSerializeService::FetchCodeObjectFromCache(uint64_t GuestRIP) {
// XXX: Actually fetch code objects from cache
return nullptr;
}
void CodeObjectSerializeService::ExecutionThread() {
// Set our thread name so we can see its relation
char ThreadName[16] = "ObjectCodeSeri\0";
pthread_setname_np(pthread_self(), ThreadName);
while (WorkerThreadShuttingDown.load() != true) {
// Wait for work
WorkAvailable.Wait();
// Handle named region async jobs first. Highest priority
NamedRegionHandler.HandleNamedRegionObjectJobs();
// XXX: Handle code serialization jobs second.
}
// Do final code region closures on thread shutdown
for (auto &it : AddressToEntryMap) {
DoCodeRegionClosure(it.first, it.second.get());
}
// Safely clear our maps now
AddressToEntryMap.clear();
UnrelocatedAddressToEntryMap.clear();
}
}
@@ -0,0 +1,459 @@
#pragma once
#include "Interface/Context/Context.h"
#include "Interface/Core/ObjectCache/Relocations.h"
#include "Interface/Core/ObjectCache/CodeObjectSerializationConfig.h"
#include "Interface/IR/AOTIR.h"
#include <FEXCore/Utils/Event.h>
#include <FEXCore/Utils/Threads.h>
#include <map>
#include <memory>
#include <shared_mutex>
#include <string>
#include <vector>
#include <tsl/robin_map.h>
namespace FEXCore::CodeSerialize {
// XXX: Does this need to be signal safe?
using CodeSerializationMutex = std::shared_mutex;
struct CodeSerializationData {
};
struct CodeObjectFileSection {
bool Serialized;
bool Invalid;
const CodeSerializationData *Data;
const char *HostCode;
uint64_t NumRelocations;
const char *Relocations;
};
/**
* @brief This is the file header that lives at the start of an object cache file
*
* This header is updated from multiple processes!
* Care must be taken to use OS locks when updating the file backing including this header
*/
struct CodeObjectSerializationHeader {
// The configuration that this file has
CodeObjectSerializationConfig Config;
// The original RIP that this object section was mapped at
uint64_t OriginalBase{};
// The original offset in to the file that this object section was loaded from
uint64_t OriginalOffset{};
// Total amount of code that should be in this file
uint64_t TotalCodeSize{};
// Used to reserve the TSL map
uint64_t NumCodeEntries{};
// The number of relocations that point to this section
uint64_t NumRelocationsTo{};
// Total relocations in this file
uint64_t TotalRelocationsCount{};
};
struct CodeRegionEntry {
/**
* @name Threaded initialization objects for the initial object creation
* @{ */
// Base address in memory where the code region is at
uint64_t Base{};
// Size of this code entry
uint64_t Size{};
// The offset inside the file that is mapped to Base
uint64_t Offset{};
// Filename of the object
std::string Filename{};
CodeObjectSerializationHeader EntryHeader{};
/** @} */
// The filename of the object cache for this entry
std::string ObjectEntrySourceFilename{};
// In the case of file corruption that we can detect, we can disable serialization early for an entry
// We should be resiliant to corruption but things happen
bool StillSerializing {true};
// Long lived FD for serialization if we have multiple jobs to serialize
// Bursts of code entries are common and this reduces file lock overhead
//
// Especially useful over network mounts where file locks are very slow
int CurrentSerializedFD {-1};
/**
* @name Objects required to sync objects between threads
* @{ */
// Refcount for the number of outstanding code entries waiting to be written for this object section
CodeSerializationMutex ObjectJobRefCountMutex;
// Refcount for outstanding named object region entry loading itself
// Will block JIT code cache look up when this has a unique_lock held
CodeSerializationMutex NamedJobRefCountMutex;
/** @} */
/**
* @name Object Entry data management
* @{ */
/**
* @name This is the raw file data that we loaded from the code region entry file
* @{ */
char *CodeData{};
size_t FileSize{};
std::vector<CodeObjectFileSection> FileCodeSections;
/** @} */
// This per section map takes the most time to load and needs to be quick
// This is the map of all code segments for this entry
tsl::robin_map<uint64_t, CodeObjectFileSection*> SectionLookupMap{};
/** @} */
// Default initialization
CodeRegionEntry() = default;
// Initializer specifically for threaded loading
CodeRegionEntry(uint64_t Base,
uint64_t Size,
uint64_t Offset,
std::string const &Filename,
CodeObjectSerializationHeader const &DefaultHeader)
: Base {Base}
, Size {Size}
, Offset {Offset}
, Filename {Filename}
, EntryHeader {DefaultHeader} {
}
};
// Map type must use an interator that isn't invalidation on erase/insert
using CodeRegionMapType = std::map<uint64_t, std::unique_ptr<CodeRegionEntry>>;
using CodeRegionPtrMapType = std::map<uint64_t, CodeRegionEntry*>;
class NamedRegionObjectHandler;
class CodeObjectSerializeService;
class AsyncJobHandler final {
public:
/**
* @brief Structure containing all the data required to async serialize code objects
*/
struct SerializationJobData {
uint64_t GuestRIP; ///< The RIP for the guest
// XXX: Support multiblock
uint64_t GuestCodeLength; ///< The Guest's code length
uint64_t GuestCodeHash; ///< Hash of the guest code
void *HostCodeBegin; ///< Host JIT code starting memory address
size_t HostCodeLength; ///< Host JIT code length
uint64_t HostCodeHash; ///< Host JIT code hash before any backpatching
// This is the thread specific ref counter for outstanding jobs.
// This shared mutex is incremented when the job is added, then decremented when the job is complete.
// If a thread is shutting down or clearing code cache then the thread will pull a unique lock on this mutex.
// This way it will wait until the async job handler is complete with it.
CodeSerializationMutex *ThreadJobRefCount;
// These are the reolocations for this serialization job
// Relatively small number of entries most of the time
std::vector<FEXCore::CPU::Relocation> Relocations;
/**
* @name Objects filled in from the Code Object Serialization service when a job is added
* @{ */
// This is the code region's ref counter for outstanding jobs.
// This shared mutex is incremented when the job is added, then decremented when the job is complete.
// If a named region is being removed then a unique lock will be pulled to wait for all jobs to complete and no new jobs to be added.
CodeSerializationMutex *ObjectJobRefCountMutexPtr;
// This is the code region iterator to reduce the number of map lookups
// This will remain valid while jobs are outstanding for this region
CodeRegionMapType::iterator CodeRegionIterator;
/** @} */
};
AsyncJobHandler(NamedRegionObjectHandler *NamedRegionHandler, CodeObjectSerializeService *CodeObjectCacheService)
: NamedRegionHandler {NamedRegionHandler}
, CodeObjectCacheService {CodeObjectCacheService} {}
protected:
friend class CodeObjectSerializeService;
friend class NamedRegionObjectHandler;
/**
* @name Async job submission functions
* @{ */
void AsyncAddNamedRegionJob(uintptr_t Base, uintptr_t Size, uintptr_t Offset, const std::string &filename);
void AsyncRemoveNamedRegionJob(uintptr_t Base, uintptr_t Size);
void AsyncAddSerializationJob(std::unique_ptr<SerializationJobData> Data);
/** @} */
/**
* @name Async named region handling
* @{ */
/**
* @brief The async named region jobs to handle.
*
* Only two, Code serialization goes in to a different queue.
*/
enum class NamedRegionJobType {
JOB_ADD_NAMED_REGION,
JOB_REMOVE_NAMED_REGION,
};
class NamedRegionWorkItem {
public:
NamedRegionJobType GetType() const { return Type; }
protected:
friend class WorkItemAddNamedRegion;
NamedRegionWorkItem(NamedRegionJobType type)
: Type {type} {}
private:
NamedRegionJobType Type;
};
class WorkItemAddNamedRegion : public NamedRegionWorkItem {
public:
WorkItemAddNamedRegion(const std::string &base, const std::string &filename, bool executable, CodeRegionMapType::iterator entry)
: NamedRegionWorkItem {NamedRegionJobType::JOB_ADD_NAMED_REGION}
, BaseFilename {base}
, Filename {filename}
, Executable {executable}
, Entry {entry}
{}
const std::string BaseFilename;
const std::string Filename;
bool Executable;
CodeRegionMapType::iterator Entry;
};
class WorkItemRemoveNamedRegion : public NamedRegionWorkItem {
public:
WorkItemRemoveNamedRegion(uint64_t base, uint64_t size, std::unique_ptr<CodeRegionEntry> entry)
: NamedRegionWorkItem {NamedRegionJobType::JOB_REMOVE_NAMED_REGION}
, Base {base}
, Size {size}
, Entry {std::move(entry)} {}
uint64_t Base;
uint64_t Size;
std::unique_ptr<CodeRegionEntry> Entry;
};
/** @} */
private:
NamedRegionObjectHandler *NamedRegionHandler;
CodeObjectSerializeService *CodeObjectCacheService;
};
class NamedRegionObjectHandler final {
public:
NamedRegionObjectHandler(FEXCore::Context::Context *ctx);
void HandleNamedRegionObjectJobs();
CodeObjectSerializationConfig const &GetDefaultSerializationConfig() const {
return DefaultSerializationConfig;
}
protected:
friend class AsyncJobHandler;
// Return a default code header based off the default serialization config
CodeObjectSerializationHeader DefaultCodeHeader(uint64_t Base, uint64_t Offset) const {
return CodeObjectSerializationHeader {
.Config = DefaultSerializationConfig,
.OriginalBase = Base,
.OriginalOffset = Offset,
.NumCodeEntries = 0,
.NumRelocationsTo = 0,
.TotalRelocationsCount = 0,
};
}
/**
* @brief Adds an asynchronous add named region work item to the object queue
*
* This adds the job that will do the loading of file resources and data tracking.
*/
void AsyncAddNamedRegionWorkItem(const std::string &base, const std::string &filename, bool executable, CodeRegionMapType::iterator entry) {
std::unique_lock lk {NamedWorkQueueMutex};
WorkQueue.emplace(std::make_unique<AsyncJobHandler::WorkItemAddNamedRegion> (
base,
filename,
executable,
entry
));
++NamedWorkQueueJobs;
}
void AsyncRemoveNamedRegionWorkItem(uint64_t Base, uint64_t Size, std::unique_ptr<CodeRegionEntry> Entry) {
std::unique_lock lk {NamedWorkQueueMutex};
WorkQueue.emplace(std::make_unique<AsyncJobHandler::WorkItemRemoveNamedRegion> (
Base,
Size,
std::move(Entry)
));
++NamedWorkQueueJobs;
}
private:
// Code version. If the code emission changes then this needs to increment
constexpr static uint32_t CODE_VERSION = 0x0;
// Default cookie header for the file header
constexpr static uint64_t CODE_COOKIE = FEXCore::IR::COOKIE_VERSION("FEXC", CODE_VERSION);
// Code serialization config for our current process configuration
CodeObjectSerializationConfig DefaultSerializationConfig;
// Atomic counter for number of jobs in the queue without needing to pull the mutex to check
std::atomic<uint64_t> NamedWorkQueueJobs{};
// Mutex for ading new jobs to the work queue
std::mutex NamedWorkQueueMutex{};
// The job queue itself
// Jobs get consumed as a FIFO
// Jobs always get appended to the end
std::queue<std::unique_ptr<AsyncJobHandler::NamedRegionWorkItem>> WorkQueue{};
/**
* @name Named Region object handling
* @{ */
void AddNamedRegionObject(CodeRegionMapType::iterator Entry, const std::string &base_filename, const std::string &filename, bool Executable);
void RemoveNamedRegionObject(uintptr_t Base, uintptr_t Size, std::unique_ptr<CodeRegionEntry> Entry);
/** @} */
};
/**
* @brief Context specific code object serialization class
*
* Contains everything required for FEXCore to serialize code objects
*/
class CodeObjectSerializeService final {
public:
CodeObjectSerializeService(FEXCore::Context::Context *ctx);
/**
* @brief Initialize the internal interface
*
* Is a public interface to allow the service to reinitialize after forking
*/
void Initialize();
/**
* @brief Safely shut down the Code Object serialization service.
*
* This service needs to be resiliant to application crashes, but shutting down safely is still preferred.
*/
void Shutdown();
/**
* @name Async interface
* @{ */
/**
* @brief Loads a named region in to the code serialization service. As async as possible.
*
* @param Base - Virtual address that this named region is loaded
* @param Size - The size of the region
* @param Offset - The offset from the file
* @param filename - The filename itself
*/
void AsyncAddNamedRegionJob(uintptr_t Base, uintptr_t Size, uintptr_t Offset, const std::string &filename) {
AsyncHandler.AsyncAddNamedRegionJob(Base, Size, Offset, filename);
}
/**
* @brief Unloads a named region from the code serialization service. As async as possible.
*
* @param Base - Virtual address of the named region
* @param Size - The size of the region
*/
void AsyncRemoveNamedRegionJob(uintptr_t Base, uintptr_t Size) {
AsyncHandler.AsyncRemoveNamedRegionJob(Base, Size);
}
/**
* @brief Adds a code object serialization job. As async as possible.
* Code hashing happens prior to async job serialization to catch invalidations due to backpatching.
*
* @param Data - A fully filled out struct containing all the code serialization
*/
void AsyncAddSerializationJob(std::unique_ptr<AsyncJobHandler::SerializationJobData> Data) {
AsyncHandler.AsyncAddSerializationJob(std::move(Data));
}
/** @} */
/**
* @name Synchronous interface
* @{ */
/**
* @brief Synchronously waits for this thread's job queue to become empty.
*
* This is necessary for when a thread is shutting down
*
* @param ThreadJobRefCount - The shared mutex to wait on until to be empty
*/
static void WaitForEmptyJobQueue(CodeSerializationMutex *ThreadJobRefCount) {
// Once the shared mutex is empty this unique lock will be gained
std::unique_lock lk {*ThreadJobRefCount};
}
/**
* @brief Fetches object code from the Code Object Cache for JIT.
*
* @param GuestRIP - Which GuestRIP to search the cache for
*
* @return Data required for the JIT to relocate the Object code.
*/
CodeObjectFileSection const *FetchCodeObjectFromCache(uint64_t GuestRIP);
/** @} */
// Public for threading
void ExecutionThread();
protected:
friend class AsyncJobHandler;
/**
* @brief Safely closes out code object regions from the map
*
* @param it - iterator to do a closure on
*/
void DoCodeRegionClosure(uint64_t Base, CodeRegionEntry *it);
CodeSerializationMutex &GetEntryMapMutex() { return EntryMapMutex; }
CodeSerializationMutex &GetUnrelocatedEntryMapMutex() { return EntryMapMutex; }
CodeRegionMapType &GetEntryMap() { return AddressToEntryMap; }
CodeRegionPtrMapType &GetUnrelocatedEntryMap() { return UnrelocatedAddressToEntryMap; }
/**
* @brief Notify the async thread that it has work to do
*/
void NotifyWork() { WorkAvailable.NotifyOne(); }
private:
FEXCore::Context::Context *CTX;
Event WorkAvailable{};
std::unique_ptr<FEXCore::Threads::Thread> WorkerThread;
std::atomic_bool WorkerThreadShuttingDown {false};
AsyncJobHandler AsyncHandler;
NamedRegionObjectHandler NamedRegionHandler;
// Mutex to hold when modifying the entry maps
CodeSerializationMutex EntryMapMutex;
CodeSerializationMutex UnrelocatedEntryMapMutex;
// Entry maps
CodeRegionMapType AddressToEntryMap;
CodeRegionPtrMapType UnrelocatedAddressToEntryMap;
};
}
@@ -0,0 +1,78 @@
#pragma once
#include <FEXCore/IR/IR.h>
namespace FEXCore::CPU {
enum class RelocationTypes : uint8_t {
// 8 byte literal in memory for symbol
// Aligned to struct RelocNamedSymbolLiteral
RELOC_NAMED_SYMBOL_LITERAL,
// Fixed size named thunk move
// 4 instruction constant generation on AArch64
// 64-bit mov on x86-64
// Aligned to struct RelocNamedThunkMove
RELOC_NAMED_THUNK_MOVE,
// Fixed size guest RIP move
// 4 instruction constant generation on AArch64
// 64-bit mov on x86-64
// Aligned to struct RelocGuestRIPMove
RELOC_GUEST_RIP_MOVE,
};
struct RelocationTypeHeader final {
RelocationTypes Type;
};
struct RelocNamedSymbolLiteral final {
enum class NamedSymbol : uint8_t {
///< Thread specific relocations
// JIT Literal pointers
SYMBOL_LITERAL_EXITFUNCTION_LINKER,
};
RelocationTypeHeader Header{};
NamedSymbol Symbol;
// Offset in to the code section to begin the relocation
uint64_t Offset{};
};
struct RelocNamedThunkMove final {
RelocationTypeHeader Header{};
// GPR index the constant is being moved to
uint8_t RegisterIndex;
// The thunk SHA256 hash
IR::SHA256Sum Symbol;
// Offset in to the code section to begin the relocation
uint64_t Offset{};
};
struct RelocGuestRIPMove final {
RelocationTypeHeader Header{};
// GPR index the constant is being moved to
uint8_t RegisterIndex;
// Offset in to the code section to begin the relocation
uint64_t Offset{};
// The unrelocated RIP that is being moved
uint64_t GuestRIP;
};
union Relocation {
RelocationTypeHeader Header{};
RelocNamedSymbolLiteral NamedSymbolLiteral;
// This makes our union of relocations at least 48 bytes
// It might be more efficient to not use a union
RelocNamedThunkMove NamedThunkMove;
RelocGuestRIPMove GuestRIPMove;
};
}
File diff suppressed because it is too large. Load diff
+677 -29
View File
@@ -14,6 +14,7 @@
#include <FEXCore/Utils/LogManager.h>
#include <cstdint>
#include <fmt/format.h>
#include <map>
#include <stddef.h>
#include <utility>
@@ -35,6 +36,38 @@ enum class SelectionFlag {
};
public:
enum class FlagsGenerationType : uint8_t {
TYPE_NONE,
TYPE_ADC,
TYPE_SBB,
TYPE_SUB,
TYPE_ADD,
TYPE_MUL,
TYPE_UMUL,
TYPE_LOGICAL,
TYPE_LSHL,
TYPE_LSHLI,
TYPE_LSHR,
TYPE_LSHRI,
TYPE_ASHR,
TYPE_ASHRI,
TYPE_ROR,
TYPE_RORI,
TYPE_ROL,
TYPE_ROLI,
TYPE_FCMP,
TYPE_BEXTR,
TYPE_BLSI,
TYPE_BLSMSK,
TYPE_BLSR,
TYPE_POPCOUNT,
TYPE_BZHI,
TYPE_TZCNT,
TYPE_LZCNT,
TYPE_BITSELECT,
TYPE_RDRAND,
};
SelectionFlag flagsOp{};
uint8_t flagsOpSize{};
OrderedNode* flagsOpDest{};
@@ -87,6 +120,9 @@ public:
// cmp qword [rdi-8], 0
// jne .label
if (LastOp && !BlockSetRIP) {
// Calculate flags first
CalculateDeferredFlags();
auto it = JumpTargets.find(NextRIP);
if (it == JumpTargets.end()) {
@@ -101,12 +137,18 @@ public:
return true;
}
}
if (LastOp) {
LOGMAN_THROW_A_FMT(IsDeferredFlagsStored(), "FinishOp: Deferred flags weren't generated at end of block");
}
BlockSetRIP = false;
return false;
}
OpDispatchBuilder(FEXCore::Context::Context *ctx);
OpDispatchBuilder(FEXCore::Utils::IntrusivePooledAllocator &Allocator);
void ResetWorkingList();
void ResetDecodeFailure() { DecodeFailure = false; }
@@ -120,7 +162,9 @@ public:
void UnhandledOp(OpcodeArgs);
template<uint32_t SrcIndex>
void MOVGPROp(OpcodeArgs);
void MOVGPRNTOp(OpcodeArgs);
void MOVVectorOp(OpcodeArgs);
void MOVVectorNTOp(OpcodeArgs);
void ALUOp(OpcodeArgs);
void INTOp(OpcodeArgs);
void SyscallOp(OpcodeArgs);
@@ -233,6 +277,8 @@ public:
void XADDOp(OpcodeArgs);
void PopcountOp(OpcodeArgs);
void XLATOp(OpcodeArgs);
template<bool Reseed>
void RDRANDOp(OpcodeArgs);
enum class Segment {
FS,
@@ -256,9 +302,14 @@ public:
template<FEXCore::IR::IROps IROp, size_t ElementSize>
void VectorALUOp(OpcodeArgs);
template<FEXCore::IR::IROps IROp, size_t ElementSize>
void VectorALUROp(OpcodeArgs);
template<FEXCore::IR::IROps IROp, size_t ElementSize>
void VectorScalarALUOp(OpcodeArgs);
template<FEXCore::IR::IROps IROp, size_t ElementSize, bool Scalar>
void VectorUnaryOp(OpcodeArgs);
template<FEXCore::IR::IROps IROp, size_t ElementSize>
void VectorUnaryDuplicateOp(OpcodeArgs);
void MOVQOp(OpcodeArgs);
template<size_t ElementSize>
void PADDQOp(OpcodeArgs);
@@ -415,6 +466,57 @@ public:
template<size_t width, bool Integer, FCOMIFlags whichflags, bool poptwice>
void FCOMI(OpcodeArgs);
// F64 X87 Ops
template<size_t width>
void FLDF64(OpcodeArgs);
template<uint64_t num>
void FLDF64_Const(OpcodeArgs);
void FBLDF64(OpcodeArgs);
void FBSTPF64(OpcodeArgs);
void FILDF64(OpcodeArgs);
template<size_t width>
void FSTF64(OpcodeArgs);
void FSTF64(OpcodeArgs);
template<bool Truncate>
void FISTF64(OpcodeArgs);
template<size_t width, bool Integer, OpResult ResInST0>
void FADDF64(OpcodeArgs);
template<size_t width, bool Integer, OpResult ResInST0>
void FMULF64(OpcodeArgs);
template<size_t width, bool Integer, bool reverse, OpResult ResInST0>
void FDIVF64(OpcodeArgs);
template<size_t width, bool Integer, bool reverse, OpResult ResInST0>
void FSUBF64(OpcodeArgs);
void FCHSF64(OpcodeArgs);
void FABSF64(OpcodeArgs);
void FTSTF64(OpcodeArgs);
void FRNDINTF64(OpcodeArgs);
void FXTRACTF64(OpcodeArgs);
void FNINITF64(OpcodeArgs);
void FSQRTF64(OpcodeArgs);
template<FEXCore::IR::IROps IROp>
void X87UnaryOpF64(OpcodeArgs);
template<FEXCore::IR::IROps IROp>
void X87BinaryOpF64(OpcodeArgs);
void X87SinCosF64(OpcodeArgs);
void X87FLDCWF64(OpcodeArgs);
void X87FYL2XF64(OpcodeArgs);
void X87TANF64(OpcodeArgs);
void X87ATANF64(OpcodeArgs);
void X87FNSAVEF64(OpcodeArgs);
void X87FRSTORF64(OpcodeArgs);
void X87FXAMF64(OpcodeArgs);
void X87LDENVF64(OpcodeArgs);
template<size_t width, bool Integer, FCOMIFlags whichflags, bool poptwice>
void FCOMIF64(OpcodeArgs);
void FXSaveOp(OpcodeArgs);
void FXRStoreOp(OpcodeArgs);
@@ -445,6 +547,17 @@ public:
template<size_t ElementSize>
void ADDSUBPOp(OpcodeArgs);
void PFNACCOp(OpcodeArgs);
void PFPNACCOp(OpcodeArgs);
void PSWAPDOp(OpcodeArgs);
template<uint8_t CompType>
void VPFCMPOp(OpcodeArgs);
void PI2FWOp(OpcodeArgs);
void PF2IWOp(OpcodeArgs);
void PMULHRWOp(OpcodeArgs);
void PMADDWD(OpcodeArgs);
void PMADDUBSW(OpcodeArgs);
@@ -476,6 +589,15 @@ public:
void PSADBW(OpcodeArgs);
void SHA1NEXTEOp(OpcodeArgs);
void SHA1MSG1Op(OpcodeArgs);
void SHA1MSG2Op(OpcodeArgs);
void SHA1RNDS4Op(OpcodeArgs);
void SHA256MSG1Op(OpcodeArgs);
void SHA256MSG2Op(OpcodeArgs);
void SHA256RNDS2Op(OpcodeArgs);
void AESImcOp(OpcodeArgs);
void AESEncOp(OpcodeArgs);
void AESEncLastOp(OpcodeArgs);
@@ -500,6 +622,8 @@ public:
void MPSADBWOp(OpcodeArgs);
void CRC32(OpcodeArgs);
void UnimplementedOp(OpcodeArgs);
void InvalidOp(OpcodeArgs);
@@ -519,12 +643,22 @@ private:
OrderedNode *AppendSegmentOffset(OrderedNode *Value, uint32_t Flags, uint32_t DefaultPrefix = 0, bool Override = false);
OrderedNode *GetDynamicPC(FEXCore::X86Tables::DecodedOp const& Op, int64_t Offset = 0);
OrderedNode *LoadSource(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp const& Op, FEXCore::X86Tables::DecodedOperand const& Operand, uint32_t Flags, int8_t Align, bool LoadData = true, bool ForceLoad = false);
OrderedNode *LoadSource_WithOpSize(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp const& Op, FEXCore::X86Tables::DecodedOperand const& Operand, uint8_t OpSize, uint32_t Flags, int8_t Align, bool LoadData = true, bool ForceLoad = false);
void StoreResult_WithOpSize(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp Op, FEXCore::X86Tables::DecodedOperand const& Operand, OrderedNode *const Src, uint8_t OpSize, int8_t Align);
void StoreResult(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp Op, FEXCore::X86Tables::DecodedOperand const& Operand, OrderedNode *const Src, int8_t Align);
void StoreResult(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp Op, OrderedNode *const Src, int8_t Align);
enum class MemoryAccessType {
// Choose TSO or Non-TSO depending on access type
ACCESS_DEFAULT,
// TSO access behaviour
ACCESS_TSO,
// Non-TSO access behaviour
ACCESS_NONTSO,
// Non-temporal streaming
ACCESS_STREAM,
};
OrderedNode *GetRelocatedPC(FEXCore::X86Tables::DecodedOp const& Op, int64_t Offset = 0);
OrderedNode *LoadSource(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp const& Op, FEXCore::X86Tables::DecodedOperand const& Operand, uint32_t Flags, int8_t Align, bool LoadData = true, bool ForceLoad = false, MemoryAccessType AccessType = MemoryAccessType::ACCESS_DEFAULT);
OrderedNode *LoadSource_WithOpSize(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp const& Op, FEXCore::X86Tables::DecodedOperand const& Operand, uint8_t OpSize, uint32_t Flags, int8_t Align, bool LoadData = true, bool ForceLoad = false, MemoryAccessType AccessType = MemoryAccessType::ACCESS_DEFAULT);
void StoreResult_WithOpSize(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp Op, FEXCore::X86Tables::DecodedOperand const& Operand, OrderedNode *const Src, uint8_t OpSize, int8_t Align, MemoryAccessType AccessType = MemoryAccessType::ACCESS_DEFAULT);
void StoreResult(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp Op, FEXCore::X86Tables::DecodedOperand const& Operand, OrderedNode *const Src, int8_t Align, MemoryAccessType AccessType = MemoryAccessType::ACCESS_DEFAULT);
void StoreResult(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp Op, OrderedNode *const Src, int8_t Align, MemoryAccessType AccessType = MemoryAccessType::ACCESS_DEFAULT);
[[nodiscard]] static uint32_t GPROffset(X86State::X86Reg reg) {
LOGMAN_THROW_A_FMT(reg <= X86State::X86Reg::REG_R15, "Invalid reg used");
@@ -557,23 +691,524 @@ private:
OrderedNode *SelectCC(uint8_t OP, OrderedNode *TrueValue, OrderedNode *FalseValue);
void GenerateFlags_ADC(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, OrderedNode *CF);
void GenerateFlags_SBB(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, OrderedNode *CF);
void GenerateFlags_SUB(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, bool UpdateCF = true);
void GenerateFlags_ADD(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, bool UpdateCF = true);
void GenerateFlags_MUL(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *High);
void GenerateFlags_UMUL(FEXCore::X86Tables::DecodedOp Op, OrderedNode *High);
void GenerateFlags_Logical(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2);
void GenerateFlags_ShiftLeft(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2);
void GenerateFlags_ShiftLeftImmediate(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift);
void GenerateFlags_ShiftRight(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2);
void GenerateFlags_ShiftRightImmediate(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift);
void GenerateFlags_SignShiftRight(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2);
void GenerateFlags_SignShiftRightImmediate(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift);
void GenerateFlags_RotateRight(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2);
void GenerateFlags_RotateLeft(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2);
void GenerateFlags_RotateRightImmediate(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift);
void GenerateFlags_RotateLeftImmediate(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift);
/**
* @name Deferred RFLAG calculation and generation.
*
* Only handles the six flags that ALU ops typically generate.
* Specifically: CF, PF, AF, ZF, SF, OF
* These six flags are heavily generated through basic ALU ops and balloon the IR if not early eliminated.
* This tracking structure only tracks single blocks and requires RFLAGS calculation at block-ending ops.
* Some flags generating ALU ops only touch part of the registers, In these cases it will do calculation up front.
* This means we still need our IR passes to eliminate all redundant flags accesses but this light OpcodeDispatcher optimization
* doesn't take it to that level.
* @{ */
// Deferred flag generation tracking structure.
// This structure is used to track RFlags from ALU ops for invalidation.
//
// Future ideas: Use an invalidation mask to do partial generation of flags.
// Particularly for the instructions that don't do the full set of flags calculations.
// These instructions currently calculate the deferred RFLAGS immediately then overwrite rflags state.
// RCLSE IR pass will catch and remove redundant rflags stores like this currently.
struct DeferredFlagData {
// What type of flags to generate
FlagsGenerationType Type {FlagsGenerationType::TYPE_NONE};
// Source size of the op
uint8_t SrcSize;
// Every flag generation type has a result
OrderedNode *Res{};
union {
// UMUL, BEXTR, BLSI, BLSMSK, POPCOUNT, TZCNT, LZCNT, BITSELECT, RDRAND
struct {
} NoSource;
// MUL, BLSR, BZHI
struct {
OrderedNode *Src1;
} OneSource;
// Logical, LSHL, LSHR, ASHR, ROR, ROL
struct {
OrderedNode *Src1;
OrderedNode *Src2;
} TwoSource;
// ADC, SBB
struct {
OrderedNode *Src1;
OrderedNode *Src2;
OrderedNode *Src3;
} ThreeSource;
// LSHLI, LSHRI, ASHRI, RORI, ROLI
struct {
OrderedNode *Src1;
uint64_t Imm;
} OneSrcImmediate;
// ADD, SUB
struct {
OrderedNode *Src1;
OrderedNode *Src2;
bool UpdateCF;
} TwoSrcImmediate;
} Sources{};
};
DeferredFlagData CurrentDeferredFlags{};
/**
* @brief Takes the current deferred flag state and stores the result in to RFLAGS.
*
* Once executed there will no longer be any deferred flag state and RFLAGS will have the correct flags in it.
* Necessary to do when leaving a IR block, or if an instruction is doing a partial overwrite of the flags.
*/
void CalculateDeferredFlags(uint32_t FlagsToCalculateMask = ~0U);
/**
* @brief Invalidates the current deferred flags structure.
*
* If the emulated instruction is going to overwrite all of the flags but isn't tracked using the deferred flag system
* then use this function to stop tracking the current active deferred flags.
*/
void InvalidateDeferredFlags() {
CurrentDeferredFlags.Type = FlagsGenerationType::TYPE_NONE;
}
/**
* @brief Checks if there is any deferred flag state active.
*
* @return True if RFLAGs contains the flags. False if deferred flags is tracking the data.
*/
bool IsDeferredFlagsStored() const {
return CurrentDeferredFlags.Type == FlagsGenerationType::TYPE_NONE;
}
/**
* @name These functions are used by the deferred flag handling while it is calculating and storing flags in to RFLAGs.
* @{ */
void CalculcateFlags_ADC(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, OrderedNode *CF);
void CalculcateFlags_SBB(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, OrderedNode *CF);
void CalculcateFlags_SUB(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, bool UpdateCF = true);
void CalculcateFlags_ADD(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, bool UpdateCF = true);
void CalculcateFlags_MUL(uint8_t SrcSize, OrderedNode *Res, OrderedNode *High);
void CalculcateFlags_UMUL(OrderedNode *High);
void CalculcateFlags_Logical(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2);
void CalculcateFlags_ShiftLeft(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2);
void CalculcateFlags_ShiftLeftImmediate(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift);
void CalculcateFlags_ShiftRight(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2);
void CalculcateFlags_ShiftRightImmediate(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift);
void CalculcateFlags_SignShiftRight(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2);
void CalculcateFlags_SignShiftRightImmediate(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift);
void CalculcateFlags_RotateRight(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2);
void CalculcateFlags_RotateLeft(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2);
void CalculcateFlags_RotateRightImmediate(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift);
void CalculcateFlags_RotateLeftImmediate(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift);
void CalculcateFlags_FCMP(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2);
void CalculcateFlags_BEXTR(OrderedNode *Src);
void CalculcateFlags_BLSI(uint8_t SrcSize, OrderedNode *Src);
void CalculcateFlags_BLSMSK(OrderedNode *Src);
void CalculcateFlags_BLSR(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src);
void CalculcateFlags_POPCOUNT(OrderedNode *Src);
void CalculcateFlags_BZHI(uint8_t SrcSize, OrderedNode *Result, OrderedNode *Src);
void CalculcateFlags_TZCNT(OrderedNode *Src);
void CalculcateFlags_LZCNT(uint8_t SrcSize, OrderedNode *Src);
void CalculcateFlags_BITSELECT(OrderedNode *Src);
void CalculcateFlags_RDRAND(OrderedNode *Src);
/** @} */
/**
* @name These functions generated deferred RFLAGs tracking.
*
* Depending on the operation it may force a RFLAGs calculation before storing the new deferred state.
* @{ */
void GenerateFlags_ADC(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, OrderedNode *CF) {
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_ADC,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.ThreeSource = {
.Src1 = Src1,
.Src2 = Src2,
.Src3 = CF,
},
},
};
}
void GenerateFlags_SBB(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, OrderedNode *CF) {
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_SBB,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.ThreeSource = {
.Src1 = Src1,
.Src2 = Src2,
.Src3 = CF,
},
},
};
}
void GenerateFlags_SUB(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, bool UpdateCF = true) {
if (!UpdateCF) {
// If we aren't updating CF then we need to calculate flags. Invalidation mask would make this not required.
CalculateDeferredFlags();
}
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_SUB,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.TwoSrcImmediate = {
.Src1 = Src1,
.Src2 = Src2,
.UpdateCF = UpdateCF,
},
},
};
}
void GenerateFlags_ADD(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, bool UpdateCF = true) {
if (!UpdateCF) {
// If we aren't updating CF then we need to calculate flags. Invalidation mask would make this not required.
CalculateDeferredFlags();
}
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_ADD,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.TwoSrcImmediate = {
.Src1 = Src1,
.Src2 = Src2,
.UpdateCF = UpdateCF,
},
},
};
}
void GenerateFlags_MUL(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *High) {
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_MUL,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.OneSource = {
.Src1 = High,
},
},
};
}
void GenerateFlags_UMUL(FEXCore::X86Tables::DecodedOp Op, OrderedNode *High) {
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_UMUL,
.SrcSize = GetSrcSize(Op),
.Res = High,
};
}
void GenerateFlags_Logical(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_LOGICAL,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.TwoSource = {
.Src1 = Src1,
.Src2 = Src2,
},
},
};
}
void GenerateFlags_ShiftLeft(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
// Flags need to be used, generate incoming flags first.
CalculateDeferredFlags();
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_LSHL,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.TwoSource = {
.Src1 = Src1,
.Src2 = Src2,
},
},
};
}
void GenerateFlags_ShiftRight(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
// Flags need to be used, generate incoming flags first.
CalculateDeferredFlags();
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_LSHR,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.TwoSource = {
.Src1 = Src1,
.Src2 = Src2,
},
},
};
}
void GenerateFlags_SignShiftRight(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
// Flags need to be used, generate incoming flags first.
CalculateDeferredFlags();
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_ASHR,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.TwoSource = {
.Src1 = Src1,
.Src2 = Src2,
},
},
};
}
void GenerateFlags_ShiftLeftImmediate(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift) {
// No flags changed if shift is zero.
if (Shift == 0) return;
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_LSHLI,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.OneSrcImmediate = {
.Src1 = Src1,
.Imm = Shift,
},
},
};
}
void GenerateFlags_SignShiftRightImmediate(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift) {
// No flags changed if shift is zero.
if (Shift == 0) return;
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_ASHRI,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.OneSrcImmediate = {
.Src1 = Src1,
.Imm = Shift,
},
},
};
}
void GenerateFlags_ShiftRightImmediate(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift) {
// No flags changed if shift is zero.
if (Shift == 0) return;
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_LSHRI,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.OneSrcImmediate = {
.Src1 = Src1,
.Imm = Shift,
},
},
};
}
void GenerateFlags_RotateRight(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
// Doesn't set all the flags, needs to calculate.
CalculateDeferredFlags();
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_ROR,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.TwoSource = {
.Src1 = Src1,
.Src2 = Src2,
},
},
};
}
void GenerateFlags_RotateLeft(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
// Doesn't set all the flags, needs to calculate.
CalculateDeferredFlags();
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_ROL,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.TwoSource = {
.Src1 = Src1,
.Src2 = Src2,
},
},
};
}
void GenerateFlags_RotateRightImmediate(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift) {
if (Shift == 0) return;
// Doesn't set all the flags, needs to calculate.
CalculateDeferredFlags();
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_RORI,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.OneSrcImmediate = {
.Src1 = Src1,
.Imm = Shift,
},
},
};
}
void GenerateFlags_RotateLeftImmediate(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift) {
if (Shift == 0) return;
// Doesn't set all the flags, needs to calculate.
CalculateDeferredFlags();
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_ROLI,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.OneSrcImmediate = {
.Src1 = Src1,
.Imm = Shift,
},
}
};
}
void GenerateFlags_FCMP(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_FCMP,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.TwoSource = {
.Src1 = Src1,
.Src2 = Src2,
},
}
};
}
void GenerateFlags_BEXTR(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Src) {
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_BEXTR,
.SrcSize = GetSrcSize(Op),
.Res = Src,
};
}
void GenerateFlags_BLSI(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Src) {
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_BLSI,
.SrcSize = GetSrcSize(Op),
.Res = Src,
};
}
void GenerateFlags_BLSMSK(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Src) {
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_BLSMSK,
.SrcSize = GetSrcSize(Op),
.Res = Src,
};
}
void GenerateFlags_BLSR(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src) {
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_BLSR,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.OneSource = {
.Src1 = Src,
},
},
};
}
void GenerateFlags_POPCOUNT(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Src) {
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_POPCOUNT,
.SrcSize = GetSrcSize(Op),
.Res = Src,
};
}
void GenerateFlags_BZHI(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Result, OrderedNode *Src) {
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_BZHI,
.SrcSize = GetSrcSize(Op),
.Res = Result,
.Sources = {
.OneSource = {
.Src1 = Src,
},
},
};
}
void GenerateFlags_TZCNT(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Src) {
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_TZCNT,
.SrcSize = GetSrcSize(Op),
.Res = Src,
};
}
void GenerateFlags_LZCNT(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Src) {
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_LZCNT,
.SrcSize = GetSrcSize(Op),
.Res = Src,
};
}
void GenerateFlags_BITSELECT(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Src) {
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_BITSELECT,
.SrcSize = GetSrcSize(Op),
.Res = Src,
};
}
void GenerateFlags_RDRAND(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Src) {
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_RDRAND,
.SrcSize = GetSrcSize(Op),
.Res = Src,
};
}
/** @} */
/** @} */
OrderedNode * GetX87Top();
enum class X87Tag {
@@ -600,24 +1235,37 @@ private:
bool Multiblock{};
uint64_t Entry;
OrderedNode* _StoreMemAutoTSO(FEXCore::IR::RegisterClassType Class, uint8_t Size, OrderedNode *ssa0, OrderedNode *ssa1, uint8_t Align = 1) {
OrderedNode* _StoreMemAutoTSO(FEXCore::IR::RegisterClassType Class, uint8_t Size, OrderedNode *Addr, OrderedNode *Value, uint8_t Align = 1) {
if (CTX->Config.TSOEnabled)
return _StoreMemTSO(ssa0, ssa1, Invalid(), Align, Class, MEM_OFFSET_SXTX, 1, Size);
return _StoreMemTSO(Class, Size, Value, Addr, Invalid(), Align, MEM_OFFSET_SXTX, 1);
else
return _StoreMem(ssa0, ssa1, Invalid(), Align, Class, MEM_OFFSET_SXTX, 1, Size);
return _StoreMem(Class, Size, Value, Addr, Invalid(), Align, MEM_OFFSET_SXTX, 1);
}
OrderedNode* _LoadMemAutoTSO(FEXCore::IR::RegisterClassType Class, uint8_t Size, OrderedNode *ssa0, uint8_t Align = 1) {
if (CTX->Config.TSOEnabled)
return _LoadMemTSO(ssa0, Invalid(), Align, Class, MEM_OFFSET_SXTX, 1, Size);
return _LoadMemTSO(Class, Size, ssa0, Invalid(), Align, MEM_OFFSET_SXTX, 1);
else
return _LoadMem(ssa0, Invalid(), Align, Class, MEM_OFFSET_SXTX, 1, Size);
return _LoadMem(Class, Size, ssa0, Invalid(), Align, MEM_OFFSET_SXTX, 1);
}
void InstallHostSpecificOpcodeHandlers();
};
void InstallOpcodeHandlers(Context::OperatingMode Mode);
}
template <>
struct fmt::formatter<FEXCore::IR::OpDispatchBuilder::FlagsGenerationType> : fmt::formatter<int> {
using Base = fmt::formatter<int>;
// Pass-through the underlying value, so IDs can
// be formatted like any integral value.
template <typename FormatContext>
auto format(const FEXCore::IR::OpDispatchBuilder::FlagsGenerationType& ID, FormatContext& ctx) {
return Base::format(static_cast<int>(ID), ctx);
}
};
@@ -10,13 +10,259 @@ $end_info$
#include <FEXCore/Utils/LogManager.h>
#include "Interface/Core/OpcodeDispatcher.h"
#include <stdint.h>
#include <array>
#include <cstdint>
#include <tuple>
#include <utility>
namespace FEXCore::IR {
class OrderedNode;
#define OpcodeArgs [[maybe_unused]] FEXCore::X86Tables::DecodedOp Op
void OpDispatchBuilder::SHA1NEXTEOp(OpcodeArgs) {
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
auto Tmp = _Ror(_VExtractToGPR(16, 4, Dest, 3), _Constant(32, 2));
auto Top = _Add(_VExtractToGPR(16, 4, Src, 3), Tmp);
auto Result = _VInsGPR(16, 4, 3, Src, Top);
StoreResult(FPRClass, Op, Result, -1);
}
void OpDispatchBuilder::SHA1MSG1Op(OpcodeArgs) {
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
auto W0 = _VExtractToGPR(16, 4, Dest, 3);
auto W1 = _VExtractToGPR(16, 4, Dest, 2);
auto W2 = _VExtractToGPR(16, 4, Dest, 1);
auto W3 = _VExtractToGPR(16, 4, Dest, 0);
auto W4 = _VExtractToGPR(16, 4, Src, 3);
auto W5 = _VExtractToGPR(16, 4, Src, 2);
auto D3 = _VInsGPR(16, 4, 3, Dest, _Xor(W2, W0));
auto D2 = _VInsGPR(16, 4, 2, D3, _Xor(W3, W1));
auto D1 = _VInsGPR(16, 4, 1, D2, _Xor(W4, W2));
auto D0 = _VInsGPR(16, 4, 0, D1, _Xor(W5, W3));
StoreResult(FPRClass, Op, D0, -1);
}
void OpDispatchBuilder::SHA1MSG2Op(OpcodeArgs) {
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
// ROR by 31 is equivalent to a ROL by 1
auto ThirtyOne = _Constant(32, 31);
auto W13 = _VExtractToGPR(16, 4, Src, 2);
auto W14 = _VExtractToGPR(16, 4, Src, 1);
auto W15 = _VExtractToGPR(16, 4, Src, 0);
auto W16 = _Ror(_Xor(_VExtractToGPR(16, 4, Dest, 3), W13), ThirtyOne);
auto W17 = _Ror(_Xor(_VExtractToGPR(16, 4, Dest, 2), W14), ThirtyOne);
auto W18 = _Ror(_Xor(_VExtractToGPR(16, 4, Dest, 1), W15), ThirtyOne);
auto W19 = _Ror(_Xor(_VExtractToGPR(16, 4, Dest, 0), W16), ThirtyOne);
auto D3 = _VInsGPR(16, 4, 3, Dest, W16);
auto D2 = _VInsGPR(16, 4, 2, D3, W17);
auto D1 = _VInsGPR(16, 4, 1, D2, W18);
auto D0 = _VInsGPR(16, 4, 0, D1, W19);
StoreResult(FPRClass, Op, D0, -1);
}
void OpDispatchBuilder::SHA1RNDS4Op(OpcodeArgs) {
LOGMAN_THROW_A_FMT(Op->Src[1].IsLiteral(),
"Src1 needs to be literal here to indicate function and constants");
using FnType = OrderedNode* (*)(OpDispatchBuilder&, OrderedNode*, OrderedNode*, OrderedNode*);
const auto f0 = [](OpDispatchBuilder &Self, OrderedNode *B, OrderedNode *C, OrderedNode *D) -> OrderedNode* {
return Self._Xor(Self._And(B, C), Self._And(Self._Not(B), D));
};
const auto f1 = [](OpDispatchBuilder &Self, OrderedNode *B, OrderedNode *C, OrderedNode *D) -> OrderedNode* {
return Self._Xor(Self._Xor(B, C), D);
};
const auto f2 = [](OpDispatchBuilder &Self, OrderedNode *B, OrderedNode *C, OrderedNode *D) -> OrderedNode* {
return Self._Xor(Self._Xor(Self._And(B, C), Self._And(B, D)), Self._And(C, D));
};
const auto f3 = [](OpDispatchBuilder &Self, OrderedNode *B, OrderedNode *C, OrderedNode *D) -> OrderedNode* {
return Self._Xor(Self._Xor(B, C), D);
};
constexpr std::array<uint32_t, 4> k_array{
0x5A827999U,
0x6ED9EBA1U,
0x8F1BBCDCU,
0xCA62C1D6U,
};
constexpr std::array<FnType, 4> fn_array{
f0, f1, f2, f3,
};
const uint64_t Imm8 = Op->Src[1].Data.Literal.Value & 0b11;
const FnType Fn = fn_array[Imm8];
auto K = _Constant(32, k_array[Imm8]);
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
auto W0E = _VExtractToGPR(16, 4, Src, 3);
auto W1 = _VExtractToGPR(16, 4, Src, 2);
auto W2 = _VExtractToGPR(16, 4, Src, 1);
auto W3 = _VExtractToGPR(16, 4, Src, 0);
using RoundResult = std::tuple<OrderedNode*, OrderedNode*, OrderedNode*, OrderedNode*, OrderedNode*>;
const auto Round0 = [&]() -> RoundResult {
auto A = _VExtractToGPR(16, 4, Dest, 3);
auto B = _VExtractToGPR(16, 4, Dest, 2);
auto C = _VExtractToGPR(16, 4, Dest, 1);
auto D = _VExtractToGPR(16, 4, Dest, 0);
auto A1 = _Add(_Add(_Add(Fn(*this, B, C, D), _Ror(A, _Constant(32, 27))), W0E), K);
auto B1 = A;
auto C1 = _Ror(B, _Constant(32, 2));
auto D1 = C;
auto E1 = D;
return {A1, B1, C1, D1, E1};
};
const auto Round1To3 = [&](OrderedNode *A, OrderedNode *B, OrderedNode *C,
OrderedNode *D, OrderedNode *E, OrderedNode *W) -> RoundResult {
auto ANext = _Add(_Add(_Add(_Add(Fn(*this, B, C, D), _Ror(A, _Constant(32, 27))), W), E), K);
auto BNext = A;
auto CNext = _Ror(B, _Constant(32, 2));
auto DNext = C;
auto ENext = D;
return {ANext, BNext, CNext, DNext, ENext};
};
auto [A1, B1, C1, D1, E1] = Round0();
auto [A2, B2, C2, D2, E2] = Round1To3(A1, B1, C1, D1, E1, W1);
auto [A3, B3, C3, D3, E3] = Round1To3(A2, B2, C2, D2, E2, W2);
auto Final = Round1To3(A3, B3, C3, D3, E3, W3);
auto Dest3 = _VInsGPR(16, 4, 3, Dest, std::get<0>(Final));
auto Dest2 = _VInsGPR(16, 4, 2, Dest3, std::get<1>(Final));
auto Dest1 = _VInsGPR(16, 4, 1, Dest2, std::get<2>(Final));
auto Dest0 = _VInsGPR(16, 4, 0, Dest1, std::get<3>(Final));
StoreResult(FPRClass, Op, Dest0, -1);
}
void OpDispatchBuilder::SHA256MSG1Op(OpcodeArgs) {
const auto Sigma0 = [this](OrderedNode* W) -> OrderedNode* {
return _Xor(_Xor(_Ror(W, _Constant(32, 7)), _Ror(W, _Constant(32, 18))), _Lshr(W, _Constant(32, 3)));
};
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
auto W4 = _VExtractToGPR(16, 4, Src, 0);
auto W3 = _VExtractToGPR(16, 4, Dest, 3);
auto W2 = _VExtractToGPR(16, 4, Dest, 2);
auto W1 = _VExtractToGPR(16, 4, Dest, 1);
auto W0 = _VExtractToGPR(16, 4, Dest, 0);
auto Sig3 = _Add(W3, Sigma0(W4));
auto Sig2 = _Add(W2, Sigma0(W3));
auto Sig1 = _Add(W1, Sigma0(W2));
auto Sig0 = _Add(W0, Sigma0(W1));
auto D3 = _VInsGPR(16, 4, 3, Dest, Sig3);
auto D2 = _VInsGPR(16, 4, 2, D3, Sig2);
auto D1 = _VInsGPR(16, 4, 1, D2, Sig1);
auto D0 = _VInsGPR(16, 4, 0, D1, Sig0);
StoreResult(FPRClass, Op, D0, -1);
}
void OpDispatchBuilder::SHA256MSG2Op(OpcodeArgs) {
const auto Sigma1 = [this](OrderedNode* W) -> OrderedNode* {
return _Xor(_Xor(_Ror(W, _Constant(32, 17)), _Ror(W, _Constant(32, 19))), _Lshr(W, _Constant(32, 10)));
};
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
auto W14 = _VExtractToGPR(16, 4, Src, 2);
auto W15 = _VExtractToGPR(16, 4, Src, 3);
auto W16 = _Add(_VExtractToGPR(16, 4, Dest, 0), Sigma1(W14));
auto W17 = _Add(_VExtractToGPR(16, 4, Dest, 1), Sigma1(W15));
auto W18 = _Add(_VExtractToGPR(16, 4, Dest, 2), Sigma1(W16));
auto W19 = _Add(_VExtractToGPR(16, 4, Dest, 3), Sigma1(W17));
auto D3 = _VInsGPR(16, 4, 3, Dest, W19);
auto D2 = _VInsGPR(16, 4, 2, D3, W18);
auto D1 = _VInsGPR(16, 4, 1, D2, W17);
auto D0 = _VInsGPR(16, 4, 0, D1, W16);
StoreResult(FPRClass, Op, D0, -1);
}
void OpDispatchBuilder::SHA256RNDS2Op(OpcodeArgs) {
const auto Ch = [this](OrderedNode *E, OrderedNode *F, OrderedNode *G) -> OrderedNode* {
return _Xor(_And(E, F), _And(_Not(E), G));
};
const auto Major = [this](OrderedNode *A, OrderedNode *B, OrderedNode *C) -> OrderedNode* {
return _Xor(_Xor(_And(A, B), _And(A, C)), _And(B, C));
};
const auto Sigma0 = [this](OrderedNode *A) -> OrderedNode* {
return _Xor(_Xor(_Ror(A, _Constant(32, 2)), _Ror(A, _Constant(32, 13))), _Ror(A, _Constant(32, 22)));
};
const auto Sigma1 = [this](OrderedNode *E) -> OrderedNode* {
return _Xor(_Xor(_Ror(E, _Constant(32, 6)), _Ror(E, _Constant(32, 11))), _Ror(E, _Constant(32, 25)));
};
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *XMM0 = _LoadContext(16, FPRClass, offsetof(FEXCore::Core::CPUState, xmm[0]));
auto A0 = _VExtractToGPR(16, 4, Src, 3);
auto B0 = _VExtractToGPR(16, 4, Src, 2);
auto C0 = _VExtractToGPR(16, 4, Dest, 3);
auto D0 = _VExtractToGPR(16, 4, Dest, 2);
auto E0 = _VExtractToGPR(16, 4, Src, 1);
auto F0 = _VExtractToGPR(16, 4, Src, 0);
auto G0 = _VExtractToGPR(16, 4, Dest, 1);
auto H0 = _VExtractToGPR(16, 4, Dest, 0);
auto WK0 = _VExtractToGPR(16, 4, XMM0, 0);
auto WK1 = _VExtractToGPR(16, 4, XMM0, 1);
using RoundResult = std::tuple<OrderedNode*, OrderedNode*, OrderedNode*, OrderedNode*,
OrderedNode*, OrderedNode*, OrderedNode*, OrderedNode*>;
const auto Round = [&](OrderedNode *A, OrderedNode *B, OrderedNode *C, OrderedNode *D,
OrderedNode *E, OrderedNode *F, OrderedNode *G, OrderedNode *H,
OrderedNode* WK) -> RoundResult {
auto ANext = _Add(_Add(_Add(_Add(_Add(Ch(E, F, G), Sigma1(E)), WK), H), Major(A, B, C)), Sigma0(A));
auto BNext = A;
auto CNext = B;
auto DNext = C;
auto ENext = _Add(_Add(_Add(_Add(Ch(E, F, G), Sigma1(E)), WK), H), D);
auto FNext = E;
auto GNext = F;
auto HNext = G;
return {ANext, BNext, CNext, DNext, ENext, FNext, GNext, HNext};
};
auto [A1, B1, C1, D1, E1, F1, G1, H1] = Round(A0, B0, C0, D0, E0, F0, G0, H0, WK0);
auto Final = Round(A1, B1, C1, D1, E1, F1, G1, H1, WK1);
auto Res3 = _VInsGPR(16, 4, 3, Dest, std::get<0>(Final));
auto Res2 = _VInsGPR(16, 4, 2, Res3, std::get<1>(Final));
auto Res1 = _VInsGPR(16, 4, 1, Res2, std::get<4>(Final));
auto Res0 = _VInsGPR(16, 4, 0, Res1, std::get<5>(Final));
StoreResult(FPRClass, Op, Res0, -1);
}
void OpDispatchBuilder::AESImcOp(OpcodeArgs) {
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
auto Res = _VAESImc(Src);
@@ -41,8 +41,17 @@ constexpr std::array<uint32_t, 17> FlagOffsets = {
void OpDispatchBuilder::SetPackedRFLAG(bool Lower8, OrderedNode *Src) {
size_t NumFlags = FlagOffsets.size();
if (Lower8) {
// Calculate flags early.
// Could use InvalidateDeferredFlags() if we had masked invalidation.
// This is only a partial overwrite of flags since OF isn't stored here.
CalculateDeferredFlags();
NumFlags = 5;
}
else {
// We are overwriting all RFLAGS. Invalidate the deferred flag state.
InvalidateDeferredFlags();
}
auto OneConst = _Constant(1);
for (size_t i = 0; i < NumFlags; ++i) {
const auto FlagOffset = FlagOffsets[i];
@@ -52,6 +61,9 @@ void OpDispatchBuilder::SetPackedRFLAG(bool Lower8, OrderedNode *Src) {
}
OrderedNode *OpDispatchBuilder::GetPackedRFLAG(bool Lower8) {
// Calculate flags early.
CalculateDeferredFlags();
OrderedNode *Original = _Constant(2);
size_t NumFlags = FlagOffsets.size();
if (Lower8) {
@@ -68,8 +80,188 @@ OrderedNode *OpDispatchBuilder::GetPackedRFLAG(bool Lower8) {
return Original;
}
void OpDispatchBuilder::GenerateFlags_ADC(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, OrderedNode *CF) {
auto Size = GetSrcSize(Op) * 8;
void OpDispatchBuilder::CalculateDeferredFlags(uint32_t FlagsToCalculateMask) {
if (CurrentDeferredFlags.Type == FlagsGenerationType::TYPE_NONE) {
// Nothing to do
return;
}
switch (CurrentDeferredFlags.Type) {
case FlagsGenerationType::TYPE_ADC:
CalculcateFlags_ADC(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.ThreeSource.Src1,
CurrentDeferredFlags.Sources.ThreeSource.Src2,
CurrentDeferredFlags.Sources.ThreeSource.Src3);
break;
case FlagsGenerationType::TYPE_SBB:
CalculcateFlags_SBB(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.ThreeSource.Src1,
CurrentDeferredFlags.Sources.ThreeSource.Src2,
CurrentDeferredFlags.Sources.ThreeSource.Src3);
break;
case FlagsGenerationType::TYPE_SUB:
CalculcateFlags_SUB(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.TwoSrcImmediate.Src1,
CurrentDeferredFlags.Sources.TwoSrcImmediate.Src2,
CurrentDeferredFlags.Sources.TwoSrcImmediate.UpdateCF);
break;
case FlagsGenerationType::TYPE_ADD:
CalculcateFlags_ADD(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.TwoSrcImmediate.Src1,
CurrentDeferredFlags.Sources.TwoSrcImmediate.Src2,
CurrentDeferredFlags.Sources.TwoSrcImmediate.UpdateCF);
break;
case FlagsGenerationType::TYPE_MUL:
CalculcateFlags_MUL(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.OneSource.Src1);
break;
case FlagsGenerationType::TYPE_UMUL:
CalculcateFlags_UMUL(CurrentDeferredFlags.Res);
break;
case FlagsGenerationType::TYPE_LOGICAL:
CalculcateFlags_Logical(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.TwoSource.Src1,
CurrentDeferredFlags.Sources.TwoSource.Src2);
break;
case FlagsGenerationType::TYPE_LSHL:
CalculcateFlags_ShiftLeft(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.TwoSource.Src1,
CurrentDeferredFlags.Sources.TwoSource.Src2);
break;
case FlagsGenerationType::TYPE_LSHLI:
CalculcateFlags_ShiftLeftImmediate(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.OneSrcImmediate.Src1,
CurrentDeferredFlags.Sources.OneSrcImmediate.Imm);
break;
case FlagsGenerationType::TYPE_LSHR:
CalculcateFlags_ShiftRight(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.TwoSource.Src1,
CurrentDeferredFlags.Sources.TwoSource.Src2);
break;
case FlagsGenerationType::TYPE_LSHRI:
CalculcateFlags_ShiftRightImmediate(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.OneSrcImmediate.Src1,
CurrentDeferredFlags.Sources.OneSrcImmediate.Imm);
break;
case FlagsGenerationType::TYPE_ASHR:
CalculcateFlags_SignShiftRight(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.TwoSource.Src1,
CurrentDeferredFlags.Sources.TwoSource.Src2);
break;
case FlagsGenerationType::TYPE_ASHRI:
CalculcateFlags_SignShiftRightImmediate(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.OneSrcImmediate.Src1,
CurrentDeferredFlags.Sources.OneSrcImmediate.Imm);
break;
case FlagsGenerationType::TYPE_ROR:
CalculcateFlags_RotateRight(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.TwoSource.Src1,
CurrentDeferredFlags.Sources.TwoSource.Src2);
break;
case FlagsGenerationType::TYPE_RORI:
CalculcateFlags_RotateRightImmediate(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.OneSrcImmediate.Src1,
CurrentDeferredFlags.Sources.OneSrcImmediate.Imm);
break;
case FlagsGenerationType::TYPE_ROL:
CalculcateFlags_RotateLeft(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.TwoSource.Src1,
CurrentDeferredFlags.Sources.TwoSource.Src2);
break;
case FlagsGenerationType::TYPE_ROLI:
CalculcateFlags_RotateLeftImmediate(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.OneSrcImmediate.Src1,
CurrentDeferredFlags.Sources.OneSrcImmediate.Imm);
break;
case FlagsGenerationType::TYPE_FCMP:
CalculcateFlags_FCMP(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.TwoSource.Src1,
CurrentDeferredFlags.Sources.TwoSource.Src2);
break;
case FlagsGenerationType::TYPE_BEXTR:
CalculcateFlags_BEXTR(CurrentDeferredFlags.Res);
break;
case FlagsGenerationType::TYPE_BLSI:
CalculcateFlags_BLSI(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res);
break;
case FlagsGenerationType::TYPE_BLSMSK:
CalculcateFlags_BLSMSK(CurrentDeferredFlags.Res);
break;
case FlagsGenerationType::TYPE_BLSR:
CalculcateFlags_BLSR(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.OneSource.Src1);
break;
case FlagsGenerationType::TYPE_POPCOUNT:
CalculcateFlags_POPCOUNT(CurrentDeferredFlags.Res);
break;
case FlagsGenerationType::TYPE_BZHI:
CalculcateFlags_BZHI(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.OneSource.Src1);
break;
case FlagsGenerationType::TYPE_TZCNT:
CalculcateFlags_TZCNT(CurrentDeferredFlags.Res);
break;
case FlagsGenerationType::TYPE_LZCNT:
CalculcateFlags_LZCNT(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res);
break;
case FlagsGenerationType::TYPE_BITSELECT:
CalculcateFlags_BITSELECT(CurrentDeferredFlags.Res);
break;
case FlagsGenerationType::TYPE_RDRAND:
CalculcateFlags_RDRAND(CurrentDeferredFlags.Res);
break;
case FlagsGenerationType::TYPE_NONE:
default: ERROR_AND_DIE_FMT("Unhandled flags type {}", CurrentDeferredFlags.Type);
}
// Done calculating
CurrentDeferredFlags.Type = FlagsGenerationType::TYPE_NONE;
}
void OpDispatchBuilder::CalculcateFlags_ADC(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, OrderedNode *CF) {
auto Size = SrcSize * 8;
// AF
{
OrderedNode *AFRes = _Xor(_Xor(Src1, Src2), Res);
@@ -79,7 +271,7 @@ void OpDispatchBuilder::GenerateFlags_ADC(FEXCore::X86Tables::DecodedOp Op, Orde
// SF
{
auto SignBitConst = _Constant(GetSrcSize(Op) * 8 - 1);
auto SignBitConst = _Constant(Size - 1);
auto LshrOp = _Lshr(Res, SignBitConst);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(LshrOp);
@@ -140,9 +332,7 @@ void OpDispatchBuilder::GenerateFlags_ADC(FEXCore::X86Tables::DecodedOp Op, Orde
}
}
void OpDispatchBuilder::GenerateFlags_SBB(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, OrderedNode *CF) {
const auto SrcSize = GetSrcSize(Op);
void OpDispatchBuilder::CalculcateFlags_SBB(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, OrderedNode *CF) {
// AF
{
OrderedNode *AFRes = _Xor(_Xor(Src1, Src2), Res);
@@ -212,7 +402,7 @@ void OpDispatchBuilder::GenerateFlags_SBB(FEXCore::X86Tables::DecodedOp Op, Orde
}
}
void OpDispatchBuilder::GenerateFlags_SUB(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, bool UpdateCF) {
void OpDispatchBuilder::CalculcateFlags_SUB(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, bool UpdateCF) {
// AF
{
OrderedNode *AFRes = _Xor(_Xor(Src1, Src2), Res);
@@ -222,7 +412,7 @@ void OpDispatchBuilder::GenerateFlags_SUB(FEXCore::X86Tables::DecodedOp Op, Orde
// SF
{
auto SignBitConst = _Constant(GetSrcSize(Op) * 8 - 1);
auto SignBitConst = _Constant(SrcSize * 8 - 1);
auto LshrOp = _Lshr(Res, SignBitConst);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(LshrOp);
@@ -263,15 +453,13 @@ void OpDispatchBuilder::GenerateFlags_SUB(FEXCore::X86Tables::DecodedOp Op, Orde
auto XorOp2 = _Xor(Res, Src1);
OrderedNode *FinalAnd = _And(XorOp1, XorOp2);
FinalAnd = _Bfe(1, GetSrcSize(Op) * 8 - 1, FinalAnd);
FinalAnd = _Bfe(1, SrcSize * 8 - 1, FinalAnd);
SetRFLAG<FEXCore::X86State::RFLAG_OF_LOC>(FinalAnd);
}
}
void OpDispatchBuilder::GenerateFlags_ADD(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, bool UpdateCF) {
const auto SrcSize = GetSrcSize(Op);
void OpDispatchBuilder::CalculcateFlags_ADD(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, bool UpdateCF) {
// AF
{
OrderedNode *AFRes = _Xor(_Xor(Src1, Src2), Res);
@@ -339,7 +527,7 @@ void OpDispatchBuilder::GenerateFlags_ADD(FEXCore::X86Tables::DecodedOp Op, Orde
}
}
void OpDispatchBuilder::GenerateFlags_MUL(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *High) {
void OpDispatchBuilder::CalculcateFlags_MUL(uint8_t SrcSize, OrderedNode *Res, OrderedNode *High) {
// PF/AF/ZF/SF
// Undefined
{
@@ -354,7 +542,7 @@ void OpDispatchBuilder::GenerateFlags_MUL(FEXCore::X86Tables::DecodedOp Op, Orde
// CF and OF are set if the result of the operation can't be fit in to the destination register
// If the value can fit then the top bits will be zero
auto SignBit = _Sbfe(1, GetSrcSize(Op) * 8 - 1, Res);
auto SignBit = _Sbfe(1, SrcSize * 8 - 1, Res);
auto SelectOp = _Select(FEXCore::IR::COND_EQ, High, SignBit, _Constant(0), _Constant(1));
@@ -363,7 +551,7 @@ void OpDispatchBuilder::GenerateFlags_MUL(FEXCore::X86Tables::DecodedOp Op, Orde
}
}
void OpDispatchBuilder::GenerateFlags_UMUL(FEXCore::X86Tables::DecodedOp Op, OrderedNode *High) {
void OpDispatchBuilder::CalculcateFlags_UMUL(OrderedNode *High) {
// AF/SF/PF/ZF
// Undefined
{
@@ -385,7 +573,7 @@ void OpDispatchBuilder::GenerateFlags_UMUL(FEXCore::X86Tables::DecodedOp Op, Ord
}
}
void OpDispatchBuilder::GenerateFlags_Logical(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
void OpDispatchBuilder::CalculcateFlags_Logical(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
// AF
{
// Undefined
@@ -395,7 +583,7 @@ void OpDispatchBuilder::GenerateFlags_Logical(FEXCore::X86Tables::DecodedOp Op,
// SF
{
auto SignBitConst = _Constant(GetSrcSize(Op) * 8 - 1);
auto SignBitConst = _Constant(SrcSize * 8 - 1);
auto LshrOp = _Lshr(Res, SignBitConst);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(LshrOp);
@@ -430,11 +618,11 @@ auto oldflag = GetRFLAG(FEXCore::X86State::flag);\
auto newval = _Select(FEXCore::IR::COND_EQ, cond, _Constant(0), oldflag, newflag);\
SetRFLAG<FEXCore::X86State::flag>(newval);
void OpDispatchBuilder::GenerateFlags_ShiftLeft(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
void OpDispatchBuilder::CalculcateFlags_ShiftLeft(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
// CF
{
// Extract the last bit shifted in to CF
auto Size = _Constant(GetSrcSize(Op) * 8);
auto Size = _Constant(SrcSize * 8);
auto ShiftAmt = _Sub(Size, Src2);
auto LastBit = _And(_Lshr(Src1, ShiftAmt), _Constant(1));
COND_FLAG_SET(Src2, RFLAG_CF_LOC, LastBit);
@@ -466,7 +654,7 @@ void OpDispatchBuilder::GenerateFlags_ShiftLeft(FEXCore::X86Tables::DecodedOp Op
// SF
{
auto val = _Bfe(1, GetSrcSize(Op) * 8 - 1, Res);
auto val = _Bfe(1, SrcSize * 8 - 1, Res);
COND_FLAG_SET(Src2, RFLAG_SF_LOC, val);
}
@@ -474,12 +662,12 @@ void OpDispatchBuilder::GenerateFlags_ShiftLeft(FEXCore::X86Tables::DecodedOp Op
{
// In the case of left shift. OF is only set from the result of <Top Source Bit> XOR <Top Result Bit>
// When Shift > 1 then OF is undefined
auto val = _Bfe(1, GetSrcSize(Op) * 8 - 1, _Xor(Src1, Res));
auto val = _Bfe(1, SrcSize * 8 - 1, _Xor(Src1, Res));
COND_FLAG_SET(Src2, RFLAG_OF_LOC, val);
}
}
void OpDispatchBuilder::GenerateFlags_ShiftRight(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
void OpDispatchBuilder::CalculcateFlags_ShiftRight(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
// CF
{
// Extract the last bit shifted in to CF
@@ -514,7 +702,7 @@ void OpDispatchBuilder::GenerateFlags_ShiftRight(FEXCore::X86Tables::DecodedOp O
// SF
{
auto val =_Bfe(1, GetSrcSize(Op) * 8 - 1, Res);
auto val =_Bfe(1, SrcSize * 8 - 1, Res);
COND_FLAG_SET(Src2, RFLAG_SF_LOC, val);
}
@@ -522,12 +710,12 @@ void OpDispatchBuilder::GenerateFlags_ShiftRight(FEXCore::X86Tables::DecodedOp O
{
// Only defined when Shift is 1 else undefined
// OF flag is set if a sign change occurred
auto val = _Bfe(1, GetSrcSize(Op) * 8 - 1, _Xor(Src1, Res));
auto val = _Bfe(1, SrcSize * 8 - 1, _Xor(Src1, Res));
COND_FLAG_SET(Src2, RFLAG_OF_LOC, val);
}
}
void OpDispatchBuilder::GenerateFlags_SignShiftRight(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
void OpDispatchBuilder::CalculcateFlags_SignShiftRight(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
// CF
{
// Extract the last bit shifted in to CF
@@ -562,7 +750,7 @@ void OpDispatchBuilder::GenerateFlags_SignShiftRight(FEXCore::X86Tables::Decoded
// SF
{
auto SignBitConst = _Constant(GetSrcSize(Op) * 8 - 1);
auto SignBitConst = _Constant(SrcSize * 8 - 1);
auto LshrOp = _Lshr(Res, SignBitConst);
COND_FLAG_SET(Src2, RFLAG_SF_LOC, LshrOp);
@@ -574,14 +762,14 @@ void OpDispatchBuilder::GenerateFlags_SignShiftRight(FEXCore::X86Tables::Decoded
}
}
void OpDispatchBuilder::GenerateFlags_ShiftLeftImmediate(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift) {
void OpDispatchBuilder::CalculcateFlags_ShiftLeftImmediate(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift) {
// No flags changed if shift is zero
if (Shift == 0) return;
// CF
{
// Extract the last bit shifted in to CF
SetRFLAG<FEXCore::X86State::RFLAG_CF_LOC>(_Bfe(1, GetSrcSize(Op) * 8 - Shift, Src1));
SetRFLAG<FEXCore::X86State::RFLAG_CF_LOC>(_Bfe(1, SrcSize * 8 - Shift, Src1));
}
// PF
@@ -610,20 +798,20 @@ void OpDispatchBuilder::GenerateFlags_ShiftLeftImmediate(FEXCore::X86Tables::Dec
// SF
{
auto LshrOp = _Bfe(1, GetSrcSize(Op) * 8 - 1, Res);
auto LshrOp = _Bfe(1, SrcSize * 8 - 1, Res);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(LshrOp);
// OF
// In the case of left shift. OF is only set from the result of <Top Source Bit> XOR <Top Result Bit>
if (Shift == 1) {
auto SourceBit = _Bfe(1, GetSrcSize(Op) * 8 - 1, Src1);
auto SourceBit = _Bfe(1, SrcSize * 8 - 1, Src1);
SetRFLAG<FEXCore::X86State::RFLAG_OF_LOC>(_Xor(SourceBit, LshrOp));
}
}
}
void OpDispatchBuilder::GenerateFlags_SignShiftRightImmediate(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift) {
void OpDispatchBuilder::CalculcateFlags_SignShiftRightImmediate(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift) {
// No flags changed if shift is zero
if (Shift == 0) return;
@@ -659,7 +847,7 @@ void OpDispatchBuilder::GenerateFlags_SignShiftRightImmediate(FEXCore::X86Tables
// SF
{
auto SignBitConst = _Constant(GetSrcSize(Op) * 8 - 1);
auto SignBitConst = _Constant(SrcSize * 8 - 1);
auto LshrOp = _Lshr(Res, SignBitConst);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(LshrOp);
@@ -674,7 +862,7 @@ void OpDispatchBuilder::GenerateFlags_SignShiftRightImmediate(FEXCore::X86Tables
}
}
void OpDispatchBuilder::GenerateFlags_ShiftRightImmediate(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift) {
void OpDispatchBuilder::CalculcateFlags_ShiftRightImmediate(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift) {
// No flags changed if shift is zero
if (Shift == 0) return;
@@ -710,7 +898,7 @@ void OpDispatchBuilder::GenerateFlags_ShiftRightImmediate(FEXCore::X86Tables::De
// SF
{
auto SignBitConst = _Constant(GetSrcSize(Op) * 8 - 1);
auto SignBitConst = _Constant(SrcSize * 8 - 1);
auto LshrOp = _Lshr(Res, SignBitConst);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(LshrOp);
@@ -721,13 +909,13 @@ void OpDispatchBuilder::GenerateFlags_ShiftRightImmediate(FEXCore::X86Tables::De
// Only defined when Shift is 1 else undefined
// Is set to the MSB of the original value
if (Shift == 1) {
SetRFLAG<FEXCore::X86State::RFLAG_OF_LOC>(_Bfe(1, GetSrcSize(Op) * 8 - 1, Src1));
SetRFLAG<FEXCore::X86State::RFLAG_OF_LOC>(_Bfe(1, SrcSize * 8 - 1, Src1));
}
}
}
void OpDispatchBuilder::GenerateFlags_RotateRight(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
auto OpSize = GetSrcSize(Op) * 8;
void OpDispatchBuilder::CalculcateFlags_RotateRight(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
auto OpSize = SrcSize * 8;
// Extract the last bit shifted in to CF
auto NewCF = _Bfe(1, OpSize - 1, Res);
@@ -755,8 +943,8 @@ void OpDispatchBuilder::GenerateFlags_RotateRight(FEXCore::X86Tables::DecodedOp
}
}
void OpDispatchBuilder::GenerateFlags_RotateLeft(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
auto OpSize = GetSrcSize(Op) * 8;
void OpDispatchBuilder::CalculcateFlags_RotateLeft(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
auto OpSize = SrcSize * 8;
// Extract the last bit shifted in to CF
//auto Size = _Constant(GetSrcSize(Res) * 8);
@@ -785,10 +973,10 @@ void OpDispatchBuilder::GenerateFlags_RotateLeft(FEXCore::X86Tables::DecodedOp O
}
}
void OpDispatchBuilder::GenerateFlags_RotateRightImmediate(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift) {
void OpDispatchBuilder::CalculcateFlags_RotateRightImmediate(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift) {
if (Shift == 0) return;
auto OpSize = GetSrcSize(Op) * 8;
auto OpSize = SrcSize * 8;
auto NewCF = _Bfe(1, OpSize - Shift, Src1);
@@ -807,10 +995,10 @@ void OpDispatchBuilder::GenerateFlags_RotateRightImmediate(FEXCore::X86Tables::D
}
}
void OpDispatchBuilder::GenerateFlags_RotateLeftImmediate(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift) {
void OpDispatchBuilder::CalculcateFlags_RotateLeftImmediate(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift) {
if (Shift == 0) return;
auto OpSize = GetSrcSize(Op) * 8;
auto OpSize = SrcSize * 8;
// CF
{
@@ -827,4 +1015,264 @@ void OpDispatchBuilder::GenerateFlags_RotateLeftImmediate(FEXCore::X86Tables::De
}
}
void OpDispatchBuilder::CalculcateFlags_FCMP(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
OrderedNode *HostFlag_CF = _GetHostFlag(Res, FCMP_FLAG_LT);
OrderedNode *HostFlag_ZF = _GetHostFlag(Res, FCMP_FLAG_EQ);
OrderedNode *HostFlag_Unordered = _GetHostFlag(Res, FCMP_FLAG_UNORDERED);
SetRFLAG<FEXCore::X86State::RFLAG_CF_LOC>(HostFlag_CF);
SetRFLAG<FEXCore::X86State::RFLAG_ZF_LOC>(HostFlag_ZF);
SetRFLAG<FEXCore::X86State::RFLAG_PF_LOC>(HostFlag_Unordered);
auto ZeroConst = _Constant(0);
SetRFLAG<FEXCore::X86State::RFLAG_AF_LOC>(ZeroConst);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(ZeroConst);
SetRFLAG<FEXCore::X86State::RFLAG_OF_LOC>(ZeroConst);
}
void OpDispatchBuilder::CalculcateFlags_BEXTR(OrderedNode *Src) {
// Handle flag setting.
//
// All that matters primarily for this instruction is
// that we only set the ZF flag properly.
//
// CF and OF are defined as being set to zero
//
SetRFLAG<X86State::RFLAG_CF_LOC>(_Constant(0));
SetRFLAG<X86State::RFLAG_OF_LOC>(_Constant(0));
// Every other flag is considered undefined after a
// BEXTR instruction, but we opt to reliably clear them.
//
SetRFLAG<X86State::RFLAG_AF_LOC>(_Constant(0));
SetRFLAG<X86State::RFLAG_SF_LOC>(_Constant(0));
// PF
if (CTX->Config.ABINoPF) {
_InvalidateFlags(1UL << X86State::RFLAG_PF_LOC);
} else {
SetRFLAG<X86State::RFLAG_PF_LOC>(_Constant(0));
}
// ZF
auto ZeroOp = _Select(IR::COND_EQ,
Src, _Constant(0),
_Constant(1), _Constant(0));
SetRFLAG<X86State::RFLAG_ZF_LOC>(ZeroOp);
}
void OpDispatchBuilder::CalculcateFlags_BLSI(uint8_t SrcSize, OrderedNode *Src) {
// Now for the flags:
//
// Only CF, SF, ZF and OF are defined as being updated
// CF is cleared if Src is zero, otherwise it's set.
// SF is set to the value of the most significant operand bit of Result.
// OF is always cleared
// ZF is set, as usual, if Result is zero or not.
//
// AF and PF are documented as being in an undefined state after
// a BLSI operation, however, we choose to reliably clear them.
auto Zero = _Constant(0);
auto One = _Constant(1);
SetRFLAG<X86State::RFLAG_OF_LOC>(Zero);
SetRFLAG<X86State::RFLAG_AF_LOC>(Zero);
if (CTX->Config.ABINoPF) {
_InvalidateFlags(1UL << X86State::RFLAG_PF_LOC);
} else {
SetRFLAG<X86State::RFLAG_PF_LOC>(Zero);
}
// ZF
{
auto ZFOp = _Select(IR::COND_EQ,
Src, Zero,
One, Zero);
SetRFLAG<X86State::RFLAG_ZF_LOC>(ZFOp);
}
// CF
{
auto CFOp = _Select(IR::COND_EQ,
Src, Zero,
Zero, One);
SetRFLAG<X86State::RFLAG_CF_LOC>(CFOp);
}
// SF
{
auto SignBit = _Constant(SrcSize * 8 - 1);
auto SFOp = _Lshr(Src, SignBit);
SetRFLAG<X86State::RFLAG_SF_LOC>(SFOp);
}
}
void OpDispatchBuilder::CalculcateFlags_BLSMSK(OrderedNode *Src) {
// Now for the flags.
auto Zero = _Constant(0);
auto One = _Constant(1);
SetRFLAG<X86State::RFLAG_ZF_LOC>(Zero);
SetRFLAG<X86State::RFLAG_OF_LOC>(Zero);
SetRFLAG<X86State::RFLAG_AF_LOC>(Zero);
if (CTX->Config.ABINoPF) {
_InvalidateFlags(1UL << X86State::RFLAG_PF_LOC);
} else {
SetRFLAG<X86State::RFLAG_PF_LOC>(Zero);
}
auto CFOp = _Select(IR::COND_EQ,
Src, Zero,
Zero, One);
SetRFLAG<X86State::RFLAG_CF_LOC>(CFOp);
}
void OpDispatchBuilder::CalculcateFlags_BLSR(uint8_t SrcSize, OrderedNode *Result, OrderedNode *Src) {
// Now for flags.
auto Zero = _Constant(0);
auto One = _Constant(1);
SetRFLAG<X86State::RFLAG_OF_LOC>(Zero);
SetRFLAG<X86State::RFLAG_AF_LOC>(Zero);
if (CTX->Config.ABINoPF) {
_InvalidateFlags(1UL << X86State::RFLAG_PF_LOC);
} else {
SetRFLAG<X86State::RFLAG_PF_LOC>(Zero);
}
// ZF
{
auto ZFOp = _Select(IR::COND_EQ,
Result, Zero,
One, Zero);
SetRFLAG<X86State::RFLAG_ZF_LOC>(ZFOp);
}
// CF
{
auto CFOp = _Select(IR::COND_EQ,
Src, Zero,
Zero, One);
SetRFLAG<X86State::RFLAG_CF_LOC>(CFOp);
}
// SF
{
auto SignBit = _Constant(SrcSize * 8 - 1);
auto SFOp = _Lshr(Result, SignBit);
SetRFLAG<X86State::RFLAG_SF_LOC>(SFOp);
}
}
void OpDispatchBuilder::CalculcateFlags_POPCOUNT(OrderedNode *Src) {
// Set ZF
auto Zero = _Constant(0);
auto ZFResult = _Select(FEXCore::IR::COND_EQ,
Src, Zero,
_Constant(1), Zero);
// Set flags
SetRFLAG<FEXCore::X86State::RFLAG_CF_LOC>(Zero);
SetRFLAG<FEXCore::X86State::RFLAG_PF_LOC>(Zero);
SetRFLAG<FEXCore::X86State::RFLAG_AF_LOC>(Zero);
SetRFLAG<FEXCore::X86State::RFLAG_ZF_LOC>(ZFResult);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(Zero);
SetRFLAG<FEXCore::X86State::RFLAG_OF_LOC>(Zero);
}
void OpDispatchBuilder::CalculcateFlags_BZHI(uint8_t SrcSize, OrderedNode *Result, OrderedNode *Src) {
// Now for the flags
auto Bounds = _Constant(SrcSize * 8- 1);
auto Zero = _Constant(0);
auto One = _Constant(1);
SetRFLAG<X86State::RFLAG_OF_LOC>(Zero);
SetRFLAG<X86State::RFLAG_AF_LOC>(Zero);
if (CTX->Config.ABINoPF) {
_InvalidateFlags(1UL << X86State::RFLAG_PF_LOC);
} else {
SetRFLAG<X86State::RFLAG_PF_LOC>(Zero);
}
// ZF
{
auto ZFOp = _Select(IR::COND_EQ,
Result, Zero,
One, Zero);
SetRFLAG<X86State::RFLAG_ZF_LOC>(ZFOp);
}
// CF
{
auto CFOp = _Select(IR::COND_UGT,
Src, Bounds,
One, Zero);
SetRFLAG<X86State::RFLAG_CF_LOC>(CFOp);
}
// SF
{
auto SFOp = _Lshr(Result, Bounds);
SetRFLAG<X86State::RFLAG_SF_LOC>(SFOp);
}
}
void OpDispatchBuilder::CalculcateFlags_TZCNT(OrderedNode *Src) {
// OF, SF, AF, PF all undefined
auto Zero = _Constant(0);
auto ZFResult = _Select(FEXCore::IR::COND_EQ,
Src, Zero,
_Constant(1), Zero);
// Set flags
SetRFLAG<FEXCore::X86State::RFLAG_CF_LOC>(ZFResult);
SetRFLAG<FEXCore::X86State::RFLAG_ZF_LOC>(_Bfe(1, 0, Src));
}
void OpDispatchBuilder::CalculcateFlags_LZCNT(uint8_t SrcSize, OrderedNode *Src) {
// OF, SF, AF, PF all undefined
auto Zero = _Constant(0);
auto ZFResult = _Select(FEXCore::IR::COND_EQ,
Src, Zero,
_Constant(1), Zero);
// Set flags
SetRFLAG<FEXCore::X86State::RFLAG_CF_LOC>(ZFResult);
SetRFLAG<FEXCore::X86State::RFLAG_ZF_LOC>(_Bfe(1, SrcSize * 8 - 1, Src));
}
void OpDispatchBuilder::CalculcateFlags_BITSELECT(OrderedNode *Src) {
// OF, SF, AF, PF, CF all undefined
auto ZeroConst = _Constant(0);
auto OneConst = _Constant(1);
// ZF is set to 1 if the source was zero
auto ZFSelectOp = _Select(FEXCore::IR::COND_EQ,
Src, ZeroConst,
OneConst, ZeroConst);
SetRFLAG<FEXCore::X86State::RFLAG_ZF_LOC>(ZFSelectOp);
}
void OpDispatchBuilder::CalculcateFlags_RDRAND(OrderedNode *Src) {
// OF, SF, ZF, AF, PF all zero
// CF is set to the incoming source
auto ZeroConst = _Constant(0);
SetRFLAG<X86State::RFLAG_OF_LOC>(ZeroConst);
SetRFLAG<X86State::RFLAG_SF_LOC>(ZeroConst);
SetRFLAG<X86State::RFLAG_ZF_LOC>(ZeroConst);
SetRFLAG<X86State::RFLAG_AF_LOC>(ZeroConst);
SetRFLAG<X86State::RFLAG_PF_LOC>(ZeroConst);
SetRFLAG<FEXCore::X86State::RFLAG_CF_LOC>(Src);
}
}
@@ -28,6 +28,11 @@ void OpDispatchBuilder::MOVVectorOp(OpcodeArgs) {
StoreResult(FPRClass, Op, Src, 1);
}
void OpDispatchBuilder::MOVVectorNTOp(OpcodeArgs) {
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, 1, true, false, MemoryAccessType::ACCESS_STREAM);
StoreResult(FPRClass, Op, Src, 1, MemoryAccessType::ACCESS_STREAM);
}
void OpDispatchBuilder::MOVAPSOp(OpcodeArgs) {
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
StoreResult(FPRClass, Op, Src, -1);
@@ -241,6 +246,8 @@ void OpDispatchBuilder::VectorALUOp<IR::OP_VCMPGT, 2>(OpcodeArgs);
template
void OpDispatchBuilder::VectorALUOp<IR::OP_VCMPGT, 4>(OpcodeArgs);
template
void OpDispatchBuilder::VectorALUOp<IR::OP_VCMPGT, 8>(OpcodeArgs);
template
void OpDispatchBuilder::VectorALUOp<IR::OP_VCMPEQ, 1>(OpcodeArgs);
template
void OpDispatchBuilder::VectorALUOp<IR::OP_VCMPEQ, 2>(OpcodeArgs);
@@ -295,6 +302,24 @@ void OpDispatchBuilder::VectorALUOp<IR::OP_VUQSUB, 1>(OpcodeArgs);
template
void OpDispatchBuilder::VectorALUOp<IR::OP_VUQSUB, 2>(OpcodeArgs);
template<FEXCore::IR::IROps IROp, size_t ElementSize>
void OpDispatchBuilder::VectorALUROp(OpcodeArgs) {
auto Size = GetSrcSize(Op);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
auto ALUOp = _VAdd(Size, ElementSize, Src, Dest);
// Overwrite our IR's op type
ALUOp.first->Header.Op = IROp;
StoreResult(FPRClass, Op, ALUOp, -1);
}
template
void OpDispatchBuilder::VectorALUROp<IR::OP_VFSUB, 4>(OpcodeArgs);
template
void OpDispatchBuilder::VectorALUROp<IR::OP_VFSUB, 8>(OpcodeArgs);
template<FEXCore::IR::IROps IROp, size_t ElementSize>
void OpDispatchBuilder::VectorScalarALUOp(OpcodeArgs) {
auto Size = GetSrcSize(Op);
@@ -390,14 +415,35 @@ void OpDispatchBuilder::VectorUnaryOp<IR::OP_VABS, 2, false>(OpcodeArgs);
template
void OpDispatchBuilder::VectorUnaryOp<IR::OP_VABS, 4, false>(OpcodeArgs);
template<FEXCore::IR::IROps IROp, size_t ElementSize>
void OpDispatchBuilder::VectorUnaryDuplicateOp(OpcodeArgs) {
auto Size = GetSrcSize(Op);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
auto ALUOp = _VFSqrt(ElementSize, ElementSize, Src);
// Overwrite our IR's op type
ALUOp.first->Header.Op = IROp;
// Duplicate the lower bits
auto Result = _VDupElement(Size, ElementSize, ALUOp, 0);
StoreResult(FPRClass, Op, Result, -1);
}
template
void OpDispatchBuilder::VectorUnaryDuplicateOp<IR::OP_VFRSQRT, 4>(OpcodeArgs);
template
void OpDispatchBuilder::VectorUnaryDuplicateOp<IR::OP_VFRECP, 4>(OpcodeArgs);
void OpDispatchBuilder::MOVQOp(OpcodeArgs) {
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
// This instruction is a bit special that if the destination is a register then it'll ZEXT the 64bit source to 128bit
if (Op->Dest.IsGPR()) {
const auto gpr = Op->Dest.Data.GPR.GPR;
_StoreContext(FPRClass, 8, offsetof(FEXCore::Core::CPUState, xmm[gpr - FEXCore::X86State::REG_XMM_0][0]), Src);
_StoreContext(8, FPRClass, Src, offsetof(FEXCore::Core::CPUState, xmm[gpr - FEXCore::X86State::REG_XMM_0][0]));
auto Const = _Constant(0);
_StoreContext(GPRClass, 8, offsetof(FEXCore::Core::CPUState, xmm[gpr - FEXCore::X86State::REG_XMM_0][1]), Const);
_StoreContext(8, GPRClass, Const, offsetof(FEXCore::Core::CPUState, xmm[gpr - FEXCore::X86State::REG_XMM_0][1]));
}
else {
// This is simple, just store the result
@@ -438,14 +484,15 @@ void OpDispatchBuilder::MOVMSKOpOne(OpcodeArgs) {
//TODO: We could remove this VCastFromGOR + VInsGPR pair if we had a VDUPFromGPR instruction that maps directly to AArch64.
auto M = _Constant(0x80'40'20'10'08'04'02'01ULL);
OrderedNode *VMask = _VCastFromGPR(16, 8, M);
VMask = _VInsGPR(16, 8, VMask, M, 1);
auto VCMP = _VCMPLTZ(Src, 16, 1);
auto VAnd = _VAnd(VCMP, VMask, 16, 1);
VMask = _VInsGPR(16, 8, 1, VMask, M);
auto VAdd1 = _VAddP(VAnd, VAnd, 16, 1);
auto VAdd2 = _VAddP(VAdd1, VAdd1, 8, 1);
auto VAdd3 = _VAddP(VAdd2, VAdd2, 8, 1);
auto VCMP = _VCMPLTZ(16, 1, Src);
auto VAnd = _VAnd(16, 1, VCMP, VMask);
auto VAdd1 = _VAddP(16, 1, VAnd, VAnd);
auto VAdd2 = _VAddP(8, 1, VAdd1, VAdd1);
auto VAdd3 = _VAddP(8, 1, VAdd2, VAdd2);
StoreResult(GPRClass, Op, _VExtractToGPR(16, 2, VAdd3, 0), -1);
}
@@ -502,11 +549,11 @@ void OpDispatchBuilder::PSHUFBOp(OpcodeArgs) {
// Bits [6:4] is reserved for 128bit
// Bits [6:3] is reserved for 64bit
if (Size == 8) {
auto MaskVector = _VectorImm(0b1000'0111, Size, 1);
auto MaskVector = _VectorImm(Size, 1, 0b1000'0111);
Src = _VAnd(Size, Size, Src, MaskVector);
}
else {
auto MaskVector = _VectorImm(0b1000'1111, Size, 1);
auto MaskVector = _VectorImm(Size, 1, 0b1000'1111);
Src = _VAnd(Size, Size, Src, MaskVector);
}
auto Res = _VTBL1(Size, Dest, Src);
@@ -627,7 +674,7 @@ void OpDispatchBuilder::PINSROp(OpcodeArgs) {
Index &= NumElements - 1;
// This maps 1:1 to an AArch64 NEON Op
auto ALUOp = _VInsGPR(Size, ElementSize, Dest, Src, Index);
auto ALUOp = _VInsGPR(Size, ElementSize, Index, Dest, Src);
StoreResult(FPRClass, Op, ALUOp, -1);
}
@@ -670,10 +717,10 @@ void OpDispatchBuilder::InsertPSOp(OpcodeArgs) {
// ZMask happens after insert
if (ZMask == 0xF) {
Dest = _VectorImm(0, 16, 4);
Dest = _VectorImm(16, 4, 0);
}
else if (ZMask) {
auto Zero = _VectorImm(0, 16, 4);
auto Zero = _VectorImm(16, 4, 0);
for (size_t i = 0; i < 4; ++i) {
if (ZMask & (1 << i)) {
Dest = _VInsElement(GetDstSize(Op), 4, i, 0, Dest, Zero);
@@ -708,7 +755,7 @@ void OpDispatchBuilder::PExtrOp(OpcodeArgs) {
}
else {
// If we are storing to memory then we store the size of the element extracted
StoreResult(GPRClass, Op, Result, -1);
StoreResult_WithOpSize(GPRClass, Op, Op->Dest, Result, ElementSize, -1);
}
}
@@ -766,7 +813,7 @@ void OpDispatchBuilder::PSRLDOp(OpcodeArgs) {
OrderedNode *Result{};
// Incoming element size for the shift source is always 8
auto MaxShift = _VectorImm(ElementSize * 8, 8, 8);
auto MaxShift = _VectorImm(8, 8, ElementSize * 8);
Src = _VUMin(8, 8, MaxShift, Src);
Result = _VUShrS(Size, ElementSize, Dest, Src);
@@ -830,7 +877,7 @@ void OpDispatchBuilder::PSLL(OpcodeArgs) {
OrderedNode *Result{};
// Incoming element size for the shift source is always 8
auto MaxShift = _VectorImm(ElementSize * 8, 8, 8);
auto MaxShift = _VectorImm(8, 8, ElementSize * 8);
Src = _VUMin(8, 8, MaxShift, Src);
Result = _VUShlS(Size, ElementSize, Dest, Src);
@@ -854,7 +901,7 @@ void OpDispatchBuilder::PSRAOp(OpcodeArgs) {
OrderedNode *Result{};
// Incoming element size for the shift source is always 8
auto MaxShift = _VectorImm(ElementSize * 8, 8, 8);
auto MaxShift = _VectorImm(8, 8, ElementSize * 8);
Src = _VUMin(8, 8, MaxShift, Src);
Result = _VSShrS(Size, ElementSize, Dest, Src);
@@ -957,12 +1004,11 @@ void OpDispatchBuilder::CVTFPR_To_GPR(OpcodeArgs) {
// Source Element size is determined by instruction
size_t GPRSize = GetDstSize(Op);
size_t ElementSize = SrcElementSize;
if constexpr (HostRoundingMode) {
Src = _Float_ToGPR_S(Src, ElementSize, GPRSize);
Src = _Float_ToGPR_S(GPRSize, SrcElementSize, Src);
}
else {
Src = _Float_ToGPR_ZS(Src, ElementSize, GPRSize);
Src = _Float_ToGPR_ZS(GPRSize, SrcElementSize, Src);
}
StoreResult_WithOpSize(GPRClass, Op, Op->Dest, Src, GPRSize, -1);
@@ -985,11 +1031,11 @@ void OpDispatchBuilder::Vector_CVT_Int_To_Float(OpcodeArgs) {
size_t ElementSize = SrcElementSize;
size_t Size = GetDstSize(Op);
if constexpr (Widen) {
Src = _VSXTL(Src, Size, ElementSize);
Src = _VSXTL(Size, ElementSize, Src);
ElementSize <<= 1;
}
Src = _Vector_SToF(Src, Size, ElementSize);
Src = _Vector_SToF(Size, ElementSize, Src);
StoreResult(FPRClass, Op, Src, -1);
}
@@ -1007,15 +1053,15 @@ void OpDispatchBuilder::Vector_CVT_Float_To_Int(OpcodeArgs) {
size_t Size = GetDstSize(Op);
if constexpr (Narrow) {
Src = _Vector_FToF(Size, SrcElementSize >> 1, SrcElementSize, Src);
Src = _Vector_FToF(Size, SrcElementSize >> 1, Src, SrcElementSize);
ElementSize >>= 1;
}
if constexpr (HostRoundingMode) {
Src = _Vector_FToS(Src, Size, ElementSize);
Src = _Vector_FToS(Size, ElementSize, Src);
}
else {
Src = _Vector_FToZS(Src, Size, ElementSize);
Src = _Vector_FToZS(Size, ElementSize, Src);
}
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, Src, Size, -1);
@@ -1025,6 +1071,8 @@ template
void OpDispatchBuilder::Vector_CVT_Float_To_Int<4, false, false>(OpcodeArgs);
template
void OpDispatchBuilder::Vector_CVT_Float_To_Int<4, false, true>(OpcodeArgs);
template
void OpDispatchBuilder::Vector_CVT_Float_To_Int<4, true, false>(OpcodeArgs);
template
void OpDispatchBuilder::Vector_CVT_Float_To_Int<8, true, true>(OpcodeArgs);
@@ -1053,10 +1101,10 @@ void OpDispatchBuilder::Vector_CVT_Float_To_Float(OpcodeArgs) {
size_t Size = GetDstSize(Op);
if constexpr (DstElementSize > SrcElementSize) {
Src = _Vector_FToF(Size, SrcElementSize << 1, SrcElementSize, Src);
Src = _Vector_FToF(Size, SrcElementSize << 1, Src, SrcElementSize);
}
else {
Src = _Vector_FToF(Size, SrcElementSize >> 1, SrcElementSize, Src);
Src = _Vector_FToF(Size, SrcElementSize >> 1, Src, SrcElementSize);
}
StoreResult(FPRClass, Op, Src, -1);
@@ -1074,12 +1122,12 @@ void OpDispatchBuilder::MMX_To_XMM_Vector_CVT_Int_To_Float(OpcodeArgs) {
size_t ElementSize = SrcElementSize;
size_t DstSize = GetDstSize(Op);
if constexpr (Widen) {
Src = _VSXTL(Src, DstSize, ElementSize);
Src = _VSXTL(DstSize, ElementSize, Src);
ElementSize <<= 1;
}
// Always signed
Src = _Vector_SToF(Src, DstSize, ElementSize);
Src = _Vector_SToF(DstSize, ElementSize, Src);
OrderedNode *Dest{};
if constexpr (Widen) {
@@ -1107,14 +1155,14 @@ void OpDispatchBuilder::XMM_To_MMX_Vector_CVT_Float_To_Int(OpcodeArgs) {
size_t Size = GetDstSize(Op);
// Always narrows
Src = _Vector_FToF(Size, SrcElementSize >> 1, SrcElementSize, Src);
Src = _Vector_FToF(Size, SrcElementSize >> 1, Src, SrcElementSize);
ElementSize >>= 1;
if constexpr (HostRoundingMode) {
Src = _Vector_FToS(Src, Size, ElementSize);
Src = _Vector_FToS(Size, ElementSize, Src);
}
else {
Src = _Vector_FToZS(Src, Size, ElementSize);
Src = _Vector_FToZS(Size, ElementSize, Src);
}
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, Src, Size, -1);
@@ -1133,7 +1181,7 @@ void OpDispatchBuilder::MASKMOVOp(OpcodeArgs) {
OrderedNode *Src = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *Dest = LoadSource(GPRClass, Op, Op->Dest, Op->Flags, -1);
OrderedNode *MemDest = _LoadContext(GPRSize, GPROffset(X86State::REG_RDI), GPRClass);
OrderedNode *MemDest = _LoadContext(GPRSize, GPRClass, GPROffset(X86State::REG_RDI));
const size_t NumElements = Size / 64;
for (size_t Element = 0; Element < NumElements; ++Element) {
@@ -1151,7 +1199,8 @@ void OpDispatchBuilder::MASKMOVOp(OpcodeArgs) {
{
auto DestByte = _Bfe(8, 8 * Select, DestElement);
auto MemLocation = _Add(MemDest, _Constant(Element * 8 + Select));
_StoreMemAutoTSO(GPRClass, 1, MemLocation, DestByte, 1);
// MASKMOVDQU/MASKMOVQ is explicitly weakly-ordered on its store
_StoreMem(GPRClass, 1, MemLocation, DestByte, 1);
}
auto Jump = _Jump();
auto NextJumpTarget = CreateNewCodeBlockAfter(StoreBlock);
@@ -1187,13 +1236,6 @@ void OpDispatchBuilder::VFCMPOp(OpcodeArgs) {
auto Size = GetSrcSize(Op);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *Dest = LoadSource_WithOpSize(FPRClass, Op, Op->Dest, GetDstSize(Op), Op->Flags, -1);
OrderedNode *Src2{};
if constexpr (Scalar) {
Src2 = _VExtractElement(GetDstSize(Op), Size, Dest, 0);
}
else {
Src2 = Dest;
}
uint8_t CompType = Op->Src[1].Data.Literal.Value;
OrderedNode *Result{};
@@ -1201,30 +1243,30 @@ void OpDispatchBuilder::VFCMPOp(OpcodeArgs) {
//auto ALUOp = _VCMPGT(Size, ElementSize, Dest, Src);
switch (CompType) {
case 0x00: case 0x08: case 0x10: case 0x18: // EQ
Result = _VFCMPEQ(Size, ElementSize, Src2, Src);
Result = _VFCMPEQ(Size, ElementSize, Dest, Src);
break;
case 0x01: case 0x09: case 0x11: case 0x19: // LT, GT(Swapped operand)
Result = _VFCMPLT(Size, ElementSize, Src2, Src);
Result = _VFCMPLT(Size, ElementSize, Dest, Src);
break;
case 0x02: case 0x0A: case 0x12: case 0x1A: // LE, GE(Swapped operand)
Result = _VFCMPLE(Size, ElementSize, Src2, Src);
Result = _VFCMPLE(Size, ElementSize, Dest, Src);
break;
case 0x03: case 0x0B: case 0x13: case 0x1B: // Unordered
Result = _VFCMPUNO(Size, ElementSize, Src2, Src);
Result = _VFCMPUNO(Size, ElementSize, Dest, Src);
break;
case 0x04: case 0x0C: case 0x14: case 0x1C: // NEQ
Result = _VFCMPNEQ(Size, ElementSize, Src2, Src);
Result = _VFCMPNEQ(Size, ElementSize, Dest, Src);
break;
case 0x05: case 0x0D: case 0x15: case 0x1D: // NLT, NGT(Swapped operand)
Result = _VFCMPLT(Size, ElementSize, Src2, Src);
Result = _VFCMPLT(Size, ElementSize, Dest, Src);
Result = _VNot(Size, ElementSize, Result);
break;
case 0x06: case 0x0E: case 0x16: case 0x1E: // NLE, NGE(Swapped operand)
Result = _VFCMPLE(Size, ElementSize, Src2, Src);
Result = _VFCMPLE(Size, ElementSize, Dest, Src);
Result = _VNot(Size, ElementSize, Result);
break;
case 0x07: case 0x0F: case 0x17: case 0x1F: // Ordered
Result = _VFCMPORD(Size, ElementSize, Src2, Src);
Result = _VFCMPORD(Size, ElementSize, Dest, Src);
break;
default:
LOGMAN_MSG_A_FMT("Unknown Comparison type: {}", CompType);
@@ -1268,7 +1310,7 @@ void OpDispatchBuilder::FXSaveOp(OpcodeArgs) {
}
{
auto FCW = _LoadContext(2, offsetof(FEXCore::Core::CPUState, FCW), GPRClass);
auto FCW = _LoadContext(2, GPRClass, offsetof(FEXCore::Core::CPUState, FCW));
_StoreMem(GPRClass, 2, Mem, FCW, 2);
}
@@ -1294,7 +1336,7 @@ void OpDispatchBuilder::FXSaveOp(OpcodeArgs) {
{
// FTW
OrderedNode *MemLocation = _Add(Mem, _Constant(4));
auto FTW = _LoadContext(2, offsetof(FEXCore::Core::CPUState, FTW), GPRClass);
auto FTW = _LoadContext(2, GPRClass, offsetof(FEXCore::Core::CPUState, FTW));
_StoreMem(GPRClass, 2, MemLocation, FTW, 2);
}
@@ -1343,7 +1385,7 @@ void OpDispatchBuilder::FXSaveOp(OpcodeArgs) {
// If OSFXSR bit in CR4 is not set than FXSAVE /may/ not save the XMM registers
// This is implementation dependent
for (unsigned i = 0; i < 8; ++i) {
OrderedNode *MMReg = _LoadContext(16, offsetof(FEXCore::Core::CPUState, mm[i]), FPRClass);
OrderedNode *MMReg = _LoadContext(16, FPRClass, offsetof(FEXCore::Core::CPUState, mm[i]));
OrderedNode *MemLocation = _Add(Mem, _Constant(i * 16 + 32));
_StoreMem(FPRClass, 16, MemLocation, MMReg, 16);
@@ -1351,7 +1393,7 @@ void OpDispatchBuilder::FXSaveOp(OpcodeArgs) {
unsigned NumRegs = CTX->Config.Is64BitMode ? 16 : 8;
for (unsigned i = 0; i < NumRegs; ++i) {
OrderedNode *XMMReg = _LoadContext(16, offsetof(FEXCore::Core::CPUState, xmm[i]), FPRClass);
OrderedNode *XMMReg = _LoadContext(16, FPRClass, offsetof(FEXCore::Core::CPUState, xmm[i]));
OrderedNode *MemLocation = _Add(Mem, _Constant(i * 16 + 160));
_StoreMem(FPRClass, 16, MemLocation, XMMReg, 16);
@@ -1364,7 +1406,7 @@ void OpDispatchBuilder::FXRStoreOp(OpcodeArgs) {
auto NewFCW = _LoadMem(GPRClass, 2, Mem, 2);
_F80LoadFCW(NewFCW);
_StoreContext(GPRClass, 2, offsetof(FEXCore::Core::CPUState, FCW), NewFCW);
_StoreContext(2, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
{
OrderedNode *MemLocation = _Add(Mem, _Constant(2));
@@ -1389,20 +1431,20 @@ void OpDispatchBuilder::FXRStoreOp(OpcodeArgs) {
// FTW
OrderedNode *MemLocation = _Add(Mem, _Constant(4));
auto NewFTW = _LoadMem(GPRClass, 2, MemLocation, 2);
_StoreContext(GPRClass, 2, offsetof(FEXCore::Core::CPUState, FTW), NewFTW);
_StoreContext(2, GPRClass, NewFTW, offsetof(FEXCore::Core::CPUState, FTW));
}
for (unsigned i = 0; i < 8; ++i) {
OrderedNode *MemLocation = _Add(Mem, _Constant(i * 16 + 32));
auto MMReg = _LoadMem(FPRClass, 16, MemLocation, 16);
_StoreContext(FPRClass, 16, offsetof(FEXCore::Core::CPUState, mm[i]), MMReg);
_StoreContext(16, FPRClass, MMReg, offsetof(FEXCore::Core::CPUState, mm[i]));
}
unsigned NumRegs = CTX->Config.Is64BitMode ? 16 : 8;
for (unsigned i = 0; i < NumRegs; ++i) {
OrderedNode *MemLocation = _Add(Mem, _Constant(i * 16 + 160));
auto XMMReg = _LoadMem(FPRClass, 16, MemLocation, 16);
_StoreContext(FPRClass, 16, offsetof(FEXCore::Core::CPUState, xmm[i]), XMMReg);
_StoreContext(16, FPRClass, XMMReg, offsetof(FEXCore::Core::CPUState, xmm[i]));
}
}
@@ -1427,23 +1469,12 @@ template<size_t ElementSize>
void OpDispatchBuilder::UCOMISxOp(OpcodeArgs) {
OrderedNode *Src1 = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
OrderedNode *Src2 = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *Res = _FCmp(Src1, Src2, ElementSize,
OrderedNode *Res = _FCmp(ElementSize, Src1, Src2,
(1 << FCMP_FLAG_EQ) |
(1 << FCMP_FLAG_LT) |
(1 << FCMP_FLAG_UNORDERED));
OrderedNode *HostFlag_CF = _GetHostFlag(Res, FCMP_FLAG_LT);
OrderedNode *HostFlag_ZF = _GetHostFlag(Res, FCMP_FLAG_EQ);
OrderedNode *HostFlag_Unordered = _GetHostFlag(Res, FCMP_FLAG_UNORDERED);
SetRFLAG<FEXCore::X86State::RFLAG_CF_LOC>(HostFlag_CF);
SetRFLAG<FEXCore::X86State::RFLAG_ZF_LOC>(HostFlag_ZF);
SetRFLAG<FEXCore::X86State::RFLAG_PF_LOC>(HostFlag_Unordered);
auto ZeroConst = _Constant(0);
SetRFLAG<FEXCore::X86State::RFLAG_AF_LOC>(ZeroConst);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(ZeroConst);
SetRFLAG<FEXCore::X86State::RFLAG_OF_LOC>(ZeroConst);
GenerateFlags_FCMP(Op, Res, Src1, Src2);
flagsOp = SelectionFlag::FCMP;
flagsOpDest = Src1;
@@ -1559,8 +1590,8 @@ void OpDispatchBuilder::MOVQ2DQ(OpcodeArgs) {
// This instruction is a bit special in that if the source is MMX then it zexts to 128bit
if constexpr (ToXMM) {
Src = _VMov(Src, 16);
_StoreContext(FPRClass, 16, offsetof(FEXCore::Core::CPUState, xmm[Op->Dest.Data.GPR.GPR - FEXCore::X86State::REG_XMM_0][0]), Src);
Src = _VMov(16, Src);
_StoreContext(16, FPRClass, Src, offsetof(FEXCore::Core::CPUState, xmm[Op->Dest.Data.GPR.GPR - FEXCore::X86State::REG_XMM_0][0]));
}
else {
// This is simple, just store the result
@@ -1649,11 +1680,150 @@ void OpDispatchBuilder::ADDSUBPOp(OpcodeArgs) {
StoreResult(FPRClass, Op, ResAdd, -1);
}
void OpDispatchBuilder::PFNACCOp(OpcodeArgs) {
auto Size = GetSrcSize(Op);
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *ResSubSrc{};
OrderedNode *ResSubDest{};
auto UpperSubDest = _VExtractElement(Size, 4, Dest, 1);
auto UpperSubSrc = _VExtractElement(Size, 4, Src, 1);
ResSubDest = _VFSub(4, 4, Dest, UpperSubDest);
ResSubSrc = _VFSub(4, 4, Src, UpperSubSrc);
auto Result = _VInsElement(8, 4, 1, 0, ResSubDest, ResSubSrc);
StoreResult(FPRClass, Op, Result, -1);
}
void OpDispatchBuilder::PFPNACCOp(OpcodeArgs) {
auto Size = GetSrcSize(Op);
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *ResAdd{};
OrderedNode *ResSub{};
auto UpperSubDest = _VExtractElement(Size, 4, Dest, 1);
ResSub = _VFSub(4, 4, Dest, UpperSubDest);
ResAdd = _VFAddP(Size, 4, Src, Src);
auto Result = _VInsElement(8, 4, 1, 0, ResSub, ResAdd);
StoreResult(FPRClass, Op, Result, -1);
}
void OpDispatchBuilder::PSWAPDOp(OpcodeArgs) {
auto Size = GetSrcSize(Op);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
auto Result = _VRev64(Size, 4, Src);
StoreResult(FPRClass, Op, Result, -1);
}
template
void OpDispatchBuilder::ADDSUBPOp<4>(OpcodeArgs);
template
void OpDispatchBuilder::ADDSUBPOp<8>(OpcodeArgs);
void OpDispatchBuilder::PI2FWOp(OpcodeArgs) {
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
size_t Size = GetDstSize(Op);
// We now need to transpose the lower 16-bits of each element together
// Only needing to move the upper element down in this case
Src = _VInsElement(Size, 2, 1, 2, Src, Src);
// Now we need to sign extend the 16bit value to 32-bit
Src = _VSXTL(Size, 2, Src);
// int32_t to float
Src = _Vector_SToF(Size, 4, Src);
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, Src, Size, -1);
}
void OpDispatchBuilder::PF2IWOp(OpcodeArgs) {
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
size_t Size = GetDstSize(Op);
// Float to int32_t
Src = _Vector_FToZS(Size, 4, Src);
// We now need to transpose the lower 16-bits of each element together
// Only needing to move the upper element down in this case
Src = _VInsElement(Size, 2, 1, 2, Src, Src);
// Now we need to sign extend the 16bit value to 32-bit
Src = _VSXTL(Size, 2, Src);
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, Src, Size, -1);
}
void OpDispatchBuilder::PMULHRWOp(OpcodeArgs) {
auto Size = GetSrcSize(Op);
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *Res{};
// Implementation is more efficient for 8byte registers
// Multiplies 4 16bit values in to 4 32bit values
Res = _VSMull(Size * 2, 2, Dest, Src);
//TODO: We could remove this VCastFromGOR + VInsGPR pair if we had a VDUPFromGPR instruction that maps directly to AArch64.
auto M = _Constant(0x0000'8000'0000'8000ULL);
OrderedNode *VConstant = _VCastFromGPR(16, 8, M);
VConstant = _VInsGPR(16, 8, 1, VConstant, M);
Res = _VAdd(Size * 2, 4, Res, VConstant);
// Now shift and narrow to convert 32-bit values to 16bit, storing the top 16bits
Res = _VUShrNI(Size * 2, 4, Res, 16);
StoreResult(FPRClass, Op, Res, -1);
}
template<uint8_t CompType>
void OpDispatchBuilder::VPFCMPOp(OpcodeArgs) {
auto Size = GetSrcSize(Op);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *Dest = LoadSource_WithOpSize(FPRClass, Op, Op->Dest, GetDstSize(Op), Op->Flags, -1);
OrderedNode *Result{};
// This maps 1:1 to an AArch64 NEON Op
//auto ALUOp = _VCMPGT(Size, 4, Dest, Src);
LogMan::Msg::DFmt("CompType: {} Size: {}", CompType, Size);
switch (CompType) {
case 0x00: // EQ
Result = _VFCMPEQ(Size, 4, Dest, Src);
break;
case 0x01: // GE(Swapped operand)
Result = _VFCMPLE(Size, 4, Src, Dest);
break;
case 0x02: // GT
Result = _VFCMPGT(Size, 4, Dest, Src);
break;
default:
LOGMAN_MSG_A_FMT("Unknown Comparison type: {}", CompType);
break;
}
StoreResult(FPRClass, Op, Result, -1);
}
template
void OpDispatchBuilder::VPFCMPOp<0>(OpcodeArgs);
template
void OpDispatchBuilder::VPFCMPOp<1>(OpcodeArgs);
template
void OpDispatchBuilder::VPFCMPOp<2>(OpcodeArgs);
void OpDispatchBuilder::PMADDWD(OpcodeArgs) {
// This is a pretty curious operation
// Does two MADD operations across 4 16bit signed integers and accumulates to 32bit integers in the destination
@@ -1806,7 +1976,7 @@ void OpDispatchBuilder::PMULHRSW(OpcodeArgs) {
// Implementation is more efficient for 8byte registers
Res = _VSMull(Size * 2, 2, Dest, Src);
Res = _VSShrI(Size * 2, 4, Res, 14);
auto OneVector = _VectorImm(1, Size * 2, 4);
auto OneVector = _VectorImm(Size * 2, 4, 1);
Res = _VAdd(Size * 2, 4, Res, OneVector);
Res = _VUShrNI(Size * 2, 4, Res, 1);
}
@@ -1820,7 +1990,7 @@ void OpDispatchBuilder::PMULHRSW(OpcodeArgs) {
ResultLow = _VSShrI(Size, 4, ResultLow, 14);
ResultHigh = _VSShrI(Size, 4, ResultHigh, 14);
auto OneVector = _VectorImm(1, Size, 4);
auto OneVector = _VectorImm(Size, 4, 1);
ResultLow = _VAdd(Size, 4, ResultLow, OneVector);
ResultHigh = _VAdd(Size, 4, ResultHigh, OneVector);
@@ -2075,10 +2245,10 @@ void OpDispatchBuilder::ExtendVectorElements(OpcodeArgs) {
CurrentElementSize != DstElementSize;
CurrentElementSize <<= 1) {
if constexpr (Signed) {
Result = _VSXTL(Result, Size, CurrentElementSize);
Result = _VSXTL(Size, CurrentElementSize, Result);
}
else {
Result = _VUXTL(Result, Size, CurrentElementSize);
Result = _VUXTL(Size, CurrentElementSize, Result);
}
}
StoreResult(FPRClass, Op, Result, -1);
@@ -2132,7 +2302,7 @@ void OpDispatchBuilder::VectorRound(OpcodeArgs) {
FEXCore::IR::Round_Host,
};
Src = _Vector_FToI(Src, SourceModes[(RoundControlSource << 2) | RoundControl], Size, ElementSize);
Src = _Vector_FToI(Size, ElementSize, Src, SourceModes[(RoundControlSource << 2) | RoundControl]);
if constexpr (Scalar) {
// Insert the lower bits
@@ -2186,13 +2356,14 @@ void OpDispatchBuilder::VectorVariableBlend(OpcodeArgs) {
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
// The mask is hardcoded to be xmm0 in this instruction
OrderedNode *Mask = _LoadContext(16, offsetof(FEXCore::Core::CPUState, xmm[0]), FPRClass);
OrderedNode *Mask = _LoadContext(16, FPRClass, offsetof(FEXCore::Core::CPUState, xmm[0]));
// Each element is selected by the high bit of that element size
// Dest[ElementIdx] = Xmm0[ElementIndex][HighBit] ? Src : Dest;
//
// To emulate this on AArch64
// Arithmetic shift right by the element size, then use BSL to select the registers
Mask = _VSShrI(Size, ElementSize, Mask, (ElementSize * 8) - 1);
auto Result = _VBSL(Mask, Src, Dest);
StoreResult(FPRClass, Op, Result, -1);
@@ -2205,13 +2376,16 @@ template
void OpDispatchBuilder::VectorVariableBlend<8>(OpcodeArgs);
void OpDispatchBuilder::PTestOp(OpcodeArgs) {
// Invalidate deferred flags early
InvalidateDeferredFlags();
auto Size = GetSrcSize(Op);
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *Test1 = _VAnd(Dest, Src, Size, 1);
OrderedNode *Test2 = _VBic(Src, Dest, Size, 1);
OrderedNode *Test1 = _VAnd(Size, 1, Dest, Src);
OrderedNode *Test2 = _VBic(Size, 1, Src, Dest);
Test1 = _VPopcount(Size, 1, Test1);
Test2 = _VPopcount(Size, 1, Test2);
@@ -2233,8 +2407,14 @@ void OpDispatchBuilder::PTestOp(OpcodeArgs) {
Test2 = _Select(FEXCore::IR::COND_EQ,
Test2, ZeroConst, OneConst, ZeroConst);
// Careful, these flags are different between {V,}PTEST and VTESTP{S,D}
SetRFLAG<FEXCore::X86State::RFLAG_ZF_LOC>(Test1);
SetRFLAG<FEXCore::X86State::RFLAG_CF_LOC>(Test2);
SetRFLAG<FEXCore::X86State::RFLAG_AF_LOC>(ZeroConst);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(ZeroConst);
SetRFLAG<FEXCore::X86State::RFLAG_OF_LOC>(ZeroConst);
SetRFLAG<FEXCore::X86State::RFLAG_PF_LOC>(ZeroConst);
}
void OpDispatchBuilder::PHMINPOSUWOp(OpcodeArgs) {
@@ -2267,10 +2447,10 @@ void OpDispatchBuilder::PHMINPOSUWOp(OpcodeArgs) {
}
// Insert the minimum in to bits [15:0]
OrderedNode *Result = _VMov(Min, 2);
OrderedNode *Result = _VMov(2, Min);
// Insert position in to bits [18:16]
Result = _VInsGPR(16, 2, Result, Pos, 1);
Result = _VInsGPR(16, 2, 1, Result, Pos);
StoreResult(FPRClass, Op, Result, -1);
}
@@ -25,12 +25,12 @@ class OrderedNode;
OrderedNode *OpDispatchBuilder::GetX87Top() {
// Yes, we are storing 3 bits in a single flag register.
// Deal with it
return _LoadContext(1, offsetof(FEXCore::Core::CPUState, flags) + FEXCore::X86State::X87FLAG_TOP_LOC, GPRClass);
return _LoadContext(1, GPRClass, offsetof(FEXCore::Core::CPUState, flags) + FEXCore::X86State::X87FLAG_TOP_LOC);
}
void OpDispatchBuilder::SetX87TopTag(OrderedNode *Value, X87Tag Tag) {
// if we are popping then we must first mark this location as empty
auto FTW = _LoadContext(2, offsetof(FEXCore::Core::CPUState, FTW), GPRClass);
auto FTW = _LoadContext(2, GPRClass, offsetof(FEXCore::Core::CPUState, FTW));
OrderedNode *Mask = _Constant(0b11);
auto TopOffset = _Lshl(Value, _Constant(1));
Mask = _Lshl(Mask, TopOffset);
@@ -40,11 +40,11 @@ void OpDispatchBuilder::SetX87TopTag(OrderedNode *Value, X87Tag Tag) {
NewFTW = _Or(NewFTW, TagVal);
}
_StoreContext(GPRClass, 2, offsetof(FEXCore::Core::CPUState, FTW), NewFTW);
_StoreContext(2, GPRClass, NewFTW, offsetof(FEXCore::Core::CPUState, FTW));
}
OrderedNode *OpDispatchBuilder::GetX87FTW(OrderedNode *Value) {
auto FTW = _LoadContext(2, offsetof(FEXCore::Core::CPUState, FTW), GPRClass);
auto FTW = _LoadContext(2, GPRClass, offsetof(FEXCore::Core::CPUState, FTW));
OrderedNode *Mask = _Constant(0b11);
auto TopOffset = _Lshl(Value, _Constant(1));
auto NewFTW = _Lshr(FTW, TopOffset);
@@ -52,7 +52,7 @@ OrderedNode *OpDispatchBuilder::GetX87FTW(OrderedNode *Value) {
}
void OpDispatchBuilder::SetX87Top(OrderedNode *Value) {
_StoreContext(GPRClass, 1, offsetof(FEXCore::Core::CPUState, flags) + FEXCore::X86State::X87FLAG_TOP_LOC, Value);
_StoreContext(1, GPRClass, Value, offsetof(FEXCore::Core::CPUState, flags) + FEXCore::X86State::X87FLAG_TOP_LOC);
}
template<size_t width>
@@ -136,7 +136,7 @@ void OpDispatchBuilder::FLD_Const(OpcodeArgs) {
auto low = _Constant(Lower);
auto high = _Constant(Upper);
OrderedNode *data = _VCastFromGPR(16, 8, low);
data = _VInsGPR(16, 8, data, high, 1);
data = _VInsGPR(16, 8, 1, data, high);
// Write to ST[TOP]
_StoreContextIndexed(data, top, 16, MMBaseOffset(), 16, FPRClass);
}
@@ -202,7 +202,7 @@ void OpDispatchBuilder::FST(OpcodeArgs) {
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, data, 10, 1);
}
else if constexpr (width == 32 || width == 64) {
auto result = _F80CVT(data, width / 8);
auto result = _F80CVT(width / 8, data);
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, result, width / 8, 1);
}
@@ -228,7 +228,7 @@ void OpDispatchBuilder::FIST(OpcodeArgs) {
auto orig_top = GetX87Top();
OrderedNode *data = _LoadContextIndexed(orig_top, 16, MMBaseOffset(), 16, FPRClass);
data = _F80CVTInt(data, Truncate, Size);
data = _F80CVTInt(Size, data, Truncate);
StoreResult_WithOpSize(GPRClass, Op, Op->Dest, data, Size, 1);
@@ -547,9 +547,9 @@ void OpDispatchBuilder::FCHS(OpcodeArgs) {
auto low = _Constant(0);
auto high = _Constant(0b1'000'0000'0000'0000ULL);
OrderedNode *data = _VCastFromGPR(16, 8, low);
data = _VInsGPR(16, 8, data, high, 1);
data = _VInsGPR(16, 8, 1, data, high);
auto result = _VXor(a, data, 16, 1);
auto result = _VXor(16, 1, a, data);
// Write to ST[TOP]
_StoreContextIndexed(result, top, 16, MMBaseOffset(), 16, FPRClass);
@@ -562,9 +562,9 @@ void OpDispatchBuilder::FABS(OpcodeArgs) {
auto low = _Constant(~0ULL);
auto high = _Constant(0b0'111'1111'1111'1111ULL);
OrderedNode *data = _VCastFromGPR(16, 8, low);
data = _VInsGPR(16, 8, data, high, 1);
data = _VInsGPR(16, 8, 1, data, high);
auto result = _VAnd(a, data, 16, 1);
auto result = _VAnd(16, 1, a, data);
// Write to ST[TOP]
_StoreContextIndexed(result, top, 16, MMBaseOffset(), 16, FPRClass);
@@ -621,10 +621,10 @@ void OpDispatchBuilder::FXTRACT(OpcodeArgs) {
}
void OpDispatchBuilder::FNINIT(OpcodeArgs) {
// Init FCW to 0x037
auto NewFCW = _Constant(16, 0x037);
// Init FCW to 0x037F
auto NewFCW = _Constant(16, 0x037F);
_F80LoadFCW(NewFCW);
_StoreContext(GPRClass, 2, offsetof(FEXCore::Core::CPUState, FCW), NewFCW);
_StoreContext(2, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
// Init FSW to 0
SetX87Top(_Constant(0));
@@ -635,7 +635,7 @@ void OpDispatchBuilder::FNINIT(OpcodeArgs) {
SetRFLAG<FEXCore::X86State::X87FLAG_C3_LOC>(_Constant(0));
// Tags all get set to 0b11
_StoreContext(GPRClass, 2, offsetof(FEXCore::Core::CPUState, FTW), _Constant(0xFFFF));
_StoreContext(2, GPRClass, _Constant(0xFFFF), offsetof(FEXCore::Core::CPUState, FTW));
}
template<size_t width, bool Integer, OpDispatchBuilder::FCOMIFlags whichflags, bool poptwice>
@@ -685,12 +685,15 @@ void OpDispatchBuilder::FCOMI(OpcodeArgs) {
SetRFLAG<FEXCore::X86State::X87FLAG_C3_LOC>(HostFlag_ZF);
}
else {
// Invalidate deferred flags early
// OF, SF, AF, PF all undefined
InvalidateDeferredFlags();
SetRFLAG<FEXCore::X86State::RFLAG_CF_LOC>(HostFlag_CF);
SetRFLAG<FEXCore::X86State::RFLAG_ZF_LOC>(HostFlag_ZF);
SetRFLAG<FEXCore::X86State::RFLAG_PF_LOC>(HostFlag_Unordered);
}
if constexpr (poptwice) {
// if we are popping then we must first mark this location as empty
SetX87TopTag(top, X87Tag::Empty);
@@ -872,7 +875,7 @@ void OpDispatchBuilder::X87FYL2X(OpcodeArgs) {
auto low = _Constant(0x8000'0000'0000'0000ULL);
auto high = _Constant(0b0'011'1111'1111'1111);
OrderedNode *data = _VCastFromGPR(16, 8, low);
data = _VInsGPR(16, 8, data, high, 1);
data = _VInsGPR(16, 8, 1, data, high);
st0 = _F80Add(st0, data);
}
@@ -895,7 +898,7 @@ void OpDispatchBuilder::X87TAN(OpcodeArgs) {
auto low = _Constant(0x8000'0000'0000'0000ULL);
auto high = _Constant(0b0'011'1111'1111'1111ULL);
OrderedNode *data = _VCastFromGPR(16, 8, low);
data = _VInsGPR(16, 8, data, high, 1);
data = _VInsGPR(16, 8, 1, data, high);
// Write to ST[TOP]
_StoreContextIndexed(result, orig_top, 16, MMBaseOffset(), 16, FPRClass);
@@ -925,7 +928,7 @@ void OpDispatchBuilder::X87LDENV(OpcodeArgs) {
auto NewFCW = _LoadMem(GPRClass, 2, Mem, 2);
_F80LoadFCW(NewFCW);
_StoreContext(GPRClass, 2, offsetof(FEXCore::Core::CPUState, FCW), NewFCW);
_StoreContext(2, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
OrderedNode *MemLocation = _Add(Mem, _Constant(Size * 1));
auto NewFSW = _LoadMem(GPRClass, Size, MemLocation, Size);
@@ -948,7 +951,7 @@ void OpDispatchBuilder::X87LDENV(OpcodeArgs) {
// FTW
OrderedNode *MemLocation = _Add(Mem, _Constant(Size * 2));
auto NewFTW = _LoadMem(GPRClass, Size, MemLocation, Size);
_StoreContext(GPRClass, 2, offsetof(FEXCore::Core::CPUState, FTW), NewFTW);
_StoreContext(2, GPRClass, NewFTW, offsetof(FEXCore::Core::CPUState, FTW));
}
}
@@ -977,7 +980,7 @@ void OpDispatchBuilder::X87FNSTENV(OpcodeArgs) {
Mem = AppendSegmentOffset(Mem, Op->Flags);
{
auto FCW = _LoadContext(2, offsetof(FEXCore::Core::CPUState, FCW), GPRClass);
auto FCW = _LoadContext(2, GPRClass, offsetof(FEXCore::Core::CPUState, FCW));
_StoreMem(GPRClass, Size, Mem, FCW, Size);
}
@@ -1005,7 +1008,7 @@ void OpDispatchBuilder::X87FNSTENV(OpcodeArgs) {
{
// FTW
OrderedNode *MemLocation = _Add(Mem, _Constant(Size * 2));
auto FTW = _LoadContext(2, offsetof(FEXCore::Core::CPUState, FTW), GPRClass);
auto FTW = _LoadContext(2, GPRClass, offsetof(FEXCore::Core::CPUState, FTW));
_StoreMem(GPRClass, Size, MemLocation, FTW, Size);
}
@@ -1037,11 +1040,11 @@ void OpDispatchBuilder::X87FNSTENV(OpcodeArgs) {
void OpDispatchBuilder::X87FLDCW(OpcodeArgs) {
OrderedNode *NewFCW = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags, -1);
_F80LoadFCW(NewFCW);
_StoreContext(GPRClass, 2, offsetof(FEXCore::Core::CPUState, FCW), NewFCW);
_StoreContext(2, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
}
void OpDispatchBuilder::X87FSTCW(OpcodeArgs) {
auto FCW = _LoadContext(2, offsetof(FEXCore::Core::CPUState, FCW), GPRClass);
auto FCW = _LoadContext(2, GPRClass, offsetof(FEXCore::Core::CPUState, FCW));
StoreResult(GPRClass, Op, FCW, -1);
}
@@ -1108,7 +1111,7 @@ void OpDispatchBuilder::X87FNSAVE(OpcodeArgs) {
OrderedNode *Top = GetX87Top();
{
auto FCW = _LoadContext(2, offsetof(FEXCore::Core::CPUState, FCW), GPRClass);
auto FCW = _LoadContext(2, GPRClass, offsetof(FEXCore::Core::CPUState, FCW));
_StoreMem(GPRClass, Size, Mem, FCW, Size);
}
@@ -1135,7 +1138,7 @@ void OpDispatchBuilder::X87FNSAVE(OpcodeArgs) {
{
// FTW
OrderedNode *MemLocation = _Add(Mem, _Constant(Size * 2));
auto FTW = _LoadContext(2, offsetof(FEXCore::Core::CPUState, FTW), GPRClass);
auto FTW = _LoadContext(2, GPRClass, offsetof(FEXCore::Core::CPUState, FTW));
_StoreMem(GPRClass, Size, MemLocation, FTW, Size);
}
@@ -1196,7 +1199,7 @@ void OpDispatchBuilder::X87FRSTOR(OpcodeArgs) {
auto NewFCW = _LoadMem(GPRClass, 2, Mem, 2);
_F80LoadFCW(NewFCW);
_StoreContext(GPRClass, 2, offsetof(FEXCore::Core::CPUState, FCW), NewFCW);
_StoreContext(2, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
OrderedNode *MemLocation = _Add(Mem, _Constant(Size * 1));
auto NewFSW = _LoadMem(GPRClass, Size, MemLocation, Size);
@@ -1219,7 +1222,7 @@ void OpDispatchBuilder::X87FRSTOR(OpcodeArgs) {
// FTW
OrderedNode *MemLocation = _Add(Mem, _Constant(Size * 2));
auto NewFTW = _LoadMem(GPRClass, Size, MemLocation, Size);
_StoreContext(GPRClass, 2, offsetof(FEXCore::Core::CPUState, FTW), NewFTW);
_StoreContext(2, GPRClass, NewFTW, offsetof(FEXCore::Core::CPUState, FTW));
}
OrderedNode *ST0Location = _Add(Mem, _Constant(Size * 7));
@@ -1231,7 +1234,7 @@ void OpDispatchBuilder::X87FRSTOR(OpcodeArgs) {
auto low = _Constant(~0ULL);
auto high = _Constant(0xFFFF);
OrderedNode *Mask = _VCastFromGPR(16, 8, low);
Mask = _VInsGPR(16, 8, Mask, high, 1);
Mask = _VInsGPR(16, 8, 1, Mask, high);
for (int i = 0; i < 7; ++i) {
OrderedNode *Reg = _LoadMem(FPRClass, 16, ST0Location, 1);
@@ -1359,7 +1362,7 @@ void OpDispatchBuilder::X87FCMOV(OpcodeArgs) {
SrcCond = _Sbfe(1, 0, SrcCond);
OrderedNode *VecCond = _VCastFromGPR(16, 8, SrcCond);
VecCond = _VInsGPR(16, 8, VecCond, SrcCond, 1);
VecCond = _VInsGPR(16, 8, 1, VecCond, SrcCond);
auto top = GetX87Top();
OrderedNode* arg;
@@ -1380,7 +1383,7 @@ void OpDispatchBuilder::X87FCMOV(OpcodeArgs) {
void OpDispatchBuilder::X87EMMS(OpcodeArgs) {
// Tags all get set to 0b11
_StoreContext(GPRClass, 2, offsetof(FEXCore::Core::CPUState, FTW), _Constant(0xFFFF));
_StoreContext(2, GPRClass, _Constant(0xFFFF), offsetof(FEXCore::Core::CPUState, FTW));
}
void OpDispatchBuilder::X87FFREE(OpcodeArgs) {
File diff suppressed because it is too large. Load diff
+5 -4
View File
@@ -95,10 +95,11 @@ namespace FEXCore {
LogMan::Msg::EFmt("[{}] Thread has received a signal and hasn't registered itself with the delegate! Programming error!", FHU::Syscalls::gettid());
}
else {
if (Handler.Handler &&
Handler.Handler(Thread, Signal, Info, UContext)) {
// If the host handler handled the fault then we can continue now
return;
for (auto &Handler : Handler.Handlers) {
if (Handler(Thread, Signal, Info, UContext)) {
// If the host handler handled the fault then we can continue now
return;
}
}
if (Handler.FrontendHandler &&
@@ -167,7 +167,7 @@ void InitializeBaseTables(Context::OperatingMode Mode) {
{0xC2, 1, X86InstInfo{"RET", TYPE_INST, FLAGS_SETS_RIP | FLAGS_BLOCK_END, 2, nullptr}},
{0xC3, 1, X86InstInfo{"RET", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_BLOCK_END , 0, nullptr}},
{0xC8, 1, X86InstInfo{"ENTER", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_DEBUG_MEM_ACCESS , 3, nullptr}},
{0xC9, 1, X86InstInfo{"LEAVE", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_BLOCK_END , 0, nullptr}},
{0xC9, 1, X86InstInfo{"LEAVE", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_DEBUG_MEM_ACCESS , 0, nullptr}},
{0xCA, 2, X86InstInfo{"RETF", TYPE_PRIV, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_BLOCK_END, 0, nullptr}},
{0xCC, 1, X86InstInfo{"INT3", TYPE_INST, FLAGS_DEBUG, 0, nullptr}},
{0xCD, 1, X86InstInfo{"INT", TYPE_INST, FLAGS_DEBUG , 1, nullptr}},
@@ -15,37 +15,42 @@ using namespace InstFlags;
void InitializeDDDTables() {
static constexpr U8U8InfoStruct DDDNowOpTable[] = {
{0x0C, 1, X86InstInfo{"PI2FW", TYPE_3DNOW_INST, FLAGS_MODRM, 0, nullptr}},
{0x0D, 1, X86InstInfo{"PI2FD", TYPE_3DNOW_INST, FLAGS_MODRM, 0, nullptr}},
{0x1C, 1, X86InstInfo{"PF2IW", TYPE_3DNOW_INST, FLAGS_MODRM, 0, nullptr}},
{0x1D, 1, X86InstInfo{"PF2ID", TYPE_3DNOW_INST, FLAGS_MODRM, 0, nullptr}},
{0x0C, 1, X86InstInfo{"PI2FW", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0x0D, 1, X86InstInfo{"PI2FD", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0x1C, 1, X86InstInfo{"PF2IW", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0x1D, 1, X86InstInfo{"PF2ID", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0x8A, 1, X86InstInfo{"PFNACC", TYPE_3DNOW_INST, FLAGS_MODRM, 0, nullptr}},
{0x8E, 1, X86InstInfo{"PFPNACC", TYPE_3DNOW_INST, FLAGS_MODRM, 0, nullptr}},
// Inverse 3DNow! These two instructions are Geode product line specific
// No CPUID for these, you're expected to read ID_CONFIG_MSR (1250h) bit 1
{0x86, 1, X86InstInfo{"PFRCPV", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0x87, 1, X86InstInfo{"PFRSQRTV", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0x9A, 1, X86InstInfo{"PFSUB", TYPE_3DNOW_INST, FLAGS_MODRM, 0, nullptr}},
{0x9E, 1, X86InstInfo{"PFADD", TYPE_3DNOW_INST, FLAGS_MODRM, 0, nullptr}},
{0x8A, 1, X86InstInfo{"PFNACC", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0x8E, 1, X86InstInfo{"PFPNACC", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0xAA, 1, X86InstInfo{"PFSUBR", TYPE_3DNOW_INST, FLAGS_MODRM, 0, nullptr}},
{0xAE, 1, X86InstInfo{"PFACC", TYPE_3DNOW_INST, FLAGS_MODRM, 0, nullptr}},
{0x90, 1, X86InstInfo{"PFCMPGE", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0x94, 1, X86InstInfo{"PFMIN", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0x96, 1, X86InstInfo{"PFRCP", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0x97, 1, X86InstInfo{"PFRSQRT", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0xBB, 1, X86InstInfo{"PSWAPD", TYPE_3DNOW_INST, FLAGS_MODRM, 0, nullptr}},
{0xBF, 1, X86InstInfo{"PAVGUSB", TYPE_3DNOW_INST, FLAGS_MODRM, 0, nullptr}},
{0x9A, 1, X86InstInfo{"PFSUB", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0x9E, 1, X86InstInfo{"PFADD", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0x90, 1, X86InstInfo{"PFCMPGE", TYPE_3DNOW_INST, FLAGS_MODRM, 0, nullptr}},
{0x94, 1, X86InstInfo{"PFMIN", TYPE_3DNOW_INST, FLAGS_MODRM, 0, nullptr}},
{0x96, 1, X86InstInfo{"PFRCP", TYPE_3DNOW_INST, FLAGS_MODRM, 0, nullptr}},
{0x97, 1, X86InstInfo{"PFRSQRT", TYPE_3DNOW_INST, FLAGS_MODRM, 0, nullptr}},
{0xA0, 1, X86InstInfo{"PFCMPGT", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0xA4, 1, X86InstInfo{"PFMAX", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0xA6, 1, X86InstInfo{"PFRCPIT1", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0xA7, 1, X86InstInfo{"PFRSQIT1", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0xA0, 1, X86InstInfo{"PFCMPGT", TYPE_3DNOW_INST, FLAGS_MODRM, 0, nullptr}},
{0xA4, 1, X86InstInfo{"PFMAX", TYPE_3DNOW_INST, FLAGS_MODRM, 0, nullptr}},
{0xA6, 1, X86InstInfo{"PFRCPIT1", TYPE_3DNOW_INST, FLAGS_MODRM, 0, nullptr}},
{0xA7, 1, X86InstInfo{"PFRSQIT1", TYPE_3DNOW_INST, FLAGS_MODRM, 0, nullptr}},
{0xAA, 1, X86InstInfo{"PFSUBR", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0xAE, 1, X86InstInfo{"PFACC", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0xB0, 1, X86InstInfo{"PFCMPEQ", TYPE_3DNOW_INST, FLAGS_MODRM, 0, nullptr}},
{0xB4, 1, X86InstInfo{"PFMUL", TYPE_3DNOW_INST, FLAGS_MODRM, 0, nullptr}},
{0xB6, 1, X86InstInfo{"PFRCPIT2", TYPE_3DNOW_INST, FLAGS_MODRM, 0, nullptr}},
{0xB7, 1, X86InstInfo{"PMULHRW", TYPE_3DNOW_INST, FLAGS_MODRM, 0, nullptr}},
{0xB0, 1, X86InstInfo{"PFCMPEQ", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0xB4, 1, X86InstInfo{"PFMUL", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0xB6, 1, X86InstInfo{"PFRCPIT2", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0xB7, 1, X86InstInfo{"PMULHRW", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0xBB, 1, X86InstInfo{"PSWAPD", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0xBF, 1, X86InstInfo{"PAVGUSB", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
};
GenerateTable(&DDDNowOps.at(0), DDDNowOpTable, std::size(DDDNowOpTable));
@@ -14,11 +14,11 @@ namespace FEXCore::X86Tables {
using namespace InstFlags;
void InitializeH0F38Tables() {
#define OPD(prefix, opcode) ((prefix << 8) | opcode)
#define OPD(prefix, opcode) (((prefix) << 8) | opcode)
constexpr uint16_t PF_38_NONE = 0;
constexpr uint16_t PF_38_66 = 1;
constexpr uint16_t PF_38_F2 = 2;
constexpr uint16_t PF_38_F3 = 3;
constexpr uint16_t PF_38_66 = (1U << 0);
constexpr uint16_t PF_38_F2 = (1U << 1);
constexpr uint16_t PF_38_F3 = (1U << 2);
static constexpr U16U8InfoStruct H0F38Table[] = {
{OPD(PF_38_NONE, 0x00), 1, X86InstInfo{"PSHUFB", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
@@ -74,6 +74,7 @@ void InitializeH0F38Tables() {
{OPD(PF_38_66, 0x33), 1, X86InstInfo{"PMOVZXWD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x34), 1, X86InstInfo{"PMOVZXWQ", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x35), 1, X86InstInfo{"PMOVZXDQ", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x37), 1, X86InstInfo{"PCMPGTQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x38), 1, X86InstInfo{"PMINSB", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x39), 1, X86InstInfo{"PMINSD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x3A), 1, X86InstInfo{"PMINUW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
@@ -86,6 +87,14 @@ void InitializeH0F38Tables() {
{OPD(PF_38_66, 0x40), 1, X86InstInfo{"PMULLD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0x41), 1, X86InstInfo{"PHMINPOSUW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_NONE, 0xC8), 1, X86InstInfo{"SHA1NEXTE", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_NONE, 0xC9), 1, X86InstInfo{"SHA1MSG1", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_NONE, 0xCA), 1, X86InstInfo{"SHA1MSG2", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_NONE, 0xCB), 1, X86InstInfo{"SHA256RNDS2", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_NONE, 0xCC), 1, X86InstInfo{"SHA256MSG1", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_NONE, 0xCD), 1, X86InstInfo{"SHA256MSG2", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0xDB), 1, X86InstInfo{"AESIMC", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0xDC), 1, X86InstInfo{"AESENC", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_66, 0xDD), 1, X86InstInfo{"AESENCLAST", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
@@ -98,8 +107,10 @@ void InitializeH0F38Tables() {
{OPD(PF_38_66, 0xF0), 1, X86InstInfo{"MOVBE", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(PF_38_66, 0xF1), 1, X86InstInfo{"MOVBE", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(PF_38_F2, 0xF0), 1, X86InstInfo{"CRC32", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(PF_38_F2, 0xF1), 1, X86InstInfo{"CRC32", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(PF_38_F2, 0xF0), 1, X86InstInfo{"CRC32", TYPE_INST, GenFlagsSizes(SIZE_DEF, SIZE_8BIT) | FLAGS_MODRM, 0, nullptr}},
{OPD(PF_38_F2, 0xF1), 1, X86InstInfo{"CRC32", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{OPD(PF_38_66 | PF_38_F2, 0xF0), 1, X86InstInfo{"CRC32", TYPE_INST, GenFlagsSizes(SIZE_DEF, SIZE_8BIT) | FLAGS_MODRM, 0, nullptr}},
{OPD(PF_38_66 | PF_38_F2, 0xF1), 1, X86InstInfo{"CRC32", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{OPD(PF_38_66, 0xF6), 1, X86InstInfo{"ADCX", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{OPD(PF_38_F3, 0xF6), 1, X86InstInfo{"ADOX", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
@@ -31,7 +31,7 @@ void InitializeH0F3ATables(Context::OperatingMode Mode) {
{OPD(0, PF_3A_66, 0x0E), 1, X86InstInfo{"PBLENDW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x0F), 1, X86InstInfo{"PALIGNR", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x14), 1, X86InstInfo{"PEXTRB", TYPE_INST, GenFlagsSizes(SIZE_8BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_DST_GPR | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x14), 1, X86InstInfo{"PEXTRB", TYPE_INST, GenFlagsSizes(SIZE_32BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_DST_GPR | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x15), 1, X86InstInfo{"PEXTRW", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_DST_GPR | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x16), 1, X86InstInfo{"PEXTRD", TYPE_INST, GenFlagsSizes(SIZE_32BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_DST_GPR | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x17), 1, X86InstInfo{"EXTRACTPS", TYPE_INST, GenFlagsSizes(SIZE_32BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_DST_GPR | FLAGS_XMM_FLAGS, 1, nullptr}},
@@ -49,6 +49,8 @@ void InitializeH0F3ATables(Context::OperatingMode Mode) {
{OPD(0, PF_3A_66, 0x62), 1, X86InstInfo{"PCMPISTRM", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(0, PF_3A_66, 0x63), 1, X86InstInfo{"PCMPISTRI", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_NONE, 0xCC), 1, X86InstInfo{"SHA1RNDS4", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0xDF), 1, X86InstInfo{"AESKEYGENASSIST", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
};
@@ -144,38 +144,38 @@ void InitializeSecondaryGroupTables() {
// AMD documentation is a bit broken for Group 9
// Claims the entire group has n/a applied for the prefix (Implies that the prefix is ignored)
// RDRAND/RDSEED only work with no prefix
// RDRAND/RDSEED only work with no prefix (Other than 66h)
// CMPXCHG8B/16B works with all prefixes
// Tooling fails to decode CMPXCHG with prefix
{OPD(TYPE_GROUP_9, PF_NONE, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_NONE, 1), 1, X86InstInfo{"CMPXCHG16B", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_NONE, 1), 1, X86InstInfo{"CMPXCHG8B/16B", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_NONE, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_NONE, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_NONE, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_NONE, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_NONE, 6), 1, X86InstInfo{"RDRAND", TYPE_UNDEC, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_NONE, 7), 1, X86InstInfo{"RDSEED", TYPE_UNDEC, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_NONE, 6), 1, X86InstInfo{"RDRAND", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_NONE, 7), 1, X86InstInfo{"RDSEED", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_F3, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_F3, 1), 1, X86InstInfo{"CMPXCHG16B", TYPE_INVALID, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_F3, 1), 1, X86InstInfo{"CMPXCHG8B/16B", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_F3, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_F3, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_F3, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_F3, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_F3, 6), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_F3, 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_F3, 7), 1, X86InstInfo{"RDPID", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_66, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_66, 1), 1, X86InstInfo{"CMPXCHG16B", TYPE_INVALID, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_66, 1), 1, X86InstInfo{"CMPXCHG8B/16B", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_66, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_66, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_66, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_66, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_66, 6), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_66, 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_66, 6), 1, X86InstInfo{"RDRAND", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_66, 7), 1, X86InstInfo{"RDSEED", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_F2, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_F2, 1), 1, X86InstInfo{"CMPXCHG16B", TYPE_INVALID, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_F2, 1), 1, X86InstInfo{"CMPXCHG8B/16B", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_F2, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_F2, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_F2, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
@@ -343,8 +343,8 @@ void InitializeSecondaryGroupTables() {
{OPD(TYPE_GROUP_15, PF_F3, 0), 1, X86InstInfo{"RDFSBASE", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_F3, 1), 1, X86InstInfo{"RDGSBASE", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_F3, 2), 1, X86InstInfo{"WRFSBASE", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_REG_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_F3, 3), 1, X86InstInfo{"WRGSBASE", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_REG_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_F3, 2), 1, X86InstInfo{"WRFSBASE", TYPE_INST, GenFlagsDstSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_REG_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_F3, 3), 1, X86InstInfo{"WRGSBASE", TYPE_INST, GenFlagsDstSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_REG_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_F3, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_F3, 5), 1, X86InstInfo{"INCSSPQ", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_15, PF_F3, 6), 1, X86InstInfo{"CLRSSBSY", TYPE_INST, FLAGS_NONE, 0, nullptr}},
@@ -120,9 +120,9 @@ void InitializeSecondaryTables(Context::OperatingMode Mode) {
{0x6F, 1, X86InstInfo{"MOVQ", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0x70, 1, X86InstInfo{"PSHUFW", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 1, nullptr}},
{0x71, 1, X86InstInfo{"", TYPE_GROUP_12, FLAGS_NONE, 0, nullptr}},
{0x72, 1, X86InstInfo{"", TYPE_GROUP_13, FLAGS_NONE, 0, nullptr}},
{0x73, 1, X86InstInfo{"", TYPE_GROUP_14, FLAGS_NONE, 0, nullptr}},
{0x71, 1, X86InstInfo{"", TYPE_GROUP_12, FLAGS_NO_OVERLAY, 0, nullptr}},
{0x72, 1, X86InstInfo{"", TYPE_GROUP_13, FLAGS_NO_OVERLAY, 0, nullptr}},
{0x73, 1, X86InstInfo{"", TYPE_GROUP_14, FLAGS_NO_OVERLAY, 0, nullptr}},
{0x74, 1, X86InstInfo{"PCMPEQB", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0x75, 1, X86InstInfo{"PCMPEQW", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0x76, 1, X86InstInfo{"PCMPEQD", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
@@ -186,8 +186,8 @@ void InitializeSecondaryTables(Context::OperatingMode Mode) {
{0xB6, 1, X86InstInfo{"MOVZX", TYPE_INST, GenFlagsSrcSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_NO_OVERLAY, 0, nullptr}},
{0xB7, 1, X86InstInfo{"MOVZX", TYPE_INST, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_MODRM | FLAGS_NO_OVERLAY, 0, nullptr}},
{0xB8, 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0xB9, 1, X86InstInfo{"", TYPE_GROUP_10, FLAGS_NONE, 0, nullptr}},
{0xBA, 1, X86InstInfo{"", TYPE_GROUP_8, FLAGS_NONE, 0, nullptr}},
{0xB9, 1, X86InstInfo{"", TYPE_GROUP_10, FLAGS_NO_OVERLAY, 0, nullptr}},
{0xBA, 1, X86InstInfo{"", TYPE_GROUP_8, FLAGS_NO_OVERLAY, 0, nullptr}},
{0xBB, 1, X86InstInfo{"BTC", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_NO_OVERLAY, 0, nullptr}},
{0xBC, 1, X86InstInfo{"BSF", TYPE_INST, FLAGS_MODRM | FLAGS_NO_OVERLAY66, 0, nullptr}},
{0xBD, 1, X86InstInfo{"BSR", TYPE_INST, FLAGS_MODRM | FLAGS_NO_OVERLAY66, 0, nullptr}},
@@ -201,7 +201,7 @@ void InitializeSecondaryTables(Context::OperatingMode Mode) {
{0xC4, 1, X86InstInfo{"PINSRW", TYPE_INST, GenFlagsSizes(SIZE_64BIT, SIZE_16BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX | FLAGS_SF_SRC_GPR, 1, nullptr}},
{0xC5, 1, X86InstInfo{"PEXTRW", TYPE_INST, GenFlagsSizes(SIZE_32BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_REG_ONLY | FLAGS_SF_DST_GPR | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 1, nullptr}},
{0xC6, 1, X86InstInfo{"SHUFPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{0xC7, 1, X86InstInfo{"", TYPE_GROUP_9, FLAGS_NONE, 0, nullptr}},
{0xC7, 1, X86InstInfo{"", TYPE_GROUP_9, FLAGS_NO_OVERLAY, 0, nullptr}},
{0xC8, 8, X86InstInfo{"BSWAP", TYPE_INST, FLAGS_SF_REX_IN_BYTE | FLAGS_NO_OVERLAY, 0, nullptr}},
{0xD0, 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
+74 -77
View File
@@ -4,6 +4,8 @@
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/IR/RegisterAllocationData.h>
#include <FEXCore/Utils/Allocator.h>
#include <FEXCore/HLE/SyscallHandler.h>
#include <Interface/Core/LookupCache.h>
#include <cstddef>
#include <cstdint>
@@ -77,7 +79,7 @@ namespace FEXCore::IR {
return true;
}
bool LoadAOTIRCache(AOTCacheType *AOTIRCache, int streamfd) {
static bool LoadAOTIRCache(AOTIRCacheEntry *Entry, int streamfd) {
uint64_t tag;
if (!readAll(streamfd, (char*)&tag, sizeof(tag)) || tag != FEXCore::IR::AOTIR_COOKIE)
@@ -99,6 +101,10 @@ namespace FEXCore::IR {
if (!readAll(streamfd, (char*)&Module[0], Module.size()))
return false;
if (Entry->FileId != Module) {
return false;
}
lseek(streamfd, -sizeof(ModSize) - ModSize - sizeof(IndexSize), SEEK_END);
if (!readAll(streamfd, (char*)&IndexSize, sizeof(IndexSize)))
@@ -119,19 +125,16 @@ namespace FEXCore::IR {
auto Array = (AOTIRInlineIndex *)((char*)FilePtr + IndexOffset);
AOTIRCache->insert({Module, {Array, FilePtr, Size}});
LOGMAN_THROW_A_FMT(Entry->Array == nullptr && Entry->FilePtr == nullptr, "Entry must not be initialized here");
Entry->Array = Array;
Entry->FilePtr = FilePtr;
Entry->Size = Size;
LogMan::Msg::DFmt("AOTIR: Module {} has {} functions", Module, Array->Count);
return true;
}
AOTIRCaptureCache::~AOTIRCaptureCache() {
for (auto &Mod: AOTIRCache) {
FEXCore::Allocator::munmap(Mod.second.mapping, Mod.second.size);
}
}
void AOTIRCaptureCache::FinalizeAOTIRCache() {
AOTIRCaptureCacheWriteoutQueue_Flush();
@@ -230,39 +233,27 @@ namespace FEXCore::IR {
void AOTIRCaptureCache::WriteFilesWithCode(std::function<void(const std::string& fileid, const std::string& filename)> Writer) {
std::shared_lock lk(AOTIRCacheLock);
for( const auto &File: FilesWithCode) {
Writer(File.first, File.second);
for( const auto &Entry: AOTIRCache) {
if (Entry.second.ContainsCode) {
Writer(Entry.second.FileId, Entry.second.Filename);
}
}
}
AOTIRCaptureCache::PreGenerateIRFetchResult AOTIRCaptureCache::PreGenerateIRFetch(uint64_t GuestRIP, FEXCore::IR::IRListView *IRList) {
{
std::shared_lock lk(AOTIRCacheLock);
auto file = AddrToFile.lower_bound(GuestRIP);
if (file != AddrToFile.begin()) {
--file;
if (!file->second.ContainsCode) {
file->second.ContainsCode = true;
FilesWithCode[file->second.fileid] = file->second.filename;
}
}
}
auto AOTIRCacheEntry = CTX->SyscallHandler->LookupAOTIRCacheEntry(GuestRIP);
PreGenerateIRFetchResult Result{};
if (IRList == nullptr && CTX->Config.AOTIRLoad()) {
std::shared_lock lk(AOTIRCacheLock);
auto file = AddrToFile.lower_bound(GuestRIP);
if (file != AddrToFile.begin()) {
--file;
auto Mod = (FEXCore::IR::AOTIRInlineIndex*)file->second.CachedFileEntry;
if (AOTIRCacheEntry.Entry) {
AOTIRCacheEntry.Entry->ContainsCode = true;
if (Mod == nullptr) {
file->second.CachedFileEntry = Mod = AOTIRCache[file->second.fileid].Array;
}
if (IRList == nullptr && CTX->Config.AOTIRLoad()) {
auto Mod = AOTIRCacheEntry.Entry->Array;
if (Mod != nullptr)
{
auto AOTEntry = Mod->Find(GuestRIP - file->second.Start + file->second.Offset);
auto AOTEntry = Mod->Find(GuestRIP - AOTIRCacheEntry.Offset);
if (AOTEntry) {
// verify hash
@@ -299,18 +290,15 @@ namespace FEXCore::IR {
FEXCore::IR::RegisterAllocationData *RAData,
FEXCore::IR::IRListView *IRList,
FEXCore::Core::DebugData *DebugData,
bool GeneratedIR,
bool DecrementRefCount) {
bool GeneratedIR) {
// Both generated ir and LibraryJITName need a named region lookup
if (GeneratedIR || CTX->Config.LibraryJITNaming()) {
std::shared_lock lk(AOTIRCacheLock);
auto file = FindAddrForFile(StartAddr, Length);
auto AOTIRCacheEntry = CTX->SyscallHandler->LookupAOTIRCacheEntry(GuestRIP);
// Only go down this path if we actually found a library region
if (file != AddrToFile.end()) {
if (AOTIRCacheEntry.Entry) {
if (DebugData && CTX->Config.LibraryJITNaming()) {
CTX->Symbols.RegisterNamedRegion(CodePtr, DebugData->HostCodeSize, file->second.filename);
CTX->Symbols.RegisterNamedRegion(CodePtr, DebugData->HostCodeSize, AOTIRCacheEntry.Entry->Filename);
}
// Add to AOT cache if aot generation is enabled
@@ -319,29 +307,31 @@ namespace FEXCore::IR {
auto hash = XXH3_64bits((void*)StartAddr, Length);
auto LocalRIP = GuestRIP - file->second.Start + file->second.Offset;
auto LocalStartAddr = StartAddr - file->second.Start + file->second.Offset;
auto fileid = file->second.fileid;
AOTIRCaptureCacheWriteoutQueue_Append([this, LocalRIP, LocalStartAddr, Length, hash, IRList, RAData, fileid]() {
auto *AotFile = &AOTIRCaptureCacheMap[fileid];
auto LocalRIP = GuestRIP - AOTIRCacheEntry.Offset;
auto LocalStartAddr = StartAddr - AOTIRCacheEntry.Offset;
auto FileId = AOTIRCacheEntry.Entry->FileId;
auto RADataCopy = RAData->CreateCopy();
auto IRListCopy = IRList->CreateCopy();
AOTIRCaptureCacheWriteoutQueue_Append([this, LocalRIP, LocalStartAddr, Length, hash, IRListCopy, RADataCopy, FileId]() {
// It is guaranteed via AOTIRCaptureCacheWriteoutLock and AOTIRCaptureCacheWriteoutFlusing that this will not run concurrently
// Memory coherency is guaranteed via AOTIRCaptureCacheWriteoutLock
auto *AotFile = &AOTIRCaptureCacheMap[FileId];
if (!AotFile->Stream) {
AotFile->Stream = AOTIRWriter(fileid);
AotFile->Stream = AOTIRWriter(FileId);
uint64_t tag = FEXCore::IR::AOTIR_COOKIE;
AotFile->Stream->write((char*)&tag, sizeof(tag));
}
AotFile->AppendAOTIRCaptureCache(LocalRIP, LocalStartAddr, Length, hash, IRList, RAData);
AotFile->AppendAOTIRCaptureCache(LocalRIP, LocalStartAddr, Length, hash, IRListCopy, RADataCopy);
FEXCore::Allocator::free(RADataCopy);
delete IRListCopy;
});
if (CTX->Config.AOTIRGenerate()) {
// cleanup memory and early exit here -- we're not running the application
if (DecrementRefCount) {
--Thread->CompileBlockReentrantRefCount;
}
Thread->CPUBackend->ClearCache();
return true;
}
}
@@ -349,30 +339,26 @@ namespace FEXCore::IR {
// Insert to caches if we generated IR
if (GeneratedIR) {
// Add to thread local ir cache
Core::LocalIREntry Entry = {StartAddr, Length, decltype(Entry.IR)(IRList), decltype(Entry.RAData)(RAData), decltype(Entry.DebugData)(DebugData)};
Thread->LocalIRCache.insert({GuestRIP, std::move(Entry)});
if (Thread->CPUBackend->NeedsRetainedIRCopy()) {
// Add to thread local ir cache
Core::LocalIREntry Entry = {StartAddr, Length, decltype(Entry.IR)(IRList), decltype(Entry.RAData)(RAData), decltype(Entry.DebugData)(DebugData)};
std::lock_guard<std::recursive_mutex> lk(Thread->LookupCache->WriteLock);
Thread->LocalIRCache.insert({GuestRIP, std::move(Entry)});
}
else {
// If the IR doesn't need to be retained then we can just delete it now
delete DebugData;
delete RAData;
delete IRList;
}
}
}
return false;
}
AOTIRCaptureCache::AddrToFileMapType::iterator AOTIRCaptureCache::FindAddrForFile(uint64_t Entry, uint64_t Length) {
// Thread safety here! We are returning an iterator to the map object
// This needs the AOTIRCacheLock locked prior to coming in to the function
auto file = AddrToFile.lower_bound(Entry);
if (file != AddrToFile.begin()) {
--file;
if (file->second.Start <= Entry && (file->second.Start + file->second.Len) >= (Entry + Length)) {
return file;
}
}
return AddrToFile.end();
}
void AOTIRCaptureCache::AddNamedRegion(uintptr_t Base, uintptr_t Size, uintptr_t Offset, const std::string &filename) {
// TODO: Support overlapping maps and region splitting
AOTIRCacheEntry *AOTIRCaptureCache::LoadAOTIRCacheEntry(const std::string &filename) {
auto base_filename = std::filesystem::path(filename).filename().string();
if (!base_filename.empty()) {
@@ -388,21 +374,32 @@ namespace FEXCore::IR {
std::unique_lock lk(AOTIRCacheLock);
AddrToFile.insert({ Base, { Base, Size, Offset, fileid, filename, nullptr, false} });
auto Inserted = AOTIRCache.insert({fileid, AOTIRCacheEntry{0, 0, 0, fileid, filename, false}});
auto Entry = &(Inserted.first->second);
if (CTX->Config.AOTIRLoad && !AOTIRCache.contains(fileid) && AOTIRLoader) {
LOGMAN_THROW_A_FMT(Entry->Array == nullptr, "Duplicate LoadAOTIRCacheEntry");
if (CTX->Config.AOTIRLoad && AOTIRLoader) {
auto streamfd = AOTIRLoader(fileid);
if (streamfd != -1) {
FEXCore::IR::LoadAOTIRCache(&AOTIRCache, streamfd);
FEXCore::IR::LoadAOTIRCache(Entry, streamfd);
close(streamfd);
}
}
return Entry;
}
return nullptr;
}
void AOTIRCaptureCache::RemoveNamedRegion(uintptr_t Base, uintptr_t Size) {
std::unique_lock lk(AOTIRCacheLock);
// TODO: Support partial removing
AddrToFile.erase(Base);
void AOTIRCaptureCache::UnloadAOTIRCacheEntry(AOTIRCacheEntry *Entry) {
LOGMAN_THROW_A_FMT(Entry != nullptr, "Removing not existing entry");
if (Entry->Array) {
FEXCore::Allocator::munmap(Entry->FilePtr, Entry->Size);
Entry->Array = nullptr;
Entry->FilePtr = nullptr;
Entry->Size = 0;
}
}
}
+8 -24
View File
@@ -72,18 +72,19 @@ namespace FEXCore::IR {
struct AOTIRCacheEntry {
AOTIRInlineIndex *Array;
void *mapping;
size_t size;
void *FilePtr;
size_t Size;
std::string FileId;
std::string Filename;
bool ContainsCode;
};
using AOTCacheType = std::unordered_map<std::string, FEXCore::IR::AOTIRCacheEntry>;
bool LoadAOTIRCache(AOTCacheType *AOTIRCache, int streamfd);
class AOTIRCaptureCache final {
public:
AOTIRCaptureCache(FEXCore::Context::Context *ctx) : CTX {ctx} {}
~AOTIRCaptureCache();
void FinalizeAOTIRCache();
void AOTIRCaptureCacheWriteoutQueue_Flush();
@@ -108,11 +109,10 @@ namespace FEXCore::IR {
FEXCore::IR::RegisterAllocationData *RAData,
FEXCore::IR::IRListView *IRList,
FEXCore::Core::DebugData *DebugData,
bool GeneratedIR,
bool DecrementRefCount);
bool GeneratedIR);
void AddNamedRegion(uintptr_t Base, uintptr_t Size, uintptr_t Offset, const std::string &filename);
void RemoveNamedRegion(uintptr_t Base, uintptr_t Size);
AOTIRCacheEntry *LoadAOTIRCacheEntry(const std::string &filename);
void UnloadAOTIRCacheEntry(AOTIRCacheEntry *Entry);
// Callbacks
void SetAOTIRLoader(std::function<int(const std::string&)> CacheReader) {
@@ -136,27 +136,11 @@ namespace FEXCore::IR {
std::queue<std::function<void()>> AOTIRCaptureCacheWriteoutQueue;
std::map<std::string, std::string> FilesWithCode;
struct AddrToFileEntry {
uint64_t Start;
uint64_t Len;
uint64_t Offset;
std::string fileid;
std::string filename;
void *CachedFileEntry;
bool ContainsCode;
};
using AddrToFileMapType = std::map<uint64_t, AddrToFileEntry>;
AddrToFileMapType AddrToFile;
FEXCore::IR::AOTCacheType AOTIRCache;
std::function<int(const std::string&)> AOTIRLoader;
std::function<std::unique_ptr<std::ofstream>(const std::string&)> AOTIRWriter;
std::function<void(const std::string&)> AOTIRRenamer;
std::unordered_map<std::string, FEXCore::IR::AOTIRCaptureCacheEntry> AOTIRCaptureCacheMap;
AddrToFileMapType::iterator FindAddrForFile(uint64_t Entry, uint64_t Length);
};
}
File diff suppressed because it is too large. Load diff
Loaded 100 of 498 files, more files were not shown because too many files have changed in this diff. Show more