Compare commits

..
878 Commits
Author SHA1 Message Date
Ryan Houdek e869aa644a Docs: Update for release FEX-2608 2026-08-04 16:55:43 -07:00
Ryan Houdek 68740b3c65 Merge pull request #5792 from Sonicadvance1/192
CI: Disable ranges-v3 from trying to build native
2026-08-03 15:30:52 -07:00
Ryan Houdek 4caad9bf55 Merge pull request #5800 from Sonicadvance1/195
#5795 but with clang_format
2026-08-03 15:30:14 -07:00
Iaying 7629323548 Fix an XMM register bug in SpillSRA, along with adding a test that reproduces the bug 2026-08-03 15:15:41 -07:00
Ryan Houdek 2fdbff3d1c Merge pull request #5790 from mstorsjo/libc++23
Fix building for Windows with libc++ 23
2026-07-30 14:29:13 -07:00
Ryan Houdek 5c7df98768 CI: Disable ranges-v3 from trying to build native
We don't want this.
2026-07-30 13:59:44 -07:00
Martin Storsjö efe38f2ee2 Add more function stubs for libc++ 23 on Windows
These are needed by libc++ when targeting Windows since
https://github.com/llvm/llvm-project/commit/8a531c3608c722ad529be448d6ecef06ba107228,
which is included in libc++ 23.

In a very brief test, it seems like we don't need to actually
implement them.
2026-07-28 23:37:53 +03:00
Martin Storsjö 08031a2767 Add missing includes
This fixes compilation with libc++ 23, which has removed a number
of unnecessary transitive includes in its headers.

Include <cstdlib> in StringConv.h for std::strtoll and std::strtoull.

Include <cstdlib> for the declarations of malloc/free/realloc/calloc
in Alloc.cpp. (Without this, the functions we define end up with
C++ name mangling.)

Include <stdarg.h> in IO.cpp for va_start/va_end.
2026-07-28 23:17:43 +03:00
Ryan Houdek d295d9f08e Merge pull request #5789 from lioncash/sig
SignalDelegator: Remove unused Required parameter in handler setting
2026-07-27 13:14:14 -07:00
Ryan Houdek 61d033f21e Merge pull request #5788 from lioncash/config
Config: Minor cleanup
2026-07-27 13:09:21 -07:00
Ryan Houdek 27315e33ab Merge pull request #5784 from FrontMage/fix/instruction-fetch-fault-priority
Frontend: Prioritize instruction fetch faults
2026-07-27 12:54:53 -07:00
LC e8b5cd18bd SignalDelegator: Remove unused Required parameter in handler setting
The required flag is set by the subsequent frontend functions that
follow the calls to these functions.
2026-07-27 15:06:24 -04:00
LC 1a77f41846 Config: Add missing override specifier 2026-07-27 14:01:11 -04:00
LC d1947f715c Config: Move strings in constructor where applicable
Same thing, just a little less memory churn.
2026-07-27 14:00:06 -04:00
LC 3f7971b7d2 Config: Mark internally linked where applicable
Makes it obvious these aren't supposed to be exposed.
2026-07-27 13:56:53 -04:00
LC 2cbd5cb3e6 Config: Amend prototypes where applicable
Previously, these didn't match up with the implementation (luckily it's
only used internally at the moment).
2026-07-27 13:54:25 -04:00
FrontMage 151b4d4c2d Frontend: Prioritize instruction fetch faults 2026-07-25 09:01:21 +08:00
Ryan Houdek 7a3fdefafb Merge pull request #5786 from OFFTKP/fist
Extend FIST tests to check for indefinite value
2026-07-24 09:27:22 -07:00
Paris Oplopoios 86c20d0519 Extend FIST tests to check for indefinite value 2026-07-24 17:04:10 +03:00
Ryan Houdek 464ec9d0bc Merge pull request #5783 from FrontMage/fix/inactive-jit-guard-range
FEXCore: Ignore inactive JIT guard ranges
2026-07-23 17:50:51 -07:00
FrontMage 35518a5fa0 FEXCore: Ignore inactive JIT guard ranges 2026-07-24 07:57:48 +08:00
Ryan Houdek d028c7942b Merge pull request #5782 from lioncash/validation
IRValidation: Minor cleanups
2026-07-23 15:44:48 -07:00
LC 585286a617 IRValidation: Remove unused members from BlockInfo
HasExit is assigned to but never used, but we check this condition a
different way right after leaving the main loop anyway.
2026-07-24 16:00:20 -04:00
LC 4904fd43e9 IRValidation: Make BlockInfo private
This isn't used outside the context of the pass.
2026-07-24 16:00:20 -04:00
LC 9e3f287c1f IRValidation: Use C instead of CW
This op isn't mutated anywhere in the pass.
2026-07-24 16:00:20 -04:00
LC f5e4e26e08 IRValidation: Move var closer to usage
Same behavior, just more compact.
2026-07-24 16:00:20 -04:00
LC 6c5a39e164 IRValidation: Turn ORs with true into assignment
These are just unconditional setting to true anyway.
2026-07-24 16:00:17 -04:00
Ryan Houdek 7469fdb0d6 Merge pull request #5781 from FrontMage/fix/multiblock-block-local-errors
FEXCore: Isolate multiblock error state per block
2026-07-23 15:11:13 -07:00
FrontMage fcf9fd77d7 FEXCore: Isolate multiblock error state per block 2026-07-23 16:52:19 +08:00
Ryan Houdek 0589d9b872 Merge pull request #5779 from lioncash/x87
x87StackOptimizationPass: Minor cleanup
2026-07-22 18:58:17 -07:00
LC 6d4c80adff x87StackOptimizationPass: Remove IR member
This is only used in the store helpers, so we can just pass it in
directly
2026-07-23 02:26:46 -04:00
LC 56dd470528 x87StackOptimizationPass: Remove unnecesary return in Run()
It's a void function, so we don't need this at the end
2026-07-23 02:21:26 -04:00
LC 6741f53d87 x87StackOptimizationPass: Mark getValidMask()/getInvalidMask() as const
These don't modify instance state.
2026-07-23 02:18:33 -04:00
LC 19550c5417 x87StackOptimizationPass: Pass by const reference in setTop()
Avoids redundant copies. Just a minor codegen saving.
2026-07-23 02:17:02 -04:00
Ryan Houdek cfa3dfaac7 Merge pull request #5780 from mrpippy/unicode
Windows: Fixes around using Unicode functions
2026-07-22 18:48:52 -07:00
Brendan Shanks 0ed0bc1dd5 CMake: Define UNICODE when building for Windows 2026-07-22 15:17:17 -07:00
Brendan Shanks 074743d6e8 Windows: Explicitly use *A/*W Win32 functions 2026-07-22 15:16:44 -07:00
Brendan Shanks 8b446de059 Windows: Use GetModuleHandleW() to avoid unnecessary string conversions 2026-07-22 15:16:44 -07:00
Ryan Houdek f2e35f336f Merge pull request #5778 from lioncash/buf
SharedCodeBufferManager: Minor header tidying
2026-07-21 09:18:25 -07:00
LC 10a0fe2e71 SharedCodeBufferManager: Make AllocateNew() signature consistent with declaration 2026-07-23 01:42:41 -04:00
LC c9add0d292 SharedCodeBufferManager: Hoist prctl define into util header
Same behavior, but just moves the potential define to be alongside all
of the others in the wrapper header.
2026-07-23 01:40:49 -04:00
LC c239d09ea0 SharedCodeBufferManager: Add missing header
Ensures the page size define is always visible.
2026-07-23 00:16:56 -04:00
LC a652a5811b Merge pull request #5777 from Sonicadvance1/191
FEXCore: Split out CodeBuffer management to its own file
2026-07-21 07:31:13 -04:00
Ryan Houdek d2c92808f5 FEXCore: Split out CodeBuffer management to its own file
NFC

- Renames CodeBufferManager to SharedCodeBufferManager to be more
  explicit about it being shared between threads
- Renames `CodeBuffers` to `SharedCodeBuffers` to make it more explicit
  about sharing these buffers between threads.
- Separates the Manager to its own file so it is distinct from the rest
  of the CPUBackend code

Makes it easier to parse ownership and lifetime semantics of these
buffers.
2026-07-20 18:09:29 -07:00
LC 2464633431 Merge pull request #5776 from Sonicadvance1/190
JIT: Remove JIT detection string
2026-07-20 21:07:23 -04:00
LC 15e76e88b4 Merge pull request #5775 from Sonicadvance1/189
JIT: Rename temporary CPU buffer allocator
2026-07-20 20:50:56 -04:00
Ryan Houdek fe1ac1bc1d JIT: Remove JIT detection string
Now that we have VMA region naming enabled on JIT buffers, this is no
longer used. Confirming a region is a JIT buffer is now just a case of
comparing the name that shows up in `/procfs/maps` rather than dumping
the first bytes of an unknown region.
2026-07-20 17:44:15 -07:00
Ryan Houdek 9edd27b214 JIT: Rename temporary CPU buffer allocator
`TempAllocator` was a bit too opaque as to what the allocator was for,
so I kept needing to lookup its usage every couple of months. Rename it
to `TempCodeBufferAllocator` so I can remember that it is a temporary
allocator for the staging JIT code buffer more easily.

NFC
2026-07-20 17:32:03 -07:00
Ryan Houdek eb7e02ea1d Merge pull request #5772 from lioncash/pass
PassManager: Simplify initialization interface
2026-07-19 16:28:07 -07:00
LC 53befc68c9 PassManager: Ensure GetPass() only queries the underlying pass mappings
Previously this would create an entry in the map if it didn't exist.
2026-07-21 12:10:26 -04:00
LC 19f95d89ec PassManager: Add basic documentation 2026-07-21 12:10:26 -04:00
LC ecb9b3b7f8 PassManager: Constrain GetPass() template to Pass-derived objects
Makes the particular conversion types constrained to catch any trivial
misuses.
2026-07-21 12:10:26 -04:00
LC d619e36523 PassManager: Pass string by const reference where applicable
Gets rid of potential extraneous copies. We also add handling for cases
where two passes with the same name are unintentionally added.
Previously we'd blindly overwrite the mapping.
2026-07-21 12:09:16 -04:00
LC fa80d11960 PassManager: Remove SyscallHandler member
This isn't used anymore, so we can get rid of it to further simplify
initialization.
2026-07-21 11:28:12 -04:00
LC 2893d2b64f PassManager: Simplify pass initialization
We don't conditionally add any passes, so we can simplify the interface
so that we just add all existing passes at once. Makes the core
initialization process a little more straightforward.
2026-07-21 11:28:09 -04:00
Ryan Houdek 99b8df4e6f Merge pull request #5773 from lioncash/fdres
ThreadManager: Fix error return values in FrontendAllocateSlots()
2026-07-19 16:14:50 -07:00
LC 331655182d ThreadManager: Fix error return values in FrontendAllocateSlots()
If ftruncate or the mmap ever fail for whatever reason, then we need to
return the current size, rather than the new size.
2026-07-21 15:12:41 -04:00
Ryan Houdek ec95330dcd Merge pull request #5765 from lioncash/signal
SignalDelegator: Group members together
2026-07-19 16:13:26 -07:00
LC 3959863462 SignalDelegator: Group members together
Hides all public members and situates all of them together to make for
an easier overview.
2026-07-19 21:10:35 -04:00
LC c00b2c9584 Merge pull request #5774 from Sonicadvance1/93
gitlab: Fixes CI
2026-07-19 18:43:37 -04:00
Ryan Houdek 6c354f3987 gitlab: Fixes CI 2026-07-19 12:23:46 -07:00
Ryan Houdek 3bd4d244a4 Merge pull request #5771 from lioncash/fmt
Externals: Update fmt to 12.2.0
2026-07-19 11:50:58 -07:00
LC 9888de25fe Externals: Update fmt to 12.2.0
Keeps fmt up to date.
2026-07-21 09:32:23 -04:00
Ryan Houdek 58c247b30c Merge pull request #5770 from lioncash/cast
CPUBackend: Remove unnecessary reinterpret_casts
2026-07-19 00:57:29 -07:00
LC 22bd10f3b1 CPUBackend: Remove unnecessary reinterpret_casts
This both take a void*, so the casting is unnecessary to begin with,
since this would occur anyway without it. We can also avoid a
duplication to reduce line noise.
2026-07-21 08:22:17 -04:00
Ryan Houdek f374b4775a Merge pull request #5769 from lioncash/bound
Core: Remove unnecessary bounds check in GenerateIR()
2026-07-18 21:35:34 -07:00
LC ff213bbc5e Core: Move vars closer to usage scope in GenerateIR()
Makes it so their purpose is more easily seen
2026-07-21 04:44:47 -04:00
LC 56a4ca6e6a Core: Remove unnecessary bounds check in GenerateIR()
We already check the bounds in the loop prior to calling at().
2026-07-21 04:40:22 -04:00
Ryan Houdek d6b38b6b1c Merge pull request #5768 from lioncash/stream
IRDumper: stringstream -> ostringstream
2026-07-18 21:33:42 -07:00
LC 04d06d386f IRDumper: stringstream -> ostringstream
These are purely output operations, so we don't need to use the more
heavyweight class.
2026-07-21 04:29:03 -04:00
LC aa26a780ed Merge pull request #5767 from Sonicadvance1/188
AVX128: Optimize 256-bit vmovmaskpd as well
2026-07-17 16:56:23 -04:00
Ryan Houdek f73b93dbc2 InstcountCI: Update 2026-07-17 13:17:39 -07:00
Ryan Houdek c4a5ac892f AVX128: Optimize 256-bit vmovmaskpd as well
Similar to #5757, but once the elements have been zipped together, we
can treat it identically to the 128-bit 32-bit element path.

Closes #3782
2026-07-17 13:15:47 -07:00
Ryan Houdek f129ca0c61 unittests/vmovmskpd: Extend test to have different lower and upper results between 128-bit lanes. 2026-07-17 13:12:27 -07:00
Ryan Houdek 941f0fbf8d Merge pull request #5764 from lioncash/alloc
LinuxAllocator: Reduce MemAllocator32Bit size by 16 bytes
2026-07-17 08:00:25 -07:00
LC 6e37dca566 LinuxAllocator: Reduce MemAllocator32Bit size by 16 bytes
These constants don't need to be member vars.
2026-07-19 18:37:38 -04:00
Ryan Houdek 1cffa009fe Merge pull request #5763 from lioncash/const
IREmitter: Mark some helpers as const
2026-07-17 07:59:47 -07:00
Tony Wasserka 83f4c9d101 Merge pull request #5760 from neobrain/fix_emitter_constants
Arm64Emitter: Fix incorrect condition for constant NOP padding
2026-07-17 13:05:50 +02:00
Tony Wasserka c0c95da796 Arm64Emitter: Fix incorrect condition for constant NOP padding
This needs to be enabled when *generating* caches, not at runtime when we're
loading them (unless we're compiling for validation).

Previous code would incorrectly disable NOP padding in FEXOfflineCompiler and
instead enable it at runtime when it wasn't needed.
2026-07-17 12:47:51 +02:00
LC 8b612a87c6 IREmitter: Mark some helpers as const
These don't modify internal state.
2026-07-17 03:20:36 -04:00
Ryan Houdek 2ac95f446b Merge pull request #5762 from lioncash/core
FEXCore: Resolve missing prototype warnings
2026-07-17 00:04:47 -07:00
LC c5eddd922d FEXCore: Resolve missing prototype warnings
Makes sure we mark everything internally linked as necessary, or make
declarations visible to their implementation.
2026-07-17 02:48:45 -04:00
Ryan Houdek b58be2c073 Merge pull request #5761 from lioncash/sys
LinuxEmulation: Resolve missing prototype warnings
2026-07-16 22:00:10 -07:00
LC a3d0b67777 LinuxEmulation: Resolve missing prototype warnings
Ensures all functions are marked whether they're intended to be
internally linked or not.
2026-07-17 00:12:28 -04:00
Ryan Houdek 0467d523c0 Merge pull request #5759 from neobrain/fix_codebuffer_max_size
CodeCache: Use maximal code buffer size when generating code caches, too
2026-07-16 13:53:42 -07:00
Ryan Houdek b0af054a95 Merge pull request #5757 from MoonFlowww/avx128-vmovmsk-256
AVX_128: Optimize VMOVMSKPS from 11 to 7 instructions
2026-07-16 13:53:00 -07:00
Ryan Houdek 6846f10510 Merge pull request #5758 from lioncash/x87
x87StackOptimizationPass: Make use of std::array for FixedSizeStack
2026-07-16 12:39:08 -07:00
Ryan Houdek e24f232f52 Merge pull request #5756 from lioncash/fill
Arm64Emitter: Pull FillSpecialRegs bools into a struct
2026-07-16 12:36:31 -07:00
Ryan Houdek a1cc5d034e Merge pull request #5755 from lioncash/host
HostRunner: Tidy up interface
2026-07-16 12:35:18 -07:00
Ryan Houdek b478f54aea Merge pull request #5754 from lioncash/vdso
VDSO_Emulation: Mark relevant members as internally linked
2026-07-16 12:34:27 -07:00
Ryan Houdek 7bb380a086 Merge pull request #5753 from lioncash/pipe
FEXServer: Fix some missing declaration warnings
2026-07-16 12:33:54 -07:00
Tony Wasserka 228c351396 CodeCache: Use maximal code buffer size when generating code caches, too
This is less likely to happen, but will still be required for very large libraries.
2026-07-16 16:37:18 +02:00
LC 4bb675a530 x87StackOptimizationPass: Reduce noise in slow push/pop paths
Deduplicates the repeated rotate behavior.
2026-07-16 10:31:40 -04:00
LC 8d62773570 x87StackOptimizationPass: Make helpers internally linked
Makes it obvious they're only used in this TU and allows the compiler to
warn if they ever become unused.
2026-07-16 09:56:14 -04:00
LC 9e26c57643 x87StackOptimizationPass: Fix isValid()
Previously this wouldn't have worked, since .first isn't a valid member.
The only reason it wasn't caught is because the function is never
instantiated.
2026-07-16 09:56:14 -04:00
LC 35bc502062 x87StackOptimizationPass: Make use of std::array for FixedSizeStack
Reduces the overall generated code for state management.

Drops the overall text size from 11447294 to 11441918
2026-07-16 09:56:05 -04:00
moonfloww 7189e1e280 InstcountCI: Update 2026-07-16 14:44:04 +02:00
LC 4b4aa1cdbe Arm64Emitter: Pull FillSpecialRegs bools into a struct
Makes this easily expandable over time without modifying the prototype,
and lets us be a little more informative at call sites.
2026-07-16 08:07:57 -04:00
moonfloww 3680282b30 new vmovmsk from 11 to 7 ins. 2026-07-16 13:57:12 +02:00
LC a4d5fbae63 HostRunner: Tidy up interface
We've accumulated a bunch of forward declarations that are no longer
necessary. We also don't need to pass the signal delegator as a
reference, since we're not modifying the pointer itself, it's just
passed in to register a signal handler.
2026-07-16 07:39:03 -04:00
LC 2943b87c83 VDSO_Emulation: Mark relevant members as internally linked
Silences missing prototype warnings and makes it obvious they're only
used within the translation unit.
2026-07-16 07:28:50 -04:00
LC 9575506391 FEXServer: Fix some missing declaration warnings
Makes sure prototypes are visible to their implementation. Also marks
functions internally linked where applicable.

Also makes it a little more visibly obvious which bits are exposed for
use elsewhere.
2026-07-16 07:10:32 -04:00
Ryan Houdek 31c2449d6e Merge pull request #5752 from lioncash/config
FEXGetConfig: Add convenience option for dumping system/tso info
2026-07-15 09:48:49 -07:00
LC 70eafc4f6f FEXGetConfig: Mark helpers as static where applicable
Makes them internally linked, and also lets them be caught by the
compiler when they're unused.
2026-07-15 12:22:12 -04:00
LC 38cbd2aeb2 FEXGetConfig: Add convenience option for dumping system/tso info
Just lets you get a broad overview all at once instead of needing to
type out every long command.

Now it's easier to be lazy and just pass "-e", or "--all-emu-info".
2026-07-15 12:22:10 -04:00
LC ba762a9326 Merge pull request #5747 from Sonicadvance1/187
HostFeatures: Pull MMFR3 identification register
2026-07-15 12:20:52 -04:00
Ryan Houdek 921ce59054 Merge pull request #5750 from lioncash/pred
VectorOps: Make use of unpredicated shifts
2026-07-15 09:15:04 -07:00
Ryan Houdek a7627ba39a Merge pull request #5751 from lioncash/sq
VectorOps: Add trivial case handling in VSQXTN2
2026-07-15 09:06:59 -07:00
Ryan Houdek 34ef28dea5 Merge pull request #5749 from lioncash/calc
RedundantFlagCalculationElimination: Minor tidying
2026-07-15 09:05:45 -07:00
Ryan Houdek 0f2463dda2 Merge pull request #5748 from lioncash/invariant
RegisterAllocationPass: Ensure pair reg invariant
2026-07-15 09:04:38 -07:00
Ryan Houdek 5ec852637c Merge pull request #5745 from lioncash/addv
VectorOps: Simplify 256-bit VAddV
2026-07-15 09:03:57 -07:00
LC ba8b0afe7a VectorOps: Use unpredicated shifts where applicable for 256-bit scalar shifts
Lets us trim some output
2026-07-15 09:10:57 -04:00
LC 98d45a6a9b VectorOps: Make use of unpredicated immediate shifts
Same behavior, just without introducing a predicate register dependency.
2026-07-15 08:00:58 -04:00
LC 6bc808cb06 VectorOps: Add trivial case handling in VSQXTN2
Lets us generate much more optimal code in the event the destination and
lower source are the same.
2026-07-15 07:49:50 -04:00
LC 1325fef703 RFCE: Prefer accessing ops with C instead of CW
CW is only intended when the op needs to be writable, but most of these
are only reading data.
2026-07-15 07:06:23 -04:00
LC d2d0ef1803 RFCE: Remove unnecessary std::invoke()
We can just call this normally (and also make the constituent helper
function internally linked).
2026-07-15 07:03:09 -04:00
LC d6fb60d512 RegisterAllocationPass: Ensure pair reg invariant
Allows us to actually catch if this requirement ever gets broken in
the future.
2026-07-15 06:46:24 -04:00
LC ef35474f88 VectorOps: Simplify 256-bit VAddV
Didn't read the manual close enough on the first read award.
2026-07-15 04:43:47 -04:00
Ryan Houdek 372891361c HostFeatures: Pull MMFR3 identification register
This has the S1POE flag that we will want to use in the future.
2026-07-14 20:09:31 -07:00
Ryan Houdek 50c75d1f43 Move Linux version calculation to common code 2026-07-14 20:06:37 -07:00
Ryan Houdek 30f2a7b23b Merge pull request #5746 from lioncash/shadow
x87StackOptimizationPass: Remove shadowing variable in PUSHSTACK case
2026-07-14 13:32:13 -07:00
LC f2212a497b x87StackOptimizationPass: Remove shadowing variable in PUSHSTACK case
No behavioral change, since the one in the outer scope does the same thing.
2026-07-14 08:01:30 -04:00
Ryan Houdek 12e8cf008a Merge pull request #5744 from lioncash/telem
AtomicOps: Avoid constrained unpredictable case in TelemetrySetValue()
2026-07-13 16:37:51 -07:00
LC 9e8e87bbb2 AtomicOps: Avoid constrained unpredictable case in TelemetrySetValue()
STLXR cannot use the same register as both the status register and the
value register, otherwise it's architecturally unpredictable
behavior.

Only applies to hardware without FEAT_LSE, so this only meaningfully
affects hardware using the v8.0 spec, since FEAT_LSE becomes mandatory
in v8.1 and newer.
2026-07-13 19:08:17 -04:00
Ryan Houdek 76c4ebb36f Merge pull request #5743 from lioncash/str
StringUtils: Handle strings entirely composed of whitespace in trims
2026-07-13 15:40:50 -07:00
Ryan Houdek 9ff322eeed Merge pull request #5742 from lioncash/sema
x32/Semaphore: Fix storing of message type in msgrcv
2026-07-13 15:40:07 -07:00
LC 4afa49824e StringUtils: Handle strings entirely composed of whitespace in trims
Previously this wouldn't handle fully whitespaced strings.
2026-07-13 17:37:09 -04:00
LC 23402bf31b x32/Semaphore: Fix storing of message type in msgrcv
This was previously storing into the local compat handler, not the
actual managed message.
2026-07-13 17:30:27 -04:00
LC 50be718b72 x32/Semaphore: Mark _ipc as static
This isn't used outside of the translation unit.
2026-07-13 17:30:24 -04:00
Ryan Houdek 903e7db427 Merge pull request #5741 from lioncash/file 2026-07-13 14:13:13 -07:00
LC 5e5e9e0803 Utils/File: Fix handle releasing
ShouldClose was never being set in the event we opened a regular file.
The only time it was set (to false) is when it's used to encapsulate
stderr and stdout.

So anything opened by a File instance was essentially held open.
2026-07-13 16:01:28 -04:00
Ryan Houdek ce27754b9d Merge pull request #5740 from lioncash/ra
RegisterAllocationPass: Function cleanup
2026-07-13 12:42:45 -07:00
LC 4254c0f5a9 RegisterAllocationPass: Function cleanup
Marks a few functions const or static to clarify usage a little more.
2026-07-13 15:26:51 -04:00
Ryan Houdek 7efc3ecaba Merge pull request #5739 from lioncash/zero
Vector: Indicate 128-bit zero vector in DefaultX87State()
2026-07-13 11:59:22 -07:00
LC 2934b01d58 Vector: Indicate 128-bit zero vector in DefaultX87State()
Same functional behavior, just makes it visually match the store size
below. Technically also avoids delegating off to the 64-bit element
path if a 128-bit constant zero is already loaded.
2026-07-13 14:07:12 -04:00
LC 192e363701 Merge pull request #5738 from Sonicadvance1/186
64BitAllocator: Removes unused additional size argument
2026-07-13 13:44:11 -04:00
Ryan Houdek 1287365616 64BitAllocator: Removes unused additional size argument
This used to be used for the intrusively allocated `LiveVMARegion` but
that is all handled internally to the object now, making this
unnecessary. It was always receiving zero and doing nothing so just
remove it.
2026-07-13 10:18:46 -07:00
Ryan Houdek 4fa539fbb2 Merge pull request #5733 from lioncash/alloc
64BitAllocator: Avoid madvising more than necessary in InitializeVMARegionsUsed()
2026-07-13 10:16:57 -07:00
Ryan Houdek 28cdae4687 Merge pull request #5737 from lioncash/bsl
VectorOps: Simplify SVE 256-bit VOrn with BSL2N
2026-07-13 10:05:08 -07:00
LC 842e22915c VectorOps: Simplify SVE 256-bit VOrn with BSL2N
Lets us shave off an instruction and also avoid using a temporary
register in some cases. We can also tweak our worst case that requires a
predicate to eliminate the temporary as well.

We can also expand our cmpps cases, so that we can reflect the
BSL2N usages in instcountci.
2026-07-13 12:27:54 -04:00
Ryan Houdek bd150233ce Merge pull request #5736 from lioncash/ushrni
VectorOps: Make SVE shift==0 case symmetric with ASIMD
2026-07-13 07:56:25 -07:00
LC 24720b67da VectorOps: Make SVE shift==0 case symmetric with ASIMD
Ensures that we have consistent behavior.
2026-07-13 10:26:42 -04:00
Ryan Houdek 356d461123 Merge pull request #5734 from lioncash/bytes
Common/BitSet: Amend byte size retrieval
2026-07-13 07:07:00 -07:00
Ryan Houdek 3cbcc7b9f8 Merge pull request #5735 from lioncash/ir
IR: Enclose straggler Desc comments in brackets
2026-07-13 07:06:06 -07:00
LC 329f12a888 json_ir_generator: Join successive write calls together for allocator helpers
We can just write these out as cohesive units. Also makes adding to them
less annoying.
2026-07-13 09:22:55 -04:00
LC c0b2eec5de IR: Enclose straggler Desc comments in brackets
Ensures the comments get rendered properly in output. We can also
make sure that the IR generation script catches this in the future.
2026-07-13 08:59:45 -04:00
LC 9b8ae25491 Common/BitSet: Amend byte size retrieval
This needs to divide by 8 to get a proper byte size for all type sizes.
The only usage of this is currently a uint64_t, so it worked by
coincidence, since sizeof(uint64_t) == 8.
2026-07-13 08:17:27 -04:00
LC c0ee865e3e 64BitAllocator: Avoid madvising more than necessary in InitializeVMARegionsUsed
Because our bitset type is uint64_t, then that means Memory + ManagedSize
is more like: Memory + (ManagedSize * 8), which is way larger of a base
than we need.
2026-07-12 20:38:19 -04:00
Ryan Houdek 46ec2797ff Merge pull request #5732 from lioncash/vec
Crypto: Clarify zero vector size in SHA1RNDS4Op()
2026-07-12 16:14:26 -07:00
LC fb2cdc8541 Crypto: Clarify zero vector size in SHA1RNDS4Op()
This ends up zeroing out the whole 128-bit vector.
2026-07-12 18:49:27 -04:00
Ryan Houdek 850ef70496 Merge pull request #5731 from lioncash/xar
Crypto: Make use of XAR in SHA1NEXTE when available
2026-07-12 14:26:14 -07:00
LC 9d3c388664 Crypto: Make use of XAR in SHA1NEXTE when available
Lets us shave an instruction off on hardware that supports XAR.

Closes #5730
2026-07-12 16:02:12 -04:00
Ryan Houdek f2b679f602 Merge pull request #5728 from lioncash/halves
x32/FD: Combine offset halves directly
2026-07-12 11:15:10 -07:00
Ryan Houdek b9d97dffe7 Merge pull request #5727 from lioncash/vmsplice
x32/FD: Make use of SanitizeIOCount for vector construction in vmsplice
2026-07-12 11:14:31 -07:00
Ryan Houdek 7ff0466c2c Merge pull request #5726 from lioncash/file
Utils/File: Handle dual read/write case
2026-07-12 11:09:00 -07:00
Ryan Houdek a135325185 Merge pull request #5725 from lioncash/dead
Signals: Preprocessor disable intentional dead code
2026-07-12 11:06:30 -07:00
LC a6c8f0d300 x32/FD: Combine offset halves directly
Shortens these up a little.
2026-07-12 13:31:48 -04:00
LC 046750354e x32/FD: Make use of SanitizeIOCount for vector construction in vmsplice
Makes this consistent with the other fd syscalls that make temporary
buffers.
2026-07-12 13:01:58 -04:00
LC e23d703873 Utils/File: Handle dual read/write case
According to POSIX open docs, this is a completely separate flag that
isn't a combination of O_RDONLY and O_WRONLY, so we need to handle this
separately.

Makes the codepath behaviorally symmetric with the Windows one.
2026-07-12 12:41:02 -04:00
LC ca2d2520d2 Signals: Preprocessor disable intentional dead code in userfaultfd
Noticed this when going through the syscalls. Avoids potential warnings.
2026-07-12 12:16:07 -04:00
Ryan Houdek bc16f902d1 Merge pull request #5724 from lioncash/file
WinAPI/IO: Fix handling of end of file offset in SetFilePointerEx
2026-07-11 22:44:55 -07:00
LC f5ae888597 WinAPI/IO: Fix handling of end of file offset in SetFilePointerEx
This just means the end of the file is being used as the base offset.

Also note that according to the documentation for SetFilePositionEx,
that setting the position beyond the current file size is not considered
an error as far as the API is concerned.
2026-07-12 01:26:23 -04:00
Ryan Houdek 376e3af058 Merge pull request #5723 from lioncash/gdb 2026-07-11 21:54:57 -07:00
Ryan Houdek af63c0a9e1 Merge pull request #5722 from lioncash/container 2026-07-11 21:54:20 -07:00
LC 7d149ebec4 GdbServer: Add missing log format argument 2026-07-12 00:28:29 -04:00
LC 37c809471b ElfContainer: Amend entry iteration in GetDynamicLibs()
These were using i in the termination condition, which is for section
headers, not entries.
2026-07-12 00:19:28 -04:00
Ryan Houdek 12baceb859 Merge pull request #5721 from lioncash/win
AllocatorHooks: Amend VirtualProtect for Windows
2026-07-11 20:49:16 -07:00
Ryan Houdek 3e6b3c8c40 Merge pull request #5720 from lioncash/hdr
64BitAllocator: Remove duplicate headers
2026-07-11 20:33:09 -07:00
LC 25f2021711 AllocatorHooks: Amend VirtualProtect for Windows
VirtualProtect returns non-zero on success, also the old protection flag
parameter isn't allowed to be null.
2026-07-11 23:24:30 -04:00
LC 8b30f7dbb5 64BitAllocator: Remove duplicate headers
These are already included.
2026-07-11 23:08:29 -04:00
Ryan Houdek 76b35dfeb0 Merge pull request #5719 from lioncash/small
64BitAllocator: Avoid overwriting Region[0] in Create64BitAllocatorWithRegions
2026-07-11 19:30:08 -07:00
Ryan Houdek 4518d5831b Merge pull request #5718 from lioncash/absolute
Filesystem: Fix Absolute() on Windows
2026-07-11 17:11:18 -07:00
Ryan Houdek 066250851f Merge pull request #5717 from lioncash/reg
RegisterAllocationPass: Amend type cast in DecodeSRANode()
2026-07-11 17:10:27 -07:00
Ryan Houdek 341a5196ba Merge pull request #5701 from lioncash/thread
Thread: Amend new thread handling in HandleNewClone()
2026-07-11 17:09:59 -07:00
LC c878a89e95 Filesystem: Fix Absolute() on Windows
sizeof(*Fill) will only ever be 1, so we wouldn't actually copy much of
anything.
2026-07-11 17:25:26 -04:00
LC c72d59a3f1 64BitAllocator: Avoid overwriting Region[0] in Create64BitAllocatorWithRegions
Since this was a reference, this would end up overwriting Region[0] with
whatever the smallest region was instead of just being a running pointer
to what happened to be the current smallest region.

We can switch over to a pointer to avoid obliterating the first memory
region.
2026-07-11 16:52:01 -04:00
LC 995e2657bb RegisterAllocationPass: Amend type cast in DecodeSRANode()
Uses the proper type for StoreRegister. Same behavior though, due to
layout.
2026-07-11 16:39:35 -04:00
Ryan Houdek fe6d6397d6 Merge pull request #5716 from lioncash/vec 2026-07-11 13:29:47 -07:00
LC c7f52bcee4 Vector: Centralize masking in InsertScalarFCMPOp
Ensures that even if someone threw bogus constants in the upper bits of
the immediate, that the special-cased comparison types would still be
handled properly.

We can move the masking in the AVX variant too, just to be consistent.
2026-07-11 16:07:19 -04:00
Ryan Houdek fb4d8a6d14 Merge pull request #5715 from lioncash/singlestep
Dispatcher: Avoid loading unnecessary reg in vixl single step
2026-07-11 12:25:13 -07:00
LC 929f9a7ad2 Dispatcher: Avoid loading unnecessary reg in vixl single step
CompileSingleStep only takes one uint64_t, not two.
2026-07-11 15:03:29 -04:00
Ryan Houdek 48ce5bf6e6 Merge pull request #5714 from lioncash/readahead
x32/FD: Fix readahead upper offset type
2026-07-11 11:20:49 -07:00
Ryan Houdek 617a518714 Merge pull request #5713 from lioncash/select
x32/FD: Correct total word calculation in select() variants
2026-07-11 11:19:29 -07:00
Ryan Houdek 480f45f2f3 Merge pull request #5712 from lioncash/bpf
BPFEmitter: Amend instruction class checking in HandleStore()
2026-07-11 11:08:58 -07:00
Ryan Houdek 0e8b01c1b5 Merge pull request #5711 from lioncash/fault
x64/Thread: Amend faulting copy handling related to LDTs
2026-07-11 11:04:47 -07:00
Ryan Houdek b750d6772f Merge pull request #5710 from lioncash/close
Common/Async: Handle fd closing a little better
2026-07-11 11:03:10 -07:00
Ryan Houdek 7ce172a497 Merge pull request #5709 from lioncash/size
FlexBitSet: Simplify MemClear/MemSet
2026-07-11 10:57:58 -07:00
Ryan Houdek 821dfe5b98 Merge pull request #5708 from lioncash/mrs
MiscOps: Fix round mode clearing for RP/RM modes in PushRoundingMode
2026-07-11 10:56:27 -07:00
Ryan Houdek 6c57b2f8f9 Merge pull request #5707 from lioncash/offset
MemoryOps: Avoid double application of base offset in {Load,Store}ContextIndexed case
2026-07-11 10:49:27 -07:00
Ryan Houdek 6e0c9d159d Merge pull request #5706 from lioncash/odd
SignalDelegator: Remove odd double negation in GuestSigProcMask
2026-07-11 10:48:42 -07:00
LC a90a7dbec7 x32/FD: Fix readahead upper offset type
This should be a uint32_t
2026-07-11 13:33:11 -04:00
LC 2250bd58a2 x32/FD: Deduplicate guest and host fd set management
Same behavior, but less copy pastey
2026-07-11 13:23:19 -04:00
LC 126be45ed1 x32/FD: Correct total word calculation in select() variants
Previously this would result in a larger amount of words specified than
necessary.

e.g. Given nfds = 1:

With AlignUp(1, 32) / 4, we'd end up with supposedly eight words, when it
should only be one word.

On the other extreme, given a full fd set of 1024 fds, then we'd end up
with 256 words, when it should only be 32.
2026-07-11 12:53:08 -04:00
LC e021a55abd BPFEmitter: Amend instruction class checking in HandleStore()
This was previously checking for a load class, which would result in ST
clobbering the index register.
2026-07-11 12:26:11 -04:00
LC 2ad6254894 x64/Thread: Amend faulting copy handling related to LDTs
CopyToUser doesn't return the number of bytes copied, but rather returns
0 to indicate success, otherwise a fault has occurred (and the SIGSEGV
handler has set X0 to EFAULT)
2026-07-11 11:57:38 -04:00
LC 5bd97c2f63 Common/Async: Handle fd closing a little better
We should be checking against -1, rather than just anything non-zero.
2026-07-11 10:49:08 -04:00
LC 570d1f2271 FlexBitSet: Simplify MemClear/MemSet
We can just make use of the helpers already in the interface.
2026-07-11 10:33:47 -04:00
LC 613e9ef701 MiscOps: Fix round mode clearing for RP/RM modes in PushRoundingMode
Previously this had the potential to not clear rounding bits properly
depending on incoming FPCR state.
2026-07-11 10:24:07 -04:00
LC 5080c6ffc5 MemoryOps: Avoid double application of base offset in {Load,Store}ContextIndexed unaligned case 2026-07-11 09:56:56 -04:00
LC f16bc12b7d SignalDelegator: Remove odd double negation in GuestSigProcMask
We can just reduce it to normal null comparisons.
2026-07-11 05:42:48 -04:00
Ryan Houdek d2f096187c Merge pull request #5705 from lioncash/thread3 2026-07-10 22:34:04 -07:00
Ryan Houdek bd1e61befd Merge pull request #5704 from lioncash/host 2026-07-10 22:33:30 -07:00
Ryan Houdek b97165bec9 Merge pull request #5703 from lioncash/size 2026-07-10 22:32:45 -07:00
LC 8b0f07b2dc Syscalls: Remove unnecessary usages of namespace FEXCore::IR
These aren't necessary.
2026-07-10 21:02:12 -04:00
LC c0ba45f6de HostFeatures: Shrink feature setting in FillFeatureFlags
Allows us to unify most of the flag setting, so the flag name only needs
to be stated once, reducing likelihood of typos.
2026-07-10 20:34:46 -04:00
LC 017c898ed0 x64/Signals: Amend set size in rt_sigtimedwait
We should be checking the size passed in, not the sizeof of it.
2026-07-10 19:40:05 -04:00
Ryan Houdek b193a0c9fd Merge pull request #5702 from lioncash/sbss
HostFeatures: Amend SSBS2 signifying
2026-07-10 16:38:07 -07:00
LC db0c7b0562 HostFeatures: Amend SSBS2 signifying 2026-07-10 19:13:29 -04:00
LC 85b8e91b57 Thread: Amend new thread handling in HandleNewClone()
Ensures that newly cloned threads get tracked properly.
2026-07-10 18:35:58 -04:00
Ryan Houdek 87301ca154 Merge pull request #5700 from lioncash/bitset
Common/Bitset: Minor API changes
2026-07-10 10:35:45 -07:00
Ryan Houdek 650f5b784d Merge pull request #5699 from lioncash/spillops
Arm64Emitter: Wire up conditional FPR spilling in SpillForPreserveAllABICall
2026-07-10 10:35:19 -07:00
Ryan Houdek c8dd9eefa6 Merge pull request #5698 from lioncash/dead
ConversionOps: Remove redundant code in Vector_FToS
2026-07-10 10:35:06 -07:00
Ryan Houdek 0bdf15977f Merge pull request #5697 from lioncash/thread
x32/Thread: Move writability check around in waitpid
2026-07-10 10:34:56 -07:00
Ryan Houdek 8a844f9cca Merge pull request #5695 from lioncash/socket
Socket: Pass size by reference in getsockopt
2026-07-10 10:34:33 -07:00
Ryan Houdek 8fbde84380 Merge pull request #5696 from lioncash/time
x32/Time: Correct sizeof expression in utimensat
2026-07-10 10:31:47 -07:00
Ryan Houdek 00826d3327 Merge pull request #5694 from lioncash/rlimit
x32/Info: Only modify output in getrlimit/ugetrlimit if successful
2026-07-10 10:27:51 -07:00
Ryan Houdek d5572e322f Merge pull request #5693 from lioncash/ir
IR: Use begin block type in operator--
2026-07-10 10:27:13 -07:00
Ryan Houdek 37814111de Merge pull request #5692 from lioncash/unary
Vector: Use unary handler for scalar unary insertions
2026-07-10 10:26:44 -07:00
Ryan Houdek 1417888a89 Merge pull request #5691 from lioncash/fpr
Vector: LoadSourceGPR -> LoadSourceFPR for MASKMOVOp
2026-07-10 10:25:26 -07:00
Ryan Houdek 5d846c3ca9 Merge pull request #5690 from lioncash/cache
MemoryOps: Make use of current working reg for cache operations
2026-07-10 10:23:55 -07:00
Ryan Houdek 5dd0477440 Merge pull request #5689 from lioncash/msg
x32/Msg: Fix result comparison in mq_getsetattr
2026-07-10 10:15:15 -07:00
Ryan Houdek c6f823f855 Merge pull request #5688 from lioncash/pidfd
Thread: Amend pidfd_open check
2026-07-10 10:14:49 -07:00
Ryan Houdek 2af7c24e79 Merge pull request #5687 from lioncash/limit
x32/Thread: Fix off-by-one in get_thread_area
2026-07-10 10:14:29 -07:00
Ryan Houdek 94ccfefa84 Merge pull request #5686 from lioncash/fd
x32/FD: Ensure sendfile updates offset if set
2026-07-10 10:14:18 -07:00
LC 46f3bec37e Common/BitSet: Ensure internal pointer is always initialized
Provides deterministic state.
2026-07-10 10:31:45 -04:00
LC f5d2e0db29 Common/BitSet: Mark getters as const
These don't modify internal state.
2026-07-10 10:31:05 -04:00
LC 88ee56f471 Common/BitSet: Amend Clear() behavior
Ensures the bits are actually being unset.
2026-07-10 10:26:07 -04:00
LC 8436154276 Arm64Emitter: Wire up conditional FPR spilling in SpillForPreserveAllABICall
Technically, this parameter wasn't being used at all. It was wired up
for filling, but not spilling.
2026-07-10 10:10:25 -04:00
LC 150b25f29e ConversionOps: Remove redundant code in Vector_FToS
These are already defined in an outer scope.
2026-07-10 09:16:35 -04:00
LC f4c50105ba x32/Thread: Move writability check around in waitpid
Same behavior, but catches the write before it actually occurs.
2026-07-10 08:08:45 -04:00
LC 3395bedc2b x32/Time: Correct sizeof expression in utimensat
Ensures we check the proper type.
2026-07-10 08:04:56 -04:00
LC 7a14210e2b Socket: Pass size by reference in getsockopt
Previously this was passing by value.
2026-07-10 08:01:29 -04:00
LC d206b67ca8 x32/Info: Only modify output in getrlimit/ugetrlimit if successful
Avoids trampling over input data.
2026-07-10 07:49:44 -04:00
LC 0bfef9008b IR: Use begin block type in operator--
Same behavior, just more correct from a descriptive PoV
2026-07-10 07:31:25 -04:00
LC 401e542fcd Vector: Use unary handler for scalar unary insertions
Same behavior, but just uses a more proper handler.
2026-07-10 07:27:54 -04:00
LC 210514f74b Vector: LoadSourceGPR -> LoadSourceFPR for MASKMOVOp 2026-07-10 07:24:39 -04:00
LC dd838d4ad3 MemoryOps: Make use of current working reg for cache operations
TMP1 technically isn't initialized properly here until after the first
iteration.
2026-07-10 07:13:17 -04:00
Ryan Houdek 5f2d19c7aa Merge pull request #5685 from lioncash/faddv 2026-07-10 03:33:43 -07:00
LC d5ae87b5ca Thread: Amend pidfd_open check
Checks for success.
2026-07-10 06:25:43 -04:00
Ryan Houdek 95e7c866cf Merge pull request #5682 from lioncash/pid 2026-07-10 03:20:23 -07:00
LC 02028eb1ad x32/Msg: Fix result comparison in mq_getsetattr
Checks against failure instead of 1.
2026-07-10 06:19:11 -04:00
Ryan Houdek 2effdb04bd Merge pull request #5684 from lioncash/midr 2026-07-10 03:18:25 -07:00
Ryan Houdek ab31e3e3bb Merge pull request #5683 from lioncash/mul 2026-07-10 03:18:09 -07:00
LC 37b1432514 x32/Thread: Fix off-by-one in get_thread_area
12, 13, and 14 are the only valid TLS areas.
2026-07-10 06:06:30 -04:00
LC 6e8bc337aa x32/FD: Ensure sendfile updates offset if set 2026-07-10 05:58:49 -04:00
Ryan Houdek 8347566815 Merge pull request #5681 from lioncash/timer 2026-07-10 02:51:20 -07:00
LC ede09a03db VectorOps: Fix 256-bit FADDV path
Avoids falling down to the SVE-128 path.
2026-07-10 05:50:22 -04:00
LC 849d60253c CPUID: Fix MIDR walking in SetupHostHybridFlag() 2026-07-10 05:46:16 -04:00
LC b5660c8a92 MiscOps: Avoid stack misalignment in ProcessorID
This needs to be an add.
2026-07-10 05:41:54 -04:00
LC 54263bb5a7 ALUOps: Fix 32-bit MulH case
These need to be 64-bit multiply and ubfx. Thankfully this case wasn't
actually hit in practice.
2026-07-10 05:39:14 -04:00
LC 4069a9f7d5 Timer: Fix typo in timer_gettime
This should be passed by reference rather than by value.
2026-07-10 05:33:06 -04:00
Ryan Houdek 77467eaf0c Merge pull request #5680 from lioncash/usrai
IR: Remove unused VUShraI
2026-07-10 01:50:51 -07:00
LC caa030714b IR: Remove unused VUShraI
Given that this is currently unused and that we don't have the signed
equivalent implemented, we can just remove this for now.
2026-07-10 04:15:52 -04:00
Ryan Houdek facbc78e0a Merge pull request #5679 from lioncash/macro
ALUOps: Remove unused macros
2026-07-10 00:50:52 -07:00
LC 6488dcbb01 ALUOps: Remove unused macros
These are now unused.
2026-07-10 03:30:03 -04:00
LC 37265b109a Merge pull request #5676 from Sonicadvance1/185
FEX: Remove FEXInterpreter binary
2026-07-09 18:39:36 -04:00
Ryan Houdek ebe7342d10 Merge pull request #5677 from mrpippy/protontso
Windows/UnixLib: Fix enabling TSO through legacy Proton codepath
2026-07-09 15:33:27 -07:00
Ryan Houdek 523bbef034 FEX: Remove FEXInterpreter binary
It's been ten months, a bit longer than than I was expecting to keep
this around. Go ahead and remove it now.
2026-07-09 15:11:30 -07:00
Brendan Shanks 84e127a637 Windows/UnixLib: Fix enabling TSO through legacy Proton codepath 2026-07-09 15:00:50 -07:00
Ryan Houdek 4a091df8cc Merge pull request #5672 from neobrain/feature_woa_code_cache_bitness
CodeCache/WoA: Support mixed WoW64/ARM64EC processing
2026-07-09 14:55:21 -07:00
Ryan Houdek 92a171ce53 Merge pull request #5675 from lioncash/branch
VectorOps: Join identical branches in VFMLS/VFNMLS
2026-07-09 13:58:54 -07:00
LC 1b1e46ff6c VectorOps: Join identical branches in VFMLS/VFNMLS
Same thing, just a little less redundant.
2026-07-09 16:33:26 -04:00
Ryan Houdek 3370d9af15 Merge pull request #5670 from simon902/MOVDoverride
Fix movd when prefixed with 0x66
2026-07-09 13:02:08 -07:00
Ryan Houdek 9306de79ad Merge pull request #5667 from simon902/CVTTSS2SIOverride
Fix cvttss2si when prefixed with 0x66
2026-07-09 12:52:00 -07:00
Ryan Houdek ff7a54add8 Merge pull request #5668 from OFFTKP/inf
Fix element getting overwritten in 66_5B test
2026-07-09 12:26:46 -07:00
Ryan Houdek c3d1157696 Merge pull request #5669 from OFFTKP/lzcnt
Fix LZCNT tests reading out of bounds
2026-07-09 12:24:43 -07:00
Ryan Houdek d0f03cb148 Merge pull request #5674 from lioncash/insertq
Vector: Trim one instruction off insertq
2026-07-09 12:24:06 -07:00
LC 9b7c9f0fb6 Vector: Trim one instruction off insertq
We can fold a bitwise not and and pair into a bic
2026-07-09 14:58:04 -04:00
Ryan Houdek e508b6df0d Merge pull request #5673 from lioncash/vbitwise
IR: Remove need to specify element size for vector bitwise ops
2026-07-09 11:42:37 -07:00
LC 710b85be70 IR: Remove need to specify element size for vector bitwise ops
Element size doesn't really mean anything here, considering all bits are
acted upon independently of segmentation.

Makes using these ops a little bit less noisy.
2026-07-09 13:26:51 -04:00
Tony Wasserka 66455b708a CodeCache: Switch between 32-/64-bit compilers during cache generation 2026-07-09 17:17:19 +02:00
Tony Wasserka 8b8000b98a CodeCache/WoA: Run cache generation in a subprocess to improve robustness 2026-07-09 17:16:16 +02:00
Tony Wasserka 54236df6e0 CodeCache: Record main executable bitness in code map
Code maps already contain the main executable they were recorded from, so
it's convenient to capture the executable's bitness along the way.
2026-07-09 17:12:45 +02:00
LC 6cd2a48910 Merge pull request #5666 from Sonicadvance1/184
FEXCore: Fixes a crash with multiblock if `ProcessorID` IR op is encountered
2026-07-09 10:16:06 -04:00
Simon Scherer fe08b96844 OpcodeDispatcher: Fix cvttss2si when prefixed with 0x66 2026-07-09 11:58:18 +02:00
Simon Scherer 4d78901420 OpcodeDispatcher: Fix movd when prefixed with 0x66 2026-07-09 11:46:48 +02:00
Simon Scherer c06468025c unittests/ASM: Test movd prefixed with 0x66 2026-07-09 11:46:10 +02:00
Paris Oplopoios 148e539025 Fix LZCNT tests reading out of bounds 2026-07-09 12:36:18 +03:00
Paris Oplopoios 92b96ff30d Fix element getting overwritten in 66_5B test 2026-07-09 11:54:38 +03:00
Simon Scherer 778df0c93b unittests/ASM: Test cvttss2si prefixed with 0x66 2026-07-09 09:24:59 +02:00
Ryan Houdek 9d18ecc5cb FEXCore: Fixes a crash with multiblock if ProcessorID IR op is encountered
If during multiblock code discovery a RDTSCP/RDPID instruction was
encountered then ProcessorID has an assert at JIT compile time. Make
sure to early exit with an illegal instruction encoding early instead.
Also make sure to correctly report RDPID support in CPUID, it's
technically a different bit than RDTSCP.

Fixes a crash in Crusader Kings 3's Paradox Launcher installer. Although
the installer seems to fail otherwise for some reason.
2026-07-08 17:41:03 -07:00
Ryan Houdek 5f2455c502 Merge pull request #5665 from lioncash/blendop
[SVE256] Handle 256-bit blend operations much more efficiently
2026-07-08 16:43:20 -07:00
LC 6bc67609a3 [SVE256] Handle 256-bit blend operations much more efficiently
We can massage a given selector into a valid predicate register bitmask
and then simply perform a merging move, which eliminates most busywork
around optimizing 256-bit blends.

In the future, once we drop SVE2.1 support in, we can use PMOV to
eliminate the load from memory and related constant management.
2026-07-08 17:35:12 -04:00
LC 1bd3945dd9 Merge pull request #5577 from Sonicadvance1/168
Context: Add support for single-step RIP ranges
2026-07-08 15:47:28 -04:00
Ryan Houdek 168f4b1e6b Context: Add support for single-step RIP ranges
Useful when debugging a range.
2026-07-08 12:21:16 -07:00
LC 8a8827c980 Merge pull request #5664 from simon902/CMPXCHGZeroing
OpcodeDispatcher: Fix 32bit cmpxchg zero extension with eax as first operand
2026-07-08 15:14:12 -04:00
Ryan Houdek 21a968b84d Merge pull request #5663 from simon902/PDEPoverlap
JIT/ALUOps: Fix operand overlapping bug for pdep
2026-07-08 11:08:12 -07:00
Simon Scherer 84fab84b3f InstcountCI: Update 2026-07-08 15:00:18 +02:00
Simon Scherer 3d65c030a8 OpcodeDispatcher: Fix 32bit cmpxchg zero extension with eax as destination operand and remove incorrect comment. 2026-07-08 14:58:32 +02:00
Simon Scherer 59097bab20 unittests/ASM: Test cmpxchg with eax as destination 2026-07-08 14:52:19 +02:00
Simon Scherer 4cbacd9261 InstcountCI: Update 2026-07-08 10:10:30 +02:00
Simon Scherer 655102fc7d JIT/ALUOps: Fix operand overlapping bug for pdep 2026-07-08 10:09:47 +02:00
Simon Scherer f718f46545 unittests/ASM: Test overlapping operands for pdep 2026-07-08 09:47:12 +02:00
LC 71d4e2c320 Merge pull request #5660 from Sonicadvance1/183
Wow64: Spin loop on atomic with WFE
2026-07-07 22:53:50 -04:00
Ryan Houdek 41241d7500 Wow64: Spin loop on atomic with WFE
Instead of burning roughly a million watts, put this spinloop on a WFE.
This tends to occur on a crash during shutdown that isn't fully able to
be avoided. The least we can do is not consume all the power in the
world.
2026-07-07 16:09:00 -07:00
Ryan Houdek b90c9836cb Merge pull request #5662 from lioncash/alias
OpcodeDispatcher: Remove asterisk from BMI source args
2026-07-07 11:22:59 -07:00
Ryan Houdek aff3fcf76d Merge pull request #5658 from neobrain/fix_woa_code_cache_ec
CodeCache: Mark executable memory as EC code on ARM64EC
2026-07-07 11:22:15 -07:00
Ryan Houdek ec2aa4063a Merge pull request #5661 from lioncash/blend
unittests: Add stress tests for VBLEND{PD, PS}
2026-07-07 11:13:56 -07:00
LC 718f2e01f7 OpcodeDispatcher: Remove asterisk from BMI source args
Keeps it consistent with the rest of the code and prevents breakages
whenever the Ref alias gets turned into its own value type.
2026-07-07 14:07:38 -04:00
LC ba9f7fb1b5 unittests: Add stress tests for VBLEND{PD, PS}
Forgot about these two
2026-07-07 13:49:09 -04:00
Tony Wasserka 0ac6b3e8f3 CodeCache: Mark executable memory as EC code on ARM64EC
See bd5b817c3a.
2026-07-07 15:55:27 +02:00
Ryan Houdek dddad1c2ca Merge pull request #5659 from lioncash/shuf
unittests: Add stress tests for VSHUF{PD, PS}
2026-07-06 15:50:08 -07:00
LC 95bfff20a4 unittests: Add stress tests for VSHUF{PD, PS}
Covers the remaining shuffle paths
2026-07-06 18:01:52 -04:00
LC 7a6f0def85 Merge pull request #5653 from Sonicadvance1/182
Config: Fixes AppOverrides with FEX_APP_CONFIG
2026-07-06 17:01:55 -04:00
Ryan Houdek 103d4d76be Config: Fixes AppOverrides with FEX_APP_CONFIG 2026-07-06 12:43:14 -07:00
Ryan Houdek db9414a756 Merge pull request #5655 from neobrain/feature_woa_cache_loading
Windows/ImageTracker: Adapt code cache loading logic to FEXOfflineCompiler
2026-07-06 12:40:40 -07:00
Ryan Houdek 5a0bf1bb5f Merge pull request #5657 from neobrain/fix_foc_syscall_abi_woa
FEXOfflineCompiler: Fix improper syscall ABI on WoA
2026-07-06 12:37:52 -07:00
LC 444c37fe2e Merge pull request #5656 from neobrain/fix_invalid_iterator_deref
Core: Fix dereference of invalid iterator
2026-07-06 10:13:04 -04:00
Tony Wasserka d3d735370f FEXOfflineCompiler: Fix improper syscall ABI on WoA 2026-07-06 15:16:30 +02:00
Tony Wasserka 852e93aa74 Core: Fix dereference of invalid iterator 2026-07-06 15:07:37 +02:00
Tony Wasserka dc1be2efe8 Windows/ImageTracker: Unindent refactored code 2026-07-06 14:58:16 +02:00
Tony Wasserka f2734ac608 Windows/ImageTracker: Adapt code cache loading logic to FEXOfflineCompiler 2026-07-06 14:58:16 +02:00
Tony Wasserka 57ca49dc5c Merge pull request #5501 from bylaws/finishloadwin
Windows/ImageTracker: Wire up LoadCache and EnableLoadedSection
2026-07-06 14:58:01 +02:00
Billy Laws 5ef3134a4d Windows/ImageTracker: Wire up LoadCache and EnableLoadedSection
The lazy code loading refactor replaced LoadData with the new
LoadCache/EnableLoadedSection API but left the Windows path as TODOs.
Implement the wiring: LoadAOTImages now calls LoadCache +
RegisterMappedCodeBuffer for each mapped cache file, and HandleImageMap
calls EnableLoadedSection (with nullptr thread since lazy mapping is not
yet implemented on Windows).
2026-07-06 14:44:59 +02:00
Ryan Houdek 5b91642883 Merge pull request #5654 from ShadowCurse/cache_va_size
Allocator: fix the caching of host va size
2026-07-05 18:32:10 -07:00
Egor Lazarchuk 1eb5abb8db Allocator: rename DetermineVASize to GetHostVABits
`DetermineVASize` does not return the size of VA, but the number of bits
it can use. Change the naming to make it more self explanatory.
In the mean time also move `HostVASize` global into `GetHostVABits`
since it is not and should not be used directly.
2026-07-05 12:46:41 +01:00
Egor Lazarchuk a65e1bf7e5 Allocator: fix the caching of host va size
Commit abf9724 ("Allocator: Fix and optimize VA range detection")
removed assignment to the `HostVASize` global thus making each call to
`DetermineVASize` redo all the work with potential to produce incorrect
results. Set the global again to fix this.
2026-07-05 12:46:29 +01:00
Ryan Houdek 3f98202b0f Merge pull request #5651 from lioncash/perm
unittests: Add stress tests for VPERMIL{PD, PS}
2026-07-04 12:13:24 -07:00
LC e7727c39f3 unittests: Add stress tests for VPERMIL{PD, PS}
While unlikely to be used in practice over other kind of
shuffling and blending, these should also have stress tests
to make sure they do the right thing.
2026-07-04 15:00:30 -04:00
Ryan Houdek 17e5637664 Merge pull request #5650 from lioncash/aes256
[SVE256] Handle 256-bit AES operations
2026-07-03 19:06:36 -07:00
LC 7fd9b897c2 [SVE256] Handle 256-bit AES operations
Currently we split these into two 128-bit operations since VIXL doesn't
have support for the unified SVE operations yet.

Now we fully support VAES on SVE256.
2026-07-03 21:52:59 -04:00
Ryan Houdek f8967aa207 Merge pull request #5649 from lioncash/pclmul256
[SVE256] EncryptionOps: Handle 256-bit VPCLMULQDQ
2026-07-03 18:33:53 -07:00
LC 5145324806 [SVE256] EncryptionOps: Handle 256-bit VPCLMULQDQ
Since vixl now handles this, we can drop this support right in.
2026-07-03 21:22:05 -04:00
Ryan Houdek 4233fb6270 Merge pull request #5648 from lioncash/aes
[SVE256] Ensure SSE insertion behavior for AES/SHA/PCLMUL operations
2026-07-03 17:33:19 -07:00
LC d11b19fd2b [SVE256] Ensure insertion behavior for PCLMUL SSE operations
Also includes accompanying test to ensure it never breaks.
2026-07-03 20:15:34 -04:00
LC 684c568033 [SVE256] Ensure insertion behavior for SHA SSE operations
These slipped through, so now we can add tests for them to prevent that
from happening again.
2026-07-03 20:09:50 -04:00
LC ab4fb7b3ad [SVE256] Ensure insertion behavior for AES operations on SSE
These slipped through, so now we can add tests for them to prevent that
from happening again.
2026-07-03 19:35:31 -04:00
Ryan Houdek 91017dbedb Merge pull request #5647 from lioncash/perm128
unittests: Add stress test for vperm2f128/vperm2i128
2026-07-03 15:32:34 -07:00
LC f5e8a051a8 unittests: Add stress test for vperm2f128/vperm2i128
Lets us test all possible immediate encodings for proper behavior.
2026-07-03 16:17:12 -04:00
LC 1d695f6db4 Merge pull request #5646 from Sonicadvance1/181
Github: More dependabot things
2026-07-02 22:22:45 -04:00
Ryan Houdek d4bdfd0592 Github: More dependabot things
They just never stop.
2026-07-02 19:05:12 -07:00
Ryan Houdek 1cc4b93e7a Docs: Update for release FEX-2607 2026-07-02 17:47:31 -07:00
LC b4e2f5118a Merge pull request #5645 from Sonicadvance1/180
Windows: Fixes SHM stats reallocation
2026-07-02 19:33:37 -04:00
Ryan Houdek 812b6398e5 Windows: Fixes SHM stats reallocation
This was accidentally setting `CurrentSize` instead of just returning
the newly allocated size to the frontend. This was causing the frontend
to then fail to detect the reallocation actually occured and no longer
get stats for new threads.

Also happened to not use `NewSize` but instead `CurrentSize * 2` which
didn't matter as it matched the growth pattern, but was technically
incorrect.

Fixes SHM stats since the introduction of the unixlib, ezpz.
2026-07-02 13:38:14 -07:00
LC 1db45e2a70 Merge pull request #5639 from Sonicadvance1/179
FEXServerClient: Workaround sun_path 108 byte limit
2026-07-02 04:03:48 -04:00
Ryan Houdek 6bcadde658 Merge pull request #5644 from lioncash/vmov
VectorOps: Eliminate unnecessary moves in VMov if applicable
2026-07-01 17:28:42 -07:00
Ryan Houdek 5f1c8efe0e Merge pull request #5643 from lioncash/rec
VectorOps: Avoid temporary if able in 256-bit VFRecp
2026-07-01 17:25:33 -07:00
Ryan Houdek b9aeccf13b Merge pull request #5642 from lioncash/feature
HostFeatures: Put SVE support querying into single function
2026-07-01 15:29:02 -07:00
Ryan Houdek 16f90b33f3 Merge pull request #5641 from lioncash/minmax
VectorOps: Reduce temporary usage in 64-bit AdvSIMD min max paths
2026-07-01 14:17:18 -07:00
Ryan Houdek 5d8d052a77 Merge pull request #5640 from lioncash/move
VectorOps: Avoid move in 256-bit VAddP/VFAddP if possible
2026-07-01 12:41:03 -07:00
Tony Wasserka 7fa4d78269 Merge pull request #5635 from Sonicadvance1/178
FEXOfflineCompiler: Fixes HostFeature detection under Win32
2026-07-01 17:15:13 +02:00
Ryan Houdek 6bf0db7df6 FEXServerClient: Workaround sun_path 108 byte limit
We really don't want to do this, but in the case that the AF_UNIX path
is longer than the 108-byte limit that sun_path provides we don't really
have a choice. The alternative choice would be to switch /entirely/ away
from AF_UNIX and instead use pipes. We need a bandage fix for now, so
throw the socket in to a temp folder if the path is too long.
2026-06-30 17:31:23 -07:00
Ryan Houdek 110313e7de Merge pull request #5638 from lioncash/swap
IR: Add constant for swapping midsections of 256-bit vectors around
2026-06-30 16:38:56 -07:00
Ryan Houdek 44e24c9e6b Merge pull request #5637 from mrpippy/unicodestring
Windows/UnixLib: Fix loading with new MemoryWineLoadUnixLibByName mechanism
2026-06-30 16:23:08 -07:00
Brendan Shanks 6c58fef220 Windows/UnixLib: Fix loading with new MemoryWineLoadUnixLibByName mechanism. 2026-06-30 15:37:55 -07:00
LC 3d593ce87d Merge pull request #5626 from Sonicadvance1/177
Re-optimize FYL2X for reduced precision x87 path
2026-06-30 18:15:17 -04:00
Ryan Houdek 36e1b5107e InstcountCI: Update 2026-06-30 15:01:46 -07:00
Ryan Houdek f89123f489 unittests/ASM: Allow up to 3-bits of precision loss 2026-06-30 15:00:16 -07:00
Paulo Matos b7280a765d asm_tests: Re-optimize FYL2X for reduced precision x87 path 2026-06-30 15:00:16 -07:00
Paulo Matos 37b010795e Re-optimize FYL2X for reduced precision x87 path 2026-06-30 15:00:16 -07:00
Ryan Houdek a4f89b79a3 Merge pull request #5636 from lioncash/move
AVX: Remove unnecessary moves from VPSHUF{D, HW, LW}
2026-06-30 12:35:13 -07:00
Ryan Houdek 5bde4d875a FEXOfflineCompiler: Fixes HostFeature detection under Win32
CPUFeature detection is marginally different between Linux and Windows.
FOC was only using the Linux path which had two broken things happening
to it.
- Feature detection was incorrect and enabling/disabling features
  differently from wow64/arm64ec .dll files
- HostType was being set as Linux even though it was generating code for
  WINE

Ensure that when built for Win32 that it uses the correct feature
fetching.

One thing that is still incorrect is that 64-bit or 32-bit is determined
at compile time on win32, whereas the Linux side parses an ELF and
determines bitness at runtime. This doesn't fix that remaining problem
there.
2026-06-30 11:39:02 -07:00
Ryan Houdek a79c471c31 Merge pull request #5634 from lioncash/shuffle
unittests: Add selector tests for VPSHUF{D, HW, LW}
2026-06-30 11:26:09 -07:00
LC c1e29f9013 VectorOps: Eliminate unnecessary moves in VMov if applicable
If the destination and source don't match, then we can just zero
and insert directly into the destination instead of a temporary.
2026-06-30 05:06:26 -04:00
LC 4c27dfd5eb VectorOps: Avoid temporary if able in 256-bit VFRecp
If we're non-aliasing, we can make use of the destination reg directly.
Makes the non-RPRES path a little nicer.
2026-06-30 04:36:20 -04:00
LC 201216ba54 HostFeatures: Put SVE support querying into single function
Lets us avoid open-coding long checks for the existence of either
SVE-128 or SVE-256.
2026-06-30 03:34:17 -04:00
LC d3a85e14d9 VectorOps: Reduce temporary usage in 64-bit AdvSIMD min max paths
We can reorganize these such that they only use one temporary in the
worst case instead of two.
2026-06-30 01:30:40 -04:00
Ryan Houdek 417bd8604c Merge pull request #5632 from lioncash/vpblendd_test
unittests: Add test for stress-testing VPBLENDD selectors
2026-06-29 22:13:05 -07:00
LC 3d289f4489 VectorOps: Avoid move in 256-bit VAddP/VFAddP if possible
If the first source and destination alias, then it's fine to use the
register in destructive operations, since the source data doesn't need
to be preserved.

Tiny saving, but reduces overall register use in some cases.
2026-06-30 00:23:30 -04:00
Ryan Houdek d138c854f3 Merge pull request #5631 from lioncash/vpblendd
instcountci/VEX_map3: Add missing third param to VPBLENDD
2026-06-29 20:11:57 -07:00
Ryan Houdek e26a792b70 Merge pull request #5630 from lioncash/comiss
Vector: Only signify 128-bit vector loads in UCOMISxOp
2026-06-29 14:40:44 -07:00
Ryan Houdek 7e2d3b07c0 Merge pull request #5629 from lioncash/comment
instcountci/VEX_map1: Remove obsolete comments
2026-06-29 13:55:23 -07:00
Ryan Houdek 394a6f28db Merge pull request #5628 from lioncash/movmsk
AVX: Reduce codegen for 256-bit VMOVMSKPD/VMOVMSKPS
2026-06-29 13:42:17 -07:00
Ryan Houdek 72e01274c9 Merge pull request #5627 from lioncash/vpblendw
instcountci: Add a few more cases for VPBLENDW/VPBLENDD and VSHUFPD/VSHUFPS
2026-06-29 13:03:49 -07:00
Ryan Houdek 9f2e982944 Merge pull request #5617 from simon902/vcvtps2ph
OpcodeDispatcher: Fix missing zeroing for vcvtps2ph
2026-06-29 10:40:45 -07:00
LC c6b0f360fe AVX: Make use of table swapping constant to trim down relevant ops
Now we can get rid of excessive overhead, with the ability to improve
this further in the future.
2026-06-29 12:51:30 -04:00
LC ef19242be3 IR: Add constant for swapping midsections of 256-bit vectors around
This'll let us eliminate some pretty gnarly inserts while
ZIP1Q/UZP1Q/etc are still unavailable.
2026-06-29 12:50:53 -04:00
LC 2cb4f8b6f5 AVX: Remove unnecessary moves from VPSHUF{D, HW, LW}
These aren't necessary anymore, since these use operations
that already zero extend.
2026-06-29 08:43:11 -04:00
LC 59919a0b0c unittests: Expand VPBLENDW selector test to cover 128-bit paths
Ensures both codepaths are tested. Technically also acts as a good
test for selector masking as well.
2026-06-29 07:03:16 -04:00
LC bf1857ecc7 unittests: Expand VPBLENDD selector test to cover 128-bit paths
Ensures both codepaths are tested. Technically also acts as a good
test for selector masking as well.
2026-06-29 06:56:37 -04:00
LC b40920596a unittests: Add selector stress tests for VPSHUFD
Ensures any added optimization paths result in the same output.
2026-06-29 06:56:34 -04:00
LC 3b14c322e8 unittests: Add selector stress tests for VPSHUFLW/VPSHUFHW
Ensures any added optimization paths result in the same output.
2026-06-29 06:45:18 -04:00
LC 3e60aa5738 Merge pull request #5610 from Sonicadvance1/171
FEXCore: Pass host type that changes codegen to FEXCore
2026-06-29 02:04:31 -04:00
Simon Scherer 6ca2d27c82 unittests/ASM: Add vcvtps2ph_zeroing test to Disabled_Tests_Simulator 2026-06-29 08:00:58 +02:00
Simon Scherer d2f26d1969 InstcountCI: Update 2026-06-29 07:55:05 +02:00
LC fc25443827 unittests: Move full_vpblendw_imm test into the VEX folder
Keeps all of the selector tests in the same location.
2026-06-28 23:17:19 -04:00
LC 986800885a unittests: Add test for stress-testing VPBLENDD selectors
Drops a test in like the one for VPERMQ to ensure that, even if different
optimization paths are introduced, the behavior remains consistent.
2026-06-28 23:13:39 -04:00
LC 52c6ab1cec instcountci/VEX_map3: Add missing third param to VPBLENDD
Ensures all registers are non-aliasing, which makes for better unideal
output for observation.
2026-06-28 21:38:43 -04:00
Ryan Houdek 83a989c6dc Merge pull request #5625 from lioncash/perm
AVX: Handle two field insertions in VPERMQ
2026-06-28 17:08:48 -07:00
LC 32b96c259b Merge pull request #5623 from Sonicadvance1/176
Windows/UnixLib: Adds remaining helpers
2026-06-28 18:44:35 -04:00
Ryan Houdek 954581c750 Windows/UnixLib: Adds remaining helpers
Centralizes all the nasty behaviour that will end up breaking when WINE
eventually turns on userspace syscall dispatch. Pushes all of the logic
in to the UnixLib. Support both paths until everyone is migrated to
supporting the UnixLib, then we can delete the bit of code duplication
between the PE side and UnixLib side.

Helpful that everything that gets punched through the UnixLib is
optional, so worst case some optional bits can break for a while.
2026-06-28 12:58:24 -07:00
LC 24da43f823 Merge pull request #5613 from Sonicadvance1/175
Windows/UnixLib: Adds support for Hardware TSO support
2026-06-28 15:56:09 -04:00
LC 600e2ddecf Vector: Only signify 128-bit vector loads in UCOMISxOp
Mainly a correctness change more than anything. The COMISX and
UCOMISX group of operations only ever load 128 bits when given
a vector source.

No change in codegen, but ensures this doesn't change if any backing
handling changes.
2026-06-28 15:40:14 -04:00
Ryan Houdek 9740f488cf Merge pull request #5624 from lioncash/psad
AVX: Slightly trim codegen for VPSADW 256-bit case
2026-06-28 12:26:37 -07:00
LC f81467fd30 instcountci/VEX_map1: Remove obsolete comments
Since the registers are non-aliasing, this is about the best we can do
now. These are just holdovers from the initial bring-up of the 256-bit
SVE path where optimization wasn't as strong a concern as getting everything
in place and running properly.
2026-06-28 14:49:24 -04:00
LC 2ec2c39cf1 AVX: Lessen codegen for VMOVMSKPD
Performs the same thing as what we've done for VMOVMSKPS. Though,
the effect isn't as drastic, given we're only operating on a max
of four elements as opposed to 8.
2026-06-28 14:06:32 -04:00
Ryan Houdek c85e426cc5 Merge pull request #5622 from lioncash/perm
AVX: Handle trivial UZP/ZIP operations in VPERMQ
2026-06-28 11:01:49 -07:00
LC 41fc57f46c AVX: Lessen codegen for 256-bit VMOVMSKPS
Currently we can trivially split this up and join the results, which
is much nicer than iterating all the elements individually and shifting
their sign bit over.
2026-06-28 13:52:19 -04:00
LC a215bb9709 instcountci: Add a few more cases for VPBLENDW/VPBLENDD and VSHUFPD/VSHUFPS
Just provides a little more comprehensive output
2026-06-28 13:08:03 -04:00
Ryan Houdek b23fa30099 Merge pull request #5618 from wsxarcher/fixsmcfullvector
Core: Start a new block for the next op in full SMC check
2026-06-28 09:25:23 -07:00
Ryan Houdek 848c4b2d68 Merge pull request #5621 from wsxarcher/fixfexfolder
Do not exec FEX if it is a folder in FEXBash
2026-06-28 09:22:14 -07:00
Ryan Houdek 4f995cbc1c Merge pull request #5620 from wsxarcher/fixuafbash
Fix UAF of PS1
2026-06-28 09:08:51 -07:00
wsxarcher b21c49352e Core: Start a new block for the next op in full SMC check 2026-06-28 18:07:28 +02:00
Ryan Houdek 3470dd1e7b Merge pull request #5619 from lioncash/dpps
AVX: Handle trivial cases better for VDPPS
2026-06-28 08:55:34 -07:00
wsxarcher 2ba35286ef Do not exec FEX if it is a folder 2026-06-28 17:28:41 +02:00
wsxarcher bbd8212ced Fix UAF of PS1 2026-06-28 17:18:05 +02:00
Simon Scherer 4564325bc7 OpcodeDispatcher: Fix upper 128 bit zeroing for vcvtps2ph 2026-06-28 13:17:06 +02:00
Simon Scherer b9052ed7f0 unittests/ASM: Test zeroing of vcvtps2ph 2026-06-28 13:14:26 +02:00
Ryan Houdek 4160a92621 Windows: Fixes duplicated hardware TSO handling
This was handled in both the Module.cpp files and also the common
TSOHandlerConfig on accident. Wouldn't have caused an issue but it was
definitely a bit weird.
2026-06-27 20:59:15 -07:00
Ryan Houdek 201bb73980 Windows/UnixLib: Adds support for Hardware TSO support
Including fallback to non-unixlib path because we need to support both.

Showcases how these are going to be implemented without throwing the
entire world at it right away. Next PR will be implementing the
remaining four necessary unixlib handlers that we will require:

- Kernel unaligned atomic control
- shm_stats thing
- madvise operation
- prctl vma naming
2026-06-27 19:59:35 -07:00
LC 78832cc0d0 Merge pull request #5612 from Sonicadvance1/174
Windows: Load unixlib if possible
2026-06-27 22:25:43 -04:00
LC a5ecb71993 AVX: Handle two field insertions in VPERMQ
Lets us trivially handle fields like 0baa'aa'bb'bb
as broadcasts and an insert.
2026-06-27 19:35:53 -04:00
LC 6d86dca20b AVX: Slightly trim codegen for VPSADW 256-bit case
We can massage this a little bit to be slightly better. At least
gets rid of the heavyweight inserts.
2026-06-27 14:16:48 -04:00
LC e24f84504f AVX: Handle trivial UZP/ZIP operations in VPERMQ
Handles cases where a permutation can be simplified into a single
zip/unzip operation.
2026-06-27 12:04:13 -04:00
LC ee2fb57f4e AVX: Handle full broadcast in VDPPS
Another trivial case that can be handled without crazy codegen.
2026-06-27 10:11:27 -04:00
LC 11fe95d8ea AVX: Simplify trivial case of VDPPS
Just a silly case where we only need to return the zero vector
2026-06-27 09:55:35 -04:00
Ryan Houdek 70fe9a4405 Merge pull request #5614 from lioncash/whoops
Vector: Fix typo in VPERMQOp
2026-06-26 21:45:33 -07:00
Ryan Houdek c09225f868 Windows: Load unixlib if possible
Currently does nothing other than load it (as the library also doesn't
do anything yet). Ensured it was working by temporarily creating a test
entrypoint and doing `Call` on to it.

Next step after this is to reimplement some of the nasty hacks FEX is
doing inside the unixlib code itself.
2026-06-26 20:50:24 -07:00
Ryan Houdek 126bcd365d winternl: Update enums
Newer WINE has a better mechanism for asking to load unix libraries.
Older WINE like what is in Proton doesn't have this yet. Add definitions
for both so we can try either one.
2026-06-26 20:50:19 -07:00
LC fe4d2bc6c5 Merge pull request #5611 from Sonicadvance1/173
Windows: Adds empty Linux side unix library
2026-06-26 23:48:51 -04:00
Ryan Houdek dbaf22372c Windows: Adds empty Linux side unix library
We are going to need a unix library. Going to take this one step at a
time without AI/ML so I fully understand all the pieces of the puzzle,
and to ensure we don't lose any functionality before we're ready.

This only ensures that we are building the Linux facing .so files for
arm64ec and wow64, but they are empty today. Next PR will be
initializing it on the PE side.
2026-06-26 19:05:08 -07:00
LC 43bd243457 Merge pull request #5609 from Sonicadvance1/170
OpcodeDispatcher: Fixes CRC32 with high 8-bit register
2026-06-26 15:30:00 -04:00
Ryan Houdek c0251dc8be FEXCore: Pass host type that changes codegen to FEXCore
Because these compile options change codegen, we need to make sure these
are runtime selected rather than compile-time selected. Will reduce
code-cache variance.
2026-06-26 12:06:30 -07:00
Ryan Houdek 26266c6a94 InstcountCI: Adds CRC32 with high 8-bit register 2026-06-26 11:38:08 -07:00
Ryan Houdek 64392b2d45 OpcodeDispatcher: Fixes CRC32 with high 8-bit register
Assertion failure in `_Bfe` IR operation when encountering this
instruction. Ensure the GPR source is sized appropriately.
2026-06-26 11:34:45 -07:00
LC d555ee8bcc Vector: Fix typo in VPERMQOp
Noticed this in my own writing and it bothered me.
2026-06-26 14:16:54 -04:00
Simon Scherer 0c1a35f297 unittests/ASM: Add unit test for crc32 with 8bit register operand 2026-06-26 11:10:06 -07:00
Tony Wasserka 9ac608ca43 Merge pull request #5606 from Sonicadvance1/169
LibraryForwarding/cuda: Convert constexpr to const
2026-06-26 11:17:49 +02:00
Ryan Houdek 7ae55d73c1 Merge pull request #5607 from lioncash/broadcast
AVX: Handle easily broadcastable permutations in VPERMQ
2026-06-25 20:23:55 -07:00
LC 8102a0974a AVX: Handle easily broadcastable permutations in VPERMQ
When we have a 3 element identical permutation followed by a single
unique outlier, we can simplify the whole operation into a single
broadcast followed by an insert.

e.g.

0b00'00'00'01
0b00'01'01'01
0b11'11'00'11

are all examples of cases where we can broadcast and then insert.
2026-06-25 21:48:56 -04:00
Ryan Houdek f8491794d3 thunks/cuda: Convert constexpr to const
Apparently some compilers or libstdc++ or libc++ takes offence to
constexpr std::array that gets filled by GOT. My compiler this generates
the same code regardless but I guess this'll probably fix #5582.
2026-06-25 16:53:59 -07:00
Ryan Houdek d5be15c90e Merge pull request #5605 from lioncash/permq
AVX: Skip identity insertions in VPERMQ
2026-06-25 11:51:25 -07:00
LC e91efc6694 AVX: Skip identity insertions in VPERMQ
In the slower case, if our iteration index and the selector index match,
then all that means is that we'd be inserting the same data that already exists
at that location, so we can skip the insertion in that case.
2026-06-25 14:33:49 -04:00
Ryan Houdek 3d66be9e5a Merge pull request #5604 from lioncash/pcmpstr
instcountci: Add 16-bit pcmpxstrx variants
2026-06-25 10:23:59 -07:00
Ryan Houdek ec1b24d05b Merge pull request #5603 from lioncash/permq
AVX: Handle transpose cases in VPERMQ
2026-06-25 01:18:01 -07:00
Ryan Houdek d93997c1cb Merge pull request #5602 from lioncash/same
AVX: Remove unnecessary dup if sources are the same in SHUFOpImpl
2026-06-24 23:38:29 -07:00
Ryan Houdek e19aa975c8 Merge pull request #5597 from simon902/BTOpTypo
OpcodeDispatcher: Fix incorrect comment for BTOp. ZF must be preserved.
2026-06-24 23:31:09 -07:00
Simon Scherer 3f02dd0a36 OpcodeDispatcher: Update BTOp comment to clarify AMD vs Intel flag behavior 2026-06-25 07:36:32 +02:00
Ryan Houdek ab4e0f653a Merge pull request #5600 from lioncash/palign
AVX: Shave some moves off 256-bit VPALIGNR
2026-06-24 11:18:51 -07:00
Ryan Houdek fda023e7dc Merge pull request #5601 from lioncash/pshufb
AVX: Reduce moves in 256-bit VPSHUFB
2026-06-24 10:44:07 -07:00
Ryan Houdek ad618be979 Merge pull request #5599 from lioncash/unused
Vector: Remove unused OpcodeArgs parameter from SHUFOpImpl
2026-06-24 09:47:38 -07:00
Ryan Houdek 5881266256 Merge pull request #5598 from lioncash/ilpd
AVX: Wire up helper to VPERMILPD
2026-06-24 09:03:22 -07:00
Simon Scherer e03187852b OpcodeDispatcher: Fix incorrect comment for BTOp. ZF must be preserved. 2026-06-24 16:11:13 +02:00
Ryan Houdek b846cb6c2c Merge pull request #5596 from lioncash/ilps
AVX: Wire up lane helper for VPERMILPS imm variant
2026-06-24 00:56:36 -07:00
Ryan Houdek 62ecdd650f Merge pull request #5595 from lioncash/pspd
AVX: Wire up lane helper for VSHUFPD/VSHUFPS
2026-06-24 00:22:12 -07:00
Ryan Houdek 9f195ff377 Merge pull request #5594 from lioncash/pshuf
AVX: Wire up lane helper for VPSHUFD/VPSHUFLW/VPSHUFHW
2026-06-23 20:31:04 -07:00
Ryan Houdek 27a5f09185 Merge pull request #5592 from ShadowCurse/fixes
JIT: Arm64: fix the loop in CacheLineClear/Clean
2026-06-23 17:15:07 -07:00
Ryan Houdek 01b0b4e653 Merge pull request #5593 from lioncash/dup
VectorOps: Avoid dup if able in VInsElement 128-bit element path
2026-06-23 17:10:35 -07:00
Egor Lazarchuk ff5dfff5bb IR: fix typo in CacheLineClean description 2026-06-24 00:38:59 +01:00
Egor Lazarchuk e3e9777ee6 JIT: Arm64: fix the loop in CacheLineClear/Clean
These functions need to clean at least 64 bytes of cache since this is
the default on x86_64, but previously they could clean less if
DCacheLineSize was smaller than 64 bytes.
2026-06-24 00:38:48 +01:00
Ryan Houdek 7dc2dc8749 Merge pull request #5591 from lioncash/typo
OpcodeDispatcher: Fix typo in comment
2026-06-23 16:16:29 -07:00
Ryan Houdek 4eb5694872 Merge pull request #5590 from lioncash/insert
AVX: Reduce inserts in VBLENDPS/VPBLENDD/VPBLENDW
2026-06-23 15:21:45 -07:00
Ryan Houdek 681c5e8097 Merge pull request #5589 from lioncash/selector
OpcodeDispatcher: Sanitize selectors for VBLENDPD/VPBLENDD
2026-06-23 13:01:41 -07:00
Ryan Houdek 5ad06b0255 Merge pull request #5588 from lioncash/pinsr
AVX: Remove unnecessary moves from PINSRX ops
2026-06-23 12:34:03 -07:00
LC a2889a09e3 Vector: Move zero constant closer to use in PCMPXSTRXOpImpl
Same behavior, but constrains the only scope it's used in.
2026-06-23 09:35:41 -04:00
LC a4cd5f7584 instcountci: Add 16-bit pcmpxstrx variants
Also adds expanded mask variants. Lets us get a better whole picture on
all the main paths of these instructions.
2026-06-23 09:25:37 -04:00
LC cf098a0de6 AVX: Handle transpose cases in VPERMQ
These can be single instruction operations.
2026-06-23 00:39:42 -04:00
Ryan Houdek 1619374252 Merge pull request #5587 from lioncash/pd
[SVE256] Add fast paths for trivial VSHUFPD flags (0b0000, and 0b1111)
2026-06-22 21:04:42 -07:00
Ryan Houdek c13064e201 Merge pull request #5586 from lioncash/mov
[SVE256] Remove unnecessary move in VCVTPS2PD
2026-06-22 20:44:23 -07:00
LC 20647f2287 AVX: Remove unnecessary dup if sources are the same in SHUFOpImpl
Eliminates a trivial move.
2026-06-22 22:15:34 -04:00
Ryan Houdek 37e32fbcb9 Merge pull request #5585 from lioncash/cmp-move
[SVE256] Remove heavy handed moves from scalar compares
2026-06-22 17:26:37 -07:00
Ryan Houdek 27acbba52e Merge pull request #5584 from lioncash/cmp_scalar
[SVE256] More comprehensively test SSE insertions for scalar comparisons
2026-06-22 15:49:57 -07:00
Ryan Houdek 6f33d2b4c4 Merge pull request #5583 from lioncash/scalar
[SVE256] Add more SSE scalar variant unit tests
2026-06-22 12:41:52 -07:00
LC a2e4209f4f AVX: Reduce moves in 256-bit VPSHUFB
Just a minor reduction by avoiding insertion overhead.
2026-06-22 10:47:49 -04:00
LC 73d6716828 AVX: Shave some moves off 256-bit VPALIGNR
Arbitrary insertion of an element requires the use of a predicate
register. Since we only care about a particular element in the vector,
being replicated, we can broadcast that element instead of doing an
insert, which is effectively the same thing without excessive busywork.
2026-06-22 10:09:13 -04:00
LC 1fb2419be2 Vector: Remove unused OpcodeArgs parameter from SHUFOpImpl
No behavior change, just a reduction in noise.
2026-06-22 09:47:46 -04:00
LC 3394808c06 AVX: Wire up helper to VPERMILPD
We can just leverage the shuffle handler for this, since VPSHUFD
essentially functions like VPERMILPD
2026-06-22 09:03:22 -04:00
LC c252b58a15 AVX: Wire up lane helper for VPERMILPS imm variant
Makes for some more trivial savings. Will need handling for VPERMILPD
added separately, since selector behavior is different.
2026-06-22 05:21:10 -04:00
LC 39ae8c3ea0 AVX: Wire up lane helper for VSHUFPD/VSHUFPS
Also allows collapsing quite a bit of emitted code, like with
the shuffles in #5594
2026-06-22 04:46:22 -04:00
LC eb8c2d964c Vector: Factor out 128-bit path in SHUFOpImpl
We can leverage this for the 256-bit path
2026-06-22 03:47:20 -04:00
LC bd9cf9ca11 AVX: Wire up lane helper for VPSHUFD/VPSHUFLW/VPSHUFHW
Lets the AVX implementation get all the optimizations that the SSE
variant has, reducing the overhead a little.

Even with the individual lane handling, this is still leagues better
than all of the individual inserts that are pretty beefy with SVE.

For example:

vpshufd ymm0, ymm1, 0b00000011

drops from 50 instructions to 9
2026-06-22 00:16:39 -04:00
Ryan Houdek 280568df2f Merge pull request #5581 from lioncash/fwd
Passes: Trim unnecessary forward declarations
2026-06-21 20:05:23 -07:00
Ryan Houdek b87ff1e2dc Merge pull request #5580 from lioncash/list
IntrusiveIRList: Amend signature for PostRA()
2026-06-21 20:04:51 -07:00
Ryan Houdek 55c90cfc38 Merge pull request #5579 from lioncash/typo
JIT: Amend op typos in implementations
2026-06-21 20:04:19 -07:00
Ryan Houdek 500d2374a5 Merge pull request #5578 from lioncash/tidy
Arm64Emitter: Tidy up load/stores in Push/PopCalleeSavedRegisters
2026-06-21 20:03:47 -07:00
LC d78963c021 VectorOps: Avoid dup if able in VInsElement 128-bit element path
We don't need to broadcast if we're inserting across registers into the
equivalent position, since we already have a predicate around that can
satisfy that.
2026-06-21 20:32:37 -04:00
LC 9ab0920f01 OpcodeDispatcher: Fix typo in comment
It's the bits in general, not just the even ones (whoops).
2026-06-21 19:16:38 -04:00
LC a6e7fba433 AVX: Reduce inserts in VBLENDPS/VPBLENDD/VPBLENDW
Lets us reduce inserts by seeing which bits in the selector mask
indicates a particular source is used more than the other one, and
then just uses that as the base to be inserted into, cutting down
on overall insertion overhead.

In some cases, this can be quite drastic, like with:

vpblendw ymm0, ymm1, ymm2, 0b00000001

being cut down from 98 instructions to 14.
2026-06-21 18:24:16 -04:00
LC 8905e39439 OpcodeDispatcher: Sanitize selectors for VBLENDPD/VPBLENDD
Ensures that junk values don't make their way through
2026-06-21 16:27:29 -04:00
LC df41b85827 OpcodeDispatcher: Merge VPINSRB/VPINSRW handling
We can just pass the size through Bind instead of having two functions
that effectively do the same thing, only differing on element size.
2026-06-21 15:33:37 -04:00
LC 9fa3b9345e AVX: Remove unnecessary moves from PINSRX ops
These are old paths still around from when StoreResult used to
automatically perform truncating moves.

These aren't necessary anymore, since the AdvSIMD operation already
ensures zero-extension.
2026-06-21 15:21:37 -04:00
LC 42af6c8508 [SVE256] Add fast paths for trivial VSHUFPD flags (0b0000, and 0b1111)
Lets us at least flatten down two paths from 20 instructions to 1.
2026-06-21 06:22:56 -04:00
LC b7df1bc259 [SVE256] Remove unnecessary move in VCVTPS2PD
FCVTL will already perform the truncation, so the subsequent move
isn't necessary.
2026-06-21 05:03:59 -04:00
LC b40f9db735 [SVE256] Remove heavy handed moves from scalar compares
(See #3799)

I had a feeling #5569 was a little overkill, but was just getting
everything up to a functional baseline at the time. Now, with the tests
added in #5584 to test all SSE paths, I was able to see which comparisons
in particular were the ones that would have deviating behavior (NLT and NLE)

This lets us safely restore the behavior without the excessive moves on
hardware that makes use of FEAT_AFP.
2026-06-21 02:27:53 -04:00
LC 938841c658 [SVE256] More comprehensively test SSE insertions for scalar comparisons
See #3799

Drops in the facilities to ensure all of the available SSE scalar comparison
paths are tested for proper insertion behavior.
2026-06-21 00:35:47 -04:00
LC 34344e5769 Merge pull request #5576 from Sonicadvance1/167
Proton: Fixes Mafia 3
2026-06-20 23:18:33 -04:00
Ryan Houdek 538a9624ec ArchHelpers/Arm64: Fixes zero register usage
The compiler is smart enough to use the zero register for atomic
operations. Our JIT never generated code like this so it was unexpected.
Make sure handle zero register in all the cases where it matters.
2026-06-20 19:56:18 -07:00
Ryan Houdek c4c69ca8de ArchHelpers/Arm64: Support CAS/CASP in non-JIT SIGBUS handler
Proton was using this
2026-06-20 19:55:39 -07:00
LC 90330ab3e4 [SVE256] Add more SSE scalar variant unit tests
See #3799

These were technically already covered when the work was done to make
SSE insertion behavior conform to hardware, so this just adds tests
that ensure that behavior holds over time.
2026-06-20 19:25:19 -04:00
Ryan Houdek c5880e7618 Merge pull request #5572 from lioncash/str
[SVE256] Handle SSE insertions for PCMPESTRM/PCMPISTRM ops
2026-06-20 15:13:21 -07:00
Ryan Houdek 9d0c05d9cc Merge pull request #5575 from lioncash/movq2dq
[SVE256] Handle SSE insertions for MOVQ2DQ
2026-06-20 12:48:31 -07:00
Ryan Houdek f6a68cb7fd Merge pull request #5574 from lioncash/sse4a
[SVE256] Handle SSE insertions for EXTRQ/INSERTQ
2026-06-20 12:45:29 -07:00
Ryan Houdek 462c785418 Merge pull request #5573 from lioncash/cvtpi
[SVE256] Handle SSE insertions for CVTPI2PD
2026-06-20 12:44:21 -07:00
LC a1f90dd8d3 Passes: Trim unnecessary forward declarations
Less visual noise and lingering types left in the header.
2026-06-20 13:59:18 -04:00
LC 2a67261eac IntrusiveIRList: Amend signature for PostRA()
PostRA is a bool, not an unsigned value. We can also adjust SpillSlots()
to use uint32_t like its returned data member.
2026-06-20 13:39:35 -04:00
LC 1d3403fdc2 JIT: Amend op typos in implementations
Mostly benign, but ensures that they're correct in the event any of
their IR definitions change.
2026-06-20 13:16:12 -04:00
LC 53301b0f56 Arm64Emitter: Tidy up load/stores in Push/PopCalleeSavedRegisters
Same thing, just a little less verbose.
2026-06-20 12:32:23 -04:00
LC 8989ce1766 [SVE256] Handle SSE insertions for PCMPESTRM/PCMPISTRM ops
See #3799
2026-06-19 18:53:50 -04:00
LC faa121e9ef Vector: Move MOVQ2DQ over to Bind
Now all vector instruction implementations are consistently using Bind.
2026-06-19 17:45:41 -04:00
LC f48759e83a [SVE256] Handle SSE insertions for MOVQ2DQ
See #3799
2026-06-19 17:42:21 -04:00
LC 2eca733603 [SVE256] Handle SSE insertions for EXTRQ/INSERTQ
See #3799
2026-06-19 17:05:59 -04:00
LC 14b65cec43 [SVE256] Handle SSE insertions for CVTPI2PD
See #3799

CVTPI2PS is technically already handled, but we can add a test for it
as well, just to cover our bases.
2026-06-19 16:27:42 -04:00
Ryan Houdek f5477039fa Merge pull request #5571 from simon902/16bitleave
Fix incorrect RSP update for 16bit leave
2026-06-19 10:27:33 -07:00
Simon Scherer 9fa3221687 FEXCore: Fix incorrect RSP update for 16bit leave 2026-06-19 14:23:16 +02:00
Simon Scherer c8c63faf15 unittests/ASM: Adds unit test for 16bit leave 2026-06-19 14:16:15 +02:00
Ryan Houdek ee4794c99e Merge pull request #5570 from lioncash/pmadd
[SVE256] Handle SSE insertions for PMADDWD
2026-06-17 22:21:16 -07:00
Ryan Houdek 3a23bb4b73 Merge pull request #5569 from lioncash/cmp
[SVE256] Handle SSE insertions for CMPSD/CMPSS
2026-06-17 21:32:24 -07:00
Ryan Houdek 46ffb25f84 Merge pull request #5568 from lioncash/mov3
[SVE256] Handle SSE insertions for MOVSHDUP/MOVSLDUP
2026-06-17 21:21:50 -07:00
Ryan Houdek 069a4025c5 Merge pull request #5567 from lioncash/mov2
[SVE256] Handle SSE insertion for aligned/unaligned loads and non-temporal loads
2026-06-17 20:18:20 -07:00
LC edd044752d [SVE256] Handle SSE insertions for PMADDWD
See #3799
2026-06-17 22:15:51 -04:00
Ryan Houdek 3929d25dcc Merge pull request #5566 from lioncash/mov
[SVE256] Handle SSE insertions for MOVH(PD, PD, LPS) and MOVL(PD, PS, HPS)
2026-06-17 18:55:22 -07:00
LC 99baa4f3d9 [SVE256] Handle SSE insertions for CMPSD/CMPSS
See #3799
2026-06-17 21:05:31 -04:00
LC f64d4c571b [SVE256] Handle SSE insertions for MOVSHDUP/MOVSLDUP
See #3799
2026-06-17 20:38:58 -04:00
LC 9372fa169a [SVE256] Handle SSE insertions for MOVSD/MOVSS
See #3799
2026-06-17 20:09:17 -04:00
Ryan Houdek f98ac7f268 Merge pull request #5565 from lioncash/xor
[SVE256] Handle SSE insertions for XOR special case
2026-06-17 16:47:50 -07:00
LC daaa6ec129 [SVE256] Handle SSE insertions for aligned and unaligned moves
See #3799
2026-06-17 19:39:06 -04:00
LC 88afc22d5b [SVE256] Handle SSE insertions for MOVNTDQA
See #3799
2026-06-17 19:38:57 -04:00
Ryan Houdek 23099100b6 Merge pull request #5564 from lioncash/misc2
[SVE256] Handle SSE insertions for INSERTPS, PSIGN, PINSR, and shuffles
2026-06-17 16:24:26 -07:00
LC b7ea9e30df [SVE256] Handle SSE insertions for MOVH(PD, PD, LPS) and MOVL(PD, PS, HPS)
See #3799

Gets a few of the moves out of the way.
2026-06-17 17:39:13 -04:00
LC 844c3bb197 [SVE256] Handle SSE insertions for XOR special case
See #3799

Ensures that our special case maintains insertion behavior
2026-06-17 16:33:05 -04:00
LC 208c6d3eac [SVE256] Handle SSE insertions for vector unary ops
See #3799
2026-06-17 14:43:56 -04:00
LC 886a2e74ac [SVE256] Handle SSE insertions for pack ops 2026-06-17 13:36:18 -04:00
Ryan Houdek adad3c27dd Merge pull request #5563 from lioncash/shift
[SVE256] Handle SSE insertions for shifts
2026-06-17 08:02:41 -07:00
Ryan Houdek 470aeab215 Merge pull request #5562 from lioncash/misc 2026-06-17 05:54:59 -07:00
LC db1d90ec9d [SVE256] Handle SSE insertions for shuffles
See #3799
2026-06-17 08:44:45 -04:00
LC 124ce8420a [SVE256] Handle SSE insertions for PINSR(B,D,Q,W)
See #3799
2026-06-17 08:22:26 -04:00
LC 4433eaf242 [SVE256] Handle SSE insertions for INSERTPS
See #3799
2026-06-17 08:11:49 -04:00
LC 5ba070f600 [SVE256] Handle SSE insertions for PSIGN(B,D,W)
See #3799
2026-06-17 07:59:33 -04:00
LC 2d1a42aa00 [SVE256] Handle SSE insertions for shifts
See #3799
2026-06-17 06:06:08 -04:00
LC 74d9f5a3e2 [SVE256] Handle SSE insertions for MOVDDUP
See #3799
2026-06-17 05:13:08 -04:00
LC 634fbb5a73 [SVE256] Handle SSE insertions for Float->Int/Int->Float conversions
See #3799
2026-06-17 04:44:39 -04:00
LC e16948bf80 [SVE256] Handle SSE insertions for CVTPD2PS/CVTPS2PD
See #3799
2026-06-17 03:51:36 -04:00
LC 76c9833ee5 [SVE256] Handle SSE insertions for CMPPD/CMPPS
See #3799
2026-06-17 03:33:26 -04:00
Ryan Houdek 99662b70ff Merge pull request #5560 from lioncash/psad
[SVE256] Handle SSE insertions for more misc ops
2026-06-17 00:14:30 -07:00
LC fcde9eabbf [SVE256] Handle SSE insertions for PALIGNR
See #3799
2026-06-17 02:13:37 -04:00
LC 4ef15951a8 [SVE256] Handle SSE insertions for PACKSS/PACKUS ops
See #3799
2026-06-17 02:08:24 -04:00
LC 1bc51c2290 [SVE256] Handle SSE insertions for PMULUDQ
See #3799
2026-06-17 02:00:11 -04:00
LC be1025901a [SVE256] Handle SSE insertions for ADDSUBPD/ADDSUBPS
See #3799
2026-06-17 01:54:55 -04:00
Ryan Houdek 1a606de29f Merge pull request #5559 from lioncash/phmin
[SVE256] Handle SSE insertions for PHMINPOSUW, DPPD, and DPPS
2026-06-16 22:50:53 -07:00
LC 454c0b31cb [SVE256] Handle SSE insertions for MPSADBW
See #3799
2026-06-17 01:47:21 -04:00
Ryan Houdek 223e0f4e53 Merge pull request #5558 from lioncash/blend
[SVE256] Handle SSE insertions for blends
2026-06-16 22:34:32 -07:00
LC 3a84091945 [SVE256] Handle SSE insertions for DPPD/DPPS
See #3799
2026-06-17 01:32:36 -04:00
LC a0e8f1097f [SVE256] Handle SSE insertions for PHMINPOSUW
See #3799
2026-06-17 01:21:19 -04:00
Ryan Houdek 08ed4fb983 Merge pull request #5546 from neobrain/feature_woa_fexofflinecompiler
CodeCache: Support targeting WOW64/ARM64EC in FEXOfflineCompiler
2026-06-16 22:15:49 -07:00
Ryan Houdek 0b1f336e03 Merge pull request #5557 from lioncash/round 2026-06-16 22:09:20 -07:00
LC 0258fcb116 [SVE256] Handle SSE insertions for blends
See #3799
2026-06-17 01:05:46 -04:00
LC 3272aa3f08 [SVE256] Handle SSE insertions for ROUNDPD/ROUNDPS
See #3799
2026-06-17 00:42:20 -04:00
Ryan Houdek d9ea6651f8 Merge pull request #5556 from lioncash/madd
[SVE256] Handle more SSE insertions for some one-off instructions
2026-06-16 21:22:29 -07:00
LC 32b11603d8 [SVE256] Handle SSE insertions for PMOVSX/PMOVZX ops 2026-06-17 00:04:03 -04:00
LC 538fd2672d [SVE256] Handle SSE insertions for PSADBW
See #3799
2026-06-17 00:04:03 -04:00
LC 3e37724e3e [SVE256] Handle SSE insertions for PHADDSW
See #379
2026-06-17 00:04:03 -04:00
LC 1b249ba76b [SVE256] Handle SSE insertion for PHSUBD/PHSUBW/PHSUBSW 2026-06-17 00:04:03 -04:00
LC df1295fbd0 [SVE256] Handle SSE insertions for HSUBPD/HSUBPS
See #3799
2026-06-17 00:04:03 -04:00
LC dcd71fe126 [SVE256] Handle SSE insertions for PMULHW/PMULHRSW
See #3799
2026-06-17 00:04:00 -04:00
LC 610ee5db76 [SVE256] Handle SSE insertions for PMADDUBSW
See #3799
2026-06-16 23:02:25 -04:00
Ryan Houdek 9d5494d9f0 Merge pull request #5555 from lioncash/alu
[SVE256] Vector: Handle SSE insertion properly for various ALU operations
2026-06-16 19:56:38 -07:00
Ryan Houdek dd44bc8d00 Merge pull request #5554 from lioncash/vmov
OpcodeDispatcher: Eliminate redundant moves in VMOVHPOp
2026-06-16 19:47:01 -07:00
Ryan Houdek 110c7cb62b Merge pull request #5553 from lioncash/bind
OpcodeDispatcher: Make use of Bind consistently
2026-06-16 19:44:44 -07:00
Ryan Houdek 09aa5abbdf Merge pull request #5552 from lioncash/literal
OpcodeDispatcher: Move a few stray literal accesses to Literal()
2026-06-16 19:37:28 -07:00
LC 1d8b6df630 [SVE256] Vector: Handle SSE insertion properly for various ALU operations
See #3799 for the bulk of the issue explanation. Ensures that emulated
SSE operation on aarch64 don't end up obliterating the upper 128-bit
lane when SVE-256 is present (Adv. SIMD operations zero-extend)

Knocks out quite a few SSE instructions right off the jump.
2026-06-16 22:26:28 -04:00
LC 195058752e OpcodeDispatcher: Eliminate redundant moves in VMOVHPOp
Will make removing the TODO in LoadSource regarding partial loads a
little easier.
2026-06-16 17:47:01 -04:00
LC c62805e86d OpcodeDispatcher: Make use of Bind consistently
We had a few places that were using Bind, and a few other places
that were using specializations as a means to composing the instruction
tables. Instead, we can just use Bind consistently, which lets us tidy
up a bunch of the implementations (and gets rid of some unnecessary
codegen).
2026-06-16 15:07:59 -04:00
LC d15b175c33 OpcodeDispatcher: Move a few stray literal accesses to Literal()
Same core behavior, but ensures that the immediates are valid literals
when assertions are enabled.
2026-06-16 11:50:48 -04:00
Tony Wasserka 6cb73adfd5 CodeCache: Integrate FEXOfflineCompiler backend for WoA 2026-06-16 17:33:16 +02:00
Billy Laws 6d1cd67900 Windows: Implement WOW64/ARM64EC offline compiler backend
Implements an offline JIT compiler backend that compiles x86 code blocks
from a given code map into an ARM64 code cache on Windows. Separate
binaries are built for WOW64 (32-bit) and ARM64EC (64-bit) targets.

NOTE: This patch originally added a separate binary; instead the new
      functionality is added to the existing FEXOfflineCompiler and will
      be properly integrated in the next patches.
2026-06-16 17:30:32 +02:00
Ryan Houdek e12bd27106 Merge pull request #5536 from Sonicadvance1/105
LibraryForwarding: Implement support for CUDA
2026-06-12 15:14:38 -07:00
Ryan Houdek cb018257cf Merge pull request #5547 from neobrain/feature_fexofflinecompiler_process_all
FEXOfflineCompiler: Add "process-all" verb
2026-06-12 02:11:33 -07:00
LC e02953dc17 Merge pull request #5548 from Sonicadvance1/166
Frontend: Fix vsyscall page tracking.
2026-06-03 21:20:20 -04:00
Ryan Houdek ba56f8e0c5 unittests/FEXLinuxTests: Adds 64-bit vsyscall test
We never actually had a unittest to ensure these keep working, so add
one now.
2026-06-03 18:03:31 -07:00
Ryan Houdek ac12dd55c3 Frontend: Fix vsyscall page tracking.
Now that NX is tracked in the frontend, we need to ensure that adjusted
RIP pages are tracked correctly. Keep around both instruction stream
pointers, validate the the original RIP is executable, and read from the
adjusted RIP as appropriate.

Fixes #5544
2026-06-03 18:03:31 -07:00
Ryan Houdek 24647820d7 Linux: Ensure 64-bit vsyscall page is tracked
It's a purely virtual page even on x86, so we need to manually add it to
tracking.
2026-06-03 18:03:31 -07:00
Ryan Houdek 7e8aa711ef Linux: Pass gettimeofday through glibc
This ensures that it hits the VDSO path if possible.
2026-06-03 17:45:41 -07:00
Tony Wasserka 329e561eff FEXOfflineCompiler: Add "process-all" verb
This operation will take care of any pending code cache operations:
* import new code maps from `CACHE_DIR/codemap/new` and process them to `CACHE_DIR/codemap/ready`
* generate caches for updated code maps with new blocks
* ensure caches already exist for all other code maps (and generate them if needed)
2026-06-03 18:42:37 +02:00
LC d848cbbc0f Merge pull request #5543 from Sonicadvance1/165
FEXCore: Ensure LOCK prefix instructions are handled correctly
2026-06-02 23:39:02 -04:00
Ryan Houdek cd46e43c20 unittests/FEXLinuxTests: Add a LOCK prefix test
Ensures we handle lock prefixing correctly.
2026-06-02 19:41:04 -07:00
Ryan Houdek 98674c1cc8 unittests/FEXLinuxTests: Support redirecting RIP entirely 2026-06-02 19:41:04 -07:00
Ryan Houdek 8bfae631b1 Frontend: Support raising unimplemented instruction on LOCK failure
When the LOCK prefix is on an instruction that doesn't support LOCK then
it raises a SIGILL. Make sure to pass that up.

Additionally if the instruction does support lock prefix, has a lock
prefix, but the destination is not memory then that is also invalid.
2026-06-02 19:41:03 -07:00
Ryan Houdek d00c7cf3a3 Frontend: Support passing the decode failure type through decoding
Only used for invalid inst currently
2026-06-02 19:41:03 -07:00
Ryan Houdek b45665fee4 OpcodeDispatcher: Support instruction type of unimplement operation 2026-06-02 19:41:03 -07:00
Ryan Houdek 1b58664541 X86Tables: Describe instructions that support LOCK prefix 2026-06-02 19:41:02 -07:00
LC ed6a178ae3 Merge pull request #5542 from Sonicadvance1/164
OpcodeDispatcher: Fixes 64-bit LODs with address size override
2026-06-02 22:40:53 -04:00
Ryan Houdek e925ca509d unittests/ASM: Adds lods tests with address size override
If the override isn't handled then it'll fall back to either using the
wrong address and getting the wrong data or crashing depending on how it
is broken.
2026-06-02 18:21:00 -07:00
Ryan Houdek ca310cf815 OpcodeDispatcher: Fixes 64-bit LODs with address size override
Fairly trivial but just need to be careful with address size wraparound
as usual.
2026-06-02 18:21:00 -07:00
Ryan Houdek 6fa27aac42 Thunks: Implement support for CUDA
This is enough to get less complex cuda applications running, and is a
good starting spot to slowly finish off the remaining implementation.

Some information:
- 429 functions in total
- 153 only compiled for 64-bit (35.6%)
- 7 functions disabled entirely (1.6%)

The main thing /not/ working with this initial implementation is .cu
files compiled in to an ELF using the static cuda runtime. This is due
to the `cuGetExportTable` function being stubbed out and the static cuda
RT requires at least two interfaces from that function before it
continues.

This function isn't publicly documented by NVIDIA but has been publicly
reverse engineered to be fairly trivial. It's just a jump table with the
first element being the size of the table in bytes.

That will be the next step of the implementation.
2026-06-01 16:32:45 -07:00
Tony Wasserka a5c3fc4751 Merge pull request #5514 from peppergrayxyz/snd_htimestamp_t
LibraryForwarding: Add annotation for snd_htimestamp_t
2026-06-01 12:53:42 +02:00
LC 65b05fa8c1 Merge pull request #5540 from Sonicadvance1/163
arm64ec: Single instruction optimization in EC map lookup
2026-06-01 05:41:34 -04:00
Pepper Gray 3ee556d858 add template for snd_htimestamp_t
building on musl fails with:
`error: Unsupported parameter type 'snd_htimestamp_t *' (aka 'timespec *')`

due to empty padding members in `alltypes.h`:
```
STRUCT timespec {
  time_t tv_sec;
  int :8*(sizeof(time_t)-sizeof(long))*(__BYTE_ORDER==4321);
  long tv_nsec;
  int :8*(sizeof(time_t)-sizeof(long))*(__BYTE_ORDER!=4321);
};
```
add (missing) annotation to libasound interface

fix: #5513
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-31 17:38:25 +02:00
Ryan Houdek fd3546a999 arm64ec: Single instruction optimization in EC map lookup
We can merge the lsr+and in to a single lsr by 18 and then use ldr with
LSL of 3 to accomplish the same result. Modern Cortex doesn't even
generate an additional integer pipeline uop for this ldr+lsl instruction
anymore.
2026-05-29 19:17:13 -07:00
Ryan Houdek f5fafa5b96 Merge pull request #5492 from FrontMage/codex/int29-failfast-probe
Windows: Trace interrupt translation and prototype INT 0x29 fail-fast mapping
2026-05-29 16:59:16 -07:00
Ryan Houdek 7ff2069e60 Merge pull request #5539 from Sonicadvance1/162
arm64ec: Fixes some FEX allocations that were missing TOP_DOWN
2026-05-29 16:49:14 -07:00
Ryan Houdek 154ff43d7f arm64ec: Fixes some FEX allocations that were missing TOP_DOWN
We were accidentally allocating some things without this flag and it was
causing us to dump memory in to the lower 32-bits on arm64ec.

This was causing the game
[Below](https://store.steampowered.com/app/250680/BELOW/) to run out of
memory to allocate for its LUA JIT and causes it to crash.
2026-05-29 16:17:04 -07:00
Ryan Houdek b754fe4810 External/rpmalloc: update 2026-05-29 16:16:46 -07:00
Tony Wasserka a1071ec01a Merge pull request #5538 from peppergrayxyz/syscall_headers
LinuxSyscalls: add missing thread header
2026-05-29 14:35:43 +02:00
Pepper Gray 92dce9a2ea add missing header <thread>
building using musl fails due to missing defintions:

```
Source/Tools/LinuxEmulation/LinuxSyscalls/Syscalls.cpp:913:23: error: no member named 'sleep_for' in namespace 'std::this_thread'
  913 |     std::this_thread::sleep_for(std::chrono::milliseconds {10});
      |                       ^~~~~~~~~
1 error generated.
```

Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-29 13:19:24 +02:00
LC 5fd917ec2f Merge pull request #5537 from Sonicadvance1/161
Format: Fix missed clang-format
2026-05-29 00:52:34 -04:00
Ryan Houdek d7cda23b25 Format: Fix missed clang-format
Minor clang-format version differences causing differing behaviour
again.
2026-05-28 12:12:24 -07:00
Ryan Houdek cae5da5777 Merge pull request #5530 from neobrain/fix_ccache_time_macros
Build: Enable ccache sloppiness for time macros
2026-05-25 10:55:27 -07:00
Ryan Houdek 97f1f47fa5 Merge pull request #5531 from neobrain/refactor_drop_cmake_settings
Drop unused CMakeSettings.json
2026-05-25 10:53:44 -07:00
Tony Wasserka 53c269ee25 Merge pull request #5516 from peppergrayxyz/thunk_rootfs
set sysroot to X86_DEV_ROOTFS for guest toolchain
2026-05-25 18:08:45 +02:00
Pepper Gray 1420d3cc10 set sysroot to X86_DEV_ROOTFS for guest toolchain
**Faulty Behaviour:**
When not building on Ubuntu `unittests/ThunkLibs` fails to find
c++ header and fails:

```
gen_input.cpp:2:10: fatal error: 'cstddef' file not found
    2 | #include <cstddef>
      |          ^~~~~~~~~
1 error generated.
```

This also includes building with nix-shell:
```
nix-shell ../Data/nix/LibraryForwarding/shell.nix --run "cmake --build . --target thunkgen_tests"
```

**Root Cause**:
`X86_DEV_ROOTFS` is not used, but  header paths are hard
coded to Ubuntu's multilib layout (`/usr/i686-linux-gnu/include/`,
`/usr/x86_64-linux-gnu/include/`). It works on Ubuntu but the
mechanism to use to another directory is broken and the include
paths always point to the host.

**Solution**:
set `--sysroot ${X86_DEV_ROOTFS}` for guest builds and remove
hard coded include paths.

Fix: #5515
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-25 17:18:47 +02:00
Tony Wasserka ce401b5ca1 Build: Enable ccache sloppiness for time macros
Ccache won't attempt to cache files that use __DATE__/__TIME__ since their
contents would always be out of date. Setting sloppiness disables this
behavior, which works for us since we don't use __DATE__/__TIME__ for anything
that needs accurate values.
2026-05-25 12:03:33 +02:00
Tony Wasserka cd412bd0f5 Drop unused CMakeSettings.json
Visual Studio (but not VS Code) used to read this file, but this convention is
discouraged nowadays.
2026-05-25 11:27:29 +02:00
FrontMage f33f88072e Windows: Fix INT exception formatting 2026-05-24 11:42:51 +08:00
LC 1240a00fa5 Merge pull request #5522 from Sonicadvance1/160
New CPL0 instructions from #5510 but with unittests
2026-05-23 22:57:53 -04:00
LC 83f325de0d Merge pull request #5519 from Sonicadvance1/157
meta: Add CONTRIBUTING.md
2026-05-23 22:55:26 -04:00
LC 0b871bf54e Merge pull request #5518 from Sonicadvance1/156
code-format-helper: More dependabot changes
2026-05-23 22:54:36 -04:00
LC 07f7aa3c8f Merge pull request #5521 from Sonicadvance1/159
HostFeatures: Don't capture CTR/MIDR under simulator
2026-05-22 22:10:38 -04:00
Ryan Houdek 7208bc6cdd unittests: Extend unittests for CPL0 instructions 2026-05-22 15:33:42 -07:00
Ryan Houdek fef5a98602 Fix build failure. 2026-05-22 15:33:41 -07:00
Daniel Lu 5cce65cdfa OpcodeDispatcher: Decode INVD and WBINVD through privileged op handling 2026-05-22 15:25:22 -07:00
Ryan Houdek 268081e5d0 Merge pull request #5520 from Sonicadvance1/158
Cherry-pick #5508 with instcountci changes
2026-05-22 14:46:20 -07:00
Ryan Houdek f5f179117e HostFeatures: Don't capture CTR/MIDR under simulator
If the simulator was selected, we would still capture CTR and MIDR on
the host AArch64 system. Potentially modifying codegen in unexpected
ways.

Ensure we return 0/0 like under x86 with simulator to simulate
"unknown".
2026-05-22 14:04:52 -07:00
Ryan Houdek 8c85096f98 meta: Add CONTRIBUTING.md 2026-05-22 14:03:14 -07:00
Ryan Houdek c5e7675c4b InstcountCI: Update 2026-05-22 13:58:10 -07:00
Daniel Lu 03009912ac JIT: Avoid clobbering guest rdx while raising generated faults 2026-05-22 13:56:40 -07:00
Ryan Houdek e4a1138291 code-format-helper: More dependabot changes 2026-05-22 13:44:22 -07:00
FrontMage a5bc54d2d5 Windows: Preserve INT 0x2D exception parameter source 2026-05-22 12:56:44 +08:00
LC df73e84725 Merge pull request #5507 from Sonicadvance1/155
CodeCache: Resolve comment https://github.com/FEX-Emu/FEX/pull/5441#discussion_r3255384588
2026-05-21 22:39:14 -04:00
Ryan Houdek 5bf07c2e77 CodeCache: Resolve comment https://github.com/FEX-Emu/FEX/pull/5441#discussion_r3255384588 2026-05-21 18:04:29 -07:00
Ryan Houdek bb0d142a65 Merge pull request #5441 from neobrain/feature_mmap_code_cache
CodeCache: Implement lazy code loading
2026-05-20 17:19:11 -07:00
LC c98cef0da1 Merge pull request #5506 from Sonicadvance1/154
HostFeatures: Only enable `dc zva` optimization on Ampere CPUs
2026-05-19 22:29:11 -04:00
Ryan Houdek 1d9c52be02 InstcountCI: Update 2026-05-19 17:38:10 -07:00
Ryan Houdek a6c9df1a64 HostFeatures: Only enable dc zva optimization on Ampere CPUs
This optimization was only written for Ampere1A where it showed a
noticable performance improvement in #5321. On Cortex it didn't matter.
Turns out this actually hits a bad case on Oryon CPUs where `dc zva` is
actually dramatically slower in the face of memory barriers and
overlapping stores in flight.

So now just detect Ampere and only use the optimization on that hardware
and send everyone else down the regular path.

microbench A1A:
```
Cycle counter frequency: 1000000000
Cycle counter granularity: 20
ns in cycle: 1
suite: memory
Test, Total Cycles, Iterations, Cycles Average, Iter Time Average, iterations/Second
dc zva - vzeroupper, 723390880, 363855872, 1.99, 1.99 nanosecond, 502986534.75
dc zva - vzeroall, 571708060, 161742848, 3.53, 3.53 nanosecond, 282911610.52
dc zva (stp emu) - vzeroupper, 541543980, 107872256, 5.02, 5.02 nanosecond, 199193897.42
dc zva (stp emu) - vzeroall, 722548940, 71958528, 10.04, 10.04 nanosecond, 99589832.63
```

microbench X2E:
```
Cycle counter frequency: 19200000
Cycle counter granularity: 1
ns in cycle: 52.083333333333336
suite: memory
Test, Total Cycles, Iterations, Cycles Average, Iter Time Average, iterations/Second
dc zva - memset 0, 12065162, 49, 246227.80, 12.82 millisecond, 77.98
dc zva - vzeroupper, 12098598, 4325376, 2.80, 145.68 nanosecond, 6864201.89
dc zva - vzeroall, 12031459, 4325376, 2.78, 144.87 nanosecond, 6902506.11
dc zva (stp emu) - vzeroupper, 13899441, 363855872, 0.04, 1.99 nanosecond, 502612496.60
dc zva (stp emu) - vzeroall, 12389283, 161742848, 0.08, 3.99 nanosecond, 250657175.37
```
2026-05-19 17:28:39 -07:00
Ryan Houdek f66368b191 Merge pull request #5505 from fixedcat/main
Arm64: Fix byte-size handling in unaligned STLXR emulation
2026-05-19 12:28:23 -07:00
LC e4daea406e Merge pull request #5503 from Sonicadvance1/154
FEXCore: Allow InterruptFaultPage to be significantly further away
2026-05-19 12:22:18 -04:00
fixedcat 6216f22cb9 Arm64: Fix byte-size handling in unaligned STLXR emulation 2026-05-19 18:53:44 +08:00
LC d4c80d9094 Merge pull request #5422 from Sonicadvance1/138
FEXCore: Add support for developer single stepping, read/write watching.
2026-05-19 01:11:04 -04:00
Ryan Houdek 27324ded87 Merge pull request #5499 from bylaws/winstuff
Windows additions for code caching
2026-05-18 16:05:59 -07:00
Ryan Houdek 7d1c625e32 FEXCore: Allow InterruptFaultPage to be significantly further away
We are actually quite close to a single page of CPU state per thread and
any additional changes are likely to cause it to overflow which would
hit these asserts. As we saw with the libc++ implementation of mutexes,
just one object type changing size could push it over the edge.

Future proof this by ensuring we can have this be sixteen pages per
thread before needing to hit more complex implementations. Which I don't
see us getting that large of CPU context tracking.
2026-05-18 15:57:25 -07:00
Ryan Houdek b4fe65f2c0 Merge pull request #5496 from neobrain/fix_libfwd_findpkg
Library Forwarding: Various build system improvements
2026-05-18 14:55:53 -07:00
Ryan Houdek f5efdac2e5 Merge pull request #5423 from bylaws/depenencey
WOW64: Support disabling DEP
2026-05-18 14:39:44 -07:00
Billy Laws 23de875516 Windows: Add NtUnmapViewOfSection prototype 2026-05-17 23:06:02 +01:00
Billy Laws 82030b8286 Windows: Declare winternl relocation APIs 2026-05-17 23:03:11 +01:00
Ryan Houdek af4da43bb8 Merge pull request #5494 from neobrain/fix_determine_va
Allocator: Fix and optimize VA range detection
2026-05-17 14:51:59 -07:00
Ryan Houdek 33f3b8659c Merge pull request #5495 from neobrain/fix_portable_config
Config: Use more sensible default for portable config location
2026-05-17 14:51:04 -07:00
Ryan Houdek ed724a61a7 Merge pull request #5498 from bylaws/evmd
ImageTracker: Support using image IDs as an extended volatile metadata key
2026-05-17 14:49:16 -07:00
Billy Laws 053bd74aa3 Windows/Common: Add ScopedHandle::reset() 2026-05-17 22:35:06 +01:00
Billy Laws 3c0410e59f ImageTracker: Support using image IDs as an extended volatile metadata key 2026-05-17 19:34:41 +01:00
Tony Wasserka e621f6c753 LibraryForwarding/Build: Explicitly look up LLVM headers
This could previously set up incorrect header paths when clang and LLVM were
installed in different directories (such as when using nix).
2026-05-15 16:07:13 +02:00
Tony Wasserka b05f000f42 LibraryForwarding/Build: Try harder to properly discover header locations 2026-05-15 16:07:13 +02:00
Tony Wasserka e60bfc6d23 LibraryForwarding/Build: Allow specifying system header location externally 2026-05-15 16:07:09 +02:00
Tony Wasserka 9a3d3201f9 Config: Use more sensible default for portable config location 2026-05-15 15:59:22 +02:00
Tony Wasserka abf9724424 Allocator: Fix and optimize VA range detection 2026-05-15 15:47:38 +02:00
LC ab9a8c62ab Merge pull request #5493 from Sonicadvance1/153
FEXGetConfig: Even more correctness changes for X2E
2026-05-14 09:39:31 -04:00
FrontMage e2fe936152 Windows: Handle INT 0x29 as fast-fail 2026-05-14 12:05:08 +08:00
Ryan Houdek 1d71650379 FEXGetConfig: Even more correctness changes for X2E
Some of the information was incorrect, so make sure it shows the
hardware correctly.
2026-05-13 20:00:57 -07:00
Tony Wasserka a040740974 CodeCache: Ensure atomicity of code page finalization 2026-05-13 22:53:09 +02:00
Tony Wasserka 5be0dc9fc5 CodeCache: Implement lazy code loading 2026-05-13 21:25:47 +02:00
Tony Wasserka 52ad434d24 LinuxSyscalls: Defer MappedResource deletion until after code invalidation
This ensures that any code buffer memory owned by the MappedResource is
invalidated before being deallocated.
2026-05-13 21:25:47 +02:00
Tony Wasserka d69d111bb6 CodeCache: Align code section within cache files
This allows mapping the code directly into memory for execution.
2026-05-13 21:25:47 +02:00
LC 50f4494875 Merge pull request #5490 from Sonicadvance1/152
FEXGetConfig: Showcase RMW versus loadstore atomic differences
2026-05-12 16:32:04 -04:00
Ryan Houdek 8a4982383a FEXGetConfig: Showcase RMW versus loadstore atomic differences
This differs on X2E, so it's good to showcase it.
2026-05-12 12:38:44 -07:00
Ryan Houdek 0d72890482 Merge pull request #5432 from pmatos/f64-fprem
JIT-inline FPREM/FPREM1 for reduced precision x87 path
2026-05-11 14:31:31 -07:00
Ryan Houdek 9be7d6d112 Merge pull request #5489 from neobrain/feature_better_fexbash
FEXBash: Drop implicit -c and add colored PS1
2026-05-11 12:22:00 -07:00
Tony Wasserka 3c4121ba07 FEXBash: Use a shiny rainbow for PS1 2026-05-11 20:30:21 +02:00
Tony Wasserka b6e44b04d6 FEXBash: Clean up path handling 2026-05-11 20:30:21 +02:00
Tony Wasserka 02c11afbc6 FEXBash: Don't imply "-c" to behave more closely like bash
Implicitly adding "-c" breaks argument passing for scripts. For example, the
command "FEXBash ./steam.sh -silent" will process steam.sh but the script
wouldn't see the "-silent" argument previously.
2026-05-11 19:49:15 +02:00
Paulo Matos 2b8f5b57eb instcountci: JIT-inline FPREM/FPREM1 for reduced precision x87 path 2026-05-11 10:27:08 +02:00
Paulo Matos 0519c9467c asm_tests: JIT-inline FPREM/FPREM1 for reduced precision x87 path 2026-05-11 10:27:08 +02:00
Paulo Matos 84d968c7d2 JIT-inline FPREM/FPREM1 for reduced precision x87 path 2026-05-11 10:27:08 +02:00
Ryan Houdek a04b0241c2 Docs: Update for release FEX-2605 2026-05-08 19:28:30 -07:00
Ryan Houdek 670fd19d33 Merge pull request #5486 from Sonicadvance1/151
Allocator: Mark large unmapped regions as DONTDUMP
2026-05-08 19:28:13 -07:00
Ryan Houdek a66544f3f4 Allocator: Mark large unmapped regions as DONTDUMP
coredump applications aren't smart enough to only dump resident pages,
so explicitly mark our 128TB and other mapped VA ranges as DONTDUMP.

This will speed up coredumps.
2026-05-08 16:27:04 -07:00
Ryan Houdek 1bfb3aefcc Merge pull request #5485 from Sonicadvance1/150
Windows: Setup `tu_override_uncached_as_cache_coherent` inside of dlls
2026-05-08 15:58:43 -07:00
Ryan Houdek b7bfbc3fcd Windows: Setup tu_override_uncached_as_cache_coherent inside of dlls
To not have this environment variable accidently be enabled on arm64
native Wine games, we need to set it from inside of FEX.

Requires the FEX dlls to set them directly rather than launch scripts.
2026-05-08 13:39:51 -07:00
Ryan Houdek e517f3259c Merge pull request #5484 from neobrain/fix_code_cache_no_guest_wrappers
CodeCache: Fix crash when guest library wrappers aren't installed
2026-05-07 10:59:32 -07:00
Tony Wasserka 8afda92a64 CodeCache: Fix crash when guest library wrappers aren't installed 2026-05-07 18:56:35 +02:00
Ryan Houdek ed216c8d4d Merge pull request #5449 from neobrain/opt_code_cache_writing
CodeCache: Slightly optimize cache file writing
2026-05-06 18:06:51 -07:00
Ryan Houdek 7506cb4ea1 Merge pull request #5483 from neobrain/fix_guest_wrapper_code_cache
CodeCache: Delay cache loading for guest library wrappers until after LoadLib
2026-05-06 18:04:55 -07:00
Tony Wasserka 60bc5944db CodeCache: Slightly optimize cache file writing
ftruncate only requires one call (and one extra seek) instead up to 64 manual
zero writes.
2026-05-06 17:08:15 +02:00
Tony Wasserka 8f0572283a Windows/CRT: Implement ftruncate and _chsize 2026-05-06 17:08:04 +02:00
Tony Wasserka b13b46eefe CodeCache: Delay cache loading for guest library wrappers until after LoadLib
These libraries need to be initialized before relocating their caches,
since the guest function hashes won't be registered before.
2026-05-05 16:29:39 +02:00
LC 4db2a98d7f Merge pull request #5481 from Sonicadvance1/148
win32: Query DCZID_EL0 so clzero works
2026-05-05 08:42:31 -04:00
LC 05ebb07753 Merge pull request #5482 from Sonicadvance1/149
OpcodeDispatcher: Optimize MMX pshufw
2026-05-05 08:41:22 -04:00
Ryan Houdek 694e68b838 Merge pull request #5468 from peppergrayxyz/proc_self_stat
read /proc/self/stat using %lu
2026-05-04 20:21:16 -07:00
Pepper Gray 78320e1433 read /proc/self/stat using %lu
building on clang/musl causes warnings:

```
FEX/Source/Tools/FEXInterpreter/ELFCodeLoader.h:782:29: warning: format specifies type 'unsigned long long *' but the argument has type 'uint64_t *' (aka 'unsigned long *') [-Wformat]
  776 |                             "%llu %llu %llu %*u %*u "   // 26 to 30
      |                              ~~~~
      |                              %lu
  777 |                             "%*u %*u %*u %*u %*u "      // 31 to 35
  778 |                             "%*u %*u %*d %*d %*u "      // 36 to 40
  779 |                             "%*u %*u %*u %*d %llu "     // 40 to 45
  780 |                             "%llu %llu %llu %llu %llu " // 46 to 50
  781 |                             "%llu",                     // 51
  782 |                             &map.start_code, &map.end_code, &map.start_stack, &map.start_data, &map.end_data, &map.start_brk,
      |                             ^~~~~~~~~~~~~~~
```

according to the [man page](https://man7.org/linux/man-pages/man5/proc_pid_stat.5.html)
`/proc/self/stat` uses `%lu`:

read the values as unsigned long (%lu) and then write them to
prctl_mm_map (platform specific format).

fixes: #5467
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-05 04:24:31 +02:00
Ryan Houdek 082e7b2695 InstcountCI: Update 2026-05-04 17:52:54 -07:00
Ryan Houdek cdbffb80c7 unittests: Adds a full coverage pshufw test 2026-05-04 17:52:53 -07:00
Ryan Houdek cb6c8cce55 OpcodeDispatcher: Optimize MMX pshufw
Found through writing a shuffle solver rather than an LLM.

Fixes #3785
2026-05-04 17:52:53 -07:00
Ryan Houdek 47e173e549 Merge pull request #5472 from peppergrayxyz/format
use portable format specifiers
2026-05-04 14:18:28 -07:00
Ryan Houdek 162bd4be97 win32: Query DCZID_EL0 so clzero works
EL0 registers are readable without going through the registry, but we
were failing to populate this register, which was causing clzero to not
be supported.
2026-05-04 12:42:18 -07:00
Ryan Houdek f0764aeafe Merge pull request #5480 from neobrain/refactor_musl_sigmask
SignalDelegator: Simplify support for musl's sigset_t
2026-05-04 10:39:42 -07:00
Ryan Houdek 1efed71696 Merge pull request #5479 from neobrain/refactor_drop_compile_service
FEXCore: Drop unused CompileService
2026-05-04 10:36:44 -07:00
Tony Wasserka 06d77c1c19 SignalDelegator: Simplify support for musl's sigset_t 2026-05-04 16:28:47 +02:00
Tony Wasserka 942d0c631a FEXCore: Drop unused CompileService 2026-05-04 15:52:34 +02:00
Pepper Gray abae5dd93b use portable format specifiers
building with clang/musl causes these warnings:

```
FEX/Source/Tools/FEXServer/ProcessPipe.cpp:100:96: warning: format specifies type 'ssize_t' (aka 'long') but the argument has type 'rlim_t' (aka 'unsigned long long') [-Wformat]
```

- cast platform specific MaxFDs members to uintmax_t and print as PRIuMAX
- use %zu for GetNumFilesOpen (size_t)

fix: #5471
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-04 11:20:55 +02:00
Ryan Houdek 93015a0266 Merge pull request #5458 from peppergrayxyz/largefile64
make 64bit symbols visible to enhance portability (musl)
2026-05-03 23:36:59 -07:00
Ryan Houdek d238db69d3 Merge pull request #5477 from peppergrayxyz/unistd
include missing header unistd.h
2026-05-03 15:18:41 -07:00
Ryan Houdek e0ead236b6 Merge pull request #5473 from peppergrayxyz/ObjectCacheRefCounter
remove dead code (ObjectCacheRefCounter)
2026-05-03 15:16:57 -07:00
Pepper Gray c548262664 include missing header unistd.h
build on clang/musl fails with:

```
FEX/unittests/APITests/Allocator.cpp:17:5: error: use of undeclared identifier 'close'
FEX/unittests/APITests/Allocator.cpp:23:5: error: use of undeclared identifier 'lseek'; did you mean 'fseek'?
FEX/unittests/APITests/Allocator.cpp:23:11: error: cannot initialize a parameter of type 'FILE *' (aka 'struct _IO_FILE *') with an lvalue of type 'int'
FEX/unittests/APITests/Allocator.cpp:24:5: error: use of undeclared identifier 'write'; did you mean '_IO_cookie_io_functions_t::write'?
FEX/unittests/APITests/Allocator.cpp:24:5: error: invalid use of non-static data member 'write'
```

include header to provide defintions

fix: #5476
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-03 15:49:40 +02:00
Pepper Gray bfc51e577f remove dead code (ObjectCacheRefCounter)
building using libc++ failes due to shared_mutex ObjectCacheRefCounter
inflating InternalThreadState beyond FEX_PAGE_SIZE, thus triggering
the static assert:

```
FEXCore/Debug/InternalThreadState.h:133:15: error: static assertion failed
FEX/FEXCore/include/FEXCore/Debug/InternalThreadState.h:133:145: note: expression evaluates to '7680 < 4096'
FEX/FEXCore/include/FEXCore/Debug/InternalThreadState.h:136:58: note: expression evaluates to '12288 == 8192'
```

remove `ObjectCacheRefCounter` as it is not used anywhere.

fix: #5456
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-03 15:13:47 +02:00
Ryan Houdek 85773995e1 Merge pull request #5470 from peppergrayxyz/header_redirect
fix include redirect for <poll.h> and <signal.h>
2026-05-03 04:41:18 -07:00
Pepper Gray 00b4777290 fix include redirect for <poll.h> and <signal.h>
building on musl/clang causes redirecting incorrect #includes warnings:

```
warning: redirecting incorrect #include <sys/poll.h> to <poll.h> [-W#warnings]
warning: redirecting incorrect #include <sys/signal.h> to <signal.h> [-W#warnings]
```

include headers instead of sys/headers.

fix: #5469
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-03 13:22:32 +02:00
Ryan Houdek 197e6de194 Merge pull request #5462 from peppergrayxyz/uc_sigmask
determine sigset_t fieldname to enhance portability (musl)
2026-05-03 03:48:53 -07:00
Ryan Houdek 9908ea4c2f Merge pull request #5466 from peppergrayxyz/libgen_h
include <libgen.h> for basename
2026-05-03 03:46:43 -07:00
Ryan Houdek c402b15bd3 Merge pull request #5464 from peppergrayxyz/tgkill
add header and classpath for tgkill
2026-05-03 03:46:07 -07:00
Ryan Houdek a3f3118ac5 Merge pull request #5460 from peppergrayxyz/sigset_t
use <signal.h> instead of glibc header to enhance portability (musl)
2026-05-03 03:18:20 -07:00
Pepper Gray b1aab0e498 determine sigset_t fieldname to enhance portability (musl)
musl build fails due to access to internal glibc member:

```
FEX/Source/Tools/LinuxEmulation/LinuxSyscalls/SignalDelegator.cpp:655:39: error: no member named '__val' in '__sigset_t'
  655 |       .SigMask = _context->uc_sigmask.__val[0],
      |                  ~~~~~~~~~~~~~~~~~~~~ ^
1 error generated.
```

add check to determine private glibc or musl member name or throw an error.

fixes: #5461
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-03 12:17:09 +02:00
Pepper Gray 927c0ce54a include <libgen.h> for basename
building on clang/musl build fails due to missing symbol:

```
/run/host/home/pepper/dev/ghrepos/FEX/Source/Tools/FEXOfflineCompiler/Main.cpp:321:35: error: use of undeclared identifier 'basename'
  321 |   auto CommandName = std::string {basename(argv[0])} + " " + (argc > 1 ? argv[1] : "");
      |                                   ^~~~~~~~
/run/host/home/pepper/dev/ghrepos/FEX/Source/Tools/FEXOfflineCompiler/Main.cpp:327:43: error: use of undeclared identifier 'basename'
  327 |     fmt::print("Usage: {} <command>\n\n", basename(argv[0]));
      |                                           ^~~~~~~~
```

include missing header.

fixes: #5465
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-03 12:10:07 +02:00
Pepper Gray 57d9dc037d add header and classpath for tgkill
building on clang/musl fails with:

```
FEX/Source/Tools/LinuxEmulation/LinuxSyscalls/GdbServer.cpp:1174:5: error: use of undeclared identifier 'tgkill'
 1174 |     tgkill(::getpid(), ::getpid(), SIGKILL);
      |     ^~~~~~
1 error generated.
```

include and use `FHU::Syscalls::tgkill`.

fixes: #5463
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-03 11:50:23 +02:00
Pepper Gray e3cfe28848 use <signal.h> instead of glibc header to enhance portability (musl)
building using musl fails with:

```
FEX/Source/Tools/LinuxEmulation/LinuxSyscalls/ThreadManager.h:33:10: fatal error: 'bits/types/sigset_t.h' file not found
   33 | #include <bits/types/sigset_t.h>
      |          ^~~~~~~~~~~~~~~~~~~~~~~
1 error generated.
```

`<bits/types/sigset_t.h>` is a glibc internal header, use <signal.h> instead.

fixes: #5459
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-03 11:03:38 +02:00
Pepper Gray ac91f583b8 make 64bit symbols visible to enhance portability (musl)
building using clang/musl fails with:

```
FEX/Source/Tools/LinuxEmulation/LinuxSyscalls/x32/Types.h:548:5: error: member access into incomplete type 'const struct statfs64'
  548 |     COPY(f_bsize);
      |     ^
```

add `_LARGEFILE64_SOURCE` to define large-file feature macros

fix: #5457
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-03 10:52:09 +02:00
Ryan Houdek 215658bf29 Merge pull request #5455 from peppergrayxyz/sys_prctl
use <sys/prctl.h> to enhance portability (clang)
2026-05-02 15:11:33 -07:00
Pepper Gray 92b1a6ea8a use <sys/prctl.h> to enhance portability (clang)
using <sys/prctl.h> and <linux/prctl.h> simultaneously causes clang to
fail:

```
In file included from FEX/FEXCore/Source/Utils/AllocatorHooks.cpp:6:
/usr/include/sys/prctl.h:88:8: error: redefinition of 'prctl_mm_map'
   88 | struct prctl_mm_map {
      |        ^
/usr/include/linux/prctl.h:134:8: note: previous definition is here
  134 | struct prctl_mm_map {
      |        ^
1 error generated.
```

prefer <sys/prctl.h> and do not include <linux/prctl.h>

fix: #5454
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-02 15:36:40 +02:00
Ryan Houdek 8ab00758be Merge pull request #5448 from bylaws/claudefix6
X87: Fix FXTRACT with Inf and NaN inputs
2026-04-30 14:42:08 -07:00
Ryan Houdek e91bda7765 Merge pull request #5447 from bylaws/claudefix5
SoftFloat: Fix FSCALE(0, +Inf) to raise IE and return a quiet NaN
2026-04-30 14:41:24 -07:00
Ryan Houdek 9db211ac97 Merge pull request #5446 from bylaws/claudefix4
OpcodeDispatcher: Fix FIST 16-bit IE for denormal inputs
2026-04-30 14:40:41 -07:00
Ryan Houdek b1381fd3b7 Merge pull request #5444 from bylaws/claudefix2
VectorOps: Fix MAXPS/MAXPD NaN and signed-zero tie selection
2026-04-30 14:39:56 -07:00
Ryan Houdek e5f6a7d85e Merge pull request #5450 from neobrain/fix_code_cache_portable
CodeCache: Make FEXServer call FEXOfflineCompiler by absolute path
2026-04-30 14:17:16 -07:00
Ryan Houdek feae76fc4f Merge pull request #5451 from Sonicadvance1/146
arm64ec: Fixes crash in many games with SDL+Dualsense
2026-04-30 14:16:06 -07:00
Ryan Houdek 015f3cffb9 arm64ec: Fixes crash in many games with SDL+Dualsense
We were pointing to an incorrect function pointer and exploding when a
pending suspend doorbell had occured.
2026-04-30 12:59:53 -07:00
Tony Wasserka 3e5c17ae80 CodeCache: Make FEXServer call FEXOfflineCompiler by absolute path
In portable mode, FEXOfflineCompiler may not be in PATH (and if it is, it's
most likely not a compatible version). Instead, use the executable next to
the FEXServer binary.
2026-04-30 17:14:22 +02:00
Billy Laws 9a1d06c6ab WOW64: Support disabling DEP
Required for older 32-bit games that assumes the execute bit is implicit
from read.
2026-04-29 03:17:03 +00:00
Billy Laws 3e278b42f8 InstcountCI: Update 2026-04-29 02:54:34 +00:00
Billy Laws fa953445e9 InstcountCI: Update 2026-04-29 02:53:07 +00:00
Billy Laws a412b1d3b7 InstcountCI: Update 2026-04-29 02:43:19 +00:00
Billy Laws fc680388e3 X87: Fix FXTRACT with Inf and NaN inputs
Both the default F80 softfloat wrappers (FXTRACT_SIG/FXTRACT_EXP) and
the reduced-precision F64 dispatcher fell through to the generic
exponent/significand extraction for Inf and NaN inputs, producing
finite garbage. The F80 wrapper returns input unchanged in the
significand slot and +Inf (or NaN) in the exponent slot; the F64
dispatcher detects the exponent-all-ones case and selects the proper
Inf/NaN result before the existing zero-case fold.
2026-04-29 02:34:41 +00:00
Billy Laws 7c260b45e1 unittests/ASM: Adds tests for FXTRACT Inf/NaN 2026-04-29 02:29:23 +00:00
Ryan Houdek 098c4c57b4 Merge pull request #5443 from bylaws/claudefix1
X86Features: Fix AVX/RAND/PCMLULQDQ feature detection on intel
2026-04-28 19:25:15 -07:00
Billy Laws 7c826e35b4 SoftFloat: Fix FSCALE(0, +Inf) to raise IE
The lhs==0 short-circuit in X80SoftFloat::FSCALE returned lhs
unchanged without calling extF80_mul, so the 0*Inf invalid-operation
case never set softfloat_flag_invalid. Detect +Inf rhs explicitly
in the zero-lhs path and raise the flag, returning QNaN to match
hardware.
2026-04-29 02:17:03 +00:00
Billy Laws 1bd2ff3fc3 unittests/ASM: Adds test for FSCALE(0, +Inf) raising IE 2026-04-29 02:16:57 +00:00
Billy Laws e2fc91fc15 OpcodeDispatcher: Fix FIST 16-bit IE for denormal inputs
The special-value check masked the biased exponent with 0x7fff and
then used TestNZ, so it triggered on Exp==0 (denormals) instead of
Exp==0x7fff (NaN/Inf). Replace the TestNZ with SubWithFlags against
0x7fff so EQ only fires for genuine NaN/Inf.
2026-04-29 02:10:35 +00:00
Billy Laws cf20647b25 unittests/ASM: Adds test for 16-bit FIST with denormal input not setting IE 2026-04-29 02:09:54 +00:00
Billy Laws 8d7071e549 VectorOps: Fix MAXPS/MAXPD NaN and signed-zero tie selection
fcmgt returns false on NaN, so the existing polarity in the non-SVE
fcmgt+bit sequences and in the SVE predicate-merge picked the wrong
source on NaN/tie. Swap the compare operands and flip bit<->bif / add
a predicate not to match x86 second source wins behaviour.
2026-04-29 02:00:58 +00:00
Billy Laws fb006b2c6d unittests/ASM: Adds test for MAXPS/MAXPD NaN and signed-zero tie 2026-04-29 01:57:41 +00:00
Billy Laws 9039eeb3cd X86Features: Fix AVX/RAND/PCMLULQDQ feature detection on intel 2026-04-29 01:57:18 +00:00
Billy Laws dd0702d30f unittests/ASM: Mark SSE4a/CLZERO as required for tests using them 2026-04-29 01:56:35 +00:00
LC 886faf0bd4 Merge pull request #5442 from neobrain/refactor_code_cache_check
CodeCache: Move bounds check to FEXOfflineCompiler
2026-04-28 19:47:52 -04:00
Tony Wasserka 86e28c6d34 CodeCache: Move bounds check to FEXOfflineCompiler
The previous check site would easily fail when loading caches for binaries
with multiple executable sections.

It makes much more sense to refuse generating caches anyway: The condition
effectively checked for invalid code map entries, so FEXOfflineCompiler
should reject them as bad inputs.
2026-04-28 17:31:20 +02:00
LC 821efab8aa Merge pull request #5440 from Sonicadvance1/145
OpcodeDispatcher: Ensure FIST* operations avoid TSO/GPR
2026-04-28 07:55:10 -04:00
Ryan Houdek 5295365dd0 InstcountCI: Update 2026-04-27 17:55:49 -07:00
Ryan Houdek 788959a98c OpcodeDispatcher: Ensure FIST* operations avoid TSO/GPR
Noticed this while benchmarking that the FIST* operations were
converting to a GPR, and then storing to memory using an atomic TSO
operation. This should be instead listening to the vector TSO
configuration option. This gives a 3.8x - 6.05x improvement in my
microbench.

Additionally when possible, make sure to use vector conversion
instructions when possible. It's lower cost to avoid the FPR->GPR
transfer, but we can only use it for 64-bit FIST operations. Microbench
couldn't show a difference for that on my platform, but that's because
it's float pipeline bounded regardless. Should help X-class Cortex and
newer Cortex-A.
2026-04-27 17:48:45 -07:00
LC dd145aaa88 Merge pull request #5439 from Sonicadvance1/144
Fix push/pop fs/gs segments and unittests
2026-04-27 19:43:44 -04:00
Ryan Houdek 34b3adc23d unittests/ASM: Adds unit test to ensure push/pop segment of o16 works
Only ensures we are pushing and popping the correct size, not any of the
selector data within it, as 64-bit systems with the FSGSBase extension
don't use them selectors anyway.

Can't test the 32-bit side currently because we would corrupt FS/GS in
CI and the host testharnessrunner can't fix that right now.
2026-04-27 15:21:47 -07:00
Simon Scherer ab14882761 FEXCore: Fix 2byte stack access for 0x66 PUSH/POP FS/GS 2026-04-27 15:21:11 -07:00
LC 7dc1f54fb6 Merge pull request #5435 from Sonicadvance1/143
unittests/ASM: Add unittests to ensure correct cmpxchg8b/16b flag setting
2026-04-25 10:34:07 -04:00
Ryan Houdek c09fb03eda unittests/ASM: Add unittests to ensure correct cmpxchg8b/16b flag setting 2026-04-25 00:46:40 -07:00
Ryan Houdek dbf2761fb7 InstcountCI: Update 2026-04-25 00:45:05 -07:00
Ryan Houdek fd1378f778 InstcountCI: Fix incorrect instruction 2026-04-25 00:43:50 -07:00
Simon Scherer 819dcee3ad FEXCore: Fix wrong shift value to extract NZCV in CmpPairZ 2026-04-25 00:41:24 -07:00
LC 4b02c04afc Merge pull request #5429 from Sonicadvance1/142
Steam/CompatTool: Fixes Graphics Provider path handling
2026-04-23 20:13:13 -04:00
Ryan Houdek deed99e7a3 Merge pull request #5425 from pmatos/f64-atan-fyl2x
JIT-inline F64ATAN and F64FYL2X for reduced precision x87 path
2026-04-23 15:16:04 -07:00
Ryan Houdek 7bffc4a177 Steam/CompatTool: Fixes Graphics Provider path handling
Graphics provider needs to be a path to a json file in the root of the
rootfs. Make sure to strip the filepath off to get the directory.

Misunderstood the assignment before.
2026-04-23 15:13:42 -07:00
Ryan Houdek 701555e400 Merge pull request #5428 from Sonicadvance1/141
Steam/CompatTool: Support `STEAM_COMPAT_GRAPHICS_PROVIDER` for rootfs path
2026-04-21 12:50:22 -07:00
Ryan Houdek 49fa86d0b5 Steam/CompatTool: Support STEAM_COMPAT_GRAPHICS_PROVIDER for rootfs path
If we have been provided a graphics provider path through an environment
variable, then use that path directly rather than searching.
2026-04-21 12:37:40 -07:00
Paulo Matos adbace8810 instcountci: JIT-inline F64ATAN and F64FYL2X for reduced precision x87 path 2026-04-21 20:23:45 +02:00
Paulo Matos 050138bcea asm_tests: JIT-inline F64ATAN and F64FYL2X for reduced precision x87 path 2026-04-21 20:23:45 +02:00
Paulo Matos db50b04663 JIT-inline F64ATAN and F64FYL2X for reduced precision x87 path 2026-04-21 12:09:55 +02:00
LC 59755ec115 Merge pull request #5426 from Sonicadvance1/139
Snapdragon X2 Elite fixes
2026-04-20 11:53:51 -04:00
Ryan Houdek 856fb1e636 Snapdragon X2 Elite fixes
- RNDRRS is still broken on this CPU
- Fault granularity checking needs to use loads
  - stlxp does monitor check before alignment check, use loads for all
    for consistency
  - Fault granularity is still 16B which matches LSE2 requirements.
- X2E still only ships a 19.2Mhz cycle counter, so still only
  ARMv9.0-a hardware
- Adds product name to CPUID
  - Oryon-3 being CPU PartID 2 isn't a mistake.
  - No distinction between Oryon-1 and Oryon-2, both are partid 1.

cpuinfo:
```
processor   : 0
BogoMIPS    : 38.40
Features    : fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm uscat ilrcpc flagm ssbs sb paca pacg dcpodp sve2 sveaes svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 rng ecv afp rpres
CPU implementer   : 0x51
CPU architecture: 8
CPU variant : 0x1
CPU part    : 0x002
CPU revision      : 1
```
2026-04-19 14:34:44 -07:00
Ryan Houdek d41d52b889 Merge pull request #5419 from pmatos/f64-scale-f2xm1
JIT-inline F64SCALE and F64F2XM1 for reduced precision x87 path
2026-04-17 14:38:22 -07:00
Paulo Matos 18f69fb16d instcountci: JIT-inline F64SCALE and F64F2XM1 for reduced precision x87 path 2026-04-17 17:28:13 +02:00
Ryan Houdek 739e85032b FEXCore: Add support for developer single stepping, read/write watching. 2026-04-16 14:13:19 -07:00
Paulo Matos d165711f2e asm_tests: JIT-inline F64SCALE and F64F2XM1 for reduced precision x87 path 2026-04-16 11:24:24 +02:00
Paulo Matos 6bf19ffe37 JIT-inline F64SCALE and F64F2XM1 for reduced precision x87 path 2026-04-16 11:24:19 +02:00
LC 441116e1e6 Merge pull request #5421 from Sonicadvance1/136
Scripts: Move arch check first in InstallFEX
2026-04-15 16:24:04 -04:00
Ryan Houdek ce97ef0ab1 Scripts: Move arch check first in InstallFEX
Don't give people false hope that the script might work on distros that
aren't Ubuntu.

Fixes #5420
2026-04-15 13:01:37 -07:00
LC 2ea0de92f4 Merge pull request #5418 from Sonicadvance1/135
ArchHelpers: Allow atomic memory operations in non-JIT handler
2026-04-15 07:22:25 -04:00
Ryan Houdek 14580c4675 ArchHelpers: Allow atomic memory operations in non-JIT handler
`Detroit: Become Human` decided to use unaligned CriticalSections. So
this workarounds that.
2026-04-14 13:38:41 -07:00
Ryan Houdek 9681559d56 Docs: Update for release FEX-2604 2026-04-09 13:45:35 -07:00
Ryan Houdek b478e4845f Merge pull request #5417 from tiopex/main
FEXRootFSFetcher: clear Unknown when distro is set on the CLI
2026-04-09 13:42:55 -07:00
tpietrus 1fa5104076 FEXRootFSFetcher: clear Unknown when distro is set on the CLI 2026-04-09 07:53:53 +02:00
LC ce65f5376f Merge pull request #5415 from Sonicadvance1/133
FEX: Workaround Docker seccomp bug
2026-04-08 11:39:56 -04:00
LC 0695249fc8 Merge pull request #5413 from Sonicadvance1/132
Arm64EC: Invert suspend doorbell and move out of hot path
2026-04-06 19:49:49 -04:00
LC 51144c99a7 Merge pull request #5408 from Sonicadvance1/129
FEXCore: Use WritePriorityMutex for Code Invalidation Mutex
2026-04-06 19:49:07 -04:00
LC 251398a7cb Merge pull request #5416 from Sonicadvance1/134
OpcodeDispatcher: Fixes nop encoded prefetch instruction
2026-04-06 19:45:52 -04:00
Ryan Houdek 2e6a7f869c OpcodeDispatcher: Fixes nop encoded prefetch instruction
We had a bug where nop encoded prefetch instructions were getting
flagged as illegal instructions erroneously. Fix that and add a unittest
for ensuring execution.

Fixes `Devil May Cry 4`
2026-04-06 11:04:46 -07:00
Ryan Houdek f308162334 FEX: Workaround Docker seccomp bug
Docker's seccomp filter fails to follow AAPCS64 and SysV zero-extension
rules.  For values smaller than 64-bit they were required in their
seccomp filters to truncate the value to the specific size but do not.
Instead they do a 64-bit comparison operation against smaller arguments
(in this case 32-bit). This means 64-bit -1 and 32-bit -1 passed through
have different values for this `personality` syscall.

The real fix would be for Docker to audit their seccomp filter rules and
ensure they zero-extend every argument that is smaller than 64-bit, but
we don't control that. So there is likely to be more bugs in their
filter that we encounter, this is just an easy one to resolve.
2026-04-06 09:51:15 -07:00
Ryan Houdek db4867839c Arm64EC: Invert suspend doorbell and move out of hot path
This was causing a surprisingly high amount of branch mispredicts in
Death Stranding. Suspend doorbell is fairly rare so just invert the
check and move the target down out of the hot path. Then the doorbell
handling code will trampoline to the correct location still.
2026-04-03 20:07:37 -07:00
LC 73ffff7d22 Merge pull request #5409 from Sonicadvance1/130
OpcodeDispatcher: Special case optimize a broadcast
2026-04-03 20:23:39 -04:00
LC dc48a4f73c Merge pull request #5406 from Sonicadvance1/128
FEXRootFSFetcher: Improve hashing performance
2026-04-03 00:26:44 -04:00
Ryan Houdek efbccccdc0 InstcountCI: Update 2026-04-02 18:45:22 -07:00
Ryan Houdek 3e7cd88dcc OpcodeDispatcher: Special case optimize a broadcast
Death Stranding 2 is using this instruction instead of `vbroadcastss`
for some reason. Optimize its specific case and add a note that when we
know sources match that we can optimize more patterns easily.
2026-04-02 18:45:22 -07:00
Ryan Houdek 12cfe8fc37 InstcountCI: Add instruction found in Death Stranding 2 2026-04-02 18:25:19 -07:00
Tony Wasserka c6d2ce043f Merge pull request #5388 from Sonicadvance1/123
Config: Finish wiring up Regex app overrides
2026-04-02 11:00:37 +02:00
Tony Wasserka 5c34c574c8 Merge pull request #5364 from Sonicadvance1/110
Win32: Enable support for virtual naming and THP control
2026-04-02 10:58:30 +02:00
Ryan Houdek d3cfdcb431 Win32: Enable support for virtual naming and THP control
Allows WTF to work (mostly) with Wine by letting us VirtualName things,
and also allows madvise control of THP, which significantly cuts back
memory usage.

This works around the problem of Wine not giving us control of this by
using raw syscalls when wine is detected.

Based on top of #5362 so the THP disable controls are in.
2026-04-01 10:50:40 -07:00
Ryan Houdek 4018d23c39 Config: Finish wiring up Regex app overrides
This wasn't quite wired up exactly how we wanted it. It was previously
matching against the opaque file config handle, which can be anything.

Instead compare it to the appname that now gets passed over to it for
matching.

This allows us to do the following:
```
{
    "Config": {
        "ProfileStats": "1",
        "X87ReducedPrecision": "1",
        "TSOEnabled": "1",
        "VectorTSOEnabled": "0",
        "MemcpySetTSOEnabled": "0",
        "HalfBarrierTSOEnabled":"1",
        "MaxInst": "500",
        "Multiblock": "1"
    },
    "AppOverrides" : {
        "setup*" : {
            "Comment": [
                "292030 - The Witcher 3: Wild Hunt"
            ],
            "X87ReducedPrecision": "0"
        }
     }
}
```

Based on #121 which needs to get merged first.

Code Review

Code Review: Class deletion
2026-04-01 10:46:02 -07:00
badumbatish 81d4e8fe9d Initial implementation for regex engine
Add support for question mark and plus mark in regex, supply testing for star

Added more characters to the regex alphabets, add more test case

Added support for regex matching of configs, awaiting reviews

Rename variable to CamelCase

Addresses PR reviews

Remove unnecessary features and test cases

Rewrite to naive regex with dp

Addresses PR reviews

Build fixes

Code Review
2026-04-01 10:45:59 -07:00
Tony Wasserka 34b48c4069 Merge pull request #5383 from Sonicadvance1/119
FEXGetConfig: Test for showing fault granularity
2026-04-01 10:48:18 +02:00
Ryan Houdek 01a3ab6ca7 FEXCore: Use WritePriorityMutex for Code Invalidation Mutex
The previous `ForkableSharedMutex` using `std::shared_mutex` was showing
up as significant CPU time on arm64ec. In particular it was showing up
upwards of 700ms/S of CPU time for read-contented workloads at only
~2700 locks per second. The libc++ implementation for arm64ec must be
particularly gnarly for this to be so slow.

With this swapped over, it's now only spending around 56ms/S on the
contended shared lock, but at ~8000 locks per second. So a significant
uplift.
2026-03-31 19:02:51 -07:00
Ryan Houdek 476c242d7f FEXCore: Moves SpinWaitLock and WritePriorityMutex to frontend visible includes
This will be used in a moment.
2026-03-31 19:02:51 -07:00
LC ae3fa6a836 Merge pull request #5403 from Sonicadvance1/127
IR: Adds support for printing strings
2026-03-31 16:23:42 -04:00
Ryan Houdek ba93bdd66d FEXRootFSFetcher: Improve hashing performance
Don't use pread, instead map the file and madvise larger blocks. This
removes copying overhead as its just mapping file pages in instead.
Also splits the implementation of file reading from hashing to make
tinkering less involved, as if I want more performance out of this (say
due to live hashing) then it's easier to tinker.

Improves hashing performance from ~2.2GB/s to ~3.6GB/s on my system,
which is CPU bounded by xxhash here.
2026-03-30 18:16:46 -07:00
Ryan Houdek 1fa0b37fac FEXGetConfig: Test for showing fault granularity
Useful for seeing if behaviour has changed. Useful with the
`--tso-emulation-info` option to show hardware behaviour
2026-03-30 12:32:00 -07:00
LC 5b4a5969cc Merge pull request #5401 from Sonicadvance1/126
Linux/GuestFrames: Ensure MXCSR is saved and restored on signal
2026-03-28 22:25:59 -04:00
Ryan Houdek 941a7934ef InstcountCI: Update 2026-03-28 18:06:54 -07:00
Ryan Houdek 474439ab4b InstcountCI: Update 2026-03-28 17:56:23 -07:00
Ryan Houdek 4428aea5ee IR: Adds support for printing strings
I utilize this functionality quite heavily when debugging and I need
bread crumbs spread around. Instead of reimplementing it a dozen times,
just have it upstreamed.
2026-03-28 17:53:33 -07:00
Ryan Houdek e9a9cc5bc3 unittests/FEXLinuxTests: Adds an MXCSR signal test
Ensures that the MXCSR value stays the same with a signal inbetween that
modifies it.
2026-03-28 17:49:12 -07:00
Ryan Houdek 2291c5b230 OpcodeDispatcher/Vector: Make sure MXCSR is masked
We don't support the exception bits, make sure these are masked off so
spurious exception checks don't break.
2026-03-28 17:47:55 -07:00
Ryan Houdek f78e194cf7 Linux/GuestFrames: Ensure MXCSR is saved and restored on signal
This was causing an unfortunate set of circumstances where Dark Souls
III was modifying MXCSR and we weren't saving it, cause the value to change
from 0x9fc0 to 0.

This "enabled" float exceptions by unmasking the exception masks in
MXCSR. This in turn had Dark Souls III's `expf` function to fault out,
as it checks if the MXCSR exception masks are set or not for determining
if underflow should assert or not.

Wow64/arm64ec has a similar problem where it always sets back to default
on signal. Which means game lose DAZ, but I'm not fixing that bug right
now.

Fixes #5391
2026-03-28 17:44:45 -07:00
LC b77ddcf1a7 Merge pull request #5398 from Sonicadvance1/125
Allocators: Remove legacy NOREPLACE handling
2026-03-27 08:14:05 -04:00
Ryan Houdek 3ef677537d Merge pull request #5397 from neobrain/fix_elfreads
LinuxSyscalls: Skip reading ELF files when code caching is disabled
2026-03-26 14:33:26 -07:00
Tony Wasserka 1df1265ed1 LinuxSyscalls: Skip reading ELF files when code caching is disabled
ELF headers were read unconditionally because doing so was assumed to be cheap
(as the guest app would read them anyway shortly after). However, relocation
parsing was added since then, which has less predictable performance due to
crossing page boundaries and reading larger amounts of memory. It might be
possible to make the underlying code more efficient, but until that's done
it's better to skip this logic unless needed.

Closes #5390
2026-03-26 21:40:20 +01:00
Ryan Houdek f0854a16fe Allocators: Remove legacy NOREPLACE handling
We needed this handling on old kernels that didn't understand the
NOREPLACE flag. We no longer support kernels this old, so remove some of
this vestigial code.
2026-03-26 13:36:29 -07:00
Ryan Houdek 6bd476fb03 Merge pull request #5395 from lioncash/cpuid
CPUID: Add basic stub handling for AVX10 info
2026-03-25 17:16:24 -07:00
Lioncache b9c0af7c3b CPUID: Add basic handling for AVX10 info
Just gets the feature bit handling stuff in place for various
facilities, so it can be easily expanded in the future.
2026-03-25 19:18:05 -04:00
Tony Wasserka 5149ebc70e Merge pull request #5389 from Sonicadvance1/124
SMCTracking: Remove relocation log
2026-03-24 11:12:10 +01:00
Ryan Houdek e92a6a5803 SMCTracking: Remove relocation log
Holy jeez does this thing spam.
2026-03-23 19:55:11 -07:00
Ryan Houdek 6da963a695 Merge pull request #5394 from neobrain/fix_eager_relocation_parsing
LinuxSyscalls: Don't try to parse relocations unless ELF parsing succeeded
2026-03-23 14:01:20 -07:00
Tony Wasserka 2a74489858 LinuxSyscalls: Don't try to parse relocations unless ELF parsing succeeded 2026-03-23 17:47:47 +01:00
Ryan Houdek 65a436ca98 Merge pull request #5392 from neobrain/fix_code_cache_dupfd
LinuxSyscalls: Re-open file descriptors for parsing ELF headers
2026-03-23 09:42:45 -07:00
LC bc533c8050 Merge pull request #5393 from neobrain/fix_code_cache_glibc_assert
CodeCache: Fix glibc debug mode assertion
2026-03-22 14:07:27 -04:00
Tony Wasserka fd6cea4698 CodeCache: Fix glibc debug mode assertion
If begin == end, the first vector::erase() call would invalidate the begin
iterator.
2026-03-22 10:11:53 +01:00
Tony Wasserka a1aa1658ec LinuxSyscalls: Re-open file descriptors for parsing ELF headers
File descriptors returned by dup() share state with the original FD, so we
need to use open() to create a fully independent object.

Fixes #5379
2026-03-22 10:09:37 +01:00
LC 5c4c468d13 Merge pull request #5387 from Sonicadvance1/122
github: Stop running unittests always on build failure
2026-03-20 00:58:09 -04:00
Ryan Houdek 69ef1658cc github: Stop running unittests always on build failure
This was taking too much time.
2026-03-19 20:08:10 -07:00
Ryan Houdek 8c72aa76a0 Merge pull request #5385 from lioncash/ilog
MemoryOps: Collapse duplicate add/sub in Memset
2026-03-19 18:59:17 -07:00
Lioncache 928a932a43 MemoryOps: Collapse duplicate add/sub in Memset
We can just use ilog2 to deduplicate this a bit.
2026-03-19 21:34:03 -04:00
Ryan Houdek 83601055dc Merge pull request #5368 from neobrain/feature_cc_elf_relocations
CodeCache: Support ELF relocations
2026-03-19 18:10:09 -07:00
Ryan Houdek 68480f6e43 Merge pull request #5356 from lioncash/mops
MemoryOps: Drop MOPS handling into place for MemSet/MemCpy
2026-03-19 18:04:14 -07:00
Ryan Houdek 42291540ab Merge pull request #5343 from pmatos/f64-sin-cos-tan
JIT-inline F64SIN, F64COS, and F64TAN for reduced precision x87 path
2026-03-19 17:51:50 -07:00
LC 194eb69838 Merge pull request #5384 from Sonicadvance1/120
Cmake: Default to release builds with a message
2026-03-19 19:42:39 -04:00
Ryan Houdek c1d27fa453 Cmake: Default to release builds with a message
People keep forgetting to set this and have a worse experience.
Default to a Release build, which ensures optimizations are enabled and
assertions are disabled.
2026-03-19 16:25:40 -07:00
Paulo Matos 9d5f7caa79 instcountci: JIT-inline F64SIN, F64COS, and F64TAN for reduced precision x87 path 2026-03-19 19:18:02 +01:00
Paulo Matos a1d78dceb0 asm_tests: JIT-inline F64SIN, F64COS, and F64TAN for reduced precision x87 path 2026-03-19 19:18:02 +01:00
Lioncache 2bcf435e0a MemoryOps: Handle overlapping memcpy 2026-03-19 14:13:36 -04:00
Paulo Matos 30e853305d JIT-inline F64SIN, F64COS, and F64TAN for reduced precision x87 path 2026-03-19 18:58:46 +01:00
Lioncache 600f4bb2b6 unittests: Add overlapping tests for MemCpy 2026-03-19 11:55:02 -04:00
Lioncache 5868814c91 MemoryOps: Drop in MOPS handling for MemCpy
With the MOPS featureset dropped in, we can also accelerate memcpy paths
on hardware that supports it.
2026-03-19 11:55:02 -04:00
Lioncache e862f8f86c unittests: Add specific paths for MOPS 2026-03-19 11:55:02 -04:00
Lioncache 85c1ecd035 MemoryOps: Handle inline values in MemSet() MOPS path
Lets us handle potential inline memset values.

Also fixes up the STOS tests to actually ensure all values
in the verification step pass.
2026-03-19 11:55:02 -04:00
Lioncache 68ad448672 MemoryOps: Drop 8-bit memset support into MemSet()
Can be further expanded to handle other optimization cases, but this
kicks it off for forward direction memsets at least.
2026-03-19 11:55:02 -04:00
LC c18fb3cb78 Merge pull request #5382 from Sonicadvance1/118
FEXpidof: Fixes another missing std::filesystem throw
2026-03-18 18:07:08 -04:00
Tony Wasserka 53702f989c Merge pull request #5362 from Sonicadvance1/108
FEX: Disable THP on key allocations that consume memory
2026-03-18 21:04:15 +01:00
Ryan Houdek c547b1bec3 FEX: Disable THP on key allocations that consume memory
Disables THP on some key locations that are fairly sparse
- rpmalloc
  - This is the big one as this allocates some heavy sparse buffers.
- CallRet stacks
  - These get in the hundreds of megabytes, while not being sparse they
    trend towards only using a handful of pages and ballooning to 2MB
    per thread is quite heavy.
- Lookup cache
  - L1 specifically gets hit here which adds a decent chunk of overhead
    due to sparsity.

Win32 for all of these also aren't handled, but that will need to be a
followup.
2026-03-18 12:15:39 -07:00
Ryan Houdek 740350c8ea FEXpidof: Fixes another missing std::filesystem throw
Turns out std::filesystem::exists throws as well if there was an
underlying OS API failure.
2026-03-18 12:10:22 -07:00
Tony Wasserka 9f9b20eac0 Merge pull request #5378 from Sonicadvance1/117
SMCTracking: Move read check up for ELF parsing
2026-03-18 12:04:23 +01:00
Tony Wasserka c67ffb82a8 Core: Support reporting blocks that are uncacheable due to unhandled ELF relocations 2026-03-18 11:59:44 +01:00
Tony Wasserka 1ea24f3d6c LinuxSyscalls: Enable delayed code cache load for ELF files
Specifically this is needed if any ELF relocations cover read-only code
sections, which is indicated in the ELF headers via DT_TEXTREL/DF_TEXTREL.
2026-03-18 11:59:41 +01:00
Tony Wasserka 226bd51afe LinuxSyscalls: Implement delayed cache load for binaries that require ELF/PE relocations 2026-03-18 11:58:26 +01:00
Tony Wasserka 293568be36 FEXOfflineCompiler: Apply relocations to loaded ELF binaries 2026-03-18 11:57:38 +01:00
Tony Wasserka 152fe81d16 LinuxSyscalls: Parse and provide ELF relocation information to the JIT 2026-03-18 11:57:28 +01:00
Ryan Houdek 494dd64c50 Merge pull request #5372 from CxnYusuf/add-fisttp-tests
Tests/ASM: Add FISTTP unit tests for 16, 32, 64-bit as well as negatives
2026-03-17 16:55:56 -07:00
Ryan Houdek 56de0d1ab4 Merge pull request #5377 from Sonicadvance1/116
FEXServer: Try both fusermount and fusermount3
2026-03-17 16:55:42 -07:00
Ryan Houdek 73c1f4cc54 Merge pull request #5374 from neobrain/fix_gcc_build
Fix most GCC build issues
2026-03-17 16:55:23 -07:00
Ryan Houdek 6a6a82385e Merge pull request #5381 from neobrain/fix_jit_restarts
JIT: Reset relocations on restart
2026-03-17 13:47:42 -07:00
Tony Wasserka fc8ef0e723 JIT: Reset relocations on restart 2026-03-17 21:37:01 +01:00
Ryan Houdek de11c05d2a Merge pull request #5380 from neobrain/fix_cc_32bit_constants
Arm64Emitter: Disable 32-bit constant optimization when NOP-padding is requested
2026-03-17 13:34:20 -07:00
Tony Wasserka addbc8cad8 Arm64Emitter: Disable 32-bit constant optimization when NOP-padding is requested
Code caching requires this even for simple libraries like libdl.so (as observed
in the 32-bit build of Super Meat Boy).
2026-03-17 21:06:13 +01:00
Ryan Houdek 63e37b7cbb SMCTracking: Move read check up for ELF parsing 2026-03-17 12:43:54 -07:00
Ryan Houdek 9ee329034f FEXServer: Try both fusermount and fusermount3
Apparently some distros don't symlink these, so try both with the newer
fusermount3 going first as its the common path now.

Fixes #5375
2026-03-17 12:36:49 -07:00
Ryan Houdek f3e904207b Merge pull request #5376 from OFFTKP/flag
Add test for shifts preserving flags
2026-03-17 09:40:59 -07:00
LC e4ae6ce635 Merge pull request #5373 from Sonicadvance1/115
code-format-helper: Another dependabot upgrade
2026-03-17 10:10:46 -04:00
Paris Oplopoios f7d76255ad Add test for shift preserving flags
Signed-off-by: Paris Oplopoios <21157395+OFFTKP@users.noreply.github.com>
2026-03-17 15:50:33 +02:00
Tony Wasserka ea45f9c694 FEXCore/VectorRegType: Use vector_size on GCC
GCC does not support neon_vector_type and silently ignores that attribute,
but vector_size(16) seems to have the same effect.
2026-03-16 19:15:05 +01:00
Tony Wasserka 1d449c0f58 FEXCore/Utils: Add quotes around preprocessor errors 2026-03-16 19:15:05 +01:00
Tony Wasserka 9ecc991043 LibraryForwarding: Remove unnecessary const qualifier 2026-03-16 19:15:05 +01:00
Tony Wasserka 4ef834859c SignalDelegator: Don't use the same name for two different symbols 2026-03-16 19:15:05 +01:00
Tony Wasserka fa082bc5c4 FileManagement: Fix ambiguous name reference 2026-03-16 19:15:05 +01:00
Tony Wasserka 67caab026a OpcodeDispatcher: Fix inconsistent types in ternary conditional 2026-03-16 19:15:05 +01:00
Tony Wasserka 91c55facb2 IRDumper: Avoid passing packed member data as references
Fixes "Cannot bind packed field to reference type" build errors on GCC.
2026-03-16 19:15:05 +01:00
Tony Wasserka 321d4d84d7 LinuxSyscalls: Don't cast away qualifiers 2026-03-16 19:15:05 +01:00
Tony Wasserka 22faa58e0b X86Tables: Use explicit type for SecondInstGroupOps definition
GCC considers it a "conflicting declaration" to use auto for a variable that
was already declared before.
2026-03-16 19:15:05 +01:00
Tony Wasserka ebd559f662 Core: Fix offsetof with runtime array indexes
GCC does not support this clang-specific language extension.
2026-03-16 19:15:05 +01:00
Tony Wasserka d27c9d3f98 CodeEmitter: Fix ambigious ExtendedType declaration 2026-03-16 18:50:01 +01:00
Tony Wasserka fbef482265 CMake: Link against libatomic if compiling with GCC 2026-03-16 18:50:01 +01:00
Tony Wasserka a57926ac57 CMake: Explicitly demote -Wchanges-meaning diagnostics to warnings on GCC 2026-03-16 18:50:01 +01:00
Ryan Houdek 70a7137e62 code-format-helper: Another dependabot upgrade 2026-03-15 20:45:17 -07:00
CxnYusuf 2d3a08a362 Tests/ASM: Add FISTTP unit tests for 16, 32, 64-bit as well as negatives 2026-03-16 03:10:38 +01:00
LC cc02edb3f6 Merge pull request #5371 from Sonicadvance1/114
Syscalls: Fixes crash in ELF parsing code
2026-03-15 20:36:00 -04:00
Ryan Houdek f894cd90f3 Merge pull request #5369 from Sonicadvance1/113
FEXCore: Update CPU frequency to be 64-bit
2026-03-15 15:20:28 -07:00
Ryan Houdek 24675969cc Merge pull request #5366 from Sonicadvance1/112
JIT: Use struct for Spill/Fill default arguments
2026-03-15 15:20:13 -07:00
Ryan Houdek a0cba1194e Syscalls: Fixes crash in ELF parsing code
When an application maps a file as PROT_NONE, we can't check if it is an
ELF. Was causing a crash in `Cisco Packet Tracer`.
2026-03-15 15:16:29 -07:00
Ryan Houdek 0952fa95f4 FEXCore: Update CPU frequency to be 64-bit
This annoyed by by trying to do math on a 32-bit value and it
overflowing. Just make it 64-bit.
2026-03-13 17:04:25 -07:00
Ryan Houdek 6177ab957b Merge pull request #5363 from Sonicadvance1/109
External/code-format-helper: Update dependencies
2026-03-13 11:52:07 -07:00
Ryan Houdek 5c1300a2b9 Merge pull request #5365 from Sonicadvance1/111
Config: Fix issue with config overrides
2026-03-13 11:51:51 -07:00
Ryan Houdek af9dd0827a JIT: Use struct for Spill/Fill default arguments
Cleans up the interface and makes the arguments explicit about what
they're setting. As promised from #5317
2026-03-12 19:25:12 -07:00
Ryan Houdek dc0162122f Config: Fix issue with config overrides
Accidentally was checking for Config override in the combination of
portable config and `FEX_APP_CONFIG_LOCATION`.

Fixes an early crash in PV.
2026-03-12 17:46:05 -07:00
Ryan Houdek ae491fb15b External/code-format-helper: Update dependencies
Removes dependabot alert.
2026-03-12 15:31:02 -07:00
Ryan Houdek 957c1fc420 Merge pull request #5357 from Sonicadvance1/106
Config: Enable Dynamic L1 and Disabled L2 caches by default
2026-03-11 14:12:38 -07:00
Ryan Houdek 86acfb35aa Config: Enable Dynamic L1 and Disabled L2 caches by default
Dramatically reduces memory consumption of FEX's per-thread lookup
structures. Primarily because L2 cache entirely goes away which can end
up reaching hundreds of megabytes or over a gigabyte of memory in some
cases, but also because L1 cache dynamically scales based on load.

Useful for conserving memory on systems with less than 16GB of RAM and
are UMA, like Asahi users inside of muvm.
2026-03-11 13:42:32 -07:00
LC e27d12ee5e Merge pull request #5359 from Sonicadvance1/107
OpcodeDispatcher: Fixes memory leak in IREmitter pool allocator
2026-03-11 08:40:04 -04:00
Ryan Houdek 63d52c2a1a OpcodeDispatcher: Fixes memory leak in IREmitter pool allocator
While not a leak in the traditional sense, we were causing pool
allocations to never become free until the thread was closed.

This meant in the case of a game running with >200 threads or so, these
would add up very quickly. So some minor reworking so the IREmitter
doesn't allocate a buffer until first JIT, and making sure to actually
disown the buffer on dispatch error resolved the problems.

Fixes an edge case where Ender Lilies was consuming 409MB with THP
enabled on my desktop, and now it is something like 6MB once idling for
a bit to have the pool allocations do its magic.
2026-03-10 20:55:09 -07:00
Ryan Houdek 9d4a71b57a Merge pull request #5355 from wsxarcher/patch-1
Handle zero length in ChangeProtectionFlags
2026-03-10 08:40:16 -07:00
Marco Bartoli 3462dc3e14 Handle zero length in ChangeProtectionFlags
Add a no-op for zero length in ChangeProtectionFlags.

This fixes AMD Vivado 2025.2 which tries to mprotect with 0 as size and merge strategies fails:

```
Unexpected ChangeProtectionFlags Merge strategy! [0x400000, 0x401000) Versus [0x0, 0x0)
```
2026-03-10 11:34:52 +01:00
LC d21351e66e Merge pull request #5354 from Sonicadvance1/104
Config: Fixes
2026-03-09 21:51:05 -04:00
Ryan Houdek 3204d20335 Config: Fix priorities of config paths
Fixes b4a87d8c0b

`FEX_APP_CONFIG_LOCATION` wasn't overriding paths properly anymore once
that commit landed. Instead legacy `~/.fex-emu/` path would get returned
if it existed first.

Ensures that it returns first, before `STEAM_COMPAT_DATA_PATH` even.
2026-03-09 18:20:56 -07:00
Ryan Houdek a519489d80 CMake: Make sure not to compile Steam tools on mingw 2026-03-09 17:25:59 -07:00
Ryan Houdek 5558c3a35a Merge pull request #5353 from lioncash/group
HostFeatures: Group feature ifdefs together more
2026-03-09 15:45:55 -07:00
Ryan Houdek 58d9755314 Merge pull request #5352 from Sonicadvance1/103
gitlab-ci: Update requirements
2026-03-09 15:45:48 -07:00
Lioncache 498ba0a384 HostFeatures: Group feature ifdefs together more
Makes it a little nicer to see everything grouped together.
2026-03-09 18:27:56 -04:00
Ryan Houdek 5a5477e895 gitlab-ci: Update requirements 2026-03-09 14:00:08 -07:00
Ryan Houdek a17d7ce6ba Merge pull request #5351 from lioncash/hostmops
HostFeatures: Drop in feature testing for FEAT_MOPS
2026-03-09 12:09:31 -07:00
LC afc7248912 Merge pull request #5347 from Sonicadvance1/102
CPUBackend: Enable Transparent Huge Pages on JIT buffers
2026-03-09 14:54:08 -04:00
Lioncache 6bb578fea8 HostFeatures: Drop in feature testing for FEAT_MOPS 2026-03-09 13:03:45 -04:00
Ryan Houdek bed5f293dc CPUBackend: Enable Transparent Huge Pages on JIT buffers
If/When this works, the amount of iTLB misses drop dramatically, which
reduces L2 TLB pressure, which just improves performance for our CPU
cores that have itty-bitty L1 iTLB entry counts.

We can't use `mmap(MAP_HUGETLB)` directly for terrible reasons, so we
are required to lean on madvise instead.
2026-03-06 11:51:17 -08:00
636 changed files with 51947 additions and 23025 deletions

No files matched your search

+14 -13
View File
@@ -49,6 +49,7 @@ jobs:
run: cmake --build build --target asm_files 32bit_asm_files JemallocLibs Catch2 vixl cephes_128bit
- name: Build
id: build
run: cmake --build build
- name: Install
@@ -56,40 +57,40 @@ jobs:
# GCC tests
- name: GCC64 Target Tests
if: ${{ always() }}
if: steps.build.outcome == 'success'
uses: ./.github/workflows/test
with:
target: gcc_target_tests_64
- name: GCC32 Target Tests
if: ${{ always() }}
if: steps.build.outcome == 'success'
uses: ./.github/workflows/test
with:
target: gcc_target_tests_32
# API tests
- name: API Tests
if: ${{ always() }}
if: steps.build.outcome == 'success'
uses: ./.github/workflows/test
with:
target: api_tests
- name: FEXCore API Tests
if: ${{ always() }}
if: steps.build.outcome == 'success'
uses: ./.github/workflows/test
with:
target: fexcore_apitests
# ARM emission tests
- name: ARM Emitter Tests
if: ${{ always() }}
if: steps.build.outcome == 'success'
uses: ./.github/workflows/test
with:
target: emitter_tests
# Linux tests
- name: FEX Linux Tests
if: ${{ always() }}
if: steps.build.outcome == 'success'
uses: ./.github/workflows/test
with:
target: fex_linux_tests_all
@@ -98,13 +99,13 @@ jobs:
# Thunking
- name: Thunkgen tests
if: ${{ always() }}
if: steps.build.outcome == 'success'
uses: ./.github/workflows/test
with:
target: thunkgen_tests
- name: Test GL No-Thunks
if: ${{ always() && matrix.arch[1] == 'x64' }}
if: ${{ steps.build.outcome == 'success' && matrix.arch[1] == 'x64' }}
uses: ./.github/workflows/test
with:
target: thunk_functional_tests_nothunks
@@ -112,7 +113,7 @@ jobs:
DISPLAY: ':0'
- name: Test GL Thunks
if: ${{ always() && matrix.arch[1] == 'x64' }}
if: ${{ steps.build.outcome == 'success' && matrix.arch[1] == 'x64' }}
uses: ./.github/workflows/test
with:
target: thunk_functional_tests_thunks
@@ -121,28 +122,28 @@ jobs:
# ASM tests
- name: ASM Tests
if: ${{ always() }}
if: steps.build.outcome == 'success'
uses: ./.github/workflows/test
with:
target: asm_tests
# POSIX tests
- name: POSIX Tests
if: ${{ always() }}
if: steps.build.outcome == 'success'
uses: ./.github/workflows/test
with:
target: posix_tests
# GVisor tests
- name: GVisor Tests
if: ${{ always() }}
if: steps.build.outcome == 'success'
uses: ./.github/workflows/test
with:
target: gvisor_tests
# Struct verifier tests
- name: Struct verifier tests
if: ${{ always() }}
if: steps.build.outcome == 'success'
uses: ./.github/workflows/test
with:
target: struct_verifier
+1 -1
View File
@@ -43,7 +43,7 @@ jobs:
distrobox upgrade steamrt4
distrobox enter --name steamrt4 -- sudo apt-get install -y \
git cmake ninja-build ccache \
lld clang \
lld clang clang-tools \
libclang-dev llvm-dev \
libstdc++-14-dev-i386-cross libgcc-14-dev-i386-cross \
libstdc++-14-dev-amd64-cross libgcc-14-dev-amd64-cross
+15 -1
View File
@@ -24,7 +24,7 @@ runs:
cmake -S . -B build_${{ inputs.target }} -DCMAKE_BUILD_TYPE=$BUILD_TYPE -DCMAKE_TOOLCHAIN_FILE=Data/CMake/toolchain_mingw.cmake \
-DMINGW_TRIPLE=${_cc}-w64-mingw32 -DCMAKE_INSTALL_LIBDIR=/usr/lib/wine/aarch64-windows -G Ninja \
-DENABLE_LTO=False -DENABLE_ASSERTIONS=False -DENABLE_JEMALLOC_GLIBC_ALLOC=False \
-DBUILD_TESTING=False -DCMAKE_INSTALL_PREFIX=/usr -DTUNE_ARCH=generic -DTUNE_CPU=none
-DBUILD_TESTING=False -DCMAKE_INSTALL_PREFIX=/usr -DTUNE_ARCH=generic -DTUNE_CPU=none -DRANGES_NATIVE=OFF
- name: Build
shell: bash
@@ -33,3 +33,17 @@ runs:
- name: Install
shell: bash
run: DESTDIR="$PWD"/install cmake --build build_${{ inputs.target }} -t install
- name: Configure UnixLib
shell: bash
run: |
cmake -S Source/Windows/UnixLib -B build_unixlib_${{ inputs.target }} -DCMAKE_BUILD_TYPE=$BUILD_TYPE \
-G Ninja -DCMAKE_INSTALL_LIBDIR=/usr/lib/wine/aarch64-unix -DCMAKE_INSTALL_PREFIX=/usr
- name: Build UnixLib
shell: bash
run: cmake --build build_unixlib_${{ inputs.target }}
- name: Install UnixLib
shell: bash
run: DESTDIR="$PWD"/install cmake --build build_unixlib_${{ inputs.target }} -t install
+3 -1
View File
@@ -50,6 +50,8 @@ jobs:
with:
overwrite: true
name: wine_dll_artifacts
path: ${{ github.workspace }}/install/usr/lib/wine/aarch64-windows/lib*.dll
path: |
${{ github.workspace }}/install/usr/lib/wine/aarch64-windows/lib*.dll
${{ github.workspace }}/install/usr/lib/wine/aarch64-unix/lib*.so
retention-days: 60
compression-level: 9
+42 -3
View File
@@ -1,3 +1,17 @@
spec:
inputs:
PROMOTE_BRANCH:
description: "Branch to promote the build to. Empty means no promotion."
default: "bleeding-edge"
---
workflow:
rules:
- when: always
variables:
PROMOTE_BRANCH: $[[ inputs.PROMOTE_BRANCH ]]
variables:
DEBIAN_FRONTEND: noninteractive
GIT_SUBMODULE_STRATEGY: recursive
@@ -5,7 +19,8 @@ variables:
CC: clang
CXX: clang++
aarch64:
build:
stage: build
image: registry.gitlab.steamos.cloud/steamrt/steamrt4/sdk/arm64:4.0.20251117.183306
tags:
- docker
@@ -16,12 +31,12 @@ aarch64:
- apt-get -y update
- apt-get install -y
git cmake ninja-build ccache
lld clang
lld clang clang-tools
libclang-dev llvm-dev
libstdc++-14-dev-i386-cross libgcc-14-dev-i386-cross
libstdc++-14-dev-amd64-cross libgcc-14-dev-amd64-cross
- cmake -E make_directory build/
- cmake -DCMAKE_BUILD_TYPE=Release -G Ninja -DBUILD_STEAM_SUPPORT=True -DENABLE_LTO=True -DENABLE_ASSERTIONS=False -DBUILD_THUNKS=True -DBUILD_FEXCONFIG=False -DBUILD_TESTING=False -DENABLE_CLANG_THUNKS=True -DUSE_LINKER=lld -DCMAKE_INSTALL_PREFIX=/usr -DTUNE_ARCH=armv8.2-a -DTUNE_CPU=none . -B build/
- cmake -DCMAKE_BUILD_TYPE=Release -G Ninja -DBUILD_STEAM_SUPPORT=True -DENABLE_LTO=True -DENABLE_ASSERTIONS=False -DBUILD_THUNKS=True -DBUILD_FEXCONFIG=False -DBUILD_TESTING=False -DENABLE_CLANG_THUNKS=True -DUSE_LINKER=lld -DCMAKE_INSTALL_PREFIX=/usr -DTUNE_ARCH=armv8.2-a -DTUNE_CPU=none -DRANGES_NATIVE=OFF . -B build/
- cmake --build build/ --config Release
- DESTDIR=$(pwd)/install/ cmake --build build/ --config Release -t install
@@ -30,3 +45,27 @@ aarch64:
untracked: false
paths:
- install/
promote:
stage: deploy
variables:
GIT_STRATEGY: none
image: registry.gitlab.steamos.cloud/steamrt/steamrt4/sdk/arm64:4.0.20251117.183306
tags:
- docker
- linux
- arm64
- aarch64
rules:
- if: '$PROMOTE_BRANCH'
before_script:
- apt-get -y update
- apt-get install -y tmux curl
script:
# comment out to debug: SSH in via GCP, go down the container and attach to the session (with `tmux attach -t debug`)
# - tmux new-session -d -s debug
# - while tmux has-session -t debug 2>/dev/null; do sleep 1; done
# ref controls which fex-depot code runs the pipeline, while VERSION_PARAM controls which fex branch's artifacts that pipeline downloads.
- >
curl --fail --location --request POST --form token=${FEX_DEPOT_TRIGGER_TOKEN} --form ref=master --form "variables[PROMOTE_BRANCH]=${PROMOTE_BRANCH}" --form "variables[VERSION_PARAM]=${CI_COMMIT_REF_NAME}" "${CI_API_V4_URL}/projects/fex%2Ffex-depot/trigger/pipeline"
+1
View File
@@ -0,0 +1 @@
AI must not be used to generate code for contributions to this project.
+1
View File
@@ -0,0 +1 @@
AI must not be used to generate code for contributions to this project.
+27 -3
View File
@@ -172,6 +172,13 @@ set(TUNE_ARCH "generic" CACHE STRING "Override the Arch the build is tuned for")
set(OVERRIDE_VERSION "detect" CACHE STRING "Override the FEX version")
set(OVERRIDE_HASH "detect" CACHE STRING "Override the FEX git hash")
get_property(IS_MULTI_CONFIG GLOBAL PROPERTY GENERATOR_IS_MULTI_CONFIG)
if (NOT IS_MULTI_CONFIG AND NOT CMAKE_BUILD_TYPE)
set(CMAKE_BUILD_TYPE Release
CACHE STRING "Choose the type of build." FORCE)
message(STATUS "No build type set, defaulting to a Release build")
endif()
string(TOUPPER "${CMAKE_BUILD_TYPE}" CMAKE_BUILD_TYPE)
if (CMAKE_BUILD_TYPE MATCHES "DEBUG")
set(ENABLE_ASSERTIONS TRUE)
@@ -187,6 +194,11 @@ if (ENABLE_GDB_SYMBOLS)
add_compile_definitions(GDB_SYMBOLS_ENABLED=1)
endif()
add_compile_definitions(_LARGEFILE64_SOURCE)
if (WIN32)
add_compile_definitions(UNICODE _UNICODE)
endif()
set(CMAKE_CXX_STANDARD 20)
set(CMAKE_EXPORT_COMPILE_COMMANDS ON)
set(CMAKE_RUNTIME_OUTPUT_DIRECTORY ${CMAKE_BINARY_DIR}/Bin)
@@ -244,8 +256,15 @@ endif()
if (ENABLE_CCACHE)
find_program(CCACHE_PROGRAM ccache)
if(CCACHE_PROGRAM)
message(STATUS "CCache enabled")
set_property(GLOBAL PROPERTY RULE_LAUNCH_COMPILE "${CCACHE_PROGRAM}")
execute_process(COMMAND "${CCACHE_PROGRAM}" --print-version
OUTPUT_VARIABLE CCACHE_VERSION OUTPUT_STRIP_TRAILING_WHITESPACE)
message(STATUS "Enabling ccache ${CCACHE_VERSION}")
if (CCACHE_VERSION VERSION_GREATER_EQUAL "4.8")
# Set sloppiness to enable caching even for files that use __DATE__/__TIME__ macros
set_property(GLOBAL PROPERTY RULE_LAUNCH_COMPILE "${CCACHE_PROGRAM} sloppiness=time_macros")
else()
set_property(GLOBAL PROPERTY RULE_LAUNCH_COMPILE "${CCACHE_PROGRAM}")
endif()
endif()
endif()
@@ -446,6 +465,11 @@ if(ENUM_ENUM_WARNING)
add_compile_options(-Wno-deprecated-enum-enum-conversion)
endif()
# GCC enables -Wchanges-meaning by default and treats some cases as an error
if(CMAKE_CXX_COMPILER_ID STREQUAL "GNU")
add_compile_options(-Wno-error=changes-meaning)
endif()
if(ENABLE_WERROR OR ENABLE_STRICT_WERROR)
add_compile_options(-Werror)
if (NOT ENABLE_STRICT_WERROR)
@@ -681,6 +705,6 @@ if (BUILD_THUNKS)
add_dependencies(uninstall uninstall_guest-libs-32)
endif()
if (BUILD_STEAM_SUPPORT)
if (NOT MINGW AND BUILD_STEAM_SUPPORT)
add_subdirectory(Source/Steam/)
endif()
-132
View File
@@ -1,132 +0,0 @@
{
"environments": [
{
"BuildPath": "${projectDir}\\out\\build\\${name}",
"InstallPath": "${projectDir}\\out\\install\\${name}",
"clangcl": "clang-cl.exe",
"cc": "clang",
"cxx": "clang++"
}
],
"configurations": [
{
"name": "WSL-Clang-Debug",
"generator": "Ninja",
"configurationType": "Debug",
"buildRoot": "${env.BuildPath}",
"installRoot": "${env.InstallPath}",
"cmakeExecutable": "/usr/bin/cmake",
"cmakeCommandArgs": "",
"buildCommandArgs": "-v",
"ctestCommandArgs": "",
"wslPath": "${defaultWSLPath}",
"inheritEnvironments": [ "linux_clang_x64" ],
"addressSanitizerRuntimeFlags": "detect_leaks=0",
"variables": [
{
"name": "WSL",
"value": "TRUE",
"type": "BOOL"
}
]
},
{
"name": "WSL-Clang-Release",
"generator": "Ninja",
"configurationType": "RelWithDebInfo",
"buildRoot": "${env.BuildPath}",
"installRoot": "${env.InstallPath}",
"cmakeExecutable": "/usr/bin/cmake",
"cmakeCommandArgs": "",
"buildCommandArgs": "-v",
"ctestCommandArgs": "",
"wslPath": "${defaultWSLPath}",
"inheritEnvironments": [ "linux_clang_x64" ],
"addressSanitizerRuntimeFlags": "detect_leaks=0",
"variables": [
{
"name": "WSL",
"value": "TRUE",
"type": "BOOL"
}
]
},
{
"name": "x86-Clang-Cross-Debug",
"generator": "Ninja",
"configurationType": "Debug",
"buildRoot": "${env.BuildPath}",
"installRoot": "${env.InstallPath}",
"cmakeCommandArgs": "",
"buildCommandArgs": "-v",
"ctestCommandArgs": "",
"inheritEnvironments": [ "clang_cl_x86" ],
"variables": [
{
"name": "CMAKE_C_COMPILER",
"value": "${env.cc}",
"type": "STRING"
},
{
"name": "CMAKE_CXX_COMPILER",
"value": "${env.cxx}",
"type": "STRING"
},
{
"name": "CMAKE_SYSROOT",
"value": "${env.fexsysroot}",
"type": "STRING"
}
]
},
{
"name": "x64-Clang-Cross-Release",
"generator": "Ninja",
"configurationType": "RelWithDebInfo",
"buildRoot": "${env.BuildPath}",
"installRoot": "${env.InstallPath}",
"cmakeCommandArgs": "",
"buildCommandArgs": "-v",
"ctestCommandArgs": "",
"inheritEnvironments": [ "clang_cl_x86" ],
"variables": [
{
"name": "CMAKE_C_COMPILER",
"value": "${env.cc}",
"type": "STRING"
},
{
"name": "CMAKE_CXX_COMPILER",
"value": "${env.cxx}",
"type": "STRING"
},
{
"name": "CMAKE_SYSROOT",
"value": "${env.fexsysroot}",
"type": "STRING"
}
]
},
{
"name": "Linux-Clang-Remote-Debug",
"generator": "Ninja",
"configurationType": "Debug",
"cmakeExecutable": "/usr/bin/cmake",
"remoteCopySourcesExclusionList": [ ".vs", ".vscode", ".git", ".github", "build", "out", "bin" ],
"cmakeCommandArgs": "",
"buildCommandArgs": "-v",
"ctestCommandArgs": "",
"inheritEnvironments": [ "linux_clang_x64" ],
"remoteMachineName": "${env.fexremote}",
"remoteCMakeListsRoot": "$HOME/projects/.vs/${projectDirName}/src",
"remoteBuildRoot": "$HOME/projects/.vs/${projectDirName}/build/${name}",
"remoteInstallRoot": "$HOME/projects/.vs/${projectDirName}/install/${name}",
"remoteCopySources": true,
"rsyncCommandArgs": "-t --delete --delete-excluded",
"remoteCopyBuildOutput": false,
"remoteCopySourcesMethod": "rsync",
"addressSanitizerRuntimeFlags": "detect_leaks=0",
"variables": []
}
]
}
+1
View File
@@ -0,0 +1 @@
No AI/ML/LLM/etc code contributions.
+2 -2
View File
@@ -311,7 +311,7 @@ class ExtendedMemOperand final {
public:
ExtendedMemOperand(XRegister rn, XRegister rm = XReg::zr, ExtendedType Option = ExtendedType::LSL_64, uint32_t Shift = 0)
: rn {rn}
, MetaType {.ExtendedType {
, MetaType {.Extended {
.Header = {.MemType = TYPE_EXTENDED},
.rm = rm,
.Option = Option,
@@ -340,7 +340,7 @@ public:
Register rm;
ExtendedType Option;
uint32_t Shift;
} ExtendedType;
} Extended;
struct {
HeaderStruct Header;
IndexType Index;
+50 -50
View File
@@ -3627,8 +3627,8 @@ public:
void strb(ARMEmitter::Register rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
strb(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
strb(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
strb(rt, MemSrc.rn);
} else {
@@ -3650,8 +3650,8 @@ public:
}
void ldrb(ARMEmitter::Register rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrb(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrb(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldrb(rt, MemSrc.rn);
} else {
@@ -3673,8 +3673,8 @@ public:
}
void ldrsb(ARMEmitter::XRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrsb(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrsb(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldrsb(rt, MemSrc.rn);
} else {
@@ -3696,8 +3696,8 @@ public:
}
void ldrsb(ARMEmitter::WRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrsb(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrsb(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldrsb(rt, MemSrc.rn);
} else {
@@ -3719,8 +3719,8 @@ public:
}
void strh(ARMEmitter::Register rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
strh(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
strh(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
strh(rt, MemSrc.rn);
} else {
@@ -3742,8 +3742,8 @@ public:
}
void ldrh(ARMEmitter::Register rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrh(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrh(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldrh(rt, MemSrc.rn);
} else {
@@ -3765,8 +3765,8 @@ public:
}
void ldrsh(ARMEmitter::XRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrsh(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrsh(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldrsh(rt, MemSrc.rn);
} else {
@@ -3788,8 +3788,8 @@ public:
}
void ldrsh(ARMEmitter::WRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrsh(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrsh(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldrsh(rt, MemSrc.rn);
} else {
@@ -3811,8 +3811,8 @@ public:
}
void str(ARMEmitter::WRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
str(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
str(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
str(rt, MemSrc.rn);
} else {
@@ -3834,8 +3834,8 @@ public:
}
void ldr(ARMEmitter::WRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldr(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldr(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldr(rt, MemSrc.rn);
} else {
@@ -3857,8 +3857,8 @@ public:
}
void ldrsw(ARMEmitter::XRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrsw(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrsw(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldrsw(rt, MemSrc.rn);
} else {
@@ -3880,8 +3880,8 @@ public:
}
void str(ARMEmitter::XRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
str(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
str(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
str(rt, MemSrc.rn);
} else {
@@ -3903,8 +3903,8 @@ public:
}
void ldr(ARMEmitter::XRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldr(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldr(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldr(rt, MemSrc.rn);
} else {
@@ -3926,8 +3926,8 @@ public:
}
void prfm(ARMEmitter::Prefetch prfop, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
prfm(prfop, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
prfm(prfop, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
prfm(prfop, MemSrc.rn);
} else {
@@ -3946,9 +3946,9 @@ public:
void strb(ARMEmitter::VRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
LOGMAN_THROW_A_FMT(MemSrc.MetaType.ExtendedType.Shift == false, "Can't shift byte");
strb(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
LOGMAN_THROW_A_FMT(MemSrc.MetaType.Extended.Shift == false, "Can't shift byte");
strb(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
strb(rt, MemSrc.rn);
} else {
@@ -3970,9 +3970,9 @@ public:
}
void ldrb(ARMEmitter::VRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
LOGMAN_THROW_A_FMT(MemSrc.MetaType.ExtendedType.Shift == false, "Can't shift byte");
ldrb(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
LOGMAN_THROW_A_FMT(MemSrc.MetaType.Extended.Shift == false, "Can't shift byte");
ldrb(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldrb(rt, MemSrc.rn);
} else {
@@ -3994,8 +3994,8 @@ public:
}
void strh(ARMEmitter::VRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
strh(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
strh(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
strh(rt, MemSrc.rn);
} else {
@@ -4017,8 +4017,8 @@ public:
}
void ldrh(ARMEmitter::VRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrh(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrh(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldrh(rt, MemSrc.rn);
} else {
@@ -4040,8 +4040,8 @@ public:
}
void str(ARMEmitter::SRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
str(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
str(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
str(rt, MemSrc.rn);
} else {
@@ -4063,8 +4063,8 @@ public:
}
void ldr(ARMEmitter::SRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldr(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldr(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldr(rt, MemSrc.rn);
} else {
@@ -4086,8 +4086,8 @@ public:
}
void str(ARMEmitter::DRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
str(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
str(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
str(rt, MemSrc.rn);
} else {
@@ -4109,8 +4109,8 @@ public:
}
void ldr(ARMEmitter::DRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldr(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldr(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldr(rt, MemSrc.rn);
} else {
@@ -4132,8 +4132,8 @@ public:
}
void str(ARMEmitter::QRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
str(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
str(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
str(rt, MemSrc.rn);
} else {
@@ -4155,8 +4155,8 @@ public:
}
void ldr(ARMEmitter::QRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldr(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldr(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldr(rt, MemSrc.rn);
} else {
+7
View File
@@ -46,6 +46,13 @@
"@PREFIX_LIB@/libwayland-client.so.0",
"@PREFIX_LIB@/libwayland-client.so.0.20.0"
]
},
"cuda": {
"Library" : "libcuda-guest.so",
"Overlay": [
"@PREFIX_LIB@/libcuda.so",
"@PREFIX_LIB@/libcuda.so.1"
]
}
}
}
+138 -90
View File
@@ -1,32 +1,37 @@
#
# This file is autogenerated by pip-compile with Python 3.13
# This file is autogenerated by pip-compile with Python 3.14
# by the following command:
#
# pip-compile --generate-hashes --output-file=requirements_formatting.txt --strip-extras requirements_formatting.txt.in
#
black==25.1.0 \
--hash=sha256:030b9759066a4ee5e5aca28c3c77f9c64789cdd4de8ac1df642c40b708be6171 \
--hash=sha256:055e59b198df7ac0b7efca5ad7ff2516bca343276c466be72eb04a3bcc1f82d7 \
--hash=sha256:0e519ecf93120f34243e6b0054db49c00a35f84f195d5bce7e9f5cfc578fc2da \
--hash=sha256:172b1dbff09f86ce6f4eb8edf9dede08b1fce58ba194c87d7a4f1a5aa2f5b3c2 \
--hash=sha256:1e2978f6df243b155ef5fa7e558a43037c3079093ed5d10fd84c43900f2d8ecc \
--hash=sha256:33496d5cd1222ad73391352b4ae8da15253c5de89b93a80b3e2c8d9a19ec2666 \
--hash=sha256:3b48735872ec535027d979e8dcb20bf4f70b5ac75a8ea99f127c106a7d7aba9f \
--hash=sha256:4b60580e829091e6f9238c848ea6750efed72140b91b048770b64e74fe04908b \
--hash=sha256:759e7ec1e050a15f89b770cefbf91ebee8917aac5c20483bc2d80a6c3a04df32 \
--hash=sha256:8f0b18a02996a836cc9c9c78e5babec10930862827b1b724ddfe98ccf2f2fe4f \
--hash=sha256:95e8176dae143ba9097f351d174fdaf0ccd29efb414b362ae3fd72bf0f710717 \
--hash=sha256:96c1c7cd856bba8e20094e36e0f948718dc688dba4a9d78c3adde52b9e6c2299 \
--hash=sha256:a1ee0a0c330f7b5130ce0caed9936a904793576ef4d2b98c40835d6a65afa6a0 \
--hash=sha256:a22f402b410566e2d1c950708c77ebf5ebd5d0d88a6a2e87c86d9fb48afa0d18 \
--hash=sha256:a39337598244de4bae26475f77dda852ea00a93bd4c728e09eacd827ec929df0 \
--hash=sha256:afebb7098bfbc70037a053b91ae8437c3857482d3a690fefc03e9ff7aa9a5fd3 \
--hash=sha256:bacabb307dca5ebaf9c118d2d2f6903da0d62c9faa82bd21a33eecc319559355 \
--hash=sha256:bce2e264d59c91e52d8000d507eb20a9aca4a778731a08cfff7e5ac4a4bb7096 \
--hash=sha256:d9e6827d563a2c820772b32ce8a42828dc6790f095f441beef18f96aa6f8294e \
--hash=sha256:db8ea9917d6f8fc62abd90d944920d95e73c83a5ee3383493e35d271aca872e9 \
--hash=sha256:ea0213189960bda9cf99be5b8c8ce66bb054af5e9e861249cd23471bd7b0b3ba \
--hash=sha256:f3df5f1bf91d36002b0a75389ca8663510cf0531cca8aa5c1ef695b46d98655f
black==26.3.1 \
--hash=sha256:0126ae5b7c09957da2bdbd91a9ba1207453feada9e9fe51992848658c6c8e01c \
--hash=sha256:0f76ff19ec5297dd8e66eb64deda23631e642c9393ab592826fd4bdc97a4bce7 \
--hash=sha256:28ef38aee69e4b12fda8dba75e21f9b4f979b490c8ac0baa7cb505369ac9e1ff \
--hash=sha256:2bd5aa94fc267d38bb21a70d7410a89f1a1d318841855f698746f8e7f51acd1b \
--hash=sha256:2c50f5063a9641c7eed7795014ba37b0f5fa227f3d408b968936e24bc0566b07 \
--hash=sha256:2d6bfaf7fd0993b420bed691f20f9492d53ce9a2bcccea4b797d34e947318a78 \
--hash=sha256:41cd2012d35b47d589cb8a16faf8a32ef7a336f56356babd9fcf70939ad1897f \
--hash=sha256:474c27574d6d7037c1bc875a81d9be0a9a4f9ee95e62800dab3cfaadbf75acd5 \
--hash=sha256:5602bdb96d52d2d0672f24f6ffe5218795736dd34807fd0fd55ccd6bf206168b \
--hash=sha256:5e9d0d86df21f2e1677cc4bd090cd0e446278bcbbe49bf3659c308c3e402843e \
--hash=sha256:5ed0ca58586c8d9a487352a96b15272b7fa55d139fc8496b519e78023a8dab0a \
--hash=sha256:6c54a4a82e291a1fee5137371ab488866b7c86a3305af4026bdd4dc78642e1ac \
--hash=sha256:6e131579c243c98f35bce64a7e08e87fb2d610544754675d4a0e73a070a5aa3a \
--hash=sha256:855822d90f884905362f602880ed8b5df1b7e3ee7d0db2502d4388a954cc8c54 \
--hash=sha256:86a8b5035fce64f5dcd1b794cf8ec4d31fe458cf6ce3986a30deb434df82a1d2 \
--hash=sha256:8a33d657f3276328ce00e4d37fe70361e1ec7614da5d7b6e78de5426cb56332f \
--hash=sha256:92c0ec1f2cc149551a2b7b47efc32c866406b6891b0ee4625e95967c8f4acfb1 \
--hash=sha256:9a5e9f45e5d5e1c5b5c29b3bd4265dcc90e8b92cf4534520896ed77f791f4da5 \
--hash=sha256:afc622538b430aa4c8c853f7f63bc582b3b8030fd8c80b70fb5fa5b834e575c2 \
--hash=sha256:b07fc0dab849d24a80a29cfab8d8a19187d1c4685d8a5e6385a5ce323c1f015f \
--hash=sha256:b5e6f89631eb88a7302d416594a32faeee9fb8fb848290da9d0a5f2903519fc1 \
--hash=sha256:bf9bf162ed91a26f1adba8efda0b573bc6924ec1408a52cc6f82cb73ec2b142c \
--hash=sha256:c7e72339f841b5a237ff14f7d3880ddd0fc7f98a1199e8c4327f9a4f478c1839 \
--hash=sha256:ddb113db38838eb9f043623ba274cfaf7d51d5b0c22ecb30afe58b1bb8322983 \
--hash=sha256:dfdd51fc3e64ea4f35873d1b3fb25326773d55d2329ff8449139ebaad7357efb \
--hash=sha256:f1cd08e99d2f9317292a311dfe578fd2a24b15dbce97792f9c4d752275c1fa56 \
--hash=sha256:f89f2ab047c76a9c03f78d0d66ca519e389519902fa27e7a91117ef7611c0568
# via
# -r requirements_formatting.txt.in
# darker
@@ -205,56 +210,53 @@ click==8.1.7 \
--hash=sha256:ae74fb96c20a0277a1d615f1e4d73c8414f5a98db8b799a7931d1582f3390c28 \
--hash=sha256:ca9853ad459e787e2192211578cc907e7594e294c7ccc834310722b41b9ca6de
# via black
cryptography==46.0.5 \
--hash=sha256:02f547fce831f5096c9a567fd41bc12ca8f11df260959ecc7c3202555cc47a72 \
--hash=sha256:039917b0dc418bb9f6edce8a906572d69e74bd330b0b3fea4f79dab7f8ddd235 \
--hash=sha256:1abfdb89b41c3be0365328a410baa9df3ff8a9110fb75e7b52e66803ddabc9a9 \
--hash=sha256:2ae6971afd6246710480e3f15824ed3029a60fc16991db250034efd0b9fb4356 \
--hash=sha256:2b7a67c9cd56372f3249b39699f2ad479f6991e62ea15800973b956f4b73e257 \
--hash=sha256:351695ada9ea9618b3500b490ad54c739860883df6c1f555e088eaf25b1bbaad \
--hash=sha256:38946c54b16c885c72c4f59846be9743d699eee2b69b6988e0a00a01f46a61a4 \
--hash=sha256:3b4995dc971c9fb83c25aa44cf45f02ba86f71ee600d81091c2f0cbae116b06c \
--hash=sha256:3ce58ba46e1bc2aac4f7d9290223cead56743fa6ab94a5d53292ffaac6a91614 \
--hash=sha256:3ee190460e2fbe447175cda91b88b84ae8322a104fc27766ad09428754a618ed \
--hash=sha256:4108d4c09fbbf2789d0c926eb4152ae1760d5a2d97612b92d508d96c861e4d31 \
--hash=sha256:420d0e909050490d04359e7fdb5ed7e667ca5c3c402b809ae2563d7e66a92229 \
--hash=sha256:47fb8a66058b80e509c47118ef8a75d14c455e81ac369050f20ba0d23e77fee0 \
--hash=sha256:4c3341037c136030cb46e4b1e17b7418ea4cbd9dd207e4a6f3b2b24e0d4ac731 \
--hash=sha256:4d7e3d356b8cd4ea5aff04f129d5f66ebdc7b6f8eae802b93739ed520c47c79b \
--hash=sha256:4d8ae8659ab18c65ced284993c2265910f6c9e650189d4e3f68445ef82a810e4 \
--hash=sha256:4e817a8920bfbcff8940ecfd60f23d01836408242b30f1a708d93198393a80b4 \
--hash=sha256:50bfb6925eff619c9c023b967d5b77a54e04256c4281b0e21336a130cd7fc263 \
--hash=sha256:556e106ee01aa13484ce9b0239bca667be5004efb0aabbed28d353df86445595 \
--hash=sha256:582f5fcd2afa31622f317f80426a027f30dc792e9c80ffee87b993200ea115f1 \
--hash=sha256:5be7bf2fb40769e05739dd0046e7b26f9d4670badc7b032d6ce4db64dddc0678 \
--hash=sha256:60ee7e19e95104d4c03871d7d7dfb3d22ef8a9b9c6778c94e1c8fcc8365afd48 \
--hash=sha256:61aa400dce22cb001a98014f647dc21cda08f7915ceb95df0c9eaf84b4b6af76 \
--hash=sha256:68f68d13f2e1cb95163fa3b4db4bf9a159a418f5f6e7242564fc75fcae667fd0 \
--hash=sha256:7d1f30a86d2757199cb2d56e48cce14deddf1f9c95f1ef1b64ee91ea43fe2e18 \
--hash=sha256:7d731d4b107030987fd61a7f8ab512b25b53cef8f233a97379ede116f30eb67d \
--hash=sha256:803812e111e75d1aa73690d2facc295eaefd4439be1023fefc4995eaea2af90d \
--hash=sha256:80a8d7bfdf38f87ca30a5391c0c9ce4ed2926918e017c29ddf643d0ed2778ea1 \
--hash=sha256:8293f3dea7fc929ef7240796ba231413afa7b68ce38fd21da2995549f5961981 \
--hash=sha256:8456928655f856c6e1533ff59d5be76578a7157224dbd9ce6872f25055ab9ab7 \
--hash=sha256:890bcb4abd5a2d3f852196437129eb3667d62630333aacc13dfd470fad3aaa82 \
--hash=sha256:94a76daa32eb78d61339aff7952ea819b1734b46f73646a07decb40e5b3448e2 \
--hash=sha256:9f16fbdf4da055efb21c22d81b89f155f02ba420558db21288b3d0035bafd5f4 \
--hash=sha256:a3d1fae9863299076f05cb8a778c467578262fae09f9dc0ee9b12eb4268ce663 \
--hash=sha256:a3d507bb6a513ca96ba84443226af944b0f7f47dcc9a399d110cd6146481d24c \
--hash=sha256:abace499247268e3757271b2f1e244b36b06f8515cf27c4d49468fc9eb16e93d \
--hash=sha256:ba2a27ff02f48193fc4daeadf8ad2590516fa3d0adeeb34336b96f7fa64c1e3a \
--hash=sha256:bc84e875994c3b445871ea7181d424588171efec3e185dced958dad9e001950a \
--hash=sha256:bfd56bb4b37ed4f330b82402f6f435845a5f5648edf1ad497da51a8452d5d62d \
--hash=sha256:c18ff11e86df2e28854939acde2d003f7984f721eba450b56a200ad90eeb0e6b \
--hash=sha256:c3bcce8521d785d510b2aad26ae2c966092b7daa8f45dd8f44734a104dc0bc1a \
--hash=sha256:c4143987a42a2397f2fc3b4d7e3a7d313fbe684f67ff443999e803dd75a76826 \
--hash=sha256:c69fd885df7d089548a42d5ec05be26050ebcd2283d89b3d30676eb32ff87dee \
--hash=sha256:ced80795227d70549a411a4ab66e8ce307899fad2220ce5ab2f296e687eacde9 \
--hash=sha256:d66e421495fdb797610a08f43b05269e0a5ea7f5e652a89bfd5a7d3c1dee3648 \
--hash=sha256:d861ee9e76ace6cf36a6a89b959ec08e7bc2493ee39d07ffe5acb23ef46d27da \
--hash=sha256:e9251e3be159d1020c4030bd2e5f84d6a43fe54b6c19c12f51cde9542a2817b2 \
--hash=sha256:f145bba11b878005c496e93e257c1e88f154d278d2638e6450d17e0f31e558d2 \
--hash=sha256:fe346b143ff9685e40192a4960938545c699054ba11d4f9029f94751e3f71d87
cryptography==49.0.0 \
--hash=sha256:026ac7423e6fa66872d3bf889be5974507da3944f866f704fa200eadacd00001 \
--hash=sha256:07cab27cc7b7e0fd28e5e26bb9eeedde5c135c868b46de4a27845abe94af6122 \
--hash=sha256:084ef1af862eb07ec46d25f68689f2102a9fc0e05ce7b80f14f5fe51e4eef0f6 \
--hash=sha256:0b82e28ee398a386f0807bba7884d30f25218855690f45115831bcce5d90822c \
--hash=sha256:0e959b578856a3924bc0cbb710fc12c387b9412a951389f3ca61704a9e25f325 \
--hash=sha256:0f21641cf4b30fca7aee061ced0ec7ad7b073518088b7c9969a297c0ae796c69 \
--hash=sha256:196ecd6a36e4e9aa10270393bb98d8df88fccee0bf1e5128b91ae4eb4375896d \
--hash=sha256:2400ef9c9e2299a25614eb1dea3db54a69b1349efd043bfac9c67630d136df36 \
--hash=sha256:28d8b15e6275f12c8a207dc309dfa957903c927d08d0cc937ee3f63f200693cc \
--hash=sha256:2afe9051da7ae7bd5905da5a949280c7d2bb75682e188f650a9d0f2756b834c6 \
--hash=sha256:2eda353d8a27bcbcaa4cbed18994a74ab4d19a2ca897db188ea269ab9b71419b \
--hash=sha256:32703d93296f5c1f4b53349ad3a250c2cae0fdecd3a3dd5d47e616d8d616af27 \
--hash=sha256:33cd0565932807baddb67b96dbee92f2c374b5c89dee09fd74079aeb8c8dba61 \
--hash=sha256:35b151772baff2c74cba7fa290ceaff4c3b11c0c881eb93eb5dbc05a7cfbba18 \
--hash=sha256:36d1709f992593689b45bda411498d62c6e365f2ca00b84657d4dadd24de16db \
--hash=sha256:42b0684e0e40cf26122427802486f6d93aea593612603a94fbf260c7eb1e9c1b \
--hash=sha256:4ae387c9cb68ea569ca17e490d66d8142b81c3cc814bf179974b7d146e490bbb \
--hash=sha256:53ecee2e23f7169b6117e99fc8a944e5e50f79e69758a83b52a00cb98ab2b2d2 \
--hash=sha256:66ec79c3904820572d7e987abdf304281f141d37ad9a489b8e97066e7b9b6459 \
--hash=sha256:67e1d20ad9ef3a563c59ef22e7a8a0b8210bd26604369ea4a30a7c66aefe504e \
--hash=sha256:6f2debedf9ca60cf1d5bd466475638af5130f89965605cd818484d19987d3a21 \
--hash=sha256:6fc361c34fb6aac015ce19435876635e5c6d21db31998b0920f675f131e043b8 \
--hash=sha256:73a205dce83953d131a4aa1e0fd917a2fd1c5b1eef251e9d7152efefcbf5caf7 \
--hash=sha256:7abcee80084cda3f7691f3eb1ce480d8df49cec637b429aa35986c1de71738aa \
--hash=sha256:8c25ceb16df5b9435f3f6a9829204985b0e0cbee3b48aacd432c7d2c850b44d9 \
--hash=sha256:966fe0e9c67490071f14c0d2b1cb2dfb3023c5ce39457343931415f08382f2db \
--hash=sha256:9e82dcc8e56052715fb18b2429e3bca4823b1629136a2084fc45a9a5cecb9b64 \
--hash=sha256:b20133d204d2bb56ba047642199603876c872026ca53e79c35b83772ab2cc505 \
--hash=sha256:b39efa323140595abd3ecca8529d321ae50f55f3aa3ba9cc81ea56a6011953d5 \
--hash=sha256:b47db11c2c3525083296069b98ac5221907455e989ae0c2e3008bde851921615 \
--hash=sha256:b87e65d263b3e5d3bb92a57e2a6638e2f31110fa7aa890c7b2dbba42248d0a3f \
--hash=sha256:b970c6da94d5bb18629db453d14f2a1300f6bf59b61e9b82377931ef95504866 \
--hash=sha256:be9fcb48a55f023493482827d4f459bd263cc20efde64f204b97c123201850c6 \
--hash=sha256:c2bc30226390d60ea19d9f82b19db005fe0452154a23c1c410c12ea801e43561 \
--hash=sha256:c83782480a4a9da4d0feb51950131ba32e12e70813848b3343f6e18c28a66838 \
--hash=sha256:cbc77da8c523d5abd028635ba850a6966fcee2c82e2bf65a41d1d8afe0f98be9 \
--hash=sha256:ccac2bfebc306b862133e3bb71f3f6ee8bb525240089b2d952e4144b3a6d5da7 \
--hash=sha256:d0527ce944105f257f605a827d6ebead966c752038b6e8656abb9c5edee6fc68 \
--hash=sha256:d8ecde755e2e91bf773fc94e8c9d730cd7f2007004cb492263a794ec3899a1c8 \
--hash=sha256:e3fb64c420688e5319ae25113a354015abbd8dffbfbc41781a1ea66fc7622ac3 \
--hash=sha256:e5dfc1e64de5677cec922ffa8da89c546d0415bf6efdf081842e5d44c84e1f0e \
--hash=sha256:ec5e529fb80935c94fe7b729f9972b50e351a0e6b50aa294fd5cabb109fcc29a \
--hash=sha256:f37d847238971164fdbc68ade6f6574aecc9c0af714190e2083429ff68f4ce9d \
--hash=sha256:f78ff2c9ed8dc2d036b0f4d640e22522213d047c1b14e61205a7e55c80a494d4 \
--hash=sha256:f89660a348f4f78a92366240a61404e337586ef7f5909a2fef59ca88ef505493 \
--hash=sha256:fc1e275c2f1d97b1a6450b8b0ea3ebfa6e087a611c2b26cb2404d48588abab7b
# via
# -r requirements_formatting.txt.in
# pyjwt
@@ -276,9 +278,9 @@ graylint==1.1.1 \
--hash=sha256:0fd8e02972ca03d0ef2bf0adea76b5343efcd492d7afb5f658f3e3a724f55a36 \
--hash=sha256:b7e0eab6c159684dbf5ef84e942c3340f6a6549b02a3d11b1a1763cc4f8f0593
# via darker
idna==3.10 \
--hash=sha256:12f65c9b470abda6dc35cf8e63cc574b1c52b11df2c86030af0ac09b01b13ea9 \
--hash=sha256:946d195a0d259cbba61165e88e65941f16e9b36ea6ddb97f00452bae8b1287d3
idna==3.16 \
--hash=sha256:cc246e3a3f89580c3a951b5ad298ca4638078b2cdd4f115654332b5c26daded5 \
--hash=sha256:d7a6da03db833450fca25d2358ac9ff06cd624577a4aea3a596d5c0f77b8e03d
# via
# -r requirements_formatting.txt.in
# requests
@@ -290,9 +292,9 @@ packaging==23.1 \
--hash=sha256:994793af429502c4ea2ebf6bf664629d07c1a9fe974af92966e4b8d2df7edc61 \
--hash=sha256:a392980d2b6cffa644431898be54b0045151319d1e7ec34f0cfed48767dd334f
# via black
pathspec==0.11.2 \
--hash=sha256:1d6ed233af05e679efb96b1851550ea95bbb64b7c490b0f5aa52996c11e92a20 \
--hash=sha256:e0d8d0ac2f12da61956eb2306b69f9469b42f4deb0f3cb6ed47b9cce9996ced3
pathspec==1.0.4 \
--hash=sha256:0210e2ae8a21a9137c0d470578cb0e595af87edaa6ebf12ff176f14a02e0e645 \
--hash=sha256:fb6ae2fd4e7c921a165808a552060e722767cfa526f99ca5156ed2ce45a5c723
# via black
platformdirs==3.10.0 \
--hash=sha256:b45696dab2d7cc691a3226759c0d3b00c47c8b6e293d96f6436f733303f77f6d \
@@ -306,10 +308,12 @@ pygithub==2.6.1 \
--hash=sha256:6f2fa6d076ccae475f9fc392cc6cdbd54db985d4f69b8833a28397de75ed6ca3 \
--hash=sha256:b5c035392991cca63959e9453286b41b54d83bf2de2daa7d7ff7e4312cebf3bf
# via -r requirements_formatting.txt.in
pyjwt==2.8.0 \
--hash=sha256:57e28d156e3d5c10088e0c68abb90bfac3df82b40a71bd0daa20c65ccd5c23de \
--hash=sha256:59127c392cc44c2da5bb3192169a91f429924e17aff6534d70fdc02ab3e04320
# via pygithub
pyjwt==2.13.0 \
--hash=sha256:41571c89ca91598c79e8ef18a2d07367d4810fbbd6f637794879baf1b7703423 \
--hash=sha256:66adcc2aff09b3f1bbd95fc1e1577df8ac8723c978552fd43304c8a290ac5728
# via
# -r requirements_formatting.txt.in
# pygithub
pynacl==1.6.2 \
--hash=sha256:018494d6d696ae03c7e656e5e74cdfd8ea1326962cc401bcf018f1ed8436811c \
--hash=sha256:04316d1fc625d860b6c162fff704eb8426b1a8bcd3abacea11142cbd99a6b574 \
@@ -339,9 +343,53 @@ pynacl==1.6.2 \
# via
# -r requirements_formatting.txt.in
# pygithub
requests==2.32.4 \
--hash=sha256:27babd3cda2a6d50b30443204ee89830707d396671944c998b5975b031ac2b2c \
--hash=sha256:27d0316682c8a29834d3264820024b62a36942083d52caf2f14c0591336d3422
pytokens==0.4.1 \
--hash=sha256:0fc71786e629cef478cbf29d7ea1923299181d0699dbe7c3c0f4a583811d9fc1 \
--hash=sha256:11edda0942da80ff58c4408407616a310adecae1ddd22eef8c692fe266fa5009 \
--hash=sha256:140709331e846b728475786df8aeb27d24f48cbcf7bcd449f8de75cae7a45083 \
--hash=sha256:24afde1f53d95348b5a0eb19488661147285ca4dd7ed752bbc3e1c6242a304d1 \
--hash=sha256:26cef14744a8385f35d0e095dc8b3a7583f6c953c2e3d269c7f82484bf5ad2de \
--hash=sha256:27b83ad28825978742beef057bfe406ad6ed524b2d28c252c5de7b4a6dd48fa2 \
--hash=sha256:292052fe80923aae2260c073f822ceba21f3872ced9a68bb7953b348e561179a \
--hash=sha256:29d1d8fb1030af4d231789959f21821ab6325e463f0503a61d204343c9b355d1 \
--hash=sha256:2a44ed93ea23415c54f3face3b65ef2b844d96aeb3455b8a69b3df6beab6acc5 \
--hash=sha256:30f51edd9bb7f85c748979384165601d028b84f7bd13fe14d3e065304093916a \
--hash=sha256:34bcc734bd2f2d5fe3b34e7b3c0116bfb2397f2d9666139988e7a3eb5f7400e3 \
--hash=sha256:3ad72b851e781478366288743198101e5eb34a414f1d5627cdd585ca3b25f1db \
--hash=sha256:3f901fe783e06e48e8cbdc82d631fca8f118333798193e026a50ce1b3757ea68 \
--hash=sha256:42f144f3aafa5d92bad964d471a581651e28b24434d184871bd02e3a0d956037 \
--hash=sha256:4a14d5f5fc78ce85e426aa159489e2d5961acf0e47575e08f35584009178e321 \
--hash=sha256:4a58d057208cb9075c144950d789511220b07636dd2e4708d5645d24de666bdc \
--hash=sha256:4e691d7f5186bd2842c14813f79f8884bb03f5995f0575272009982c5ac6c0f7 \
--hash=sha256:5502408cab1cb18e128570f8d598981c68a50d0cbd7c61312a90507cd3a1276f \
--hash=sha256:584c80c24b078eec1e227079d56dc22ff755e0ba8654d8383b2c549107528918 \
--hash=sha256:5ad948d085ed6c16413eb5fec6b3e02fa00dc29a2534f088d3302c47eb59adf9 \
--hash=sha256:670d286910b531c7b7e3c0b453fd8156f250adb140146d234a82219459b9640c \
--hash=sha256:682fa37ff4d8e95f7df6fe6fe6a431e8ed8e788023c6bcc0f0880a12eab80ad1 \
--hash=sha256:6d6c4268598f762bc8e91f5dbf2ab2f61f7b95bdc07953b602db879b3c8c18e1 \
--hash=sha256:79fc6b8699564e1f9b521582c35435f1bd32dd06822322ec44afdeba666d8cb3 \
--hash=sha256:8bdb9d0ce90cbf99c525e75a2fa415144fd570a1ba987380190e8b786bc6ef9b \
--hash=sha256:8fcb9ba3709ff77e77f1c7022ff11d13553f3c30299a9fe246a166903e9091eb \
--hash=sha256:941d4343bf27b605e9213b26bfa1c4bf197c9c599a9627eb7305b0defcfe40c1 \
--hash=sha256:967cf6e3fd4adf7de8fc73cd3043754ae79c36475c1c11d514fc72cf5490094a \
--hash=sha256:970b08dd6b86058b6dc07efe9e98414f5102974716232d10f32ff39701e841c4 \
--hash=sha256:97f50fd18543be72da51dd505e2ed20d2228c74e0464e4262e4899797803d7fa \
--hash=sha256:9bd7d7f544d362576be74f9d5901a22f317efc20046efe2034dced238cbbfe78 \
--hash=sha256:add8bf86b71a5d9fb5b89f023a80b791e04fba57960aa790cc6125f7f1d39dfe \
--hash=sha256:b35d7e5ad269804f6697727702da3c517bb8a5228afa450ab0fa787732055fc9 \
--hash=sha256:b49750419d300e2b5a3813cf229d4e5a4c728dae470bcc89867a9ad6f25a722d \
--hash=sha256:d31b97b3de0f61571a124a00ffe9a81fb9939146c122c11060725bd5aea79975 \
--hash=sha256:d70e77c55ae8380c91c0c18dea05951482e263982911fc7410b1ffd1dadd3440 \
--hash=sha256:d9907d61f15bf7261d7e775bd5d7ee4d2930e04424bab1972591918497623a16 \
--hash=sha256:da5baeaf7116dced9c6bb76dc31ba04a2dc3695f3d9f74741d7910122b456edc \
--hash=sha256:dc74c035f9bfca0255c1af77ddd2d6ae8419012805453e4b0e7513e17904545d \
--hash=sha256:dcafc12c30dbaf1e2af0490978352e0c4041a7cde31f4f81435c2a5e8b9cabb6 \
--hash=sha256:ee44d0f85b803321710f9239f335aafe16553b39106384cef8e6de40cb4ef2f6 \
--hash=sha256:f66a6bbe741bd431f6d741e617e0f39ec7257ca1f89089593479347cc4d13324
# via black
requests==2.34.2 \
--hash=sha256:2a0d60c172f83ac6ab31e4554906c0f3b3588d37b5cb939b1c061f4907e278e0 \
--hash=sha256:f288924cae4e29463698d6d60bc6a4da69c89185ad1e0bcc4104f584e960b9ed
# via
# -r requirements_formatting.txt.in
# pygithub
@@ -355,9 +403,9 @@ typing-extensions==4.14.1 \
--hash=sha256:38b39f4aeeab64884ce9f74c94263ef78f3c22467c8724005483154c26648d36 \
--hash=sha256:d1e1e3b58374dc93031d6eda2420a48ea44a36c2b4766a4fdeb3710755731d76
# via pygithub
urllib3==2.6.3 \
--hash=sha256:1b62b6884944a57dbe321509ab94fd4d3b307075e0c2eae991ac71ee15ad38ed \
--hash=sha256:bf272323e553dfb2e87d9bfd225ca7b0f467b919d7bbd355436d3fd37cb0acd4
urllib3==2.7.0 \
--hash=sha256:231e0ec3b63ceb14667c67be60f2f2c40a518cb38b03af60abc813da26505f4c \
--hash=sha256:9fb4c81ebbb1ce9531cce37674bbc6f1360472bc18ca9a553ede278ef7276897
# via
# -r requirements_formatting.txt.in
# pygithub
+6 -5
View File
@@ -1,9 +1,10 @@
black~=25.1
black>=26.3.1
darker==2.1.1
PyGithub==2.6.1
cryptography>=46.0.5
urllib3>=2.6.3
requests>=2.32.4
idna>=3.7
cryptography>=48.0.1
urllib3>=2.7.0
requests>=2.33.0
idna>=3.15
certifi>=2024.7.4
PyNaCl>=1.6.2
PyJWT>=2.13.0
+1 -1
+64 -60
View File
@@ -251,6 +251,10 @@ def parse_ops(ops):
if "Desc" in op_val:
OpDef.Desc = op_val["Desc"]
if not isinstance(OpDef.Desc, list):
ExitError(f"Desc field for op {OpDef.Name} must be an array of strings")
if not all(isinstance(item, str) for item in OpDef.Desc):
ExitError(f"Desc field for op {OpDef.Name} must only contain strings")
if "DynamicDispatch" in op_val:
OpDef.DynamicDispatch = bool(op_val["DynamicDispatch"])
@@ -603,77 +607,77 @@ def print_validation(op):
def print_ir_allocator_helpers():
output_file.write("#ifdef IROP_ALLOCATE_HELPERS\n")
output_file.write("\ttemplate <class T>\n")
output_file.write("\tstruct Wrapper final {\n")
output_file.write("\t\tT *first;\n")
output_file.write("\t\tOrderedNode *Node; ///< Actual offset of this IR in ths list\n")
output_file.write("\n")
output_file.write("\t\toperator Wrapper<IROp_Header>() const { return Wrapper<IROp_Header> {reinterpret_cast<IROp_Header*>(first), Node}; }\n")
output_file.write("\t\toperator OrderedNode *() { return Node; }\n")
output_file.write("\t\toperator const OrderedNode *() const { return Node; }\n")
output_file.write("\t\toperator OpNodeWrapper () const { return Node->Header.Value; }\n")
output_file.write("\t};\n")
output_file.write("\ttemplate <class T>\n"
"\tstruct Wrapper final {\n"
"\t\tT *first;\n"
"\t\tOrderedNode *Node; ///< Actual offset of this IR in ths list\n"
"\n"
"\t\toperator Wrapper<IROp_Header>() const { return Wrapper<IROp_Header> {reinterpret_cast<IROp_Header*>(first), Node}; }\n"
"\t\toperator OrderedNode *() { return Node; }\n"
"\t\toperator const OrderedNode *() const { return Node; }\n"
"\t\toperator OpNodeWrapper () const { return Node->Header.Value; }\n"
"\t};\n")
output_file.write("\ttemplate <class T>\n")
output_file.write("\tusing IRPair = Wrapper<T>;\n\n")
output_file.write("\ttemplate <class T>\n"
"\tusing IRPair = Wrapper<T>;\n\n")
output_file.write("\tIRPair<IROp_Header> AllocateRawOp(size_t HeaderSize) {\n")
output_file.write("\t\tauto Op = reinterpret_cast<IROp_Header*>(DualListData.DataAllocate(HeaderSize));\n")
output_file.write("\t\tmemset(Op, 0, HeaderSize);\n")
output_file.write("\t\tOp->Op = IROps::OP_DUMMY;\n")
output_file.write("\t\treturn IRPair<IROp_Header>{Op, CreateNode(Op)};\n")
output_file.write("\t}\n\n")
output_file.write("\tIRPair<IROp_Header> AllocateRawOp(size_t HeaderSize) {\n"
"\t\tauto Op = reinterpret_cast<IROp_Header*>(DualListData.DataAllocate(HeaderSize));\n"
"\t\tmemset(Op, 0, HeaderSize);\n"
"\t\tOp->Op = IROps::OP_DUMMY;\n"
"\t\treturn IRPair<IROp_Header>{Op, CreateNode(Op)};\n"
"\t}\n\n")
output_file.write("\ttemplate<class T, IROps T2>\n")
output_file.write("\tT *AllocateOrphanOp() {\n")
output_file.write("\t\tsize_t Size = FEXCore::IR::GetSize(T2);\n")
output_file.write("\t\tauto Op = reinterpret_cast<T*>(DualListData.DataAllocate(Size));\n")
output_file.write("\t\tmemset(Op, 0, Size);\n")
output_file.write("\t\tOp->Header.Op = T2;\n")
output_file.write("\t\treturn Op;\n")
output_file.write("\t}\n\n")
output_file.write("\ttemplate<class T, IROps T2>\n"
"\tT *AllocateOrphanOp() {\n"
"\t\tsize_t Size = FEXCore::IR::GetSize(T2);\n"
"\t\tauto Op = reinterpret_cast<T*>(DualListData.DataAllocate(Size));\n"
"\t\tmemset(Op, 0, Size);\n"
"\t\tOp->Header.Op = T2;\n"
"\t\treturn Op;\n"
"\t}\n\n")
output_file.write("\ttemplate<class T, IROps T2>\n")
output_file.write("\tIRPair<T> AllocateOp() {\n")
output_file.write("\t\tsize_t Size = FEXCore::IR::GetSize(T2);\n")
output_file.write("\t\tauto Op = reinterpret_cast<T*>(DualListData.DataAllocate(Size));\n")
output_file.write("\t\tmemset(Op, 0, Size);\n")
output_file.write("\t\tOp->Header.Op = T2;\n")
output_file.write("\t\treturn IRPair<T>{Op, CreateNode(&Op->Header)};\n")
output_file.write("\t}\n\n")
output_file.write("\ttemplate<class T, IROps T2>\n"
"\tIRPair<T> AllocateOp() {\n"
"\t\tsize_t Size = FEXCore::IR::GetSize(T2);\n"
"\t\tauto Op = reinterpret_cast<T*>(DualListData.DataAllocate(Size));\n"
"\t\tmemset(Op, 0, Size);\n"
"\t\tOp->Header.Op = T2;\n"
"\t\treturn IRPair<T>{Op, CreateNode(&Op->Header)};\n"
"\t}\n\n")
output_file.write("\tIR::OpSize GetOpSize(const OrderedNode *Op) const {\n")
output_file.write("\t\tauto HeaderOp = Op->Header.Value.GetNode(DualListData.DataBegin());\n")
output_file.write("\t\treturn HeaderOp->Size;\n")
output_file.write("\t}\n\n")
output_file.write("\tIR::OpSize GetOpSize(const OrderedNode *Op) const {\n"
"\t\tauto HeaderOp = Op->Header.Value.GetNode(DualListData.DataBegin());\n"
"\t\treturn HeaderOp->Size;\n"
"\t}\n\n")
output_file.write("\tIR::OpSize GetOpElementSize(const OrderedNode *Op) const {\n")
output_file.write("\t\tauto HeaderOp = Op->Header.Value.GetNode(DualListData.DataBegin());\n")
output_file.write("\t\treturn HeaderOp->ElementSize;\n")
output_file.write("\t}\n\n")
output_file.write("\tIR::OpSize GetOpElementSize(const OrderedNode *Op) const {\n"
"\t\tauto HeaderOp = Op->Header.Value.GetNode(DualListData.DataBegin());\n"
"\t\treturn HeaderOp->ElementSize;\n"
"\t}\n\n")
output_file.write("\tuint8_t GetOpElements(const OrderedNode *Op) const {\n")
output_file.write("\t\tLOGMAN_THROW_A_FMT(OpHasDest(Op), \"Op {} has no dest\\n\", GetOpName(Op));\n")
output_file.write("\t\treturn IR::OpSizeToSize(GetOpSize(Op)) / IR::OpSizeToSize(GetOpElementSize(Op));\n")
output_file.write("\t}\n\n")
output_file.write("\tuint8_t GetOpElements(const OrderedNode *Op) const {\n"
"\t\tLOGMAN_THROW_A_FMT(OpHasDest(Op), \"Op {} has no dest\\n\", GetOpName(Op));\n"
"\t\treturn IR::OpSizeToSize(GetOpSize(Op)) / IR::OpSizeToSize(GetOpElementSize(Op));\n"
"\t}\n\n")
output_file.write("\tbool OpHasDest(const OrderedNode *Op) const {\n")
output_file.write("\t\tauto HeaderOp = Op->Header.Value.GetNode(DualListData.DataBegin());\n")
output_file.write("\t\treturn GetHasDest(HeaderOp->Op);\n")
output_file.write("\t}\n\n")
output_file.write("\tbool OpHasDest(const OrderedNode *Op) const {\n"
"\t\tauto HeaderOp = Op->Header.Value.GetNode(DualListData.DataBegin());\n"
"\t\treturn GetHasDest(HeaderOp->Op);\n"
"\t}\n\n")
output_file.write("\tIROps GetOpType(const OrderedNode *Op) const {\n")
output_file.write("\t\tauto HeaderOp = Op->Header.Value.GetNode(DualListData.DataBegin());\n")
output_file.write("\t\treturn HeaderOp->Op;\n")
output_file.write("\t}\n\n")
output_file.write("\tIROps GetOpType(const OrderedNode *Op) const {\n"
"\t\tauto HeaderOp = Op->Header.Value.GetNode(DualListData.DataBegin());\n"
"\t\treturn HeaderOp->Op;\n"
"\t}\n\n")
output_file.write("\tFEXCore::IR::RegClass GetOpRegClass(const OrderedNode *Op) const {\n")
output_file.write("\t\treturn GetRegClass(GetOpType(Op));\n")
output_file.write("\t}\n\n")
output_file.write("\tFEXCore::IR::RegClass GetOpRegClass(const OrderedNode *Op) const {\n"
"\t\treturn GetRegClass(GetOpType(Op));\n"
"\t}\n\n")
output_file.write("\tstd::string_view const& GetOpName(const OrderedNode *Op) const {\n")
output_file.write("\t\treturn IR::GetName(GetOpType(Op));\n")
output_file.write("\t}\n\n")
output_file.write("\tstd::string_view const& GetOpName(const OrderedNode *Op) const {\n"
"\t\treturn IR::GetName(GetOpType(Op));\n"
"\t}\n\n")
# Generate helpers with operands
for op in IROps:
+8 -1
View File
@@ -6,7 +6,8 @@ set(FEXCORE_BASE_SRCS
Utils/FileLoading.cpp
Utils/ForcedAssert.cpp
Utils/LogManager.cpp
Utils/SpinWaitLock.cpp)
Utils/SpinWaitLock.cpp
Utils/WildcardMatcher.cpp)
if (NOT MINGW)
list(APPEND FEXCORE_BASE_SRCS
@@ -23,6 +24,7 @@ set(SRCS
Interface/Core/Addressing.cpp
Interface/Core/CPUID.cpp
Interface/Core/Frontend.cpp
Interface/Core/SharedCodeBufferManager.cpp
Interface/Core/OpcodeDispatcher/AVX_128.cpp
Interface/Core/OpcodeDispatcher/Crypto.cpp
Interface/Core/OpcodeDispatcher/Flags.cpp
@@ -123,6 +125,11 @@ else()
endif()
endif()
if (CMAKE_CXX_COMPILER_ID STREQUAL "GNU")
# GCC requires libatomic to use 128-bit atomics
list(APPEND LIBS atomic)
endif()
# Generate config
configure_file(${CMAKE_CURRENT_SOURCE_DIR}/Interface/Config/Config.json.in
${CMAKE_BINARY_DIR}/generated/Config/Config.json)
+19 -14
View File
@@ -18,7 +18,7 @@ struct BitSet final {
constexpr static size_t MinimumSize = sizeof(ElementType);
constexpr static size_t MinimumSizeBits = sizeof(ElementType) * 8;
ElementType* Memory;
ElementType* Memory {};
void Allocate(size_t Elements) {
size_t AllocateSize = ToBytes(Elements);
LOGMAN_THROW_A_FMT((AllocateSize * MinimumSize) >= Elements, "Fail");
@@ -33,14 +33,15 @@ struct BitSet final {
FEXCore::Allocator::free(Memory);
Memory = nullptr;
}
bool Get(T Element) {
[[nodiscard]]
bool Get(T Element) const {
return (Memory[Element / MinimumSizeBits] & (1ULL << (Element % MinimumSizeBits))) != 0;
}
void Set(T Element) {
Memory[Element / MinimumSizeBits] |= (1ULL << (Element % MinimumSizeBits));
}
void Clear(T Element) {
Memory[Element / MinimumSizeBits] &= (1ULL << (Element % MinimumSizeBits));
Memory[Element / MinimumSizeBits] &= ~(1ULL << (Element % MinimumSizeBits));
}
void MemClear(size_t Elements) {
memset(Memory, 0, ToBytes(Elements));
@@ -48,13 +49,15 @@ struct BitSet final {
void MemSet(size_t Elements) {
memset(Memory, 0xFF, ToBytes(Elements));
}
uint32_t ToBytes(size_t Elements) {
return AlignUp(Elements, MinimumSizeBits) / MinimumSize;
[[nodiscard]]
static size_t ToBytes(size_t Elements) {
return AlignUp(Elements, MinimumSizeBits) / 8;
}
// This very explicitly doesn't let you take an address
// Is only a getter
bool operator[](T Element) {
[[nodiscard]]
bool operator[](T Element) const {
return Get(Element);
}
};
@@ -62,35 +65,37 @@ struct BitSet final {
template<typename T>
struct BitSetView final {
using ElementType = T;
constexpr static size_t MinimumSize = sizeof(ElementType);
constexpr static size_t MinimumSizeBits = sizeof(ElementType) * 8;
constexpr static size_t MinimumSize = BitSet<T>::MinimumSize;
constexpr static size_t MinimumSizeBits = BitSet<T>::MinimumSizeBits;
ElementType* Memory;
ElementType* Memory {};
void GetView(BitSet<T>& Set, uint64_t ElementOffset) {
LOGMAN_THROW_A_FMT((ElementOffset % MinimumSize) == 0, "Bitset view offset needs to be aligned to size of backing element");
Memory = &Set.Memory[ElementOffset / MinimumSizeBits];
}
bool Get(T Element) {
[[nodiscard]]
bool Get(T Element) const {
return (Memory[Element / MinimumSizeBits] & (1ULL << (Element % MinimumSizeBits))) != 0;
}
void Set(T Element) {
Memory[Element / MinimumSizeBits] |= (1ULL << (Element % MinimumSizeBits));
}
void Clear(T Element) {
Memory[Element / MinimumSizeBits] &= (1ULL << (Element % MinimumSizeBits));
Memory[Element / MinimumSizeBits] &= ~(1ULL << (Element % MinimumSizeBits));
}
void MemClear(size_t Elements) {
memset(Memory, 0, AlignUp(Elements / MinimumSizeBits, MinimumSizeBits));
memset(Memory, 0, BitSet<T>::ToBytes(Elements));
}
void MemSet(size_t Elements) {
memset(Memory, 0xFF, AlignUp(Elements / MinimumSizeBits, MinimumSizeBits));
memset(Memory, 0xFF, BitSet<T>::ToBytes(Elements));
}
// This very explicitly doesn't let you take an address
// Is only a getter
bool operator[](T Element) {
[[nodiscard]]
bool operator[](T Element) const {
return Get(Element);
}
};
+19
View File
@@ -260,6 +260,10 @@ struct FEX_PACKED X80SoftFloat {
if (lhs.Top.Exponent == 0x0 && lhs.Significand == 0x0) {
return lhs;
}
// Inf/NaN pass through unchanged in the significand slot.
if (lhs.Top.Exponent == 0x7FFF) {
return lhs;
}
X80SoftFloat Tmp = lhs;
Tmp.Top.Exponent = 0x3FFF;
Tmp.Top.Sign = lhs.Top.Sign;
@@ -288,6 +292,14 @@ struct FEX_PACKED X80SoftFloat {
X80SoftFloat Result(1, 0x7FFFUL, 0x8000'0000'0000'0000UL);
return Result;
}
// +/-Inf returns +Inf in the exponent slot; NaN propagates.
if (lhs.Top.Exponent == 0x7FFF) {
if ((lhs.Significand & 0x7FFFFFFFFFFFFFFFULL) == 0) {
X80SoftFloat Result(0, 0x7FFFUL, 0x8000'0000'0000'0000UL);
return Result;
}
return lhs;
}
int32_t TrueExp = lhs.Top.Exponent - ExponentBias;
return i32_to_extF80(TrueExp);
@@ -324,6 +336,13 @@ struct FEX_PACKED X80SoftFloat {
#else
extFloat80_t Zero {0, 0};
if (extF80_eq(state, lhs, Zero)) {
// FSCALE(0, +Inf) is 0 * Inf, which is invalid. FSCALE(0, anything
// else) is still 0.
if (rhs.Top.Exponent == 0x7FFF && rhs.Top.Sign == 0 && (rhs.Significand & 0x7FFFFFFFFFFFFFFFULL) == 0) {
state->exceptionFlags |= softfloat_flag_invalid;
X80SoftFloat QNaN(0, 0x7FFFUL, 0xC000000000000000ULL);
return QNaN;
}
return lhs;
}
X80SoftFloat Int = FRNDINT(state, rhs, softfloat_round_minMag);
+1
View File
@@ -4,6 +4,7 @@
#include <concepts>
#include <string_view>
#include <cstdlib>
namespace FEXCore::StrConv {
template<std::integral T>
+4
View File
@@ -16,7 +16,11 @@ struct VectorScalarF64Pair {
#ifdef ARCHITECTURE_arm64
// Can't use uint8x16_t directly from arm_neon.h here.
// Overrides softfloat-3e's defines which causes problems.
#ifdef __clang__
using VectorRegType = __attribute__((neon_vector_type(16))) uint8_t;
#else
using VectorRegType = __attribute__((vector_size(16))) uint8_t;
#endif
struct VectorRegPairType {
VectorRegType val[2];
};
+3 -2
View File
@@ -1,9 +1,10 @@
// SPDX-License-Identifier: MIT
#include "Common/StringConv.h"
#include "FEXCore/Utils/EnumUtils.h"
#include "Utils/Config.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/Utils/Allocator.h>
#include <FEXCore/Utils/EnumUtils.h>
#include <FEXCore/Utils/FileLoading.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/StringUtils.h>
@@ -252,7 +253,7 @@ void Load() {
}
}
fextl::string ExpandPath(const fextl::string& ContainerPrefix, const fextl::string& PathName) {
static fextl::string ExpandPath(const fextl::string& ContainerPrefix, const fextl::string& PathName) {
if (PathName.empty()) {
return {};
}
+14 -4
View File
@@ -23,6 +23,13 @@
"Enable the code caching subsystem"
]
},
"EnableLazyCodeCachingWIP": {
"Type": "bool",
"Default": "false",
"Desc": [
"Enable lazy loading of chunks in code caches"
]
},
"EnableCodeCacheValidation": {
"Type": "bool",
"Default": "false",
@@ -75,7 +82,9 @@
"ENABLE3DNOW": "enable3dnow",
"DISABLE3DNOW": "disable3dnow",
"ENABLESSE4A": "enablesse4a",
"DISABLESSE4A": "disablesse4a"
"DISABLESSE4A": "disablesse4a",
"ENABLEMOPS": "enablemops",
"DISABLEMOPS": "disablemops"
},
"Desc": [
"Allows controlling of the CPU features in the JIT.",
@@ -99,7 +108,8 @@
"\t{enable,disable}preserveallabi: Will force enable or disable preserve_all abi even if the host doesn't support it",
"\t{enable,disable}wfxt: Will force enable or disable wfxt even if the host doesn't support it",
"\t{enable,disable}3dnow: Will force enable or disable 3DNow! even if the host doesn't support it",
"\t{enable,disable}sse4a: Will force enable or disable SSE4a even if the host doesn't support it"
"\t{enable,disable}sse4a: Will force enable or disable SSE4a even if the host doesn't support it",
"\t{enable,disable}mops: Will force enable or disable FEAT_MOPS even if the host doesn't support it"
]
},
"SmallTSCScale": {
@@ -192,7 +202,7 @@
},
"DisableL2Cache": {
"Type": "bool",
"Default": "false",
"Default": "true",
"Desc": [
"Disables FEXCore's JIT L2 cache lookup. Saving memory.",
"Can potentially introduce more stutters."
@@ -200,7 +210,7 @@
},
"DynamicL1Cache": {
"Type": "bool",
"Default": "false",
"Default": "true",
"Desc": [
"Switches FEXCore's JIT L1 cache to be dynamically sized. Saving memory.",
"Can potentially introduce more stutters."
+7 -1
View File
@@ -53,6 +53,12 @@ FEXCore::CPUID::FunctionResults FEXCore::Context::ContextImpl::RunCPUIDFunctionN
}
bool FEXCore::Context::ContextImpl::IsAddressInCodeBuffer(FEXCore::Core::InternalThreadState* Thread, uintptr_t Address) const {
return Thread->CPUBackend->IsAddressInCodeBuffer(Address);
return Thread->CPUBackend->IsAddressInCodeBuffer(Address) || CodeCache.IsAddressInMappedCodeBuffer(Address);
}
bool FEXCore::Context::ContextImpl::RequiresRelocatableConstants() const {
// Support relocation when generating a cache or when generating reference code for validation
return CodeCache.IsGeneratingCache || FEXCore::Config::Get_ENABLECODECACHEVALIDATION();
}
} // namespace FEXCore::Context
+109 -5
View File
@@ -4,6 +4,7 @@
#include "Common/JitSymbols.h"
#include "Interface/Core/CPUBackend.h"
#include "Interface/Core/CPUID.h"
#include "Interface/Core/SharedCodeBufferManager.h"
#include <Interface/IR/IntrusiveIRList.h>
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/Context.h>
@@ -64,6 +65,8 @@ struct CustomIRResult {
using BlockDelinkerFunc = void (*)(FEXCore::Context::ExitFunctionLinkData* Record);
constexpr uint32_t TSC_SCALE_MAXIMUM = 1'000'000'000; ///< 1Ghz
constexpr static bool BLOCK_DEBUGGING = false;
class CodeCache : public AbstractCodeCache {
public:
CodeCache(ContextImpl&);
@@ -76,11 +79,17 @@ public:
bool IsGeneratingCache = false;
FEX_CONFIG_OPT(EnableCodeCaching, ENABLECODECACHINGWIP);
FEX_CONFIG_OPT(EnableLazyCodeCaching, ENABLELAZYCODECACHINGWIP);
FEX_CONFIG_OPT(EnableCodeCacheValidation, ENABLECODECACHEVALIDATION);
uint64_t ComputeCodeMapId(std::string_view Filename, int FD) override;
bool SaveData(Core::InternalThreadState&, int TargetFD, const ExecutableFileSectionInfo&, uint64_t SerializedBaseAddress) override;
bool LoadData(Core::InternalThreadState*, std::byte* MappedCacheFile, const ExecutableFileSectionInfo&) override;
fextl::unique_ptr<MappedCodeCacheFile> LoadCache(std::span<std::byte> CacheFile, const ExecutableFileInfo&, uint64_t FileStartVA) override;
bool EnableLoadedSection(Core::InternalThreadState*, MappedCodeCacheFile&, const ExecutableFileSectionInfo&) override;
void FinalizeCodePages(MappedCodeCacheFile&, std::span<std::byte> CodeRange) override;
/**
* Performs expensive extra validation on the loaded code cache data.
@@ -112,21 +121,24 @@ public:
* Note that FEX relocations are unrelated to ELF/PE relocations.
*
* @param GuestDelta Guest address offset to apply to RIP-relative data
* @param RelocationOffset Offset to subtract from relocation target offsets
* @param ForStorage True for serializing data (producing deterministic output); false for de-serializing it (resolving dynamic symbols)
*
* @return Returns true on success
*/
[[nodiscard]]
bool ApplyCodeRelocations(uint64_t GuestDelta, std::span<std::byte> Code, std::span<const CPU::Relocation> Relocations, bool ForStorage);
bool ApplyCodeRelocations(uint64_t GuestDelta, std::span<std::byte> Code, std::span<const CPU::Relocation> Relocations,
uint32_t RelocationOffset, bool ForStorage);
};
class ContextImpl final : public FEXCore::Context::Context, public CPU::CodeBufferManager {
class ContextImpl final : public FEXCore::Context::Context, public CPU::SharedCodeBufferManager {
public:
// Context base class implementation.
bool InitCore() override;
void ExecuteThread(FEXCore::Core::InternalThreadState* Thread) override;
bool CheckIfBlockIsCacheable(FEXCore::Core::InternalThreadState&, uint64_t GuestRIP, uint64_t MaxInst) override;
void CompileRIP(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP) override;
void CompileRIPCount(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP, uint64_t MaxInst) override;
@@ -208,7 +220,7 @@ public:
void ClearCodeCache(FEXCore::Core::InternalThreadState* Thread, bool NewCodeBuffer = true) override;
void InvalidateCodeBuffersCodeRange(uint64_t Start, uint64_t Length) override;
void InvalidateThreadCachedCodeRange(FEXCore::Core::InternalThreadState* Thread, uint64_t Start, uint64_t Length) override;
FEXCore::ForkableSharedMutex& GetCodeInvalidationMutex() override {
FEXCore::Utils::WritePriorityMutex::Mutex& GetCodeInvalidationMutex() override {
return CodeInvalidationMutex;
}
@@ -232,6 +244,96 @@ public:
void MarkMonoBackpatcherBlock(uint64_t BlockEntry) override;
// Manual debugging tooling which is useful for developers.
struct TrackingEmpty {
// RIP stepping handling
virtual void AddSingleStepTarget(uint64_t GuestRIP) {}
virtual void AddSingleStepTargetRange(uint64_t RIPBegin, uint64_t RipEnd) {}
virtual void AllTargetSingleStep() {}
virtual void RemoveSingleStepTarget(uint64_t GuestRIP) {}
virtual bool IsSingleStepTarget(uint64_t GuestRIP) {
return false;
}
// Watchpoints
virtual void AddWriteWatchPoint(uint64_t Ptr) {}
virtual void AddReadWatchPoint(uint64_t Ptr) {}
virtual bool ContainsWriteWatchPoint(uint64_t Ptr, size_t Size) {
return false;
}
virtual bool ContainsReadWatchPoint(uint64_t Ptr, size_t Size) {
return false;
}
};
struct TrackingPossible final : public TrackingEmpty {
void AddSingleStepTarget(uint64_t GuestRIP) override {
SingleStepTargets.emplace(GuestRIP);
}
virtual void AddSingleStepTargetRange(uint64_t RIPBegin, uint64_t RIPEnd) override {
SingleStepRanges.emplace_back(Range {RIPBegin, RIPEnd});
}
void RemoveSingleStepTarget(uint64_t GuestRIP) override {
SingleStepTargets.erase(GuestRIP);
}
void AllTargetSingleStep() override {
SingleStepEverything = true;
}
bool IsSingleStepTarget(uint64_t GuestRIP) override {
return SingleStepEverything || SingleStepTargets.contains(GuestRIP) || IsInRange(GuestRIP);
}
void AddWriteWatchPoint(uint64_t Ptr) override {
WatchWriteTargets.emplace(Ptr);
}
void AddReadWatchPoint(uint64_t Ptr) override {
WatchReadTargets.emplace(Ptr);
}
bool ContainsWriteWatchPoint(uint64_t Ptr, size_t Size) override {
return ContainsRange(WatchWriteTargets, Ptr, Size);
}
bool ContainsReadWatchPoint(uint64_t Ptr, size_t Size) override {
return ContainsRange(WatchReadTargets, Ptr, Size);
}
private:
bool SingleStepEverything {};
fextl::set<uint64_t> SingleStepTargets {};
fextl::set<uint64_t> WatchWriteTargets {};
fextl::set<uint64_t> WatchReadTargets {};
struct Range {
uint64_t Begin, End;
};
fextl::vector<Range> SingleStepRanges {};
bool IsInRange(uint64_t RIP) const {
return std::ranges::any_of(SingleStepRanges, [RIP](const auto& range) { return RIP >= range.Begin && RIP <= range.End; });
}
static bool ContainsRange(const fextl::set<uint64_t>& Set, uint64_t Ptr, size_t Size) {
for (auto it = Set.lower_bound(Ptr); it != Set.end(); --it) {
auto Watch = *it;
if (Watch < Ptr) {
break;
}
if (Watch >= Ptr && Watch < (Ptr + Size)) {
return true;
}
}
return false;
}
};
using TrackingStructure = std::conditional<BLOCK_DEBUGGING, TrackingPossible, TrackingEmpty>::type;
TrackingStructure BlockDebuggerTracker {};
public:
struct {
uint64_t VirtualMemSize {1ULL << 36};
@@ -262,7 +364,7 @@ public:
FEX_CONFIG_OPT(MonoHacks, MONOHACKS);
} Config;
FEXCore::ForkableSharedMutex CodeInvalidationMutex;
FEXCore::Utils::WritePriorityMutex::Mutex CodeInvalidationMutex {};
uint32_t StrictSplitLockMutex {};
@@ -350,6 +452,8 @@ public:
return Config.MonoHacks && MonoDetected;
}
bool RequiresRelocatableConstants() const;
protected:
void UpdateAtomicTSOEmulationConfig() {
if (SupportsHardwareTSO) {
@@ -360,6 +360,7 @@ namespace x32 {
Arm64Emitter::Arm64Emitter(FEXCore::Context::ContextImpl* ctx, void* EmissionPtr, size_t size)
: Emitter(static_cast<uint8_t*>(EmissionPtr), size)
, EmitterCTX {ctx}
, SupportCodeRelocations {ctx->RequiresRelocatableConstants()}
#ifdef VIXL_SIMULATOR
, Simulator {&SimDecoder, stdout, vixl::aarch64::SimStack(SimulatorStackSize).Allocate()}
#endif
@@ -425,7 +426,7 @@ void Arm64Emitter::LoadConstant(ARMEmitter::Size s, ARMEmitter::Register Reg, ui
NOPPad = false;
} else if (Pad == PadType::AUTOPAD) {
// Force NOP padding to ensure relocated constants always have enough encoding space available
NOPPad = EnableCodeCaching;
NOPPad = SupportCodeRelocations;
}
bool Is64Bit = s == ARMEmitter::Size::i64Bit;
@@ -448,8 +449,10 @@ void Arm64Emitter::LoadConstant(ARMEmitter::Size s, ARMEmitter::Register Reg, ui
return;
}
if ((Constant >> 32) == 0) {
if ((Constant >> 32) == 0 && !NOPPad) {
// If the upper 32-bits is all zero, we can now switch to a 32-bit move.
// NOTE: The NOP padding code does not appropriately adjust to this yet,
// so we skip this optimization in that case
s = ARMEmitter::Size::i32Bit;
Is64Bit = false;
Segments = std::min(Segments, 2);
@@ -584,8 +587,8 @@ void Arm64Emitter::PushCalleeSavedRegisters() {
{ARMEmitter::XReg::x29, ARMEmitter::XReg::x30},
}};
for (auto& RegPair : CalleeSaved) {
stp<ARMEmitter::IndexType::PRE>(RegPair.first, RegPair.second, ARMEmitter::Reg::rsp, -16);
for (const auto& [rt, rt2] : CalleeSaved) {
stp<ARMEmitter::IndexType::PRE>(rt, rt2, ARMEmitter::Reg::rsp, -16);
}
// Additionally we need to store the lower 64bits of v8-v15
@@ -602,9 +605,8 @@ void Arm64Emitter::PushCalleeSavedRegisters() {
// We just saved x19 so it is safe
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r19, ARMEmitter::Reg::rsp, 0);
for (auto& RegQuad : FPRs) {
st4(ARMEmitter::SubRegSize::i64Bit, std::get<0>(RegQuad), std::get<1>(RegQuad), std::get<2>(RegQuad), std::get<3>(RegQuad), 0,
ARMEmitter::Reg::r19, 32);
for (const auto& [rt, rt2, rt3, rt4] : FPRs) {
st4(ARMEmitter::SubRegSize::i64Bit, rt, rt2, rt3, rt4, 0, ARMEmitter::Reg::r19, 32);
}
}
@@ -614,9 +616,8 @@ void Arm64Emitter::PopCalleeSavedRegisters() {
{ARMEmitter::DReg::d12, ARMEmitter::DReg::d13, ARMEmitter::DReg::d14, ARMEmitter::DReg::d15},
}};
for (auto& RegQuad : FPRs) {
ld4(ARMEmitter::SubRegSize::i64Bit, std::get<0>(RegQuad), std::get<1>(RegQuad), std::get<2>(RegQuad), std::get<3>(RegQuad), 0,
ARMEmitter::Reg::rsp, 32);
for (const auto& [rt, rt2, rt3, rt4] : FPRs) {
ld4(ARMEmitter::SubRegSize::i64Bit, rt, rt2, rt3, rt4, 0, ARMEmitter::Reg::rsp, 32);
}
constexpr static std::array<std::pair<ARMEmitter::XRegister, ARMEmitter::XRegister>, 6> CalleeSaved = {{
@@ -628,12 +629,12 @@ void Arm64Emitter::PopCalleeSavedRegisters() {
{ARMEmitter::XReg::x19, ARMEmitter::XReg::x20},
}};
for (auto& RegPair : CalleeSaved) {
ldp<ARMEmitter::IndexType::POST>(RegPair.first, RegPair.second, ARMEmitter::Reg::rsp, 16);
for (const auto& [rt, rt2] : CalleeSaved) {
ldp<ARMEmitter::IndexType::POST>(rt, rt2, ARMEmitter::Reg::rsp, 16);
}
}
void Arm64Emitter::FillSpecialRegs(ARMEmitter::Register TmpReg, ARMEmitter::Register TmpReg2, bool SetFIZ, bool SetPredRegs) {
void Arm64Emitter::FillSpecialRegs(ARMEmitter::Register TmpReg, ARMEmitter::Register TmpReg2, const FillSpecialRegsOptions& Options) {
#ifndef VIXL_SIMULATOR
if (EmitterCTX->HostFeatures.SupportsAFP) {
// Enable AFP features when filling JIT state.
@@ -649,7 +650,7 @@ void Arm64Emitter::FillSpecialRegs(ARMEmitter::Register TmpReg, ARMEmitter::Regi
(1U << 2) | // NEP
(1U << 1)); // AH
if (SetFIZ) {
if (Options.SetFIZ) {
// Insert MXCSR.DAZ in to FIZ
ldr(TmpReg2.W(), STATE.R(), offsetof(FEXCore::Core::CPUState, mxcsr));
bfxil(ARMEmitter::Size::i64Bit, TmpReg, TmpReg2, 6, 1);
@@ -659,7 +660,7 @@ void Arm64Emitter::FillSpecialRegs(ARMEmitter::Register TmpReg, ARMEmitter::Regi
}
#endif
if (SetPredRegs && (EmitterCTX->HostFeatures.SupportsSVE256 || EmitterCTX->HostFeatures.SupportsSVE128)) {
if (Options.SetPredRegs && EmitterCTX->HostFeatures.SupportsSVE()) {
// Set up predicate registers.
// We don't bother spilling these in SpillStaticRegs,
// since all that matters is we restore them on a fill.
@@ -677,7 +678,7 @@ void Arm64Emitter::FillSpecialRegs(ARMEmitter::Register TmpReg, ARMEmitter::Regi
}
}
void Arm64Emitter::SpillStaticRegs(ARMEmitter::Register TmpReg, bool FPRs, uint32_t GPRSpillMask, uint32_t FPRSpillMask, bool NZCV) {
void Arm64Emitter::SpillStaticRegs(ARMEmitter::Register TmpReg, SpillStaticRegOptions Options) {
#ifndef VIXL_SIMULATOR
if (EmitterCTX->HostFeatures.SupportsAFP) {
// Disable AFP features when spilling registers.
@@ -698,7 +699,7 @@ void Arm64Emitter::SpillStaticRegs(ARMEmitter::Register TmpReg, bool FPRs, uint3
}
#endif
if (NZCV) {
if (Options.NZCV) {
// Regardless of what GPRs/FPRs we're spilling, we need to spill NZCV since it
// is always static and almost certainly clobbered by the subsequent code.
//
@@ -710,25 +711,25 @@ void Arm64Emitter::SpillStaticRegs(ARMEmitter::Register TmpReg, bool FPRs, uint3
// PF/AF are special, remove them from the mask
uint32_t PFAFMask = ((1u << REG_PF.Idx()) | ((1u << REG_AF.Idx())));
unsigned PFAFSpillMask = GPRSpillMask & PFAFMask;
GPRSpillMask &= ~PFAFSpillMask;
unsigned PFAFSpillMask = Options.GPRSpillMask & PFAFMask;
Options.GPRSpillMask &= ~PFAFSpillMask;
str(REG_CALLRET_SP, STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.callret_sp));
for (size_t i = 0; i < StaticRegisters.size(); i += 2) {
auto Reg1 = StaticRegisters[i];
auto Reg2 = StaticRegisters[i + 1];
if (((1U << Reg1.Idx()) & GPRSpillMask) && ((1U << Reg2.Idx()) & GPRSpillMask)) {
stp<ARMEmitter::IndexType::OFFSET>(Reg1.X(), Reg2.X(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i]));
} else if (((1U << Reg1.Idx()) & GPRSpillMask)) {
str(Reg1.X(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i]));
} else if (((1U << Reg2.Idx()) & GPRSpillMask)) {
str(Reg2.X(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i + 1]));
if (((1U << Reg1.Idx()) & Options.GPRSpillMask) && ((1U << Reg2.Idx()) & Options.GPRSpillMask)) {
stp<ARMEmitter::IndexType::OFFSET>(Reg1.X(), Reg2.X(), STATE.R(), ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, State.gregs, i));
} else if (((1U << Reg1.Idx()) & Options.GPRSpillMask)) {
str(Reg1.X(), STATE.R(), ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, State.gregs, i));
} else if (((1U << Reg2.Idx()) & Options.GPRSpillMask)) {
str(Reg2.X(), STATE.R(), ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, State.gregs, i + 1));
}
}
// Now handle PF/AF
if (NZCV && PFAFSpillMask) {
if (Options.NZCV && PFAFSpillMask) {
auto PFOffset = offsetof(FEXCore::Core::CpuStateFrame, State.pf_raw);
auto AFOffset = offsetof(FEXCore::Core::CpuStateFrame, State.af_raw);
LOGMAN_THROW_A_FMT(PFAFSpillMask == PFAFMask, "PF/AF not spilled together");
@@ -737,21 +738,21 @@ void Arm64Emitter::SpillStaticRegs(ARMEmitter::Register TmpReg, bool FPRs, uint3
stp<ARMEmitter::IndexType::OFFSET>(REG_PF.W(), REG_AF.W(), STATE.R(), PFOffset);
}
if (FPRs) {
if (Options.FPRs) {
if (EmitterCTX->HostFeatures.SupportsAVX && EmitterCTX->HostFeatures.SupportsSVE256) {
for (size_t i = 0; i < StaticFPRegisters.size(); i++) {
const auto Reg = StaticFPRegisters[i];
if (((1U << Reg.Idx()) & FPRSpillMask) != 0) {
mov(ARMEmitter::Size::i64Bit, TmpReg, offsetof(Core::CpuStateFrame, State.xmm.avx.data[i][0]));
if (((1U << Reg.Idx()) & Options.FPRSpillMask) != 0) {
mov(ARMEmitter::Size::i64Bit, TmpReg, ARRAY_OFFSETOF(Core::CpuStateFrame, State.xmm.avx.data, i));
st1b<ARMEmitter::SubRegSize::i8Bit>(Reg.Z(), PRED_TMP_32B, STATE.R(), TmpReg);
}
}
} else {
if (GPRSpillMask && FPRSpillMask == ~0U) {
if (Options.GPRSpillMask && Options.FPRSpillMask == ~0U) {
// Optimize the common case where we can spill four registers per instruction
// Load the sse offset in to the temporary register
add(ARMEmitter::Size::i64Bit, TmpReg, STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[0][0]));
add(ARMEmitter::Size::i64Bit, TmpReg, STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data));
for (size_t i = 0; i < StaticFPRegisters.size(); i += 4) {
const auto Reg1 = StaticFPRegisters[i];
const auto Reg2 = StaticFPRegisters[i + 1];
@@ -764,12 +765,12 @@ void Arm64Emitter::SpillStaticRegs(ARMEmitter::Register TmpReg, bool FPRs, uint3
const auto Reg1 = StaticFPRegisters[i];
const auto Reg2 = StaticFPRegisters[i + 1];
if (((1U << Reg1.Idx()) & FPRSpillMask) && ((1U << Reg2.Idx()) & FPRSpillMask)) {
stp<ARMEmitter::IndexType::OFFSET>(Reg1.Q(), Reg2.Q(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[i][0]));
} else if (((1U << Reg1.Idx()) & FPRSpillMask)) {
str(Reg1.Q(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[i][0]));
} else if (((1U << Reg2.Idx()) & FPRSpillMask)) {
str(Reg2.Q(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[i + 1][0]));
if (((1U << Reg1.Idx()) & Options.FPRSpillMask) && ((1U << Reg2.Idx()) & Options.FPRSpillMask)) {
stp<ARMEmitter::IndexType::OFFSET>(Reg1.Q(), Reg2.Q(), STATE.R(), ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, State.xmm.sse.data, i));
} else if (((1U << Reg1.Idx()) & Options.FPRSpillMask)) {
str(Reg1.Q(), STATE.R(), ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, State.xmm.sse.data, i));
} else if (((1U << Reg2.Idx()) & Options.FPRSpillMask)) {
str(Reg2.Q(), STATE.R(), ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, State.xmm.sse.data, i + 1));
}
}
}
@@ -777,8 +778,7 @@ void Arm64Emitter::SpillStaticRegs(ARMEmitter::Register TmpReg, bool FPRs, uint3
}
}
void Arm64Emitter::FillStaticRegs(bool FPRs, uint32_t GPRFillMask, uint32_t FPRFillMask, std::optional<ARMEmitter::Register> OptionalReg,
std::optional<ARMEmitter::Register> OptionalReg2, bool NZCV) {
void Arm64Emitter::FillStaticRegs(FillStaticRegOptions Options) {
auto FindTempReg = [this](uint32_t* GPRFillMask) -> std::optional<ARMEmitter::Register> {
for (auto Reg : StaticRegisters) {
if (((1U << Reg.Idx()) & *GPRFillMask)) {
@@ -789,20 +789,21 @@ void Arm64Emitter::FillStaticRegs(bool FPRs, uint32_t GPRFillMask, uint32_t FPRF
return std::nullopt;
};
LOGMAN_THROW_A_FMT(GPRFillMask != 0, "Must fill at least 2 GPRs for a temp");
uint32_t TempGPRFillMask = GPRFillMask;
if (!OptionalReg.has_value()) {
OptionalReg = FindTempReg(&TempGPRFillMask);
LOGMAN_THROW_A_FMT(Options.GPRFillMask != 0, "Must fill at least 2 GPRs for a temp");
uint32_t TempGPRFillMask = Options.GPRFillMask;
if (!Options.OptionalReg.has_value()) {
Options.OptionalReg = FindTempReg(&TempGPRFillMask);
}
if (!OptionalReg2.has_value()) {
OptionalReg2 = FindTempReg(&TempGPRFillMask);
if (!Options.OptionalReg2.has_value()) {
Options.OptionalReg2 = FindTempReg(&TempGPRFillMask);
}
LOGMAN_THROW_A_FMT(OptionalReg.has_value() && OptionalReg2.has_value(), "Didn't have an SRA register to use as a temporary while "
"spilling!");
LOGMAN_THROW_A_FMT(Options.OptionalReg.has_value() && Options.OptionalReg2.has_value(), "Didn't have an SRA register to use as a "
"temporary while "
"spilling!");
auto TmpReg = *OptionalReg;
auto TmpReg2 = *OptionalReg2;
auto TmpReg = *Options.OptionalReg;
auto TmpReg2 = *Options.OptionalReg2;
#ifdef ARCHITECTURE_arm64ec
// Load STATE in from the CPU area as x28 is not callee saved in the ARM64EC ABI.
@@ -812,7 +813,7 @@ void Arm64Emitter::FillStaticRegs(bool FPRs, uint32_t GPRFillMask, uint32_t FPRF
ldr(REG_CALLRET_SP, STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.callret_sp));
if (NZCV) {
if (Options.NZCV) {
// Regardless of what GPRs/FPRs we're filling, we need to fill NZCV since it
// is always static and was almost certainly clobbered.
//
@@ -822,23 +823,23 @@ void Arm64Emitter::FillStaticRegs(bool FPRs, uint32_t GPRFillMask, uint32_t FPRF
msr(ARMEmitter::SystemRegister::NZCV, TmpReg);
}
FillSpecialRegs(TmpReg, TmpReg2, true, FPRs);
FillSpecialRegs(TmpReg, TmpReg2, {.SetFIZ = true, .SetPredRegs = Options.FPRs});
if (FPRs) {
if (Options.FPRs) {
if (EmitterCTX->HostFeatures.SupportsAVX && EmitterCTX->HostFeatures.SupportsSVE256) {
for (size_t i = 0; i < StaticFPRegisters.size(); i++) {
const auto Reg = StaticFPRegisters[i];
if (((1U << Reg.Idx()) & FPRFillMask) != 0) {
mov(ARMEmitter::Size::i64Bit, TmpReg, offsetof(Core::CpuStateFrame, State.xmm.avx.data[i][0]));
if (((1U << Reg.Idx()) & Options.FPRFillMask) != 0) {
mov(ARMEmitter::Size::i64Bit, TmpReg, ARRAY_OFFSETOF(Core::CpuStateFrame, State.xmm.avx.data, i));
ld1b<ARMEmitter::SubRegSize::i8Bit>(Reg.Z(), PRED_TMP_32B.Zeroing(), STATE.R(), TmpReg);
}
}
} else {
if (GPRFillMask && FPRFillMask == ~0U) {
if (Options.GPRFillMask && Options.FPRFillMask == ~0U) {
// Optimize the common case where we can fill four registers per instruction.
// Use one of the filling static registers before we fill it.
// Load the sse offset in to the temporary register
add(ARMEmitter::Size::i64Bit, TmpReg, STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[0][0]));
add(ARMEmitter::Size::i64Bit, TmpReg, STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data));
for (size_t i = 0; i < StaticFPRegisters.size(); i += 4) {
const auto Reg1 = StaticFPRegisters[i];
const auto Reg2 = StaticFPRegisters[i + 1];
@@ -851,12 +852,12 @@ void Arm64Emitter::FillStaticRegs(bool FPRs, uint32_t GPRFillMask, uint32_t FPRF
const auto Reg1 = StaticFPRegisters[i];
const auto Reg2 = StaticFPRegisters[i + 1];
if (((1U << Reg1.Idx()) & FPRFillMask) && ((1U << Reg2.Idx()) & FPRFillMask)) {
ldp<ARMEmitter::IndexType::OFFSET>(Reg1.Q(), Reg2.Q(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[i][0]));
} else if (((1U << Reg1.Idx()) & FPRFillMask)) {
ldr(Reg1.Q(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[i][0]));
} else if (((1U << Reg2.Idx()) & FPRFillMask)) {
ldr(Reg2.Q(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[i + 1][0]));
if (((1U << Reg1.Idx()) & Options.FPRFillMask) && ((1U << Reg2.Idx()) & Options.FPRFillMask)) {
ldp<ARMEmitter::IndexType::OFFSET>(Reg1.Q(), Reg2.Q(), STATE.R(), ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, State.xmm.sse.data, i));
} else if (((1U << Reg1.Idx()) & Options.FPRFillMask)) {
ldr(Reg1.Q(), STATE.R(), ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, State.xmm.sse.data, i));
} else if (((1U << Reg2.Idx()) & Options.FPRFillMask)) {
ldr(Reg2.Q(), STATE.R(), ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, State.xmm.sse.data, i + 1));
}
}
}
@@ -865,23 +866,23 @@ void Arm64Emitter::FillStaticRegs(bool FPRs, uint32_t GPRFillMask, uint32_t FPRF
// PF/AF are special, remove them from the mask
uint32_t PFAFMask = ((1u << REG_PF.Idx()) | ((1u << REG_AF.Idx())));
uint32_t PFAFFillMask = GPRFillMask & PFAFMask;
GPRFillMask &= ~PFAFMask;
uint32_t PFAFFillMask = Options.GPRFillMask & PFAFMask;
Options.GPRFillMask &= ~PFAFMask;
for (size_t i = 0; i < StaticRegisters.size(); i += 2) {
auto Reg1 = StaticRegisters[i];
auto Reg2 = StaticRegisters[i + 1];
if (((1U << Reg1.Idx()) & GPRFillMask) && ((1U << Reg2.Idx()) & GPRFillMask)) {
ldp<ARMEmitter::IndexType::OFFSET>(Reg1.X(), Reg2.X(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i]));
} else if ((1U << Reg1.Idx()) & GPRFillMask) {
ldr(Reg1.X(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i]));
} else if ((1U << Reg2.Idx()) & GPRFillMask) {
ldr(Reg2.X(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i + 1]));
if (((1U << Reg1.Idx()) & Options.GPRFillMask) && ((1U << Reg2.Idx()) & Options.GPRFillMask)) {
ldp<ARMEmitter::IndexType::OFFSET>(Reg1.X(), Reg2.X(), STATE.R(), ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, State.gregs, i));
} else if ((1U << Reg1.Idx()) & Options.GPRFillMask) {
ldr(Reg1.X(), STATE.R(), ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, State.gregs, i));
} else if ((1U << Reg2.Idx()) & Options.GPRFillMask) {
ldr(Reg2.X(), STATE.R(), ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, State.gregs, i + 1));
}
}
// Now handle PF/AF
if (NZCV && PFAFFillMask) {
if (Options.NZCV && PFAFFillMask) {
LOGMAN_THROW_A_FMT(PFAFFillMask == PFAFMask, "PF/AF not filled together");
ldp<ARMEmitter::IndexType::OFFSET>(REG_PF.W(), REG_AF.W(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.pf_raw));
@@ -1056,7 +1057,11 @@ size_t Arm64Emitter::SpillForPreserveAllABICall(ARMEmitter::Register TmpReg, boo
const uint64_t SPOffset = AlignUp(GPRSize + FPRSize, 16);
// Spill the static registers.
SpillStaticRegs(TmpReg, true, PreserveSRAMask, PreserveSRAFPRMask);
SpillStaticRegs(TmpReg, {
.GPRSpillMask = PreserveSRAMask,
.FPRSpillMask = PreserveSRAFPRMask,
.FPRs = FPRs,
});
sub(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, ARMEmitter::Reg::rsp, SPOffset);
@@ -1103,7 +1108,11 @@ void Arm64Emitter::FillForPreserveAllABICall(bool FPRs) {
}
// Fill the static registers.
FillStaticRegs(FPRs, PreserveSRAMask, PreserveSRAFPRMask);
FillStaticRegs({
.GPRFillMask = PreserveSRAMask,
.FPRFillMask = PreserveSRAFPRMask,
.FPRs = FPRs,
});
// Pop the vector registers.
PopVectorRegisters(CanUseSVE256, DynamicFPRs);
@@ -129,16 +129,55 @@ protected:
std::span<const ARMEmitter::VRegister> GeneralFPRegisters {};
uint32_t PairRegisters = 0;
void FillSpecialRegs(ARMEmitter::Register TmpReg, ARMEmitter::Register TmpReg2, bool SetFIZ, bool SetPredRegs);
bool SupportCodeRelocations;
struct FillSpecialRegsOptions {
// Whether or not to set the FPCR.FIZ (flush inputs to zero) bit in the FPCR to
// the current value of the emulated MXCSR.DAZ bit.
// Will only attempt to do so, even when set to true, if and only if the host system
// supports FEAT_AFP.
bool SetFIZ {};
// Whether or not FillSpecialRegs should load our SVE predicate temporaries
// with certain canned values that accelerate some operations. Will (obviously)
// not load predicates, even if set to true, on host systems that do not support SVE.
bool SetPredRegs {};
};
void FillSpecialRegs(ARMEmitter::Register TmpReg, ARMEmitter::Register TmpReg2, const FillSpecialRegsOptions& Options);
// Correlate an ARM register back to an x86 register index.
// Returning REG_INVALID if there was no mapping.
FEXCore::X86State::X86Reg GetX86RegRelationToARMReg(ARMEmitter::Register Reg);
void SpillStaticRegs(ARMEmitter::Register TmpReg, bool FPRs = true, uint32_t GPRSpillMask = ~0U, uint32_t FPRSpillMask = ~0U, bool NZCV = true);
void FillStaticRegs(bool FPRs = true, uint32_t GPRFillMask = ~0U, uint32_t FPRFillMask = ~0U,
std::optional<ARMEmitter::Register> OptionalReg = std::nullopt,
std::optional<ARMEmitter::Register> OptionalReg2 = std::nullopt, bool NZCV = true);
struct SpillStaticRegOptions final {
uint32_t GPRSpillMask {~0U};
uint32_t FPRSpillMask {~0U};
bool FPRs {true};
bool NZCV {true};
};
struct FillStaticRegOptions final {
std::optional<ARMEmitter::Register> OptionalReg {std::nullopt};
std::optional<ARMEmitter::Register> OptionalReg2 {std::nullopt};
uint32_t GPRFillMask {~0U};
uint32_t FPRFillMask {~0U};
bool FPRs {true};
bool NZCV {true};
};
void SpillStaticRegs(ARMEmitter::Register TmpReg, SpillStaticRegOptions Options);
void FillStaticRegs(FillStaticRegOptions Options);
void SpillStaticRegs(ARMEmitter::Register TmpReg) {
// Work around a clang bug: https://bugs.llvm.org/show_bug.cgi?id=36684
SpillStaticRegs(TmpReg, {});
}
void FillStaticRegs() {
// Work around a clang bug: https://bugs.llvm.org/show_bug.cgi?id=36684
FillStaticRegs({});
}
// Register 0-18 + 29 + 30 are caller saved
static constexpr uint32_t CALLER_GPR_MASK = 0b0110'0000'0000'0111'1111'1111'1111'1111U;
@@ -178,7 +217,9 @@ protected:
if (SupportsPreserveAllABI) {
return SpillForPreserveAllABICall(TmpReg, FPRs);
} else {
SpillStaticRegs(TmpReg, FPRs);
SpillStaticRegs(TmpReg, {
.FPRs = FPRs,
});
return PushDynamicRegs(TmpReg);
}
}
@@ -188,7 +229,7 @@ protected:
FillForPreserveAllABICall(FPRs);
} else {
PopDynamicRegs();
FillStaticRegs(FPRs);
FillStaticRegs({.FPRs = FPRs});
}
}
@@ -281,8 +322,6 @@ protected:
FEX_CONFIG_OPT(Disassemble, DISASSEMBLE);
#endif
FEX_CONFIG_OPT(EnableCodeCaching, ENABLECODECACHINGWIP);
};
} // namespace FEXCore::CPU
+13 -105
View File
@@ -11,18 +11,8 @@
#include <cstdint>
#ifndef _WIN32
#include <linux/prctl.h>
#include <sys/prctl.h>
#endif
namespace FEXCore {
namespace CPU {
static constexpr size_t INITIAL_CODE_SIZE = 1024 * 1024 * 16;
// We don't want to move above 128MB atm because that means we will have to encode longer jumps
static constexpr size_t MAX_CODE_SIZE = 1024 * 1024 * 128;
constexpr static uint64_t NamedVectorConstants[FEXCore::IR::NamedVectorConstant::NAMED_VECTOR_CONST_POOL_MAX][2] = {
{0x0003'0002'0001'0000ULL, 0x0007'0006'0005'0004ULL}, // NAMED_VECTOR_INCREMENTAL_U16_INDEX
{0x000B'000A'0009'0008ULL, 0x000F'000E'000D'000CULL}, // NAMED_VECTOR_INCREMENTAL_U16_INDEX_UPPER
@@ -44,6 +34,8 @@ namespace CPU {
{0x0706'0504'FFFF'FFFFULL, 0x0F0E'0D0C'0B0A'0908ULL}, // NAMED_VECTOR_BLENDPS_1110B
{0x8040'2010'0804'0201ULL, 0x8040'2010'0804'0201ULL}, // NAMED_VECTOR_MOVMASKB
{0x8040'2010'0804'0201ULL, 0x8040'2010'0804'0201ULL}, // NAMED_VECTOR_MOVMASKB_UPPER
{0x0706'0504'0302'0100ULL, 0x1716'1514'1312'1110ULL}, // NAMED_VECTOR_256_MID_ELEMENT_SWAP
{0x0F0E'0D0C'0B0A'0908ULL, 0x1F1E'1D1C'1B1A'1918ULL}, // NAMED_VECTOR_256_MID_ELEMENT_SWAP_UPPER
{0x8000'0000'0000'0000ULL, 0x0000'0000'0000'3FFFULL}, // NAMED_VECTOR_X87_ONE
{0xD49A'784B'CD1B'8AFEULL, 0x0000'0000'0000'4000ULL}, // NAMED_VECTOR_X87_LOG2_10
{0xB8AA'3B29'5C17'F0BCULL, 0x0000'0000'0000'3FFFULL}, // NAMED_VECTOR_X87_LOG2_E
@@ -274,9 +266,9 @@ namespace CPU {
return TotalLUT;
}()};
CPUBackend::CPUBackend(CodeBufferManager& CodeBuffers, FEXCore::Core::InternalThreadState* ThreadState)
CPUBackend::CPUBackend(SharedCodeBufferManager& SharedCodeBuffers, FEXCore::Core::InternalThreadState* ThreadState)
: ThreadState(ThreadState)
, CodeBuffers(CodeBuffers) {
, SharedCodeBuffers(SharedCodeBuffers) {
auto& Ptrs = ThreadState->CurrentFrame->Pointers;
@@ -315,11 +307,11 @@ namespace CPU {
CPUBackend::~CPUBackend() = default;
auto CPUBackend::GetEmptyCodeBuffer() -> CodeBuffer* {
auto CPUBackend::GetEmptySharedCodeBuffer() -> CodeBuffer* {
auto PrevCodeBuffer = CurrentCodeBuffer;
// Resize the code buffer and reallocate our code size
CurrentCodeBuffer = CodeBuffers.StartLargerCodeBuffer();
CurrentCodeBuffer = SharedCodeBuffers.StartLargerCodeBuffer();
RegisterForSignalHandler(std::move(PrevCodeBuffer));
return CurrentCodeBuffer.get();
@@ -337,7 +329,7 @@ namespace CPU {
}
fextl::shared_ptr<CodeBuffer> CPUBackend::CheckCodeBufferUpdate() {
auto NewCodeBuffer = CodeBuffers.GetLatest();
auto NewCodeBuffer = SharedCodeBuffers.GetLatest();
if (CurrentCodeBuffer != NewCodeBuffer) {
RegisterForSignalHandler(CurrentCodeBuffer);
return std::exchange(CurrentCodeBuffer, NewCodeBuffer);
@@ -345,104 +337,20 @@ namespace CPU {
return nullptr;
}
GuestToHostMap& GetLookupCache(const CodeBuffer& Buffer) {
return *Buffer.LookupCache;
}
CodeBuffer::CodeBuffer(size_t Size)
: AllocatedSize(Size) {
Ptr = static_cast<uint8_t*>(FEXCore::Allocator::VirtualAlloc(Size, true));
LOGMAN_THROW_A_FMT(!!Ptr, "Couldn't allocate code buffer");
// Protect the last page of the allocated buffer to trigger SIGSEGV on write access
uintptr_t LastPageAddr = AlignDown(reinterpret_cast<uintptr_t>(Ptr) + Size - 1, FEXCore::Utils::FEX_PAGE_SIZE);
if (!FEXCore::Allocator::VirtualProtect(reinterpret_cast<void*>(LastPageAddr), FEXCore::Utils::FEX_PAGE_SIZE,
FEXCore::Allocator::ProtectOptions::None)) {
LogMan::Msg::EFmt("Failed to mprotect last page of code buffer.");
}
FEXCore::Allocator::VirtualName("FEXMemJIT", reinterpret_cast<void*>(Ptr), Size);
LookupCache = fextl::make_unique<GuestToHostMap>();
}
CodeBuffer::~CodeBuffer() {
FEXCore::Allocator::VirtualFree(Ptr, AllocatedSize);
}
auto CodeBufferManager::AllocateNew(size_t Size) -> fextl::shared_ptr<CodeBuffer> {
#ifndef _WIN32
// MDWE (Memory-Deny-Write-Execute) is a new Linux 6.3 feature.
// It's equivalent to systemd's `MemoryDenyWriteExecute` but implemented entirely in the kernel.
//
// MDWE prevents applications from creating RWX memory mappings.
// This prevents FEX from doing anything JIT related, as FEX uses RWX for JIT memory mappings.
//
// A potential workaround to make FEX work with MDWE is to call mprotect every time we need to write or modify code.
// Alternatively, FEX could use a memory mirror where one half is mapped as RW and the other is RX.
//
// Once MDWE is enabled with the prctl, the feature is sealed and it can /NOT/ be turned off.
//
// Status of MDWE is queried through prctl using `PR_GET_MDWE`:
// -1: The kernel doesn't support MDWE
// 0: MDWE is supported but disabled
// >0: MDWE is enabled, hence prohibiting RWX mappings
#ifndef PR_GET_MDWE
#define PR_GET_MDWE 66
#endif
int MDWE = ::prctl(PR_GET_MDWE, 0, 0, 0, 0);
if (MDWE != -1 && MDWE != 0) {
LogMan::Msg::EFmt("MDWE was set to 0x{:x} which means FEX can't allocate executable memory", MDWE);
}
#endif
auto Buffer = fextl::make_shared<CodeBuffer>(Size);
Latest = Buffer;
LatestOffset = 0;
OnCodeBufferAllocated(Buffer);
return Buffer;
}
fextl::shared_ptr<CodeBuffer> CodeBufferManager::GetLatest() {
if (!Latest) {
if (FEXCore::Config::Get_ENABLECODECACHINGWIP()) {
// Start with a larger code buffer to avoid resizes that would discard
// code loaded from caches
AllocateNew(MAX_CODE_SIZE);
} else {
AllocateNew(INITIAL_CODE_SIZE);
}
}
return Latest;
}
fextl::shared_ptr<CodeBuffer> CodeBufferManager::StartLargerCodeBuffer() {
if (!Latest) {
// Allocate initial CodeBuffer and return it
return GetLatest();
}
auto NewCodeBufferSize = GetLatest()->AllocatedSize;
NewCodeBufferSize = std::min<size_t>(NewCodeBufferSize * 2, MAX_CODE_SIZE);
return AllocateNew(NewCodeBufferSize);
}
bool CPUBackend::IsAddressInCodeBuffer(uintptr_t Address) const {
auto CheckCodeBuffer = [](CodeBuffer& Buffer, uintptr_t Address) {
const auto CheckCodeBuffer = [](const CodeBuffer& Buffer, uintptr_t Address) {
const auto BufferPtr = reinterpret_cast<uintptr_t>(Buffer.Ptr);
// The last page of the code buffer is protected, so we need to exclude it from the valid range
// when checking if the address is in the code buffer.
uintptr_t LastPageAddr = AlignDown(reinterpret_cast<uintptr_t>(Buffer.Ptr) + Buffer.AllocatedSize - 1, FEXCore::Utils::FEX_PAGE_SIZE);
return (Address >= reinterpret_cast<uintptr_t>(Buffer.Ptr) && Address < LastPageAddr);
const uintptr_t LastPageAddr = AlignDown(BufferPtr + Buffer.AllocatedSize - 1, FEXCore::Utils::FEX_PAGE_SIZE);
return (Address >= BufferPtr && Address < LastPageAddr);
};
if (CheckCodeBuffer(*CurrentCodeBuffer, Address)) {
return true;
}
for (auto& Buffer : SignalHandlerCodeBuffers) {
for (const auto& Buffer : SignalHandlerCodeBuffers) {
if (CheckCodeBuffer(*Buffer, Address)) {
return true;
}
+5 -56
View File
@@ -8,6 +8,8 @@ $end_info$
#pragma once
#include "Interface/Core/SharedCodeBufferManager.h"
#include <FEXCore/Utils/CompilerDefs.h>
#include <FEXCore/Utils/SignalScopeGuards.h>
#include <FEXCore/fextl/memory.h>
@@ -41,63 +43,10 @@ namespace CodeSerialize {
struct GuestToHostMap;
namespace CPU {
struct CodeBuffer {
uint8_t* Ptr;
size_t AllocatedSize; // including guard page; see UsableSize()
fextl::unique_ptr<GuestToHostMap> LookupCache;
CodeBuffer(size_t Size);
CodeBuffer(const CodeBuffer&) = delete;
CodeBuffer& operator=(const CodeBuffer&) = delete;
CodeBuffer(CodeBuffer&& oth) = delete;
CodeBuffer& operator=(CodeBuffer&&) = delete;
~CodeBuffer();
/// Returns the number of bytes available for storing code
size_t UsableSize() const {
return AllocatedSize - FEXCore::Utils::FEX_PAGE_SIZE;
}
};
/**
* A manager that coordinates access to the CodeBuffer used for compiling new code across threads.
*
* The CodeBuffer is managed as a partially persistent data structure:
* - Exactly one CodeBuffer is now designated as "active", which means data can be appended to it
* - Lossy modifications to the active CodeBuffer will not invalidate any data in use by other threads (which is what enables save CodeBuffer sharing across threads)
* - Instead, such lossy modifications trigger a new "version" of the data in the modifying thread. Old versions of the CodeBuffer persist as read-only data for use by the other threads.
* - The other threads can update their version of the CodeBuffer. This will decrease the reference count and eventually trigger deallocation of the old version
*/
class CodeBufferManager {
public:
// Get the CodeBuffer that was most recently allocated.
// This is the only CodeBuffer that data may be written to.
fextl::shared_ptr<CodeBuffer> GetLatest();
// Allocate a new CodeBuffer with geometric growth up to an internal maximum.
// Subsequent calls to GetLatest will point to the returned buffer.
fextl::shared_ptr<CodeBuffer> StartLargerCodeBuffer();
// Write offset into the latest CodeBuffer
std::size_t LatestOffset {};
// Protects writes to the latest CodeBuffer and changes to LatestOffset
FEXCore::ForkableUniqueMutex CodeBufferWriteMutex;
virtual void OnCodeBufferAllocated(const std::shared_ptr<CodeBuffer>&) {};
private:
fextl::shared_ptr<CodeBuffer> Latest;
fextl::shared_ptr<CodeBuffer> AllocateNew(size_t Size);
};
class CPUBackend {
public:
CPUBackend(CodeBufferManager&, FEXCore::Core::InternalThreadState*);
CPUBackend(SharedCodeBufferManager&, FEXCore::Core::InternalThreadState*);
virtual ~CPUBackend();
@@ -190,7 +139,7 @@ namespace CPU {
FEXCore::Core::InternalThreadState* ThreadState;
[[nodiscard]]
CodeBuffer* GetEmptyCodeBuffer();
CodeBuffer* GetEmptySharedCodeBuffer();
// This is the code buffer containing the main code under execution by this thread.
// CheckCodeBufferUpdate must be used before compiling new code.
@@ -199,7 +148,7 @@ namespace CPU {
// Old CodeBuffer generations required to be valid until returning from signal handlers
fextl::vector<fextl::shared_ptr<CodeBuffer>> SignalHandlerCodeBuffers;
CodeBufferManager& CodeBuffers;
SharedCodeBufferManager& SharedCodeBuffers;
private:
void RegisterForSignalHandler(fextl::shared_ptr<CodeBuffer>);
+127 -9
View File
@@ -92,6 +92,7 @@ namespace ProductNames {
static const char ARM_AppleSilicon[] = "Apple Silicon";
static const char ARM_ORYON_1[] = "Oryon-1";
static const char ARM_ORYON_3[] = "Oryon-3";
static const char ARM_Ampere_1[] = "AmpereOne";
static const char ARM_Ampere_1A[] = "AmpereOneA";
static const char ARM_Ampere_1B[] = "AmpereOneB";
@@ -99,7 +100,7 @@ namespace ProductNames {
#endif
} // namespace ProductNames
uint32_t GetCPUID_Syscall() {
static uint32_t GetCPUID_Syscall() {
uint32_t CPU {};
FHU::Syscalls::getcpu(&CPU, nullptr);
return CPU;
@@ -141,13 +142,13 @@ constexpr uint32_t FAMILY_IDENTIFIER = GenerateFamily(CPUFamily {
#endif
#ifdef ARCHITECTURE_arm64
uint32_t GetCycleCounterFrequency() {
uint64_t GetCycleCounterFrequency() {
uint64_t Result {};
__asm("mrs %[Res], CNTFRQ_EL0" : [Res] "=r"(Result));
return Result;
}
uint32_t GetCPUID_TPIDRRO() {
static uint32_t GetCPUID_TPIDRRO() {
uint64_t Result {};
__asm("mrs %[Res], TPIDRRO_EL0" : [Res] "=r"(Result));
return Result;
@@ -186,8 +187,9 @@ void CPUIDEmu::SetupHostHybridFlag() {
// CPU priority order
// This is mostly arbitrary but will sort by some sort of CPU priority by performance
// Relative list so things they will commonly end up in big.little configurations sort of relate
static constexpr std::array<CPUMIDR, 67> CPUMIDRs = {{
static constexpr std::array<CPUMIDR, 68> CPUMIDRs = {{
// Typically big CPU cores
{0x51, 0x002, 1, ProductNames::ARM_ORYON_3}, // Qualcomm Oryon-3
{0x51, 0x001, 1, ProductNames::ARM_ORYON_1}, // Qualcomm Oryon-1
{0x61, 0x039, 1, ProductNames::ARM_Avalanche_M2Max}, // Apple Avalanche (M2 Max)
@@ -314,9 +316,8 @@ void CPUIDEmu::SetupHostHybridFlag() {
// Walk our list of CPUMIDRs to find the most little core
for (size_t j = LowestMIDRIdx; j < CPUMIDRs.size(); ++j) {
auto& MIDROption = CPUMIDRs[i];
const auto& MIDROption = CPUMIDRs[j];
if ((MIDROption.Implementer == Implementer && MIDROption.Part == Part) || (MIDROption.Implementer == 0 && MIDROption.Part == 0)) {
LowestMIDRIdx = j;
LowestMIDR = MIDR;
break;
@@ -408,7 +409,7 @@ void CPUIDEmu::SetupHostHybridFlag() {
}
#else
uint32_t GetCycleCounterFrequency() {
uint64_t GetCycleCounterFrequency() {
return 0;
}
@@ -648,6 +649,13 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_06h(uint32_t Leaf) const {
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_07h(uint32_t Leaf) const {
FEXCore::CPUID::FunctionResults Res {};
if (Leaf == 0) {
#ifndef _WIN32
constexpr uint32_t SUPPORTS_RDPID = 1;
#else
// RDPID under WIN32 is only supported if CPUIndex is available in TPIDRRO.
const uint32_t SUPPORTS_RDPID = SupportsCPUIndexInTPIDRRO;
#endif
// Disable Enhanced REP MOVS when TSO is enabled.
// vcruntime140 memmove will use `rep movsb` in this case which completely destroys perf in Hades(appId 1145360)
// This is due to LRCPC performance on Cortex being abysmal.
@@ -713,7 +721,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_07h(uint32_t Leaf) const {
(0 << 19) | // MPX MAWAU
(0 << 20) | // MPX MAWAU
(0 << 21) | // MPX MAWAU
(1 << 22) | // RDPID Read Processor ID
(SUPPORTS_RDPID << 22) | // RDPID Read Processor ID
(0 << 23) | // AES Key Locker
(1 << 24) | // bus-lock-detect
(0 << 25) | // CLDEMOTE
@@ -756,6 +764,95 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_07h(uint32_t Leaf) const {
(0 << 29) | // Arch capabilities - Speculative side channel mitigations
(0 << 30) | // Arch capabilities - MSR module specific
(0 << 31); // SSBD - Speculative Store Bypass Disable
} else if (Leaf == 1) {
Res.eax = (0U << 0) | // SHA512
(0U << 1) | // SM3
(0U << 2) | // SM4
(0U << 3) | // RAO_INT
(0U << 4) | // AVX_VNNI
(0U << 5) | // AVX512_BF16
(0U << 6) | // LASS (Linear Address Space Separation)
(0U << 7) | // CMPCCXADD
(0U << 8) | // ARCH_PERFMON_EXT
(0U << 9) | // Reserved
(0U << 10) | // FAST_REP_MOVSB
(0U << 11) | // FAST_REP_STOSB
(0U << 12) | // FAST_REP_CMPSB_SCASB
(0U << 13) | // Reserved
(0U << 14) | // Reserved
(0U << 15) | // Reserved
(0U << 16) | // Reserved
(0U << 17) | // FRED (Flexible Return and Event Delivery)
(0U << 18) | // LKGS (Load into Kernel GS Base)
(0U << 19) | // WRMSRNS
(0U << 20) | // NMI_SRC
(0U << 21) | // AMX_FP16
(0U << 22) | // HRESET
(0U << 23) | // AVX_IFMA
(0U << 24) | // Reserved
(0U << 25) | // Reserved
(0U << 26) | // LAM (Linear Address Masking)
(0U << 27) | // MSRLIST
(0U << 28) | // Reserved
(0U << 29) | // Reserved
(0U << 30) | // INVD_DISABLE_POST_BIOS_DONE
(0U << 31); // MOVRS
// Bits 4-31 currently reserved.
Res.ebx = (0U << 0) | // PPIN
(0U << 1) | // PBNDKB
(0U << 2) | // Reserved
(0U << 3); // CPUIDMAXVAL_LIM_RMV
// Bits 6-31 also reserved.
Res.ecx = (0U << 0) | // RDT_M_ASYM
(0U << 1) | // RDT_A_ASYM
(0U << 2) | // Reserved
(0U << 3) | // Reserved
(0U << 4) | // Reserved
(0U << 5); // MSR_IMM
// Bits 25-31 also reserved.
Res.edx = (0U << 0) | // Reserved
(0U << 1) | // Reserved
(0U << 2) | // Reserved
(0U << 3) | // Reserved
(0U << 4) | // AVX_VNNI_INT8
(0U << 5) | // AVX_NE_CONVERT
(0U << 6) | // Reserved
(0U << 7) | // Reserved
(0U << 8) | // AMX_COMPLEX
(0U << 9) | // Reserved
(0U << 10) | // AVX_VNNI_INT16
(0U << 11) | // Reserved
(0U << 12) | // Reserved
(0U << 13) | // UTMR (User-timer events)
(0U << 14) | // PREFETCHI
(0U << 15) | // USER_MSR
(0U << 16) | // Reserved
(0U << 17) | // UIRET_UIF
(0U << 18) | // CET_SSS
(0U << 19) | // AVX10
(0U << 20) | // Reserved
(0U << 21) | // APX_F
(0U << 22) | // SEC-TEE_ATTESTATION
(0U << 23) | // MWAIT
(0U << 24); // SLSM (Static LSM)
} else if (Leaf == 2) {
// All bits are reserved except for EDX
Res.eax = 0;
Res.ebx = 0;
Res.ecx = 0;
// Bits 8-31 are reserved.
Res.edx = (0U << 0) | // PSFD
(0U << 1) | // IPRED_CTRL
(0U << 2) | // RRSBA_CTRL
(0U << 3) | // DDPD_U
(0U << 4) | // BHI_CTRL
(0U << 5) | // MCDT_NO
(0U << 6) | // UC_LOCK_DISABLE
(0U << 7); // MONITOR_MITG_NO
}
return Res;
@@ -813,7 +910,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_0Dh(uint32_t Leaf) const {
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_15h(uint32_t Leaf) const {
FEXCore::CPUID::FunctionResults Res {};
// TSC frequency = ECX * EBX / EAX
uint32_t FrequencyHz = GetCycleCounterFrequency();
uint64_t FrequencyHz = GetCycleCounterFrequency();
if (FrequencyHz) {
Res.eax = 1;
Res.ebx = 1U << CTX->Config.TSCScale;
@@ -834,6 +931,27 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_1Ah(uint32_t Leaf) const {
return Res;
}
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_24h(uint32_t Leaf) const {
FEXCore::CPUID::FunctionResults Res {};
if (Leaf == 0) {
// EAX indicates the maximum number of subleaves.
Res.eax = 0;
// Bits 19-31 reserved
// NOTE: We return all zero here until we have a CPU with AVX10
// even if some of the fields otherwise have fixed values.
Res.ebx = (0U << 0) | // (bits 0-7 specify the vector ISA version)
(0U << 16); // Defined as always 0b111
// All bits reserved
Res.ecx = 0;
Res.edx = 0;
}
return Res;
}
// Hypervisor CPUID information leaf
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_4000_0000h(uint32_t Leaf) const {
FEXCore::CPUID::FunctionResults Res {};
+84 -2
View File
@@ -14,7 +14,7 @@ namespace Context {
class ContextImpl;
}
uint32_t GetCycleCounterFrequency();
uint64_t GetCycleCounterFrequency();
// Debugging define to switch what family of CPU we execute as.
// Might be useful if an application makes an assumption about a CPU.
@@ -176,6 +176,7 @@ private:
FEXCore::CPUID::FunctionResults Function_0Dh(uint32_t Leaf) const;
FEXCore::CPUID::FunctionResults Function_15h(uint32_t Leaf) const;
FEXCore::CPUID::FunctionResults Function_1Ah(uint32_t Leaf) const;
FEXCore::CPUID::FunctionResults Function_24h(uint32_t Leaf) const;
FEXCore::CPUID::FunctionResults Function_4000_0000h(uint32_t Leaf) const;
FEXCore::CPUID::FunctionResults Function_4000_0001h(uint32_t Leaf) const;
FEXCore::CPUID::FunctionResults Function_8000_0000h(uint32_t Leaf) const;
@@ -200,7 +201,7 @@ private:
void SetupHostHybridFlag();
void SetupFeatures();
static constexpr size_t PRIMARY_FUNCTION_COUNT = 27;
static constexpr size_t PRIMARY_FUNCTION_COUNT = 37;
static constexpr size_t HYPERVISOR_FUNCTION_COUNT = 2;
static constexpr size_t EXTENDED_FUNCTION_COUNT = 32;
static constexpr std::array<FunctionHandler, PRIMARY_FUNCTION_COUNT> Primary = {
@@ -268,7 +269,48 @@ private:
#ifndef CPUID_AMD
// 0x1A: Hybrid Information Sub-leaf
&CPUIDEmu::Function_1Ah,
// 0x1B: PCONFIG info
&CPUIDEmu::Function_Reserved,
// 0x1C: Last Branch Records (LBR) info
&CPUIDEmu::Function_Reserved,
// 0x1D: Tile info
&CPUIDEmu::Function_Reserved,
// 0x1E: TMUL info
&CPUIDEmu::Function_Reserved,
// 0x1F: V2 Extended topology
&CPUIDEmu::Function_Reserved,
// 0x20: Processor History Reset info
&CPUIDEmu::Function_Reserved,
// 0x21: Unimplemented
&CPUIDEmu::Function_Reserved,
// 0x22: Reserved
&CPUIDEmu::Function_Reserved,
// 0x23: Architectural Performance Monitoring Extended
&CPUIDEmu::Function_Reserved,
// 0x24: Converged Vector ISA
&CPUIDEmu::Function_24h,
#else
// 0x1A: Reserved
&CPUIDEmu::Function_Reserved,
// 0x1B: Reserved
&CPUIDEmu::Function_Reserved,
// 0x1C: Reserved
&CPUIDEmu::Function_Reserved,
// 0x1D: Reserved
&CPUIDEmu::Function_Reserved,
// 0x1E: Reserved
&CPUIDEmu::Function_Reserved,
// 0x1F: Reserved
&CPUIDEmu::Function_Reserved,
// 0x20: Reserved
&CPUIDEmu::Function_Reserved,
// 0x21: Reserved
&CPUIDEmu::Function_Reserved,
// 0x22: Reserved
&CPUIDEmu::Function_Reserved,
// 0x23: Reserved
&CPUIDEmu::Function_Reserved,
// 0x24: Reserved
&CPUIDEmu::Function_Reserved,
#endif
};
@@ -340,9 +382,49 @@ private:
#ifndef CPUID_AMD
// 0x1A: Hybrid Information Sub-leaf
{SupportsConstant::NONCONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x1B: PCONFIG info
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x1C: Last Branch Records (LBR) info
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x1D: Tile info
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x1E: TMUL info
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x1F: V2 Extended topology
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x20: Processor History Reset info
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x21: Unimplemented/Reserved
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x22: Reserved
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x23: Architectural Performance Monitoring Extended
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x24: Converged Vector ISA
{SupportsConstant::CONSTANT, NeedsLeafConstant::NEEDSLEAFCONSTANT},
#else
// 0x1A: Reserved
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x1B: Reserved
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x1C: Reserved
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x1D: Reserved
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x1E: Reserved
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x1F: Reserved
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x20: Reserved
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x21: Reserved
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x22: Reserved
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x23: Reserved
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x24: Reserved
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
#endif
}};
+431 -192
View File
@@ -1,5 +1,10 @@
// SPDX-License-Identifier: MIT
#include "Utils/SpinWaitLock.h"
#include "FEXCore/Utils/LogManager.h"
#include "FEXCore/Utils/MathUtils.h"
#include "FEXCore/Utils/TypeDefines.h"
#include "FEXCore/fextl/memory.h"
#include <FEXCore/Utils/Profiler.h>
#include <FEXCore/Utils/SpinWaitLock.h>
#include <Interface/Context/Context.h>
#include <Interface/Core/ArchHelpers/Arm64Emitter.h>
@@ -16,10 +21,14 @@
#include <FEXHeaderUtils/Filesystem.h>
#include <algorithm>
#include <git_version.h>
#include <span>
#include <xxhash.h>
#include <FEXCore/Utils/AllocatorHooks.h>
#include <fstream>
namespace FEXCore {
@@ -32,6 +41,42 @@ ExecutableFileInfo::ExecutableFileInfo(fextl::unique_ptr<HLE::SourcecodeMap> Map
#endif
ExecutableFileInfo::~ExecutableFileInfo() = default;
MappedCodeCacheFile::~MappedCodeCacheFile() {
if (CacheManager) {
CacheManager->UnregisterMappedCodeBuffer(*this);
}
#ifndef _WIN32
if (!CodeBuffer.empty()) {
FEXCore::Allocator::munmap(CodeBuffer.data(), CodeBuffer.size_bytes());
}
#elif defined(_M_ARM64EC)
if (!CodeBuffer.empty()) {
FEXCore::Allocator::VirtualFree(CodeBuffer.data(), CodeBuffer.size_bytes());
}
#endif
}
void AbstractCodeCache::RegisterMappedCodeBuffer(MappedCodeCacheFile& Code) {
MappedCodeBuffers.push_back(Code.CodeBuffer);
// Unregister on destruction of Code
Code.CacheManager = this;
}
void AbstractCodeCache::UnregisterMappedCodeBuffer(MappedCodeCacheFile& Code) {
std::erase_if(MappedCodeBuffers, [&](const auto& Elem) { return Elem.data() == Code.CodeBuffer.data(); });
}
bool AbstractCodeCache::IsAddressInMappedCodeBuffer(uintptr_t Address) const {
for (const auto& Range : MappedCodeBuffers) {
auto Start = reinterpret_cast<uintptr_t>(Range.data());
if (Address >= Start && Address < Start + Range.size_bytes()) {
return true;
}
}
return false;
}
fextl::string CodeMap::GetBaseFilename(const ExecutableFileInfo& MainExecutable, bool AddNombSuffix) {
auto FileId = MainExecutable.FileId;
@@ -66,16 +111,22 @@ fextl::map<CodeMapFileId, CodeMap::ParsedContents> CodeMap::ParseCodeMap(std::if
break;
}
Ret[Info.ExternalFileId].Filename = std::move(Filename);
} else if (Entry.FileId == SetExecutableFileId {}.Marker.FileId && Entry.BlockOffset == SetExecutableFileId {}.Marker.BlockOffset) {
} else if ((Entry.FileId == SetExecutableFileId::Marker32.FileId && Entry.BlockOffset == SetExecutableFileId::Marker32.BlockOffset) ||
(Entry.FileId == SetExecutableFileId::Marker64.FileId && Entry.BlockOffset == SetExecutableFileId::Marker64.BlockOffset)) {
CodeMapFileId ExecutableFileId;
File.read(reinterpret_cast<char*>(&ExecutableFileId), sizeof(ExecutableFileId));
if (!File) {
break;
}
Ret[ExecutableFileId].IsExecutable = true;
Ret[ExecutableFileId].ExecutableBitness =
(Entry.FileId == SetExecutableFileId::Marker32.FileId && Entry.BlockOffset == SetExecutableFileId::Marker32.BlockOffset) ? 32 : 64;
} else {
if (!Ret.contains(Entry.FileId)) {
LogMan::Msg::EFmt("Code map referenced unknown file id {:016x}", Entry.FileId);
if (Entry.FileId == 0xffff'ffff'ffff'ffff) {
ERROR_AND_DIE_FMT("Malformed code map");
} else {
LogMan::Msg::EFmt("Code map referenced unknown file id {:016x}", Entry.FileId);
}
} else {
Ret[Entry.FileId].Blocks.insert(Entry.BlockOffset);
}
@@ -181,8 +232,8 @@ void CodeMapWriter::AppendLibraryLoad(const FEXCore::ExecutableFileInfo& FileInf
AppendData(std::as_bytes(std::span {Data, TotalSize}));
}
void CodeMapWriter::AppendSetMainExecutable(const FEXCore::ExecutableFileInfo& FileInfo) {
CodeMap::SetExecutableFileId Data {.ExecutableFileId = FileInfo.FileId};
void CodeMapWriter::AppendSetMainExecutable(const FEXCore::ExecutableFileInfo& FileInfo, bool Is64Bit) {
CodeMap::SetExecutableFileId Data {Is64Bit ? CodeMap::SetExecutableFileId::Marker64 : CodeMap::SetExecutableFileId::Marker32, FileInfo.FileId};
AppendData(std::span {reinterpret_cast<const std::byte*>(&Data), sizeof(Data)});
}
@@ -233,7 +284,10 @@ uint64_t CodeCache::ComputeCodeMapId(std::string_view Filename, int FD) {
struct CodeCacheHeader {
std::array<char, 4> Magic = ExpectedMagic;
uint32_t FormatVersion = 1;
// Version history:
// 1: Initial version
// 2: Padding code buffer data to enable direct mapping
uint32_t FormatVersion = 2;
uint8_t FEXVersion[20] = {};
uint32_t NumBlocks;
uint32_t NumCodePages;
@@ -260,7 +314,7 @@ bool CodeCache::SaveData(Core::InternalThreadState& Thread, int fd, const Execut
std::ranges::copy(GIT_HASH, header.FEXVersion);
header.NumBlocks = LookupCache.BlockList.size();
header.NumCodePages = LookupCache.CodePages.size();
header.CodeBufferSize = CTX.LatestOffset;
header.CodeBufferSize = FEXCore::AlignUp(CTX.LatestOffset, Utils::FEX_PAGE_SIZE);
header.NumRelocations = Relocations.size();
header.SerializedBaseAddress = SerializedBaseAddress;
::write(fd, &header, sizeof(header));
@@ -300,21 +354,25 @@ bool CodeCache::SaveData(Core::InternalThreadState& Thread, int fd, const Execut
::write(fd, Relocations.data(), Relocations.size() * sizeof(Relocations[0]));
// Pad to next page in file so that the CodeBuffer can be mmap'ed into process on load
char Zero[64] {};
auto Off = lseek(fd, 0, SEEK_CUR);
while (Off != AlignUp(Off, Utils::FEX_PAGE_SIZE)) {
auto BytesToWrite = std::min(AlignUp(Off, Utils::FEX_PAGE_SIZE) - Off, sizeof(Zero));
::write(fd, Zero, BytesToWrite);
Off += BytesToWrite;
{
auto AlignedSize = AlignUp(lseek(fd, 0, SEEK_CUR), Utils::FEX_PAGE_SIZE);
::ftruncate(fd, AlignedSize);
lseek(fd, AlignedSize, SEEK_SET);
}
// Dump the host code (relocated for position-independent serialization)
std::span CodeBufferData(reinterpret_cast<std::byte*>(CodeBuffer->Ptr), reinterpret_cast<std::byte*>(CodeBuffer->Ptr) + CTX.LatestOffset);
if (!ApplyCodeRelocations(SerializedBaseAddress, CodeBufferData, Relocations, true)) {
if (!ApplyCodeRelocations(SerializedBaseAddress, CodeBufferData, Relocations, 0, true)) {
LOGMAN_THROW_A_FMT(false, "Failed to apply code relocations");
return false;
}
::write(fd, CodeBufferData.data(), CodeBufferData.size());
// Pad to next page in file for mmap
{
auto PaddedSize = AlignUp(lseek(fd, 0, SEEK_CUR), Utils::FEX_PAGE_SIZE);
::ftruncate(fd, PaddedSize);
lseek(fd, PaddedSize, SEEK_SET);
}
// Dump code pages
static_assert(OrderedContainer<decltype(LookupCache.CodePages)>, "Non-deterministic data source");
@@ -332,178 +390,6 @@ bool CodeCache::SaveData(Core::InternalThreadState& Thread, int fd, const Execut
return true;
}
bool CodeCache::LoadData(Core::InternalThreadState* Thread, std::byte* MappedCacheFile, const ExecutableFileSectionInfo& BinarySection) {
if (!EnableCodeCaching) {
return true;
}
namespace ranges = std::ranges;
// Read file header
CodeCacheHeader header {};
::memcpy(&header, MappedCacheFile, sizeof(header));
MappedCacheFile += sizeof(header);
LogMan::Msg::IFmt("Cache load: {:5} blocks; base={:#14x}; off={:#9x}-{:#09x}; {:016x} {}", header.NumBlocks, BinarySection.FileStartVA,
BinarySection.BeginVA - BinarySection.FileStartVA, BinarySection.EndVA - BinarySection.FileStartVA,
BinarySection.FileInfo.FileId, BinarySection.FileInfo.Filename);
if (!ranges::equal(header.Magic, header.ExpectedMagic)) {
LogMan::Msg::EFmt("Invalid cache file header");
return false;
}
if (!ranges::equal(header.FEXVersion, GIT_HASH)) {
LogMan::Msg::IFmt("Cache generated from old FEX version {:02x}, current is {:02x}; skipping", fmt::join(header.FEXVersion, ""),
fmt::join(GIT_HASH, ""));
return false;
}
if (header.NumBlocks == 0) {
// Valid caches are never empty
LogMan::Msg::IFmt("Code cache empty, aborting");
return false;
}
// Read guest<->host block mappings
using BlockListEntry = decltype(GuestToHostMap::BlockList)::value_type;
fextl::vector<BlockListEntry> BlockList(header.NumBlocks);
{
for (auto& BlockPtr : BlockList) {
::memcpy(&BlockPtr.first, MappedCacheFile, sizeof(BlockPtr.first));
MappedCacheFile += sizeof(BlockPtr.first);
::memcpy(&BlockPtr.second.HostCode, MappedCacheFile, sizeof(BlockPtr.second.HostCode));
MappedCacheFile += sizeof(BlockPtr.second.HostCode);
uint64_t NumGuestPages;
::memcpy(&NumGuestPages, MappedCacheFile, sizeof(NumGuestPages));
MappedCacheFile += sizeof(NumGuestPages);
BlockPtr.second.CodePages.resize(NumGuestPages);
::memcpy(BlockPtr.second.CodePages.data(), MappedCacheFile, std::span {BlockPtr.second.CodePages}.size_bytes());
MappedCacheFile += std::span {BlockPtr.second.CodePages}.size_bytes();
}
// Consistency check: VMA regions at the top and end should belong to the same file
auto [min_val, max_val] = ranges::minmax_element(BlockList, std::less {}, &decltype(BlockList)::value_type::first);
auto MinBound = CTX.SyscallHandler->LookupExecutableFileSection(Thread, min_val->first + BinarySection.FileStartVA);
auto MaxBound = CTX.SyscallHandler->LookupExecutableFileSection(Thread, max_val->first + BinarySection.FileStartVA);
if (&MinBound->FileInfo != &BinarySection.FileInfo || &MaxBound->FileInfo != &BinarySection.FileInfo) {
ERROR_AND_DIE_FMT("Cached blocks offsets {:#x}-{:#x} out of bounds for guest library {} ({:016x} @ {:#x}) while trying to load "
"section {:#x}-{:#x}!",
min_val->first, max_val->first, BinarySection.FileInfo.Filename, BinarySection.FileInfo.FileId,
BinarySection.FileStartVA, BinarySection.BeginVA, BinarySection.EndVA);
}
// Constrain BlockList to the given ExecutableFileSectionInfo
LOGMAN_THROW_A_FMT(ranges::is_sorted(BlockList, [](auto& a, auto& b) { return a.first < b.first; }), "Expected sorted block list");
auto begin = ranges::lower_bound(BlockList, BinarySection.BeginVA - BinarySection.FileStartVA, std::less {}, &BlockListEntry::first);
auto end =
ranges::upper_bound(begin, BlockList.end(), BinarySection.EndVA - BinarySection.FileStartVA - 1, std::less {}, &BlockListEntry::first);
BlockList.erase(end, BlockList.end());
BlockList.erase(BlockList.begin(), begin);
if (BlockList.empty()) {
// Not an error since there is just no data to load
LogMan::Msg::IFmt("No blocks cached in this range, aborting");
return true;
}
}
// Read relocations
fextl::vector<FEXCore::CPU::Relocation> Relocations(header.NumRelocations, FEXCore::CPU::Relocation::Default());
::memcpy(Relocations.data(), MappedCacheFile, Relocations.size() * sizeof(Relocations[0]));
MappedCacheFile += Relocations.size() * sizeof(Relocations[0]);
// Pad to next page in file, which contains CodeBuffer data
MappedCacheFile = reinterpret_cast<std::byte*>(AlignUp(reinterpret_cast<uintptr_t>(MappedCacheFile), Utils::FEX_PAGE_SIZE));
// Prepare CodeBuffer: Page aligned and big enough to hold all cached data
auto Lock = std::unique_lock {CTX.CodeBufferWriteMutex};
if (Thread) {
if (auto Prev = Thread->CPUBackend->CheckCodeBufferUpdate()) {
Allocator::VirtualDontNeed(Thread->CallRetStackBase, FEXCore::Core::InternalThreadState::CALLRET_STACK_SIZE);
auto lk = Thread->LookupCache->AcquireWriteLock();
Thread->LookupCache->ChangeGuestToHostMapping(*Prev, *CTX.GetLatest()->LookupCache, lk);
}
}
auto CodeBuffer = CTX.GetLatest();
LOGMAN_THROW_A_FMT(reinterpret_cast<uintptr_t>(CodeBuffer->Ptr) % 0x1000 == 0, "Expected CodeBuffer base to be page-aligned");
const auto Delta = AlignUp(CTX.LatestOffset, 0x1000) - CTX.LatestOffset;
CTX.LatestOffset += Delta;
while (CTX.LatestOffset + header.CodeBufferSize > CodeBuffer->UsableSize()) {
if (Thread) {
CTX.ClearCodeCache(Thread);
CodeBuffer = CTX.GetLatest();
LogMan::Msg::IFmt("Increased code buffer size to {} MiB for cache load", CodeBuffer->AllocatedSize / 1024 / 1024);
} else {
ERROR_AND_DIE_FMT("Cannot extend codebuffer without thread!");
}
}
// Read CodeBuffer data from file. Make sure the destination is page-aligned.
// TODO: Only load the data needed for the selected section
auto CodeBufferRange =
std::as_writable_bytes(std::span {CodeBuffer->Ptr, CodeBuffer->UsableSize()}).subspan(CTX.LatestOffset, header.CodeBufferSize);
::memcpy(CodeBufferRange.data(), MappedCacheFile, header.CodeBufferSize);
MappedCacheFile += header.CodeBufferSize;
CTX.LatestOffset += header.CodeBufferSize;
// Apply FEX relocations
auto Ret = ApplyCodeRelocations(BinarySection.FileStartVA, CodeBufferRange, Relocations, false);
LOGMAN_THROW_A_FMT(Ret == true, "Failed to apply code cache relocations");
{
auto& LookupCache = *CodeBuffer->LookupCache;
auto WriteLock = LookupCache.AcquireWriteLock();
// Register blocks to LookupCache
for (auto& [Guest, Host] : BlockList) {
for (auto& CodePage : Host.CodePages) {
CodePage += BinarySection.FileStartVA;
}
auto HostCode = reinterpret_cast<void*>(Host.HostCode + reinterpret_cast<uintptr_t>(CodeBufferRange.data()));
LookupCache.AddBlockMapping(Guest + BinarySection.FileStartVA, std::move(Host.CodePages), HostCode, WriteLock);
}
// Register loaded code ranges
fextl::vector<uint64_t> Entrypoints;
for (uint32_t i = 0; i < header.NumCodePages; ++i) {
uint64_t CodePage;
memcpy(&CodePage, MappedCacheFile, sizeof(CodePage));
CodePage += BinarySection.FileStartVA;
MappedCacheFile += sizeof(CodePage);
uint64_t NumEntrypoints;
memcpy(&NumEntrypoints, MappedCacheFile, sizeof(NumEntrypoints));
MappedCacheFile += sizeof(NumEntrypoints);
Entrypoints.resize(NumEntrypoints);
memcpy(Entrypoints.data(), MappedCacheFile, NumEntrypoints * sizeof(Entrypoints[0]));
MappedCacheFile += NumEntrypoints * sizeof(Entrypoints[0]);
for (auto& Entrypoint : Entrypoints) {
Entrypoint += BinarySection.FileStartVA;
}
if (LookupCache.AddBlockExecutableRange(Entrypoints, CodePage, FEXCore::Utils::FEX_PAGE_SIZE, WriteLock)) {
CTX.SyscallHandler->MarkGuestExecutableRange(Thread, CodePage, FEXCore::Utils::FEX_PAGE_SIZE);
}
}
}
if (EnableCodeCacheValidation) {
fextl::set<uint64_t> GuestBlocks, HostBlocks;
for (auto& [Guest, Host] : BlockList) {
GuestBlocks.insert(Guest + BinarySection.FileStartVA);
HostBlocks.insert(Host.HostCode);
}
Validate(BinarySection, std::move(GuestBlocks), HostBlocks, CodeBufferRange);
}
return true;
}
void CodeCache::Validate(const ExecutableFileSectionInfo& Section, fextl::set<uint64_t> GuestBlocks, const fextl::set<uint64_t>& HostBlocks,
std::span<std::byte> CachedCode) {
LOGMAN_THROW_A_FMT(!HostBlocks.empty(), "Tried to validate without any host blocks");
@@ -558,7 +444,7 @@ void CodeCache::Validate(const ExecutableFileSectionInfo& Section, fextl::set<ui
NewRelocations.erase(std::remove_if(NewRelocations.begin(), NewRelocations.end(), [](const CPU::Relocation& Reloc) {
return Reloc.Header.Type != CPU::RelocationTypes::RELOC_NAMED_SYMBOL_LITERAL && Reloc.Header.Type != CPU::RelocationTypes::RELOC_NAMED_THUNK_MOVE;
}));
(void)ApplyCodeRelocations(Section.FileStartVA, CodeBufferRangeRef, NewRelocations, false);
(void)ApplyCodeRelocations(Section.FileStartVA, CodeBufferRangeRef, NewRelocations, 0, false);
if (ValidationCTX->LatestOffset <= CodeBufferRangeRef.size()) {
// Reference compilation produced fewer bytes than our cache, so validation is going to fail.
@@ -590,7 +476,7 @@ void CodeCache::Validate(const ExecutableFileSectionInfo& Section, fextl::set<ui
if (tail->RIP >= Section.BeginVA && tail->RIP < Section.EndVA) {
auto [IRView, TotalInstructions, TotalInstructionsLength, StartAddr, Length, _] =
ValidationCTX->GenerateIR(ValidationThread.get(), tail->RIP, false, FEXCore::Config::Get_MAXINST());
fextl::stringstream ss;
fextl::ostringstream ss;
FEXCore::IR::Dump(&ss, &*IRView);
LogMan::Msg::EFmt("IR:\n{}", ss.str());
} else {
@@ -616,15 +502,17 @@ void CodeCache::Validate(const ExecutableFileSectionInfo& Section, fextl::set<ui
ValidationThread->LookupCache->ClearCache(ValidationThread->LookupCache->AcquireWriteLock());
ValidationCTX->LatestOffset = 0;
LogMan::Msg::IFmt("\tSuccessfully validated cache");
LogMan::Msg::IFmt(" successfully validated cache");
}
bool CodeCache::ApplyCodeRelocations(uint64_t GuestEntry, std::span<std::byte> Code,
std::span<const FEXCore::CPU::Relocation> EntryRelocations, bool ForStorage) {
std::span<const FEXCore::CPU::Relocation> EntryRelocations, uint32_t RelocationOffset, bool ForStorage) {
CPU::Arm64Emitter Emitter(&CTX, Code.data(), Code.size_bytes());
for (size_t j = 0; j < EntryRelocations.size(); ++j) {
const FEXCore::CPU::Relocation& Reloc = EntryRelocations[j];
Emitter.SetCursorOffset(Reloc.Header.Offset);
LOGMAN_THROW_A_FMT(Reloc.Header.Offset >= RelocationOffset, "Invalid relocation offset");
LOGMAN_THROW_A_FMT(Reloc.Header.Offset - RelocationOffset < Code.size_bytes(), "Invalid relocation offset");
Emitter.SetCursorOffset(Reloc.Header.Offset - RelocationOffset);
switch (Reloc.Header.Type) {
case FEXCore::CPU::RelocationTypes::RELOC_NAMED_SYMBOL_LITERAL: {
@@ -663,4 +551,355 @@ bool CodeCache::ApplyCodeRelocations(uint64_t GuestEntry, std::span<std::byte> C
return true;
}
fextl::unique_ptr<MappedCodeCacheFile>
CodeCache::LoadCache(std::span<std::byte> CacheFile, const ExecutableFileInfo& FileInfo, uint64_t FileStartVA) {
if (!EnableCodeCaching) {
return nullptr;
}
FEXCORE_PROFILE_SCOPED("LoadCache");
// Read file header
CodeCacheHeader header {};
::memcpy(&header, CacheFile.data(), sizeof(header));
if (!std::ranges::equal(header.Magic, header.ExpectedMagic)) {
LogMan::Msg::EFmt("Invalid cache file header");
return nullptr;
}
if (!std::ranges::equal(header.FEXVersion, GIT_HASH)) {
LogMan::Msg::IFmt("Cache generated from old FEX version {:02x}, current is {:02x}; skipping", fmt::join(header.FEXVersion, ""),
fmt::join(GIT_HASH, ""));
return nullptr;
}
if (header.NumBlocks == 0) {
// Valid caches are never empty
LogMan::Msg::IFmt("Code cache empty, aborting");
return nullptr;
}
// Skip over BlockEntry data since it won't be used until EnableLoadedSection
// TODO: Store direct offset to relocations in the header
auto* BlockListStart = CacheFile.data() + sizeof(header);
auto* Cursor = BlockListStart;
for (uint32_t i = 0; i < header.NumBlocks; ++i) {
Cursor += sizeof(uint64_t); // guest address
Cursor += sizeof(uint64_t); // host code address
uint64_t NumGuestCodePages;
::memcpy(&NumGuestCodePages, Cursor, sizeof(NumGuestCodePages));
Cursor += sizeof(NumGuestCodePages);
Cursor += NumGuestCodePages * sizeof(uint64_t);
}
auto Relocations = std::span {reinterpret_cast<const FEXCore::CPU::Relocation*>(Cursor), header.NumRelocations};
Cursor += Relocations.size_bytes();
// Pad to next page to get the code buffer data
Cursor = reinterpret_cast<std::byte*>(AlignUp(reinterpret_cast<uintptr_t>(Cursor), Utils::FEX_PAGE_SIZE));
auto CodeDataInFile = std::span {Cursor, header.CodeBufferSize};
#ifndef _WIN32
// Allocate target memory for post-relocation code. This is PROT_NONE until
// the first execution, so that contents can be lazily populated in a
// frontend-provided segfault handler.
void* CodeBufferAllocation = Allocator::mmap(nullptr, header.CodeBufferSize, PROT_NONE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
if (CodeBufferAllocation == MAP_FAILED) {
LogMan::Msg::EFmt("Failed to reserve target memory for code cache");
return nullptr;
}
auto CodeBuffer = std::span {static_cast<std::byte*>(CodeBufferAllocation), header.CodeBufferSize};
#elif defined(_M_ARM64EC)
// TODO: Implement lazy mapping on Windows
// NOTE: The executed code must have MEM_EXTENDED_PARAMETER_EC_CODE set, so we can't operate on the mapped cache file directly
void* CodeBufferAllocation = Allocator::VirtualAlloc(header.CodeBufferSize, true);
if (!CodeBufferAllocation) {
LogMan::Msg::EFmt("Failed to allocate code cache memory");
return nullptr;
}
auto CodeBuffer = std::span {reinterpret_cast<std::byte*>(CodeBufferAllocation), header.CodeBufferSize};
#else // WoW64
// TODO: Implement lazy mapping on Windows
auto CodeBuffer = CodeDataInFile;
#endif
// Group relocations by page
size_t NumPages = header.CodeBufferSize / Utils::FEX_PAGE_SIZE;
fextl::vector<MappedCodeCacheFile::PageRelocationRange> PageRelocationRanges(NumPages, {0, 0});
auto RelocBaseOffset = std::as_bytes(Relocations).data() - CacheFile.data();
auto RelocIt = Relocations.begin();
for (size_t Page = 0; Page < NumPages; ++Page) {
auto EndRelocIt = std::upper_bound(RelocIt, Relocations.end(), Page,
[](auto& Page, auto& Reloc) { return Page < Reloc.Header.Offset / Utils::FEX_PAGE_SIZE; });
PageRelocationRanges.at(Page) = {static_cast<uint32_t>(RelocBaseOffset + (RelocIt - Relocations.begin()) * sizeof(CPU::Relocation)),
static_cast<uint32_t>(EndRelocIt - RelocIt)};
RelocIt = EndRelocIt;
}
auto Storage = FEXCore::Allocator::aligned_alloc(alignof(MappedCodeCacheFile), sizeof(MappedCodeCacheFile));
return fextl::unique_ptr<MappedCodeCacheFile>(
new (Storage) MappedCodeCacheFile {this, CacheFile, CodeDataInFile, CodeBuffer, BlockListStart, header.NumBlocks, header.NumCodePages,
std::move(PageRelocationRanges), fextl::vector<bool>(NumPages), FileStartVA});
}
bool CodeCache::EnableLoadedSection(Core::InternalThreadState* Thread, MappedCodeCacheFile& Code, const ExecutableFileSectionInfo& BinarySection) {
if (!EnableCodeCaching) {
return true;
}
namespace ranges = std::ranges;
FEXCORE_PROFILE_SCOPED("EnableLoadedSection");
// Read block list from cache file
// TODO: Store section-ized BlockLists in cache file
using BlockListEntry = decltype(GuestToHostMap::BlockList)::value_type;
fextl::vector<BlockListEntry> BlockList(Code.NumBlocks);
{
auto* Cursor = Code.BlockListInFile;
for (auto& BlockPtr : BlockList) {
::memcpy(&BlockPtr.first, Cursor, sizeof(BlockPtr.first));
Cursor += sizeof(BlockPtr.first);
::memcpy(&BlockPtr.second.HostCode, Cursor, sizeof(BlockPtr.second.HostCode));
Cursor += sizeof(BlockPtr.second.HostCode);
uint64_t NumGuestPages;
::memcpy(&NumGuestPages, Cursor, sizeof(NumGuestPages));
Cursor += sizeof(NumGuestPages);
BlockPtr.second.CodePages.resize(NumGuestPages);
::memcpy(BlockPtr.second.CodePages.data(), Cursor, std::span {BlockPtr.second.CodePages}.size_bytes());
Cursor += std::span {BlockPtr.second.CodePages}.size_bytes();
}
// Constrain BlockList to the given ExecutableFileSectionInfo
LOGMAN_THROW_A_FMT(ranges::is_sorted(BlockList, [](auto& a, auto& b) { return a.first < b.first; }), "Expected sorted block list");
auto begin = ranges::lower_bound(BlockList, BinarySection.BeginVA - BinarySection.FileStartVA, std::less {}, &BlockListEntry::first);
auto end =
ranges::upper_bound(begin, BlockList.end(), BinarySection.EndVA - BinarySection.FileStartVA - 1, std::less {}, &BlockListEntry::first);
if (begin == end) {
LogMan::Msg::IFmt("No blocks cached in this range, aborting");
return true;
}
BlockList.erase(end, BlockList.end());
BlockList.erase(BlockList.begin(), begin);
}
LogMan::Msg::IFmt("Cache load: {:5} blocks; base={:#14x}; off={:#9x}-{:#09x}; {:016x} {}", BlockList.size(), BinarySection.FileStartVA,
BinarySection.BeginVA - BinarySection.FileStartVA, BinarySection.EndVA - BinarySection.FileStartVA,
BinarySection.FileInfo.FileId, BinarySection.FileInfo.Filename);
if (EnableLazyCodeCaching) {
LogMan::Msg::IFmt(" lazy mapping: base={:#14x} -> host={}; cache_source={}", BinarySection.FileStartVA,
fmt::ptr(Code.CodeBuffer.data()), fmt::ptr(Code.MappedFile.data()));
}
// Register blocks to LookupCache.
// The host addresses will point into the protected code buffer, so that FEX
// can lazily apply relocations on first execution of each page.
auto CodeBuffer = CTX.GetLatest();
{
FEXCORE_PROFILE_SCOPED("Decode");
auto& LookupCache = *CodeBuffer->LookupCache;
auto WriteLock = LookupCache.AcquireWriteLock();
for (auto& [Guest, Block] : BlockList) {
for (auto& CodePage : Block.CodePages) {
CodePage += BinarySection.FileStartVA;
}
LOGMAN_THROW_A_FMT(Block.HostCode < Code.CodeBuffer.size_bytes(), "Host offset {:#x} out of range ({:#x})", Block.HostCode,
Code.CodeBuffer.size_bytes());
auto HostCode = &Code.CodeBuffer[Block.HostCode];
LookupCache.AddBlockMapping(Guest + BinarySection.FileStartVA, std::move(Block.CodePages), HostCode, WriteLock);
}
// Guest code pages
auto* Cursor = Code.CodeBufferInFile.data() + Code.CodeBufferInFile.size_bytes();
fextl::vector<uint64_t> Entrypoints;
for (uint32_t i = 0; i < Code.NumCodePages; ++i) {
uint64_t CodePage;
memcpy(&CodePage, Cursor, sizeof(CodePage));
CodePage += BinarySection.FileStartVA;
Cursor += sizeof(CodePage);
uint64_t NumEntrypoints;
memcpy(&NumEntrypoints, Cursor, sizeof(NumEntrypoints));
Cursor += sizeof(NumEntrypoints);
Entrypoints.resize(NumEntrypoints);
memcpy(Entrypoints.data(), Cursor, std::span {Entrypoints}.size_bytes());
Cursor += std::span {Entrypoints}.size_bytes();
for (auto& Entrypoint : Entrypoints) {
Entrypoint += BinarySection.FileStartVA;
}
if (LookupCache.AddBlockExecutableRange(Entrypoints, CodePage, FEXCore::Utils::FEX_PAGE_SIZE, WriteLock)) {
CTX.SyscallHandler->MarkGuestExecutableRange(Thread, CodePage, FEXCore::Utils::FEX_PAGE_SIZE);
}
}
}
#ifndef _WIN32
if (!EnableLazyCodeCaching || EnableCodeCacheValidation) {
#else
// TODO: Implement lazy mapping on Windows
if (true) {
#endif
auto Range = SelectCodeRangeToFinalize(Code, 0, Code.CodeBuffer.size_bytes() / Utils::FEX_PAGE_SIZE);
FinalizeCodePages(Code, Range);
}
if (EnableCodeCacheValidation) {
fextl::set<uint64_t> GuestBlocks, HostBlocks;
for (auto& [Guest, Host] : BlockList) {
GuestBlocks.insert(Guest + BinarySection.FileStartVA);
HostBlocks.insert(Host.HostCode);
}
Validate(BinarySection, std::move(GuestBlocks), HostBlocks, Code.CodeBuffer);
}
return true;
}
} // namespace FEXCore::Context
namespace FEXCore {
static std::span<CPU::Relocation> SpanPageRelocations(const MappedCodeCacheFile& Code, size_t PageIndex) {
auto [Offset, Count] = Code.PageRelocationRanges.at(PageIndex);
return std::span {reinterpret_cast<FEXCore::CPU::Relocation*>(Code.MappedFile.data() + Offset), Count};
}
std::span<std::byte> AbstractCodeCache::SelectCodeRangeToFinalize(MappedCodeCacheFile& Code, size_t StartPage, size_t EndPage) {
// First, check if we were racing another thread in loading this range
if (std::find(Code.LoadedPages.begin() + StartPage, Code.LoadedPages.begin() + EndPage, false) == Code.LoadedPages.begin() + EndPage) {
return {};
}
LOGMAN_THROW_A_FMT(StartPage < EndPage, "Invalid page range [{}, {})", StartPage, EndPage);
LOGMAN_THROW_A_FMT(EndPage <= Code.NumPages(), "End page {} out of range ({})", EndPage, Code.NumPages());
// Include any pages that have relocations or block link records crossing
// into the current page range. This ensures we don't attempt to finalize
// any page twice, partially apply FEX relocations, or trigger page loads
// during block linking.
while (EndPage < Code.NumPages()) {
auto PageRelocs = SpanPageRelocations(Code, EndPage - 1);
if (!PageRelocs.empty()) {
auto It = std::prev(PageRelocs.end());
size_t RelocEnd = It->Header.Offset + 16 /* Upper bound for relocation size */;
if (RelocEnd > EndPage * Utils::FEX_PAGE_SIZE) {
++EndPage;
continue;
}
}
// Check for trailing block link
{
auto PageRelocs = SpanPageRelocations(Code, EndPage);
if (!PageRelocs.empty() && PageRelocs.begin()->Header.Offset < EndPage * Utils::FEX_PAGE_SIZE + 0x18) {
++EndPage;
continue;
}
}
break;
};
while (StartPage != 0) {
auto PageRelocs = SpanPageRelocations(Code, StartPage - 1);
if (!PageRelocs.empty()) {
auto It = std::prev(PageRelocs.end());
size_t RelocEnd = It->Header.Offset + 16 /* Upper bound for relocation size */;
if (RelocEnd > StartPage * Utils::FEX_PAGE_SIZE) {
--StartPage;
continue;
}
}
// Check for trailing block link
{
auto PageRelocs = SpanPageRelocations(Code, StartPage);
if (!PageRelocs.empty() && PageRelocs.begin()->Header.Offset < StartPage * Utils::FEX_PAGE_SIZE + 0x18) {
--StartPage;
continue;
}
}
break;
};
return Code.CodeBuffer.subspan(StartPage * Utils::FEX_PAGE_SIZE, (EndPage - StartPage) * Utils::FEX_PAGE_SIZE);
}
} // namespace FEXCore
namespace FEXCore::Context {
void CodeCache::FinalizeCodePages(MappedCodeCacheFile& Code, std::span<std::byte> CodeRange) {
const size_t StartOffset = CodeRange.data() - Code.CodeBuffer.data();
const auto StartPage = StartOffset / Utils::FEX_PAGE_SIZE;
const auto EndPage = StartPage + CodeRange.size_bytes() / Utils::FEX_PAGE_SIZE;
const size_t Size = CodeRange.size_bytes();
// None of the selected pages should be loaded at all; otherwise, SelectCodeRangeToFinalize returned inconsistent ranges
LOGMAN_THROW_A_FMT(std::find(Code.LoadedPages.begin() + StartPage, Code.LoadedPages.begin() + EndPage, true) == Code.LoadedPages.begin() + EndPage,
"Inconsistent page load state");
FEXCORE_PROFILE_SCOPED("FinalizeCodePages");
#ifndef _WIN32
// Atomicity is critical when making the finalized code data visible.
// We ensure this by remapping a temporary buffer onto the PROT_NONE
// placeholder page in CodeBuffer. Some constraints to keep in mind are:
// 1. Pages can't be write-only (readability is implicitly added), so
// we can't change CodeBuffer from PROT_NONE to PROT_WRITE even for just
// a short duration
// 2. Naive mremap from CodeBufferInFile to CodeBuffer would leave a gap in
// the former, which would make cleanup overly complicated
//
// Due to (1), we can't apply relocations in place (CodeBufferInFile); at
// least a secondary buffer is needed for execution (CodeBuffer).
// Due to (2), a third buffer is temporarily allocated here and freed on
// completion. The final code data is computed here and then the memory
// is remapped onto CodeBuffer.
auto* Staging = reinterpret_cast<std::byte*>(Allocator::VirtualAlloc(nullptr, Size, true));
if (!Staging) {
ERROR_AND_DIE_FMT("Failed to allocate {} bytes of staging memory for code-cache finalization", Size);
}
// Copy code from the cache file to the staging buffer
memcpy(Staging, Code.CodeBufferInFile.data() + StartOffset, Size);
// Apply relocations
auto StagingSpan = std::span {Staging, Size};
for (size_t i = StartPage; i < EndPage; ++i) {
auto PageRelocations = SpanPageRelocations(Code, i);
(void)ApplyCodeRelocations(Code.GuestBase, StagingSpan, PageRelocations, static_cast<uint32_t>(StartOffset), false);
Code.LoadedPages[i] = true;
}
// Atomically make the finalized code data visible by remapping the staging
// buffer onto the requested CodeBuffer window. MREMAP_DONTUNMAP is used to
// leave the old VA range reserved so that we can cleanly deallocate it
// through Allocator.
void* RemapResult = ::mremap(Staging, Size, Size, MREMAP_FIXED | MREMAP_MAYMOVE | MREMAP_DONTUNMAP, CodeRange.data());
if (RemapResult == MAP_FAILED) {
ERROR_AND_DIE_FMT("{}: mremap failed: {}", __FUNCTION__, errno);
}
Allocator::VirtualFree(Staging, Size);
// Release resident file pages that will no longer be needed. The VA range is left allocated to allow cleanup with a single VirtualFree.
Allocator::VirtualDontNeed(Code.CodeBufferInFile.data() + StartOffset, Size);
#else
// TODO: Implement lazy mapping on Windows
#ifdef _M_ARM64EC
memcpy(Code.CodeBuffer.data() + StartOffset, Code.CodeBufferInFile.data() + StartOffset, Size);
#endif
for (size_t i = StartPage; i < EndPage; ++i) {
auto PageRelocations = SpanPageRelocations(Code, i);
(void)ApplyCodeRelocations(Code.GuestBase, Code.CodeBuffer, PageRelocations, 0, false);
Code.LoadedPages[i] = true;
}
#endif
ARMEmitter::Emitter::ClearICache(CodeRange.data(), Size);
}
} // namespace FEXCore::Context
+61 -31
View File
@@ -30,7 +30,7 @@ $end_info$
#include "Interface/IR/RegisterAllocationData.h"
#include "Utils/Allocator.h"
#include "Utils/Allocator/HostAllocator.h"
#include "Utils/SpinWaitLock.h"
#include <FEXCore/Utils/SpinWaitLock.h>
#include "Utils/variable_length_integer.h"
#include <FEXCore/Config/Config.h>
@@ -342,6 +342,11 @@ void ContextImpl::SetFlagsFromCompactedEFLAGS(FEXCore::Core::InternalThreadState
}
bool ContextImpl::InitCore() {
if (CodeCache.IsGeneratingCache || FEXCore::Config::Get_ENABLECODECACHINGWIP()) {
// Start with a larger code buffer to avoid resizes that would discard code
StartMaximalCodeBuffer();
}
// Initialize the CPU core signal handlers & DispatcherConfig
Dispatcher = FEXCore::CPU::Dispatcher::Create(this);
@@ -358,6 +363,16 @@ bool ContextImpl::InitCore() {
Config.NeedsPendingInterruptFaultCheck = true;
}
if constexpr (BLOCK_DEBUGGING) {
// If the developer wants to do any single-stepping points or watch points.
// Add them here.
//
// eg:
// BlockDebuggerTracker.AllTargetSingleStep();
// BlockDebuggerTracker.AddSingleStepTarget(0x14000'0000ULL);
// BlockDebuggerTracker.AddWriteWatchPoint(0x420BA5ED);
}
return true;
}
@@ -376,11 +391,11 @@ void ContextImpl::ExecuteThread(FEXCore::Core::InternalThreadState* Thread) {
}
void ContextImpl::InitializeCompiler(FEXCore::Core::InternalThreadState* Thread) {
Thread->OpDispatcher = fextl::make_unique<FEXCore::IR::OpDispatchBuilder>(this);
Thread->OpDispatcher = fextl::make_unique<FEXCore::IR::OpDispatchBuilder>(this, Thread);
Thread->OpDispatcher->SetMultiblock(Config.Multiblock);
Thread->LookupCache = fextl::make_unique<FEXCore::LookupCache>(this);
Thread->FrontendDecoder = fextl::make_unique<FEXCore::Frontend::Decoder>(Thread);
Thread->PassManager = fextl::make_unique<FEXCore::IR::PassManager>();
Thread->PassManager = fextl::make_unique<FEXCore::IR::PassManager>(this);
Thread->CurrentFrame->State.L1Pointer = Thread->LookupCache->GetL1Pointer();
Thread->CurrentFrame->State.L1Mask = Thread->LookupCache->GetScaledL1PointerMask();
@@ -389,15 +404,11 @@ void ContextImpl::InitializeCompiler(FEXCore::Core::InternalThreadState* Thread)
Dispatcher->InitThreadPointers(Thread);
Thread->PassManager->AddDefaultPasses(this);
Thread->PassManager->AddDefaultValidationPasses();
Thread->PassManager->RegisterSyscallHandler(SyscallHandler);
// Create CPU backend
Thread->PassManager->InsertRegisterAllocationPass(this);
Thread->CPUBackend = FEXCore::CPU::CreateArm64JITCore(this, Thread);
// We finalize *after* the CPU backend is initialized, as the CPU backend will
// provide necessary register information to the register allocation pass.
Thread->PassManager->Finalize();
}
@@ -456,7 +467,6 @@ void ContextImpl::UnlockAfterFork(FEXCore::Core::InternalThreadState* LiveThread
if (Config.StrictInProcessSplitLocks) {
FEXCore::Utils::SpinWaitLock::unlock(&StrictSplitLockMutex);
}
return;
}
}
@@ -496,17 +506,20 @@ void ContextImpl::ClearCodeCache(FEXCore::Core::InternalThreadState* Thread, boo
static void IRDumper(FEXCore::Core::InternalThreadState* Thread, IR::IREmitter* IREmitter, uint64_t GuestRIP) {
FEXCore::File::File FD = FEXCore::File::File::GetStdERR();
fextl::stringstream out;
fextl::ostringstream out;
auto NewIR = IREmitter->ViewIR();
FEXCore::IR::Dump(&out, &NewIR);
fextl::fmt::print(FD, "IR-ShouldDump-{} 0x{:x}:\n{}\n@@@@@\n", NewIR.PostRA() ? "post" : "pre", GuestRIP, out.str());
};
}
bool ContextImpl::CheckIfBlockIsCacheable(FEXCore::Core::InternalThreadState& Thread, uint64_t GuestRIP, uint64_t MaxInst) {
return Thread.FrontendDecoder->CheckIfCacheable(Thread, reinterpret_cast<const uint8_t*>(GuestRIP), GuestRIP, MaxInst);
}
ContextImpl::GenerateIRResult
ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP, bool ExtendedDebugInfo, uint64_t MaxInst) {
FEXCORE_PROFILE_SCOPED("GenerateIR");
Thread->OpDispatcher->ReownOrClaimBuffer();
Thread->OpDispatcher->ResetWorkingList();
uint64_t TotalInstructions {0};
@@ -526,18 +539,14 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
}
if (!HasCustomIR) {
const uint8_t* GuestCode {};
GuestCode = reinterpret_cast<const uint8_t*>(GuestRIP);
bool HadDispatchError {false};
bool HadInvalidInst {false};
const auto* GuestCode = reinterpret_cast<const uint8_t*>(GuestRIP);
Thread->FrontendDecoder->DecodeInstructionsAtEntry(Thread, GuestCode, GuestRIP, MaxInst);
auto BlockInfo = Thread->FrontendDecoder->GetDecodedBlockInfo();
auto CodeBlocks = &BlockInfo->Blocks;
const auto* BlockInfo = Thread->FrontendDecoder->GetDecodedBlockInfo();
const auto& CodeBlocks = BlockInfo->Blocks;
Thread->OpDispatcher->BeginFunction(GuestRIP, CodeBlocks, BlockInfo->TotalInstructionCount, BlockInfo->Is64BitMode,
Thread->OpDispatcher->BeginFunction(GuestRIP, &CodeBlocks, BlockInfo->TotalInstructionCount, BlockInfo->Is64BitMode,
AreMonoHacksActive() && MonoBackpatcherBlock.load(std::memory_order_relaxed) == GuestRIP);
const auto GPRSize = Thread->OpDispatcher->GetGPROpSize();
@@ -551,11 +560,17 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
}
#endif
for (size_t j = 0; j < CodeBlocks->size(); ++j) {
const FEXCore::Frontend::Decoder::DecodedBlocks& Block = CodeBlocks->at(j);
for (size_t j = 0; j < CodeBlocks.size(); ++j) {
const auto& Block = CodeBlocks[j];
// Dispatch failures and invalid instructions terminate only the decoded
// block that contains them. Other block targets in the same multiblock
// compilation unit are independent entry paths.
bool HadDispatchError {false};
bool HadInvalidInst {false};
#ifdef ZYDIS_DISASSEMBLER
if (FEXCore::Config::Get_X86DISASSEMBLE() && CodeBlocks->size() > 1) {
if (FEXCore::Config::Get_X86DISASSEMBLE() && CodeBlocks.size() > 1) {
LogMan::Msg::IFmt(" Block {} Entry={:#x} NumInsts={}", j, Block.Entry, Block.NumInstructions);
}
#endif
@@ -563,7 +578,7 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
bool BlockInForceTSOValidRange = false;
auto InstForceTSOIt = ForceTSOInstructions.end();
if (ForceTSOValidRanges.Contains({Block.Entry, Block.Entry + Block.Size})) {
if (auto It = ForceTSOInstructions.lower_bound(Block.Entry); *It < Block.Entry + Block.Size) {
if (auto It = ForceTSOInstructions.lower_bound(Block.Entry); It != ForceTSOInstructions.end() && *It < Block.Entry + Block.Size) {
InstForceTSOIt = It;
BlockInForceTSOValidRange = true;
}
@@ -572,18 +587,16 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
// Set the block entry point
Thread->OpDispatcher->SetNewBlockIfChanged(Block.Entry);
uint64_t BlockInstructionsLength {};
// Reset any block-specific state
Thread->OpDispatcher->StartNewBlock();
uint64_t InstsInBlock = Block.NumInstructions;
const uint64_t InstsInBlock = Block.NumInstructions;
if (InstsInBlock == 0) {
// Special case for an empty instruction block.
Thread->OpDispatcher->ExitFunction(Thread->OpDispatcher->_InlineEntrypointOffset(GPRSize, Block.Entry - GuestRIP));
}
uint64_t BlockInstructionsLength {};
for (size_t i = 0; i < InstsInBlock; ++i) {
uint64_t InstAddress = Block.Entry + BlockInstructionsLength;
const FEXCore::X86Tables::X86InstInfo* TableInfo {nullptr};
@@ -638,6 +651,7 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
Thread->OpDispatcher->SetTrueJumpTarget(InvalidateCodeCond, CodeWasChangedBlock);
Thread->OpDispatcher->SetCurrentCodeBlock(CodeWasChangedBlock);
Thread->OpDispatcher->StartNewBlock();
Thread->OpDispatcher->_ThreadRemoveCodeEntry();
Thread->OpDispatcher->ExitFunction(Thread->OpDispatcher->_InlineEntrypointOffset(GPRSize, InstAddress - GuestRIP));
@@ -645,6 +659,7 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
Thread->OpDispatcher->SetFalseJumpTarget(InvalidateCodeCond, NextOpBlock);
Thread->OpDispatcher->SetCurrentCodeBlock(NextOpBlock);
Thread->OpDispatcher->StartNewBlock();
}
if (TableInfo && TableInfo->OpcodeDispatcher.OpDispatch) {
@@ -692,6 +707,8 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
if (Block.BlockStatus == Frontend::Decoder::DecodedBlockStatus::INVALID_INST ||
Block.BlockStatus == Frontend::Decoder::DecodedBlockStatus::BAD_RELOCATION) {
Thread->OpDispatcher->InvalidOp(DecodedInfo);
} else if (Block.BlockStatus == Frontend::Decoder::DecodedBlockStatus::UNIMPLEMENTED_INST) {
Thread->OpDispatcher->UnimplementedOp(DecodedInfo);
} else {
Thread->OpDispatcher->NoExecOp(DecodedInfo);
}
@@ -706,8 +723,8 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
// If we had a dispatch error then leave early
if (HadDispatchError && TotalInstructions == 0) {
// Couldn't handle any instruction in op dispatcher
Thread->OpDispatcher->ResetWorkingList();
return {{}, 0, 0, 0, 0};
Thread->OpDispatcher->DelayedDisownBuffer();
return {std::nullopt, 0, 0, 0, 0};
}
if (NeedsBlockEnd) {
@@ -774,6 +791,7 @@ ContextImpl::CompileCodeResult ContextImpl::CompileCode(FEXCore::Core::InternalT
auto [IRView, TotalInstructions, TotalInstructionsLength, StartAddr, Length, NeedsAddGuestCodeRanges] =
GenerateIR(Thread, GuestRIP, Config.GDBSymbols(), MaxInst);
if (!IRView) {
// OpDispatcher IR already released in this case.
return {{}, nullptr, 0, 0, false};
}
@@ -784,6 +802,7 @@ ContextImpl::CompileCodeResult ContextImpl::CompileCode(FEXCore::Core::InternalT
// as expensive and are easily reverted.
if (MaxInst != 1) {
if (auto Block = Thread->LookupCache->FindBlock(Thread, GuestRIP)) {
// Raced to compile, release the OpDispatcher IR.
Thread->OpDispatcher->DelayedDisownBuffer();
return {.CompiledCode = {.BlockBegin = reinterpret_cast<uint8_t*>(Block), .EntryPoints = {{GuestRIP, reinterpret_cast<uint8_t*>(Block)}}},
.DebugData = nullptr,
@@ -813,6 +832,17 @@ ContextImpl::CompileCodeResult ContextImpl::CompileCode(FEXCore::Core::InternalT
}
uintptr_t ContextImpl::CompileBlock(FEXCore::Core::CpuStateFrame* Frame, uint64_t GuestRIP, uint64_t MaxInst) {
if constexpr (BLOCK_DEBUGGING) {
// Block debugging logic is hand-written and needs to be handled with care.
// Force MaxInst to only be one in this case.
MaxInst = 1;
// If the entrypoint is part of the single step targets then single step it.
if (BlockDebuggerTracker.IsSingleStepTarget(GuestRIP)) {
return CompileSingleStep(Frame, GuestRIP);
}
}
auto Thread = Frame->Thread;
FEXCORE_PROFILE_SCOPED("CompileBlock");
FEXCORE_PROFILE_ACCUMULATION(Thread, AccumulatedJITTime);
File diff suppressed because it is too large. Load diff
@@ -28,6 +28,10 @@ class ContextImpl;
namespace FEXCore::CPU {
#define STATE_PTR(STATE_TYPE, FIELD) STATE.R(), offsetof(FEXCore::Core::STATE_TYPE, FIELD)
#define STATE_PTR_IDX(STATE_TYPE, FIELD, INDEX) STATE.R(), ARRAY_OFFSETOF(FEXCore::Core::STATE_TYPE, FIELD, INDEX)
#define FALLBACK_HANDLER_OFFSET(INDEX, FIELD) \
STATE.R(), \
(ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, Pointers.FallbackHandlerPointers, INDEX) + offsetof(FEXCore::Core::FallbackABIInfo, FIELD))
class Dispatcher final : public Arm64Emitter {
public:
@@ -95,6 +99,18 @@ private:
uint64_t LUDIVHandlerAddress {};
uint64_t LDIVHandlerAddress {};
// F64 reduced-precision shared handlers
uint64_t F64SinHandlerAddress {};
uint64_t F64CosHandlerAddress {};
uint64_t F64TanHandlerAddress {};
uint64_t F64F2XM1HandlerAddress {};
uint64_t F64ScaleHandlerAddress {};
uint64_t F64AtanHandlerAddress {};
uint64_t F64FYL2XHandlerAddress {};
uint64_t F64FYL2XP1HandlerAddress {};
uint64_t F64FPREMHandlerAddress {};
uint64_t F64FPREM1HandlerAddress {};
void EmitDispatcher();
uint64_t GenerateABICall(FallbackABI ABI);
@@ -105,6 +121,26 @@ private:
void EmitF32ToExtF80();
void EmitF64ToExtF80();
// Shared label set for the LUT-based F64 log2 path used by both FYL2X and
// FYL2XP1. The pool is emitted once via EmitF64Log2Constants.
struct F64Log2Constants {
ARMEmitter::ForwardLabel One;
ARMEmitter::ForwardLabel A0, A1, A2, A3, A4, A5, A6, A7;
ARMEmitter::ForwardLabel Table;
};
void EmitF64Sin();
void EmitF64Cos();
void EmitF64Tan();
void EmitF64F2XM1();
void EmitF64Scale();
void EmitF64Atan();
void EmitF64FYL2X(F64Log2Constants& C);
void EmitF64FYL2XP1(F64Log2Constants& C);
void EmitF64Log2Constants(F64Log2Constants& C);
void EmitF64FPREM();
void EmitF64FPREM1();
FEX_CONFIG_OPT(DisableL2Cache, DISABLEL2CACHE);
};
+85 -47
View File
@@ -124,9 +124,9 @@ uint8_t Decoder::ReadByte() {
}
std::optional<uint8_t> Decoder::PeekByte(uint8_t Offset) {
uint64_t ByteAddress = reinterpret_cast<uint64_t>(InstStream + InstructionSize + Offset);
uint64_t ByteAddress = reinterpret_cast<uint64_t>(InstStream.InstStream + InstructionSize + Offset);
if (CheckRangeExecutable(ByteAddress, 1)) {
return InstStream[InstructionSize + Offset];
return InstStream.AdjustedInstStream[InstructionSize + Offset];
} else {
return std::nullopt;
}
@@ -136,9 +136,9 @@ std::pair<uint64_t, bool> Decoder::ReadData(uint8_t Size) {
LOGMAN_THROW_A_FMT(Size != 0 && Size <= sizeof(uint64_t), "Unknown data size to read");
uint64_t Res = 0;
uint64_t Address = reinterpret_cast<uint64_t>(InstStream + InstructionSize);
uint64_t Address = reinterpret_cast<uint64_t>(InstStream.InstStream + InstructionSize);
if (CheckRangeExecutable(Address, Size)) {
std::memcpy(&Res, &InstStream[InstructionSize], Size);
std::memcpy(&Res, &InstStream.AdjustedInstStream[InstructionSize], Size);
} else {
HitNonExecutableRange = true;
// See PeekByte, this specific case may cause some executable memory to read as 0 but it doesn't matter as the entire instruction will be rolled back anyway.
@@ -342,7 +342,7 @@ void Decoder::DecodeModRM_64(X86Tables::DecodedOperand* Operand, X86Tables::ModR
}
}
bool Decoder::NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op, DecodedHeader Options) {
Decoder::DecodedBlockStatus Decoder::NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op, DecodedHeader Options) {
if (Info->Type == FEXCore::X86Tables::TYPE_ARCH_DISPATCHER) [[unlikely]] {
// Dispatcher Op.
// TODO: Move this in to `NormalOpHeader`, Dispatch tables have a bug currently where some subtables don't inherit flags correctly.
@@ -354,11 +354,16 @@ bool Decoder::NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op,
DecodeInst->TableInfo = Info;
if (Info->Type == FEXCore::X86Tables::TYPE_UNKNOWN) {
return false;
return DecodedBlockStatus::INVALID_INST;
}
if (Info->Type == FEXCore::X86Tables::TYPE_INVALID) {
return false;
return DecodedBlockStatus::INVALID_INST;
}
if (!(Info->Flags & FEXCore::X86Tables::InstFlags::FLAGS_SUPPORTS_LOCK) && (DecodeInst->Flags & DecodeFlags::FLAG_LOCK)) {
// Instruction has lock prefix but doesn't support lock.
return DecodedBlockStatus::UNIMPLEMENTED_INST;
}
LOGMAN_THROW_A_FMT(!(Info->Type >= FEXCore::X86Tables::TYPE_GROUP_1 && Info->Type <= FEXCore::X86Tables::TYPE_GROUP_P), "Group Ops "
@@ -390,15 +395,15 @@ bool Decoder::NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op,
const bool Has16BitAddressing = !BlockInfo.Is64BitMode && DecodeInst->Flags & DecodeFlags::FLAG_ADDRESS_SIZE;
if (Options.w && (Info->Flags & InstFlags::FLAGS_REX_W_0)) {
return false;
return DecodedBlockStatus::INVALID_INST;
} else if (!Options.w && (Info->Flags & InstFlags::FLAGS_REX_W_1)) {
return false;
return DecodedBlockStatus::INVALID_INST;
}
if (Options.L && (Info->Flags & InstFlags::FLAGS_VEX_L_0)) {
return false;
return DecodedBlockStatus::INVALID_INST;
} else if (!Options.L && (Info->Flags & InstFlags::FLAGS_VEX_L_1)) {
return false;
return DecodedBlockStatus::INVALID_INST;
}
const bool UseVEXL = Options.L && !(Info->Flags & InstFlags::FLAGS_VEX_L_IGNORE);
@@ -507,7 +512,7 @@ bool Decoder::NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op,
MapModRMToReg(DecodeInst->Flags & DecodeFlags::FLAG_REX_XGPR_B ? 1 : 0, Op & 0b111, Is8BitDest, HasREX, false, false);
if (CurrentDest->Data.GPR.GPR == FEXCore::X86State::REG_INVALID) {
return false;
return DecodedBlockStatus::INVALID_INST;
}
}
@@ -576,7 +581,7 @@ bool Decoder::NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op,
const auto VEXOperand = Info->Flags & FEXCore::X86Tables::InstFlags::FLAGS_VEX_SRC_MASK;
if (VEXOperand == FEXCore::X86Tables::InstFlags::FLAGS_VEX_NO_OPERAND && Options.vvvv) {
return false;
return DecodedBlockStatus::INVALID_INST;
}
if (VEXOperand == FEXCore::X86Tables::InstFlags::FLAGS_VEX_1ST_SRC) {
@@ -594,11 +599,11 @@ bool Decoder::NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op,
if (Info->Flags & FEXCore::X86Tables::InstFlags::FLAGS_MODRM) {
if (Info->Flags & FEXCore::X86Tables::InstFlags::FLAGS_SF_MOD_DST) {
if (!ModRMOperand(DecodeInst->Src[CurrentSrc], DecodeInst->Dest, HasXMMSrc, HasXMMDst, HasMMSrc, HasMMDst, Is8BitSrc, Is8BitDest)) {
return false;
return DecodedBlockStatus::INVALID_INST;
}
} else {
if (!ModRMOperand(DecodeInst->Dest, DecodeInst->Src[CurrentSrc], HasXMMDst, HasXMMSrc, HasMMDst, HasMMSrc, Is8BitDest, Is8BitSrc)) {
return false;
return DecodedBlockStatus::INVALID_INST;
}
}
++CurrentSrc;
@@ -660,22 +665,27 @@ bool Decoder::NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op,
Bytes = 0;
}
if ((DecodeInst->Flags & DecodeFlags::FLAG_LOCK) && DecodeInst->Dest.IsGPR()) {
// Instruction has lock prefix, but the destination isn't memory, this is invalid.
return DecodedBlockStatus::UNIMPLEMENTED_INST;
}
LOGMAN_THROW_A_FMT(Bytes == 0, "Inst at 0x{:x}: 0x{:04x} '{}' Had an instruction of size {} with {} remaining", DecodeInst->PC,
DecodeInst->OP, DecodeInst->TableInfo->Name ?: "UND", InstructionSize, Bytes);
DecodeInst->InstSize = InstructionSize;
return true;
return DecodedBlockStatus::SUCCESS;
}
bool Decoder::NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op) {
Decoder::DecodedBlockStatus Decoder::NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op) {
DecodeInst->OPRaw = DecodeInst->OP = Op;
DecodeInst->TableInfo = Info;
if (Info->Type == FEXCore::X86Tables::TYPE_UNKNOWN) {
return false;
return DecodedBlockStatus::INVALID_INST;
}
if (Info->Type == FEXCore::X86Tables::TYPE_INVALID) {
return false;
return DecodedBlockStatus::INVALID_INST;
}
LOGMAN_THROW_A_FMT(Info->Type != FEXCore::X86Tables::TYPE_REX_PREFIX, "REX PREFIX should have been decoded before this!");
@@ -732,7 +742,7 @@ bool Decoder::NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16
};
uint8_t Field = RegToField[ModRM.reg];
if (Field == 255) {
return false;
return DecodedBlockStatus::INVALID_INST;
}
LocalOp = (Field << 3) | ModRM.rm;
@@ -751,7 +761,7 @@ bool Decoder::NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16
} else if (Info->Type == FEXCore::X86Tables::TYPE_VEX_TABLE_PREFIX) {
if (!VEXTable) {
// AVX not enabled.
return false;
return DecodedBlockStatus::INVALID_INST;
}
uint16_t map_select = 1;
@@ -761,7 +771,7 @@ bool Decoder::NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16
if ((Byte1 & 0b10000000) == 0) {
if (!BlockInfo.Is64BitMode) {
return false;
return DecodedBlockStatus::INVALID_INST;
}
DecodeInst->Flags |= DecodeFlags::FLAG_REX_XGPR_R;
@@ -772,7 +782,7 @@ bool Decoder::NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16
const uint8_t vvvv = ((Byte1 & 0b01111000) >> 3);
if (!BlockInfo.Is64BitMode && vvvv <= 0b0111) {
// Invalid on 32-bit, can't use the high registers.
return false;
return DecodedBlockStatus::INVALID_INST;
}
options.vvvv = 15 - vvvv;
options.L = (Byte1 & 0b100) != 0;
@@ -783,14 +793,14 @@ bool Decoder::NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16
const uint8_t vvvv = ((Byte2 & 0b01111000) >> 3);
if (!BlockInfo.Is64BitMode && vvvv <= 0b0111) {
// Invalid on 32-bit, can't use the high registers.
return false;
return DecodedBlockStatus::INVALID_INST;
}
options.vvvv = 15 - vvvv;
options.w = (Byte2 & 0b10000000) != 0;
options.L = (Byte2 & 0b100) != 0;
if ((Byte1 & 0b01000000) == 0) {
if (!BlockInfo.Is64BitMode) {
return false;
return DecodedBlockStatus::INVALID_INST;
}
DecodeInst->Flags |= DecodeFlags::FLAG_REX_XGPR_X;
}
@@ -801,7 +811,7 @@ bool Decoder::NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16
DecodeInst->Flags |= DecodeFlags::FLAG_OPTION_AVX_W;
}
if (!(map_select >= 1 && map_select <= 3)) {
return false;
return DecodedBlockStatus::INVALID_INST;
}
}
@@ -831,14 +841,14 @@ bool Decoder::NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16
} else if (Info->Type == FEXCore::X86Tables::TYPE_GROUP_EVEX) {
FEXCORE_TELEMETRY_SET(TYPE_USES_EVEX_OPS, 1);
// EVEX unsupported
return false;
return DecodedBlockStatus::INVALID_INST;
}
LOGMAN_MSG_A_FMT("Invalid instruction decoding type");
FEX_UNREACHABLE;
}
bool Decoder::DecodeInstructionImpl(uint64_t PC) {
Decoder::DecodedBlockStatus Decoder::DecodeInstructionImpl(uint64_t PC) {
InstructionSize = 0;
LastEscapePrefix = 0;
Instruction.fill(0);
@@ -849,7 +859,7 @@ bool Decoder::DecodeInstructionImpl(uint64_t PC) {
for (;;) {
if (InstructionSize >= MAX_INST_SIZE) {
return false;
return DecodedBlockStatus::INVALID_INST;
}
uint8_t Op = ReadByte();
switch (Op) {
@@ -1035,10 +1045,10 @@ bool Decoder::DecodeInstructionImpl(uint64_t PC) {
}
if (DecodeInst->Dest.IsGPR()) {
return false;
return DecodedBlockStatus::INVALID_INST;
}
return true;
return DecodedBlockStatus::SUCCESS;
}
void Decoder::DecodeREXIfValid(int8_t ExpectedOffset) {
@@ -1076,18 +1086,26 @@ Decoder::DecodedBlockStatus Decoder::DecodeInstruction(uint64_t PC) {
// Will be set if DecodeInstructionImpl tries to read non-executable memory
HitNonExecutableRange = false;
HitBadRelocation = false;
bool ErrorDuringDecoding = !DecodeInstructionImpl(PC);
auto ErrorDuringDecoding = DecodeInstructionImpl(PC);
if (ErrorDuringDecoding || HitNonExecutableRange || HitBadRelocation) [[unlikely]] {
if (ErrorDuringDecoding != DecodedBlockStatus::SUCCESS || HitNonExecutableRange || HitBadRelocation) [[unlikely]] {
// Put an invalid instruction in the stream so the core can raise SIGILL if hit
// Error while decoding instruction. We don't know the table or instruction size
const auto InstSize = DecodeInst->InstSize;
DecodeInst->TableInfo = nullptr;
auto Result = ErrorDuringDecoding ? DecodedBlockStatus::INVALID_INST :
DecodeInst->InstSize ? DecodedBlockStatus::PARTIAL_DECODE_INST :
HitNonExecutableRange ? DecodedBlockStatus::NOEXEC_INST :
DecodedBlockStatus::BAD_RELOCATION;
DecodeInst->InstSize = 0;
return Result;
// A decode error can be caused by substituting zero for an inaccessible
// instruction byte, so the instruction fetch fault takes priority.
if (HitNonExecutableRange) {
return InstSize ? DecodedBlockStatus::PARTIAL_DECODE_INST : DecodedBlockStatus::NOEXEC_INST;
}
if (HitBadRelocation) {
return DecodedBlockStatus::BAD_RELOCATION;
}
return ErrorDuringDecoding;
} else if (!DecodeInst->TableInfo || (DecodeInst->TableInfo->Type == TYPE_INST && !DecodeInst->TableInfo->OpcodeDispatcher.OpDispatch)) {
// If there wasn't an error during decoding but we have no dispatcher for the instruction then claim invalid instruction.
return DecodedBlockStatus::INVALID_INST;
@@ -1331,7 +1349,7 @@ void Decoder::AddBranchTarget(uint64_t Target) {
}
}
const uint8_t* Decoder::AdjustAddrForSpecialRegion(const uint8_t* _InstStream, uint64_t EntryPoint, uint64_t RIP) {
const Decoder::DecodeStream Decoder::AdjustAddrForSpecialRegion(const uint8_t* _InstStream, uint64_t EntryPoint, uint64_t RIP) {
constexpr uint64_t VSyscall_Base = 0xFFFF'FFFF'FF60'0000ULL;
constexpr uint64_t VSyscall_End = VSyscall_Base + 0x1000;
@@ -1342,10 +1360,23 @@ const uint8_t* Decoder::AdjustAddrForSpecialRegion(const uint8_t* _InstStream, u
// Offset 0x400: vtime
// Offset 0x800: vgetcpu
uint64_t Offset = RIP - VSyscall_Base;
return VSyscallData + Offset;
return DecodeStream {
.InstStream = _InstStream - EntryPoint + RIP,
.AdjustedInstStream = VSyscallData + Offset,
};
}
return _InstStream - EntryPoint + RIP;
return DecodeStream {
.InstStream = _InstStream - EntryPoint + RIP,
.AdjustedInstStream = _InstStream - EntryPoint + RIP,
};
}
bool Decoder::CheckIfCacheable(FEXCore::Core::InternalThreadState& Thread, const uint8_t* InstStream, uint64_t PC, uint64_t MaxInst) {
DecodeInstructionsAtEntry(&Thread, InstStream, PC, MaxInst);
bool Uncacheable = HitBadRelocation;
DelayedDisownBuffer();
return !Uncacheable;
}
void Decoder::DecodeInstructionsAtEntry(FEXCore::Core::InternalThreadState* Thread, const uint8_t* _InstStream, uint64_t PC, uint64_t MaxInst) {
@@ -1366,7 +1397,6 @@ void Decoder::DecodeInstructionsAtEntry(FEXCore::Core::InternalThreadState* Thre
EntryPoint = PC;
BlockInfo.EntryPoints = {PC};
InstStream = _InstStream;
uint64_t TotalInstructions {};
@@ -1465,6 +1495,13 @@ void Decoder::DecodeInstructionsAtEntry(FEXCore::Core::InternalThreadState* Thre
}
BlockIt->BlockStatus = DecodeInstruction(OpAddress);
if (HitBadRelocation) {
BlockInfo.TotalInstructionCount = 0;
BlockInfo.Blocks = {*BlockIt};
BlockInfo.EntryPoints.clear();
BlockInfo.CodePages.clear();
return;
}
uint64_t OpEndAddress = OpAddress + DecodeInst->InstSize;
DecodedMinAddress = std::min(DecodedMinAddress, OpAddress);
@@ -1483,7 +1520,7 @@ void Decoder::DecodeInstructionsAtEntry(FEXCore::Core::InternalThreadState* Thre
// Can not continue this block at all on invalid instruction
if (BlockIt->BlockStatus != DecodedBlockStatus::SUCCESS) [[unlikely]] {
if (!EntryBlock) {
if (!EntryBlock && BlockIt->BlockStatus != DecodedBlockStatus::BAD_RELOCATION) {
// In multiblock configurations, we can early terminate any non-entrypoint blocks with the expectation that this won't get hit.
// Improves compile-times.
// Just need to undo additions that this block decoding has caused.
@@ -1493,10 +1530,11 @@ void Decoder::DecodeInstructionsAtEntry(FEXCore::Core::InternalThreadState* Thre
EraseBlock = true;
} else {
LogMan::Msg::EFmt("{} instruction in entry block: {:X}",
BlockIt->BlockStatus == DecodedBlockStatus::INVALID_INST ? "Invalid" :
BlockIt->BlockStatus == DecodedBlockStatus::NOEXEC_INST ? "NoExec" :
BlockIt->BlockStatus == DecodedBlockStatus::BAD_RELOCATION ? "BadRelocation" :
"PartialDecode",
BlockIt->BlockStatus == DecodedBlockStatus::INVALID_INST ? "Invalid" :
BlockIt->BlockStatus == DecodedBlockStatus::NOEXEC_INST ? "NoExec" :
BlockIt->BlockStatus == DecodedBlockStatus::BAD_RELOCATION ? "BadRelocation" :
BlockIt->BlockStatus == DecodedBlockStatus::UNIMPLEMENTED_INST ? "Unimplemented" :
"PartialDecode",
OpAddress);
}
break;
+28 -5
View File
@@ -32,6 +32,7 @@ public:
NOEXEC_INST,
PARTIAL_DECODE_INST,
BAD_RELOCATION,
UNIMPLEMENTED_INST,
};
// New Frontend decoding
@@ -54,6 +55,8 @@ public:
};
Decoder(FEXCore::Core::InternalThreadState* Thread);
bool CheckIfCacheable(FEXCore::Core::InternalThreadState&, const uint8_t* InstStream, uint64_t PC, uint64_t MaxInst);
void DecodeInstructionsAtEntry(FEXCore::Core::InternalThreadState* Thread, const uint8_t* InstStream, uint64_t PC, uint64_t MaxInst);
const DecodedBlockInformation* GetDecodedBlockInfo() const {
@@ -90,7 +93,7 @@ private:
FEX_CONFIG_OPT(EnableCodeCacheValidation, ENABLECODECACHEVALIDATION);
bool DecodeInstructionImpl(uint64_t PC);
DecodedBlockStatus DecodeInstructionImpl(uint64_t PC);
DecodedBlockStatus DecodeInstruction(uint64_t PC);
void BranchTargetInMultiblockRange();
@@ -109,8 +112,8 @@ private:
InstructionSize += Size;
}
bool NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op, DecodedHeader Options = {});
bool NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op);
DecodedBlockStatus NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op, DecodedHeader Options = {});
DecodedBlockStatus NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op);
void DecodeREXIfValid(int8_t ExpectedOffset = -1);
@@ -125,7 +128,27 @@ private:
bool HitNonExecutableRange {};
bool HitBadRelocation {};
const uint8_t* InstStream {};
struct DecodeStream {
// Original instruction stream RIP location.
const uint8_t* InstStream;
// Adjusted location for FEX actually decodes from.
const uint8_t* AdjustedInstStream;
DecodeStream& operator-=(size_t offset) noexcept {
InstStream -= offset;
AdjustedInstStream -= offset;
return *this;
}
DecodeStream& operator+=(size_t offset) noexcept {
InstStream += offset;
AdjustedInstStream += offset;
return *this;
}
};
DecodeStream InstStream;
IR::OpSize GetGPROpSize() const {
return BlockInfo.Is64BitMode ? IR::OpSize::i64Bit : IR::OpSize::i32Bit;
}
@@ -167,6 +190,6 @@ private:
const std::array<X86Tables::X86InstInfo, X86Tables::MAX_VEX_TABLE_SIZE>* VEXTable {};
const std::array<X86Tables::X86InstInfo, X86Tables::MAX_VEX_GROUP_TABLE_SIZE>* VEXTableGroup {};
const uint8_t* AdjustAddrForSpecialRegion(const uint8_t* _InstStream, uint64_t EntryPoint, uint64_t RIP);
const DecodeStream AdjustAddrForSpecialRegion(const uint8_t* _InstStream, uint64_t EntryPoint, uint64_t RIP);
};
} // namespace FEXCore::Frontend
@@ -302,6 +302,16 @@ struct OpHandlers<IR::OP_F80FYL2X> {
}
};
template<>
struct OpHandlers<IR::OP_F80FYL2XP1> {
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1, VectorRegType Src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
ScopedSoftFloatState State {FCW, Frame, true};
const X80SoftFloat One {&State.State, 1.0};
return X80SoftFloat::FYL2X(&State.State, X80SoftFloat::FADD(&State.State, Src1, One), Src2);
}
};
template<>
struct OpHandlers<IR::OP_F80ATAN> {
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle(uint16_t FCW, VectorRegType Src1, VectorRegType Src2, FEXCore::Core::CpuStateFrame* Frame) {
@@ -417,6 +427,14 @@ struct OpHandlers<IR::OP_F64FYL2X> {
}
};
template<>
struct OpHandlers<IR::OP_F64FYL2XP1> {
FEXCORE_PRESERVE_ALL_ATTR static double handle(double src1, double src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
return src2 * log2(1.0 + src1);
}
};
template<>
struct OpHandlers<IR::OP_F64SCALE> {
FEXCORE_PRESERVE_ALL_ATTR static double handle(double src1, double src2, FEXCore::Core::CpuStateFrame* Frame) {
@@ -72,6 +72,8 @@ void InterpreterOps::FillFallbackIndexPointers(Core::FallbackABIInfo* Info, uint
reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80DIV>::handle)};
Info[Core::OPINDEX_F80FYL2X] = {ABIHandlers[FABI_F80_I16_F80_F80_PTR],
reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80FYL2X>::handle)};
Info[Core::OPINDEX_F80FYL2XP1] = {ABIHandlers[FABI_F80_I16_F80_F80_PTR],
reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80FYL2XP1>::handle)};
Info[Core::OPINDEX_F80ATAN] = {ABIHandlers[FABI_F80_I16_F80_F80_PTR],
reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80ATAN>::handle)};
Info[Core::OPINDEX_F80FPREM1] = {ABIHandlers[FABI_F80_I16_F80_F80_PTR],
@@ -97,6 +99,8 @@ void InterpreterOps::FillFallbackIndexPointers(Core::FallbackABIInfo* Info, uint
reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F64FPREM1>::handle)};
Info[Core::OPINDEX_F64FYL2X] = {ABIHandlers[FABI_F64_F64_F64_PTR],
reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F64FYL2X>::handle)};
Info[Core::OPINDEX_F64FYL2XP1] = {ABIHandlers[FABI_F64_F64_F64_PTR],
reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F64FYL2XP1>::handle)};
Info[Core::OPINDEX_F64SCALE] = {ABIHandlers[FABI_F64_F64_F64_PTR],
reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F64SCALE>::handle)};
@@ -254,6 +258,7 @@ bool InterpreterOps::GetFallbackHandler(const IR::IROp_Header* IROp, FallbackInf
COMMON_BINARY_X87_OP(MUL)
COMMON_BINARY_X87_OP(DIV)
COMMON_BINARY_X87_OP(FYL2X)
COMMON_BINARY_X87_OP(FYL2XP1)
COMMON_BINARY_X87_OP(ATAN)
COMMON_BINARY_X87_OP(FPREM1)
COMMON_BINARY_X87_OP(FPREM)
@@ -268,6 +273,7 @@ bool InterpreterOps::GetFallbackHandler(const IR::IROp_Header* IROp, FallbackInf
// Double Precision Binary
COMMON_BINARY_F64_OP(FYL2X)
COMMON_BINARY_F64_OP(FYL2XP1)
COMMON_BINARY_F64_OP(ATAN)
COMMON_BINARY_F64_OP(FPREM1)
COMMON_BINARY_F64_OP(FPREM)
+6 -9
View File
@@ -13,9 +13,6 @@ $end_info$
namespace FEXCore::CPU {
#define GRD(Node) (IROp->Size <= 4 ? GetDst<RA_32>(Node) : GetDst<RA_64>(Node))
#define GRS(Node) (IROp->Size <= 4 ? GetReg<RA_32>(Node) : GetReg<RA_64>(Node))
#define DEF_BINOP_WITH_CONSTANT(FEXOp, VarOp, ConstOp) \
DEF_OP(FEXOp) { \
auto Op = IROp->C<IR::IROp_##FEXOp>(); \
@@ -274,7 +271,7 @@ DEF_OP(CmpPairZ) {
// Restore NzCV
if (CTX->HostFeatures.SupportsFlagM) {
rmif(TMP1, 0, 0xb /* NzCV */);
rmif(TMP1, 28, 0xb /* NzCV */);
} else {
cset(ARMEmitter::Size::i32Bit, TMP2, ARMEmitter::Condition::CC_EQ);
bfi(ARMEmitter::Size::i32Bit, TMP1, TMP2, 30 /* lsb: Z */, 1);
@@ -421,8 +418,8 @@ DEF_OP(MulH) {
if (OpSize == IR::OpSize::i32Bit) {
sxtw(TMP1, Src1.W());
sxtw(TMP2, Src2.W());
mul(ARMEmitter::Size::i32Bit, Dst, TMP1, TMP2);
ubfx(ARMEmitter::Size::i32Bit, Dst, Dst, 32, 32);
mul(ARMEmitter::Size::i64Bit, Dst, TMP1, TMP2);
ubfx(ARMEmitter::Size::i64Bit, Dst, Dst, 32, 32);
} else {
smulh(Dst.X(), Src1.X(), Src2.X());
}
@@ -523,7 +520,7 @@ DEF_OP(AndWithFlags) {
}
DEF_OP(AndShift) {
auto Op = IROp->C<IR::IROp_XorShift>();
auto Op = IROp->C<IR::IROp_AndShift>();
and_(ConvertSize48(IROp), GetReg(Node), GetReg(Op->Src1), GetReg(Op->Src2), ConvertIRShiftType(Op->Shift), Op->ShiftAmount);
}
@@ -721,7 +718,7 @@ DEF_OP(Extr) {
}
DEF_OP(PDep) {
auto Op = IROp->C<IR::IROp_PExt>();
auto Op = IROp->C<IR::IROp_PDep>();
const auto EmitSize = ConvertSize48(IROp);
const auto Dest = GetReg(Node);
@@ -774,7 +771,7 @@ DEF_OP(PDep) {
// Now, they're copied, so we can start setting Dest (even if it overlaps with
// one of them). Handle early exit case
mov(EmitSize, Dest, 0);
(void)cbz(EmitSize, OrigMask, &Done);
(void)cbz(EmitSize, Mask, &Done);
// Setup for first iteration
neg(EmitSize, T0, Mask);
@@ -329,7 +329,7 @@ DEF_OP(TelemetrySetValue) {
auto Op = IROp->C<IR::IROp_TelemetrySetValue>();
auto Src = GetReg(Op->Value);
ldr(TMP2, STATE_PTR(CpuStateFrame, Pointers.TelemetryValueAddresses[Op->TelemetryValueIndex]));
ldr(TMP2, STATE_PTR_IDX(CpuStateFrame, Pointers.TelemetryValueAddresses, Op->TelemetryValueIndex));
// Cortex fuses cmp+cset.
cmp(ARMEmitter::Size::i32Bit, Src, 0);
@@ -342,8 +342,8 @@ DEF_OP(TelemetrySetValue) {
(void)Bind(&LoopTop);
ldaxr(ARMEmitter::SubRegSize::i64Bit, TMP3, TMP2);
orr(ARMEmitter::Size::i32Bit, TMP3, TMP3, Src);
stlxr(ARMEmitter::SubRegSize::i64Bit, TMP3, TMP3, TMP2);
(void)cbnz(ARMEmitter::Size::i32Bit, TMP3, &LoopTop);
stlxr(ARMEmitter::SubRegSize::i64Bit, TMP4, TMP3, TMP2);
(void)cbnz(ARMEmitter::Size::i32Bit, TMP4, &LoopTop);
}
#endif
}
@@ -55,6 +55,29 @@ DEF_OP(ExitFunction) {
uint64_t NewRIP;
if constexpr (Context::BLOCK_DEBUGGING) {
// Skip block linking when BLOCK_DEBUGGING as it adds overhead and is unncessary.
// This is a debug only feature and doesn't need caching help.
bool IsInlineRIP = IsInlineConstant(Op->NewRIP, &NewRIP) || IsInlineEntrypointOffset(Op->NewRIP, &NewRIP);
ARMEmitter::ForwardLabel l_ExitLink;
if (IsInlineRIP) {
ldr(TMP1, &l_ExitLink);
str(TMP1, STATE, offsetof(FEXCore::Core::CpuStateFrame, State.rip));
} else {
auto RipReg = GetReg(Op->NewRIP);
str(RipReg.X(), STATE, offsetof(FEXCore::Core::CpuStateFrame, State.rip));
}
ldr(TMP2, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.DispatcherLoopTop));
br(TMP2);
if (IsInlineRIP) {
BindOrRestart(&l_ExitLink);
dc64(NewRIP);
}
return;
}
if (IsInlineConstant(Op->NewRIP, &NewRIP) || IsInlineEntrypointOffset(Op->NewRIP, &NewRIP)) {
#ifdef ARCHITECTURE_arm64ec
if (NewRIP < EC_CODE_BITMAP_MAX_ADDRESS && RtlIsEcCode(NewRIP)) {
@@ -265,7 +288,10 @@ DEF_OP(Syscall) {
uint32_t GPRSpillMask = ~0U;
uint32_t FPRSpillMask = ~0U;
SpillStaticRegs(TMP1, true, GPRSpillMask, FPRSpillMask);
SpillStaticRegs(TMP1, {
.GPRSpillMask = GPRSpillMask,
.FPRSpillMask = FPRSpillMask,
});
// Now that we are spilled, store in the state that we are in a syscall
// Still without overwriting registers that matter
@@ -299,7 +325,12 @@ DEF_OP(Syscall) {
// Result is now in x0
// Fix the stack and any values that were stepped on
FillStaticRegs(true, GPRSpillMask, FPRSpillMask, ARMEmitter::Reg::r1, ARMEmitter::Reg::r2);
FillStaticRegs({
.OptionalReg = ARMEmitter::Reg::r1,
.OptionalReg2 = ARMEmitter::Reg::r2,
.GPRFillMask = GPRSpillMask,
.FPRFillMask = FPRSpillMask,
});
// Now the registers we've spilled are back in their original host registers
// We can safely claim we are no longer in a syscall
@@ -322,7 +353,10 @@ DEF_OP(Thunk) {
// X0: CTX
// X1: Args (from guest stack)
SpillStaticRegs(TMP1, true, ~0U, ~0U, false); // spill to ctx before ra64 spill
// spill to ctx before ra64 spill
SpillStaticRegs(TMP1, {
.NZCV = false,
});
PushDynamicRegs(TMP1);
@@ -337,7 +371,10 @@ DEF_OP(Thunk) {
PopDynamicRegs();
FillStaticRegs(true, ~0U, ~0U, std::nullopt, std::nullopt, false); // load from ctx after ra64 refill
// load from ctx after ra64 refill
FillStaticRegs({
.NZCV = false,
});
}
DEF_OP(ValidateCode) {
@@ -292,8 +292,6 @@ DEF_OP(Vector_FToS) {
frinti(SubEmitSize, Dst.Z(), Mask.Merging(), Vector.Z());
fcvtzs(Dst.Z(), SubEmitSize, Mask.Merging(), Dst.Z(), SubEmitSize);
} else {
const auto Dst = GetVReg(Node);
const auto Vector = GetVReg(Op->Vector);
if (OpSize == IR::OpSize::i64Bit) {
frinti(SubEmitSize, Dst.D(), Vector.D());
fcvtzs(SubEmitSize, Dst.D(), Dst.D());
@@ -324,24 +324,62 @@ DEF_OP(PCLMUL) {
const auto Op = IROp->C<IR::IROp_PCLMUL>();
const auto OpSize = IROp->Size;
const auto Is256Bit = OpSize == IR::OpSize::i256Bit;
LOGMAN_THROW_A_FMT(!Is256Bit || HostSupportsSVE256, "Need SVE256 support in order to use {} with 256-bit operation", __func__);
const auto Dst = GetVReg(Node);
const auto Src1 = GetVReg(Op->Src1);
const auto Src2 = GetVReg(Op->Src2);
LOGMAN_THROW_A_FMT(OpSize == IR::OpSize::i128Bit, "Currently only supports 128-bit operations.");
switch (Op->Selector) {
case 0b00000000: pmull(ARMEmitter::SubRegSize::i128Bit, Dst.D(), Src1.D(), Src2.D()); break;
case 0b00000001:
dup(ARMEmitter::SubRegSize::i64Bit, VTMP1.Q(), Src1.Q(), 1);
pmull(ARMEmitter::SubRegSize::i128Bit, Dst.D(), VTMP1.D(), Src2.D());
break;
case 0b00010000:
dup(ARMEmitter::SubRegSize::i64Bit, VTMP1.Q(), Src2.Q(), 1);
pmull(ARMEmitter::SubRegSize::i128Bit, Dst.D(), VTMP1.D(), Src1.D());
break;
case 0b00010001: pmull2(ARMEmitter::SubRegSize::i128Bit, Dst.Q(), Src1.Q(), Src2.Q()); break;
default: LOGMAN_MSG_A_FMT("Unknown PCLMUL selector: {}", Op->Selector); break;
if (HostSupportsSVE256 && Is256Bit) {
switch (Op->Selector) {
case 0b00000000: {
pmullb(ARMEmitter::SubRegSize::i128Bit, Dst.Z(), Src1.Z(), Src2.Z());
break;
}
case 0b00000001: {
trn2(ARMEmitter::SubRegSize::i64Bit, VTMP1.Z(), Src1.Z(), Src1.Z());
pmullb(ARMEmitter::SubRegSize::i128Bit, Dst.Z(), VTMP1.Z(), Src2.Z());
break;
}
case 0b00010000:
trn2(ARMEmitter::SubRegSize::i64Bit, VTMP1.Z(), Src2.Z(), Src2.Z());
pmullb(ARMEmitter::SubRegSize::i128Bit, Dst.Z(), Src1.Z(), VTMP1.Z());
break;
case 0b00010001: {
pmullt(ARMEmitter::SubRegSize::i128Bit, Dst.Z(), Src1.Z(), Src2.Z());
break;
}
default: {
LOGMAN_MSG_A_FMT("Unknown PCLMUL selector: {}", Op->Selector);
break;
}
}
} else {
switch (Op->Selector) {
case 0b00000000: {
pmull(ARMEmitter::SubRegSize::i128Bit, Dst.D(), Src1.D(), Src2.D());
break;
}
case 0b00000001: {
dup(ARMEmitter::SubRegSize::i64Bit, VTMP1.Q(), Src1.Q(), 1);
pmull(ARMEmitter::SubRegSize::i128Bit, Dst.D(), VTMP1.D(), Src2.D());
break;
}
case 0b00010000: {
dup(ARMEmitter::SubRegSize::i64Bit, VTMP1.Q(), Src2.Q(), 1);
pmull(ARMEmitter::SubRegSize::i128Bit, Dst.D(), VTMP1.D(), Src1.D());
break;
}
case 0b00010001: {
pmull2(ARMEmitter::SubRegSize::i128Bit, Dst.Q(), Src1.Q(), Src2.Q());
break;
}
default: {
LOGMAN_MSG_A_FMT("Unknown PCLMUL selector: {}", Op->Selector);
break;
}
}
}
}
+63 -57
View File
@@ -68,6 +68,10 @@ PrintValue(uint64_t Value) {
LogMan::Msg::DFmt("Value: 0x{:x}", Value);
}
static void PrintMsg(const char* Value) {
LogMan::Msg::DFmt("{}", Value);
}
static void PrintVectorValue(uint64_t Value, uint64_t ValueUpper) {
LogMan::Msg::DFmt("Value: 0x{:016x}'{:016x}", ValueUpper, Value);
}
@@ -133,8 +137,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
const auto Src1 = GetVReg(IROp->Args[0]);
fmov(VTMP1.S(), Src1.S());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -151,8 +155,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
const auto Src1 = GetVReg(IROp->Args[0]);
fmov(VTMP1.D(), Src1.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -176,8 +180,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
mov(ARMEmitter::Size::i32Bit, TMP2, Src1);
}
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -194,8 +198,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
const auto Src1 = GetVReg(IROp->Args[0]);
mov(VTMP1.Q(), Src1.Q());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -212,8 +216,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
const auto Src1 = GetVReg(IROp->Args[0]);
mov(VTMP1.Q(), Src1.Q());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -230,8 +234,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
const auto Src1 = GetVReg(IROp->Args[0]);
fmov(VTMP1.D(), Src1.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -254,8 +258,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
fmov(VTMP1.D(), Src1.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -276,8 +280,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
fmov(VTMP1.D(), Src1.D());
fmov(VTMP2.D(), Src2.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -294,8 +298,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
const auto Src1 = GetVReg(IROp->Args[0]);
mov(VTMP1.Q(), Src1.Q());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -312,8 +316,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
const auto Src1 = GetVReg(IROp->Args[0]);
mov(VTMP1.Q(), Src1.Q());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -330,8 +334,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
const auto Src1 = GetVReg(IROp->Args[0]);
mov(VTMP1.Q(), Src1.Q());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -351,8 +355,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
mov(VTMP1.Q(), Src1.Q());
mov(VTMP2.Q(), Src2.Q());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -369,8 +373,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
const auto Src1 = GetVReg(IROp->Args[0]);
mov(VTMP1.Q(), Src1.Q());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -394,8 +398,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
mov(VTMP1.Q(), Src1.Q());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -416,8 +420,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
mov(VTMP1.Q(), Src1.Q());
mov(VTMP2.Q(), Src2.Q());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -434,8 +438,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
// tmp2 (x1/x11): source 2
// tmp3 (x2/x12): source 3
const auto Op = IROp->C<IR::IROp_VPCMPESTRX>();
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
stp<ARMEmitter::IndexType::PRE>(TMP1, ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
@@ -476,8 +480,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
mov(VTMP2.Q(), Src2.Q());
movz(ARMEmitter::Size::i32Bit, TMP1, Control);
ldr(TMP2, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP2, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP2);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -618,7 +622,7 @@ Arm64JITCore::Arm64JITCore(FEXCore::Context::ContextImpl* ctx, FEXCore::Core::In
, HostSupportsRPRES {ctx->HostFeatures.SupportsRPRES}
, HostSupportsAFP {ctx->HostFeatures.SupportsAFP}
, CTX {ctx}
, TempAllocator(ctx->CPUBackendAllocator, 0) {
, TempCodeBufferAllocator(ctx->CPUBackendAllocator, 0) {
RAPass = Thread->PassManager->GetPass<IR::RegisterAllocationPass>("RA");
@@ -626,7 +630,7 @@ Arm64JITCore::Arm64JITCore(FEXCore::Context::ContextImpl* ctx, FEXCore::Core::In
RAPass->AddRegisters(IR::RegClass::GPRFixed, StaticRegisters.size());
RAPass->AddRegisters(IR::RegClass::FPR, GeneralFPRegisters.size());
RAPass->AddRegisters(IR::RegClass::FPRFixed, StaticFPRegisters.size());
RAPass->PairRegs = PairRegisters;
RAPass->SetNumPairRegs(PairRegisters);
{
// Set up pointers that the JIT needs to load
@@ -636,6 +640,8 @@ Arm64JITCore::Arm64JITCore(FEXCore::Context::ContextImpl* ctx, FEXCore::Core::In
Ptrs.PrintValue = reinterpret_cast<uint64_t>(PrintValue);
Ptrs.PrintVectorValue = reinterpret_cast<uint64_t>(PrintVectorValue);
Ptrs.PrintMsgValue = reinterpret_cast<uint64_t>(PrintMsg);
Ptrs.ThreadRemoveCodeEntryFromJIT = reinterpret_cast<uintptr_t>(&Context::ContextImpl::ThreadRemoveCodeEntryFromJit);
Ptrs.MonoBackpatcherWrite = reinterpret_cast<uint64_t>(&Context::ContextImpl::MonoBackpatcherWrite);
Ptrs.CPUIDObj = reinterpret_cast<uint64_t>(&CTX->CPUID);
@@ -660,24 +666,17 @@ Arm64JITCore::Arm64JITCore(FEXCore::Context::ContextImpl* ctx, FEXCore::Core::In
Ptrs.LDIV = reinterpret_cast<uint64_t>(LDIV);
}
CurrentCodeBuffer = CodeBuffers.GetLatest();
CurrentCodeBuffer = SharedCodeBuffers.GetLatest();
ThreadState->LookupCache->Shared = CurrentCodeBuffer->LookupCache.get();
}
void Arm64JITCore::EmitDetectionString() {
const char JITString[] = "FEXJIT::Arm64JITCore::";
EmitString(JITString);
Align();
}
void Arm64JITCore::ClearCache() {
// NOTE: Holding on to the reference here is required to ensure validity of the WriteLock mutex
auto PrevCodeBuffer = CurrentCodeBuffer;
auto lk = PrevCodeBuffer->LookupCache->AcquireWriteLock();
auto CodeBuffer = GetEmptyCodeBuffer();
auto CodeBuffer = GetEmptySharedCodeBuffer();
SetBuffer(CodeBuffer->Ptr, CodeBuffer->AllocatedSize);
EmitDetectionString();
ThreadState->LookupCache->ChangeGuestToHostMapping(*PrevCodeBuffer, *CurrentCodeBuffer->LookupCache, lk);
}
@@ -770,8 +769,15 @@ void Arm64JITCore::EmitSuspendInterruptCheck() {
if (CTX->Config.NeedsPendingInterruptFaultCheck) {
// Trigger a fault if there are any pending interrupts
// Used only for suspend on WIN32 at the moment
strb(ARMEmitter::XReg::zr, STATE,
offsetof(FEXCore::Core::InternalThreadState, InterruptFaultPage) - offsetof(FEXCore::Core::InternalThreadState, BaseFrameState));
constexpr size_t InterruptPageOffset =
offsetof(FEXCore::Core::InternalThreadState, InterruptFaultPage) - offsetof(FEXCore::Core::InternalThreadState, BaseFrameState);
if constexpr (InterruptPageOffset <= 32760) {
str(ARMEmitter::XReg::zr, STATE, InterruptPageOffset);
} else {
// Need to use vector 128-bit store for this range.
// Doesn't matter which register we use to store.
str(ARMEmitter::QReg::q0, STATE, InterruptPageOffset);
}
}
#ifdef ARCHITECTURE_arm64ec
@@ -830,7 +836,7 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
case RestartOptions::Control::EnableFarARM64Jumps: RequiresFarARM64Jumps = true; break;
case RestartOptions::Control::NeedsLargerJITSpace:
// Get rid of the claimed buffer immediately, we can't fit in it at all.
TempAllocator.UnclaimBuffer();
TempCodeBufferAllocator.UnclaimBuffer();
SSANodeMultiplier *= 2;
break;
default: LOGMAN_MSG_A_FMT("Unhandled Arm64 restart condition!");
@@ -841,6 +847,7 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
CallReturnTargets.clear();
PendingJumpThunks.clear();
JumpTargets.resize(IR->GetHeader()->BlockCount, {});
Relocations.resize(PrevNumAllocations, FEXCore::CPU::Relocation::Default()); // Discard any relocations generated from a previous attempt
CodeData.EntryPoints.clear();
@@ -850,7 +857,7 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
// JIT output is first written to a temporary buffer and later relocated to the CodeBuffer.
// This minimizes lock contention of CodeBufferWriteMutex.
auto TempCodeBufferInfo = TempAllocator.ReownOrClaimBufferWithSize(DesiredBufferRange);
auto TempCodeBufferInfo = TempCodeBufferAllocator.ReownOrClaimBufferWithSize(DesiredBufferRange);
auto TempCodeBuffer = TempCodeBufferInfo.Ptr;
const uint32_t UsableBufferRange = TempCodeBufferInfo.Size - FEXCore::Utils::FEX_PAGE_SIZE;
@@ -895,7 +902,6 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
PendingCallReturnTargetLabel = nullptr;
for (auto [BlockNode, BlockHeader] : IR->GetBlocks()) {
using namespace FEXCore::IR;
auto BlockIROp = BlockHeader->CW<FEXCore::IR::IROp_CodeBlock>();
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
LOGMAN_THROW_A_FMT(BlockIROp->Header.Op == IR::OP_CODEBLOCK, "IR type failed to be a code block");
@@ -1066,7 +1072,7 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
// Migrate the compile output from temporary storage to the actual CodeBuffer.
// This can block progress in other compiling threads, so the duration of the lock should be as small as possible.
{
auto CodeBufferLock = std::unique_lock {CodeBuffers.CodeBufferWriteMutex};
auto CodeBufferLock = std::unique_lock {SharedCodeBuffers.CodeBufferWriteMutex};
// Query size of generated code
const auto TempSize = GetCursorOffset();
@@ -1083,13 +1089,13 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
// NOTE: 16-byte alignment of the new cursor offset must be preserved for block linking records
SetBuffer(CurrentCodeBuffer->Ptr, CurrentCodeBuffer->AllocatedSize);
SetCursorOffset(CodeBuffers.LatestOffset);
SetCursorOffset(SharedCodeBuffers.LatestOffset);
Align16B();
if ((GetCursorOffset() + TempSize) > CurrentCodeBuffer->UsableSize()) {
CTX->ClearCodeCache(ThreadState);
}
CodeBuffers.LatestOffset = GetCursorOffset();
SharedCodeBuffers.LatestOffset = GetCursorOffset();
}
// Adjust host addresses
@@ -1101,17 +1107,17 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
CodeBegin += Delta;
for (std::size_t Idx = PrevNumAllocations; Idx != Relocations.size(); ++Idx) {
Relocations[Idx].Header.Offset += CodeBuffers.LatestOffset;
Relocations[Idx].Header.Offset += SharedCodeBuffers.LatestOffset;
}
// Copy over CodeBuffer contents
memcpy(GetCursorAddress<uint8_t*>(), TempCodeBuffer, TempSize);
SetCursorOffset(CodeBuffers.LatestOffset + TempSize);
SetCursorOffset(SharedCodeBuffers.LatestOffset + TempSize);
CodeBuffers.LatestOffset = GetCursorOffset();
SharedCodeBuffers.LatestOffset = GetCursorOffset();
}
TempAllocator.DelayedDisownBuffer();
TempCodeBufferAllocator.DelayedDisownBuffer();
ClearICache(CodeBegin, CodeOnlySize);
+1 -3
View File
@@ -105,7 +105,7 @@ private:
};
fextl::vector<PendingJumpThunk> PendingJumpThunks;
Utils::PoolBufferWithTimedRetirement<uint8_t*, 5000, 500> TempAllocator;
Utils::PoolBufferWithTimedRetirement<uint8_t*, 5000, 500> TempCodeBufferAllocator;
static uint64_t ExitFunctionLink(FEXCore::Core::CpuStateFrame* Frame, FEXCore::Context::ExitFunctionLinkData* Record);
@@ -526,8 +526,6 @@ private:
FEXCore::UncheckedLongJump::LongJump(ThreadState->RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
// This is purely a debugging aid for developers to see if they are in JIT code space when inspecting raw memory
void EmitDetectionString();
IR::RegisterAllocationPass* RAPass {};
FEXCore::Core::DebugData* DebugData {};
+163 -88
View File
@@ -267,7 +267,7 @@ DEF_OP(LoadContextIndexed) {
ldr(Dst.Q(), TMP1, Op->BaseOffset);
} else {
add(ARMEmitter::Size::i64Bit, TMP1, TMP1, Op->BaseOffset);
ldur(Dst.Q(), TMP1, Op->BaseOffset);
ldur(Dst.Q(), TMP1);
}
break;
case IR::OpSize::i256Bit:
@@ -333,7 +333,7 @@ DEF_OP(StoreContextIndexed) {
str(Value.Q(), TMP1, Op->BaseOffset);
} else {
add(ARMEmitter::Size::i64Bit, TMP1, TMP1, Op->BaseOffset);
stur(Value.Q(), TMP1, Op->BaseOffset);
stur(Value.Q(), TMP1);
}
break;
case IR::OpSize::i256Bit:
@@ -563,12 +563,12 @@ DEF_OP(LoadDF) {
auto Flag = X86State::RFLAG_DF_RAW_LOC;
// DF needs sign extension to turn 0x1/0xFF into 1/-1
ldrsb(Dst.X(), STATE, offsetof(FEXCore::Core::CPUState, flags[Flag]));
ldrsb(Dst.X(), STATE, ARRAY_OFFSETOF(FEXCore::Core::CPUState, flags, Flag));
}
DEF_OP(ContextClear) {
auto Op = IROp->C<IR::IROp_ContextClear>();
if (CTX->HostFeatures.SupportsCLZERO) {
if (CTX->HostFeatures.PreferZVAForVZero) {
// We can use CLZero directly when hardware supports it.
// Provides a fairly generous speed-up on Ampere1A hardware.
// TODO: When FEAT_MOPS hardware ships, test memset using MOPS.
@@ -1849,13 +1849,6 @@ DEF_OP(StoreMemTSO) {
}
DEF_OP(MemSet) {
// TODO: A future looking task would be to support this with ARM's MOPS instructions.
// The 8-bit non-atomic forward path directly matches ARM's SETP/SETM/SETE instruction,
// while the backward version needs some fixup to convert it to a forward direction.
//
// Assuming non-atomicity and non-faulting behaviour, this can accelerate this implementation.
// Additionally: This is commonly used as a memset to zero. If we know up-front with an inline constant
// that the value is zero, we can optimize any operation larger than 8-bit down to 8-bit to use the MOPS implementation.
const auto Op = IROp->C<IR::IROp_MemSet>();
const bool IsAtomic = CTX->IsMemcpyAtomicTSOEnabled();
@@ -1933,8 +1926,30 @@ DEF_OP(MemSet) {
ARMEmitter::SubRegSize::i8Bit;
auto EmitMemset = [&](int32_t Direction) {
const int32_t OpSize = Size;
const int32_t SizeDirection = Size * Direction;
const bool IsBackwards = Direction == -1;
// Sets the result to the final address written depending on
// whether or not the memset is forwards or backwards.
const auto MakeFinalAddress = [&] {
if (IsBackwards) {
switch (Size) {
case 1: sub(Dst.X(), MemReg.X(), Length.X()); break;
case 2:
case 4:
case 8: sub(Dst.X(), MemReg.X(), Length.X(), ARMEmitter::ShiftType::LSL, FEXCore::ilog2(Size)); break;
default: LOGMAN_MSG_A_FMT("Unhandled MemSet size: {}", Size); break;
}
} else {
switch (Size) {
case 1: add(Dst.X(), MemReg.X(), Length.X()); break;
case 2:
case 4:
case 8: add(Dst.X(), MemReg.X(), Length.X(), ARMEmitter::ShiftType::LSL, FEXCore::ilog2(Size)); break;
default: LOGMAN_MSG_A_FMT("Unhandled MemSet size: {}", Size); break;
}
}
};
ARMEmitter::BiDirectionalLabel AgainInternal {};
ARMEmitter::ForwardLabel DoneInternal {};
@@ -1943,12 +1958,56 @@ DEF_OP(MemSet) {
(void)cbz(ARMEmitter::Size::i64Bit, TMP1, &DoneInternal);
if (!IsAtomic) {
if (CTX->HostFeatures.SupportsMOPS) {
const bool Is8Bit = SubRegSize == ARMEmitter::SubRegSize::i8Bit;
// We can handle 8-bit memsets and any other size that happens
// to be using an inlined zero value (resulting in the use of ZR).
//
// NOTE:
// Strictly speaking, this can also be trivially expanded to handle other sizes
// that happen to use any value that could fit inside a byte if the need
// arises. This does increase branching and code generation, however, since
// we'd still need to emit the fallback in the event a value for a larger size
// falls outside the range of a byte instead of only generating the MOPS code.
if (Is8Bit || Value == ARMEmitter::Reg::zr) {
// If we're performing a non-byte-sized zeroing operation then we need to
// scale the counter accordingly. (e.g. a 64-bit memset of size 2 needs to
// be turned into an 8-bit memset of size 16)
if (!Is8Bit) {
lsl(ARMEmitter::Size::i64Bit, TMP1, TMP1, FEXCore::ToUnderlying(SubRegSize));
}
// If backwards, then we need to adjust the starting address because
// set{p, m, e} memset forwards, so we need to slide this bad boy
// back like: (address - count) + 1.
//
// This lets us offset the address such that we can treat a backwards
// memset as if it were a forwards one.
if (IsBackwards) {
sub(TMP2, TMP2, TMP1);
add(ARMEmitter::Size::i64Bit, TMP2, TMP2, 1);
}
// Unfortunately set operations fiddle with NZCV, so we need to preserve it.
mrs(TMP3, ARMEmitter::SystemRegister::NZCV);
setp(TMP2, TMP1, Value.X());
setm(TMP2, TMP1, Value.X());
sete(TMP2, TMP1, Value.X());
msr(ARMEmitter::SystemRegister::NZCV, TMP3);
MakeFinalAddress();
(void)Bind(&DoneInternal);
return;
}
}
ARMEmitter::ForwardLabel AgainInternal256Exit {};
ARMEmitter::BackwardLabel AgainInternal256 {};
ARMEmitter::ForwardLabel AgainInternal128Exit {};
ARMEmitter::BackwardLabel AgainInternal128 {};
if (Direction == -1) {
if (IsBackwards) {
sub(ARMEmitter::Size::i64Bit, TMP2, TMP2, 32 - Size);
}
@@ -1986,39 +2045,23 @@ DEF_OP(MemSet) {
add(ARMEmitter::Size::i64Bit, TMP1, TMP1, 32 / Size);
(void)cbz(ARMEmitter::Size::i64Bit, TMP1, &DoneInternal);
if (Direction == -1) {
if (IsBackwards) {
add(ARMEmitter::Size::i64Bit, TMP2, TMP2, 32 - Size);
}
}
(void)Bind(&AgainInternal);
if (IsAtomic) {
MemStoreTSO(Value, OpSize, SizeDirection);
MemStoreTSO(Value, Size, SizeDirection);
} else {
MemStore(Value, OpSize, SizeDirection);
MemStore(Value, Size, SizeDirection);
}
sub(ARMEmitter::Size::i64Bit, TMP1, TMP1, 1);
(void)cbnz(ARMEmitter::Size::i64Bit, TMP1, &AgainInternal);
(void)Bind(&DoneInternal);
if (SizeDirection >= 0) {
switch (OpSize) {
case 1: add(Dst.X(), MemReg.X(), Length.X()); break;
case 2: add(Dst.X(), MemReg.X(), Length.X(), ARMEmitter::ShiftType::LSL, 1); break;
case 4: add(Dst.X(), MemReg.X(), Length.X(), ARMEmitter::ShiftType::LSL, 2); break;
case 8: add(Dst.X(), MemReg.X(), Length.X(), ARMEmitter::ShiftType::LSL, 3); break;
default: LOGMAN_MSG_A_FMT("Unhandled {} size: {}", __func__, OpSize); break;
}
} else {
switch (OpSize) {
case 1: sub(Dst.X(), MemReg.X(), Length.X()); break;
case 2: sub(Dst.X(), MemReg.X(), Length.X(), ARMEmitter::ShiftType::LSL, 1); break;
case 4: sub(Dst.X(), MemReg.X(), Length.X(), ARMEmitter::ShiftType::LSL, 2); break;
case 8: sub(Dst.X(), MemReg.X(), Length.X(), ARMEmitter::ShiftType::LSL, 3); break;
default: LOGMAN_MSG_A_FMT("Unhandled {} size: {}", __func__, OpSize); break;
}
}
MakeFinalAddress();
};
if (DirectionIsInline) {
@@ -2041,10 +2084,6 @@ DEF_OP(MemSet) {
}
DEF_OP(MemCpy) {
// TODO: A future looking task would be to support this with ARM's MOPS instructions.
// The 8-bit non-atomic path directly matches ARM's CPYP/CPYM/CPYE instruction,
//
// Assuming non-atomicity and non-faulting behaviour, this can accelerate this implementation.
const auto Op = IROp->C<IR::IROp_MemCpy>();
const bool IsAtomic = CTX->IsMemcpyAtomicTSOEnabled();
@@ -2175,8 +2214,40 @@ DEF_OP(MemCpy) {
};
auto EmitMemcpy = [&](int32_t Direction) {
const int32_t OpSize = Size;
const int32_t SizeDirection = Size * Direction;
const bool IsBackwards = Direction == -1;
const auto FinalizeAddresses = [&] {
if (IsBackwards) {
switch (Size) {
case 1:
sub(Dst0.X(), TMP1, TMP3);
sub(Dst1.X(), TMP2, TMP3);
break;
case 2:
case 4:
case 8:
sub(Dst0.X(), TMP1, TMP3, ARMEmitter::ShiftType::LSL, FEXCore::ilog2(Size));
sub(Dst1.X(), TMP2, TMP3, ARMEmitter::ShiftType::LSL, FEXCore::ilog2(Size));
break;
default: LOGMAN_MSG_A_FMT("Unhandled MemCpy size: {}", Size); break;
}
} else {
switch (Size) {
case 1:
add(Dst0.X(), TMP1, TMP3);
add(Dst1.X(), TMP2, TMP3);
break;
case 2:
case 4:
case 8:
add(Dst0.X(), TMP1, TMP3, ARMEmitter::ShiftType::LSL, FEXCore::ilog2(Size));
add(Dst1.X(), TMP2, TMP3, ARMEmitter::ShiftType::LSL, FEXCore::ilog2(Size));
break;
default: LOGMAN_MSG_A_FMT("Unhandled MemCpy size: {}", Size); break;
}
}
};
ARMEmitter::BiDirectionalLabel AgainInternal {};
ARMEmitter::ForwardLabel DoneInternal {};
@@ -2185,6 +2256,48 @@ DEF_OP(MemCpy) {
(void)cbz(ARMEmitter::Size::i64Bit, TMP1, &DoneInternal);
if (!IsAtomic) {
if (CTX->HostFeatures.SupportsMOPS) {
// In the event we have an overlap (gross), we need to fall back
// to the non-mops copy handler. Since the overlap check needs to
// make use of NZCV, we need to save it. This can be avoided with
// ARMv9.6+'s FEAT_CMPBR, but alas, we don't have access to that right now.
//
// NOTE: That we need to temporarily trash TMP1 and restore it after the
// comparison.
ARMEmitter::ForwardLabel OverlapCase;
mrs(TMP4, ARMEmitter::SystemRegister::NZCV);
sub(ARMEmitter::Size::i64Bit, TMP1, TMP2, TMP3);
cmp(ARMEmitter::Size::i64Bit, TMP1, Length.X());
mov(TMP1, Length.X());
(void)bc(ARMEmitter::Condition::CC_LT, &OverlapCase);
// If doing something larger than a byte copy, then we need to scale
// the counter value accordingly to convert it to bytes.
if (Size > 1) {
lsl(ARMEmitter::Size::i64Bit, TMP1, TMP1, FEXCore::ilog2(Size));
}
// Adjust addresses so that we treat the backward copy as a forward copy
if (IsBackwards) {
sub(ARMEmitter::Size::i64Bit, TMP2, TMP2, TMP1);
sub(ARMEmitter::Size::i64Bit, TMP3, TMP3, TMP1);
add(ARMEmitter::Size::i64Bit, TMP2, TMP2, Size);
add(ARMEmitter::Size::i64Bit, TMP3, TMP3, Size);
}
// Unfortunately copy operations fiddle with NZCV, so we need to preserve it.
cpyfp(TMP2, TMP3, TMP1);
cpyfm(TMP2, TMP3, TMP1);
cpyfe(TMP2, TMP3, TMP1);
msr(ARMEmitter::SystemRegister::NZCV, TMP4);
(void)b(&DoneInternal);
// Turns out we overlap and need to fall back. Make sure to restore NZCV.
(void)Bind(&OverlapCase);
msr(ARMEmitter::SystemRegister::NZCV, TMP4);
}
ARMEmitter::ForwardLabel AbsPos {};
ARMEmitter::ForwardLabel AgainInternal256Exit {};
ARMEmitter::ForwardLabel AgainInternal128Exit {};
@@ -2198,7 +2311,7 @@ DEF_OP(MemCpy) {
sub(ARMEmitter::Size::i64Bit, TMP4, TMP4, 32);
(void)tbnz(TMP4, 63, &AgainInternal);
if (Direction == -1) {
if (IsBackwards) {
sub(ARMEmitter::Size::i64Bit, TMP2, TMP2, 32 - Size);
sub(ARMEmitter::Size::i64Bit, TMP3, TMP3, 32 - Size);
}
@@ -2233,7 +2346,7 @@ DEF_OP(MemCpy) {
add(ARMEmitter::Size::i64Bit, TMP1, TMP1, 32 / Size);
(void)cbz(ARMEmitter::Size::i64Bit, TMP1, &DoneInternal);
if (Direction == -1) {
if (IsBackwards) {
add(ARMEmitter::Size::i64Bit, TMP2, TMP2, 32 - Size);
add(ARMEmitter::Size::i64Bit, TMP3, TMP3, 32 - Size);
}
@@ -2241,9 +2354,9 @@ DEF_OP(MemCpy) {
(void)Bind(&AgainInternal);
if (IsAtomic) {
MemCpyTSO(OpSize, SizeDirection);
MemCpyTSO(Size, SizeDirection);
} else {
MemCpy(OpSize, SizeDirection);
MemCpy(Size, SizeDirection);
}
sub(ARMEmitter::Size::i64Bit, TMP1, TMP1, 1);
(void)cbnz(ARMEmitter::Size::i64Bit, TMP1, &AgainInternal);
@@ -2255,54 +2368,14 @@ DEF_OP(MemCpy) {
mov(TMP2, MemRegSrc.X());
mov(TMP3, Length.X());
if (SizeDirection >= 0) {
switch (OpSize) {
case 1:
add(Dst0.X(), TMP1, TMP3);
add(Dst1.X(), TMP2, TMP3);
break;
case 2:
add(Dst0.X(), TMP1, TMP3, ARMEmitter::ShiftType::LSL, 1);
add(Dst1.X(), TMP2, TMP3, ARMEmitter::ShiftType::LSL, 1);
break;
case 4:
add(Dst0.X(), TMP1, TMP3, ARMEmitter::ShiftType::LSL, 2);
add(Dst1.X(), TMP2, TMP3, ARMEmitter::ShiftType::LSL, 2);
break;
case 8:
add(Dst0.X(), TMP1, TMP3, ARMEmitter::ShiftType::LSL, 3);
add(Dst1.X(), TMP2, TMP3, ARMEmitter::ShiftType::LSL, 3);
break;
default: LOGMAN_MSG_A_FMT("Unhandled {} size: {}", __func__, OpSize); break;
}
} else {
switch (OpSize) {
case 1:
sub(Dst0.X(), TMP1, TMP3);
sub(Dst1.X(), TMP2, TMP3);
break;
case 2:
sub(Dst0.X(), TMP1, TMP3, ARMEmitter::ShiftType::LSL, 1);
sub(Dst1.X(), TMP2, TMP3, ARMEmitter::ShiftType::LSL, 1);
break;
case 4:
sub(Dst0.X(), TMP1, TMP3, ARMEmitter::ShiftType::LSL, 2);
sub(Dst1.X(), TMP2, TMP3, ARMEmitter::ShiftType::LSL, 2);
break;
case 8:
sub(Dst0.X(), TMP1, TMP3, ARMEmitter::ShiftType::LSL, 3);
sub(Dst1.X(), TMP2, TMP3, ARMEmitter::ShiftType::LSL, 3);
break;
default: LOGMAN_MSG_A_FMT("Unhandled {} size: {}", __func__, OpSize); break;
}
}
FinalizeAddresses();
};
if (DirectionIsInline) {
LOGMAN_THROW_A_FMT(DirectionConstant == 1 || DirectionConstant == -1, "unexpected direction");
EmitMemcpy(DirectionConstant);
} else {
// Emit forward direction memset then backward direction memset.
// Emit forward direction memcpy then backward direction memcpy.
for (int32_t Direction : {1, -1}) {
EmitMemcpy(Direction);
if (Direction == 1) {
@@ -2327,12 +2400,13 @@ DEF_OP(CacheLineClear) {
// Clear dcache only
// icache doesn't matter here since the guest application shouldn't be calling clflush on JIT code.
// check host cacheline size again x86_64 size to ensure at least 64 bytes are cleaned
if (CTX->HostFeatures.DCacheLineSize >= 64U) {
dc(ARMEmitter::DataCacheOperation::CIVAC, MemReg);
} else {
auto CurrentWorkingReg = MemReg.X();
for (size_t i = 0; i < std::max(1U, CTX->HostFeatures.DCacheLineSize / 64U); ++i) {
dc(ARMEmitter::DataCacheOperation::CIVAC, TMP1);
for (size_t i = 0; i < std::max(1U, 64U / CTX->HostFeatures.DCacheLineSize); ++i) {
dc(ARMEmitter::DataCacheOperation::CIVAC, CurrentWorkingReg);
add(ARMEmitter::Size::i64Bit, TMP1, CurrentWorkingReg, CTX->HostFeatures.DCacheLineSize);
CurrentWorkingReg = TMP1;
}
@@ -2355,12 +2429,13 @@ DEF_OP(CacheLineClean) {
auto MemReg = GetReg(Op->Addr);
// Clean dcache only
// check host cacheline size again x86_64 size to ensure at least 64 bytes are cleaned
if (CTX->HostFeatures.DCacheLineSize >= 64U) {
dc(ARMEmitter::DataCacheOperation::CVAC, MemReg);
} else {
auto CurrentWorkingReg = MemReg.X();
for (size_t i = 0; i < std::max(1U, CTX->HostFeatures.DCacheLineSize / 64U); ++i) {
dc(ARMEmitter::DataCacheOperation::CVAC, TMP1);
for (size_t i = 0; i < std::max(1U, 64U / CTX->HostFeatures.DCacheLineSize); ++i) {
dc(ARMEmitter::DataCacheOperation::CVAC, CurrentWorkingReg);
add(ARMEmitter::Size::i64Bit, TMP1, CurrentWorkingReg, CTX->HostFeatures.DCacheLineSize);
CurrentWorkingReg = TMP1;
}
+34 -6
View File
@@ -73,8 +73,8 @@ DEF_OP(Break) {
uint64_t Constant {};
memcpy(&Constant, &State, sizeof(State));
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r1, Constant);
str(ARMEmitter::XReg::x1, STATE, offsetof(FEXCore::Core::CpuStateFrame, SynchronousFaultData));
LoadConstant(ARMEmitter::Size::i64Bit, TMP1, Constant);
str(TMP1, STATE, offsetof(FEXCore::Core::CpuStateFrame, SynchronousFaultData));
switch (Op->Reason.Signal) {
case Core::FAULT_SIGILL:
@@ -168,7 +168,7 @@ DEF_OP(PushRoundingMode) {
} else {
LOGMAN_THROW_A_FMT(Op->RoundMode == 1 || Op->RoundMode == 2, "expect a valid round mode");
and_(ARMEmitter::Size::i64Bit, TMP1, Dest, ~(Op->RoundMode << 22));
and_(ARMEmitter::Size::i64Bit, TMP1, Dest, ~(3 << 22));
orr(ARMEmitter::Size::i64Bit, TMP1, TMP1, (Op->RoundMode == 2 ? 1 : 2) << 22);
}
@@ -210,6 +210,25 @@ DEF_OP(Print) {
PopDynamicRegs();
}
DEF_OP(PrintMsg) {
auto Op = IROp->C<IR::IROp_PrintMsg>();
PushDynamicRegs(TMP1);
SpillStaticRegs(TMP1);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, reinterpret_cast<uintptr_t>(Op->Value));
ldr(ARMEmitter::XReg::x1, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.PrintMsgValue));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<void, uint64_t>(ARMEmitter::Reg::r1);
} else {
blr(ARMEmitter::Reg::r1);
}
FillStaticRegs();
PopDynamicRegs();
}
DEF_OP(ProcessorID) {
if (CTX->HostFeatures.SupportsCPUIndexInTPIDRRO) {
mrs(GetReg(Node), ARMEmitter::SystemRegister::TPIDRRO_EL0);
@@ -227,7 +246,10 @@ DEF_OP(ProcessorID) {
// Ordering is incredibly important here
// We must spill any overlapping registers first THEN claim we are in a syscall without invalidating state at all
// Only spill the registers that intersect with our usage
SpillStaticRegs(TMP1, false, SpillMask);
SpillStaticRegs(TMP1, {
.GPRSpillMask = SpillMask,
.FPRs = false,
});
// Now that we are spilled, store in the state that we are in a syscall
// Still without overwriting registers that matter
@@ -260,11 +282,17 @@ DEF_OP(ProcessorID) {
// Load the values returned by the kernel
ldp<ARMEmitter::IndexType::OFFSET>(ARMEmitter::WReg::w0, ARMEmitter::WReg::w1, ARMEmitter::Reg::rsp);
// Deallocate stack space
sub(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, ARMEmitter::Reg::rsp, 16);
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, ARMEmitter::Reg::rsp, 16);
// Now that we are done in the syscall we need to carefully peel back the state
// First unspill the registers from before
FillStaticRegs(false, SpillMask, ~0U, ARMEmitter::Reg::r8, ARMEmitter::Reg::r2);
FillStaticRegs({
.OptionalReg = ARMEmitter::Reg::r8,
.OptionalReg2 = ARMEmitter::Reg::r2,
.GPRFillMask = SpillMask,
.FPRs = false,
});
// Now the registers we've spilled are back in their original host registers
// We can safely claim we are no longer in a syscall
+406 -177
View File
@@ -977,7 +977,7 @@ DEF_OP(LoadNamedVectorConstant) {
}
// Load the pointer.
auto GenerateMemOperand = [this](IR::OpSize OpSize, uint32_t NamedConstant, ARMEmitter::Register Base) {
const auto ConstantOffset = offsetof(FEXCore::Core::CpuStateFrame, Pointers.NamedVectorConstants[NamedConstant]);
const auto ConstantOffset = ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, Pointers.NamedVectorConstants, NamedConstant);
if (ConstantOffset <= 255 || // Unscaled 9-bit signed
((ConstantOffset & (IR::OpSizeToSize(OpSize) - 1)) == 0 &&
@@ -985,13 +985,13 @@ DEF_OP(LoadNamedVectorConstant) {
return ARMEmitter::ExtendedMemOperand(Base.X(), ARMEmitter::IndexType::OFFSET, ConstantOffset);
}
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.NamedVectorConstantPointers[NamedConstant]));
ldr(TMP1, STATE_PTR_IDX(CpuStateFrame, Pointers.NamedVectorConstantPointers, NamedConstant));
return ARMEmitter::ExtendedMemOperand(TMP1, ARMEmitter::IndexType::OFFSET, 0);
};
if (OpSize == IR::OpSize::i256Bit) {
// Handle SVE 32-byte variant upfront.
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.NamedVectorConstantPointers[Op->Constant]));
ldr(TMP1, STATE_PTR_IDX(CpuStateFrame, Pointers.NamedVectorConstantPointers, Op->Constant));
ld1b<ARMEmitter::SubRegSize::i8Bit>(Dst.Z(), PRED_TMP_32B.Zeroing(), TMP1, 0);
return;
}
@@ -1013,7 +1013,7 @@ DEF_OP(LoadNamedVectorIndexedConstant) {
const auto Dst = GetVReg(Node);
// Load the pointer.
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.IndexedNamedVectorConstantPointers[Op->Constant]));
ldr(TMP1, STATE_PTR_IDX(CpuStateFrame, Pointers.IndexedNamedVectorConstantPointers, Op->Constant));
switch (OpSize) {
case IR::OpSize::i8Bit: ldrb(Dst, TMP1, Op->Index); break;
@@ -1036,24 +1036,28 @@ DEF_OP(VMov) {
const auto Dst = GetVReg(Node);
const auto Source = GetVReg(Op->Source);
const auto Sub64BitHandler = [&](ARMEmitter::SubRegSize InsertSize) {
if (Dst != Source) {
movi(ARMEmitter::SubRegSize::i64Bit, Dst.Q(), 0);
ins(InsertSize, Dst, 0, Source, 0);
} else {
movi(ARMEmitter::SubRegSize::i64Bit, VTMP1.Q(), 0);
ins(InsertSize, VTMP1, 0, Source, 0);
mov(Dst.Q(), VTMP1.Q());
}
};
switch (OpSize) {
case IR::OpSize::i8Bit: {
movi(ARMEmitter::SubRegSize::i64Bit, VTMP1.Q(), 0);
ins(ARMEmitter::SubRegSize::i8Bit, VTMP1, 0, Source, 0);
mov(Dst.Q(), VTMP1.Q());
Sub64BitHandler(ARMEmitter::SubRegSize::i8Bit);
break;
}
case IR::OpSize::i16Bit: {
movi(ARMEmitter::SubRegSize::i64Bit, VTMP1.Q(), 0);
ins(ARMEmitter::SubRegSize::i16Bit, VTMP1, 0, Source, 0);
mov(Dst.Q(), VTMP1.Q());
Sub64BitHandler(ARMEmitter::SubRegSize::i16Bit);
break;
}
case IR::OpSize::i32Bit: {
movi(ARMEmitter::SubRegSize::i64Bit, VTMP1.Q(), 0);
ins(ARMEmitter::SubRegSize::i32Bit, VTMP1, 0, Source, 0);
mov(Dst.Q(), VTMP1.Q());
Sub64BitHandler(ARMEmitter::SubRegSize::i32Bit);
break;
}
case IR::OpSize::i64Bit: {
@@ -1095,16 +1099,21 @@ DEF_OP(VAddP) {
if (HostSupportsSVE256 && Is256Bit) {
const auto Pred = PRED_TMP_32B.Merging();
// SVE ADDP is a destructive operation, so we need a temporary
movprfx(VTMP1.Z(), VectorLower.Z());
// SVE ADDP is a destructive operation, so we need a temporary if
// the destination and the lower vector don't alias.
auto LHS = Dst;
if (Dst != VectorLower) {
movprfx(VTMP1.Z(), VectorLower.Z());
LHS = VTMP1;
}
// Unlike Adv. SIMD's version of ADDP, which acts like it concats the
// upper vector onto the end of the lower vector and then performs
// pairwise addition, the SVE version actually interleaves the
// results of the pairwise addition (gross!), so we need to undo that.
addp(SubRegSize, VTMP1.Z(), Pred, VTMP1.Z(), VectorUpper.Z());
uzp1(SubRegSize, Dst.Z(), VTMP1.Z(), VTMP1.Z());
uzp2(SubRegSize, VTMP2.Z(), VTMP1.Z(), VTMP1.Z());
addp(SubRegSize, LHS.Z(), Pred, LHS.Z(), VectorUpper.Z());
uzp1(SubRegSize, Dst.Z(), LHS.Z(), LHS.Z());
uzp2(SubRegSize, VTMP2.Z(), LHS.Z(), LHS.Z());
// Merge upper half with lower half.
splice<ARMEmitter::OpType::Destructive>(ARMEmitter::SubRegSize::i64Bit, Dst.Z(), PRED_TMP_16B, Dst.Z(), VTMP2.Z());
@@ -1130,9 +1139,16 @@ DEF_OP(VOrn) {
const auto Vector2 = GetVReg(Op->Vector2);
if (HostSupportsSVE256 && Is256Bit) {
const auto Pred = PRED_TMP_32B.Merging();
not_(ARMEmitter::SubRegSize::i8Bit, VTMP1.Z(), Pred, Vector2.Z());
orr(Dst.Z(), Vector1.Z(), VTMP1.Z());
if (Dst == Vector1) {
bsl2n(Dst.Z(), Dst.Z(), Vector2.Z(), Dst.Z());
} else if (Dst == Vector2) {
const auto Pred = PRED_TMP_32B.Merging();
not_(ARMEmitter::SubRegSize::i8Bit, Dst.Z(), Pred, Dst.Z());
orr(Dst.Z(), Vector1.Z(), Dst.Z());
} else {
movprfx(Dst.Z(), Vector1.Z());
bsl2n(Dst.Z(), Dst.Z(), Vector2.Z(), Vector1.Z());
}
} else if (Is128Bit) {
orn(Dst.Q(), Vector1.Q(), Vector2.Q());
} else {
@@ -1156,8 +1172,7 @@ DEF_OP(VFAddV) {
if (HostSupportsSVE256 && Is256Bit) {
const auto Pred = PRED_TMP_32B.Merging();
faddv(SubRegSize.Vector, Dst, Pred, Vector.Z());
}
if (HostSupportsSVE128) {
} else if (HostSupportsSVE128) {
const auto Pred = PRED_TMP_16B.Merging();
faddv(SubRegSize.Vector, Dst, Pred, Vector.Z());
} else {
@@ -1184,20 +1199,16 @@ DEF_OP(VAddV) {
const auto Vector = GetVReg(Op->Vector);
if (HostSupportsSVE256 && Is256Bit) {
// SVE doesn't have an equivalent ADDV instruction, so we make do
// by performing two Adv. SIMD ADDV operations on the high and low
// 128-bit lanes and then sum them up.
const auto Mask = PRED_TMP_32B.Zeroing();
const auto CompactPred = ARMEmitter::PReg::p0;
// Select all our upper elements to run ADDV over them.
not_(CompactPred, Mask, PRED_TMP_16B);
compact(ARMEmitter::SubRegSize::i64Bit, VTMP1.Z(), CompactPred, Vector.Z());
addv(SubRegSize.Vector, VTMP2.Q(), Vector.Q());
addv(SubRegSize.Vector, VTMP1.Q(), VTMP1.Q());
add(SubRegSize.Vector, Dst.Q(), VTMP1.Q(), VTMP2.Q());
if (ElementSize == IR::OpSize::i64Bit) {
const auto Mask = PRED_TMP_32B.Zeroing();
uaddv(SubRegSize.Vector, Dst.D(), Mask, Vector.Z());
} else {
const auto Mask = ARMEmitter::PReg::p0;
uaddv(SubRegSize.Vector, VTMP1.D(), Mask, Vector.Z());
mov_imm(ARMEmitter::SubRegSize::i64Bit, Dst.Z(), 0);
ptrue(SubRegSize.Vector, Mask, ARMEmitter::PredicatePattern::SVE_VL1);
mov(SubRegSize.Vector, Dst.Z(), Mask.Merging(), VTMP1.Z());
}
} else {
if (ElementSize == IR::OpSize::i64Bit) {
addp(SubRegSize.Scalar, Dst, Vector);
@@ -1298,16 +1309,21 @@ DEF_OP(VFAddP) {
if (HostSupportsSVE256 && Is256Bit) {
const auto Pred = PRED_TMP_32B.Merging();
// SVE FADDP is a destructive operation, so we need a temporary
movprfx(VTMP1.Z(), VectorLower.Z());
// SVE FADDP is a destructive operation, so we need a temporary if
// the destination and the lower vector don't alias.
auto LHS = Dst;
if (Dst != VectorLower) {
movprfx(VTMP1.Z(), VectorLower.Z());
LHS = VTMP1;
}
// Unlike Adv. SIMD's version of FADDP, which acts like it concats the
// upper vector onto the end of the lower vector and then performs
// pairwise addition, the SVE version actually interleaves the
// results of the pairwise addition (gross!), so we need to undo that.
faddp(SubRegSize, VTMP1.Z(), Pred, VTMP1.Z(), VectorUpper.Z());
uzp1(SubRegSize, Dst.Z(), VTMP1.Z(), VTMP1.Z());
uzp2(SubRegSize, VTMP2.Z(), VTMP1.Z(), VTMP1.Z());
faddp(SubRegSize, LHS.Z(), Pred, LHS.Z(), VectorUpper.Z());
uzp1(SubRegSize, Dst.Z(), LHS.Z(), LHS.Z());
uzp2(SubRegSize, VTMP2.Z(), LHS.Z(), LHS.Z());
// Merge upper half with lower half.
splice<ARMEmitter::OpType::Destructive>(ARMEmitter::SubRegSize::i64Bit, Dst.Z(), PRED_TMP_16B, Dst.Z(), VTMP2.Z());
@@ -1434,8 +1450,8 @@ DEF_OP(VFMin) {
bif(Dst.Q(), Vector2.Q(), VTMP1.Q());
} else if (Dst == Vector2) {
// Destination is already Vector2, Invert arguments and insert Vector1 on false.
fcmgt(SubRegSize, VTMP1.Q(), Vector1.Q(), Vector2.Q());
bif(Dst.Q(), Vector1.Q(), VTMP1.Q());
fcmgt(SubRegSize, VTMP1.Q(), Vector2.Q(), Vector1.Q());
bit(Dst.Q(), Vector1.Q(), VTMP1.Q());
} else {
// Dst is not either source, need a move.
fcmgt(SubRegSize, VTMP1.Q(), Vector2.Q(), Vector1.Q());
@@ -1466,7 +1482,8 @@ DEF_OP(VFMax) {
const auto Mask = PRED_TMP_32B;
const auto ComparePred = ARMEmitter::PReg::p0;
fcmgt(SubRegSize, ComparePred, Mask.Zeroing(), Vector2.Z(), Vector1.Z());
fcmgt(SubRegSize, ComparePred, Mask.Zeroing(), Vector1.Z(), Vector2.Z());
not_(ComparePred, Mask.Zeroing(), ComparePred);
if (Dst == Vector1) {
// Trivial case where Vector1 is also the destination.
@@ -1488,17 +1505,17 @@ DEF_OP(VFMax) {
if (Dst == Vector1) {
// Destination is already Vector1, need to insert Vector2 on true.
fcmgt(SubRegSize, VTMP1.Q(), Vector2.Q(), Vector1.Q());
bit(Dst.Q(), Vector2.Q(), VTMP1.Q());
fcmgt(SubRegSize, VTMP1.Q(), Vector1.Q(), Vector2.Q());
bif(Dst.Q(), Vector2.Q(), VTMP1.Q());
} else if (Dst == Vector2) {
// Destination is already Vector2, Invert arguments and insert Vector1 on true.
fcmgt(SubRegSize, VTMP1.Q(), Vector1.Q(), Vector2.Q());
bit(Dst.Q(), Vector1.Q(), VTMP1.Q());
} else {
// Dst is not either source, need a move.
fcmgt(SubRegSize, VTMP1.Q(), Vector2.Q(), Vector1.Q());
fcmgt(SubRegSize, VTMP1.Q(), Vector1.Q(), Vector2.Q());
mov(Dst.Q(), Vector1.Q());
bit(Dst.Q(), Vector2.Q(), VTMP1.Q());
bif(Dst.Q(), Vector2.Q(), VTMP1.Q());
}
}
}
@@ -1525,9 +1542,14 @@ DEF_OP(VFRecp) {
return;
}
fmov(SubRegSize.Vector, VTMP1.Z(), 1.0);
fdiv(SubRegSize.Vector, VTMP1.Z(), Pred, VTMP1.Z(), Vector.Z());
mov(Dst.Z(), VTMP1.Z());
if (Dst != Vector) {
fmov(SubRegSize.Vector, Dst.Z(), 1.0);
fdiv(SubRegSize.Vector, Dst.Z(), Pred, Dst.Z(), Vector.Z());
} else {
fmov(SubRegSize.Vector, VTMP1.Z(), 1.0);
fdiv(SubRegSize.Vector, VTMP1.Z(), Pred, VTMP1.Z(), Vector.Z());
mov(Dst.Z(), VTMP1.Z());
}
} else {
if (IsScalar) {
if (ElementSize == IR::OpSize::i32Bit && HostSupportsRPRES) {
@@ -1779,10 +1801,14 @@ DEF_OP(VUMin) {
break;
}
case IR::OpSize::i64Bit: {
cmhi(SubRegSize, VTMP1.Q(), Vector2.Q(), Vector1.Q());
mov(VTMP2.Q(), Vector1.Q());
bif(VTMP2.Q(), Vector2.Q(), VTMP1.Q());
mov(Dst.Q(), VTMP2.Q());
if (Dst != Vector1 && Dst != Vector2) {
cmhi(SubRegSize, Dst.Q(), Vector1.Q(), Vector2.Q());
bsl(Dst.Q(), Vector2.Q(), Vector1.Q());
} else {
cmhi(SubRegSize, VTMP1.Q(), Vector1.Q(), Vector2.Q());
bsl(VTMP1.Q(), Vector2.Q(), Vector1.Q());
mov(Dst.Q(), VTMP1.Q());
}
break;
}
default: break;
@@ -1828,10 +1854,14 @@ DEF_OP(VSMin) {
break;
}
case IR::OpSize::i64Bit: {
cmgt(SubRegSize, VTMP1.Q(), Vector1.Q(), Vector2.Q());
mov(VTMP2.Q(), Vector1.Q());
bif(VTMP2.Q(), Vector2.Q(), VTMP1.Q());
mov(Dst.Q(), VTMP2.Q());
if (Dst != Vector1 && Dst != Vector2) {
cmgt(SubRegSize, Dst.Q(), Vector1.Q(), Vector2.Q());
bsl(Dst.Q(), Vector2.Q(), Vector1.Q());
} else {
cmgt(SubRegSize, VTMP1.Q(), Vector1.Q(), Vector2.Q());
bsl(VTMP1.Q(), Vector2.Q(), Vector1.Q());
mov(Dst.Q(), VTMP1.Q());
}
break;
}
default: break;
@@ -1877,10 +1907,14 @@ DEF_OP(VUMax) {
break;
}
case IR::OpSize::i64Bit: {
cmhi(SubRegSize, VTMP1.Q(), Vector2.Q(), Vector1.Q());
mov(VTMP2.Q(), Vector1.Q());
bif(VTMP2.Q(), Vector2.Q(), VTMP1.Q());
mov(Dst.Q(), VTMP2.Q());
if (Dst != Vector1 && Dst != Vector2) {
cmhi(SubRegSize, Dst.Q(), Vector1.Q(), Vector2.Q());
bsl(Dst.Q(), Vector1.Q(), Vector2.Q());
} else {
cmhi(SubRegSize, VTMP1.Q(), Vector1.Q(), Vector2.Q());
bsl(VTMP1.Q(), Vector1.Q(), Vector2.Q());
mov(Dst.Q(), VTMP1.Q());
}
break;
}
default: break;
@@ -1926,10 +1960,14 @@ DEF_OP(VSMax) {
break;
}
case IR::OpSize::i64Bit: {
cmgt(SubRegSize, VTMP1.Q(), Vector2.Q(), Vector1.Q());
mov(VTMP2.Q(), Vector1.Q());
bif(VTMP2.Q(), Vector2.Q(), VTMP1.Q());
mov(Dst.Q(), VTMP2.Q());
if (Dst != Vector1 && Dst != Vector2) {
cmgt(SubRegSize, Dst.Q(), Vector1.Q(), Vector2.Q());
bsl(Dst.Q(), Vector1.Q(), Vector2.Q());
} else {
cmgt(SubRegSize, VTMP1.Q(), Vector1.Q(), Vector2.Q());
bsl(VTMP1.Q(), Vector1.Q(), Vector2.Q());
mov(Dst.Q(), VTMP1.Q());
}
break;
}
default: break;
@@ -2749,17 +2787,17 @@ DEF_OP(VUShrSWide) {
const auto Vector = GetVReg(Op->Vector);
if (HostSupportsSVE256 && Is256Bit) {
const auto Mask = PRED_TMP_32B.Merging();
dup(ARMEmitter::SubRegSize::i64Bit, VTMP1.Z(), ShiftScalar.Z(), 0);
if (Dst != Vector) {
// NOTE: SVE LSR is a destructive operation.
movprfx(Dst.Z(), Vector.Z());
}
if (ElementSize == IR::OpSize::i64Bit) {
const auto Mask = PRED_TMP_32B.Merging();
if (Dst != Vector) {
// NOTE: SVE LSR is a destructive operation.
movprfx(Dst.Z(), Vector.Z());
}
lsr(SubRegSize, Dst.Z(), Mask, Dst.Z(), VTMP1.Z());
} else {
lsr_wide(SubRegSize, Dst.Z(), Mask, Dst.Z(), VTMP1.Z());
lsr_wide(SubRegSize, Dst.Z(), Vector.Z(), VTMP1.Z());
}
} else if (HostSupportsSVE128) {
const auto Mask = PRED_TMP_16B.Merging();
@@ -2815,17 +2853,17 @@ DEF_OP(VSShrSWide) {
const auto Vector = GetVReg(Op->Vector);
if (HostSupportsSVE256 && Is256Bit) {
const auto Mask = PRED_TMP_32B.Merging();
dup(ARMEmitter::SubRegSize::i64Bit, VTMP1.Z(), ShiftScalar.Z(), 0);
if (Dst != Vector) {
// NOTE: SVE LSR is a destructive operation.
movprfx(Dst.Z(), Vector.Z());
}
if (ElementSize == IR::OpSize::i64Bit) {
const auto Mask = PRED_TMP_32B.Merging();
if (Dst != Vector) {
// NOTE: SVE LSR is a destructive operation.
movprfx(Dst.Z(), Vector.Z());
}
asr(SubRegSize, Dst.Z(), Mask, Dst.Z(), VTMP1.Z());
} else {
asr_wide(SubRegSize, Dst.Z(), Mask, Dst.Z(), VTMP1.Z());
asr_wide(SubRegSize, Dst.Z(), Vector.Z(), VTMP1.Z());
}
} else if (HostSupportsSVE128) {
const auto Mask = PRED_TMP_16B.Merging();
@@ -2881,17 +2919,17 @@ DEF_OP(VUShlSWide) {
const auto Vector = GetVReg(Op->Vector);
if (HostSupportsSVE256 && Is256Bit) {
const auto Mask = PRED_TMP_32B.Merging();
dup(ARMEmitter::SubRegSize::i64Bit, VTMP1.Z(), ShiftScalar.Z(), 0);
if (Dst != Vector) {
// NOTE: SVE LSR is a destructive operation.
movprfx(Dst.Z(), Vector.Z());
}
if (ElementSize == IR::OpSize::i64Bit) {
const auto Mask = PRED_TMP_32B.Merging();
if (Dst != Vector) {
// NOTE: SVE LSR is a destructive operation.
movprfx(Dst.Z(), Vector.Z());
}
lsl(SubRegSize, Dst.Z(), Mask, Dst.Z(), VTMP1.Z());
} else {
lsl_wide(SubRegSize, Dst.Z(), Mask, Dst.Z(), VTMP1.Z());
lsl_wide(SubRegSize, Dst.Z(), Vector.Z(), VTMP1.Z());
}
} else if (HostSupportsSVE128) {
const auto Mask = PRED_TMP_16B.Merging();
@@ -2979,9 +3017,14 @@ DEF_OP(VInsElement) {
auto Reg = GetVReg(Op->DestVector);
if (HostSupportsSVE256 && Is256Bit) {
// Broadcast our source value across a temporary,
// then combine with the destination.
dup(SubRegSize, VTMP2.Z(), SrcVector.Z(), SrcIdx);
// Broadcast our source value across a temporary, then combine
// with the destination.
//
// We don't need to perform the dup if we're just merging a 128-bit vector into
// into an equivalent position since we have a predicate set up already.
if (!(ElementSize == IR::OpSize::i128Bit && SrcIdx == DestIdx)) {
dup(SubRegSize, VTMP2.Z(), SrcVector.Z(), SrcIdx);
}
// We don't need to move the data unnecessarily if
// DestVector just so happens to also be the IR op
@@ -2994,10 +3037,12 @@ DEF_OP(VInsElement) {
if (ElementSize == IR::OpSize::i128Bit) {
if (DestIdx == 0) {
mov(ARMEmitter::SubRegSize::i8Bit, Dst.Z(), PRED_TMP_16B.Merging(), VTMP2.Z());
const auto Source = SrcIdx == 0 ? SrcVector : VTMP2;
mov(ARMEmitter::SubRegSize::i8Bit, Dst.Z(), PRED_TMP_16B.Merging(), Source.Z());
} else {
const auto Source = SrcIdx == 1 ? SrcVector : VTMP2;
not_(Predicate, PRED_TMP_32B.Zeroing(), PRED_TMP_16B);
mov(ARMEmitter::SubRegSize::i8Bit, Dst.Z(), Predicate.Merging(), VTMP2.Z());
mov(ARMEmitter::SubRegSize::i8Bit, Dst.Z(), Predicate.Merging(), Source.Z());
}
} else {
const auto UpperBound = 16 >> FEXCore::ilog2(IR::OpSizeToSize(ElementSize));
@@ -3132,19 +3177,12 @@ DEF_OP(VUShrI) {
movi(ARMEmitter::SubRegSize::i64Bit, Dst.Q(), 0);
} else {
if (HostSupportsSVE256 && Is256Bit) {
const auto Mask = PRED_TMP_32B.Merging();
if (BitShift == 0) {
if (Dst != Vector) {
mov(Dst.Z(), Vector.Z());
}
} else {
// SVE LSR is destructive, so lets set up the destination if
// Vector doesn't already alias it.
if (Dst != Vector) {
movprfx(Dst.Z(), Vector.Z());
}
lsr(SubRegSize, Dst.Z(), Mask, Dst.Z(), BitShift);
lsr(SubRegSize, Dst.Z(), Vector.Z(), BitShift);
}
} else {
if (BitShift == 0) {
@@ -3158,48 +3196,6 @@ DEF_OP(VUShrI) {
}
}
DEF_OP(VUShraI) {
const auto Op = IROp->C<IR::IROp_VUShraI>();
const auto OpSize = IROp->Size;
const auto BitShift = Op->BitShift;
const auto SubRegSize = ConvertSubRegSize8(IROp);
const auto Is256Bit = OpSize == IR::OpSize::i256Bit;
LOGMAN_THROW_A_FMT(!Is256Bit || HostSupportsSVE256, "Need SVE256 support in order to use {} with 256-bit operation", __func__);
const auto Dst = GetVReg(Node);
const auto DestVector = GetVReg(Op->DestVector);
const auto Vector = GetVReg(Op->Vector);
if (HostSupportsSVE256 && Is256Bit) {
if (Dst == DestVector) {
usra(SubRegSize, Dst.Z(), Vector.Z(), BitShift);
} else {
if (Dst != Vector) {
mov(Dst.Z(), DestVector.Z());
usra(SubRegSize, Dst.Z(), Vector.Z(), BitShift);
} else {
mov(VTMP1.Z(), DestVector.Z());
usra(SubRegSize, Dst.Z(), Vector.Z(), BitShift);
mov(Dst.Z(), VTMP1.Z());
}
}
} else {
if (Dst == DestVector) {
usra(SubRegSize, Dst.Q(), Vector.Q(), BitShift);
} else {
if (Dst != Vector) {
mov(Dst.Q(), DestVector.Q());
usra(SubRegSize, Dst.Q(), Vector.Q(), BitShift);
} else {
mov(VTMP1.Q(), DestVector.Q());
usra(SubRegSize, VTMP1.Q(), Vector.Q(), BitShift);
mov(Dst.Q(), VTMP1.Q());
}
}
}
}
DEF_OP(VSShrI) {
const auto Op = IROp->C<IR::IROp_VSShrI>();
const auto OpSize = IROp->Size;
@@ -3215,19 +3211,12 @@ DEF_OP(VSShrI) {
const auto Vector = GetVReg(Op->Vector);
if (HostSupportsSVE256 && Is256Bit) {
const auto Mask = PRED_TMP_32B.Merging();
if (Shift == 0) {
if (Dst != Vector) {
mov(Dst.Z(), Vector.Z());
}
} else {
// SVE ASR is destructive, so lets set up the destination if
// Vector doesn't already alias it.
if (Dst != Vector) {
movprfx(Dst.Z(), Vector.Z());
}
asr(SubRegSize, Dst.Z(), Mask, Dst.Z(), Shift);
asr(SubRegSize, Dst.Z(), Vector.Z(), Shift);
}
} else {
if (Shift == 0) {
@@ -3257,19 +3246,12 @@ DEF_OP(VShlI) {
movi(ARMEmitter::SubRegSize::i64Bit, Dst.Q(), 0);
} else {
if (HostSupportsSVE256 && Is256Bit) {
const auto Mask = PRED_TMP_32B.Merging();
if (BitShift == 0) {
if (Dst != Vector) {
mov(Dst.Z(), Vector.Z());
}
} else {
// SVE LSL is destructive, so lets set up the destination if
// Vector doesn't already alias it.
if (Dst != Vector) {
movprfx(Dst.Z(), Vector.Z());
}
lsl(SubRegSize, Dst.Z(), Mask, Dst.Z(), BitShift);
lsl(SubRegSize, Dst.Z(), Vector.Z(), BitShift);
}
} else {
if (BitShift == 0) {
@@ -3296,8 +3278,13 @@ DEF_OP(VUShrNI) {
const auto Vector = GetVReg(Op->Vector);
if (HostSupportsSVE256 && Is256Bit) {
shrnb(SubRegSize, Dst.Z(), Vector.Z(), BitShift);
uzp1(SubRegSize, Dst.Z(), Dst.Z(), Dst.Z());
if (BitShift == 0) {
mov_imm(ARMEmitter::SubRegSize::i64Bit, VTMP1.Z(), 0);
uzp1(SubRegSize, Dst.Z(), Dst.Z(), VTMP1.Z());
} else {
shrnb(SubRegSize, Dst.Z(), Vector.Z(), BitShift);
uzp1(SubRegSize, Dst.Z(), Dst.Z(), Dst.Z());
}
} else {
if (BitShift == 0) {
xtn(SubRegSize, Dst.D(), Vector.D());
@@ -3550,9 +3537,13 @@ DEF_OP(VSQXTN2) {
mov(Dst.Q(), VectorLower.Q());
ins(ARMEmitter::SubRegSize::i32Bit, Dst, 1, VTMP2, 0);
} else {
mov(VTMP1.Q(), VectorLower.Q());
sqxtn2(SubRegSize, VTMP1, VectorUpper);
mov(Dst.Q(), VTMP1.Q());
if (Dst == VectorLower) {
sqxtn2(SubRegSize, VectorLower, VectorUpper);
} else {
mov(VTMP1.Q(), VectorLower.Q());
sqxtn2(SubRegSize, VTMP1, VectorUpper);
mov(Dst.Q(), VTMP1.Q());
}
}
}
}
@@ -4408,13 +4399,9 @@ DEF_OP(VFMLS) {
if (Is128Bit) {
fneg(SubRegSize, DestTmp.Q(), VectorAddend.Q());
} else {
fneg(SubRegSize, DestTmp.D(), VectorAddend.D());
}
if (Is128Bit) {
fmla(SubRegSize, DestTmp.Q(), Vector1.Q(), Vector2.Q());
} else {
fneg(SubRegSize, DestTmp.D(), VectorAddend.D());
fmla(SubRegSize, DestTmp.D(), Vector1.D(), Vector2.D());
}
@@ -4434,7 +4421,7 @@ DEF_OP(VFNMLA) {
// - SVE - FMLS
// - ASIMD - FMLS
// - Scalar - FMSUB
const auto Op = IROp->C<IR::IROp_VFMLA>();
const auto Op = IROp->C<IR::IROp_VFNMLA>();
const auto OpSize = IROp->Size;
const auto SubRegSize = ConvertSubRegSize248(IROp);
@@ -4502,7 +4489,7 @@ DEF_OP(VFNMLS) {
// - ASIMD - FMLS (With Negated addend)
// - Scalar - FNMADD
const auto Op = IROp->C<IR::IROp_VFMLS>();
const auto Op = IROp->C<IR::IROp_VFNMLS>();
const auto OpSize = IROp->Size;
const auto SubRegSize = ConvertSubRegSize248(IROp);
@@ -4568,13 +4555,9 @@ DEF_OP(VFNMLS) {
if (Is128Bit) {
fneg(SubRegSize, DestTmp.Q(), VectorAddend.Q());
} else {
fneg(SubRegSize, DestTmp.D(), VectorAddend.D());
}
if (Is128Bit) {
fmls(SubRegSize, DestTmp.Q(), Vector1.Q(), Vector2.Q());
} else {
fneg(SubRegSize, DestTmp.D(), VectorAddend.D());
fmls(SubRegSize, DestTmp.D(), Vector1.D(), Vector2.D());
}
@@ -4588,6 +4571,106 @@ DEF_OP(VFNMLS) {
}
}
DEF_OP(VBlendImm) {
LOGMAN_THROW_A_FMT(HostSupportsSVE128 || HostSupportsSVE256, "Host must support SVE to use {}", __func__);
auto Op = IROp->C<IR::IROp_VBlendImm>();
const auto OpSize = IROp->Size;
const auto Is256Bit = OpSize == IR::OpSize::i256Bit;
const auto SubRegSize = ConvertSubRegSize8(IROp);
const auto ElementSize = IROp->ElementSize;
const auto Selector = Op->Selector;
const auto GoverningPredicate = Is256Bit ? PRED_TMP_32B : PRED_TMP_16B;
const auto Dst = GetVReg(Node);
const auto LHS = GetVReg(Op->LHS);
const auto RHS = GetVReg(Op->RHS);
const auto DstIsNonAliasing = Dst != LHS && Dst != RHS;
// Silly case where two blending sources are the same.
if (LHS == RHS) {
if (DstIsNonAliasing) {
mov(SubRegSize, Dst.Z(), GoverningPredicate.Merging(), LHS.Z());
}
return;
}
// We'll need to expand our selector to match its predicate equivalent.
// The lowest bit of each predicate element being set to 1 signifies
// that it's enabled.
const auto MakePredicateMask = [ElementSize, Is256Bit, OpSize](uint16_t Imm) {
if (ElementSize == IR::OpSize::i8Bit) {
// Since we use a u16 selector, we have enough bits for every byte in a
// 128-bit lane, so we don't need to do anything here except replicate the
// bits in the event of 256-bit.
return Is256Bit ? uint32_t(Imm) << 16 | Imm : Imm;
}
uint32_t Mask = 0;
const auto DataSize = IR::OpSizeToSize(ElementSize);
const auto NumElements = IR::NumElements(OpSize, ElementSize);
for (uint32_t i = 0; i < NumElements; i++) {
if (((Imm >> i) & 1) != 0) {
Mask |= 1U << (DataSize * i);
}
}
return Mask;
};
// Our predicate that we'll be firing our constructed bitmask into.
constexpr auto Predicate = ARMEmitter::PReg::p0.Merging();
// TODO: We can completely eliminate this via PMOV in SVE2.1
ARMEmitter::ForwardLabel AfterLabel;
ARMEmitter::BackwardLabel ConstantLabel;
(void)b(&AfterLabel);
(void)Bind(&ConstantLabel);
const auto PredicateMask = MakePredicateMask(Selector);
if (Dst == RHS) {
dc32(~PredicateMask);
} else {
dc32(PredicateMask);
}
(void)Bind(&AfterLabel);
(void)adr(TMP1, &ConstantLabel);
ldr(Predicate, TMP1);
if (Dst == LHS) {
mov(SubRegSize, LHS.Z(), Predicate, RHS.Z());
} else if (Dst == RHS) {
mov(SubRegSize, RHS.Z(), Predicate, LHS.Z());
} else {
mov(SubRegSize, Dst.Z(), GoverningPredicate.Merging(), LHS.Z());
mov(SubRegSize, Dst.Z(), Predicate, RHS.Z());
}
}
DEF_OP(VXar) {
LOGMAN_THROW_A_FMT(HostSupportsSVE128 || HostSupportsSVE256, "Host must support SVE to use {}", __func__);
auto Op = IROp->C<IR::IROp_VXar>();
const auto SubRegSize = ConvertSubRegSize8(IROp);
const auto ElementSizeBits = IR::OpSizeAsBits(IROp->ElementSize);
const auto Dst = GetVReg(Node);
const auto LHS = GetVReg(Op->LHS);
const auto RHS = GetVReg(Op->RHS);
const auto Rotate = Op->Rotate;
LOGMAN_THROW_A_FMT(Rotate >= 1 && Rotate <= ElementSizeBits, "Rotate immediate must be within [1, {}]", ElementSizeBits);
if (Dst == LHS) {
xar(SubRegSize, Dst.Z(), RHS.Z(), Rotate);
} else if (Dst == RHS) {
movprfx(VTMP1.Z(), LHS.Z());
xar(SubRegSize, VTMP1.Z(), RHS.Z(), Rotate);
mov(Dst.Z(), VTMP1.Z());
} else {
movprfx(Dst.Z(), LHS.Z());
xar(SubRegSize, Dst.Z(), RHS.Z(), Rotate);
}
}
DEF_OP(VFCopySign) {
auto Op = IROp->C<IR::IROp_VFCopySign>();
const auto OpSize = IROp->Size;
@@ -4611,4 +4694,150 @@ DEF_OP(VFCopySign) {
}
}
DEF_OP(F64FPREM) {
const auto Op = IROp->C<IR::IROp_F64FPREM>();
const auto Dst = GetVReg(Node);
const auto Src1 = GetVReg(Op->Src1);
const auto Src2 = GetVReg(Op->Src2);
fmov(VTMP1.D(), Src1.D());
fmov(VTMP2.D(), Src2.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.F64FPREMHandler));
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
fmov(Dst.D(), VTMP1.D());
}
DEF_OP(F64FPREM1) {
const auto Op = IROp->C<IR::IROp_F64FPREM1>();
const auto Dst = GetVReg(Node);
const auto Src1 = GetVReg(Op->Src1);
const auto Src2 = GetVReg(Op->Src2);
fmov(VTMP1.D(), Src1.D());
fmov(VTMP2.D(), Src2.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.F64FPREM1Handler));
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
fmov(Dst.D(), VTMP1.D());
}
DEF_OP(F64SIN) {
const auto Op = IROp->C<IR::IROp_F64SIN>();
const auto Src = GetVReg(Op->Src);
const auto Dst = GetVReg(Node);
fmov(VTMP1.D(), Src.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.F64SinHandler));
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
fmov(Dst.D(), VTMP1.D());
}
DEF_OP(F64COS) {
const auto Op = IROp->C<IR::IROp_F64COS>();
const auto Src = GetVReg(Op->Src);
const auto Dst = GetVReg(Node);
fmov(VTMP1.D(), Src.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.F64CosHandler));
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
fmov(Dst.D(), VTMP1.D());
}
DEF_OP(F64TAN) {
const auto Op = IROp->C<IR::IROp_F64TAN>();
const auto Src = GetVReg(Op->Src);
const auto Dst = GetVReg(Node);
fmov(VTMP1.D(), Src.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.F64TanHandler));
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
fmov(Dst.D(), VTMP1.D());
}
// Src1=y(ST1), Src2=x(ST0). Marshal into VTMP1/VTMP2 and dispatch the shared handler.
DEF_OP(F64ATAN) {
const auto Op = IROp->C<IR::IROp_F64ATAN>();
const auto Src1 = GetVReg(Op->Src1);
const auto Src2 = GetVReg(Op->Src2);
const auto Dst = GetVReg(Node);
fmov(VTMP1.D(), Src1.D());
fmov(VTMP2.D(), Src2.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.F64AtanHandler));
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
fmov(Dst.D(), VTMP1.D());
}
// Src=x(ST0), Src2=y(ST1). Marshal into VTMP1/VTMP2 and dispatch the shared handler.
DEF_OP(F64FYL2X) {
const auto Op = IROp->C<IR::IROp_F64FYL2X>();
const auto Src = GetVReg(Op->Src);
const auto Src2 = GetVReg(Op->Src2);
const auto Dst = GetVReg(Node);
fmov(VTMP1.D(), Src.D());
fmov(VTMP2.D(), Src2.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.F64FYL2XHandler));
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
fmov(Dst.D(), VTMP1.D());
}
// Src=x(ST0), Src2=y(ST1). Marshal into VTMP1/VTMP2 and dispatch the shared handler.
DEF_OP(F64FYL2XP1) {
const auto Op = IROp->C<IR::IROp_F64FYL2XP1>();
const auto Src = GetVReg(Op->Src);
const auto Src2 = GetVReg(Op->Src2);
const auto Dst = GetVReg(Node);
fmov(VTMP1.D(), Src.D());
fmov(VTMP2.D(), Src2.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.F64FYL2XP1Handler));
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
fmov(Dst.D(), VTMP1.D());
}
DEF_OP(F64SCALE) {
const auto Op = IROp->C<IR::IROp_F64SCALE>();
const auto Src1 = GetVReg(Op->Src1);
const auto Src2 = GetVReg(Op->Src2);
const auto Dst = GetVReg(Node);
fmov(VTMP1.D(), Src1.D());
fmov(VTMP2.D(), Src2.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.F64ScaleHandler));
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
fmov(Dst.D(), VTMP1.D());
}
DEF_OP(F64F2XM1) {
const auto Op = IROp->C<IR::IROp_F64F2XM1>();
const auto Src = GetVReg(Op->Src);
const auto Dst = GetVReg(Node);
fmov(VTMP1.D(), Src.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.F64F2XM1Handler));
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
fmov(Dst.D(), VTMP1.D());
}
} // namespace FEXCore::CPU
@@ -41,6 +41,9 @@ LookupCache::LookupCache(FEXCore::Context::ContextImpl* CTX)
PagePointer = reinterpret_cast<uintptr_t>(FEXCore::Allocator::VirtualAlloc(TotalCacheSize, false, false));
LOGMAN_THROW_A_FMT(PagePointer != -1ULL, "Failed to allocate PagePointer");
// Disable THP on the Lookup cache.
FEXCore::Allocator::VirtualTHPControl(reinterpret_cast<const void*>(PagePointer), TotalCacheSize, FEXCore::Allocator::THPControl::Disable);
FEXCore::Allocator::VirtualName("FEXMem_Lookup", reinterpret_cast<void*>(PagePointer),
ctx->Config.VirtualMemSize / FEXCore::Utils::FEX_PAGE_SIZE * 8 + CODE_SIZE);
CTX->SyscallHandler->MarkOvercommitRange(PagePointer, TotalCacheSize);
@@ -84,8 +87,11 @@ void LookupCache::ClearL2Cache(const FEXCore::LookupCacheBaseLockToken& lk) {
}
void LookupCache::ClearThreadLocalCaches(const LookupCacheWriteLockToken&) {
// TODO: Preserve code cache entries?
// Clear L1 and L2 by clearing the full cache.
FEXCore::Allocator::VirtualDontNeed(reinterpret_cast<void*>(PagePointer), TotalCacheSize, false);
// TODO: Rename this member to avoid confusion with code caching
CachedCodePages.clear();
}
+1 -1
View File
@@ -3,7 +3,7 @@
#include "Interface/Context/Context.h"
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/SHMStats.h>
#include "Utils/WritePriorityMutex.h"
#include <FEXCore/Utils/WritePriorityMutex.h>
#include <FEXCore/fextl/map.h>
#include <FEXCore/fextl/memory_resource.h>
@@ -502,7 +502,7 @@ void OpDispatchBuilder::LEAVEOp(OpcodeArgs) {
auto NewGPR = Pop(OperandSize, SP);
// Store the new stack pointer
StoreGPRRegister(X86State::REG_RSP, SP, OperandSize);
StoreGPRRegister(X86State::REG_RSP, SP, GPRSize);
// Store what we loaded to RBP
StoreGPRRegister(X86State::REG_RBP, NewGPR, OperandSize);
@@ -1689,8 +1689,8 @@ void OpDispatchBuilder::RotateOp(OpcodeArgs, bool Left, bool IsImmediate, bool I
}
void OpDispatchBuilder::ANDNBMIOp(OpcodeArgs) {
auto* Src1 = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.AllowUpperGarbage = true});
auto* Src2 = LoadSourceGPR(Op, Op->Src[1], Op->Flags, {.AllowUpperGarbage = true});
auto Src1 = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.AllowUpperGarbage = true});
auto Src2 = LoadSourceGPR(Op, Op->Src[1], Op->Flags, {.AllowUpperGarbage = true});
auto Dest = _Andn(OpSizeFromSrc(Op), Src2, Src1);
@@ -1703,8 +1703,8 @@ void OpDispatchBuilder::BEXTRBMIOp(OpcodeArgs) {
// along with some edge-case handling and flag setting.
LOGMAN_THROW_A_FMT(Op->InstSize >= 4, "No masking needed");
auto* Src1 = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.AllowUpperGarbage = true});
auto* Src2 = LoadSourceGPR(Op, Op->Src[1], Op->Flags, {.AllowUpperGarbage = true});
auto Src1 = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.AllowUpperGarbage = true});
auto Src2 = LoadSourceGPR(Op, Op->Src[1], Op->Flags, {.AllowUpperGarbage = true});
const auto Size = OpSizeFromSrc(Op);
const auto SrcSize = IR::OpSizeAsBits(Size);
@@ -1746,7 +1746,7 @@ void OpDispatchBuilder::BLSIBMIOp(OpcodeArgs) {
LOGMAN_THROW_A_FMT(Op->InstSize >= 4, "No masking needed");
const auto Size = OpSizeFromSrc(Op);
auto* Src = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.AllowUpperGarbage = true});
auto Src = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.AllowUpperGarbage = true});
auto NegatedSrc = _Neg(Size, Src);
auto Result = _And(Size, Src, NegatedSrc);
@@ -1767,7 +1767,7 @@ void OpDispatchBuilder::BLSMSKBMIOp(OpcodeArgs) {
LOGMAN_THROW_A_FMT(Op->InstSize >= 4, "No masking needed");
const auto Size = OpSizeFromSrc(Op);
auto* Src = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.AllowUpperGarbage = true});
auto Src = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.AllowUpperGarbage = true});
auto Result = _Xor(Size, Sub(Size, Src, 1), Src);
StoreResultGPR(Op, Result);
@@ -1787,7 +1787,7 @@ void OpDispatchBuilder::BLSRBMIOp(OpcodeArgs) {
LOGMAN_THROW_A_FMT(Op->InstSize >= 4, "No masking needed");
const auto Size = OpSizeFromSrc(Op);
auto* Src = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.AllowUpperGarbage = true});
auto Src = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.AllowUpperGarbage = true});
auto Result = _And(Size, Sub(Size, Src, 1), Src);
StoreResultGPR(Op, Result);
@@ -1807,8 +1807,8 @@ void OpDispatchBuilder::BMI2Shift(OpcodeArgs) {
const auto Size = OpSizeFromSrc(Op);
const auto SrcSize = Op->Src[0].IsGPR() ? GPRSize : Size;
auto* Src = LoadSourceGPR_WithOpSize(Op, Op->Src[0], SrcSize, Op->Flags);
auto* Shift = LoadSourceGPR_WithOpSize(Op, Op->Src[1], GPRSize, Op->Flags, {.AllowUpperGarbage = true});
auto Src = LoadSourceGPR_WithOpSize(Op, Op->Src[0], SrcSize, Op->Flags);
auto Shift = LoadSourceGPR_WithOpSize(Op, Op->Src[1], GPRSize, Op->Flags, {.AllowUpperGarbage = true});
Ref Result;
if (Op->OP == 0x6F7) {
@@ -1831,9 +1831,9 @@ void OpDispatchBuilder::BZHI(OpcodeArgs) {
// In 32-bit mode we only look at bottom 32-bit, no 8 or 16-bit BZHI so no
// need to zero-extend sources
auto* Src = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.AllowUpperGarbage = true});
auto Src = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.AllowUpperGarbage = true});
auto* Index = LoadSourceGPR(Op, Op->Src[1], Op->Flags, {.AllowUpperGarbage = true});
auto Index = LoadSourceGPR(Op, Op->Src[1], Op->Flags, {.AllowUpperGarbage = true});
// Clear the high bits specified by the index. A64 only considers bottom bits
// of the shift, so we don't need to mask bottom 8-bits ourselves.
@@ -1878,8 +1878,8 @@ void OpDispatchBuilder::RORX(OpcodeArgs) {
return;
}
auto* Src = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.AllowUpperGarbage = true});
auto* Result = Src;
auto Src = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.AllowUpperGarbage = true});
auto Result = Src;
if (DoRotation) [[likely]] {
Result = _Ror(OpSizeFromSrc(Op), Src, _InlineConstant(Amount));
}
@@ -1916,8 +1916,8 @@ void OpDispatchBuilder::MULX(OpcodeArgs) {
void OpDispatchBuilder::PDEP(OpcodeArgs) {
LOGMAN_THROW_A_FMT(Op->InstSize >= 4, "No masking needed");
auto* Input = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.AllowUpperGarbage = true});
auto* Mask = LoadSourceGPR(Op, Op->Src[1], Op->Flags, {.AllowUpperGarbage = true});
auto Input = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.AllowUpperGarbage = true});
auto Mask = LoadSourceGPR(Op, Op->Src[1], Op->Flags, {.AllowUpperGarbage = true});
auto Result = _PDep(OpSizeFromSrc(Op), Input, Mask);
StoreResultGPR(Op, Op->Dest, Result);
@@ -1925,8 +1925,8 @@ void OpDispatchBuilder::PDEP(OpcodeArgs) {
void OpDispatchBuilder::PEXT(OpcodeArgs) {
LOGMAN_THROW_A_FMT(Op->InstSize >= 4, "No masking needed");
auto* Input = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.AllowUpperGarbage = true});
auto* Mask = LoadSourceGPR(Op, Op->Src[1], Op->Flags, {.AllowUpperGarbage = true});
auto Input = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.AllowUpperGarbage = true});
auto Mask = LoadSourceGPR(Op, Op->Src[1], Op->Flags, {.AllowUpperGarbage = true});
auto Result = _PExt(OpSizeFromSrc(Op), Input, Mask);
StoreResultGPR(Op, Op->Dest, Result);
@@ -1936,8 +1936,8 @@ void OpDispatchBuilder::ADXOp(OpcodeArgs) {
const auto OpSize = OpSizeFromSrc(Op);
// Only 32/64-bit anyway so allow garbage, we use 32-bit ops.
auto* Src = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.AllowUpperGarbage = true});
auto* Before = LoadSourceGPR(Op, Op->Dest, Op->Flags, {.AllowUpperGarbage = true});
auto Src = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.AllowUpperGarbage = true});
auto Before = LoadSourceGPR(Op, Op->Dest, Op->Flags, {.AllowUpperGarbage = true});
// Handles ADCX and ADOX
const bool IsADCX = Op->OP == 0x1F6;
@@ -2427,9 +2427,10 @@ void OpDispatchBuilder::BTOp(OpcodeArgs, uint32_t SrcIndex, BTAction Action) {
auto BitSelect = (Size == (LshrSize * 8)) ? Src : Src.And(Mask);
auto LshrOpSize = IR::SizeToOpSize(LshrSize);
// OF/SF/AF/PF undefined. ZF must be preserved. We choose to preserve OF/SF
// too since we just use an rmif to insert into CF directly. We could
// optimize perhaps.
// AMD: OF/SF/ZF/AF/PF undefined.
// Intel: OF/SF/AF/PF undefined. ZF must be preserved.
// We choose to preserve ZF/OF/SF since we just use an rmif
// to insert into CF directly. We could optimize perhaps.
//
// Set CF before the action to save a move, except for complements where we
// can reuse the invert.
@@ -2547,7 +2548,10 @@ void OpDispatchBuilder::BTOp(OpcodeArgs, uint32_t SrcIndex, BTAction Action) {
Value = _Lshr(std::max(OpSize::i32Bit, GetOpSize(Value)), Value, BitSelect.Ref());
}
// OF/SF/ZF/AF/PF undefined.
// AMD: OF/SF/ZF/AF/PF undefined.
// Intel: OF/SF/AF/PF undefined. ZF must be preserved.
// We choose to preserve ZF/OF/SF since we just use an rmif
// to insert into CF directly. We could optimize perhaps.
SetCFDirect(Value, 0, true);
}
}
@@ -2643,7 +2647,10 @@ void OpDispatchBuilder::IMULOp(OpcodeArgs) {
}
// 64-bit special cased to save a move
Ref Result = Size < OpSize::i64Bit ? _Mul(OpSize::i64Bit, Src1, Src2) : nullptr;
Ref Result {};
if (Size < OpSize::i64Bit) {
Result = _Mul(OpSize::i64Bit, Src1, Src2);
}
Ref ResultHigh {};
if (Size == OpSize::i8Bit) {
// Result is stored in AX
@@ -3446,30 +3453,38 @@ void OpDispatchBuilder::CMPSOp(OpcodeArgs) {
}
void OpDispatchBuilder::LODSOp(OpcodeArgs) {
if (Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_ADDRESS_SIZE) {
LogMan::Msg::EFmt("LODSOp: Can't handle address size override (OP: 0x{:04X}, Flags: 0x{:08X})", Op->OP, Op->Flags);
if (!Is64BitMode && (Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_ADDRESS_SIZE)) {
LogMan::Msg::EFmt("LODSOp: Address size override (0x67) not supported in 32-bit mode (OP: 0x{:04X}).", Op->OP);
DecodeFailure = true;
return;
}
const auto Size = OpSizeFromSrc(Op);
OpSize AddrSize = GetStringOpSize(Op);
const bool Repeat = (Op->Flags & (FEXCore::X86Tables::DecodeFlags::FLAG_REP_PREFIX | FEXCore::X86Tables::DecodeFlags::FLAG_REPNE_PREFIX)) != 0;
if (!Repeat) {
Ref Dest_RSI = MakeSegmentAddress(X86State::REG_RSI, Op->Flags, X86Tables::DecodeFlags::FLAG_DS_PREFIX);
Ref Src_RSI = LoadGPRRegister(X86State::REG_RSI, AddrSize);
Ref Dest_RSI = AppendSegmentOffset(Src_RSI, 0, X86Tables::DecodeFlags::FLAG_DS_PREFIX, true);
auto Src = _LoadMemGPRAutoTSO(Size, Dest_RSI, Size);
StoreResultGPR(Op, Src);
// Offset the pointer
Ref TailDest_RSI = LoadGPRRegister(X86State::REG_RSI);
StoreGPRRegister(X86State::REG_RSI, OffsetByDir(TailDest_RSI, IR::OpSizeToSize(Size)));
Ref TailDest_RSI = OffsetByDir(Src_RSI, IR::OpSizeToSize(Size));
if (Is64BitMode && AddrSize == OpSize::i32Bit) {
TailDest_RSI = _Bfe(OpSize::i64Bit, 32, 0, TailDest_RSI);
StoreGPRRegister(X86State::REG_RSI, TailDest_RSI);
} else {
StoreGPRRegister(X86State::REG_RSI, TailDest_RSI, AddrSize);
}
} else {
// Calculate flags early. because end of block
CalculateDeferredFlags();
ForeachDirection([this, Op, Size](int32_t PtrDir) {
ForeachDirection([this, Op, Size, AddrSize](int32_t PtrDir) {
// XXX: Theoretically LODS could be optimized to
// RSI += {-}(RCX * Size)
// RAX = [RSI - Size]
@@ -3497,7 +3512,8 @@ void OpDispatchBuilder::LODSOp(OpcodeArgs) {
// Working loop
{
Ref Dest_RSI = MakeSegmentAddress(X86State::REG_RSI, Op->Flags, X86Tables::DecodeFlags::FLAG_DS_PREFIX);
Ref Src_RSI = LoadGPRRegister(X86State::REG_RSI, AddrSize);
Ref Dest_RSI = AppendSegmentOffset(Src_RSI, 0, X86Tables::DecodeFlags::FLAG_DS_PREFIX, true);
auto Src = _LoadMemGPRAutoTSO(Size, Dest_RSI, Size);
@@ -3513,8 +3529,13 @@ void OpDispatchBuilder::LODSOp(OpcodeArgs) {
StoreGPRRegister(X86State::REG_RCX, TailCounter);
// Offset the pointer
TailDest_RSI = Add(OpSize::i64Bit, TailDest_RSI, PtrDir * static_cast<int32_t>(IR::OpSizeToSize(Size)));
StoreGPRRegister(X86State::REG_RSI, TailDest_RSI);
TailDest_RSI = Add(AddrSize, TailDest_RSI, PtrDir * static_cast<int32_t>(IR::OpSizeToSize(Size)));
if (Is64BitMode && AddrSize == OpSize::i32Bit) {
TailDest_RSI = _Bfe(OpSize::i64Bit, 32, 0, TailDest_RSI);
StoreGPRRegister(X86State::REG_RSI, TailDest_RSI);
} else {
StoreGPRRegister(X86State::REG_RSI, TailDest_RSI, AddrSize);
}
// Jump back to the start, we have more work to do
Jump(LoopStart);
@@ -3856,7 +3877,6 @@ void OpDispatchBuilder::CMPXCHGOp(OpcodeArgs) {
// This allows us to only hit the ZEXT case on failure
Ref RAXResult = NZCVSelect(OpSize::i64Bit, CondClass::EQ, Src3, Src1Lower);
// When the size is 4 we need to make sure not zext the GPR when the comparison fails
StoreGPRRegister(X86State::REG_RAX, RAXResult);
} else {
StoreGPRRegister(X86State::REG_RAX, Src1Lower, Size);
@@ -3870,7 +3890,7 @@ void OpDispatchBuilder::CMPXCHGOp(OpcodeArgs) {
if (GPRSize == OpSize::i64Bit && Size == OpSize::i32Bit) {
Src2Lower = _Bfe(GPRSize, IR::OpSizeAsBits(Size), 0, Src2);
}
Ref DestResult = Trivial ? Src2 : NZCVSelect(OpSize::i64Bit, CondClass::EQ, Src2Lower, Src1);
Ref DestResult = Trivial ? Src2Lower : NZCVSelect(OpSize::i64Bit, CondClass::EQ, Src2Lower, Src1);
// Store in to GPR Dest
if (GPRSize == OpSize::i64Bit && Size == OpSize::i32Bit) {
@@ -4213,6 +4233,94 @@ void OpDispatchBuilder::UpdatePrefixFromSegment(Ref Segment, uint32_t SegmentReg
}
}
uint64_t OpDispatchBuilder::CalcAddress(const X86Tables::DecodedOp& Op, const X86Tables::DecodedOperand& Operand, bool IsLoad) {
if constexpr (!Context::BLOCK_DEBUGGING) {
LOGMAN_MSG_A_FMT("Tried to calculate address without block debugging enabled!");
FEX_UNREACHABLE;
}
const auto GPRSize = GetGPROpSize();
const auto GPRMask = GPRSize == OpSize::i64Bit ? ~0ULL : ~0U;
// This makes the assumption that InternalThreadState is synchronized at the point of call!
uint64_t Ptr {};
if (Operand.IsLiteral()) {
Ptr = Operand.Literal();
if (Operand.Data.Literal.Size != 8 && IsLoad) {
// zero extend
uint64_t width = Operand.Data.Literal.Size * 8;
Ptr &= ((1ULL << width) - 1);
}
} else if (Operand.IsGPR()) {
// Not a memory source.
return ~0ULL;
} else if (Operand.IsGPRDirect()) {
Ptr = Thread->CurrentFrame->State.gregs[Operand.Data.GPR.GPR] & GPRMask;
} else if (Operand.IsGPRIndirect() || Operand.IsGPRIndirectRelocation()) {
Ptr = Thread->CurrentFrame->State.gregs[Operand.Data.GPR.GPR] & GPRMask;
Ptr += static_cast<int32_t>(Operand.Data.GPRIndirect.Displacement);
} else if (Operand.IsRIPRelative() || Operand.IsRIPRelativeRelocation()) {
// 64-bit is RIP relative, while 32-bit is absolute.
if (Is64BitMode) {
Ptr = Op->PC + Op->InstSize + static_cast<int32_t>(Operand.Data.RIPLiteral.Value) - Entry;
} else {
Ptr = Operand.Data.RIPLiteral.Value;
}
} else if (Operand.IsSIB() || Operand.IsSIBRelocation()) {
const bool IsVSIB = IsLoad && ((Op->Flags & X86Tables::DecodeFlags::FLAG_VSIB_BYTE) != 0);
if (IsVSIB) {
// TODO: Unhandled.
return ~0ULL;
}
if (Operand.Data.SIB.Base != FEXCore::X86State::REG_INVALID) {
Ptr = Thread->CurrentFrame->State.gregs[Operand.Data.SIB.Base] & GPRMask;
}
if (Operand.Data.SIB.Index != FEXCore::X86State::REG_INVALID) {
Ptr += (Thread->CurrentFrame->State.gregs[Operand.Data.SIB.Index] * Operand.Data.SIB.Scale) & GPRMask;
}
Ptr += static_cast<int32_t>(Operand.Data.SIB.Offset);
}
auto AppendSegment = [&](uint64_t Ptr, uint32_t Flags, uint32_t DefaultPrefix = FEXCore::X86Tables::DecodeFlags::FLAG_NO_PREFIX,
bool Override = false) -> uint64_t {
uint32_t Prefix = Flags & FEXCore::X86Tables::DecodeFlags::FLAG_SEGMENTS;
if (Is64BitMode) {
if (Prefix == FEXCore::X86Tables::DecodeFlags::FLAG_FS_PREFIX) {
return Ptr + Thread->CurrentFrame->State.fs_cached;
} else if (Prefix == FEXCore::X86Tables::DecodeFlags::FLAG_GS_PREFIX) {
return Ptr + Thread->CurrentFrame->State.gs_cached;
}
// If there was any other segment in 64bit then it is ignored
} else {
if (Prefix == FEXCore::X86Tables::DecodeFlags::FLAG_NO_PREFIX || Override) {
// If there was no prefix then use the default one if available
// Or the argument only uses a specific prefix (with override set)
Prefix = DefaultPrefix;
}
// With the segment register optimization we store the GDT bases directly in the segment register to remove indexed loads
switch (Prefix) {
[[likely]] case FEXCore::X86Tables::DecodeFlags::FLAG_NO_PREFIX:
return Ptr;
case FEXCore::X86Tables::DecodeFlags::FLAG_ES_PREFIX: return Ptr + Thread->CurrentFrame->State.es_cached;
case FEXCore::X86Tables::DecodeFlags::FLAG_CS_PREFIX: return Ptr + Thread->CurrentFrame->State.cs_cached;
case FEXCore::X86Tables::DecodeFlags::FLAG_SS_PREFIX: return Ptr + Thread->CurrentFrame->State.ss_cached;
case FEXCore::X86Tables::DecodeFlags::FLAG_DS_PREFIX: return Ptr + Thread->CurrentFrame->State.ds_cached;
case FEXCore::X86Tables::DecodeFlags::FLAG_FS_PREFIX: return Ptr + Thread->CurrentFrame->State.fs_cached;
case FEXCore::X86Tables::DecodeFlags::FLAG_GS_PREFIX: return Ptr + Thread->CurrentFrame->State.gs_cached;
default: FEX_UNREACHABLE;
}
}
return Ptr;
};
return AppendSegment(Ptr, Op->Flags);
};
AddressMode OpDispatchBuilder::DecodeAddress(const X86Tables::DecodedOp& Op, const X86Tables::DecodedOperand& Operand,
MemoryAccessType AccessType, bool IsLoad) {
const auto GPRSize = GetGPROpSize();
@@ -4303,6 +4411,7 @@ Ref OpDispatchBuilder::LoadSource_WithOpSize(RegClass Class, const X86Tables::De
auto [Align, LoadData, ForceLoad, AccessType, AllowUpperGarbage] = Options;
AddressMode A = DecodeAddress(Op, Operand, AccessType, true /* IsLoad */);
Ref Result {};
if (Operand.IsGPR()) {
const auto gpr = Operand.Data.GPR.GPR;
const auto highIndex = Operand.Data.GPR.HighBits ? 1 : 0;
@@ -4332,22 +4441,35 @@ Ref OpDispatchBuilder::LoadSource_WithOpSize(RegClass Class, const X86Tables::De
}
}
if ((IsOperandMem(Operand, true) && LoadData) || ForceLoad) {
const bool ShouldLoad = (IsOperandMem(Operand, true) && LoadData) || ForceLoad;
if (ShouldLoad) {
if (OpSize == OpSize::f80Bit) {
Ref MemSrc = LoadEffectiveAddress(this, A, GetGPROpSize(), true);
if (CTX->HostFeatures.SupportsSVE128 || CTX->HostFeatures.SupportsSVE256) {
return _LoadMemX87SVEOptPredicate(OpSize::i128Bit, OpSize::i16Bit, MemSrc);
if (CTX->HostFeatures.SupportsSVE()) {
Result = _LoadMemX87SVEOptPredicate(OpSize::i128Bit, OpSize::i16Bit, MemSrc);
} else {
// For X87 extended doubles, Split the load.
auto Res = _LoadMem(Class, OpSize::i64Bit, MemSrc, Align == OpSize::iInvalid ? OpSize : Align);
return _VLoadVectorElement(OpSize::i128Bit, OpSize::i16Bit, Res, 4, Add(OpSize::i64Bit, MemSrc, 8));
Result = _VLoadVectorElement(OpSize::i128Bit, OpSize::i16Bit, Res, 4, Add(OpSize::i64Bit, MemSrc, 8));
}
} else {
Result = _LoadMemAutoTSO(Class, OpSize, A, Align == OpSize::iInvalid ? OpSize : Align);
}
} else {
Result = LoadEffectiveAddress(this, A, GetGPROpSize(), false, AllowUpperGarbage);
}
if constexpr (Context::BLOCK_DEBUGGING) {
if (ShouldLoad && CTX->BlockDebuggerTracker.IsSingleStepTarget(Entry)) {
uint64_t Ptr = CalcAddress(Op, Operand, true);
if (CTX->BlockDebuggerTracker.ContainsReadWatchPoint(Ptr, OpSizeToSize(OpSize))) {
// It's up to the developer if they want more advanced debugging logic here.
LogMan::Msg::IFmt("Entrypoint 0x{:x} will hit read watch: [0x{:x}, 0x{:x})", Entry, Ptr, Ptr + OpSizeToSize(OpSize));
}
}
return _LoadMemAutoTSO(Class, OpSize, A, Align == OpSize::iInvalid ? OpSize : Align);
} else {
return LoadEffectiveAddress(this, A, GetGPROpSize(), false, AllowUpperGarbage);
}
return Result;
}
Ref OpDispatchBuilder::LoadGPRRegister(uint32_t GPR, IR::OpSize Size, uint8_t Offset, bool AllowUpperGarbage) {
@@ -4462,7 +4584,7 @@ void OpDispatchBuilder::StoreResult_WithOpSize(RegClass Class, FEXCore::X86Table
if (OpSize == OpSize::f80Bit) {
Ref MemStoreDst = LoadEffectiveAddress(this, A, GetGPROpSize(), true);
if (CTX->HostFeatures.SupportsSVE128 || CTX->HostFeatures.SupportsSVE256) {
if (CTX->HostFeatures.SupportsSVE()) {
_StoreMemX87SVEOptPredicate(OpSize::i128Bit, OpSize::i16Bit, Src, MemStoreDst);
} else {
// For X87 extended doubles, split before storing
@@ -4473,6 +4595,16 @@ void OpDispatchBuilder::StoreResult_WithOpSize(RegClass Class, FEXCore::X86Table
} else {
_StoreMemAutoTSO(Class, OpSize, A, Src, Align == OpSize::iInvalid ? OpSize : Align);
}
if constexpr (Context::BLOCK_DEBUGGING) {
if (CTX->BlockDebuggerTracker.IsSingleStepTarget(Entry)) {
uint64_t Ptr = CalcAddress(Op, Operand, false);
if (CTX->BlockDebuggerTracker.ContainsWriteWatchPoint(Ptr, OpSizeToSize(OpSize))) {
// It's up to the developer if they want more advanced debugging logic here.
LogMan::Msg::IFmt("Entrypoint 0x{:x} will hit write watch: [0x{:x}, 0x{:x})", Entry, Ptr, Ptr + OpSizeToSize(OpSize));
}
}
}
}
void OpDispatchBuilder::StoreResult(RegClass Class, X86Tables::DecodedOp Op, const X86Tables::DecodedOperand& Operand, Ref Src,
@@ -4484,11 +4616,10 @@ void OpDispatchBuilder::StoreResult(RegClass Class, X86Tables::DecodedOp Op, Ref
StoreResult(Class, Op, Op->Dest, Src, Align, AccessType);
}
OpDispatchBuilder::OpDispatchBuilder(FEXCore::Context::ContextImpl* ctx)
OpDispatchBuilder::OpDispatchBuilder(FEXCore::Context::ContextImpl* ctx, FEXCore::Core::InternalThreadState* Thread)
: IREmitter {ctx->OpDispatcherAllocator, ctx->HostFeatures.SupportsTSOImm9}
, CTX {ctx} {
ResetWorkingList();
, CTX {ctx}
, Thread {Thread} {
if (CTX->HostFeatures.SupportsAVX && CTX->HostFeatures.SupportsSVE256) {
SaveAVXStateFunc = &OpDispatchBuilder::SaveAVXState;
RestoreAVXStateFunc = &OpDispatchBuilder::RestoreAVXState;
@@ -4501,7 +4632,8 @@ OpDispatchBuilder::OpDispatchBuilder(FEXCore::Context::ContextImpl* ctx)
}
void OpDispatchBuilder::ResetWorkingList() {
IREmitter::ResetWorkingList();
IREmitter::ReownOrClaimBuffer();
JumpTargets.clear();
BlockSetRIP = false;
DecodeFailure = false;
@@ -4645,35 +4777,36 @@ void OpDispatchBuilder::INTOp(OpcodeArgs) {
case 0xCD: { // INT imm8
uint8_t Literal = Op->Src[0].Literal();
#ifndef _WIN32
constexpr uint8_t SYSCALL_LITERAL = 0x80;
if (Literal == SYSCALL_LITERAL) {
if (Is64BitMode) [[unlikely]] {
LogMan::Msg::EFmt("[Unsupported] Trying to execute 32-bit syscall from a 64-bit process.");
UnhandledOp(Op);
if (CTX->HostFeatures.HostType == FEXCore::HostFeatures::HostTypeEnum::Linux) {
constexpr uint8_t SYSCALL_LITERAL = 0x80;
if (Literal == SYSCALL_LITERAL) {
if (Is64BitMode) [[unlikely]] {
LogMan::Msg::EFmt("[Unsupported] Trying to execute 32-bit syscall from a 64-bit process.");
UnhandledOp(Op);
return;
}
// Syscall on linux
SyscallOp(Op, false);
return;
}
} else if (CTX->HostFeatures.HostType == FEXCore::HostFeatures::HostTypeEnum::Wow64 ||
CTX->HostFeatures.HostType == FEXCore::HostFeatures::HostTypeEnum::Arm64ec) {
constexpr uint8_t SYSCALL_LITERAL = 0x2E;
if (Literal == SYSCALL_LITERAL) {
// Can be used for both 64-bit and 32-bit syscalls on windows
SyscallOp(Op, false);
return;
}
// Syscall on linux
SyscallOp(Op, false);
return;
}
#else
constexpr uint8_t SYSCALL_LITERAL = 0x2E;
if (Literal == SYSCALL_LITERAL) {
// Can be used for both 64-bit and 32-bit syscalls on windows
SyscallOp(Op, false);
return;
}
#endif
#ifdef ARCHITECTURE_arm64ec
// This is used when QueryPerformanceCounter is called on recent Windows versions, it causes CNTVCT to be written into RAX.
constexpr uint8_t GET_CNTVCT_LITERAL = 0x81;
if (Literal == GET_CNTVCT_LITERAL) {
StoreGPRRegister(X86State::REG_RAX, _CycleCounter(false));
return;
if (CTX->HostFeatures.HostType == FEXCore::HostFeatures::HostTypeEnum::Arm64ec) {
// This is used when QueryPerformanceCounter is called on recent Windows versions, it causes CNTVCT to be written into RAX.
constexpr uint8_t GET_CNTVCT_LITERAL = 0x81;
if (Literal == GET_CNTVCT_LITERAL) {
StoreGPRRegister(X86State::REG_RAX, _CycleCounter(false));
return;
}
}
}
#endif
Reason.ErrorRegister = Literal << 3 | (0b010);
Reason.Signal = Core::FAULT_SIGSEGV;
@@ -4887,6 +5020,11 @@ void OpDispatchBuilder::CLZeroOp(OpcodeArgs) {
}
void OpDispatchBuilder::Prefetch(OpcodeArgs, bool ForStore, bool Stream, uint8_t Level) {
if (Op->Src[0].IsGPR()) {
// NOP instance.
return;
}
Ref DestMem = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.LoadData = false});
_Prefetch(ForStore, Stream, Level, DestMem, Invalid(), MemOffsetType::SXTX, 1);
}
@@ -4899,7 +5037,11 @@ void OpDispatchBuilder::RDTSCPOp(OpcodeArgs) {
// - Explicitly use an MFENCE before this instruction if you want this behaviour
// This instruction is not an execution fence, so subsequent instructions can execute after this
// - Explicitly use an LFENCE after RDTSCP if you want to block this behaviour
if (CTX->HostFeatures.HostType != FEXCore::HostFeatures::HostTypeEnum::Linux && !CTX->HostFeatures.SupportsCPUIndexInTPIDRRO) {
// RDTSCP is unsupported on Win32 platforms if TPIDRRO isn't supported.
UnimplementedOp(Op);
return;
}
auto Counter = CycleCounter(true);
auto ID = _ProcessorID();
@@ -4909,6 +5051,11 @@ void OpDispatchBuilder::RDTSCPOp(OpcodeArgs) {
}
void OpDispatchBuilder::RDPIDOp(OpcodeArgs) {
if (CTX->HostFeatures.HostType != FEXCore::HostFeatures::HostTypeEnum::Linux && !CTX->HostFeatures.SupportsCPUIndexInTPIDRRO) {
// RDTSCP is unsupported on Win32 platforms if TPIDRRO isn't supported.
UnimplementedOp(Op);
return;
}
StoreResultGPR(Op, _ProcessorID());
}
@@ -4918,6 +5065,7 @@ void OpDispatchBuilder::CRC32(OpcodeArgs) {
return;
}
const auto GPRSize = GetGPROpSize();
const auto SrcSize = OpSizeFromSrc(Op);
// Destination GPR size is always 4 or 8 bytes depending on widening
const auto DstSize = Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_REX_WIDENING ? OpSize::i64Bit : OpSize::i32Bit;
@@ -4926,16 +5074,15 @@ void OpDispatchBuilder::CRC32(OpcodeArgs) {
// Incoming memory is 8, 16, 32, or 64
Ref Src {};
if (Op->Src[0].IsGPR()) {
Src = LoadSourceGPR_WithOpSize(Op, Op->Src[0], GPRSize, Op->Flags);
Src = LoadSourceGPR_WithOpSize(Op, Op->Src[0], SrcSize, Op->Flags, {.AllowUpperGarbage = true});
} else {
Src = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.Align = OpSize::i8Bit});
}
auto Result = _CRC32(Dest, Src, OpSizeFromSrc(Op));
auto Result = _CRC32(Dest, Src, SrcSize);
StoreResultGPR_WithOpSize(Op, Op->Dest, Result, DstSize);
}
template<bool Reseed>
void OpDispatchBuilder::RDRANDOp(OpcodeArgs) {
void OpDispatchBuilder::RDRANDOp(OpcodeArgs, bool Reseed) {
if (!CTX->HostFeatures.SupportsRAND) {
UnimplementedOp(Op);
return;
@@ -4960,9 +5107,6 @@ void OpDispatchBuilder::RDRANDOp(OpcodeArgs) {
}
}
template void OpDispatchBuilder::RDRANDOp<true>(OpcodeArgs);
template void OpDispatchBuilder::RDRANDOp<false>(OpcodeArgs);
void OpDispatchBuilder::BreakOp(OpcodeArgs, FEXCore::IR::BreakDefinition BreakDefinition) {
const auto GPRSize = GetGPROpSize();
+79 -139
View File
@@ -303,9 +303,11 @@ public:
StartNewBlock();
}
OpDispatchBuilder(FEXCore::Context::ContextImpl* ctx);
OpDispatchBuilder(FEXCore::Context::ContextImpl* ctx, FEXCore::Core::InternalThreadState* Thread);
// Should only be called at the start of IR Emission.
void ResetWorkingList();
void ResetDecodeFailure() {
NeedsBlockEnd = DecodeFailure = false;
}
@@ -358,7 +360,7 @@ public:
void MOVGPRNTOp(OpcodeArgs);
void MOVVectorAlignedOp(OpcodeArgs);
void MOVVectorUnalignedOp(OpcodeArgs);
void MOVVectorNTOp(OpcodeArgs);
void MOVVectorNTOp(OpcodeArgs, bool IsAVX);
void ALUOp(OpcodeArgs, FEXCore::IR::IROps ALUIROp, FEXCore::IR::IROps AtomicFetchOp, unsigned SrcIdx);
void LSLOp(OpcodeArgs);
void INTOp(OpcodeArgs);
@@ -468,8 +470,7 @@ public:
void AAMOp(OpcodeArgs);
void AADOp(OpcodeArgs);
void XLATOp(OpcodeArgs);
template<bool Reseed>
void RDRANDOp(OpcodeArgs);
void RDRANDOp(OpcodeArgs, bool Reseed);
enum class Segment {
FS,
@@ -499,8 +500,7 @@ public:
void VectorALUROp(OpcodeArgs, IROps IROp, IR::OpSize ElementSize);
void VectorUnaryOp(OpcodeArgs, IROps IROp, IR::OpSize ElementSize);
void RSqrt3DNowOp(OpcodeArgs, bool Duplicate);
template<FEXCore::IR::IROps IROp, IR::OpSize ElementSize>
void VectorUnaryDuplicateOp(OpcodeArgs);
void VectorUnaryDuplicateOp(OpcodeArgs, IROps IROp, IR::OpSize ElementSize);
void MOVQOp(OpcodeArgs, VectorOpType VectorType);
void MOVQMMXOp(OpcodeArgs);
@@ -522,36 +522,24 @@ public:
void PSLLDQ(OpcodeArgs);
void PSRAIOp(OpcodeArgs, IR::OpSize ElementSize);
void MOVDDUPOp(OpcodeArgs);
template<IR::OpSize DstElementSize>
void CVTGPR_To_FPR(OpcodeArgs);
template<IR::OpSize SrcElementSize, bool HostRoundingMode>
void CVTFPR_To_GPR(OpcodeArgs);
template<IR::OpSize SrcElementSize, bool Widen>
void Vector_CVT_Int_To_Float(OpcodeArgs);
template<IR::OpSize DstElementSize, IR::OpSize SrcElementSize>
void Scalar_CVT_Float_To_Float(OpcodeArgs);
void CVTFPR_To_GPR(OpcodeArgs, IR::OpSize SrcElementSize, bool HostRoundingMode);
void Vector_CVT_Int_To_Float(OpcodeArgs, IR::OpSize SrcElementSize, bool Widen, bool IsAVX);
void Vector_CVT_Float_To_Float(OpcodeArgs, IR::OpSize DstElementSize, IR::OpSize SrcElementSize, bool IsAVX);
template<IR::OpSize SrcElementSize, bool HostRoundingMode>
void Vector_CVT_Float_To_Int(OpcodeArgs);
void Vector_CVT_Float_To_Int(OpcodeArgs, IR::OpSize SrcElementSize, bool HostRoundingMode, bool IsAVX);
void MMX_To_XMM_Vector_CVT_Int_To_Float(OpcodeArgs);
template<IR::OpSize SrcElementSize, bool HostRoundingMode>
void XMM_To_MMX_Vector_CVT_Float_To_Int(OpcodeArgs);
void XMM_To_MMX_Vector_CVT_Float_To_Int(OpcodeArgs, IR::OpSize SrcElementSize, bool HostRoundingMode);
void MASKMOVOp(OpcodeArgs);
void MOVBetweenGPR_FPR(OpcodeArgs, VectorOpType VectorType);
void TZCNT(OpcodeArgs);
void LZCNT(OpcodeArgs);
template<IR::OpSize ElementSize>
void VFCMPOp(OpcodeArgs);
void VFCMPOp(OpcodeArgs, IR::OpSize ElementSize);
void SHUFOp(OpcodeArgs, IR::OpSize ElementSize);
template<IR::OpSize ElementSize>
void PINSROp(OpcodeArgs);
void PINSROp(OpcodeArgs, IR::OpSize ElementSize);
void InsertPSOp(OpcodeArgs);
void PExtrOp(OpcodeArgs, IR::OpSize ElementSize);
template<IR::OpSize ElementSize>
void PSIGN(OpcodeArgs);
template<IR::OpSize ElementSize>
void VPSIGN(OpcodeArgs);
void PSIGN(OpcodeArgs, IR::OpSize ElementSize);
void VPSIGN(OpcodeArgs, IR::OpSize ElementSize);
// BMI1 Ops
void ANDNBMIOp(OpcodeArgs);
@@ -574,53 +562,32 @@ public:
// AVX Ops
void AVXVectorXOROp(OpcodeArgs);
template<IR::OpSize ElementSize>
void AVXVectorRound(OpcodeArgs);
void AVXVectorRound(OpcodeArgs, IR::OpSize ElementSize);
template<IR::OpSize DstElementSize, IR::OpSize SrcElementSize>
void AVXScalar_CVT_Float_To_Float(OpcodeArgs);
void VectorScalarInsertALUOp(OpcodeArgs, IROps IROp, IR::OpSize ElementSize);
void AVXVectorScalarInsertALUOp(OpcodeArgs, IROps IROp, IR::OpSize ElementSize);
template<FEXCore::IR::IROps IROp, IR::OpSize ElementSize>
void VectorScalarInsertALUOp(OpcodeArgs);
template<FEXCore::IR::IROps IROp, IR::OpSize ElementSize>
void AVXVectorScalarInsertALUOp(OpcodeArgs);
template<FEXCore::IR::IROps IROp, IR::OpSize ElementSize>
void VectorScalarUnaryInsertALUOp(OpcodeArgs);
template<FEXCore::IR::IROps IROp, IR::OpSize ElementSize>
void AVXVectorScalarUnaryInsertALUOp(OpcodeArgs);
void VectorScalarUnaryInsertALUOp(OpcodeArgs, IROps IROp, IR::OpSize ElementSize);
void AVXVectorScalarUnaryInsertALUOp(OpcodeArgs, IROps IROp, IR::OpSize ElementSize);
void InsertMMX_To_XMM_Vector_CVT_Int_To_Float(OpcodeArgs);
template<IR::OpSize DstElementSize>
void InsertCVTGPR_To_FPR(OpcodeArgs);
template<IR::OpSize DstElementSize>
void AVXInsertCVTGPR_To_FPR(OpcodeArgs);
void InsertCVTGPR_To_FPR(OpcodeArgs, IR::OpSize DstElementSize);
void AVXInsertCVTGPR_To_FPR(OpcodeArgs, IR::OpSize DstElementSize);
template<IR::OpSize DstElementSize, IR::OpSize SrcElementSize>
void InsertScalar_CVT_Float_To_Float(OpcodeArgs);
template<IR::OpSize DstElementSize, IR::OpSize SrcElementSize>
void AVXInsertScalar_CVT_Float_To_Float(OpcodeArgs);
void InsertScalar_CVT_Float_To_Float(OpcodeArgs, IR::OpSize DstElementSize, IR::OpSize SrcElementSize);
void AVXInsertScalar_CVT_Float_To_Float(OpcodeArgs, IR::OpSize DstElementSize, IR::OpSize SrcElementSize);
RoundMode TranslateRoundType(uint8_t Mode);
template<IR::OpSize ElementSize>
void InsertScalarRound(OpcodeArgs);
template<IR::OpSize ElementSize>
void AVXInsertScalarRound(OpcodeArgs);
void InsertScalarRound(OpcodeArgs, IR::OpSize ElementSize);
void AVXInsertScalarRound(OpcodeArgs, IR::OpSize ElementSize);
template<IR::OpSize ElementSize>
void InsertScalarFCMPOp(OpcodeArgs);
template<IR::OpSize ElementSize>
void AVXInsertScalarFCMPOp(OpcodeArgs);
void InsertScalarFCMPOp(OpcodeArgs, IR::OpSize ElementSize);
void AVXInsertScalarFCMPOp(OpcodeArgs, IR::OpSize ElementSize);
template<IR::OpSize DstElementSize>
void AVXCVTGPR_To_FPR(OpcodeArgs);
void AVXVFCMPOp(OpcodeArgs, IR::OpSize ElementSize);
template<IR::OpSize ElementSize>
void AVXVFCMPOp(OpcodeArgs);
template<IR::OpSize ElementSize>
void VADDSUBPOp(OpcodeArgs);
void VADDSUBPOp(OpcodeArgs, IR::OpSize ElementSize);
void VAESDecOp(OpcodeArgs);
void VAESDecLastOp(OpcodeArgs);
@@ -629,34 +596,31 @@ public:
void VANDNOp(OpcodeArgs);
Ref VBLENDOpImpl(IR::OpSize VecSize, IR::OpSize ElementSize, Ref Src1, Ref Src2, Ref ZeroRegister, uint64_t Selector);
Ref VBLENDOpImpl(IR::OpSize VecSize, IR::OpSize ElementSize, Ref Src1, Ref Src2, uint64_t Selector);
void VBLENDPDOp(OpcodeArgs);
void VPBLENDDOp(OpcodeArgs);
void VPBLENDWOp(OpcodeArgs);
void VBROADCASTOp(OpcodeArgs, IR::OpSize ElementSize);
template<IR::OpSize ElementSize>
void VDPPOp(OpcodeArgs);
void VDPPOp(OpcodeArgs, IR::OpSize ElementSize);
void VEXTRACT128Op(OpcodeArgs);
template<IROps IROp, IR::OpSize ElementSize>
void VHADDPOp(OpcodeArgs);
void VHADDPOp(OpcodeArgs, IROps IROp, IR::OpSize ElementSize);
void VHSUBPOp(OpcodeArgs, IR::OpSize ElementSize);
void VINSERTOp(OpcodeArgs);
void VINSERTPSOp(OpcodeArgs);
template<IR::OpSize ElementSize, bool IsStore>
void VMASKMOVOp(OpcodeArgs);
void VMASKMOVOp(OpcodeArgs, IR::OpSize ElementSize, bool IsStore);
void VMOVHPOp(OpcodeArgs);
void VMOVLPOp(OpcodeArgs);
void VMOVDDUPOp(OpcodeArgs);
void VMOVSHDUPOp(OpcodeArgs);
void VMOVSLDUPOp(OpcodeArgs);
void VMOVSHDUPOp(OpcodeArgs, bool IsAVX);
void VMOVSLDUPOp(OpcodeArgs, bool IsAVX);
void VMOVSDOp(OpcodeArgs);
void VMOVSSOp(OpcodeArgs);
@@ -667,15 +631,14 @@ public:
void VMPSADBWOp(OpcodeArgs);
void VPACKSSOp(OpcodeArgs, IR::OpSize ElementSize);
void VPACKUSOp(OpcodeArgs, IR::OpSize ElementSize);
void VPALIGNROp(OpcodeArgs);
void VPCMPESTRIOp(OpcodeArgs);
void VPCMPESTRMOp(OpcodeArgs);
void VPCMPISTRIOp(OpcodeArgs);
void VPCMPISTRMOp(OpcodeArgs);
void VPCMPESTRIOp(OpcodeArgs, bool IsAVX);
void VPCMPESTRMOp(OpcodeArgs, bool IsAVX);
void VPCMPISTRIOp(OpcodeArgs, bool IsAVX);
void VPCMPISTRMOp(OpcodeArgs, bool IsAVX);
void VCVTPH2PSOp(OpcodeArgs);
void VCVTPS2PHOp(OpcodeArgs);
@@ -688,36 +651,28 @@ public:
void VPERMILImmOp(OpcodeArgs, IR::OpSize ElementSize);
Ref VPERMILRegOpImpl(OpSize DstSize, IR::OpSize ElementSize, Ref Src, Ref Indices);
template<IR::OpSize ElementSize>
void VPERMILRegOp(OpcodeArgs);
void VPERMILRegOp(OpcodeArgs, IR::OpSize ElementSize);
void VPHADDSWOp(OpcodeArgs);
void VPHSUBOp(OpcodeArgs, IR::OpSize ElementSize);
void VPHSUBSWOp(OpcodeArgs);
void VPINSRBOp(OpcodeArgs);
void VPINSRBWOp(OpcodeArgs, IR::OpSize ElementSize);
void VPINSRDQOp(OpcodeArgs);
void VPINSRWOp(OpcodeArgs);
void VPMADDUBSWOp(OpcodeArgs);
void VPMADDWDOp(OpcodeArgs);
template<bool IsStore>
void VPMASKMOVOp(OpcodeArgs);
void VPMASKMOVOp(OpcodeArgs, bool IsStore);
void VPMULHRSWOp(OpcodeArgs);
template<bool Signed>
void VPMULHWOp(OpcodeArgs);
template<IR::OpSize ElementSize, bool Signed>
void VPMULLOp(OpcodeArgs);
void VPMULHWOp(OpcodeArgs, bool Signed);
void VPMULLOp(OpcodeArgs, IR::OpSize ElementSize, bool Signed);
void VPSADBWOp(OpcodeArgs);
void VPSHUFBOp(OpcodeArgs);
void VPSHUFWOp(OpcodeArgs, IR::OpSize ElementSize, bool Low);
void VPSLLOp(OpcodeArgs, IR::OpSize ElementSize);
@@ -726,7 +681,6 @@ public:
void VPSLLVOp(OpcodeArgs);
void VPSRAOp(OpcodeArgs, IR::OpSize ElementSize);
void VPSRAIOp(OpcodeArgs, IR::OpSize ElementSize);
void VPSRAVDOp(OpcodeArgs);
@@ -734,17 +688,14 @@ public:
void VPSRLDOp(OpcodeArgs, IR::OpSize ElementSize);
void VPSRLDQOp(OpcodeArgs);
void VPSRLIOp(OpcodeArgs, IR::OpSize ElementSize);
void VPUNPCKHOp(OpcodeArgs, IR::OpSize ElementSize);
void VPUNPCKLOp(OpcodeArgs, IR::OpSize ElementSize);
void VPSRLIOp(OpcodeArgs, IR::OpSize ElementSize);
void VSHUFOp(OpcodeArgs, IR::OpSize ElementSize);
template<IR::OpSize ElementSize>
void VTESTPOp(OpcodeArgs);
void VTESTPOp(OpcodeArgs, IR::OpSize ElementSize);
void VZEROOp(OpcodeArgs);
@@ -828,32 +779,24 @@ public:
void XSaveOp(OpcodeArgs);
void PAlignrOp(OpcodeArgs);
template<IR::OpSize ElementSize>
void UCOMISxOp(OpcodeArgs);
void UCOMISxOp(OpcodeArgs, IR::OpSize ElementSize);
void LDMXCSR(OpcodeArgs);
void STMXCSR(OpcodeArgs);
template<IR::OpSize ElementSize>
void PACKUSOp(OpcodeArgs);
void PACKUSOp(OpcodeArgs, IR::OpSize ElementSize);
void PACKSSOp(OpcodeArgs, IR::OpSize ElementSize);
template<IR::OpSize ElementSize>
void PACKSSOp(OpcodeArgs);
void PMULLOp(OpcodeArgs, IR::OpSize ElementSize, bool Signed);
template<IR::OpSize ElementSize, bool Signed>
void PMULLOp(OpcodeArgs);
void MOVQ2DQ(OpcodeArgs, bool ToXMM);
template<bool ToXMM>
void MOVQ2DQ(OpcodeArgs);
template<IR::OpSize ElementSize>
void ADDSUBPOp(OpcodeArgs);
void ADDSUBPOp(OpcodeArgs, IR::OpSize ElementSize);
void PFNACCOp(OpcodeArgs);
void PFPNACCOp(OpcodeArgs);
void PSWAPDOp(OpcodeArgs);
template<uint8_t CompType>
void VPFCMPOp(OpcodeArgs);
void VPFCMPOp(OpcodeArgs, uint8_t CompType);
void PI2FWOp(OpcodeArgs);
void PF2IWOp(OpcodeArgs);
@@ -862,16 +805,12 @@ public:
void PMADDWD(OpcodeArgs);
void PMADDUBSW(OpcodeArgs);
template<bool Signed>
void PMULHW(OpcodeArgs);
void PMULHW(OpcodeArgs, bool Signed);
void PMULHRSW(OpcodeArgs);
void MOVBEOp(OpcodeArgs);
template<IR::OpSize ElementSize>
void HSUBP(OpcodeArgs);
template<IR::OpSize ElementSize>
void PHSUB(OpcodeArgs);
void HSUBP(OpcodeArgs, IR::OpSize ElementSize);
void PHSUB(OpcodeArgs, IR::OpSize ElementSize);
void PHADDS(OpcodeArgs);
void PHSUBS(OpcodeArgs);
@@ -900,12 +839,12 @@ public:
void SHA256MSG2Op(OpcodeArgs);
void SHA256RNDS2Op(OpcodeArgs);
void AESImcOp(OpcodeArgs);
void AESImcOp(OpcodeArgs, bool IsAVX);
void AESEncOp(OpcodeArgs);
void AESEncLastOp(OpcodeArgs);
void AESDecOp(OpcodeArgs);
void AESDecLastOp(OpcodeArgs);
void AESKeyGenAssist(OpcodeArgs);
void AESKeyGenAssist(OpcodeArgs, bool IsAVX);
void VFMAImpl(OpcodeArgs, IROps IROp, bool Scalar, uint8_t Src1Idx, uint8_t Src2Idx, uint8_t AddendIdx);
void VFMAddSubImpl(OpcodeArgs, bool AddSub, uint8_t Src1Idx, uint8_t Src2Idx, uint8_t AddendIdx);
@@ -918,25 +857,24 @@ public:
};
RefVSIB LoadVSIB(const X86Tables::DecodedOp& Op, const X86Tables::DecodedOperand& Operand, uint32_t Flags);
template<OpSize AddrElementSize>
void VPGATHER(OpcodeArgs);
void VPGATHER(OpcodeArgs, OpSize AddrElementSize);
template<IR::OpSize ElementSize, IR::OpSize DstElementSize, bool Signed>
void ExtendVectorElements(OpcodeArgs);
template<IR::OpSize ElementSize>
void VectorRound(OpcodeArgs);
void AVXExtendVectorElements(OpcodeArgs, IR::OpSize ElementSize, IR::OpSize DstElementSize, bool Signed);
void ExtendVectorElements(OpcodeArgs, IR::OpSize ElementSize, IR::OpSize DstElementSize, bool Signed);
Ref VectorBlend(OpSize Size, IR::OpSize ElementSize, Ref Src1, Ref Src2, uint8_t Selector);
void VectorRound(OpcodeArgs, IR::OpSize ElementSize);
template<IR::OpSize ElementSize>
void VectorBlend(OpcodeArgs);
Ref VectorBlendImpl(OpSize Size, IR::OpSize ElementSize, Ref Src1, Ref Src2, uint8_t Selector);
void VectorBlend(OpcodeArgs, IR::OpSize ElementSize);
void VectorVariableBlend(OpcodeArgs, IR::OpSize ElementSize);
void PTestOpImpl(OpSize Size, Ref Dest, Ref Src);
void PTestOp(OpcodeArgs);
void AVXPHMINPOSUWOp(OpcodeArgs);
void PHMINPOSUWOp(OpcodeArgs);
template<IR::OpSize ElementSize>
void DPPOp(OpcodeArgs);
void DPPOp(OpcodeArgs, IR::OpSize ElementSize);
void MPSADBWOp(OpcodeArgs);
void PCLMULQDQOp(OpcodeArgs);
@@ -1374,6 +1312,7 @@ private:
};
FEXCore::Context::ContextImpl* CTX {};
FEXCore::Core::InternalThreadState* Thread;
constexpr static unsigned FullNZCVMask = (1U << FEXCore::X86State::RFLAG_CF_RAW_LOC) | (1U << FEXCore::X86State::RFLAG_ZF_RAW_LOC) |
(1U << FEXCore::X86State::RFLAG_SF_RAW_LOC) | (1U << FEXCore::X86State::RFLAG_OF_RAW_LOC);
@@ -1441,7 +1380,7 @@ private:
Ref PALIGNROpImpl(OpcodeArgs, const X86Tables::DecodedOperand& Src1, const X86Tables::DecodedOperand& Src2,
const X86Tables::DecodedOperand& Imm, bool IsAVX);
void PCMPXSTRXOpImpl(OpcodeArgs, bool IsExplicit, bool IsMask);
void PCMPXSTRXOpImpl(OpcodeArgs, bool IsExplicit, bool IsMask, bool IsAVX);
Ref PHADDSOpImpl(OpSize Size, Ref Src1, Ref Src2);
@@ -1479,7 +1418,7 @@ private:
Ref PSRLDOpImpl(OpcodeArgs, IR::OpSize ElementSize, Ref Src, Ref ShiftVec);
Ref SHUFOpImpl(OpcodeArgs, IR::OpSize DstSize, IR::OpSize ElementSize, Ref Src1, Ref Src2, uint8_t Shuffle);
Ref SHUFOpImpl(IR::OpSize DstSize, IR::OpSize ElementSize, Ref Src1, Ref Src2, uint8_t Shuffle);
void VMASKMOVOpImpl(OpcodeArgs, IR::OpSize ElementSize, IR::OpSize DataSize, bool IsStore, const X86Tables::DecodedOperand& MaskOp,
const X86Tables::DecodedOperand& DataOp);
@@ -1589,6 +1528,7 @@ private:
}
AddressMode DecodeAddress(const X86Tables::DecodedOp& Op, const X86Tables::DecodedOperand& Operand, MemoryAccessType AccessType, bool IsLoad);
uint64_t CalcAddress(const X86Tables::DecodedOp& Op, const X86Tables::DecodedOperand& Operand, bool IsLoad);
Ref LoadSource(RegClass Class, const X86Tables::DecodedOp& Op, const X86Tables::DecodedOperand& Operand, uint32_t Flags,
const LoadSourceOptions& Options = {});
@@ -1666,7 +1606,7 @@ private:
[[nodiscard]]
static uint32_t GPROffset(X86State::X86Reg reg) {
LOGMAN_THROW_A_FMT(reg <= X86State::X86Reg::REG_R15, "Invalid reg used");
return static_cast<uint32_t>(offsetof(Core::CPUState, gregs[static_cast<size_t>(reg)]));
return static_cast<uint32_t>(ARRAY_OFFSETOF(Core::CPUState, gregs, reg));
}
[[nodiscard]]
@@ -1885,15 +1825,15 @@ private:
// For DF, we need to transform 0/1 into 1/-1
StoreDF(_SubShift(OpSize::i64Bit, Constant(1), Value, ShiftType::LSL, 1));
} else if (BitOffset == FEXCore::X86State::RFLAG_TF_RAW_LOC) {
auto PackedTF = _LoadContextGPR(OpSize::i8Bit, offsetof(FEXCore::Core::CPUState, flags[BitOffset]));
auto PackedTF = _LoadContextGPR(OpSize::i8Bit, ARRAY_OFFSETOF(FEXCore::Core::CPUState, flags, BitOffset));
// An exception should still be raised after an instruction that unsets TF, leave the unblocked bit set but unset
// the TF bit to cause such behaviour. The handling code at the start of the next block will then unset the
// unblocked bit before raising the exception.
auto NewPackedTF =
_Select(OpSize::i64Bit, OpSize::i64Bit, CondClass::EQ, Value, Constant(0), _And(OpSize::i32Bit, PackedTF, Constant(~1)), Constant(1));
_StoreContextGPR(OpSize::i8Bit, NewPackedTF, offsetof(FEXCore::Core::CPUState, flags[BitOffset]));
_StoreContextGPR(OpSize::i8Bit, NewPackedTF, ARRAY_OFFSETOF(FEXCore::Core::CPUState, flags, BitOffset));
} else {
_StoreContextGPR(OpSize::i8Bit, Value, offsetof(FEXCore::Core::CPUState, flags[BitOffset]));
_StoreContextGPR(OpSize::i8Bit, Value, ARRAY_OFFSETOF(FEXCore::Core::CPUState, flags, BitOffset));
}
}
@@ -1948,8 +1888,8 @@ private:
[[nodiscard]]
static uint32_t CacheIndexToContextOffset(int Index) {
switch (Index) {
case MM0Index ... MM7Index: return offsetof(FEXCore::Core::CPUState, mm[Index - MM0Index]);
case AVXHigh0Index ... AVXHigh15Index: return offsetof(FEXCore::Core::CPUState, avx_high[Index - AVXHigh0Index][0]);
case MM0Index ... MM7Index: return ARRAY_OFFSETOF(FEXCore::Core::CPUState, mm, Index - MM0Index);
case AVXHigh0Index ... AVXHigh15Index: return ARRAY_OFFSETOF(FEXCore::Core::CPUState, avx_high, Index - AVXHigh0Index);
default: return ~0U;
}
}
@@ -2149,7 +2089,7 @@ private:
// Recover the sign bit, it is the logical DF value
return _Lshr(OpSize::i64Bit, LoadDF(), Constant(63));
} else {
return _LoadContextGPR(OpSize::i8Bit, offsetof(Core::CPUState, flags[BitOffset]));
return _LoadContextGPR(OpSize::i8Bit, ARRAY_OFFSETOF(Core::CPUState, flags, BitOffset));
}
}
@@ -603,7 +603,7 @@ void OpDispatchBuilder::AVX128_CVTFPR_To_GPR(OpcodeArgs, IR::OpSize SrcElementSi
void OpDispatchBuilder::AVX128_VANDN(OpcodeArgs) {
AVX128_VectorBinaryImpl(Op, OpSizeFromSrc(Op), OpSize::i128Bit,
[this](IR::OpSize _ElementSize, Ref Src1, Ref Src2) { return _VAndn(OpSize::i128Bit, _ElementSize, Src2, Src1); });
[this](IR::OpSize, Ref Src1, Ref Src2) { return _VAndn(OpSize::i128Bit, Src2, Src1); });
}
void OpDispatchBuilder::AVX128_VPACKSS(OpcodeArgs, IR::OpSize ElementSize) {
@@ -630,7 +630,7 @@ void OpDispatchBuilder::AVX128_VPSIGN(OpcodeArgs, IR::OpSize ElementSize) {
}
void OpDispatchBuilder::AVX128_UCOMISx(OpcodeArgs, IR::OpSize ElementSize) {
const auto SrcSize = Op->Src[0].IsGPR() ? GetGuestVectorLength() : ElementSize;
const auto SrcSize = Op->Src[0].IsGPR() ? OpSize::i128Bit : ElementSize;
auto Src1 = AVX128_LoadSource_WithOpSize(Op, Op->Dest, Op->Flags, false);
@@ -865,13 +865,14 @@ void OpDispatchBuilder::AVX128_MOVMSK(OpcodeArgs, IR::OpSize ElementSize) {
GPR = Mask4Byte(Src.Low);
}
} else if (ElementSize == OpSize::i32Bit) {
auto GPRLow = Mask4Byte(Src.Low);
auto GPRHigh = Mask4Byte(Src.High);
GPR = _Orlshl(OpSize::i64Bit, GPRLow, GPRHigh, 4);
Ref Fused = _VUnZip2(OpSize::i128Bit, OpSize::i16Bit, Src.Low, Src.High);
Fused = _VUShrI(OpSize::i128Bit, OpSize::i16Bit, Fused, 15);
auto ConstantUSHL = LoadAndCacheNamedVectorConstant(OpSize::i128Bit, NAMED_VECTOR_INCREMENTAL_U16_INDEX);
Fused = _VUShl(OpSize::i128Bit, OpSize::i16Bit, Fused, ConstantUSHL, false);
Fused = _VAddV(OpSize::i128Bit, OpSize::i16Bit, Fused);
GPR = _VExtractToGPR(OpSize::i128Bit, OpSize::i16Bit, Fused, 0);
} else {
auto GPRLow = Mask8Byte(Src.Low);
auto GPRHigh = Mask8Byte(Src.High);
GPR = _Orlshl(OpSize::i64Bit, GPRLow, GPRHigh, 2);
GPR = Mask4Byte(_VUnZip2(OpSize::i128Bit, OpSize::i32Bit, Src.Low, Src.High));
}
StoreResultGPR_WithOpSize(Op, Op->Dest, GPR, GetGPROpSize());
}
@@ -885,7 +886,7 @@ void OpDispatchBuilder::AVX128_MOVMSKB(OpcodeArgs) {
auto Mask1Byte = [this](Ref Src, Ref VMask) {
auto VCMP = _VCMPLTZ(OpSize::i128Bit, OpSize::i8Bit, Src);
auto VAnd = _VAnd(OpSize::i128Bit, OpSize::i8Bit, VCMP, VMask);
auto VAnd = _VAnd(OpSize::i128Bit, VCMP, VMask);
auto VAdd1 = _VAddP(OpSize::i128Bit, OpSize::i8Bit, VAnd, VAnd);
auto VAdd2 = _VAddP(OpSize::i128Bit, OpSize::i8Bit, VAdd1, VAdd1);
@@ -1260,26 +1261,26 @@ void OpDispatchBuilder::AVX128_VAESKeyGenAssist(OpcodeArgs) {
}
void OpDispatchBuilder::AVX128_VPCMPESTRI(OpcodeArgs) {
PCMPXSTRXOpImpl(Op, true, false);
PCMPXSTRXOpImpl(Op, true, false, true);
///< Does not zero anything.
}
void OpDispatchBuilder::AVX128_VPCMPESTRM(OpcodeArgs) {
PCMPXSTRXOpImpl(Op, true, true);
PCMPXSTRXOpImpl(Op, true, true, true);
///< Zero the upper 128-bits of hardcoded YMM0
AVX128_StoreXMMRegister(0, LoadZeroVector(OpSize::i128Bit), true);
}
void OpDispatchBuilder::AVX128_VPCMPISTRI(OpcodeArgs) {
PCMPXSTRXOpImpl(Op, false, false);
PCMPXSTRXOpImpl(Op, false, false, true);
///< Does not zero anything.
}
void OpDispatchBuilder::AVX128_VPCMPISTRM(OpcodeArgs) {
PCMPXSTRXOpImpl(Op, false, true);
PCMPXSTRXOpImpl(Op, false, true, true);
///< Zero the upper 128-bits of hardcoded YMM0
AVX128_StoreXMMRegister(0, LoadZeroVector(OpSize::i128Bit), true);
@@ -1399,13 +1400,13 @@ void OpDispatchBuilder::AVX128_VSHUF(OpcodeArgs, IR::OpSize ElementSize) {
auto Src2 = AVX128_LoadSource_WithOpSize(Op, Op->Src[1], Op->Flags, !Is128Bit);
RefPair Result {};
Result.Low = SHUFOpImpl(Op, OpSize::i128Bit, ElementSize, Src1.Low, Src2.Low, Shuffle);
Result.Low = SHUFOpImpl(OpSize::i128Bit, ElementSize, Src1.Low, Src2.Low, Shuffle);
if (Is128Bit) {
Result.High = LoadZeroVector(OpSize::i128Bit);
} else {
const uint8_t ShiftAmount = ElementSize == OpSize::i32Bit ? 0 : 2;
Result.High = SHUFOpImpl(Op, OpSize::i128Bit, ElementSize, Src1.High, Src2.High, Shuffle >> ShiftAmount);
Result.High = SHUFOpImpl(OpSize::i128Bit, ElementSize, Src1.High, Src2.High, Shuffle >> ShiftAmount);
}
AVX128_StoreResult_WithOpSize(Op, Op->Dest, Result);
}
@@ -1484,12 +1485,12 @@ void OpDispatchBuilder::AVX128_VBLEND(OpcodeArgs, IR::OpSize ElementSize) {
auto Src2 = AVX128_LoadSource_WithOpSize(Op, Op->Src[1], Op->Flags, !Is128Bit);
RefPair Result {};
Result.Low = VectorBlend(OpSize::i128Bit, ElementSize, Src1.Low, Src2.Low, Selector);
Result.Low = VectorBlendImpl(OpSize::i128Bit, ElementSize, Src1.Low, Src2.Low, Selector);
if (Is128Bit) {
Result = AVX128_Zext(Result.Low);
} else {
Result.High = VectorBlend(OpSize::i128Bit, ElementSize, Src1.High, Src2.High, (Selector >> SelectorShift));
Result.High = VectorBlendImpl(OpSize::i128Bit, ElementSize, Src1.High, Src2.High, (Selector >> SelectorShift));
}
AVX128_StoreResult_WithOpSize(Op, Op->Dest, Result);
@@ -1729,8 +1730,8 @@ void OpDispatchBuilder::AVX128_VTESTP(OpcodeArgs, IR::OpSize ElementSize) {
{
// Calculate ZF first.
auto AndLow = _VAnd(OpSize::i128Bit, OpSize::i8Bit, Src2.Low, Src1.Low);
auto AndHigh = _VAnd(OpSize::i128Bit, OpSize::i8Bit, Src2.High, Src1.High);
auto AndLow = _VAnd(OpSize::i128Bit, Src2.Low, Src1.Low);
auto AndHigh = _VAnd(OpSize::i128Bit, Src2.High, Src1.High);
auto ShiftLow = _VUShrI(OpSize::i128Bit, ElementSize, AndLow, ElementSizeInBits - 1);
auto ShiftHigh = _VUShrI(OpSize::i128Bit, ElementSize, AndHigh, ElementSizeInBits - 1);
@@ -1749,8 +1750,8 @@ void OpDispatchBuilder::AVX128_VTESTP(OpcodeArgs, IR::OpSize ElementSize) {
{
// Calculate CF Second
auto AndLow = _VAndn(OpSize::i128Bit, OpSize::i8Bit, Src2.Low, Src1.Low);
auto AndHigh = _VAndn(OpSize::i128Bit, OpSize::i8Bit, Src2.High, Src1.High);
auto AndLow = _VAndn(OpSize::i128Bit, Src2.Low, Src1.Low);
auto AndHigh = _VAndn(OpSize::i128Bit, Src2.High, Src1.High);
auto ShiftLow = _VUShrI(OpSize::i128Bit, ElementSize, AndLow, ElementSizeInBits - 1);
auto ShiftHigh = _VUShrI(OpSize::i128Bit, ElementSize, AndHigh, ElementSizeInBits - 1);
@@ -1788,11 +1789,11 @@ void OpDispatchBuilder::AVX128_PTest(OpcodeArgs) {
}
// For 256-bit, we need to unroll. This is nontrivial.
Ref Test1Low = _VAnd(OpSize::i128Bit, OpSize::i8Bit, Src1.Low, Src2.Low);
Ref Test2Low = _VAndn(OpSize::i128Bit, OpSize::i8Bit, Src2.Low, Src1.Low);
Ref Test1Low = _VAnd(OpSize::i128Bit, Src1.Low, Src2.Low);
Ref Test2Low = _VAndn(OpSize::i128Bit, Src2.Low, Src1.Low);
Ref Test1High = _VAnd(OpSize::i128Bit, OpSize::i8Bit, Src1.High, Src2.High);
Ref Test2High = _VAndn(OpSize::i128Bit, OpSize::i8Bit, Src2.High, Src1.High);
Ref Test1High = _VAnd(OpSize::i128Bit, Src1.High, Src2.High);
Ref Test2High = _VAndn(OpSize::i128Bit, Src2.High, Src1.High);
// Element size must be less than 32-bit for the sign bit tricks.
Ref Test1Max = _VUMax(OpSize::i128Bit, OpSize::i16Bit, Test1Low, Test1High);
@@ -2009,13 +2010,13 @@ void OpDispatchBuilder::AVX128_VFMAddSubImpl(OpcodeArgs, bool AddSub, uint8_t Sr
ConstantEOR = LoadAndCacheNamedVectorConstant(
OpSize::i128Bit, ElementSize == OpSize::i32Bit ? NAMED_VECTOR_PSUBADDPS_INVERT : NAMED_VECTOR_PSUBADDPD_INVERT);
}
auto InvertedSourceLow = _VXor(OpSize::i128Bit, ElementSize, Sources[AddendIdx - 1].Low, ConstantEOR);
auto InvertedSourceLow = _VXor(OpSize::i128Bit, Sources[AddendIdx - 1].Low, ConstantEOR);
Result.Low = _VFMLA(OpSize::i128Bit, ElementSize, Sources[Src1Idx - 1].Low, Sources[Src2Idx - 1].Low, InvertedSourceLow);
if (Is128Bit) {
Result.High = LoadZeroVector(OpSize::i128Bit);
} else {
auto InvertedSourceHigh = _VXor(OpSize::i128Bit, ElementSize, Sources[AddendIdx - 1].High, ConstantEOR);
auto InvertedSourceHigh = _VXor(OpSize::i128Bit, Sources[AddendIdx - 1].High, ConstantEOR);
Result.High = _VFMLA(OpSize::i128Bit, ElementSize, Sources[Src1Idx - 1].High, Sources[Src2Idx - 1].High, InvertedSourceHigh);
}
AVX128_StoreResult_WithOpSize(Op, Op->Dest, Result);
@@ -2293,9 +2294,8 @@ void OpDispatchBuilder::AVX128_VCVTPS2PH(OpcodeArgs) {
_PopRoundingMode(OldFPCR);
}
// We need to eliminate upper junk if we're storing into a register with
// a 256-bit source (VCVTPS2PH's destination for registers is an XMM).
if (Op->Src[0].IsGPR() && SrcSize == OpSize::i256Bit) {
// We need to zero the upper 128 bits if we're storing into a register
if (Op->Dest.IsGPR()) {
Result = AVX128_Zext(Result.Low);
}
@@ -26,17 +26,25 @@ void OpDispatchBuilder::SHA1NEXTEOp(OpcodeArgs) {
Ref Dest = LoadSourceFPR(Op, Op->Dest, Op->Flags);
Ref Src = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
// ARMv8 SHA1 extension provides a `SHA1H` instruction which does a fixed rotate by 30.
// This only operates on element 0 rather than element 3. We don't have the luxury of rewriting the x86 SHA algorithm to take advantage of this.
// Move the element to zero, rotate, and then move back (Using duplicates).
// Saves one instruction versus that path that doesn't support SHA extension.
auto Duplicated = _VDupElement(OpSize::i128Bit, OpSize::i32Bit, Dest, 3);
auto Sha1HRotated = _VSha1H(Duplicated);
auto RotatedNode = _VDupElement(OpSize::i128Bit, OpSize::i32Bit, Sha1HRotated, 0);
auto Tmp = _VAdd(OpSize::i128Bit, OpSize::i32Bit, Src, RotatedNode);
auto Result = _VInsElement(OpSize::i128Bit, OpSize::i32Bit, 3, 3, Src, Tmp);
Ref Result {};
if (CTX->HostFeatures.SupportsSVE128) {
auto ZeroVec = LoadZeroVector(OpSize::i128Bit);
auto Tmp = _VInsElement(OpSize::i128Bit, OpSize::i32Bit, 3, 3, ZeroVec, Dest);
auto Xar = _VXar(OpSize::i128Bit, OpSize::i32Bit, ZeroVec, Tmp, 2);
Result = _VAdd(OpSize::i128Bit, OpSize::i32Bit, Src, Xar);
} else {
// ARMv8 SHA1 extension provides a `SHA1H` instruction which does a fixed rotate by 30.
// This only operates on element 0 rather than element 3. We don't have the luxury of rewriting the x86 SHA algorithm to take advantage of this.
// Move the element to zero, rotate, and then move back (Using duplicates).
// Saves one instruction versus that path that doesn't support SHA extension.
auto Duplicated = _VDupElement(OpSize::i128Bit, OpSize::i32Bit, Dest, 3);
auto Sha1HRotated = _VSha1H(Duplicated);
auto RotatedNode = _VDupElement(OpSize::i128Bit, OpSize::i32Bit, Sha1HRotated, 0);
auto Tmp = _VAdd(OpSize::i128Bit, OpSize::i32Bit, Src, RotatedNode);
Result = _VInsElement(OpSize::i128Bit, OpSize::i32Bit, 3, 3, Src, Tmp);
}
StoreResultFPR(Op, Result);
StoreResult_WithAVXInsert(VectorOpType::SSE, RegClass::FPR, Op, Result);
}
void OpDispatchBuilder::SHA1MSG1Op(OpcodeArgs) {
@@ -50,9 +58,9 @@ void OpDispatchBuilder::SHA1MSG1Op(OpcodeArgs) {
Ref NewVec = _VExtr(OpSize::i128Bit, OpSize::i64Bit, Dest, Src, 1);
// [W0, W1, W2, W3] ^ [W2, W3, W4, W5]
Ref Result = _VXor(OpSize::i128Bit, OpSize::i8Bit, Dest, NewVec);
Ref Result = _VXor(OpSize::i128Bit, Dest, NewVec);
StoreResultFPR(Op, Result);
StoreResult_WithAVXInsert(VectorOpType::SSE, RegClass::FPR, Op, Result);
}
void OpDispatchBuilder::SHA1MSG2Op(OpcodeArgs) {
@@ -70,7 +78,7 @@ void OpDispatchBuilder::SHA1MSG2Op(OpcodeArgs) {
// The result is swizzled differently than expected
auto Result = SHADataShuffle(_VSha1SU1(Src1, Src2));
StoreResultFPR(Op, Result);
StoreResult_WithAVXInsert(VectorOpType::SSE, RegClass::FPR, Op, Result);
}
void OpDispatchBuilder::SHA1RNDS4Op(OpcodeArgs) {
@@ -99,7 +107,7 @@ void OpDispatchBuilder::SHA1RNDS4Op(OpcodeArgs) {
break;
}
const auto ZeroRegister = LoadZeroVector(OpSize::i32Bit);
const auto ZeroRegister = LoadZeroVector(OpSize::i128Bit);
Ref Src1 = SHADataShuffle(Dest);
Ref Src2 = SHADataShuffle(Src);
@@ -112,7 +120,7 @@ void OpDispatchBuilder::SHA1RNDS4Op(OpcodeArgs) {
case 3: Result = SHADataShuffle(_VSha1P(Src1, ZeroRegister, Src2)); break;
}
StoreResultFPR(Op, Result);
StoreResult_WithAVXInsert(VectorOpType::SSE, RegClass::FPR, Op, Result);
}
void OpDispatchBuilder::SHA256MSG1Op(OpcodeArgs) {
@@ -125,7 +133,7 @@ void OpDispatchBuilder::SHA256MSG1Op(OpcodeArgs) {
auto Result = _VSha256U0(Dest, Src);
StoreResultFPR(Op, Result);
StoreResult_WithAVXInsert(VectorOpType::SSE, RegClass::FPR, Op, Result);
}
void OpDispatchBuilder::SHA256MSG2Op(OpcodeArgs) {
@@ -142,7 +150,7 @@ void OpDispatchBuilder::SHA256MSG2Op(OpcodeArgs) {
auto Result = _VSha256U1(Src1, Src2);
StoreResultFPR(Op, Result);
StoreResult_WithAVXInsert(VectorOpType::SSE, RegClass::FPR, Op, Result);
}
void OpDispatchBuilder::SHA256RNDS2Op(OpcodeArgs) {
@@ -177,17 +185,22 @@ void OpDispatchBuilder::SHA256RNDS2Op(OpcodeArgs) {
auto B = _VSha256H2(EFGH, ABCD, Key);
auto Result = shuffle_abcd(A, B);
StoreResultFPR(Op, Result);
StoreResult_WithAVXInsert(VectorOpType::SSE, RegClass::FPR, Op, Result);
}
void OpDispatchBuilder::AESImcOp(OpcodeArgs) {
void OpDispatchBuilder::AESImcOp(OpcodeArgs, bool IsAVX) {
if (!CTX->HostFeatures.SupportsAES) {
UnimplementedOp(Op);
return;
}
Ref Src = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
Ref Result = _VAESImc(Src);
StoreResultFPR(Op, Result);
if (IsAVX) {
StoreResultFPR(Op, Result);
} else {
StoreResult_WithAVXInsert(VectorOpType::SSE, RegClass::FPR, Op, Result);
}
}
void OpDispatchBuilder::AESEncOp(OpcodeArgs) {
@@ -198,19 +211,30 @@ void OpDispatchBuilder::AESEncOp(OpcodeArgs) {
Ref Dest = LoadSourceFPR(Op, Op->Dest, Op->Flags);
Ref Src = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
Ref Result = _VAESEnc(OpSize::i128Bit, Dest, Src, LoadZeroVector(OpSize::i128Bit));
StoreResultFPR(Op, Result);
StoreResult_WithAVXInsert(VectorOpType::SSE, RegClass::FPR, Op, Result);
}
void OpDispatchBuilder::VAESEncOp(OpcodeArgs) {
const auto DstSize = OpSizeFromDst(Op);
const auto Is128Bit = DstSize == OpSize::i128Bit;
// TODO: Handle 256-bit VAESENC.
LOGMAN_THROW_A_FMT(Is128Bit, "256-bit VAESENC unimplemented");
const auto Is256Bit = DstSize == OpSize::i256Bit;
Ref State = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
Ref Key = LoadSourceFPR(Op, Op->Src[1], Op->Flags);
Ref Result = _VAESEnc(DstSize, State, Key, LoadZeroVector(DstSize));
Ref ZeroVec = LoadZeroVector(DstSize);
Ref Result {};
if (Is256Bit) {
// TODO: Handle as one operation once vixl supports it.
auto UpperState = _VDupElement(DstSize, OpSize::i128Bit, State, 1);
auto UpperKey = _VDupElement(DstSize, OpSize::i128Bit, Key, 1);
auto Lower = _VAESEnc(OpSize::i128Bit, State, Key, ZeroVec);
auto Upper = _VAESEnc(OpSize::i128Bit, UpperState, UpperKey, ZeroVec);
Result = _VInsElement(DstSize, OpSize::i128Bit, 1, 0, Lower, Upper);
} else {
Result = _VAESEnc(DstSize, State, Key, ZeroVec);
}
StoreResultFPR(Op, Result);
}
@@ -223,19 +247,30 @@ void OpDispatchBuilder::AESEncLastOp(OpcodeArgs) {
Ref Dest = LoadSourceFPR(Op, Op->Dest, Op->Flags);
Ref Src = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
Ref Result = _VAESEncLast(OpSize::i128Bit, Dest, Src, LoadZeroVector(OpSize::i128Bit));
StoreResultFPR(Op, Result);
StoreResult_WithAVXInsert(VectorOpType::SSE, RegClass::FPR, Op, Result);
}
void OpDispatchBuilder::VAESEncLastOp(OpcodeArgs) {
const auto DstSize = OpSizeFromDst(Op);
const auto Is128Bit = DstSize == OpSize::i128Bit;
// TODO: Handle 256-bit VAESENCLAST.
LOGMAN_THROW_A_FMT(Is128Bit, "256-bit VAESENCLAST unimplemented");
const auto Is256Bit = DstSize == OpSize::i256Bit;
Ref State = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
Ref Key = LoadSourceFPR(Op, Op->Src[1], Op->Flags);
Ref Result = _VAESEncLast(DstSize, State, Key, LoadZeroVector(DstSize));
Ref ZeroVec = LoadZeroVector(DstSize);
Ref Result {};
if (Is256Bit) {
// TODO: Handle as one operation once vixl supports it.
auto UpperState = _VDupElement(DstSize, OpSize::i128Bit, State, 1);
auto UpperKey = _VDupElement(DstSize, OpSize::i128Bit, Key, 1);
auto Lower = _VAESEncLast(OpSize::i128Bit, State, Key, ZeroVec);
auto Upper = _VAESEncLast(OpSize::i128Bit, UpperState, UpperKey, ZeroVec);
Result = _VInsElement(DstSize, OpSize::i128Bit, 1, 0, Lower, Upper);
} else {
Result = _VAESEncLast(DstSize, State, Key, ZeroVec);
}
StoreResultFPR(Op, Result);
}
@@ -248,19 +283,30 @@ void OpDispatchBuilder::AESDecOp(OpcodeArgs) {
Ref Dest = LoadSourceFPR(Op, Op->Dest, Op->Flags);
Ref Src = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
Ref Result = _VAESDec(OpSize::i128Bit, Dest, Src, LoadZeroVector(OpSize::i128Bit));
StoreResultFPR(Op, Result);
StoreResult_WithAVXInsert(VectorOpType::SSE, RegClass::FPR, Op, Result);
}
void OpDispatchBuilder::VAESDecOp(OpcodeArgs) {
const auto DstSize = OpSizeFromDst(Op);
const auto Is128Bit = DstSize == OpSize::i128Bit;
// TODO: Handle 256-bit VAESDEC.
LOGMAN_THROW_A_FMT(Is128Bit, "256-bit VAESDEC unimplemented");
const auto Is256Bit = DstSize == OpSize::i256Bit;
Ref State = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
Ref Key = LoadSourceFPR(Op, Op->Src[1], Op->Flags);
Ref Result = _VAESDec(DstSize, State, Key, LoadZeroVector(DstSize));
Ref ZeroVec = LoadZeroVector(DstSize);
Ref Result {};
if (Is256Bit) {
// TODO: Handle as one operation once vixl supports it.
auto UpperState = _VDupElement(DstSize, OpSize::i128Bit, State, 1);
auto UpperKey = _VDupElement(DstSize, OpSize::i128Bit, Key, 1);
auto Lower = _VAESDec(OpSize::i128Bit, State, Key, ZeroVec);
auto Upper = _VAESDec(OpSize::i128Bit, UpperState, UpperKey, ZeroVec);
Result = _VInsElement(DstSize, OpSize::i128Bit, 1, 0, Lower, Upper);
} else {
Result = _VAESDec(DstSize, State, Key, ZeroVec);
}
StoreResultFPR(Op, Result);
}
@@ -273,19 +319,30 @@ void OpDispatchBuilder::AESDecLastOp(OpcodeArgs) {
Ref Dest = LoadSourceFPR(Op, Op->Dest, Op->Flags);
Ref Src = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
Ref Result = _VAESDecLast(OpSize::i128Bit, Dest, Src, LoadZeroVector(OpSize::i128Bit));
StoreResultFPR(Op, Result);
StoreResult_WithAVXInsert(VectorOpType::SSE, RegClass::FPR, Op, Result);
}
void OpDispatchBuilder::VAESDecLastOp(OpcodeArgs) {
const auto DstSize = OpSizeFromDst(Op);
const auto Is128Bit = DstSize == OpSize::i128Bit;
// TODO: Handle 256-bit VAESDECLAST.
LOGMAN_THROW_A_FMT(Is128Bit, "256-bit VAESDECLAST unimplemented");
const auto Is256Bit = DstSize == OpSize::i256Bit;
Ref State = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
Ref Key = LoadSourceFPR(Op, Op->Src[1], Op->Flags);
Ref Result = _VAESDecLast(DstSize, State, Key, LoadZeroVector(DstSize));
Ref ZeroVec = LoadZeroVector(DstSize);
Ref Result {};
if (Is256Bit) {
// TODO: Handle as one operation once vixl supports it.
auto UpperState = _VDupElement(DstSize, OpSize::i128Bit, State, 1);
auto UpperKey = _VDupElement(DstSize, OpSize::i128Bit, Key, 1);
auto Lower = _VAESDecLast(OpSize::i128Bit, State, Key, ZeroVec);
auto Upper = _VAESDecLast(OpSize::i128Bit, UpperState, UpperKey, ZeroVec);
Result = _VInsElement(DstSize, OpSize::i128Bit, 1, 0, Lower, Upper);
} else {
Result = _VAESDecLast(DstSize, State, Key, ZeroVec);
}
StoreResultFPR(Op, Result);
}
@@ -298,14 +355,19 @@ Ref OpDispatchBuilder::AESKeyGenAssistImpl(OpcodeArgs) {
return _VAESKeyGenAssist(Src, KeyGenSwizzle, LoadZeroVector(OpSize::i128Bit), RCON);
}
void OpDispatchBuilder::AESKeyGenAssist(OpcodeArgs) {
void OpDispatchBuilder::AESKeyGenAssist(OpcodeArgs, bool IsAVX) {
if (!CTX->HostFeatures.SupportsAES) {
UnimplementedOp(Op);
return;
}
Ref Result = AESKeyGenAssistImpl(Op);
StoreResultFPR(Op, Result);
if (IsAVX) {
StoreResultFPR(Op, Result);
} else {
StoreResult_WithAVXInsert(VectorOpType::SSE, RegClass::FPR, Op, Result);
}
}
void OpDispatchBuilder::PCLMULQDQOp(OpcodeArgs) {
@@ -317,8 +379,8 @@ void OpDispatchBuilder::PCLMULQDQOp(OpcodeArgs) {
Ref Src = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
const auto Selector = static_cast<uint8_t>(Op->Src[1].Literal());
auto Res = _PCLMUL(OpSize::i128Bit, Dest, Src, Selector & 0b1'0001);
StoreResultFPR(Op, Res);
auto Result = _PCLMUL(OpSize::i128Bit, Dest, Src, Selector & 0b1'0001);
StoreResult_WithAVXInsert(VectorOpType::SSE, RegClass::FPR, Op, Result);
}
void OpDispatchBuilder::VPCLMULQDQOp(OpcodeArgs) {
@@ -5,9 +5,9 @@
namespace FEXCore::IR {
constexpr DispatchTableEntry OpDispatch_DDDTable[] = {
{0x0C, 1, &OpDispatchBuilder::PI2FWOp},
{0x0D, 1, &OpDispatchBuilder::Vector_CVT_Int_To_Float<OpSize::i32Bit, false>},
{0x0D, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Int_To_Float, OpSize::i32Bit, false, false>},
{0x1C, 1, &OpDispatchBuilder::PF2IWOp},
{0x1D, 1, &OpDispatchBuilder::Vector_CVT_Float_To_Int<OpSize::i32Bit, false>},
{0x1D, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Float_To_Int, OpSize::i32Bit, false, false>},
{0x86, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorUnaryOp, IR::OP_VFRECPPRECISION, OpSize::i32Bit>},
{0x87, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::RSqrt3DNowOp, false>},
@@ -15,15 +15,15 @@ constexpr DispatchTableEntry OpDispatch_DDDTable[] = {
{0x8A, 1, &OpDispatchBuilder::PFNACCOp},
{0x8E, 1, &OpDispatchBuilder::PFPNACCOp},
{0x90, 1, &OpDispatchBuilder::VPFCMPOp<1>},
{0x90, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPFCMPOp, 1>},
{0x94, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFMIN, OpSize::i32Bit>},
{0x96, 1, &OpDispatchBuilder::VectorUnaryDuplicateOp<IR::OP_VFRECPPRECISION, OpSize::i32Bit>},
{0x96, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorUnaryDuplicateOp, IR::OP_VFRECPPRECISION, OpSize::i32Bit>},
{0x97, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::RSqrt3DNowOp, true>},
{0x9A, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFSUB, OpSize::i32Bit>},
{0x9E, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFADD, OpSize::i32Bit>},
{0xA0, 1, &OpDispatchBuilder::VPFCMPOp<2>},
{0xA0, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPFCMPOp, 2>},
{0xA4, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFMAX, OpSize::i32Bit>},
// Can be treated as a move
{0xA6, 1, &OpDispatchBuilder::MOVVectorUnalignedOp},
@@ -32,7 +32,7 @@ constexpr DispatchTableEntry OpDispatch_DDDTable[] = {
{0xAA, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUROp, IR::OP_VFSUB, OpSize::i32Bit>},
{0xAE, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFADDP, OpSize::i32Bit>},
{0xB0, 1, &OpDispatchBuilder::VPFCMPOp<0>},
{0xB0, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPFCMPOp, 0>},
{0xB4, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFMUL, OpSize::i32Bit>},
// Can be treated as a move
{0xB6, 1, &OpDispatchBuilder::MOVVectorUnalignedOp},
@@ -20,18 +20,18 @@ constexpr DispatchTableEntry OpDispatch_H0F38Table[] = {
{OPD(PF_38_66, 0x03), 1, &OpDispatchBuilder::PHADDS},
{OPD(PF_38_NONE, 0x04), 1, &OpDispatchBuilder::PMADDUBSW},
{OPD(PF_38_66, 0x04), 1, &OpDispatchBuilder::PMADDUBSW},
{OPD(PF_38_NONE, 0x05), 1, &OpDispatchBuilder::PHSUB<OpSize::i16Bit>},
{OPD(PF_38_66, 0x05), 1, &OpDispatchBuilder::PHSUB<OpSize::i16Bit>},
{OPD(PF_38_NONE, 0x06), 1, &OpDispatchBuilder::PHSUB<OpSize::i32Bit>},
{OPD(PF_38_66, 0x06), 1, &OpDispatchBuilder::PHSUB<OpSize::i32Bit>},
{OPD(PF_38_NONE, 0x05), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PHSUB, OpSize::i16Bit>},
{OPD(PF_38_66, 0x05), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PHSUB, OpSize::i16Bit>},
{OPD(PF_38_NONE, 0x06), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PHSUB, OpSize::i32Bit>},
{OPD(PF_38_66, 0x06), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PHSUB, OpSize::i32Bit>},
{OPD(PF_38_NONE, 0x07), 1, &OpDispatchBuilder::PHSUBS},
{OPD(PF_38_66, 0x07), 1, &OpDispatchBuilder::PHSUBS},
{OPD(PF_38_NONE, 0x08), 1, &OpDispatchBuilder::PSIGN<OpSize::i8Bit>},
{OPD(PF_38_66, 0x08), 1, &OpDispatchBuilder::PSIGN<OpSize::i8Bit>},
{OPD(PF_38_NONE, 0x09), 1, &OpDispatchBuilder::PSIGN<OpSize::i16Bit>},
{OPD(PF_38_66, 0x09), 1, &OpDispatchBuilder::PSIGN<OpSize::i16Bit>},
{OPD(PF_38_NONE, 0x0A), 1, &OpDispatchBuilder::PSIGN<OpSize::i32Bit>},
{OPD(PF_38_66, 0x0A), 1, &OpDispatchBuilder::PSIGN<OpSize::i32Bit>},
{OPD(PF_38_NONE, 0x08), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSIGN, OpSize::i8Bit>},
{OPD(PF_38_66, 0x08), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSIGN, OpSize::i8Bit>},
{OPD(PF_38_NONE, 0x09), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSIGN, OpSize::i16Bit>},
{OPD(PF_38_66, 0x09), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSIGN, OpSize::i16Bit>},
{OPD(PF_38_NONE, 0x0A), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSIGN, OpSize::i32Bit>},
{OPD(PF_38_66, 0x0A), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSIGN, OpSize::i32Bit>},
{OPD(PF_38_NONE, 0x0B), 1, &OpDispatchBuilder::PMULHRSW},
{OPD(PF_38_66, 0x0B), 1, &OpDispatchBuilder::PMULHRSW},
{OPD(PF_38_66, 0x10), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorVariableBlend, OpSize::i8Bit>},
@@ -44,22 +44,22 @@ constexpr DispatchTableEntry OpDispatch_H0F38Table[] = {
{OPD(PF_38_66, 0x1D), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorUnaryOp, IR::OP_VABS, OpSize::i16Bit>},
{OPD(PF_38_NONE, 0x1E), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorUnaryOp, IR::OP_VABS, OpSize::i32Bit>},
{OPD(PF_38_66, 0x1E), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorUnaryOp, IR::OP_VABS, OpSize::i32Bit>},
{OPD(PF_38_66, 0x20), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i8Bit, OpSize::i16Bit, true>},
{OPD(PF_38_66, 0x21), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i8Bit, OpSize::i32Bit, true>},
{OPD(PF_38_66, 0x22), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i8Bit, OpSize::i64Bit, true>},
{OPD(PF_38_66, 0x23), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i16Bit, OpSize::i32Bit, true>},
{OPD(PF_38_66, 0x24), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i16Bit, OpSize::i64Bit, true>},
{OPD(PF_38_66, 0x25), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i32Bit, OpSize::i64Bit, true>},
{OPD(PF_38_66, 0x28), 1, &OpDispatchBuilder::PMULLOp<OpSize::i32Bit, true>},
{OPD(PF_38_66, 0x20), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::ExtendVectorElements, OpSize::i8Bit, OpSize::i16Bit, true>},
{OPD(PF_38_66, 0x21), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::ExtendVectorElements, OpSize::i8Bit, OpSize::i32Bit, true>},
{OPD(PF_38_66, 0x22), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::ExtendVectorElements, OpSize::i8Bit, OpSize::i64Bit, true>},
{OPD(PF_38_66, 0x23), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::ExtendVectorElements, OpSize::i16Bit, OpSize::i32Bit, true>},
{OPD(PF_38_66, 0x24), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::ExtendVectorElements, OpSize::i16Bit, OpSize::i64Bit, true>},
{OPD(PF_38_66, 0x25), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::ExtendVectorElements, OpSize::i32Bit, OpSize::i64Bit, true>},
{OPD(PF_38_66, 0x28), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PMULLOp, OpSize::i32Bit, true>},
{OPD(PF_38_66, 0x29), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPEQ, OpSize::i64Bit>},
{OPD(PF_38_66, 0x2A), 1, &OpDispatchBuilder::MOVVectorNTOp},
{OPD(PF_38_66, 0x2B), 1, &OpDispatchBuilder::PACKUSOp<OpSize::i32Bit>},
{OPD(PF_38_66, 0x30), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i8Bit, OpSize::i16Bit, false>},
{OPD(PF_38_66, 0x31), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i8Bit, OpSize::i32Bit, false>},
{OPD(PF_38_66, 0x32), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i8Bit, OpSize::i64Bit, false>},
{OPD(PF_38_66, 0x33), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i16Bit, OpSize::i32Bit, false>},
{OPD(PF_38_66, 0x34), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i16Bit, OpSize::i64Bit, false>},
{OPD(PF_38_66, 0x35), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i32Bit, OpSize::i64Bit, false>},
{OPD(PF_38_66, 0x2A), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVVectorNTOp, false>},
{OPD(PF_38_66, 0x2B), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PACKUSOp, OpSize::i32Bit>},
{OPD(PF_38_66, 0x30), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::ExtendVectorElements, OpSize::i8Bit, OpSize::i16Bit, false>},
{OPD(PF_38_66, 0x31), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::ExtendVectorElements, OpSize::i8Bit, OpSize::i32Bit, false>},
{OPD(PF_38_66, 0x32), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::ExtendVectorElements, OpSize::i8Bit, OpSize::i64Bit, false>},
{OPD(PF_38_66, 0x33), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::ExtendVectorElements, OpSize::i16Bit, OpSize::i32Bit, false>},
{OPD(PF_38_66, 0x34), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::ExtendVectorElements, OpSize::i16Bit, OpSize::i64Bit, false>},
{OPD(PF_38_66, 0x35), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::ExtendVectorElements, OpSize::i32Bit, OpSize::i64Bit, false>},
{OPD(PF_38_66, 0x37), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPGT, OpSize::i64Bit>},
{OPD(PF_38_66, 0x38), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSMIN, OpSize::i8Bit>},
{OPD(PF_38_66, 0x39), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSMIN, OpSize::i32Bit>},
@@ -79,7 +79,7 @@ constexpr DispatchTableEntry OpDispatch_H0F38Table[] = {
{OPD(PF_38_NONE, 0xCC), 1, &OpDispatchBuilder::SHA256MSG1Op},
{OPD(PF_38_NONE, 0xCD), 1, &OpDispatchBuilder::SHA256MSG2Op},
{OPD(PF_38_66, 0xDB), 1, &OpDispatchBuilder::AESImcOp},
{OPD(PF_38_66, 0xDB), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AESImcOp, false>},
{OPD(PF_38_66, 0xDC), 1, &OpDispatchBuilder::AESEncOp},
{OPD(PF_38_66, 0xDD), 1, &OpDispatchBuilder::AESEncLastOp},
{OPD(PF_38_66, 0xDE), 1, &OpDispatchBuilder::AESDecOp},
@@ -9,13 +9,13 @@ namespace FEXCore::IR {
constexpr auto OpDispatchTableGenH0F3A = []() consteval {
constexpr auto OpDispatchTableGenH0F3AREX = []<uint16_t REX>() consteval {
constexpr DispatchTableEntry Table[] = {
{OPD(REX, PF_3A_66, 0x08), 1, &OpDispatchBuilder::VectorRound<OpSize::i32Bit>},
{OPD(REX, PF_3A_66, 0x09), 1, &OpDispatchBuilder::VectorRound<OpSize::i64Bit>},
{OPD(REX, PF_3A_66, 0x0A), 1, &OpDispatchBuilder::InsertScalarRound<OpSize::i32Bit>},
{OPD(REX, PF_3A_66, 0x0B), 1, &OpDispatchBuilder::InsertScalarRound<OpSize::i64Bit>},
{OPD(REX, PF_3A_66, 0x0C), 1, &OpDispatchBuilder::VectorBlend<OpSize::i32Bit>},
{OPD(REX, PF_3A_66, 0x0D), 1, &OpDispatchBuilder::VectorBlend<OpSize::i64Bit>},
{OPD(REX, PF_3A_66, 0x0E), 1, &OpDispatchBuilder::VectorBlend<OpSize::i16Bit>},
{OPD(REX, PF_3A_66, 0x08), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorRound, OpSize::i32Bit>},
{OPD(REX, PF_3A_66, 0x09), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorRound, OpSize::i64Bit>},
{OPD(REX, PF_3A_66, 0x0A), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::InsertScalarRound, OpSize::i32Bit>},
{OPD(REX, PF_3A_66, 0x0B), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::InsertScalarRound, OpSize::i64Bit>},
{OPD(REX, PF_3A_66, 0x0C), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorBlend, OpSize::i32Bit>},
{OPD(REX, PF_3A_66, 0x0D), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorBlend, OpSize::i64Bit>},
{OPD(REX, PF_3A_66, 0x0E), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorBlend, OpSize::i16Bit>},
{OPD(REX, PF_3A_NONE, 0x0F), 1, &OpDispatchBuilder::PAlignrOp},
{OPD(REX, PF_3A_66, 0x0F), 1, &OpDispatchBuilder::PAlignrOp},
@@ -24,20 +24,20 @@ constexpr auto OpDispatchTableGenH0F3A = []() consteval {
{OPD(REX, PF_3A_66, 0x15), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PExtrOp, OpSize::i16Bit>},
{OPD(REX, PF_3A_66, 0x17), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PExtrOp, OpSize::i32Bit>},
{OPD(REX, PF_3A_66, 0x20), 1, &OpDispatchBuilder::PINSROp<OpSize::i8Bit>},
{OPD(REX, PF_3A_66, 0x20), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PINSROp, OpSize::i8Bit>},
{OPD(REX, PF_3A_66, 0x21), 1, &OpDispatchBuilder::InsertPSOp},
{OPD(REX, PF_3A_66, 0x40), 1, &OpDispatchBuilder::DPPOp<OpSize::i32Bit>},
{OPD(REX, PF_3A_66, 0x41), 1, &OpDispatchBuilder::DPPOp<OpSize::i64Bit>},
{OPD(REX, PF_3A_66, 0x40), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::DPPOp, OpSize::i32Bit>},
{OPD(REX, PF_3A_66, 0x41), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::DPPOp, OpSize::i64Bit>},
{OPD(REX, PF_3A_66, 0x42), 1, &OpDispatchBuilder::MPSADBWOp},
{OPD(REX, PF_3A_66, 0x44), 1, &OpDispatchBuilder::PCLMULQDQOp},
{OPD(REX, PF_3A_66, 0x60), 1, &OpDispatchBuilder::VPCMPESTRMOp},
{OPD(REX, PF_3A_66, 0x61), 1, &OpDispatchBuilder::VPCMPESTRIOp},
{OPD(REX, PF_3A_66, 0x62), 1, &OpDispatchBuilder::VPCMPISTRMOp},
{OPD(REX, PF_3A_66, 0x63), 1, &OpDispatchBuilder::VPCMPISTRIOp},
{OPD(REX, PF_3A_66, 0x60), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPCMPESTRMOp, false>},
{OPD(REX, PF_3A_66, 0x61), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPCMPESTRIOp, false>},
{OPD(REX, PF_3A_66, 0x62), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPCMPISTRMOp, false>},
{OPD(REX, PF_3A_66, 0x63), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPCMPISTRIOp, false>},
{OPD(REX, PF_3A_NONE, 0xCC), 1, &OpDispatchBuilder::SHA1RNDS4Op},
{OPD(REX, PF_3A_66, 0xDF), 1, &OpDispatchBuilder::AESKeyGenAssist},
{OPD(REX, PF_3A_66, 0xDF), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AESKeyGenAssist, false>},
};
return std::to_array(Table);
@@ -65,7 +65,7 @@ constexpr auto OpDispatch_H0F3ATableIgnoreREX = OpDispatchTableGenH0F3A();
constexpr DispatchTableEntry OpDispatch_H0F3ATableNeedsREX0[] = {
{OPD(0, PF_3A_66, 0x16), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PExtrOp, OpSize::i32Bit>},
{OPD(0, PF_3A_66, 0x22), 1, &OpDispatchBuilder::PINSROp<OpSize::i32Bit>},
{OPD(0, PF_3A_66, 0x22), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PINSROp, OpSize::i32Bit>},
};
#undef PF_3A_NONE
@@ -69,12 +69,12 @@ constexpr DispatchTableEntry OpDispatch_SecondaryGroupTables[] = {
// GROUP 9
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_NONE, 1), 1, &OpDispatchBuilder::CMPXCHGPairOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_NONE, 6), 1, &OpDispatchBuilder::RDRANDOp<false>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_NONE, 7), 1, &OpDispatchBuilder::RDRANDOp<true>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_NONE, 6), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::RDRANDOp, false>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_NONE, 7), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::RDRANDOp, true>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_66, 1), 1, &OpDispatchBuilder::CMPXCHGPairOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_66, 6), 1, &OpDispatchBuilder::RDRANDOp<false>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_66, 7), 1, &OpDispatchBuilder::RDRANDOp<true>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_66, 6), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::RDRANDOp, false>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_66, 7), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::RDRANDOp, true>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_F2, 1), 1, &OpDispatchBuilder::CMPXCHGPairOp},
@@ -6,8 +6,7 @@ namespace FEXCore::IR {
constexpr DispatchTableEntry OpDispatch_TwoByteOpTable[] = {
// Instructions
{0x03, 1, &OpDispatchBuilder::LSLOp},
{0x06, 1, &OpDispatchBuilder::PermissionRestrictedOp},
{0x07, 1, &OpDispatchBuilder::PermissionRestrictedOp},
{0x06, 4, &OpDispatchBuilder::PermissionRestrictedOp},
{0x0B, 1, &OpDispatchBuilder::INTOp},
{0x0E, 1, &OpDispatchBuilder::X87EMMS},
@@ -44,7 +43,7 @@ constexpr DispatchTableEntry OpDispatch_TwoByteOpTable[] = {
{0xBE, 2, &OpDispatchBuilder::MOVSXOp},
{0xC0, 2, &OpDispatchBuilder::XADDOp},
{0xC3, 1, &OpDispatchBuilder::MOVGPRNTOp},
{0xC4, 1, &OpDispatchBuilder::PINSROp<OpSize::i16Bit>},
{0xC4, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PINSROp, OpSize::i16Bit>},
{0xC5, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PExtrOp, OpSize::i16Bit>},
{0xC8, 8, &OpDispatchBuilder::BSWAPOp},
@@ -56,10 +55,10 @@ constexpr DispatchTableEntry OpDispatch_TwoByteOpTable[] = {
{0x16, 2, &OpDispatchBuilder::MOVHPDOp},
{0x28, 2, &OpDispatchBuilder::MOVVectorAlignedOp},
{0x2A, 1, &OpDispatchBuilder::InsertMMX_To_XMM_Vector_CVT_Int_To_Float},
{0x2B, 1, &OpDispatchBuilder::MOVVectorNTOp},
{0x2C, 1, &OpDispatchBuilder::XMM_To_MMX_Vector_CVT_Float_To_Int<OpSize::i32Bit, false>},
{0x2D, 1, &OpDispatchBuilder::XMM_To_MMX_Vector_CVT_Float_To_Int<OpSize::i32Bit, true>},
{0x2E, 2, &OpDispatchBuilder::UCOMISxOp<OpSize::i32Bit>},
{0x2B, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVVectorNTOp, false>},
{0x2C, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::XMM_To_MMX_Vector_CVT_Float_To_Int, OpSize::i32Bit, false>},
{0x2D, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::XMM_To_MMX_Vector_CVT_Float_To_Int, OpSize::i32Bit, true>},
{0x2E, 2, &OpDispatchBuilder::Bind<&OpDispatchBuilder::UCOMISxOp, OpSize::i32Bit>},
{0x50, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVMSKOp, OpSize::i32Bit>},
{0x51, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorUnaryOp, IR::OP_VFSQRT, OpSize::i32Bit>},
{0x52, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorUnaryOp, IR::OP_VFRSQRT, OpSize::i32Bit>},
@@ -71,7 +70,7 @@ constexpr DispatchTableEntry OpDispatch_TwoByteOpTable[] = {
{0x58, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFADD, OpSize::i32Bit>},
{0x59, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFMUL, OpSize::i32Bit>},
{0x5A, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Float_To_Float, OpSize::i64Bit, OpSize::i32Bit, false>},
{0x5B, 1, &OpDispatchBuilder::Vector_CVT_Int_To_Float<OpSize::i32Bit, false>},
{0x5B, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Int_To_Float, OpSize::i32Bit, false, false>},
{0x5C, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFSUB, OpSize::i32Bit>},
{0x5D, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFMIN, OpSize::i32Bit>},
{0x5E, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFDIV, OpSize::i32Bit>},
@@ -79,15 +78,15 @@ constexpr DispatchTableEntry OpDispatch_TwoByteOpTable[] = {
{0x60, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKLOp, OpSize::i8Bit>},
{0x61, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKLOp, OpSize::i16Bit>},
{0x62, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKLOp, OpSize::i32Bit>},
{0x63, 1, &OpDispatchBuilder::PACKSSOp<OpSize::i16Bit>},
{0x63, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PACKSSOp, OpSize::i16Bit>},
{0x64, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPGT, OpSize::i8Bit>},
{0x65, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPGT, OpSize::i16Bit>},
{0x66, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPGT, OpSize::i32Bit>},
{0x67, 1, &OpDispatchBuilder::PACKUSOp<OpSize::i16Bit>},
{0x67, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PACKUSOp, OpSize::i16Bit>},
{0x68, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKHOp, OpSize::i8Bit>},
{0x69, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKHOp, OpSize::i16Bit>},
{0x6A, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKHOp, OpSize::i32Bit>},
{0x6B, 1, &OpDispatchBuilder::PACKSSOp<OpSize::i32Bit>},
{0x6B, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PACKSSOp, OpSize::i32Bit>},
{0x70, 1, &OpDispatchBuilder::PSHUFW8ByteOp},
{0x74, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPEQ, OpSize::i8Bit>},
@@ -95,7 +94,7 @@ constexpr DispatchTableEntry OpDispatch_TwoByteOpTable[] = {
{0x76, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPEQ, OpSize::i32Bit>},
{0x77, 1, &OpDispatchBuilder::X87EMMS},
{0xC2, 1, &OpDispatchBuilder::VFCMPOp<OpSize::i32Bit>},
{0xC2, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VFCMPOp, OpSize::i32Bit>},
{0xC6, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::SHUFOp, OpSize::i32Bit>},
{0xD1, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRLDOp, OpSize::i16Bit>},
@@ -116,9 +115,9 @@ constexpr DispatchTableEntry OpDispatch_TwoByteOpTable[] = {
{0xE1, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRAOp, OpSize::i16Bit>},
{0xE2, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRAOp, OpSize::i32Bit>},
{0xE3, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VURAVG, OpSize::i16Bit>},
{0xE4, 1, &OpDispatchBuilder::PMULHW<false>},
{0xE5, 1, &OpDispatchBuilder::PMULHW<true>},
{0xE7, 1, &OpDispatchBuilder::MOVVectorNTOp},
{0xE4, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PMULHW, false>},
{0xE5, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PMULHW, true>},
{0xE7, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVVectorNTOp, false>},
{0xE8, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSQSUB, OpSize::i8Bit>},
{0xE9, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSQSUB, OpSize::i16Bit>},
{0xEA, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSMIN, OpSize::i16Bit>},
@@ -131,7 +130,7 @@ constexpr DispatchTableEntry OpDispatch_TwoByteOpTable[] = {
{0xF1, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSLL, OpSize::i16Bit>},
{0xF2, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSLL, OpSize::i32Bit>},
{0xF3, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSLL, OpSize::i64Bit>},
{0xF4, 1, &OpDispatchBuilder::PMULLOp<OpSize::i32Bit, false>},
{0xF4, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PMULLOp, OpSize::i32Bit, false>},
{0xF5, 1, &OpDispatchBuilder::PMADDWD},
{0xF6, 1, &OpDispatchBuilder::PSADBW},
{0xF7, 1, &OpDispatchBuilder::MASKMOVOp},
@@ -152,23 +151,23 @@ constexpr DispatchTableEntry OpDispatch_TwoByteOpTable[] = {
constexpr DispatchTableEntry OpDispatch_SecondaryRepModTables[] = {
{0x10, 2, &OpDispatchBuilder::MOVSSOp},
{0x12, 1, &OpDispatchBuilder::VMOVSLDUPOp},
{0x16, 1, &OpDispatchBuilder::VMOVSHDUPOp},
{0x2A, 1, &OpDispatchBuilder::InsertCVTGPR_To_FPR<OpSize::i32Bit>},
{0x2B, 1, &OpDispatchBuilder::MOVVectorNTOp},
{0x2C, 1, &OpDispatchBuilder::CVTFPR_To_GPR<OpSize::i32Bit, false>},
{0x2D, 1, &OpDispatchBuilder::CVTFPR_To_GPR<OpSize::i32Bit, true>},
{0x51, 1, &OpDispatchBuilder::VectorScalarUnaryInsertALUOp<IR::OP_VFSQRTSCALARINSERT, OpSize::i32Bit>},
{0x52, 1, &OpDispatchBuilder::VectorScalarUnaryInsertALUOp<IR::OP_VFRSQRTSCALARINSERT, OpSize::i32Bit>},
{0x53, 1, &OpDispatchBuilder::VectorScalarUnaryInsertALUOp<IR::OP_VFRECPSCALARINSERT, OpSize::i32Bit>},
{0x58, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFADDSCALARINSERT, OpSize::i32Bit>},
{0x59, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFMULSCALARINSERT, OpSize::i32Bit>},
{0x5A, 1, &OpDispatchBuilder::InsertScalar_CVT_Float_To_Float<OpSize::i64Bit, OpSize::i32Bit>},
{0x5B, 1, &OpDispatchBuilder::Vector_CVT_Float_To_Int<OpSize::i32Bit, false>},
{0x5C, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFSUBSCALARINSERT, OpSize::i32Bit>},
{0x5D, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFMINSCALARINSERT, OpSize::i32Bit>},
{0x5E, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFDIVSCALARINSERT, OpSize::i32Bit>},
{0x5F, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFMAXSCALARINSERT, OpSize::i32Bit>},
{0x12, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VMOVSLDUPOp, false>},
{0x16, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VMOVSHDUPOp, false>},
{0x2A, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::InsertCVTGPR_To_FPR, OpSize::i32Bit>},
{0x2B, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVVectorNTOp, false>},
{0x2C, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::CVTFPR_To_GPR, OpSize::i32Bit, false>},
{0x2D, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::CVTFPR_To_GPR, OpSize::i32Bit, true>},
{0x51, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorScalarUnaryInsertALUOp, IR::OP_VFSQRTSCALARINSERT, OpSize::i32Bit>},
{0x52, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorScalarUnaryInsertALUOp, IR::OP_VFRSQRTSCALARINSERT, OpSize::i32Bit>},
{0x53, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorScalarUnaryInsertALUOp, IR::OP_VFRECPSCALARINSERT, OpSize::i32Bit>},
{0x58, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorScalarInsertALUOp, IR::OP_VFADDSCALARINSERT, OpSize::i32Bit>},
{0x59, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorScalarInsertALUOp, IR::OP_VFMULSCALARINSERT, OpSize::i32Bit>},
{0x5A, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::InsertScalar_CVT_Float_To_Float, OpSize::i64Bit, OpSize::i32Bit>},
{0x5B, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Float_To_Int, OpSize::i32Bit, false, false>},
{0x5C, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorScalarInsertALUOp, IR::OP_VFSUBSCALARINSERT, OpSize::i32Bit>},
{0x5D, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorScalarInsertALUOp, IR::OP_VFMINSCALARINSERT, OpSize::i32Bit>},
{0x5E, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorScalarInsertALUOp, IR::OP_VFDIVSCALARINSERT, OpSize::i32Bit>},
{0x5F, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorScalarInsertALUOp, IR::OP_VFMAXSCALARINSERT, OpSize::i32Bit>},
{0x6F, 1, &OpDispatchBuilder::MOVVectorUnalignedOp},
{0x70, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSHUFWOp, false>},
{0x7E, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVQOp, OpDispatchBuilder::VectorOpType::SSE>},
@@ -176,36 +175,36 @@ constexpr DispatchTableEntry OpDispatch_SecondaryRepModTables[] = {
{0xB8, 1, &OpDispatchBuilder::PopcountOp},
{0xBC, 1, &OpDispatchBuilder::TZCNT},
{0xBD, 1, &OpDispatchBuilder::LZCNT},
{0xC2, 1, &OpDispatchBuilder::InsertScalarFCMPOp<OpSize::i32Bit>},
{0xD6, 1, &OpDispatchBuilder::MOVQ2DQ<true>},
{0xE6, 1, &OpDispatchBuilder::Vector_CVT_Int_To_Float<OpSize::i32Bit, true>},
{0xC2, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::InsertScalarFCMPOp, OpSize::i32Bit>},
{0xD6, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVQ2DQ, true>},
{0xE6, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Int_To_Float, OpSize::i32Bit, true, false>},
};
constexpr DispatchTableEntry OpDispatch_SecondaryRepNEModTables[] = {
{0x10, 2, &OpDispatchBuilder::MOVSDOp},
{0x12, 1, &OpDispatchBuilder::MOVDDUPOp},
{0x2A, 1, &OpDispatchBuilder::InsertCVTGPR_To_FPR<OpSize::i64Bit>},
{0x2B, 1, &OpDispatchBuilder::MOVVectorNTOp},
{0x2C, 1, &OpDispatchBuilder::CVTFPR_To_GPR<OpSize::i64Bit, false>},
{0x2D, 1, &OpDispatchBuilder::CVTFPR_To_GPR<OpSize::i64Bit, true>},
{0x51, 1, &OpDispatchBuilder::VectorScalarUnaryInsertALUOp<IR::OP_VFSQRTSCALARINSERT, OpSize::i64Bit>},
{0x2A, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::InsertCVTGPR_To_FPR, OpSize::i64Bit>},
{0x2B, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVVectorNTOp, false>},
{0x2C, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::CVTFPR_To_GPR, OpSize::i64Bit, false>},
{0x2D, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::CVTFPR_To_GPR, OpSize::i64Bit, true>},
{0x51, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorScalarUnaryInsertALUOp, IR::OP_VFSQRTSCALARINSERT, OpSize::i64Bit>},
// x52 = Invalid
{0x58, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFADDSCALARINSERT, OpSize::i64Bit>},
{0x59, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFMULSCALARINSERT, OpSize::i64Bit>},
{0x5A, 1, &OpDispatchBuilder::InsertScalar_CVT_Float_To_Float<OpSize::i32Bit, OpSize::i64Bit>},
{0x5C, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFSUBSCALARINSERT, OpSize::i64Bit>},
{0x5D, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFMINSCALARINSERT, OpSize::i64Bit>},
{0x5E, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFDIVSCALARINSERT, OpSize::i64Bit>},
{0x5F, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFMAXSCALARINSERT, OpSize::i64Bit>},
{0x58, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorScalarInsertALUOp, IR::OP_VFADDSCALARINSERT, OpSize::i64Bit>},
{0x59, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorScalarInsertALUOp, IR::OP_VFMULSCALARINSERT, OpSize::i64Bit>},
{0x5A, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::InsertScalar_CVT_Float_To_Float, OpSize::i32Bit, OpSize::i64Bit>},
{0x5C, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorScalarInsertALUOp, IR::OP_VFSUBSCALARINSERT, OpSize::i64Bit>},
{0x5D, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorScalarInsertALUOp, IR::OP_VFMINSCALARINSERT, OpSize::i64Bit>},
{0x5E, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorScalarInsertALUOp, IR::OP_VFDIVSCALARINSERT, OpSize::i64Bit>},
{0x5F, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorScalarInsertALUOp, IR::OP_VFMAXSCALARINSERT, OpSize::i64Bit>},
{0x70, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSHUFWOp, true>},
{0x78, 1, &OpDispatchBuilder::Insertq_imm},
{0x79, 1, &OpDispatchBuilder::Insertq},
{0x7C, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFADDP, OpSize::i32Bit>},
{0x7D, 1, &OpDispatchBuilder::HSUBP<OpSize::i32Bit>},
{0xD0, 1, &OpDispatchBuilder::ADDSUBPOp<OpSize::i32Bit>},
{0xD6, 1, &OpDispatchBuilder::MOVQ2DQ<false>},
{0xC2, 1, &OpDispatchBuilder::InsertScalarFCMPOp<OpSize::i64Bit>},
{0xE6, 1, &OpDispatchBuilder::Vector_CVT_Float_To_Int<OpSize::i64Bit, true>},
{0x7D, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::HSUBP, OpSize::i32Bit>},
{0xD0, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::ADDSUBPOp, OpSize::i32Bit>},
{0xD6, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVQ2DQ, false>},
{0xC2, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::InsertScalarFCMPOp, OpSize::i64Bit>},
{0xE6, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Float_To_Int, OpSize::i64Bit, true, false>},
{0xF0, 1, &OpDispatchBuilder::MOVVectorUnalignedOp},
};
@@ -217,10 +216,10 @@ constexpr DispatchTableEntry OpDispatch_SecondaryOpSizeModTables[] = {
{0x16, 2, &OpDispatchBuilder::MOVHPDOp},
{0x28, 2, &OpDispatchBuilder::MOVVectorAlignedOp},
{0x2A, 1, &OpDispatchBuilder::MMX_To_XMM_Vector_CVT_Int_To_Float},
{0x2B, 1, &OpDispatchBuilder::MOVVectorNTOp},
{0x2C, 1, &OpDispatchBuilder::XMM_To_MMX_Vector_CVT_Float_To_Int<OpSize::i64Bit, false>},
{0x2D, 1, &OpDispatchBuilder::XMM_To_MMX_Vector_CVT_Float_To_Int<OpSize::i64Bit, true>},
{0x2E, 2, &OpDispatchBuilder::UCOMISxOp<OpSize::i64Bit>},
{0x2B, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVVectorNTOp, false>},
{0x2C, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::XMM_To_MMX_Vector_CVT_Float_To_Int, OpSize::i64Bit, false>},
{0x2D, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::XMM_To_MMX_Vector_CVT_Float_To_Int, OpSize::i64Bit, true>},
{0x2E, 2, &OpDispatchBuilder::Bind<&OpDispatchBuilder::UCOMISxOp, OpSize::i64Bit>},
{0x50, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVMSKOp, OpSize::i64Bit>},
{0x51, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorUnaryOp, IR::OP_VFSQRT, OpSize::i64Bit>},
@@ -231,7 +230,7 @@ constexpr DispatchTableEntry OpDispatch_SecondaryOpSizeModTables[] = {
{0x58, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFADD, OpSize::i64Bit>},
{0x59, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFMUL, OpSize::i64Bit>},
{0x5A, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Float_To_Float, OpSize::i32Bit, OpSize::i64Bit, false>},
{0x5B, 1, &OpDispatchBuilder::Vector_CVT_Float_To_Int<OpSize::i32Bit, true>},
{0x5B, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Float_To_Int, OpSize::i32Bit, true, false>},
{0x5C, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFSUB, OpSize::i64Bit>},
{0x5D, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFMIN, OpSize::i64Bit>},
{0x5E, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFDIV, OpSize::i64Bit>},
@@ -239,15 +238,15 @@ constexpr DispatchTableEntry OpDispatch_SecondaryOpSizeModTables[] = {
{0x60, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKLOp, OpSize::i8Bit>},
{0x61, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKLOp, OpSize::i16Bit>},
{0x62, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKLOp, OpSize::i32Bit>},
{0x63, 1, &OpDispatchBuilder::PACKSSOp<OpSize::i16Bit>},
{0x63, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PACKSSOp, OpSize::i16Bit>},
{0x64, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPGT, OpSize::i8Bit>},
{0x65, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPGT, OpSize::i16Bit>},
{0x66, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPGT, OpSize::i32Bit>},
{0x67, 1, &OpDispatchBuilder::PACKUSOp<OpSize::i16Bit>},
{0x67, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PACKUSOp, OpSize::i16Bit>},
{0x68, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKHOp, OpSize::i8Bit>},
{0x69, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKHOp, OpSize::i16Bit>},
{0x6A, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKHOp, OpSize::i32Bit>},
{0x6B, 1, &OpDispatchBuilder::PACKSSOp<OpSize::i32Bit>},
{0x6B, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PACKSSOp, OpSize::i32Bit>},
{0x6C, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKLOp, OpSize::i64Bit>},
{0x6D, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUNPCKHOp, OpSize::i64Bit>},
{0x6E, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVBetweenGPR_FPR, OpDispatchBuilder::VectorOpType::SSE>},
@@ -260,15 +259,15 @@ constexpr DispatchTableEntry OpDispatch_SecondaryOpSizeModTables[] = {
{0x78, 1, nullptr}, // GROUP 17
{0x79, 1, &OpDispatchBuilder::Extrq},
{0x7C, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFADDP, OpSize::i64Bit>},
{0x7D, 1, &OpDispatchBuilder::HSUBP<OpSize::i64Bit>},
{0x7D, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::HSUBP, OpSize::i64Bit>},
{0x7E, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVBetweenGPR_FPR, OpDispatchBuilder::VectorOpType::SSE>},
{0x7F, 1, &OpDispatchBuilder::MOVVectorAlignedOp},
{0xC2, 1, &OpDispatchBuilder::VFCMPOp<OpSize::i64Bit>},
{0xC4, 1, &OpDispatchBuilder::PINSROp<OpSize::i16Bit>},
{0xC2, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VFCMPOp, OpSize::i64Bit>},
{0xC4, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PINSROp, OpSize::i16Bit>},
{0xC5, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PExtrOp, OpSize::i16Bit>},
{0xC6, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::SHUFOp, OpSize::i64Bit>},
{0xD0, 1, &OpDispatchBuilder::ADDSUBPOp<OpSize::i64Bit>},
{0xD0, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::ADDSUBPOp, OpSize::i64Bit>},
{0xD1, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRLDOp, OpSize::i16Bit>},
{0xD2, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRLDOp, OpSize::i32Bit>},
{0xD3, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRLDOp, OpSize::i64Bit>},
@@ -288,10 +287,10 @@ constexpr DispatchTableEntry OpDispatch_SecondaryOpSizeModTables[] = {
{0xE1, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRAOp, OpSize::i16Bit>},
{0xE2, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSRAOp, OpSize::i32Bit>},
{0xE3, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VURAVG, OpSize::i16Bit>},
{0xE4, 1, &OpDispatchBuilder::PMULHW<false>},
{0xE5, 1, &OpDispatchBuilder::PMULHW<true>},
{0xE6, 1, &OpDispatchBuilder::Vector_CVT_Float_To_Int<OpSize::i64Bit, false>},
{0xE7, 1, &OpDispatchBuilder::MOVVectorNTOp},
{0xE4, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PMULHW, false>},
{0xE5, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PMULHW, true>},
{0xE6, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Float_To_Int, OpSize::i64Bit, false, false>},
{0xE7, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVVectorNTOp, false>},
{0xE8, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSQSUB, OpSize::i8Bit>},
{0xE9, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSQSUB, OpSize::i16Bit>},
{0xEA, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VSMIN, OpSize::i16Bit>},
@@ -304,7 +303,7 @@ constexpr DispatchTableEntry OpDispatch_SecondaryOpSizeModTables[] = {
{0xF1, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSLL, OpSize::i16Bit>},
{0xF2, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSLL, OpSize::i32Bit>},
{0xF3, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSLL, OpSize::i64Bit>},
{0xF4, 1, &OpDispatchBuilder::PMULLOp<OpSize::i32Bit, false>},
{0xF4, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PMULLOp, OpSize::i32Bit, false>},
{0xF5, 1, &OpDispatchBuilder::PMADDWD},
{0xF6, 1, &OpDispatchBuilder::PSADBW},
{0xF7, 1, &OpDispatchBuilder::MASKMOVOp},
File diff suppressed because it is too large. Load diff
@@ -163,7 +163,7 @@ void OpDispatchBuilder::FIST(OpcodeArgs, bool Truncate) {
// Check for NaN/Infinity: exponent = 0x7fff
SaveNZCV();
_TestNZ(OpSize::i64Bit, Exponent, Constant(0x7fff));
SubWithFlags(OpSize::i64Bit, Exponent, 0x7fff);
Ref IsSpecial = _NZCVSelect01(CondClass::EQ);
// For overflow detection, check if exponent indicates a value >= 2^15
@@ -178,7 +178,8 @@ void OpDispatchBuilder::FIST(OpcodeArgs, bool Truncate) {
Data = _F80CVTInt(Size, Data, Truncate);
StoreResultGPR_WithOpSize(Op, Op->Dest, Data, Size, OpSize::i8Bit);
StoreResultGPR_WithOpSize(Op, Op->Dest, Data, Size, OpSize::i8Bit,
CTX->IsVectorAtomicTSOEnabled() ? MemoryAccessType::DEFAULT : MemoryAccessType::NONTSO);
if ((Op->TableInfo->Flags & X86Tables::InstFlags::FLAGS_POP) != 0) {
_PopStackDestroy();
@@ -574,7 +575,7 @@ void OpDispatchBuilder::X87FRSTOR(OpcodeArgs) {
for (int i = 0; i < 7; ++i) {
Ref Reg = _LoadMemFPR(OpSize::i128Bit, Mem, Constant((IR::OpSizeToSize(Size) * 7) + (10 * i)), OpSize::i8Bit, MemOffsetType::SXTX, 1);
// Mask off the top bits
Reg = _VAnd(OpSize::i128Bit, OpSize::i128Bit, Reg, Mask);
Reg = _VAnd(OpSize::i128Bit, Reg, Mask);
if (ReducedPrecisionMode) {
// Convert to double precision
Reg = _F80CVT(OpSize::i64Bit, Reg);
@@ -623,13 +624,10 @@ void OpDispatchBuilder::FXCH(OpcodeArgs) {
void OpDispatchBuilder::X87FYL2X(OpcodeArgs, bool IsFYL2XP1) {
if (IsFYL2XP1) {
// create an add between top of stack and 1.
Ref One = ReducedPrecisionMode ? _VCastFromGPR(OpSize::i64Bit, OpSize::i64Bit, Constant(0x3FF0000000000000)) :
LoadAndCacheNamedVectorConstant(OpSize::i128Bit, NamedVectorConstant::NAMED_VECTOR_X87_ONE);
_F80AddValue(0, One);
_F80FYL2XP1Stack();
} else {
_F80FYL2XStack();
}
_F80FYL2XStack();
}
void OpDispatchBuilder::FCOMI(OpcodeArgs, IR::OpSize Width, bool Integer, OpDispatchBuilder::FCOMIFlags WhichFlags, bool PopTwice) {
@@ -106,12 +106,24 @@ void OpDispatchBuilder::FISTF64(OpcodeArgs, bool Truncate) {
const auto Size = OpSizeFromSrc(Op);
Ref data = _ReadStackValue(0);
if (Truncate) {
data = _Float_ToGPR_ZS(Size == OpSize::i32Bit ? OpSize::i32Bit : OpSize::i64Bit, OpSize::i64Bit, data);
bool CanUseFloatReg = Size == OpSize::i64Bit;
if (CanUseFloatReg) {
// If possible, it's faster to keep the data in an FPR than doing a GPR transfer.
if (Truncate) {
data = _Vector_FToZS(OpSize::i128Bit, OpSize::i64Bit, data);
} else {
data = _Vector_FToS(OpSize::i128Bit, OpSize::i64Bit, data);
}
StoreResultFPR_WithOpSize(Op, Op->Dest, data, OpSize::i64Bit, OpSize::i8Bit);
} else {
data = _Float_ToGPR_S(Size == OpSize::i32Bit ? OpSize::i32Bit : OpSize::i64Bit, OpSize::i64Bit, data);
if (Truncate) {
data = _Float_ToGPR_ZS(Size == OpSize::i32Bit ? OpSize::i32Bit : OpSize::i64Bit, OpSize::i64Bit, data);
} else {
data = _Float_ToGPR_S(Size == OpSize::i32Bit ? OpSize::i32Bit : OpSize::i64Bit, OpSize::i64Bit, data);
}
StoreResultGPR_WithOpSize(Op, Op->Dest, data, Size, OpSize::i8Bit,
CTX->IsVectorAtomicTSOEnabled() ? MemoryAccessType::DEFAULT : MemoryAccessType::NONTSO);
}
StoreResultGPR_WithOpSize(Op, Op->Dest, data, Size, OpSize::i8Bit);
if ((Op->TableInfo->Flags & X86Tables::InstFlags::FLAGS_POP) != 0) {
_PopStackDestroy();
@@ -370,6 +382,8 @@ void OpDispatchBuilder::X87FXTRACTF64(OpcodeArgs) {
// Split node into SIG and EXP while handling the special zero case.
// i.e. if val == 0.0, then sig = 0.0, exp = -inf
// if val == -0.0, then sig = -0.0, exp = -inf
// if val is +/-Inf, then sig = val, exp = +inf
// if val is NaN, then sig = val, exp = val
// otherwise we just extract the 64-bit sig and exp as normal.
Ref Node = _ReadStackValue(0);
@@ -379,6 +393,11 @@ void OpDispatchBuilder::X87FXTRACTF64(OpcodeArgs) {
Ref ExpZV = _VCastFromGPR(OpSize::i64Bit, OpSize::i64Bit, Constant(0xfff0'0000'0000'0000UL));
Ref SigZV = Node;
// Inf/NaN case
Ref ExpInfOnlyV = _VCastFromGPR(OpSize::i64Bit, OpSize::i64Bit, Constant(0x7ff0'0000'0000'0000UL));
Ref ExpNanV = Node;
Ref SigInfV = Node;
// non zero case
Ref ExpNZ = _Bfe(OpSize::i64Bit, 11, 52, Gpr);
ExpNZ = Sub(OpSize::i64Bit, ExpNZ, Constant(1023));
@@ -388,12 +407,24 @@ void OpDispatchBuilder::X87FXTRACTF64(OpcodeArgs) {
SigNZ = _Or(OpSize::i64Bit, SigNZ, Constant(0x3ff0'0000'0000'0000LL));
Ref SigNZV = _VCastFromGPR(OpSize::i64Bit, OpSize::i64Bit, SigNZ);
// Comparison and select to push onto stack
SaveNZCV();
// Mantissa non-zero => NaN (exp result = input); else Inf (exp result = +Inf)
Ref Mantissa = _And(OpSize::i64Bit, Gpr, Constant(0x000f'ffff'ffff'ffffULL));
_TestNZ(OpSize::i64Bit, Mantissa, Constant(~0ULL));
Ref ExpInfV = _NZCVSelectV(OpSize::i64Bit, CondClass::EQ, ExpInfOnlyV, ExpNanV);
// Biased exponent == 0x7ff => Inf/NaN path, else non-zero-case.
Ref BiasedExp = _Bfe(OpSize::i64Bit, 11, 52, Gpr);
SubWithFlags(OpSize::i64Bit, BiasedExp, 0x7ff);
Ref ExpNZOrInf = _NZCVSelectV(OpSize::i64Bit, CondClass::EQ, ExpInfV, ExpNZV);
Ref SigNZOrInf = _NZCVSelectV(OpSize::i64Bit, CondClass::EQ, SigInfV, SigNZV);
// Zero folds on top.
_TestNZ(OpSize::i64Bit, Gpr, Constant(0x7fff'ffff'ffff'ffffUL));
Ref Sig = _NZCVSelectV(OpSize::i64Bit, CondClass::EQ, SigZV, SigNZV);
Ref Exp = _NZCVSelectV(OpSize::i64Bit, CondClass::EQ, ExpZV, ExpNZV);
Ref Sig = _NZCVSelectV(OpSize::i64Bit, CondClass::EQ, SigZV, SigNZOrInf);
Ref Exp = _NZCVSelectV(OpSize::i64Bit, CondClass::EQ, ExpZV, ExpNZOrInf);
_PopStackDestroy();
_PushStack(Exp, Invalid(), OpSize::iInvalid);
@@ -0,0 +1,97 @@
// SPDX-License-Identifier: MIT
#include "Interface/Core/LookupCache.h"
#include "Interface/Core/SharedCodeBufferManager.h"
#include <FEXCore/fextl/memory.h>
#include <FEXCore/Utils/AllocatorHooks.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/MathUtils.h>
#ifndef _WIN32
#include <FEXCore/Utils/PrctlUtils.h>
#endif
namespace FEXCore::CPU {
static constexpr size_t INITIAL_CODE_SIZE = 1024 * 1024 * 16;
// We don't want to move above 128MB atm because that means we will have to encode longer jumps
static constexpr size_t MAX_CODE_SIZE = 1024 * 1024 * 128;
CodeBuffer::CodeBuffer(size_t Size)
: AllocatedSize(Size) {
Ptr = static_cast<uint8_t*>(FEXCore::Allocator::VirtualAlloc(Size, true));
LOGMAN_THROW_A_FMT(!!Ptr, "Couldn't allocate code buffer");
// Protect the last page of the allocated buffer to trigger SIGSEGV on write access
uintptr_t LastPageAddr = AlignDown(reinterpret_cast<uintptr_t>(Ptr) + Size - 1, FEXCore::Utils::FEX_PAGE_SIZE);
if (!FEXCore::Allocator::VirtualProtect(reinterpret_cast<void*>(LastPageAddr), FEXCore::Utils::FEX_PAGE_SIZE,
FEXCore::Allocator::ProtectOptions::None)) {
LogMan::Msg::EFmt("Failed to mprotect last page of code buffer.");
}
FEXCore::Allocator::VirtualName("FEXMemJIT", Ptr, Size);
// Huge-pages reduce the amount of iTLB misses dramatically when it works.
FEXCore::Allocator::VirtualTHPControl(Ptr, Size, FEXCore::Allocator::THPControl::Enable);
LookupCache = fextl::make_unique<GuestToHostMap>();
}
CodeBuffer::~CodeBuffer() {
FEXCore::Allocator::VirtualFree(Ptr, AllocatedSize);
}
fextl::shared_ptr<CodeBuffer> SharedCodeBufferManager::AllocateNew(size_t Size) {
#ifndef _WIN32
// MDWE (Memory-Deny-Write-Execute) is a new Linux 6.3 feature.
// It's equivalent to systemd's `MemoryDenyWriteExecute` but implemented entirely in the kernel.
//
// MDWE prevents applications from creating RWX memory mappings.
// This prevents FEX from doing anything JIT related, as FEX uses RWX for JIT memory mappings.
//
// A potential workaround to make FEX work with MDWE is to call mprotect every time we need to write or modify code.
// Alternatively, FEX could use a memory mirror where one half is mapped as RW and the other is RX.
//
// Once MDWE is enabled with the prctl, the feature is sealed and it can /NOT/ be turned off.
//
// Status of MDWE is queried through prctl using `PR_GET_MDWE`:
// -1: The kernel doesn't support MDWE
// 0: MDWE is supported but disabled
// >0: MDWE is enabled, hence prohibiting RWX mappings
int MDWE = ::prctl(PR_GET_MDWE, 0, 0, 0, 0);
if (MDWE != -1 && MDWE != 0) {
LogMan::Msg::EFmt("MDWE was set to 0x{:x} which means FEX can't allocate executable memory", MDWE);
}
#endif
auto Buffer = fextl::make_shared<CodeBuffer>(Size);
Latest = Buffer;
LatestOffset = 0;
OnCodeBufferAllocated(Buffer);
return Buffer;
}
fextl::shared_ptr<CodeBuffer> SharedCodeBufferManager::GetLatest() {
if (!Latest) {
AllocateNew(INITIAL_CODE_SIZE);
}
return Latest;
}
fextl::shared_ptr<CodeBuffer> SharedCodeBufferManager::StartLargerCodeBuffer() {
if (!Latest) {
// Allocate initial CodeBuffer and return it
return GetLatest();
}
auto NewCodeBufferSize = GetLatest()->AllocatedSize;
NewCodeBufferSize = std::min<size_t>(NewCodeBufferSize * 2, MAX_CODE_SIZE);
return AllocateNew(NewCodeBufferSize);
}
fextl::shared_ptr<CodeBuffer> SharedCodeBufferManager::StartMaximalCodeBuffer() {
return AllocateNew(MAX_CODE_SIZE);
}
} // namespace FEXCore::CPU
@@ -0,0 +1,78 @@
// SPDX-License-Identifier: MIT
/*
$info$
category: Thread shared code buffer management
tags: backend|shared
$end_info$
*/
#pragma once
#include <FEXCore/fextl/memory.h>
#include <FEXCore/Utils/SignalScopeGuards.h>
#include <FEXCore/Utils/TypeDefines.h>
#include <cstddef>
#include <cstdint>
namespace FEXCore {
struct GuestToHostMap;
}
namespace FEXCore::CPU {
struct CodeBuffer {
uint8_t* Ptr;
size_t AllocatedSize; // including guard page; see UsableSize()
fextl::unique_ptr<GuestToHostMap> LookupCache;
CodeBuffer(size_t Size);
CodeBuffer(const CodeBuffer&) = delete;
CodeBuffer& operator=(const CodeBuffer&) = delete;
CodeBuffer(CodeBuffer&& oth) = delete;
CodeBuffer& operator=(CodeBuffer&&) = delete;
~CodeBuffer();
/// Returns the number of bytes available for storing code
size_t UsableSize() const {
return AllocatedSize - FEXCore::Utils::FEX_PAGE_SIZE;
}
};
/**
* A manager that coordinates access to the CodeBuffer used for compiling new code across threads.
*
* The CodeBuffer is managed as a partially persistent data structure:
* - Exactly one CodeBuffer is now designated as "active", which means data can be appended to it
* - Lossy modifications to the active CodeBuffer will not invalidate any data in use by other threads (which is what enables save CodeBuffer sharing across threads)
* - Instead, such lossy modifications trigger a new "version" of the data in the modifying thread. Old versions of the CodeBuffer persist as read-only data for use by the other threads.
* - The other threads can update their version of the CodeBuffer. This will decrease the reference count and eventually trigger deallocation of the old version
*/
class SharedCodeBufferManager {
public:
// Get the CodeBuffer that was most recently allocated.
// This is the only CodeBuffer that data may be written to.
fextl::shared_ptr<CodeBuffer> GetLatest();
// Allocate a new CodeBuffer with geometric growth up to an internal maximum.
// Subsequent calls to GetLatest will point to the returned buffer.
fextl::shared_ptr<CodeBuffer> StartLargerCodeBuffer();
// Allocate a new CodeBuffer with maximum internal size.
// Subsequent calls to GetLatest will point to the returned buffer.
fextl::shared_ptr<CodeBuffer> StartMaximalCodeBuffer();
// Write offset into the latest CodeBuffer
std::size_t LatestOffset {};
// Protects writes to the latest CodeBuffer and changes to LatestOffset
FEXCore::ForkableUniqueMutex CodeBufferWriteMutex;
virtual void OnCodeBufferAllocated(const std::shared_ptr<CodeBuffer>&) {};
private:
fextl::shared_ptr<CodeBuffer> Latest;
fextl::shared_ptr<CodeBuffer> AllocateNew(size_t Size);
};
} // namespace FEXCore::CPU
@@ -200,8 +200,8 @@ const std::array<X86InstInfo, MAX_PRIMARY_TABLE_SIZE> BaseOps = []() consteval {
{0xF3, 1, X86InstInfo{"REP", TYPE_PREFIX, FLAGS_NONE, 0}},
// Instructions
{0x00, 1, X86InstInfo{"ADD", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x01, 1, X86InstInfo{"ADD", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_DISPLACE_SIZE_DIV_2, 0}},
{0x00, 1, X86InstInfo{"ADD", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{0x01, 1, X86InstInfo{"ADD", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_SUPPORTS_LOCK, 0}},
{0x02, 1, X86InstInfo{"ADD", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0}},
{0x03, 1, X86InstInfo{"ADD", TYPE_INST, FLAGS_MODRM, 0}},
{0x04, 1, X86InstInfo{"ADD", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1}},
@@ -210,16 +210,16 @@ const std::array<X86InstInfo, MAX_PRIMARY_TABLE_SIZE> BaseOps = []() consteval {
{0x06, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_06] }}},
{0x07, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_07] }}},
{0x08, 1, X86InstInfo{"OR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x09, 1, X86InstInfo{"OR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x08, 1, X86InstInfo{"OR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{0x09, 1, X86InstInfo{"OR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{0x0A, 1, X86InstInfo{"OR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0}},
{0x0B, 1, X86InstInfo{"OR", TYPE_INST, FLAGS_MODRM, 0}},
{0x0C, 1, X86InstInfo{"OR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1}},
{0x0D, 1, X86InstInfo{"OR", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{0x0E, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_0E] }}},
{0x10, 1, X86InstInfo{"ADC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x11, 1, X86InstInfo{"ADC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_DISPLACE_SIZE_DIV_2, 0}},
{0x10, 1, X86InstInfo{"ADC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{0x11, 1, X86InstInfo{"ADC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_SUPPORTS_LOCK, 0}},
{0x12, 1, X86InstInfo{"ADC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0}},
{0x13, 1, X86InstInfo{"ADC", TYPE_INST, FLAGS_MODRM, 0}},
{0x14, 1, X86InstInfo{"ADC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1}},
@@ -227,8 +227,8 @@ const std::array<X86InstInfo, MAX_PRIMARY_TABLE_SIZE> BaseOps = []() consteval {
{0x16, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_16] }}},
{0x17, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_17] }}},
{0x18, 1, X86InstInfo{"SBB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x19, 1, X86InstInfo{"SBB", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_DISPLACE_SIZE_DIV_2, 0}},
{0x18, 1, X86InstInfo{"SBB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{0x19, 1, X86InstInfo{"SBB", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_SUPPORTS_LOCK, 0}},
{0x1A, 1, X86InstInfo{"SBB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0}},
{0x1B, 1, X86InstInfo{"SBB", TYPE_INST, FLAGS_MODRM, 0}},
{0x1C, 1, X86InstInfo{"SBB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1}},
@@ -236,24 +236,24 @@ const std::array<X86InstInfo, MAX_PRIMARY_TABLE_SIZE> BaseOps = []() consteval {
{0x1E, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_1E] }}},
{0x1F, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_1F] }}},
{0x20, 1, X86InstInfo{"AND", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x21, 1, X86InstInfo{"AND", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x20, 1, X86InstInfo{"AND", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{0x21, 1, X86InstInfo{"AND", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{0x22, 1, X86InstInfo{"AND", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0}},
{0x23, 1, X86InstInfo{"AND", TYPE_INST, FLAGS_MODRM, 0}},
{0x24, 1, X86InstInfo{"AND", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1}},
{0x25, 1, X86InstInfo{"AND", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{0x27, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_27] }}},
{0x28, 1, X86InstInfo{"SUB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x29, 1, X86InstInfo{"SUB", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x28, 1, X86InstInfo{"SUB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{0x29, 1, X86InstInfo{"SUB", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{0x2A, 1, X86InstInfo{"SUB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0}},
{0x2B, 1, X86InstInfo{"SUB", TYPE_INST, FLAGS_MODRM, 0}},
{0x2C, 1, X86InstInfo{"SUB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1}},
{0x2D, 1, X86InstInfo{"SUB", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{0x2F, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_2F] }}},
{0x30, 1, X86InstInfo{"XOR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x31, 1, X86InstInfo{"XOR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x30, 1, X86InstInfo{"XOR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{0x31, 1, X86InstInfo{"XOR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{0x32, 1, X86InstInfo{"XOR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0}},
{0x33, 1, X86InstInfo{"XOR", TYPE_INST, FLAGS_MODRM, 0}},
{0x34, 1, X86InstInfo{"XOR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1}},
@@ -310,8 +310,8 @@ const std::array<X86InstInfo, MAX_PRIMARY_TABLE_SIZE> BaseOps = []() consteval {
{0x84, 1, X86InstInfo{"TEST", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x85, 1, X86InstInfo{"TEST", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x86, 1, X86InstInfo{"XCHG", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x87, 1, X86InstInfo{"XCHG", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x86, 1, X86InstInfo{"XCHG", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{0x87, 1, X86InstInfo{"XCHG", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{0x88, 1, X86InstInfo{"MOV", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x89, 1, X86InstInfo{"MOV", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
@@ -34,7 +34,7 @@ constexpr std::array<X86InstInfo[2], ENTRY_MAX> H0F3A_ArchSelect_LUT = {{
// ENTRY_1_3A_66_22
{
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
{"PINSRQ", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_SRC_GPR, 1, { .OpDispatch = &IR::OpDispatchBuilder::PINSROp<IR::OpSize::i64Bit> }},
{"PINSRQ", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_SRC_GPR, 1, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::PINSROp, IR::OpSize::i64Bit> }},
},
}};
@@ -28,31 +28,31 @@ enum PrimaryGroup_LUT {
constexpr std::array<X86InstInfo[2], ENTRY_MAX> PrimaryGroup_ArchSelect_LUT = {{
{
{"ADD", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, { .OpDispatch = &IR::OpDispatchBuilder::SecondaryALUOp }},
{"ADD", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1, { .OpDispatch = &IR::OpDispatchBuilder::SecondaryALUOp }},
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
{
{"OR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, { .OpDispatch = &IR::OpDispatchBuilder::SecondaryALUOp }},
{"OR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1, { .OpDispatch = &IR::OpDispatchBuilder::SecondaryALUOp }},
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
{
{"ADC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::ADCOp, 1> }},
{"ADC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::ADCOp, 1> }},
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
{
{"SBB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::SBBOp, 1> }},
{"SBB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::SBBOp, 1> }},
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
{
{"AND", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, { .OpDispatch = &IR::OpDispatchBuilder::SecondaryALUOp }},
{"AND", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1, { .OpDispatch = &IR::OpDispatchBuilder::SecondaryALUOp }},
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
{
{"SUB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, { .OpDispatch = &IR::OpDispatchBuilder::SecondaryALUOp }},
{"SUB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1, { .OpDispatch = &IR::OpDispatchBuilder::SecondaryALUOp }},
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
{
{"XOR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, { .OpDispatch = &IR::OpDispatchBuilder::SecondaryALUOp }},
{"XOR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1, { .OpDispatch = &IR::OpDispatchBuilder::SecondaryALUOp }},
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
{
@@ -66,23 +66,23 @@ constexpr std::array<X86InstInfo, MAX_INST_GROUP_TABLE_SIZE> PrimaryInstGroupOps
#define OPD(group, prefix, Reg) (((group - FEXCore::X86Tables::TYPE_GROUP_1) << 6) | (prefix) << 3 | (Reg))
constexpr U16U8InfoStruct PrimaryGroupOpTable[] = {
// GROUP_1 | 0x80 | reg
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 0), 1, X86InstInfo{"ADD", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 1), 1, X86InstInfo{"OR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 2), 1, X86InstInfo{"ADC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 3), 1, X86InstInfo{"SBB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 4), 1, X86InstInfo{"AND", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 5), 1, X86InstInfo{"SUB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 6), 1, X86InstInfo{"XOR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 7), 1, X86InstInfo{"CMP", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 0), 1, X86InstInfo{"ADD", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 1), 1, X86InstInfo{"OR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 2), 1, X86InstInfo{"ADC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 3), 1, X86InstInfo{"SBB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 4), 1, X86InstInfo{"AND", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 5), 1, X86InstInfo{"SUB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 6), 1, X86InstInfo{"XOR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 7), 1, X86InstInfo{"CMP", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 0), 1, X86InstInfo{"ADD", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 1), 1, X86InstInfo{"OR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 2), 1, X86InstInfo{"ADC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 3), 1, X86InstInfo{"SBB", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 4), 1, X86InstInfo{"AND", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 5), 1, X86InstInfo{"SUB", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 6), 1, X86InstInfo{"XOR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 7), 1, X86InstInfo{"CMP", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 0), 1, X86InstInfo{"ADD", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_SUPPORTS_LOCK, 4}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 1), 1, X86InstInfo{"OR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_SUPPORTS_LOCK, 4}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 2), 1, X86InstInfo{"ADC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_SUPPORTS_LOCK, 4}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 3), 1, X86InstInfo{"SBB", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_SUPPORTS_LOCK, 4}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 4), 1, X86InstInfo{"AND", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_SUPPORTS_LOCK, 4}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 5), 1, X86InstInfo{"SUB", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_SUPPORTS_LOCK, 4}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 6), 1, X86InstInfo{"XOR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_SUPPORTS_LOCK, 4}},
{OPD(TYPE_GROUP_1, OpToIndex(0x81), 7), 1, X86InstInfo{"CMP", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_SUPPORTS_LOCK, 4}},
// Duplicates the 0x80 opcode group
{OPD(TYPE_GROUP_1, OpToIndex(0x82), 0), 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = PrimaryGroup_ArchSelect_LUT[ENTRY_1_82_0] }}},
@@ -94,14 +94,14 @@ constexpr std::array<X86InstInfo, MAX_INST_GROUP_TABLE_SIZE> PrimaryInstGroupOps
{OPD(TYPE_GROUP_1, OpToIndex(0x82), 6), 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = PrimaryGroup_ArchSelect_LUT[ENTRY_1_82_6] }}},
{OPD(TYPE_GROUP_1, OpToIndex(0x82), 7), 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = PrimaryGroup_ArchSelect_LUT[ENTRY_1_82_7] }}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 0), 1, X86InstInfo{"ADD", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 1), 1, X86InstInfo{"OR", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 2), 1, X86InstInfo{"ADC", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 3), 1, X86InstInfo{"SBB", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 4), 1, X86InstInfo{"AND", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 5), 1, X86InstInfo{"SUB", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 6), 1, X86InstInfo{"XOR", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 7), 1, X86InstInfo{"CMP", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 0), 1, X86InstInfo{"ADD", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 1), 1, X86InstInfo{"OR", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 2), 1, X86InstInfo{"ADC", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 3), 1, X86InstInfo{"SBB", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 4), 1, X86InstInfo{"AND", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 5), 1, X86InstInfo{"SUB", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 6), 1, X86InstInfo{"XOR", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_1, OpToIndex(0x83), 7), 1, X86InstInfo{"CMP", TYPE_INST, FLAGS_SRC_SEXT | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
// GROUP 2
{OPD(TYPE_GROUP_2, OpToIndex(0xC0), 0), 1, X86InstInfo{"ROL", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
@@ -161,8 +161,8 @@ constexpr std::array<X86InstInfo, MAX_INST_GROUP_TABLE_SIZE> PrimaryInstGroupOps
// GROUP 3
{OPD(TYPE_GROUP_3, OpToIndex(0xF6), 0), 1, X86InstInfo{"TEST", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF6), 1), 1, X86InstInfo{"TEST", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF6), 2), 1, X86InstInfo{"NOT", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF6), 3), 1, X86InstInfo{"NEG", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF6), 2), 1, X86InstInfo{"NOT", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF6), 3), 1, X86InstInfo{"NEG", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF6), 4), 1, X86InstInfo{"MUL", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF6), 5), 1, X86InstInfo{"IMUL", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF6), 6), 1, X86InstInfo{"DIV", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
@@ -170,21 +170,21 @@ constexpr std::array<X86InstInfo, MAX_INST_GROUP_TABLE_SIZE> PrimaryInstGroupOps
{OPD(TYPE_GROUP_3, OpToIndex(0xF7), 0), 1, X86InstInfo{"TEST", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF7), 1), 1, X86InstInfo{"TEST", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT64BIT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF7), 2), 1, X86InstInfo{"NOT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF7), 3), 1, X86InstInfo{"NEG", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF7), 2), 1, X86InstInfo{"NOT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF7), 3), 1, X86InstInfo{"NEG", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF7), 4), 1, X86InstInfo{"MUL", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF7), 5), 1, X86InstInfo{"IMUL", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF7), 6), 1, X86InstInfo{"DIV", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_3, OpToIndex(0xF7), 7), 1, X86InstInfo{"IDIV", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
// GROUP 4
{OPD(TYPE_GROUP_4, OpToIndex(0xFE), 0), 1, X86InstInfo{"INC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_4, OpToIndex(0xFE), 1), 1, X86InstInfo{"DEC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_4, OpToIndex(0xFE), 0), 1, X86InstInfo{"INC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{OPD(TYPE_GROUP_4, OpToIndex(0xFE), 1), 1, X86InstInfo{"DEC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{OPD(TYPE_GROUP_4, OpToIndex(0xFE), 2), 6, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
// GROUP 5
{OPD(TYPE_GROUP_5, OpToIndex(0xFF), 0), 1, X86InstInfo{"INC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_5, OpToIndex(0xFF), 1), 1, X86InstInfo{"DEC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{OPD(TYPE_GROUP_5, OpToIndex(0xFF), 0), 1, X86InstInfo{"INC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{OPD(TYPE_GROUP_5, OpToIndex(0xFF), 1), 1, X86InstInfo{"DEC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{OPD(TYPE_GROUP_5, OpToIndex(0xFF), 2), 1, X86InstInfo{"CALL", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_MODRM | FLAGS_BLOCK_END | FLAGS_CALL , 0}},
{OPD(TYPE_GROUP_5, OpToIndex(0xFF), 3), 1, X86InstInfo{"CALLF", TYPE_INST, FLAGS_SETS_RIP | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY | FLAGS_BLOCK_END, 0}},
{OPD(TYPE_GROUP_5, OpToIndex(0xFF), 4), 1, X86InstInfo{"JMP", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_MODRM | FLAGS_BLOCK_END , 0}},
@@ -50,7 +50,7 @@ constexpr std::array<X86InstInfo[2], ENTRY_MAX> SecondGroup_ArchSelect_LUT = {{
},
}};
constexpr auto SecondInstGroupOps = []() consteval {
constexpr std::array<X86InstInfo, MAX_INST_SECOND_GROUP_TABLE_SIZE> SecondInstGroupOps = []() consteval {
std::array<X86InstInfo, MAX_INST_SECOND_GROUP_TABLE_SIZE> Table{};
constexpr U16U8InfoStruct SecondaryExtensionOpTable[] = {
// GROUP 1
@@ -139,37 +139,37 @@ constexpr auto SecondInstGroupOps = []() consteval {
{OPD(TYPE_GROUP_8, PF_NONE, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_NONE, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_NONE, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_NONE, 4), 1, X86InstInfo{"BT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_NONE, 5), 1, X86InstInfo{"BTS", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_NONE, 6), 1, X86InstInfo{"BTR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_NONE, 7), 1, X86InstInfo{"BTC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_NONE, 4), 1, X86InstInfo{"BT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_8, PF_NONE, 5), 1, X86InstInfo{"BTS", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_8, PF_NONE, 6), 1, X86InstInfo{"BTR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_8, PF_NONE, 7), 1, X86InstInfo{"BTC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_8, PF_F3, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_F3, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_F3, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_F3, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_F3, 4), 1, X86InstInfo{"BT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_F3, 5), 1, X86InstInfo{"BTS", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_F3, 6), 1, X86InstInfo{"BTR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_F3, 7), 1, X86InstInfo{"BTC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_F3, 4), 1, X86InstInfo{"BT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_8, PF_F3, 5), 1, X86InstInfo{"BTS", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_8, PF_F3, 6), 1, X86InstInfo{"BTR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_8, PF_F3, 7), 1, X86InstInfo{"BTC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_8, PF_66, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_66, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_66, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_66, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_66, 4), 1, X86InstInfo{"BT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_66, 5), 1, X86InstInfo{"BTS", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_66, 6), 1, X86InstInfo{"BTR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_66, 7), 1, X86InstInfo{"BTC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_66, 4), 1, X86InstInfo{"BT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_8, PF_66, 5), 1, X86InstInfo{"BTS", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_8, PF_66, 6), 1, X86InstInfo{"BTR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_8, PF_66, 7), 1, X86InstInfo{"BTC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_8, PF_F2, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_F2, 1), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_F2, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_F2, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_8, PF_F2, 4), 1, X86InstInfo{"BT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_F2, 5), 1, X86InstInfo{"BTS", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_F2, 6), 1, X86InstInfo{"BTR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_F2, 7), 1, X86InstInfo{"BTC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 1}},
{OPD(TYPE_GROUP_8, PF_F2, 4), 1, X86InstInfo{"BT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_8, PF_F2, 5), 1, X86InstInfo{"BTS", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_8, PF_F2, 6), 1, X86InstInfo{"BTR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
{OPD(TYPE_GROUP_8, PF_F2, 7), 1, X86InstInfo{"BTC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 1}},
// GROUP 9
@@ -179,7 +179,7 @@ constexpr auto SecondInstGroupOps = []() consteval {
// CMPXCHG8B/16B works with all prefixes
// Tooling fails to decode CMPXCHG with prefix
{OPD(TYPE_GROUP_9, PF_NONE, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_NONE, 1), 1, X86InstInfo{"CMPXCHG8B/16B", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_9, PF_NONE, 1), 1, X86InstInfo{"CMPXCHG8B/16B", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY | FLAGS_SUPPORTS_LOCK, 0}},
{OPD(TYPE_GROUP_9, PF_NONE, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_NONE, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_NONE, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
@@ -188,7 +188,7 @@ constexpr auto SecondInstGroupOps = []() consteval {
{OPD(TYPE_GROUP_9, PF_NONE, 7), 1, X86InstInfo{"RDSEED", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY, 0}},
{OPD(TYPE_GROUP_9, PF_F3, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_F3, 1), 1, X86InstInfo{"CMPXCHG8B/16B", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_9, PF_F3, 1), 1, X86InstInfo{"CMPXCHG8B/16B", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY | FLAGS_SUPPORTS_LOCK, 0}},
{OPD(TYPE_GROUP_9, PF_F3, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_F3, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_F3, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
@@ -197,7 +197,7 @@ constexpr auto SecondInstGroupOps = []() consteval {
{OPD(TYPE_GROUP_9, PF_F3, 7), 1, X86InstInfo{"RDPID", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY, 0}},
{OPD(TYPE_GROUP_9, PF_66, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_66, 1), 1, X86InstInfo{"CMPXCHG8B/16B", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_9, PF_66, 1), 1, X86InstInfo{"CMPXCHG8B/16B", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY | FLAGS_SUPPORTS_LOCK, 0}},
{OPD(TYPE_GROUP_9, PF_66, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_66, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_66, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
@@ -206,7 +206,7 @@ constexpr auto SecondInstGroupOps = []() consteval {
{OPD(TYPE_GROUP_9, PF_66, 7), 1, X86InstInfo{"RDSEED", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY, 0}},
{OPD(TYPE_GROUP_9, PF_F2, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_F2, 1), 1, X86InstInfo{"CMPXCHG8B/16B", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_9, PF_F2, 1), 1, X86InstInfo{"CMPXCHG8B/16B", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY | FLAGS_SUPPORTS_LOCK, 0}},
{OPD(TYPE_GROUP_9, PF_F2, 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_F2, 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(TYPE_GROUP_9, PF_F2, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
@@ -402,37 +402,37 @@ constexpr auto SecondInstGroupOps = []() consteval {
// GROUP 16
// AMD documentation claims again that this entire group is n/a to prefix
// Tooling once again fails to disassemble oens with the prefix. Disable until proven otherwise
{OPD(TYPE_GROUP_16, PF_NONE, 0), 1, X86InstInfo{"PREFETCH NTA", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_NONE, 1), 1, X86InstInfo{"PREFETCH T0", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_NONE, 2), 1, X86InstInfo{"PREFETCH T1", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_NONE, 3), 1, X86InstInfo{"PREFETCH T2", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_NONE, 0), 1, X86InstInfo{"PREFETCH NTA", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_NONE, 1), 1, X86InstInfo{"PREFETCH T0", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_NONE, 2), 1, X86InstInfo{"PREFETCH T1", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_NONE, 3), 1, X86InstInfo{"PREFETCH T2", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_NONE, 4), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_NONE, 5), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_NONE, 6), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_NONE, 7), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 0), 1, X86InstInfo{"PREFETCH NTA", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 1), 1, X86InstInfo{"PREFETCH T0", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 2), 1, X86InstInfo{"PREFETCH T1", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 3), 1, X86InstInfo{"PREFETCH T2", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 0), 1, X86InstInfo{"PREFETCH NTA", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 1), 1, X86InstInfo{"PREFETCH T0", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 2), 1, X86InstInfo{"PREFETCH T1", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 3), 1, X86InstInfo{"PREFETCH T2", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 4), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 5), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 6), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 7), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_66, 0), 1, X86InstInfo{"PREFETCH NTA", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_66, 1), 1, X86InstInfo{"PREFETCH T0", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_66, 2), 1, X86InstInfo{"PREFETCH T1", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_66, 3), 1, X86InstInfo{"PREFETCH T2", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_66, 0), 1, X86InstInfo{"PREFETCH NTA", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_66, 1), 1, X86InstInfo{"PREFETCH T0", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_66, 2), 1, X86InstInfo{"PREFETCH T1", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_66, 3), 1, X86InstInfo{"PREFETCH T2", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_66, 4), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_66, 5), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_66, 6), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_66, 7), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F2, 0), 1, X86InstInfo{"PREFETCH NTA", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_F2, 1), 1, X86InstInfo{"PREFETCH T0", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_F2, 2), 1, X86InstInfo{"PREFETCH T1", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_F2, 3), 1, X86InstInfo{"PREFETCH T2", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_F2, 0), 1, X86InstInfo{"PREFETCH NTA", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F2, 1), 1, X86InstInfo{"PREFETCH T0", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F2, 2), 1, X86InstInfo{"PREFETCH T1", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F2, 3), 1, X86InstInfo{"PREFETCH T2", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F2, 4), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F2, 5), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F2, 6), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
@@ -31,19 +31,19 @@ constexpr std::array<X86InstInfo[2], ENTRY_MAX> Secondary_ArchSelect_LUT = {{
},
{
{"PUSH FS", TYPE_INST, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_FS_PREFIX> } },
{"PUSH FS", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_FS_PREFIX> } },
{"PUSH FS", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_FS_PREFIX> } },
},
{
{"POP FS", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_DEF) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::POPSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_FS_PREFIX> } },
{"POP FS", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_64BIT) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::POPSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_FS_PREFIX> } },
{"POP FS", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_64BITDEF) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::POPSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_FS_PREFIX> } },
},
{
{"PUSH GS", TYPE_INST, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_GS_PREFIX> } },
{"PUSH GS", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_GS_PREFIX> } },
{"PUSH GS", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_GS_PREFIX> } },
},
{
{"POP GS", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_DEF) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::POPSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_GS_PREFIX> } },
{"POP GS", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_64BIT) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::POPSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_GS_PREFIX> } },
{"POP GS", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_64BITDEF) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::POPSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_GS_PREFIX> } },
},
}};
@@ -61,8 +61,8 @@ constexpr std::array<X86InstInfo, MAX_SECOND_TABLE_SIZE> SecondBaseOps = []() co
{0x05, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Secondary_ArchSelect_LUT[ENTRY_05] }}},
{0x06, 1, X86InstInfo{"CLTS", TYPE_INST, FLAGS_NO_OVERLAY, 0}},
{0x07, 1, X86InstInfo{"SYSRET", TYPE_INST, FLAGS_NO_OVERLAY, 0}},
{0x08, 1, X86InstInfo{"INVD", TYPE_PRIV, FLAGS_NO_OVERLAY, 0}},
{0x09, 1, X86InstInfo{"WBINVD", TYPE_PRIV, FLAGS_NO_OVERLAY, 0}},
{0x08, 1, X86InstInfo{"INVD", TYPE_INST, FLAGS_NO_OVERLAY, 0}},
{0x09, 1, X86InstInfo{"WBINVD", TYPE_INST, FLAGS_NO_OVERLAY, 0}},
{0x0A, 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NO_OVERLAY, 0}},
{0x0B, 1, X86InstInfo{"UD2", TYPE_INST, FLAGS_BLOCK_END | FLAGS_NO_OVERLAY, 0}},
{0x0C, 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NO_OVERLAY, 0}},
@@ -205,23 +205,23 @@ constexpr std::array<X86InstInfo, MAX_SECOND_TABLE_SIZE> SecondBaseOps = []() co
{0xA0, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .Indirect = Secondary_ArchSelect_LUT[ENTRY_A0] }}},
{0xA1, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .Indirect = Secondary_ArchSelect_LUT[ENTRY_A1] }}},
{0xA2, 1, X86InstInfo{"CPUID", TYPE_INST, FLAGS_SF_SRC_RAX | FLAGS_NO_OVERLAY, 0}},
{0xA3, 1, X86InstInfo{"BT", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_NO_OVERLAY, 0}},
{0xA3, 1, X86InstInfo{"BT", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_NO_OVERLAY | FLAGS_SUPPORTS_LOCK, 0}},
{0xA4, 1, X86InstInfo{"SHLD", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_NO_OVERLAY, 1}},
{0xA5, 1, X86InstInfo{"SHLD", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_SRC_RCX | FLAGS_NO_OVERLAY, 0}},
{0xA6, 2, X86InstInfo{"", TYPE_INVALID, FLAGS_NO_OVERLAY, 0}},
{0xA8, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .Indirect = Secondary_ArchSelect_LUT[ENTRY_A8] }}},
{0xA9, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .Indirect = Secondary_ArchSelect_LUT[ENTRY_A9] }}},
{0xAA, 1, X86InstInfo{"RSM", TYPE_PRIV, FLAGS_NO_OVERLAY, 0}},
{0xAB, 1, X86InstInfo{"BTS", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_NO_OVERLAY, 0}},
{0xAB, 1, X86InstInfo{"BTS", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_NO_OVERLAY | FLAGS_SUPPORTS_LOCK, 0}},
{0xAC, 1, X86InstInfo{"SHRD", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_NO_OVERLAY, 1}},
{0xAD, 1, X86InstInfo{"SHRD", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_SRC_RCX | FLAGS_NO_OVERLAY, 0}},
{0xAE, 1, X86InstInfo{"", TYPE_GROUP_15, FLAGS_NO_OVERLAY, 0}},
{0xAF, 1, X86InstInfo{"IMUL", TYPE_INST, FLAGS_MODRM | FLAGS_NO_OVERLAY, 0}},
{0xB0, 1, X86InstInfo{"CMPXCHG", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_NO_OVERLAY, 0}},
{0xB1, 1, X86InstInfo{"CMPXCHG", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_NO_OVERLAY, 0}},
{0xB0, 1, X86InstInfo{"CMPXCHG", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_NO_OVERLAY | FLAGS_SUPPORTS_LOCK, 0}},
{0xB1, 1, X86InstInfo{"CMPXCHG", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_NO_OVERLAY | FLAGS_SUPPORTS_LOCK, 0}},
{0xB2, 1, X86InstInfo{"LSS", TYPE_INVALID, FLAGS_NO_OVERLAY, 0}},
{0xB3, 1, X86InstInfo{"BTR", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_NO_OVERLAY, 0}},
{0xB3, 1, X86InstInfo{"BTR", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_NO_OVERLAY | FLAGS_SUPPORTS_LOCK, 0}},
{0xB4, 1, X86InstInfo{"LFS", TYPE_INVALID, FLAGS_NO_OVERLAY, 0}},
{0xB5, 1, X86InstInfo{"LGS", TYPE_INVALID, FLAGS_NO_OVERLAY, 0}},
{0xB6, 1, X86InstInfo{"MOVZX", TYPE_INST, GenFlagsSrcSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_NO_OVERLAY, 0}},
@@ -229,14 +229,14 @@ constexpr std::array<X86InstInfo, MAX_SECOND_TABLE_SIZE> SecondBaseOps = []() co
{0xB8, 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{0xB9, 1, X86InstInfo{"", TYPE_GROUP_10, FLAGS_NO_OVERLAY, 0}},
{0xBA, 1, X86InstInfo{"", TYPE_GROUP_8, FLAGS_NO_OVERLAY, 0}},
{0xBB, 1, X86InstInfo{"BTC", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_NO_OVERLAY, 0}},
{0xBB, 1, X86InstInfo{"BTC", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_NO_OVERLAY | FLAGS_SUPPORTS_LOCK, 0}},
{0xBC, 1, X86InstInfo{"BSF", TYPE_INST, FLAGS_MODRM | FLAGS_NO_OVERLAY66, 0}},
{0xBD, 1, X86InstInfo{"BSR", TYPE_INST, FLAGS_MODRM | FLAGS_NO_OVERLAY66, 0}},
{0xBE, 1, X86InstInfo{"MOVSX", TYPE_INST, GenFlagsSrcSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_NO_OVERLAY, 0}},
{0xBF, 1, X86InstInfo{"MOVSX", TYPE_INST, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_MODRM | FLAGS_NO_OVERLAY, 0}},
{0xC0, 1, X86InstInfo{"XADD", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0xC1, 1, X86InstInfo{"XADD", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_NO_OVERLAY, 0}},
{0xC0, 1, X86InstInfo{"XADD", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SUPPORTS_LOCK, 0}},
{0xC1, 1, X86InstInfo{"XADD", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_NO_OVERLAY | FLAGS_SUPPORTS_LOCK, 0}},
{0xC2, 1, X86InstInfo{"CMPPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1}},
{0xC3, 1, X86InstInfo{"MOVNTI", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_SF_MOD_DST, 0}},
{0xC4, 1, X86InstInfo{"PINSRW", TYPE_INST, GenFlagsSizes(SIZE_64BIT, SIZE_16BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX | FLAGS_SF_SRC_GPR, 1}},
@@ -474,7 +474,7 @@ namespace AVX256 {
{OPD(1, 0b00, 0x12), 1, &OpDispatchBuilder::VMOVLPOp},
{OPD(1, 0b01, 0x12), 1, &OpDispatchBuilder::VMOVLPOp},
{OPD(1, 0b10, 0x12), 1, &OpDispatchBuilder::VMOVSLDUPOp},
{OPD(1, 0b10, 0x12), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VMOVSLDUPOp, true>},
{OPD(1, 0b11, 0x12), 1, &OpDispatchBuilder::VMOVDDUPOp},
{OPD(1, 0b00, 0x13), 1, &OpDispatchBuilder::VMOVLPOp},
{OPD(1, 0b01, 0x13), 1, &OpDispatchBuilder::VMOVLPOp},
@@ -487,7 +487,7 @@ namespace AVX256 {
{OPD(1, 0b00, 0x16), 1, &OpDispatchBuilder::VMOVHPOp},
{OPD(1, 0b01, 0x16), 1, &OpDispatchBuilder::VMOVHPOp},
{OPD(1, 0b10, 0x16), 1, &OpDispatchBuilder::VMOVSHDUPOp},
{OPD(1, 0b10, 0x16), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VMOVSHDUPOp, true>},
{OPD(1, 0b00, 0x17), 1, &OpDispatchBuilder::VMOVHPOp},
{OPD(1, 0b01, 0x17), 1, &OpDispatchBuilder::VMOVHPOp},
@@ -496,36 +496,36 @@ namespace AVX256 {
{OPD(1, 0b00, 0x29), 1, &OpDispatchBuilder::VMOVAPS_VMOVAPDOp},
{OPD(1, 0b01, 0x29), 1, &OpDispatchBuilder::VMOVAPS_VMOVAPDOp},
{OPD(1, 0b10, 0x2A), 1, &OpDispatchBuilder::AVXInsertCVTGPR_To_FPR<OpSize::i32Bit>},
{OPD(1, 0b11, 0x2A), 1, &OpDispatchBuilder::AVXInsertCVTGPR_To_FPR<OpSize::i64Bit>},
{OPD(1, 0b10, 0x2A), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXInsertCVTGPR_To_FPR, OpSize::i32Bit>},
{OPD(1, 0b11, 0x2A), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXInsertCVTGPR_To_FPR, OpSize::i64Bit>},
{OPD(1, 0b00, 0x2B), 1, &OpDispatchBuilder::MOVVectorNTOp},
{OPD(1, 0b01, 0x2B), 1, &OpDispatchBuilder::MOVVectorNTOp},
{OPD(1, 0b00, 0x2B), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVVectorNTOp, true>},
{OPD(1, 0b01, 0x2B), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVVectorNTOp, true>},
{OPD(1, 0b10, 0x2C), 1, &OpDispatchBuilder::CVTFPR_To_GPR<OpSize::i32Bit, false>},
{OPD(1, 0b11, 0x2C), 1, &OpDispatchBuilder::CVTFPR_To_GPR<OpSize::i64Bit, false>},
{OPD(1, 0b10, 0x2C), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::CVTFPR_To_GPR, OpSize::i32Bit, false>},
{OPD(1, 0b11, 0x2C), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::CVTFPR_To_GPR, OpSize::i64Bit, false>},
{OPD(1, 0b10, 0x2D), 1, &OpDispatchBuilder::CVTFPR_To_GPR<OpSize::i32Bit, true>},
{OPD(1, 0b11, 0x2D), 1, &OpDispatchBuilder::CVTFPR_To_GPR<OpSize::i64Bit, true>},
{OPD(1, 0b10, 0x2D), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::CVTFPR_To_GPR, OpSize::i32Bit, true>},
{OPD(1, 0b11, 0x2D), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::CVTFPR_To_GPR, OpSize::i64Bit, true>},
{OPD(1, 0b00, 0x2E), 1, &OpDispatchBuilder::UCOMISxOp<OpSize::i32Bit>},
{OPD(1, 0b01, 0x2E), 1, &OpDispatchBuilder::UCOMISxOp<OpSize::i64Bit>},
{OPD(1, 0b00, 0x2F), 1, &OpDispatchBuilder::UCOMISxOp<OpSize::i32Bit>},
{OPD(1, 0b01, 0x2F), 1, &OpDispatchBuilder::UCOMISxOp<OpSize::i64Bit>},
{OPD(1, 0b00, 0x2E), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::UCOMISxOp, OpSize::i32Bit>},
{OPD(1, 0b01, 0x2E), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::UCOMISxOp, OpSize::i64Bit>},
{OPD(1, 0b00, 0x2F), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::UCOMISxOp, OpSize::i32Bit>},
{OPD(1, 0b01, 0x2F), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::UCOMISxOp, OpSize::i64Bit>},
{OPD(1, 0b00, 0x50), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVMSKOp, OpSize::i32Bit>},
{OPD(1, 0b01, 0x50), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVMSKOp, OpSize::i64Bit>},
{OPD(1, 0b00, 0x51), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorUnaryOp, IR::OP_VFSQRT, OpSize::i32Bit>},
{OPD(1, 0b01, 0x51), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorUnaryOp, IR::OP_VFSQRT, OpSize::i64Bit>},
{OPD(1, 0b10, 0x51), 1, &OpDispatchBuilder::AVXVectorScalarUnaryInsertALUOp<IR::OP_VFSQRTSCALARINSERT, OpSize::i32Bit>},
{OPD(1, 0b11, 0x51), 1, &OpDispatchBuilder::AVXVectorScalarUnaryInsertALUOp<IR::OP_VFSQRTSCALARINSERT, OpSize::i64Bit>},
{OPD(1, 0b10, 0x51), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorScalarUnaryInsertALUOp, IR::OP_VFSQRTSCALARINSERT, OpSize::i32Bit>},
{OPD(1, 0b11, 0x51), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorScalarUnaryInsertALUOp, IR::OP_VFSQRTSCALARINSERT, OpSize::i64Bit>},
{OPD(1, 0b00, 0x52), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorUnaryOp, IR::OP_VFRSQRT, OpSize::i32Bit>},
{OPD(1, 0b10, 0x52), 1, &OpDispatchBuilder::AVXVectorScalarUnaryInsertALUOp<IR::OP_VFRSQRTSCALARINSERT, OpSize::i32Bit>},
{OPD(1, 0b10, 0x52), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorScalarUnaryInsertALUOp, IR::OP_VFRSQRTSCALARINSERT, OpSize::i32Bit>},
{OPD(1, 0b00, 0x53), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorUnaryOp, IR::OP_VFRECP, OpSize::i32Bit>},
{OPD(1, 0b10, 0x53), 1, &OpDispatchBuilder::AVXVectorScalarUnaryInsertALUOp<IR::OP_VFRECPSCALARINSERT, OpSize::i32Bit>},
{OPD(1, 0b10, 0x53), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorScalarUnaryInsertALUOp, IR::OP_VFRECPSCALARINSERT, OpSize::i32Bit>},
{OPD(1, 0b00, 0x54), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VAND, OpSize::i128Bit>},
{OPD(1, 0b01, 0x54), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VAND, OpSize::i128Bit>},
@@ -541,42 +541,42 @@ namespace AVX256 {
{OPD(1, 0b00, 0x58), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VFADD, OpSize::i32Bit>},
{OPD(1, 0b01, 0x58), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VFADD, OpSize::i64Bit>},
{OPD(1, 0b10, 0x58), 1, &OpDispatchBuilder::AVXVectorScalarInsertALUOp<IR::OP_VFADDSCALARINSERT, OpSize::i32Bit>},
{OPD(1, 0b11, 0x58), 1, &OpDispatchBuilder::AVXVectorScalarInsertALUOp<IR::OP_VFADDSCALARINSERT, OpSize::i64Bit>},
{OPD(1, 0b10, 0x58), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorScalarInsertALUOp, IR::OP_VFADDSCALARINSERT, OpSize::i32Bit>},
{OPD(1, 0b11, 0x58), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorScalarInsertALUOp, IR::OP_VFADDSCALARINSERT, OpSize::i64Bit>},
{OPD(1, 0b00, 0x59), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VFMUL, OpSize::i32Bit>},
{OPD(1, 0b01, 0x59), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VFMUL, OpSize::i64Bit>},
{OPD(1, 0b10, 0x59), 1, &OpDispatchBuilder::AVXVectorScalarInsertALUOp<IR::OP_VFMULSCALARINSERT, OpSize::i32Bit>},
{OPD(1, 0b11, 0x59), 1, &OpDispatchBuilder::AVXVectorScalarInsertALUOp<IR::OP_VFMULSCALARINSERT, OpSize::i64Bit>},
{OPD(1, 0b10, 0x59), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorScalarInsertALUOp, IR::OP_VFMULSCALARINSERT, OpSize::i32Bit>},
{OPD(1, 0b11, 0x59), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorScalarInsertALUOp, IR::OP_VFMULSCALARINSERT, OpSize::i64Bit>},
{OPD(1, 0b00, 0x5A), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Float_To_Float, OpSize::i64Bit, OpSize::i32Bit, true>},
{OPD(1, 0b01, 0x5A), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Float_To_Float, OpSize::i32Bit, OpSize::i64Bit, true>},
{OPD(1, 0b10, 0x5A), 1, &OpDispatchBuilder::AVXInsertScalar_CVT_Float_To_Float<OpSize::i64Bit, OpSize::i32Bit>},
{OPD(1, 0b11, 0x5A), 1, &OpDispatchBuilder::AVXInsertScalar_CVT_Float_To_Float<OpSize::i32Bit, OpSize::i64Bit>},
{OPD(1, 0b10, 0x5A), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXInsertScalar_CVT_Float_To_Float, OpSize::i64Bit, OpSize::i32Bit>},
{OPD(1, 0b11, 0x5A), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXInsertScalar_CVT_Float_To_Float, OpSize::i32Bit, OpSize::i64Bit>},
{OPD(1, 0b00, 0x5B), 1, &OpDispatchBuilder::Vector_CVT_Int_To_Float<OpSize::i32Bit, false>},
{OPD(1, 0b01, 0x5B), 1, &OpDispatchBuilder::Vector_CVT_Float_To_Int<OpSize::i32Bit, true>},
{OPD(1, 0b10, 0x5B), 1, &OpDispatchBuilder::Vector_CVT_Float_To_Int<OpSize::i32Bit, false>},
{OPD(1, 0b00, 0x5B), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Int_To_Float, OpSize::i32Bit, false, true>},
{OPD(1, 0b01, 0x5B), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Float_To_Int, OpSize::i32Bit, true, true>},
{OPD(1, 0b10, 0x5B), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Float_To_Int, OpSize::i32Bit, false, true>},
{OPD(1, 0b00, 0x5C), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VFSUB, OpSize::i32Bit>},
{OPD(1, 0b01, 0x5C), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VFSUB, OpSize::i64Bit>},
{OPD(1, 0b10, 0x5C), 1, &OpDispatchBuilder::AVXVectorScalarInsertALUOp<IR::OP_VFSUBSCALARINSERT, OpSize::i32Bit>},
{OPD(1, 0b11, 0x5C), 1, &OpDispatchBuilder::AVXVectorScalarInsertALUOp<IR::OP_VFSUBSCALARINSERT, OpSize::i64Bit>},
{OPD(1, 0b10, 0x5C), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorScalarInsertALUOp, IR::OP_VFSUBSCALARINSERT, OpSize::i32Bit>},
{OPD(1, 0b11, 0x5C), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorScalarInsertALUOp, IR::OP_VFSUBSCALARINSERT, OpSize::i64Bit>},
{OPD(1, 0b00, 0x5D), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VFMIN, OpSize::i32Bit>},
{OPD(1, 0b01, 0x5D), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VFMIN, OpSize::i64Bit>},
{OPD(1, 0b10, 0x5D), 1, &OpDispatchBuilder::AVXVectorScalarInsertALUOp<IR::OP_VFMINSCALARINSERT, OpSize::i32Bit>},
{OPD(1, 0b11, 0x5D), 1, &OpDispatchBuilder::AVXVectorScalarInsertALUOp<IR::OP_VFMINSCALARINSERT, OpSize::i64Bit>},
{OPD(1, 0b10, 0x5D), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorScalarInsertALUOp, IR::OP_VFMINSCALARINSERT, OpSize::i32Bit>},
{OPD(1, 0b11, 0x5D), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorScalarInsertALUOp, IR::OP_VFMINSCALARINSERT, OpSize::i64Bit>},
{OPD(1, 0b00, 0x5E), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VFDIV, OpSize::i32Bit>},
{OPD(1, 0b01, 0x5E), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VFDIV, OpSize::i64Bit>},
{OPD(1, 0b10, 0x5E), 1, &OpDispatchBuilder::AVXVectorScalarInsertALUOp<IR::OP_VFDIVSCALARINSERT, OpSize::i32Bit>},
{OPD(1, 0b11, 0x5E), 1, &OpDispatchBuilder::AVXVectorScalarInsertALUOp<IR::OP_VFDIVSCALARINSERT, OpSize::i64Bit>},
{OPD(1, 0b10, 0x5E), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorScalarInsertALUOp, IR::OP_VFDIVSCALARINSERT, OpSize::i32Bit>},
{OPD(1, 0b11, 0x5E), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorScalarInsertALUOp, IR::OP_VFDIVSCALARINSERT, OpSize::i64Bit>},
{OPD(1, 0b00, 0x5F), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VFMAX, OpSize::i32Bit>},
{OPD(1, 0b01, 0x5F), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VFMAX, OpSize::i64Bit>},
{OPD(1, 0b10, 0x5F), 1, &OpDispatchBuilder::AVXVectorScalarInsertALUOp<IR::OP_VFMAXSCALARINSERT, OpSize::i32Bit>},
{OPD(1, 0b11, 0x5F), 1, &OpDispatchBuilder::AVXVectorScalarInsertALUOp<IR::OP_VFMAXSCALARINSERT, OpSize::i64Bit>},
{OPD(1, 0b10, 0x5F), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorScalarInsertALUOp, IR::OP_VFMAXSCALARINSERT, OpSize::i32Bit>},
{OPD(1, 0b11, 0x5F), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorScalarInsertALUOp, IR::OP_VFMAXSCALARINSERT, OpSize::i64Bit>},
{OPD(1, 0b01, 0x60), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPUNPCKLOp, OpSize::i8Bit>},
{OPD(1, 0b01, 0x61), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPUNPCKLOp, OpSize::i16Bit>},
@@ -607,8 +607,8 @@ namespace AVX256 {
{OPD(1, 0b00, 0x77), 1, &OpDispatchBuilder::VZEROOp},
{OPD(1, 0b01, 0x7C), 1, &OpDispatchBuilder::VHADDPOp<IR::OP_VFADDP, OpSize::i64Bit>},
{OPD(1, 0b11, 0x7C), 1, &OpDispatchBuilder::VHADDPOp<IR::OP_VFADDP, OpSize::i32Bit>},
{OPD(1, 0b01, 0x7C), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VHADDPOp, IR::OP_VFADDP, OpSize::i64Bit>},
{OPD(1, 0b11, 0x7C), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VHADDPOp, IR::OP_VFADDP, OpSize::i32Bit>},
{OPD(1, 0b01, 0x7D), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VHSUBPOp, OpSize::i64Bit>},
{OPD(1, 0b11, 0x7D), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VHSUBPOp, OpSize::i32Bit>},
@@ -618,19 +618,19 @@ namespace AVX256 {
{OPD(1, 0b01, 0x7F), 1, &OpDispatchBuilder::VMOVAPS_VMOVAPDOp},
{OPD(1, 0b10, 0x7F), 1, &OpDispatchBuilder::VMOVUPS_VMOVUPDOp},
{OPD(1, 0b00, 0xC2), 1, &OpDispatchBuilder::AVXVFCMPOp<OpSize::i32Bit>},
{OPD(1, 0b01, 0xC2), 1, &OpDispatchBuilder::AVXVFCMPOp<OpSize::i64Bit>},
{OPD(1, 0b10, 0xC2), 1, &OpDispatchBuilder::AVXInsertScalarFCMPOp<OpSize::i32Bit>},
{OPD(1, 0b11, 0xC2), 1, &OpDispatchBuilder::AVXInsertScalarFCMPOp<OpSize::i64Bit>},
{OPD(1, 0b00, 0xC2), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVFCMPOp, OpSize::i32Bit>},
{OPD(1, 0b01, 0xC2), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVFCMPOp, OpSize::i64Bit>},
{OPD(1, 0b10, 0xC2), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXInsertScalarFCMPOp, OpSize::i32Bit>},
{OPD(1, 0b11, 0xC2), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXInsertScalarFCMPOp, OpSize::i64Bit>},
{OPD(1, 0b01, 0xC4), 1, &OpDispatchBuilder::VPINSRWOp},
{OPD(1, 0b01, 0xC4), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPINSRBWOp, OpSize::i16Bit>},
{OPD(1, 0b01, 0xC5), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PExtrOp, OpSize::i16Bit>},
{OPD(1, 0b00, 0xC6), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VSHUFOp, OpSize::i32Bit>},
{OPD(1, 0b01, 0xC6), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VSHUFOp, OpSize::i64Bit>},
{OPD(1, 0b01, 0xD0), 1, &OpDispatchBuilder::VADDSUBPOp<OpSize::i64Bit>},
{OPD(1, 0b11, 0xD0), 1, &OpDispatchBuilder::VADDSUBPOp<OpSize::i32Bit>},
{OPD(1, 0b01, 0xD0), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VADDSUBPOp, OpSize::i64Bit>},
{OPD(1, 0b11, 0xD0), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VADDSUBPOp, OpSize::i32Bit>},
{OPD(1, 0b01, 0xD1), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPSRLDOp, OpSize::i16Bit>},
{OPD(1, 0b01, 0xD2), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPSRLDOp, OpSize::i32Bit>},
@@ -653,14 +653,14 @@ namespace AVX256 {
{OPD(1, 0b01, 0xE1), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPSRAOp, OpSize::i16Bit>},
{OPD(1, 0b01, 0xE2), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPSRAOp, OpSize::i32Bit>},
{OPD(1, 0b01, 0xE3), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VURAVG, OpSize::i16Bit>},
{OPD(1, 0b01, 0xE4), 1, &OpDispatchBuilder::VPMULHWOp<false>},
{OPD(1, 0b01, 0xE5), 1, &OpDispatchBuilder::VPMULHWOp<true>},
{OPD(1, 0b01, 0xE4), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPMULHWOp, false>},
{OPD(1, 0b01, 0xE5), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPMULHWOp, true>},
{OPD(1, 0b01, 0xE6), 1, &OpDispatchBuilder::Vector_CVT_Float_To_Int<OpSize::i64Bit, false>},
{OPD(1, 0b10, 0xE6), 1, &OpDispatchBuilder::Vector_CVT_Int_To_Float<OpSize::i32Bit, true>},
{OPD(1, 0b11, 0xE6), 1, &OpDispatchBuilder::Vector_CVT_Float_To_Int<OpSize::i64Bit, true>},
{OPD(1, 0b01, 0xE6), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Float_To_Int, OpSize::i64Bit, false, true>},
{OPD(1, 0b10, 0xE6), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Int_To_Float, OpSize::i32Bit, true, true>},
{OPD(1, 0b11, 0xE6), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Float_To_Int, OpSize::i64Bit, true, true>},
{OPD(1, 0b01, 0xE7), 1, &OpDispatchBuilder::MOVVectorNTOp},
{OPD(1, 0b01, 0xE7), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVVectorNTOp, true>},
{OPD(1, 0b01, 0xE8), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VSQSUB, OpSize::i8Bit>},
{OPD(1, 0b01, 0xE9), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VSQSUB, OpSize::i16Bit>},
@@ -671,11 +671,11 @@ namespace AVX256 {
{OPD(1, 0b01, 0xEE), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VSMAX, OpSize::i16Bit>},
{OPD(1, 0b01, 0xEF), 1, &OpDispatchBuilder::AVXVectorXOROp},
{OPD(1, 0b11, 0xF0), 1, &OpDispatchBuilder::MOVVectorUnalignedOp},
{OPD(1, 0b11, 0xF0), 1, &OpDispatchBuilder::VMOVUPS_VMOVUPDOp},
{OPD(1, 0b01, 0xF1), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPSLLOp, OpSize::i16Bit>},
{OPD(1, 0b01, 0xF2), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPSLLOp, OpSize::i32Bit>},
{OPD(1, 0b01, 0xF3), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPSLLOp, OpSize::i64Bit>},
{OPD(1, 0b01, 0xF4), 1, &OpDispatchBuilder::VPMULLOp<OpSize::i32Bit, false>},
{OPD(1, 0b01, 0xF4), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPMULLOp, OpSize::i32Bit, false>},
{OPD(1, 0b01, 0xF5), 1, &OpDispatchBuilder::VPMADDWDOp},
{OPD(1, 0b01, 0xF6), 1, &OpDispatchBuilder::VPSADBWOp},
{OPD(1, 0b01, 0xF7), 1, &OpDispatchBuilder::MASKMOVOp},
@@ -689,8 +689,8 @@ namespace AVX256 {
{OPD(1, 0b01, 0xFE), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VADD, OpSize::i32Bit>},
{OPD(2, 0b01, 0x00), 1, &OpDispatchBuilder::VPSHUFBOp},
{OPD(2, 0b01, 0x01), 1, &OpDispatchBuilder::VHADDPOp<IR::OP_VADDP, OpSize::i16Bit>},
{OPD(2, 0b01, 0x02), 1, &OpDispatchBuilder::VHADDPOp<IR::OP_VADDP, OpSize::i32Bit>},
{OPD(2, 0b01, 0x01), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VHADDPOp, IR::OP_VADDP, OpSize::i16Bit>},
{OPD(2, 0b01, 0x02), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VHADDPOp, IR::OP_VADDP, OpSize::i32Bit>},
{OPD(2, 0b01, 0x03), 1, &OpDispatchBuilder::VPHADDSWOp},
{OPD(2, 0b01, 0x04), 1, &OpDispatchBuilder::VPMADDUBSWOp},
@@ -698,14 +698,14 @@ namespace AVX256 {
{OPD(2, 0b01, 0x06), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPHSUBOp, OpSize::i32Bit>},
{OPD(2, 0b01, 0x07), 1, &OpDispatchBuilder::VPHSUBSWOp},
{OPD(2, 0b01, 0x08), 1, &OpDispatchBuilder::VPSIGN<OpSize::i8Bit>},
{OPD(2, 0b01, 0x09), 1, &OpDispatchBuilder::VPSIGN<OpSize::i16Bit>},
{OPD(2, 0b01, 0x0A), 1, &OpDispatchBuilder::VPSIGN<OpSize::i32Bit>},
{OPD(2, 0b01, 0x08), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPSIGN, OpSize::i8Bit>},
{OPD(2, 0b01, 0x09), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPSIGN, OpSize::i16Bit>},
{OPD(2, 0b01, 0x0A), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPSIGN, OpSize::i32Bit>},
{OPD(2, 0b01, 0x0B), 1, &OpDispatchBuilder::VPMULHRSWOp},
{OPD(2, 0b01, 0x0C), 1, &OpDispatchBuilder::VPERMILRegOp<OpSize::i32Bit>},
{OPD(2, 0b01, 0x0D), 1, &OpDispatchBuilder::VPERMILRegOp<OpSize::i64Bit>},
{OPD(2, 0b01, 0x0E), 1, &OpDispatchBuilder::VTESTPOp<OpSize::i32Bit>},
{OPD(2, 0b01, 0x0F), 1, &OpDispatchBuilder::VTESTPOp<OpSize::i64Bit>},
{OPD(2, 0b01, 0x0C), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPERMILRegOp, OpSize::i32Bit>},
{OPD(2, 0b01, 0x0D), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPERMILRegOp, OpSize::i64Bit>},
{OPD(2, 0b01, 0x0E), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VTESTPOp, OpSize::i32Bit>},
{OPD(2, 0b01, 0x0F), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VTESTPOp, OpSize::i64Bit>},
{OPD(2, 0b01, 0x13), 1, &OpDispatchBuilder::VCVTPH2PSOp},
{OPD(2, 0b01, 0x16), 1, &OpDispatchBuilder::VPERMDOp},
@@ -717,28 +717,28 @@ namespace AVX256 {
{OPD(2, 0b01, 0x1D), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorUnaryOp, IR::OP_VABS, OpSize::i16Bit>},
{OPD(2, 0b01, 0x1E), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorUnaryOp, IR::OP_VABS, OpSize::i32Bit>},
{OPD(2, 0b01, 0x20), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i8Bit, OpSize::i16Bit, true>},
{OPD(2, 0b01, 0x21), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i8Bit, OpSize::i32Bit, true>},
{OPD(2, 0b01, 0x22), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i8Bit, OpSize::i64Bit, true>},
{OPD(2, 0b01, 0x23), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i16Bit, OpSize::i32Bit, true>},
{OPD(2, 0b01, 0x24), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i16Bit, OpSize::i64Bit, true>},
{OPD(2, 0b01, 0x25), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i32Bit, OpSize::i64Bit, true>},
{OPD(2, 0b01, 0x20), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXExtendVectorElements, OpSize::i8Bit, OpSize::i16Bit, true>},
{OPD(2, 0b01, 0x21), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXExtendVectorElements, OpSize::i8Bit, OpSize::i32Bit, true>},
{OPD(2, 0b01, 0x22), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXExtendVectorElements, OpSize::i8Bit, OpSize::i64Bit, true>},
{OPD(2, 0b01, 0x23), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXExtendVectorElements, OpSize::i16Bit, OpSize::i32Bit, true>},
{OPD(2, 0b01, 0x24), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXExtendVectorElements, OpSize::i16Bit, OpSize::i64Bit, true>},
{OPD(2, 0b01, 0x25), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXExtendVectorElements, OpSize::i32Bit, OpSize::i64Bit, true>},
{OPD(2, 0b01, 0x28), 1, &OpDispatchBuilder::VPMULLOp<OpSize::i32Bit, true>},
{OPD(2, 0b01, 0x28), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPMULLOp, OpSize::i32Bit, true>},
{OPD(2, 0b01, 0x29), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VCMPEQ, OpSize::i64Bit>},
{OPD(2, 0b01, 0x2A), 1, &OpDispatchBuilder::MOVVectorNTOp},
{OPD(2, 0b01, 0x2A), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVVectorNTOp, true>},
{OPD(2, 0b01, 0x2B), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPACKUSOp, OpSize::i32Bit>},
{OPD(2, 0b01, 0x2C), 1, &OpDispatchBuilder::VMASKMOVOp<OpSize::i32Bit, false>},
{OPD(2, 0b01, 0x2D), 1, &OpDispatchBuilder::VMASKMOVOp<OpSize::i64Bit, false>},
{OPD(2, 0b01, 0x2E), 1, &OpDispatchBuilder::VMASKMOVOp<OpSize::i32Bit, true>},
{OPD(2, 0b01, 0x2F), 1, &OpDispatchBuilder::VMASKMOVOp<OpSize::i64Bit, true>},
{OPD(2, 0b01, 0x2C), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VMASKMOVOp, OpSize::i32Bit, false>},
{OPD(2, 0b01, 0x2D), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VMASKMOVOp, OpSize::i64Bit, false>},
{OPD(2, 0b01, 0x2E), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VMASKMOVOp, OpSize::i32Bit, true>},
{OPD(2, 0b01, 0x2F), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VMASKMOVOp, OpSize::i64Bit, true>},
{OPD(2, 0b01, 0x30), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i8Bit, OpSize::i16Bit, false>},
{OPD(2, 0b01, 0x31), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i8Bit, OpSize::i32Bit, false>},
{OPD(2, 0b01, 0x32), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i8Bit, OpSize::i64Bit, false>},
{OPD(2, 0b01, 0x33), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i16Bit, OpSize::i32Bit, false>},
{OPD(2, 0b01, 0x34), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i16Bit, OpSize::i64Bit, false>},
{OPD(2, 0b01, 0x35), 1, &OpDispatchBuilder::ExtendVectorElements<OpSize::i32Bit, OpSize::i64Bit, false>},
{OPD(2, 0b01, 0x30), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXExtendVectorElements, OpSize::i8Bit, OpSize::i16Bit, false>},
{OPD(2, 0b01, 0x31), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXExtendVectorElements, OpSize::i8Bit, OpSize::i32Bit, false>},
{OPD(2, 0b01, 0x32), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXExtendVectorElements, OpSize::i8Bit, OpSize::i64Bit, false>},
{OPD(2, 0b01, 0x33), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXExtendVectorElements, OpSize::i16Bit, OpSize::i32Bit, false>},
{OPD(2, 0b01, 0x34), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXExtendVectorElements, OpSize::i16Bit, OpSize::i64Bit, false>},
{OPD(2, 0b01, 0x35), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXExtendVectorElements, OpSize::i32Bit, OpSize::i64Bit, false>},
{OPD(2, 0b01, 0x36), 1, &OpDispatchBuilder::VPERMDOp},
{OPD(2, 0b01, 0x37), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VCMPGT, OpSize::i64Bit>},
@@ -752,7 +752,7 @@ namespace AVX256 {
{OPD(2, 0b01, 0x3F), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VUMAX, OpSize::i32Bit>},
{OPD(2, 0b01, 0x40), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorALUOp, IR::OP_VMUL, OpSize::i32Bit>},
{OPD(2, 0b01, 0x41), 1, &OpDispatchBuilder::PHMINPOSUWOp},
{OPD(2, 0b01, 0x41), 1, &OpDispatchBuilder::AVXPHMINPOSUWOp},
{OPD(2, 0b01, 0x45), 1, &OpDispatchBuilder::VPSRLVOp},
{OPD(2, 0b01, 0x46), 1, &OpDispatchBuilder::VPSRAVDOp},
{OPD(2, 0b01, 0x47), 1, &OpDispatchBuilder::VPSLLVOp},
@@ -764,13 +764,13 @@ namespace AVX256 {
{OPD(2, 0b01, 0x78), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VBROADCASTOp, OpSize::i8Bit>},
{OPD(2, 0b01, 0x79), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VBROADCASTOp, OpSize::i16Bit>},
{OPD(2, 0b01, 0x8C), 1, &OpDispatchBuilder::VPMASKMOVOp<false>},
{OPD(2, 0b01, 0x8E), 1, &OpDispatchBuilder::VPMASKMOVOp<true>},
{OPD(2, 0b01, 0x8C), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPMASKMOVOp, false>},
{OPD(2, 0b01, 0x8E), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPMASKMOVOp, true>},
{OPD(2, 0b01, 0x90), 1, &OpDispatchBuilder::VPGATHER<OpSize::i32Bit>},
{OPD(2, 0b01, 0x91), 1, &OpDispatchBuilder::VPGATHER<OpSize::i64Bit>},
{OPD(2, 0b01, 0x92), 1, &OpDispatchBuilder::VPGATHER<OpSize::i32Bit>},
{OPD(2, 0b01, 0x93), 1, &OpDispatchBuilder::VPGATHER<OpSize::i64Bit>},
{OPD(2, 0b01, 0x90), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPGATHER, OpSize::i32Bit>},
{OPD(2, 0b01, 0x91), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPGATHER, OpSize::i64Bit>},
{OPD(2, 0b01, 0x92), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPGATHER, OpSize::i32Bit>},
{OPD(2, 0b01, 0x93), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPGATHER, OpSize::i64Bit>},
{OPD(2, 0b01, 0x96), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VFMAddSubImpl, true, 1, 3, 2>}, // VFMADDSUB
{OPD(2, 0b01, 0x97), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VFMAddSubImpl, false, 1, 3, 2>}, // VFMSUBADD
@@ -808,7 +808,7 @@ namespace AVX256 {
{OPD(2, 0b01, 0xB6), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VFMAddSubImpl, true, 2, 3, 1>}, // VFMADDSUB
{OPD(2, 0b01, 0xB7), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VFMAddSubImpl, false, 2, 3, 1>}, // VFMSUBADD
{OPD(2, 0b01, 0xDB), 1, &OpDispatchBuilder::AESImcOp},
{OPD(2, 0b01, 0xDB), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AESImcOp, true>},
{OPD(2, 0b01, 0xDC), 1, &OpDispatchBuilder::VAESEncOp},
{OPD(2, 0b01, 0xDD), 1, &OpDispatchBuilder::VAESEncLastOp},
{OPD(2, 0b01, 0xDE), 1, &OpDispatchBuilder::VAESDecOp},
@@ -820,10 +820,10 @@ namespace AVX256 {
{OPD(3, 0b01, 0x04), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPERMILImmOp, OpSize::i32Bit>},
{OPD(3, 0b01, 0x05), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPERMILImmOp, OpSize::i64Bit>},
{OPD(3, 0b01, 0x06), 1, &OpDispatchBuilder::VPERM2Op},
{OPD(3, 0b01, 0x08), 1, &OpDispatchBuilder::AVXVectorRound<OpSize::i32Bit>},
{OPD(3, 0b01, 0x09), 1, &OpDispatchBuilder::AVXVectorRound<OpSize::i64Bit>},
{OPD(3, 0b01, 0x0A), 1, &OpDispatchBuilder::AVXInsertScalarRound<OpSize::i32Bit>},
{OPD(3, 0b01, 0x0B), 1, &OpDispatchBuilder::AVXInsertScalarRound<OpSize::i64Bit>},
{OPD(3, 0b01, 0x08), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorRound, OpSize::i32Bit>},
{OPD(3, 0b01, 0x09), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorRound, OpSize::i64Bit>},
{OPD(3, 0b01, 0x0A), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXInsertScalarRound, OpSize::i32Bit>},
{OPD(3, 0b01, 0x0B), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXInsertScalarRound, OpSize::i64Bit>},
{OPD(3, 0b01, 0x0C), 1, &OpDispatchBuilder::VPBLENDDOp},
{OPD(3, 0b01, 0x0D), 1, &OpDispatchBuilder::VBLENDPDOp},
{OPD(3, 0b01, 0x0E), 1, &OpDispatchBuilder::VPBLENDWOp},
@@ -837,15 +837,15 @@ namespace AVX256 {
{OPD(3, 0b01, 0x18), 1, &OpDispatchBuilder::VINSERTOp},
{OPD(3, 0b01, 0x19), 1, &OpDispatchBuilder::VEXTRACT128Op},
{OPD(3, 0b01, 0x1D), 1, &OpDispatchBuilder::VCVTPS2PHOp},
{OPD(3, 0b01, 0x20), 1, &OpDispatchBuilder::VPINSRBOp},
{OPD(3, 0b01, 0x20), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPINSRBWOp, OpSize::i8Bit>},
{OPD(3, 0b01, 0x21), 1, &OpDispatchBuilder::VINSERTPSOp},
{OPD(3, 0b01, 0x22), 1, &OpDispatchBuilder::VPINSRDQOp},
{OPD(3, 0b01, 0x38), 1, &OpDispatchBuilder::VINSERTOp},
{OPD(3, 0b01, 0x39), 1, &OpDispatchBuilder::VEXTRACT128Op},
{OPD(3, 0b01, 0x40), 1, &OpDispatchBuilder::VDPPOp<OpSize::i32Bit>},
{OPD(3, 0b01, 0x41), 1, &OpDispatchBuilder::VDPPOp<OpSize::i64Bit>},
{OPD(3, 0b01, 0x40), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VDPPOp, OpSize::i32Bit>},
{OPD(3, 0b01, 0x41), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VDPPOp, OpSize::i64Bit>},
{OPD(3, 0b01, 0x42), 1, &OpDispatchBuilder::VMPSADBWOp},
{OPD(3, 0b01, 0x44), 1, &OpDispatchBuilder::VPCLMULQDQOp},
@@ -855,12 +855,12 @@ namespace AVX256 {
{OPD(3, 0b01, 0x4B), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorVariableBlend, OpSize::i64Bit>},
{OPD(3, 0b01, 0x4C), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AVXVectorVariableBlend, OpSize::i8Bit>},
{OPD(3, 0b01, 0x60), 1, &OpDispatchBuilder::VPCMPESTRMOp},
{OPD(3, 0b01, 0x61), 1, &OpDispatchBuilder::VPCMPESTRIOp},
{OPD(3, 0b01, 0x62), 1, &OpDispatchBuilder::VPCMPISTRMOp},
{OPD(3, 0b01, 0x63), 1, &OpDispatchBuilder::VPCMPISTRIOp},
{OPD(3, 0b01, 0x60), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPCMPESTRMOp, true>},
{OPD(3, 0b01, 0x61), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPCMPESTRIOp, true>},
{OPD(3, 0b01, 0x62), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPCMPISTRMOp, true>},
{OPD(3, 0b01, 0x63), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPCMPISTRIOp, true>},
{OPD(3, 0b01, 0xDF), 1, &OpDispatchBuilder::AESKeyGenAssist},
{OPD(3, 0b01, 0xDF), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AESKeyGenAssist, true>},
};
#undef OPD
@@ -392,8 +392,11 @@ namespace InstFlags {
constexpr InstFlagType FLAGS_REX_W_1 = (1ULL << 29);
constexpr InstFlagType FLAGS_CALL = (1ULL << 30);
constexpr InstFlagType FLAGS_SUPPORTS_LOCK = (1ULL << 31);
// Flags [57..32]: Undefined
// Flags [60..58]: Dst size
constexpr InstFlagType FLAGS_SIZE_DST_OFF = 58;
// Flags [63..61]: Src size
constexpr InstFlagType FLAGS_SIZE_SRC_OFF = FLAGS_SIZE_DST_OFF + 3;
constexpr InstFlagType SIZE_MASK = 0b111;
+2 -2
View File
@@ -643,7 +643,7 @@ public:
auto IROp = Node.GetNode(BaseList)->Op(IRList);
if (IROp->Op == OP_BEGINBLOCK) {
auto BeginBlock = IROp->C<IROp_EndBlock>();
auto BeginBlock = IROp->C<IROp_BeginBlock>();
Node = BeginBlock->BlockHeader;
} else if (IROp->Op == OP_CODEBLOCK) {
@@ -675,7 +675,7 @@ inline NodeID NodeWrapperBase<Type>::ID() const {
[[nodiscard]]
bool IsBlockExit(FEXCore::IR::IROps Op);
void Dump(fextl::stringstream* out, const IRListView* IR);
void Dump(fextl::ostringstream* out, const IRListView* IR);
constexpr auto format_as(FEXCore::IR::NodeID ID) {
return ID.Value;
+107 -64
View File
@@ -136,6 +136,7 @@
"u16": "uint16_t",
"u32": "uint32_t",
"u64": "uint64_t",
"c_str": "const char*",
"OpSize": "FEXCore::IR::OpSize",
"SSA": "OrderedNode*",
"GPR": "OrderedNode*",
@@ -240,6 +241,12 @@
"Desc": ["Debug operation that prints an SSA value to the console",
"May only print 64bits of the value"]
},
"PrintMsg c_str:$Value": {
"HasSideEffects": true,
"Desc": ["Debug operation that prints an string to the console.",
"This is for debug only! Will break code caching!"
]
},
"GPR = AllocateGPR i1:$ForPair": {
"Desc": ["Silly pseudo-instruction to allocate a register for a future destination",
"Note: if an instruction uses allocated destinations-as-sources,",
@@ -710,7 +717,7 @@
"HasSideEffects": true
},
"CacheLineClean GPR:$Addr": {
"Desc": ["Does a 64 byte cacheline cleanat the address specified",
"Desc": ["Does a 64 byte cacheline clean at the address specified",
"Only cleans the data cachelines. Doesn't do any zeroing",
"Skips the invalidation step of the CacheLineClear operation"
],
@@ -1857,9 +1864,9 @@
"DestSize": "RegisterSize",
"ElementSize": "ElementSize"
},
"FPR = VNot OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$Vector": {
"FPR = VNot OpSize:#RegisterSize, FPR:$Vector": {
"DestSize": "RegisterSize",
"ElementSize": "ElementSize"
"ElementSize": "OpSize::i8Bit"
},
"FPR = VAbs OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$Vector": {
@@ -1997,15 +2004,6 @@
"BitShift > 0"
]
},
"FPR = VUShraI OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$DestVector, FPR:$Vector, u8:$BitShift": {
"TiedSource": 0,
"DestSize": "RegisterSize",
"ElementSize": "ElementSize",
"EmitValidation": [
"ElementSize >= FEXCore::IR::OpSize::i8Bit && ElementSize <= FEXCore::IR::OpSize::i64Bit",
"BitShift > 0 && BitShift <= IR::OpSizeAsBits(ElementSize)"
]
},
"FPR = VSShrI OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$Vector, u8:$BitShift": {
"TiedSource": 0,
"DestSize": "RegisterSize",
@@ -2018,7 +2016,7 @@
"FPR = VUShrNI OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$Vector, u8:$BitShift": {
"TiedSource": 0,
"Desc": "Unsigned shifts right each element and then narrows to the next lower element size",
"Desc": ["Unsigned shifts right each element and then narrows to the next lower element size"],
"DestSize": "RegisterSize",
"ElementSize": "ElementSize >> 1",
"EmitValidation": [
@@ -2040,7 +2038,7 @@
]
},
"FPR = VSXTL OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$Vector": {
"Desc": "Sign extends elements from the source element size to the next size up",
"Desc": ["Sign extends elements from the source element size to the next size up"],
"DestSize": "RegisterSize",
"ElementSize": "ElementSize << 1"
},
@@ -2052,7 +2050,7 @@
"ElementSize": "ElementSize << 1"
},
"FPR = VSSHLL OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$Vector, u8:$BitShift{0}": {
"Desc": "Sign extends elements from the source element size to the next size up",
"Desc": ["Sign extends elements from the source element size to the next size up"],
"DestSize": "RegisterSize",
"ElementSize": "ElementSize << 1"
},
@@ -2064,7 +2062,7 @@
"ElementSize": "ElementSize << 1"
},
"FPR = VUXTL OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$Vector": {
"Desc": "Zero extends elements from the source element size to the next size up",
"Desc": ["Zero extends elements from the source element size to the next size up"],
"DestSize": "RegisterSize",
"ElementSize": "ElementSize << 1"
},
@@ -2146,43 +2144,55 @@
"ElementSize": "ElementSize"
},
"FPR = VAnd OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$Vector1, FPR:$Vector2": {
"FPR = VAnd OpSize:#RegisterSize, FPR:$Vector1, FPR:$Vector2": {
"DestSize": "RegisterSize",
"ElementSize": "ElementSize",
"ElementSize": "OpSize::i8Bit",
"EmitValidation": [
"RegisterSize == FEXCore::IR::OpSize::i256Bit || RegisterSize == FEXCore::IR::OpSize::i128Bit || RegisterSize == FEXCore::IR::OpSize::i64Bit"
]
},
"FPR = VAndn OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$Vector1, FPR:$Vector2": {
"FPR = VAndn OpSize:#RegisterSize, FPR:$Vector1, FPR:$Vector2": {
"DestSize": "RegisterSize",
"ElementSize": "ElementSize",
"ElementSize": "OpSize::i8Bit",
"EmitValidation": [
"RegisterSize == FEXCore::IR::OpSize::i256Bit || RegisterSize == FEXCore::IR::OpSize::i128Bit || RegisterSize == FEXCore::IR::OpSize::i64Bit"
]
},
"FPR = VOrn OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$Vector1, FPR:$Vector2": {
"FPR = VOrn OpSize:#RegisterSize, FPR:$Vector1, FPR:$Vector2": {
"DestSize": "RegisterSize",
"ElementSize": "ElementSize",
"ElementSize": "OpSize::i8Bit",
"EmitValidation": [
"RegisterSize == FEXCore::IR::OpSize::i256Bit || RegisterSize == FEXCore::IR::OpSize::i128Bit || RegisterSize == FEXCore::IR::OpSize::i64Bit"
]
},
"FPR = VOr OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$Vector1, FPR:$Vector2": {
"FPR = VOr OpSize:#RegisterSize, FPR:$Vector1, FPR:$Vector2": {
"DestSize": "RegisterSize",
"ElementSize": "ElementSize",
"ElementSize": "OpSize::i8Bit",
"EmitValidation": [
"RegisterSize == FEXCore::IR::OpSize::i256Bit || RegisterSize == FEXCore::IR::OpSize::i128Bit || RegisterSize == FEXCore::IR::OpSize::i64Bit"
]
},
"FPR = VXor OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$Vector1, FPR:$Vector2": {
"FPR = VXor OpSize:#RegisterSize, FPR:$Vector1, FPR:$Vector2": {
"DestSize": "RegisterSize",
"ElementSize": "OpSize::i8Bit",
"EmitValidation": [
"RegisterSize == FEXCore::IR::OpSize::i256Bit || RegisterSize == FEXCore::IR::OpSize::i128Bit || RegisterSize == FEXCore::IR::OpSize::i64Bit"
]
},
"FPR = VXar OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$LHS, FPR:$RHS, u8:$Rotate": {
"Desc": [
"Performs an XOR of corresponding elements and then rotates them right by",
"an amount between [1, ElementSize]"
],
"DestSize": "RegisterSize",
"ElementSize": "ElementSize",
"EmitValidation": [
"RegisterSize == FEXCore::IR::OpSize::i256Bit || RegisterSize == FEXCore::IR::OpSize::i128Bit || RegisterSize == FEXCore::IR::OpSize::i64Bit"
"RegisterSize == IR::OpSize::i256Bit || RegisterSize == IR::OpSize::i128Bit"
]
},
@@ -2207,7 +2217,7 @@
},
"FPR = VAddP OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$VectorLower, FPR:$VectorUpper": {
"Desc": "Does a horizontal pairwise add of elements across the two source vectors",
"Desc": ["Does a horizontal pairwise add of elements across the two source vectors"],
"DestSize": "RegisterSize",
"ElementSize": "ElementSize"
},
@@ -2264,7 +2274,7 @@
"ElementSize": "ElementSize"
},
"FPR = VFAddP OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$VectorLower, FPR:$VectorUpper": {
"Desc": "Does a horizontal pairwise add of elements across the two source vectors with float element types",
"Desc": ["Does a horizontal pairwise add of elements across the two source vectors with float element types"],
"DestSize": "RegisterSize",
"ElementSize": "ElementSize"
},
@@ -2314,29 +2324,27 @@
"ElementSize": "ElementSize << 1"
},
"FPR = VUMull2 OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$Vector1, FPR:$Vector2": {
"Desc": "Multiplies the high elements with size extension",
"Desc": ["Multiplies the high elements with size extension"],
"DestSize": "RegisterSize",
"ElementSize": "ElementSize << 1"
},
"FPR = VSMull2 OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$Vector1, FPR:$Vector2": {
"Desc": "Multiplies the high elements with size extension",
"Desc": ["Multiplies the high elements with size extension"],
"DestSize": "RegisterSize",
"ElementSize": "ElementSize << 1"
},
"FPR = VUMulH OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$Vector1, FPR:$Vector2": {
"Desc": "Wide unsigned multiply returning the high results",
"Desc": ["Wide unsigned multiply returning the high results"],
"DestSize": "RegisterSize",
"ElementSize": "ElementSize"
},
"FPR = VSMulH OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$Vector1, FPR:$Vector2": {
"Desc": "Wide signed multiply returning the high results",
"Desc": ["Wide signed multiply returning the high results"],
"DestSize": "RegisterSize",
"ElementSize": "ElementSize"
},
"FPR = VUABDL OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$Vector1, FPR:$Vector2": {
"Desc": ["Unsigned Absolute Difference Long"
],
"Desc": ["Unsigned Absolute Difference Long"],
"DestSize": "RegisterSize",
"ElementSize": "ElementSize << 1"
},
@@ -2557,6 +2565,24 @@
"DestSize": "RegisterSize",
"ElementSize": "ElementSize",
"TiedSource": 2
},
"FPR = VBlendImm OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$LHS, FPR:$RHS, u16:$Selector": {
"Desc": [
"Functions the same way an immediate blend operation on x86 would.",
"That is: (e.g. using 16-bit elements)",
" if (Selector[0] == 1)",
" Dst[15:0] = RHS[15:0]",
" else",
" Dst[15:0] = LHS[15:0]",
" <etc for the rest of the elements along the vector>",
"",
"Note that like x86, due to the selector size, the operation of this IR op",
"uses a 128-bit lane granularity, so each blending selector independently operates",
"on each 128-bit element that composes the vector."
],
"DestSize": "RegisterSize",
"ElementSize": "ElementSize",
"TiedSource": 0
}
},
"Conv": {
@@ -2594,7 +2620,7 @@
},
"FPR = Vector_SToF OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$Vector": {
"Desc": "Vector op: Converts signed integer to same size float",
"Desc": ["Vector op: Converts signed integer to same size float"],
"DestSize": "RegisterSize",
"ElementSize": "ElementSize"
},
@@ -2606,12 +2632,12 @@
"ElementSize": "ElementSize"
},
"FPR = Vector_FToZS OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$Vector": {
"Desc": "Vector op: Converts float to signed integer, rounding towards zero",
"Desc": ["Vector op: Converts float to signed integer, rounding towards zero"],
"DestSize": "RegisterSize",
"ElementSize": "ElementSize"
},
"FPR = Vector_FToF OpSize:#RegisterSize, OpSize:#DestElementSize, FPR:$Vector, OpSize:$SrcElementSize": {
"Desc": "Vector op: Converts float from source element size to destination size (fp32<->fp64)",
"Desc": ["Vector op: Converts float from source element size to destination size (fp32<->fp64)"],
"DestSize": "RegisterSize",
"ElementSize": "DestElementSize"
},
@@ -2666,75 +2692,74 @@
},
"Crypto": {
"FPR = VAESImc FPR:$Vector": {
"Desc": "Does a stage of the inverse mix column transformation",
"Desc": ["Does a stage of the inverse mix column transformation"],
"DestSize": "OpSize::i128Bit"
},
"FPR = VAESEnc OpSize:#RegisterSize, FPR:$State, FPR:$Key, FPR:$ZeroReg": {
"Desc": "Does a step of AES encryption",
"Desc": ["Does a step of AES encryption"],
"DestSize": "RegisterSize"
},
"FPR = VAESEncLast OpSize:#RegisterSize, FPR:$State, FPR:$Key, FPR:$ZeroReg": {
"Desc": "Does the last step of AES encryption",
"Desc": ["Does the last step of AES encryption"],
"DestSize": "RegisterSize"
},
"FPR = VAESDec OpSize:#RegisterSize, FPR:$State, FPR:$Key, FPR:$ZeroReg": {
"Desc": "Does a step of AES decryption",
"Desc": ["Does a step of AES decryption"],
"DestSize": "RegisterSize"
},
"FPR = VAESDecLast OpSize:#RegisterSize, FPR:$State, FPR:$Key, FPR:$ZeroReg": {
"Desc": "Does the last step of AES decryption",
"Desc": ["Does the last step of AES decryption"],
"DestSize": "RegisterSize"
},
"FPR = VAESKeyGenAssist FPR:$Src, FPR:$KeyGenTBLSwizzle, FPR:$ZeroReg, u8:$RCON": {
"Desc": "Assists in key generation",
"Desc": ["Assists in key generation"],
"DestSize": "OpSize::i128Bit"
},
"FPR = VSha1H FPR:$Src": {
"Desc": "Does vector scalar SHA1H instruction",
"Desc": ["Does vector scalar SHA1H instruction"],
"DestSize": "FEXCore::IR::OpSize::i32Bit"
},
"FPR = VSha1C FPR:$Src1, FPR:$Src2, FPR:$Src3": {
"Desc": "Does vector SHA1C instruction",
"Desc": ["Does vector SHA1C instruction"],
"DestSize": "FEXCore::IR::OpSize::i128Bit",
"TiedSource": 0
},
"FPR = VSha1M FPR:$Src1, FPR:$Src2, FPR:$Src3": {
"Desc": "Does vector SHA1M instruction",
"Desc": ["Does vector SHA1M instruction"],
"DestSize": "FEXCore::IR::OpSize::i128Bit",
"TiedSource": 0
},
"FPR = VSha1P FPR:$Src1, FPR:$Src2, FPR:$Src3": {
"Desc": "Does vector SHA1P instruction",
"Desc": ["Does vector SHA1P instruction"],
"DestSize": "FEXCore::IR::OpSize::i128Bit",
"TiedSource": 0
},
"FPR = VSha1SU1 FPR:$Src1, FPR:$Src2": {
"Desc": "Does vector scalar SHA1H instruction",
"Desc": ["Does vector scalar SHA1H instruction"],
"DestSize": "FEXCore::IR::OpSize::i128Bit",
"TiedSource": 0
},
"FPR = VSha256U0 FPR:$Src1, FPR:$Src2": {
"Desc": "Does vector scalar VSha256U0 instruction",
"Desc": ["Does vector scalar VSha256U0 instruction"],
"DestSize": "FEXCore::IR::OpSize::i128Bit",
"TiedSource": 0
},
"FPR = VSha256U1 FPR:$Src1, FPR:$Src2": {
"Desc": "Does vector scalar VSha256U1 instruction",
"Desc": ["Does vector scalar VSha256U1 instruction"],
"DestSize": "FEXCore::IR::OpSize::i128Bit"
},
"FPR = VSha256H FPR:$Src1, FPR:$Src2, FPR:$Src3": {
"Desc": "Does vector scalar VSha256H instruction",
"Desc": ["Does vector scalar VSha256H instruction"],
"DestSize": "FEXCore::IR::OpSize::i128Bit",
"TiedSource": 0
},
"FPR = VSha256H2 FPR:$Src1, FPR:$Src2, FPR:$Src3": {
"Desc": "Does vector scalar VSha256H2 instruction",
"Desc": ["Does vector scalar VSha256H2 instruction"],
"DestSize": "FEXCore::IR::OpSize::i128Bit",
"TiedSource": 0
},
"GPR = CRC32 GPR:$Src1, GPR:$Src2, OpSize:$SrcSize": {
"Desc": ["CRC32 using polynomial 0x1EDC6F41"
],
"Desc": ["CRC32 using polynomial 0x1EDC6F41"],
"DestSize": "OpSize::i32Bit"
},
"FPR = PCLMUL OpSize:#RegisterSize, FPR:$Src1, FPR:$Src2, u8:$Selector": {
@@ -2751,39 +2776,43 @@
"F64": {
"FPR = F64ATAN FPR:$Src1, FPR:$Src2": {
"DestSize": "OpSize::i64Bit",
"JITDispatch": false
"JITDispatch": true
},
"FPR = F64FPREM FPR:$Src1, FPR:$Src2": {
"DestSize": "OpSize::i64Bit",
"JITDispatch": false
"JITDispatch": true
},
"FPR = F64FPREM1 FPR:$Src1, FPR:$Src2": {
"DestSize": "OpSize::i64Bit",
"JITDispatch": false
"JITDispatch": true
},
"FPR = F64SCALE FPR:$Src1, FPR:$Src2": {
"DestSize": "OpSize::i64Bit",
"JITDispatch": false
"JITDispatch": true
},
"FPR = F64F2XM1 FPR:$Src": {
"DestSize": "OpSize::i64Bit",
"JITDispatch": false
"JITDispatch": true
},
"FPR = F64FYL2X FPR:$Src, FPR:$Src2": {
"DestSize": "OpSize::i64Bit",
"JITDispatch": false
"JITDispatch": true
},
"FPR = F64FYL2XP1 FPR:$Src, FPR:$Src2": {
"DestSize": "OpSize::i64Bit",
"JITDispatch": true
},
"FPR = F64TAN FPR:$Src": {
"DestSize": "OpSize::i64Bit",
"JITDispatch": false
"JITDispatch": true
},
"FPR = F64SIN FPR:$Src": {
"DestSize": "OpSize::i64Bit",
"JITDispatch": false
"JITDispatch": true
},
"FPR = F64COS FPR:$Src": {
"DestSize": "OpSize::i64Bit",
"JITDispatch": false
"JITDispatch": true
},
"FPR:$Sin, FPR:$Cos = F64SINCOS FPR:$Src": {
"DestSize": "OpSize::i64Bit",
@@ -3201,6 +3230,20 @@
"DestSize": "OpSize::i128Bit",
"JITDispatch": false
},
"FPR = F80FYL2XP1Stack": {
"Desc": [
"Computes ST1 * log2(1 + ST0)",
"Stores the result in ST1, and pops the top of the stack.",
"Returns the new value at the top of the stack, i.e. the result of the operation."
],
"HasSideEffects": true,
"DestSize": "OpSize::i128Bit",
"X87": true
},
"FPR = F80FYL2XP1 FPR:$X80Src1, FPR:$X80Src2": {
"DestSize": "OpSize::i128Bit",
"JITDispatch": false
},
"F80VBSLStack OpSize:#RegisterSize, FPR:$VectorMask, u8:$SrcStack1, u8:$SrcStack2": {
"Desc": [
"Does a vector bitwise select.",
+28 -20
View File
@@ -30,15 +30,19 @@ namespace FEXCore::IR {
#include <FEXCore/IR/IRDefines.inc>
static void PrintArg(fextl::stringstream* out, const IRListView*, const SHA256Sum& Arg) {
static void PrintArg(fextl::ostringstream* out, const IRListView*, const SHA256Sum& Arg) {
*out << fextl::fmt::format("sha256:{:02x}", fmt::join(Arg.data, ""));
}
static void PrintArg(fextl::stringstream* out, const IRListView*, uint64_t Arg) {
static void PrintArg(fextl::ostringstream* out, const IRListView*, uint64_t Arg) {
*out << fextl::fmt::format("#{:#x}", Arg);
}
static void PrintArg(fextl::stringstream* out, const IRListView*, CondClass Arg) {
static void PrintArg(fextl::ostringstream* out, const IRListView*, const char* const Arg) {
*out << fextl::fmt::format("'{}'", Arg);
}
static void PrintArg(fextl::ostringstream* out, const IRListView*, CondClass Arg) {
if (Arg == CondClass::AL) {
*out << "ALWAYS";
return;
@@ -51,7 +55,7 @@ static void PrintArg(fextl::stringstream* out, const IRListView*, CondClass Arg)
*out << CondNames[FEXCore::ToUnderlying(Arg)];
}
static void PrintArg(fextl::stringstream* out, const IRListView*, MemOffsetType Arg) {
static void PrintArg(fextl::ostringstream* out, const IRListView*, MemOffsetType Arg) {
static constexpr std::array<std::string_view, 3> Names = {
"SXTX",
"UXTW",
@@ -61,7 +65,7 @@ static void PrintArg(fextl::stringstream* out, const IRListView*, MemOffsetType
*out << Names[FEXCore::ToUnderlying(Arg)];
}
static void PrintArg(fextl::stringstream* out, const IRListView*, RegClass Arg) {
static void PrintArg(fextl::ostringstream* out, const IRListView*, RegClass Arg) {
*out << [Arg] {
switch (Arg) {
case RegClass::Invalid: return "Invalid";
@@ -75,7 +79,7 @@ static void PrintArg(fextl::stringstream* out, const IRListView*, RegClass Arg)
}();
}
static void PrintArg(fextl::stringstream* out, const IRListView* IR, OrderedNodeWrapper Arg) {
static void PrintArg(fextl::ostringstream* out, const IRListView* IR, OrderedNodeWrapper Arg) {
if (Arg.IsImmediate()) {
auto PhyReg = PhysicalRegister(Arg);
@@ -124,7 +128,7 @@ static void PrintArg(fextl::stringstream* out, const IRListView* IR, OrderedNode
}
}
static void PrintArg(fextl::stringstream* out, const IRListView*, FenceType Arg) {
static void PrintArg(fextl::ostringstream* out, const IRListView*, FenceType Arg) {
*out << [Arg] {
switch (Arg) {
case FenceType::Load: return "Loads";
@@ -136,7 +140,7 @@ static void PrintArg(fextl::stringstream* out, const IRListView*, FenceType Arg)
}();
}
static void PrintArg(fextl::stringstream* out, const IRListView*, RoundMode Arg) {
static void PrintArg(fextl::ostringstream* out, const IRListView*, RoundMode Arg) {
*out << [Arg] {
switch (Arg) {
case RoundMode::Nearest: return "Nearest";
@@ -149,7 +153,7 @@ static void PrintArg(fextl::stringstream* out, const IRListView*, RoundMode Arg)
}();
}
static void PrintArg(fextl::stringstream* out, const IRListView*, ConstPad Arg) {
static void PrintArg(fextl::ostringstream* out, const IRListView*, ConstPad Arg) {
*out << [Arg] {
switch (Arg) {
case ConstPad::NoPad: return "NoPad";
@@ -160,7 +164,7 @@ static void PrintArg(fextl::stringstream* out, const IRListView*, ConstPad Arg)
}();
}
static void PrintArg(fextl::stringstream* out, const IRListView*, NamedVectorConstant Arg) {
static void PrintArg(fextl::ostringstream* out, const IRListView*, NamedVectorConstant Arg) {
*out << [Arg] {
// clang-format off
switch (Arg) {
@@ -204,6 +208,10 @@ static void PrintArg(fextl::stringstream* out, const IRListView*, NamedVectorCon
return "movmaskb";
case NamedVectorConstant::NAMED_VECTOR_MOVMASKB_UPPER:
return "movmaskb_upper";
case NamedVectorConstant::NAMED_VECTOR_256_MID_ELEMENT_SWAP:
return "v256_mid_element_swap";
case NamedVectorConstant::NAMED_VECTOR_256_MID_ELEMENT_SWAP_UPPER:
return "v256_mid_element_swap_upper";
case NamedVectorConstant::NAMED_VECTOR_ZERO:
return "vectorzero";
case NamedVectorConstant::NAMED_VECTOR_X87_ONE:
@@ -252,7 +260,7 @@ static void PrintArg(fextl::stringstream* out, const IRListView*, NamedVectorCon
}();
}
static void PrintArg(fextl::stringstream* out, const IRListView*, IndexNamedVectorConstant Arg) {
static void PrintArg(fextl::ostringstream* out, const IRListView*, IndexNamedVectorConstant Arg) {
*out << [Arg] {
// clang-format off
switch (Arg) {
@@ -278,7 +286,7 @@ static void PrintArg(fextl::stringstream* out, const IRListView*, IndexNamedVect
}();
}
static void PrintArg(fextl::stringstream* out, const IRListView*, OpSize Arg) {
static void PrintArg(fextl::ostringstream* out, const IRListView*, OpSize Arg) {
*out << [Arg] {
switch (Arg) {
case OpSize::iUnsized: return "Unsized";
@@ -295,7 +303,7 @@ static void PrintArg(fextl::stringstream* out, const IRListView*, OpSize Arg) {
}();
}
static void PrintArg(fextl::stringstream* out, const IRListView*, FloatCompareOp Arg) {
static void PrintArg(fextl::ostringstream* out, const IRListView*, FloatCompareOp Arg) {
*out << [Arg] {
switch (Arg) {
case FloatCompareOp::EQ: return "FEQ";
@@ -309,14 +317,14 @@ static void PrintArg(fextl::stringstream* out, const IRListView*, FloatCompareOp
}();
}
static void PrintArg(fextl::stringstream* out, const IRListView*, FEXCore::IR::BreakDefinition Arg) {
static void PrintArg(fextl::ostringstream* out, const IRListView*, FEXCore::IR::BreakDefinition Arg) {
*out << "{" << Arg.ErrorRegister << ".";
*out << static_cast<uint32_t>(Arg.Signal) << ".";
*out << static_cast<uint32_t>(Arg.TrapNumber) << ".";
*out << static_cast<uint32_t>(Arg.si_code) << "}";
}
static void PrintArg(fextl::stringstream* out, const IRListView*, ShiftType Arg) {
static void PrintArg(fextl::ostringstream* out, const IRListView*, ShiftType Arg) {
*out << [Arg] {
switch (Arg) {
case ShiftType::LSL: return "LSL";
@@ -328,7 +336,7 @@ static void PrintArg(fextl::stringstream* out, const IRListView*, ShiftType Arg)
}();
}
static void PrintArg(fextl::stringstream* out, const IRListView*, BranchHint Arg) {
static void PrintArg(fextl::ostringstream* out, const IRListView*, BranchHint Arg) {
*out << [Arg] {
switch (Arg) {
case BranchHint::None: return "None";
@@ -340,11 +348,11 @@ static void PrintArg(fextl::stringstream* out, const IRListView*, BranchHint Arg
}();
}
static void PrintArg(fextl::stringstream* out, const IRListView*, const std::array<uint8_t, 0x10>& Arg) {
static void PrintArg(fextl::ostringstream* out, const IRListView*, const std::array<uint8_t, 0x10>& Arg) {
*out << fextl::fmt::format("{:02x}", fmt::join(Arg, ""));
}
void Dump(fextl::stringstream* out, const IRListView* IR) {
void Dump(fextl::ostringstream* out, const IRListView* IR) {
auto HeaderOp = IR->GetHeader();
int8_t CurrentIndent = 0;
@@ -356,8 +364,8 @@ void Dump(fextl::stringstream* out, const IRListView* IR) {
++CurrentIndent;
AddIndent();
*out << fextl::fmt::format("(%0) IRHeader %{}, #{:#x}, #{}, #{}\n", HeaderOp->Blocks.ID(), HeaderOp->OriginalRIP, HeaderOp->BlockCount,
HeaderOp->NumHostInstructions);
*out << fextl::fmt::format("(%0) IRHeader %{}, #{:#x}, #{}, #{}\n", HeaderOp->Blocks.ID(), +HeaderOp->OriginalRIP, +HeaderOp->BlockCount,
+HeaderOp->NumHostInstructions);
for (auto [BlockNode, BlockHeader] : IR->GetBlocks()) {
{
+3 -4
View File
@@ -33,7 +33,7 @@ bool IsBlockExit(FEXCore::IR::IROps Op) {
}
}
RegClass IREmitter::WalkFindRegClass(Ref Node) {
RegClass IREmitter::WalkFindRegClass(Ref Node) const {
auto Class = GetOpRegClass(Node);
switch (Class) {
case RegClass::GPR:
@@ -45,9 +45,8 @@ RegClass IREmitter::WalkFindRegClass(Ref Node) {
}
// Complex case, needs to be handled on an op by op basis
uintptr_t DataBegin = DualListData.DataBegin();
FEXCore::IR::IROp_Header* IROp = Node->Op(DataBegin);
const uintptr_t DataBegin = DualListData.DataBegin();
const auto* IROp = Node->Op(DataBegin);
switch (IROp->Op) {
case IROps::OP_LOADREGISTER: {
+16 -14
View File
@@ -21,15 +21,15 @@ class IREmitter {
public:
IREmitter(FEXCore::Utils::IntrusivePooledAllocator& ThreadAllocator, bool SupportsTSOImm9)
: DualListData {ThreadAllocator, 8 * 1024 * 1024}
, SupportsTSOImm9(SupportsTSOImm9) {
ReownOrClaimBuffer();
ResetWorkingList();
}
, SupportsTSOImm9(SupportsTSOImm9) {}
virtual ~IREmitter() = default;
void ReownOrClaimBuffer() {
DualListData.ReownOrClaimBuffer();
// Reset the working list on new buffer.
ResetWorkingList();
}
void DelayedDisownBuffer() {
@@ -39,14 +39,13 @@ public:
IRListView ViewIR() {
return IRListView(&DualListData);
}
void ResetWorkingList();
/**
* @name IR allocation routines
*
* @{ */
RegClass WalkFindRegClass(Ref Node);
RegClass WalkFindRegClass(Ref Node) const;
// These inlining helpers are used by IRDefines.inc so define first.
Ref InlineMem(OpSize Size, Ref Offset, MemOffsetType OffsetType, uint8_t& OffsetScale, bool TSO = false) {
@@ -311,14 +310,14 @@ public:
}
/** @} */
RegClass WalkFindRegClass(OrderedNodeWrapper ssa) {
Ref RealNode = ssa.GetNode(DualListData.ListBegin());
RegClass WalkFindRegClass(OrderedNodeWrapper ssa) const {
auto RealNode = ssa.GetNode(DualListData.ListBegin());
return WalkFindRegClass(RealNode);
}
bool IsValueConstant(OrderedNodeWrapper ssa, uint64_t* Constant = nullptr) {
Ref RealNode = ssa.GetNode(DualListData.ListBegin());
FEXCore::IR::IROp_Header* IROp = RealNode->Op(DualListData.DataBegin());
bool IsValueConstant(OrderedNodeWrapper ssa, uint64_t* Constant = nullptr) const {
auto RealNode = ssa.GetNode(DualListData.ListBegin());
const auto* IROp = RealNode->Op(DualListData.DataBegin());
if (IROp->Op == OP_CONSTANT) {
auto Op = IROp->C<IR::IROp_Constant>();
if (Constant) {
@@ -329,9 +328,9 @@ public:
return false;
}
bool IsValueInlineConstant(OrderedNodeWrapper ssa) {
Ref RealNode = ssa.GetNode(DualListData.ListBegin());
FEXCore::IR::IROp_Header* IROp = RealNode->Op(DualListData.DataBegin());
bool IsValueInlineConstant(OrderedNodeWrapper ssa) const {
auto RealNode = ssa.GetNode(DualListData.ListBegin());
const auto* IROp = RealNode->Op(DualListData.DataBegin());
if (IROp->Op == OP_INLINECONSTANT) {
return true;
}
@@ -512,6 +511,9 @@ protected:
fextl::vector<Ref> CodeBlocks;
uint64_t Entry {};
bool SupportsTSOImm9 {};
private:
void ResetWorkingList();
};
} // namespace FEXCore::IR
@@ -119,11 +119,7 @@ class DualIntrusiveAllocatorThreadPool final : public DualIntrusiveAllocator {
public:
DualIntrusiveAllocatorThreadPool(FEXCore::Utils::IntrusivePooledAllocator& ThreadAllocator, size_t Size)
: DualIntrusiveAllocator {Size}
, PoolObject {ThreadAllocator, Size * 2} {
// Claim a buffer on allocation
PoolObject.ReownOrClaimBuffer();
}
, PoolObject {ThreadAllocator, Size * 2} {}
void ReownOrClaimBuffer() {
Data = PoolObject.ReownOrClaimBuffer();
List = Data + MemorySize;
@@ -190,12 +186,12 @@ public:
}
[[nodiscard]]
unsigned PostRA() const {
bool PostRA() const {
return GetHeader()->PostRA;
}
[[nodiscard]]
unsigned SpillSlots() const {
uint32_t SpillSlots() const {
return GetHeader()->SpillSlots;
}
+34 -5
View File
@@ -13,6 +13,7 @@ $end_info$
#include "Interface/IR/Passes/RegisterAllocationPass.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/Profiler.h>
namespace FEXCore::IR {
@@ -66,25 +67,42 @@ void PassManager::Finalize() {
}
}
void PassManager::AddDefaultPasses(FEXCore::Context::ContextImpl* ctx) {
void PassManager::AddDefaultPasses(Context::ContextImpl* ctx) {
FEX_CONFIG_OPT(DisablePasses, O0);
// We only specifically disable optimization passes if desired, as IR output should
// still be well-formed regardless of the modifications made to it.
if (!DisablePasses()) {
InsertPass(CreateX87StackOptimizationPass(ctx->HostFeatures, ctx->Config.Is64BitMode ? IR::OpSize::i64Bit : IR::OpSize::i32Bit));
InsertPass(CreateDeadFlagCalculationEliminination());
}
}
void PassManager::AddDefaultValidationPasses() {
InsertPass(IR::CreateRegisterAllocationPass(&ctx->CPUID), "RA");
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
InsertValidationPass(Validation::CreateIRValidation(), "IRValidation");
#endif
}
void PassManager::InsertRegisterAllocationPass(FEXCore::Context::ContextImpl* ctx) {
InsertPass(IR::CreateRegisterAllocationPass(&ctx->CPUID), "RA");
Pass* PassManager::InsertPass(fextl::unique_ptr<Pass> Pass, const fextl::string& Name) {
auto* PassPtr = InsertAt(Passes.end(), std::move(Pass))->get();
AttemptNameMapping(Name, PassPtr);
return PassPtr;
}
PassManager::PassArrayType::iterator PassManager::InsertAt(PassArrayType::iterator pos, fextl::unique_ptr<Pass> Pass) {
Pass->RegisterPassManager(this);
return Passes.insert(pos, std::move(Pass));
}
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
void PassManager::InsertValidationPass(fextl::unique_ptr<Pass> Pass, const fextl::string& Name) {
Pass->RegisterPassManager(this);
auto* PassPtr = ValidationPasses.emplace_back(std::move(Pass)).get();
AttemptNameMapping(Name, PassPtr);
}
#endif
void PassManager::Run(IREmitter* IREmit) {
FEXCORE_PROFILE_SCOPED("PassManager::Run");
@@ -98,4 +116,15 @@ void PassManager::Run(IREmitter* IREmit) {
}
#endif
}
void PassManager::AttemptNameMapping(const fextl::string& Name, Pass* NewPass) {
if (Name.empty()) {
// Empty name is a 'don't care' case. e.g. Passes that just need to run,
// but don't need to be actively looked up.
return;
}
const auto Result = NameToPassMaping.emplace(Name, NewPass);
LOGMAN_THROW_A_FMT(Result.second, "Tried to insert pass with name '{}'. But name is already used", Name);
}
} // namespace FEXCore::IR
+31 -43
View File
@@ -8,23 +8,18 @@ $end_info$
#pragma once
#include <FEXCore/Config/Config.h>
#include <FEXCore/Utils/ThreadPoolAllocator.h>
#include <FEXCore/fextl/memory.h>
#include <FEXCore/fextl/string.h>
#include <FEXCore/fextl/unordered_map.h>
#include <FEXCore/fextl/vector.h>
#include <functional>
#include <concepts>
#include <utility>
namespace FEXCore::Context {
class ContextImpl;
}
namespace FEXCore::HLE {
class SyscallHandler;
}
namespace FEXCore::IR {
class PassManager;
class IREmitter;
@@ -44,64 +39,57 @@ protected:
class PassManager final {
public:
void AddDefaultPasses(FEXCore::Context::ContextImpl* ctx);
void AddDefaultValidationPasses();
Pass* InsertPass(fextl::unique_ptr<Pass> Pass, fextl::string Name = "") {
auto PassPtr = InsertAt(Passes.end(), std::move(Pass))->get();
if (!Name.empty()) {
NameToPassMaping[Name] = PassPtr;
}
return PassPtr;
explicit PassManager(Context::ContextImpl* CTX) {
AddDefaultPasses(CTX);
}
void InsertRegisterAllocationPass(FEXCore::Context::ContextImpl* ctx);
// Executes all of the passes added to the manager.
// If assertions are enabled, this will also run all validation passes.
void Run(IREmitter* IREmit);
bool HasPass(fextl::string Name) const {
// Inserts a new pass into the manager, optionally also assigning a name to it
// for use in the lookup functions,
Pass* InsertPass(fextl::unique_ptr<Pass> Pass, const fextl::string& Name = "");
// Whether or not a pass with the given name is within the manager.
bool HasPass(const fextl::string& Name) const {
return NameToPassMaping.contains(Name);
}
template<typename T>
T* GetPass(fextl::string Name) {
return dynamic_cast<T*>(NameToPassMaping[Name]);
// Retrieves a pass from the manager that has the given name assigned to it.
// Will return nullptr if the pass doesn't exist.
template<std::derived_from<Pass> T>
T* GetPass(const fextl::string& Name) {
return dynamic_cast<T*>(GetPass(Name));
}
Pass* GetPass(fextl::string Name) {
return NameToPassMaping[Name];
}
void RegisterSyscallHandler(FEXCore::HLE::SyscallHandler* Handler) {
SyscallHandler = Handler;
Pass* GetPass(const fextl::string& Name) {
const auto Iter = NameToPassMaping.find(Name);
if (Iter == NameToPassMaping.end()) {
return nullptr;
}
return Iter->second;
}
// Finalizes the pass manager state and assumes no other passes will be added after called.
// This will reorganize the execution order of the passes if necessary.
void Finalize();
protected:
FEXCore::HLE::SyscallHandler* SyscallHandler {};
private:
void AddDefaultPasses(Context::ContextImpl* ctx);
using PassArrayType = fextl::vector<fextl::unique_ptr<Pass>>;
PassArrayType::iterator InsertAt(PassArrayType::iterator pos, fextl::unique_ptr<Pass> Pass) {
Pass->RegisterPassManager(this);
return Passes.insert(pos, std::move(Pass));
}
PassArrayType::iterator InsertAt(PassArrayType::iterator pos, fextl::unique_ptr<Pass> Pass);
PassArrayType Passes;
fextl::unordered_map<fextl::string, Pass*> NameToPassMaping;
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
fextl::vector<fextl::unique_ptr<Pass>> ValidationPasses;
void InsertValidationPass(fextl::unique_ptr<Pass> Pass, fextl::string Name = "") {
Pass->RegisterPassManager(this);
auto PassPtr = ValidationPasses.emplace_back(std::move(Pass)).get();
if (!Name.empty()) {
NameToPassMaping[Name] = PassPtr;
}
}
void InsertValidationPass(fextl::unique_ptr<Pass> Pass, const fextl::string& Name = "");
#endif
void AttemptNameMapping(const fextl::string& Name, Pass* NewPass);
FEX_CONFIG_OPT(Is64BitMode, IS64BIT_MODE);
FEX_CONFIG_OPT(PassManagerDumpIR, PASSMANAGERDUMPIR);
};
+5 -10
View File
@@ -8,23 +8,18 @@ class CPUIDEmu;
struct HostFeatures;
} // namespace FEXCore
namespace FEXCore::Utils {
class IntrusivePooledAllocator;
}
namespace FEXCore::IR {
class Pass;
class RegisterAllocationPass;
fextl::unique_ptr<FEXCore::IR::Pass> CreateDeadFlagCalculationEliminination();
fextl::unique_ptr<FEXCore::IR::RegisterAllocationPass> CreateRegisterAllocationPass(const FEXCore::CPUIDEmu* CPUID);
fextl::unique_ptr<FEXCore::IR::Pass> CreateX87StackOptimizationPass(const FEXCore::HostFeatures&, OpSize GPROpSize);
fextl::unique_ptr<Pass> CreateDeadFlagCalculationEliminination();
fextl::unique_ptr<Pass> CreateRegisterAllocationPass(const CPUIDEmu* CPUID);
fextl::unique_ptr<Pass> CreateX87StackOptimizationPass(const HostFeatures&, OpSize GPROpSize);
namespace Validation {
fextl::unique_ptr<FEXCore::IR::Pass> CreateIRValidation();
fextl::unique_ptr<Pass> CreateIRValidation();
} // namespace Validation
namespace Debug {
fextl::unique_ptr<FEXCore::IR::Pass> CreateIRDumper();
fextl::unique_ptr<Pass> CreateIRDumper();
}
} // namespace FEXCore::IR
@@ -51,23 +51,23 @@ void IRDumper::Run(IREmitter* IREmit) {
// DumpIRStr might be no if not dumping but ShouldDump is set in OpDisp
if (DumpToFile) {
const auto fileName = fextl::fmt::format("{}/{:x}{}", DumpIR(), HeaderOp->OriginalRIP, IR.PostRA() ? "-post.ir" : "-pre.ir");
const auto fileName = fextl::fmt::format("{}/{:x}{}", DumpIR(), +HeaderOp->OriginalRIP, IR.PostRA() ? "-post.ir" : "-pre.ir");
FD = FEXCore::File::File(fileName.c_str(),
FEXCore::File::FileModes::WRITE | FEXCore::File::FileModes::CREATE | FEXCore::File::FileModes::TRUNCATE);
}
if (FD.IsValid() || DumpToLog) {
fextl::stringstream out;
fextl::ostringstream out;
FEXCore::IR::Dump(&out, &IR);
if (FD.IsValid()) {
fextl::fmt::print(FD, "IR-{} 0x{:x}:\n{}\n@@@@@\n", IR.PostRA() ? "post" : "pre", HeaderOp->OriginalRIP, out.str());
fextl::fmt::print(FD, "IR-{} 0x{:x}:\n{}\n@@@@@\n", IR.PostRA() ? "post" : "pre", +HeaderOp->OriginalRIP, out.str());
} else {
LogMan::Msg::IFmt("IR-{} 0x{:x}:\n{}\n@@@@@\n", IR.PostRA() ? "post" : "pre", HeaderOp->OriginalRIP, out.str());
LogMan::Msg::IFmt("IR-{} 0x{:x}:\n{}\n@@@@@\n", IR.PostRA() ? "post" : "pre", +HeaderOp->OriginalRIP, out.str());
}
}
}
fextl::unique_ptr<FEXCore::IR::Pass> CreateIRDumper() {
fextl::unique_ptr<Pass> CreateIRDumper() {
return fextl::make_unique<IRDumper>();
}
} // namespace FEXCore::IR::Debug
@@ -10,6 +10,7 @@ $end_info$
#include "Interface/IR/IREmitter.h"
#include "Interface/IR/PassManager.h"
#include "Interface/IR/RegisterAllocationData.h"
#include "Interface/IR/Passes.h"
#include "Interface/IR/Passes/IRValidation.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
@@ -46,7 +47,7 @@ void IRValidation::Run(IREmitter* IREmit) {
OffsetToBlockMap.clear();
EntryBlock = nullptr;
uint32_t Count = CurrentIR.GetSSACount();
const auto Count = CurrentIR.GetSSACount();
if (Count > MaxNodes) {
NodeIsLive.Realloc(Count);
}
@@ -59,7 +60,7 @@ void IRValidation::Run(IREmitter* IREmit) {
#endif
for (auto [BlockNode, BlockHeader] : CurrentIR.GetBlocks()) {
auto BlockIROp = BlockHeader->CW<FEXCore::IR::IROp_CodeBlock>();
auto BlockIROp = BlockHeader->C<FEXCore::IR::IROp_CodeBlock>();
LOGMAN_THROW_A_FMT(BlockIROp->Header.Op == OP_CODEBLOCK, "IR type failed to be a code block");
if (!EntryBlock) {
@@ -77,15 +78,15 @@ void IRValidation::Run(IREmitter* IREmit) {
const auto OpSize = IROp->Size;
if (GetHasDest(IROp->Op)) {
HadError |= OpSize == IR::OpSize::iInvalid;
// Does the op have a destination of size 0?
// Does the op have an unsized destination?
if (OpSize == IR::OpSize::iInvalid) {
HadError = true;
Errors << "%" << ID << ": Had destination but with no size" << std::endl;
}
// Does the node have zero uses? Should have been DCE'd
if (CodeNode->GetUses() == 0) {
HadWarning |= true;
HadWarning = true;
Warnings << "%" << ID << ": Destination created but had no uses" << std::endl;
}
@@ -98,27 +99,26 @@ void IRValidation::Run(IREmitter* IREmit) {
// If no register class was assigned
if (AssignedClass == IR::RegClass::Invalid) {
HadError |= true;
HadError = true;
Errors << "%" << ID << ": Had destination but with no register class assigned" << std::endl;
}
// If no physical register was assigned
if (PhyReg.IsInvalid()) {
HadError |= true;
HadError = true;
Errors << "%" << ID << ": Had destination but with no register assigned" << std::endl;
}
// Assigned class wasn't the expected class and it is a non-complex op
if (AssignedClass != ExpectedClass && ExpectedClass != IR::RegClass::Complex) {
HadWarning |= true;
HadWarning = true;
Warnings << "%" << ID << ": Destination had register class " << uint32_t(AssignedClass) << " When register class "
<< uint32_t(ExpectedClass) << " Was expected" << std::endl;
}
}
}
uint8_t NumArgs = IR::GetRAArgs(IROp->Op);
const uint8_t NumArgs = IR::GetRAArgs(IROp->Op);
for (uint32_t i = 0; i < NumArgs; ++i) {
OrderedNodeWrapper Arg = IROp->Args[i];
const auto ArgID = Arg.ID();
@@ -126,8 +126,6 @@ void IRValidation::Run(IREmitter* IREmit) {
continue;
}
IROps Op = CurrentIR.GetOp<IROp_Header>(Arg)->Op;
if (ArgID.IsValid()) {
Uses[ArgID.Value]++;
}
@@ -135,10 +133,11 @@ void IRValidation::Run(IREmitter* IREmit) {
// We do not validate the location of inline constants because it's
// irrelevant, they're ignored by RA and always inlined to where they
// need to be. This lets us pool inline constants globally.
bool Ignore = (Op == OP_IRHEADER || Op == OP_INLINECONSTANT);
const IROps Op = CurrentIR.GetOp<IROp_Header>(Arg)->Op;
const bool Ignore = (Op == OP_IRHEADER || Op == OP_INLINECONSTANT);
if (!Ignore && ArgID.IsValid() && !NodeIsLive.Get(ArgID.Value)) {
HadError |= true;
HadError = true;
Errors << "%" << ID << ": Arg[" << i << "] references invalid %" << ArgID << std::endl;
}
}
@@ -147,7 +146,6 @@ void IRValidation::Run(IREmitter* IREmit) {
switch (IROp->Op) {
case IR::OP_EXITFUNCTION: {
CurrentBlock->HasExit = true;
break;
}
case IR::OP_CONDJUMP: {
@@ -163,7 +161,7 @@ void IRValidation::Run(IREmitter* IREmit) {
const FEXCore::IR::IROp_Header* FalseTargetOp = CurrentIR.GetOp<IROp_Header>(FalseTargetNode);
if (TrueTargetOp->Op != OP_CODEBLOCK) {
HadError |= true;
HadError = true;
Errors << "CondJump %" << ID << ": True Target Jumps to Op that isn't the begining of a block" << std::endl;
} else {
auto Block = OffsetToBlockMap.try_emplace(Op->TrueBlock.ID()).first;
@@ -171,7 +169,7 @@ void IRValidation::Run(IREmitter* IREmit) {
}
if (FalseTargetOp->Op != OP_CODEBLOCK) {
HadError |= true;
HadError = true;
Errors << "CondJump %" << ID << ": False Target Jumps to Op that isn't the begining of a block" << std::endl;
} else {
auto Block = OffsetToBlockMap.try_emplace(Op->FalseBlock.ID()).first;
@@ -187,7 +185,7 @@ void IRValidation::Run(IREmitter* IREmit) {
const FEXCore::IR::IROp_Header* TargetOp = CurrentIR.GetOp<IROp_Header>(TargetNode);
if (TargetOp->Op != OP_CODEBLOCK) {
HadError |= true;
HadError = true;
Errors << "Jump %" << ID << ": Jump to Op that isn't the begining of a block" << std::endl;
} else {
auto Block = OffsetToBlockMap.try_emplace(Op->Header.Args[0].ID()).first;
@@ -204,7 +202,7 @@ void IRValidation::Run(IREmitter* IREmit) {
// Blocks can only have zero (Exit), 1 (Unconditional branch) or 2 (Conditional) successors
size_t NumSuccessors = CurrentBlock->Successors.size();
if (NumSuccessors > 2) {
HadError |= true;
HadError = true;
Errors << "%" << BlockID << " Has " << NumSuccessors << " successors which is too many" << std::endl;
}
@@ -220,7 +218,7 @@ void IRValidation::Run(IREmitter* IREmit) {
{
auto Op = GetOp(CodeCurrent);
if (Op != IR::OP_ENDBLOCK) {
HadError |= true;
HadError = true;
Errors << "%" << BlockID << " Failed to end block with EndBlock" << std::endl;
}
}
@@ -231,7 +229,7 @@ void IRValidation::Run(IREmitter* IREmit) {
{
auto Op = GetOp(CodeCurrent);
if (!IsBlockExit(Op)) {
HadError |= true;
HadError = true;
Errors << "%" << BlockID << " Didn't have a block exit IR op as its last instruction" << std::endl;
}
}
@@ -243,7 +241,7 @@ void IRValidation::Run(IREmitter* IREmit) {
for (uint32_t i = 0; i < CurrentIR.GetSSACount(); i++) {
auto [Node, IROp] = CurrentIR.at(IR::NodeID {i})();
if (Node->NumUses != Uses[i] && IROp->Op != OP_CODEBLOCK && IROp->Op != OP_IRHEADER) {
HadError |= true;
HadError = true;
Errors << "%" << i << " Has " << Uses[i] << " Uses, but reports " << Node->NumUses << std::endl;
}
}
@@ -251,7 +249,7 @@ void IRValidation::Run(IREmitter* IREmit) {
HadWarning = false;
if (HadError || HadWarning) {
fextl::stringstream Out;
fextl::ostringstream Out;
FEXCore::IR::Dump(&Out, &CurrentIR);
if (HadError) {
@@ -271,7 +269,7 @@ void IRValidation::Run(IREmitter* IREmit) {
}
}
fextl::unique_ptr<FEXCore::IR::Pass> CreateIRValidation() {
fextl::unique_ptr<Pass> CreateIRValidation() {
return fextl::make_unique<IRValidation>();
}
} // namespace FEXCore::IR::Validation
@@ -8,20 +8,16 @@
namespace FEXCore::IR::Validation {
struct BlockInfo {
bool HasExit;
const OrderedNode* BlockNode;
fextl::vector<OrderedNode*> Predecessors;
fextl::vector<OrderedNode*> Successors;
};
class IRValidation final : public FEXCore::IR::Pass {
public:
~IRValidation();
void Run(IREmitter* IREmit) override;
private:
struct BlockInfo {
fextl::vector<OrderedNode*> Predecessors;
fextl::vector<OrderedNode*> Successors;
};
BitSet<uint64_t> NodeIsLive {};
OrderedNode* EntryBlock {};
@@ -7,6 +7,7 @@ $end_info$
#include "Interface/IR/IR.h"
#include "Interface/IR/IREmitter.h"
#include "Interface/IR/Passes.h"
#include "Interface/IR/PassManager.h"
#include <FEXCore/Core/X86Enums.h>
@@ -196,7 +197,7 @@ unsigned DeadFlagCalculationEliminination::FlagsForCondClassType(CondClass Cond)
}
}
constexpr FlagInfo ClassifyConst(IROps Op) {
static constexpr FlagInfo ClassifyConst(IROps Op) {
switch (Op) {
case OP_ANDWITHFLAGS:
return FlagInfo::Pack({
@@ -332,15 +333,15 @@ constexpr FlagInfo ClassifyConst(IROps Op) {
}
}
constexpr auto FlagInfos = std::invoke([] {
constexpr auto FlagInfos = [] {
std::array<FlagInfo, OP_LAST> ret = {};
for (unsigned i = 0; i < OP_LAST; ++i) {
ret[i] = ClassifyConst((IROps)i);
ret[i] = ClassifyConst(IROps(i));
}
return ret;
});
}();
FlagInfo DeadFlagCalculationEliminination::Classify(IROp_Header* IROp) {
FlagInfo Info = FlagInfos[IROp->Op];
@@ -351,22 +352,22 @@ FlagInfo DeadFlagCalculationEliminination::Classify(IROp_Header* IROp) {
switch (IROp->Op) {
case OP_NZCVSELECT:
case OP_NZCVSELECTINCREMENT: {
auto Op = IROp->CW<IR::IROp_NZCVSelect>();
auto Op = IROp->C<IR::IROp_NZCVSelect>();
return FlagInfo::Pack({.Read = FlagsForCondClassType(Op->Cond)});
}
case OP_NZCVSELECTV: {
auto Op = IROp->CW<IR::IROp_NZCVSelectV>();
auto Op = IROp->C<IR::IROp_NZCVSelectV>();
return FlagInfo::Pack({.Read = FlagsForCondClassType(Op->Cond)});
}
case OP_NEG: {
auto Op = IROp->CW<IR::IROp_Neg>();
auto Op = IROp->C<IR::IROp_Neg>();
return FlagInfo::Pack({.Read = FlagsForCondClassType(Op->Cond)});
}
case OP_CONDJUMP: {
auto Op = IROp->CW<IR::IROp_CondJump>();
auto Op = IROp->C<IR::IROp_CondJump>();
if (!Op->FromNZCV) {
return FlagInfo::Pack({});
}
@@ -376,7 +377,7 @@ FlagInfo DeadFlagCalculationEliminination::Classify(IROp_Header* IROp) {
case OP_CONDSUBNZCV:
case OP_CONDADDNZCV: {
auto Op = IROp->CW<IR::IROp_CondAddNZCV>();
auto Op = IROp->C<IR::IROp_CondAddNZCV>();
return FlagInfo::Pack({
.Read = FlagsForCondClassType(Op->Cond),
.Write = FLAG_NZCV,
@@ -385,7 +386,7 @@ FlagInfo DeadFlagCalculationEliminination::Classify(IROp_Header* IROp) {
}
case OP_RMIFNZCV: {
auto Op = IROp->CW<IR::IROp_RmifNZCV>();
auto Op = IROp->C<IR::IROp_RmifNZCV>();
static_assert(FLAG_N == (1 << 3), "rmif mask lines up with our bits");
static_assert(FLAG_Z == (1 << 2), "rmif mask lines up with our bits");
@@ -399,7 +400,7 @@ FlagInfo DeadFlagCalculationEliminination::Classify(IROp_Header* IROp) {
}
case OP_INVALIDATEFLAGS: {
auto Op = IROp->CW<IR::IROp_InvalidateFlags>();
auto Op = IROp->C<IR::IROp_InvalidateFlags>();
unsigned Flags = 0;
// TODO: Make this translation less silly
@@ -536,7 +537,7 @@ bool DeadFlagCalculationEliminination::ProcessBlock(IREmitter* IREmit, IRListVie
// Initialize the FlagsRead mask according to the exit instruction.
auto [ExitNode, ExitOp] = CodeLast();
if (ExitOp->Op == IR::OP_CONDJUMP) {
auto Op = ExitOp->CW<IR::IROp_CondJump>();
auto Op = ExitOp->C<IR::IROp_CondJump>();
FlagsRead = CFG.Get(Op->TrueBlock)->Flags | CFG.Get(Op->FalseBlock)->Flags;
} else if (ExitOp->Op == IR::OP_JUMP) {
FlagsRead = CFG.Get(ExitOp->Args[0])->Flags;
@@ -643,7 +644,7 @@ void DeadFlagCalculationEliminination::OptimizeParity(IREmitter* IREmit, IRListV
for (auto [CodeNode, IROp] : CurrentIR.GetCode(Block)) {
if (IROp->Op == OP_STOREPF) {
auto Op = IROp->CW<IR::IROp_StorePF>();
auto Op = IROp->C<IR::IROp_StorePF>();
auto Generator = CurrentIR.GetOp<IR::IROp_Header>(Op->Value);
// Determine if we only write 0/1 to the parity flag.
@@ -696,7 +697,7 @@ void DeadFlagCalculationEliminination::Run(IREmitter* IREmit) {
--CodeLast;
auto [ExitNode, ExitOp] = CodeLast();
if (ExitOp->Op == IR::OP_CONDJUMP) {
auto Op = ExitOp->CW<IR::IROp_CondJump>();
auto Op = ExitOp->C<IR::IROp_CondJump>();
CFG.RecordEdge(Block->ID, Op->TrueBlock);
CFG.RecordEdge(Block->ID, Op->FalseBlock);
@@ -747,7 +748,7 @@ void DeadFlagCalculationEliminination::Run(IREmitter* IREmit) {
}
}
fextl::unique_ptr<FEXCore::IR::Pass> CreateDeadFlagCalculationEliminination() {
fextl::unique_ptr<Pass> CreateDeadFlagCalculationEliminination() {
return fextl::make_unique<DeadFlagCalculationEliminination>();
}
@@ -33,7 +33,7 @@ namespace {
Ref RegToSSA[32];
};
IR::RegClass GetRegClassFromNode(IR::IRListView* IR, IR::IROp_Header* IROp) {
IR::RegClass GetRegClassFromNode(const IR::IROp_Header* IROp) {
const auto Class = IR::GetRegClass(IROp->Op);
if (Class != IR::RegClass::Complex) {
return Class;
@@ -49,9 +49,14 @@ namespace {
case IR::OP_FILLREGISTER: return IROp->C<IR::IROp_FillRegister>()->Class;
default: return IR::RegClass::Invalid;
}
};
}
} // Anonymous namespace
void RegisterAllocationPass::SetNumPairRegs(uint32_t NumRegs) {
LOGMAN_THROW_A_FMT((NumRegs % 2) == 0, "Number of pair regs must be even. (Given: {})", NumRegs);
PairRegs = NumRegs;
}
class ConstrainedRAPass final : public RegisterAllocationPass {
public:
explicit ConstrainedRAPass(const FEXCore::CPUIDEmu* CPUID)
@@ -85,27 +90,27 @@ private:
// SourcesNextUses is read backwards, this tracks the index
int64_t SourceIndex {};
bool Rematerializable(IROp_Header* IROp) {
static bool Rematerializable(const IROp_Header* IROp) {
return IROp->Op == OP_CONSTANT;
}
Ref InsertFill(Ref Node) {
IROp_Header* IROp = IR->GetOp<IROp_Header>(Node);
const auto* IROp = IR->GetOp<IROp_Header>(Node);
// Remat if we can
if (Rematerializable(IROp)) {
const auto Op = IROp->C<IR::IROp_Constant>();
uint64_t Const = Op->Constant;
const auto* Op = IROp->C<IR::IROp_Constant>();
const uint64_t Const = Op->Constant;
return IREmit->_Constant(Const, Op->Pad, Op->MaxBytes);
}
// Otherwise fill from stack
uint32_t SlotPlusOne = SpillSlots[IR->GetID(Node).Value];
const uint32_t SlotPlusOne = SpillSlots[IR->GetID(Node).Value];
LOGMAN_THROW_A_FMT(SlotPlusOne >= 1, "Node must have been spilled");
const auto RegClass = GetRegClassFromNode(IR, IROp);
const auto RegClass = GetRegClassFromNode(IROp);
return IREmit->_FillRegister(IROp->Size, IROp->ElementSize, SlotPlusOne - 1, RegClass);
};
}
// IP of next-use of each source. IPs are measured from the end of the
// block, so we don't need to size the block up-front.
@@ -113,32 +118,35 @@ private:
bool AnySpilled {};
bool IsValidArg(OrderedNodeWrapper Arg) {
bool IsValidArg(OrderedNodeWrapper Arg) const {
if (Arg.IsInvalid()) {
return false;
}
auto Op = IR->GetOp<IROp_Header>(Arg)->Op;
return Op != OP_INLINECONSTANT && Op != OP_INLINEENTRYPOINTOFFSET;
};
}
RegisterClassData* GetClass(PhysicalRegister Reg) {
return &Classes[Reg.Class];
};
}
const RegisterClassData* GetClass(PhysicalRegister Reg) const {
return &Classes[Reg.Class];
}
uint32_t GetRegBits(PhysicalRegister Reg) {
return 1 << Reg.Reg;
};
static uint32_t GetRegBits(PhysicalRegister Reg) {
return 1U << Reg.Reg;
}
bool IsInRegisterFile(Ref Node) {
bool IsInRegisterFile(Ref Node) const {
auto ID = IR->GetID(Node).Value;
LOGMAN_THROW_A_FMT(ID < SSAToReg.size(), "Only old nodes looked up");
PhysicalRegister Reg = SSAToReg[ID];
RegisterClassData* Class = GetClass(Reg);
const PhysicalRegister Reg = SSAToReg[ID];
const RegisterClassData* Class = GetClass(Reg);
return (Class->Available & GetRegBits(Reg)) == 0 && Class->RegToSSA[Reg.Reg] == Node;
};
}
void FreeReg(PhysicalRegister Reg) {
RegisterClassData* Class = GetClass(Reg);
@@ -147,7 +155,7 @@ private:
LOGMAN_THROW_A_FMT(!(Class->Available & RegBits), "Register double-free");
Class->Available |= RegBits;
};
}
bool HasSource(IROp_Header* I, PhysicalRegister Reg) {
int NumArgs = IR::GetRAArgs(I->Op);
@@ -170,13 +178,13 @@ private:
}
return false;
};
}
Ref DecodeSRANode(const IROp_Header* IROp, Ref Node) {
if (IROp->Op == OP_LOADREGISTER || IROp->Op == OP_LOADPF || IROp->Op == OP_LOADAF) {
return Node;
} else if (IROp->Op == OP_STOREREGISTER) {
auto V = IROp->C<IR::IROp_StorePF>()->Value;
auto V = IROp->C<IR::IROp_StoreRegister>()->Value;
V.ClearKill();
return IR->GetNode(V);
} else if (IROp->Op == OP_STOREPF || IROp->Op == OP_STOREAF) {
@@ -186,9 +194,9 @@ private:
}
return nullptr;
};
}
PhysicalRegister DecodeSRAReg(const IROp_Header* IROp, Ref Node) {
PhysicalRegister DecodeSRAReg(const IROp_Header* IROp, Ref Node) const {
uint8_t FlagOffset = Classes[FEXCore::ToUnderlying(RegClass::GPRFixed)].Count - 2;
if (IROp->Op == OP_STOREREGISTER) {
@@ -207,9 +215,9 @@ private:
return PhysicalRegister {RegClass::GPRFixed, uint8_t(Op->Reg)};
}
}
};
}
bool IsTrivial(Ref Node, const IROp_Header* Header) {
bool IsTrivial(Ref Node, const IROp_Header* Header) const {
switch (Header->Op) {
case OP_ALLOCATEGPR: return true;
case OP_ALLOCATEGPRAFTER: return true;
@@ -320,7 +328,7 @@ private:
// If we already spilled the Candidate, we don't need to spill again.
// Similarly, if we can rematerialize the instruction, we don't spill it.
if (!Spilled && Header->Op != OP_CONSTANT) {
LOGMAN_THROW_A_FMT(Reg.AsRegClass() == GetRegClassFromNode(IR, Header), "Consistent");
LOGMAN_THROW_A_FMT(Reg.AsRegClass() == GetRegClassFromNode(Header), "Consistent");
// SpillSlots allocation is deferred.
if (SpillSlots.empty()) {
@@ -340,7 +348,7 @@ private:
// Now that we've spilled the value, take it out of the register file
FreeReg(Reg);
AnySpilled = true;
};
}
void RemapReg(Ref Node, PhysicalRegister Reg) {
RegisterClassData* Class = GetClass(Reg);
@@ -350,7 +358,7 @@ private:
if (Index < SSAToReg.size()) {
SSAToReg[Index] = Reg;
}
};
}
// Record a given assignment of register Reg to Node.
void SetReg(Ref Node, PhysicalRegister Reg) {
@@ -363,7 +371,7 @@ private:
RemapReg(Node, Reg);
Node->Reg = Reg.Raw;
};
}
// Assign a register for a given Node, spilling if necessary.
void AssignReg(IROp_Header* IROp, IROp_CodeBlock* Block, Ref CodeNode, IROp_Header* Pivot) {
@@ -419,7 +427,7 @@ private:
}
}
RegClass ClassType = GetRegClassFromNode(IR, IROp);
RegClass ClassType = GetRegClassFromNode(IROp);
RegisterClassData* Class = &Classes[FEXCore::ToUnderlying(ClassType)];
// Spill to make room in the register file.
@@ -432,7 +440,7 @@ private:
LOGMAN_THROW_A_FMT(Class->Available != 0, "Post-condition of spilling");
unsigned Reg = std::countr_zero(Class->Available);
SetReg(CodeNode, PhysicalRegister(ClassType, Reg));
};
}
};
void ConstrainedRAPass::AddRegisters(IR::RegClass Class, uint32_t RegisterCount) {
@@ -441,7 +449,7 @@ void ConstrainedRAPass::AddRegisters(IR::RegClass Class, uint32_t RegisterCount)
Classes[FEXCore::ToUnderlying(Class)].Count = RegisterCount;
}
inline bool KillMove(IROp_Header* LastOp, IROp_Header* IROp, Ref LastNode, Ref CodeNode) {
static bool KillMove(const IROp_Header* LastOp, IROp_Header* IROp, Ref LastNode, Ref CodeNode) {
// 32-bit moves in x86_64 are represented as a Bfe, detect them.
if (LastOp->Op == OP_BFE && LastOp->C<IR::IROp_Bfe>()->lsb == 0 && LastOp->C<IR::IROp_Bfe>()->Width == 32) {
auto Op = IROp->Op;
@@ -459,7 +467,7 @@ inline bool KillMove(IROp_Header* LastOp, IROp_Header* IROp, Ref LastNode, Ref C
return LastOp->Op == OP_STOREREGISTER;
}
inline bool IsSignext(const IROp_Header* IROp, OrderedNodeWrapper Src, OpSize Size) {
static bool IsSignext(const IROp_Header* IROp, OrderedNodeWrapper Src, OpSize Size) {
if (IROp->Op == OP_SBFE) {
auto Sbfe = IROp->C<IR::IROp_Sbfe>();
return Sbfe->Width == 1 && Sbfe->lsb == (IR::OpSizeAsBits(Size) - 1) && Sbfe->Src == Src;
@@ -468,7 +476,7 @@ inline bool IsSignext(const IROp_Header* IROp, OrderedNodeWrapper Src, OpSize Si
}
}
inline bool IsZero(const IROp_Header* IROp) {
static bool IsZero(const IROp_Header* IROp) {
return IROp->Op == OP_CONSTANT && IROp->C<IROp_Constant>()->Constant == 0;
}
@@ -781,7 +789,7 @@ void ConstrainedRAPass::Run(IREmitter* IREmit_) {
IR->GetHeader()->PostRA = true;
}
fextl::unique_ptr<IR::RegisterAllocationPass> CreateRegisterAllocationPass(const FEXCore::CPUIDEmu* CPUID) {
fextl::unique_ptr<IR::Pass> CreateRegisterAllocationPass(const CPUIDEmu* CPUID) {
return fextl::make_unique<ConstrainedRAPass>(CPUID);
}
} // namespace FEXCore::IR
@@ -6,10 +6,11 @@ $end_info$
*/
#pragma once
#include "Interface/IR/PassManager.h"
#include <cstdint>
#include <memory>
#include <stdint.h>
namespace FEXCore::IR {
enum class RegClass : uint32_t;
@@ -18,6 +19,9 @@ class RegisterAllocationPass : public FEXCore::IR::Pass {
public:
virtual void AddRegisters(RegClass Class, uint32_t RegisterCount) = 0;
void SetNumPairRegs(uint32_t NumRegs);
protected:
// Number of GPRs usable for pairs at start of GPR set. Must be even.
uint32_t PairRegs {};
};
@@ -3,6 +3,7 @@
#include "Interface/Core/Interpreter/Fallbacks/FallbackOpHandler.h"
#include "Interface/IR/IR.h"
#include "Interface/IR/IREmitter.h"
#include "Interface/IR/Passes.h"
#include "Interface/IR/PassManager.h"
#include "FEXCore/IR/IR.h"
#include "FEXCore/Utils/Profiler.h"
@@ -32,14 +33,14 @@
namespace FEXCore::IR {
// FIXME(pmatos): copy from OpcodeDispatcher.h
inline uint32_t MMBaseOffset() {
static uint32_t MMBaseOffset() {
return static_cast<uint32_t>(offsetof(Core::CPUState, mm[0][0]));
}
// Similar helper to the one in OpcodeDispatcher.h except we do not
// need to handle flags, etc.
template<typename T>
void DeriveOp(Ref& RefV, IROps NewOp, IREmitter::IRPair<T> Expr) {
static void DeriveOp(Ref& RefV, IROps NewOp, IREmitter::IRPair<T> Expr) {
Expr.first->Header.Op = NewOp;
RefV = Expr;
}
@@ -52,8 +53,8 @@ template<typename T>
class FixedSizeStack {
public:
struct StackSlotEntry final {
StackSlot Type;
T Value;
StackSlot Type = StackSlot::UNUSED;
T Value = T::Invalid;
};
static constexpr uint8_t size = 8;
@@ -64,8 +65,7 @@ public:
// If SlowPath is true, then TopOffset is always zero.
int8_t TopOffset = 0;
FixedSizeStack()
: buffer(FixedSizeStack::size, {StackSlot::UNUSED, T::Invalid}) {}
FixedSizeStack() = default;
void push(const T& Value) {
rotate();
@@ -92,25 +92,23 @@ public:
return buffer[Offset];
}
void setTop(T Value, size_t Offset = 0) {
void setTop(const T& Value, size_t Offset = 0) {
buffer[Offset] = {StackSlot::VALID, Value};
}
bool isValid(size_t Offset) const {
return buffer[Offset].first;
return buffer[Offset].Type == StackSlot::VALID;
}
void clear() {
for (auto& Elem : buffer) {
Elem = {StackSlot::UNUSED, T::Invalid};
}
buffer.fill({StackSlot::UNUSED, T::Invalid});
TopOffset = 0;
}
void dump() const {
LogMan::Msg::DFmt("-- Stack");
for (size_t i = 0; i < 8; i++) {
for (size_t i = 0; i < buffer.size(); i++) {
const auto& [Valid, Element] = buffer[i];
if (Valid == StackSlot::VALID) {
LogMan::Msg::DFmt("| ST{}: 0x{:x}", i, (uintptr_t)(Element.StackDataNode));
@@ -126,7 +124,7 @@ public:
}
// Returns a mask to set in AbridgedTagWord
uint8_t getValidMask() {
uint8_t getValidMask() const {
uint8_t Mask = 0;
for (size_t i = 0; i < buffer.size(); i++) {
if (buffer[i].Type == StackSlot::VALID) {
@@ -137,7 +135,7 @@ public:
}
// Returns a mask to set in AbridgedTagWord
uint8_t getInvalidMask() {
uint8_t getInvalidMask() const {
uint8_t Mask = 0;
for (size_t i = 0; i < buffer.size(); i++) {
if (buffer[i].Type == StackSlot::INVALID) {
@@ -148,7 +146,7 @@ public:
}
private:
fextl::vector<StackSlotEntry> buffer;
std::array<StackSlotEntry, size> buffer {};
};
class X87StackOptimization final : public Pass {
@@ -188,7 +186,7 @@ private:
void Store80BitToMem(const IROp_StoreStackMem* Op, Ref StackNode, Ref AddrNode, Ref Offset, OpSize Align, MemOffsetType OffsetType,
uint8_t OffsetScale) {
if (Features.SupportsSVE128 || Features.SupportsSVE256) {
if (Features.SupportsSVE()) {
AddressMode A {.Base = AddrNode,
.Index = Op->Offset.IsInvalid() ? nullptr : Offset,
.IndexType = MemOffsetType::SXTX,
@@ -201,11 +199,11 @@ private:
}
}
void StoreStackMem_Helper(const IROp_StoreStackMem* Op, Ref StackNode) {
void StoreStackMem_Helper(const IRListView& IR, const IROp_StoreStackMem* Op, Ref StackNode) {
LOGMAN_THROW_A_FMT(!ReducedPrecisionMode, "Full precision mode expected.");
Ref AddrNode = IR->GetNode(Op->Addr);
Ref Offset = IR->GetNode(Op->Offset);
Ref AddrNode = IR.GetNode(Op->Addr);
Ref Offset = IR.GetNode(Op->Offset);
OpSize Align = Op->Align;
MemOffsetType OffsetType = Op->OffsetType;
uint8_t OffsetScale = Op->OffsetScale;
@@ -229,11 +227,11 @@ private:
// Performs a store to memory from a value the stack passed in as StackNode.
// This is the version dealing with the reduced precision case.
void StoreStackMem_Reduced_Helper(const IROp_StoreStackMem* Op, Ref StackNode) {
void StoreStackMem_Reduced_Helper(const IRListView& IR, const IROp_StoreStackMem* Op, Ref StackNode) {
LOGMAN_THROW_A_FMT(ReducedPrecisionMode, "Reduced precision mode expected.");
Ref AddrNode = IR->GetNode(Op->Addr);
Ref Offset = IR->GetNode(Op->Offset);
Ref AddrNode = IR.GetNode(Op->Addr);
Ref Offset = IR.GetNode(Op->Offset);
OpSize Align = Op->Align;
MemOffsetType OffsetType = Op->OffsetType;
uint8_t OffsetScale = Op->OffsetScale;
@@ -292,10 +290,10 @@ private:
void Reset();
struct StackMemberInfo {
StackMemberInfo() = delete;
StackMemberInfo(Ref Data)
constexpr StackMemberInfo() = default;
constexpr StackMemberInfo(Ref Data)
: StackDataNode(Data) {}
StackMemberInfo(Ref Data, Ref Source, OpSize Size)
constexpr StackMemberInfo(Ref Data, Ref Source, OpSize Size)
: StackDataNode(Data)
, Source({Size, Source}) {}
Ref StackDataNode {}; // Reference to the data in the Stack.
@@ -357,9 +355,9 @@ private:
// On the slow path TopCache is always the last obtained version of top.
// TopOffset is ignored
bool SlowPath = false;
// Keeping IREmitter not to pass arguments around
IREmitter* IREmit = nullptr;
IRListView* IR = nullptr;
};
inline const X87StackOptimization::StackMemberInfo X87StackOptimization::StackMemberInfo::Invalid {nullptr};
@@ -576,24 +574,34 @@ void X87StackOptimization::HandleBinopStack(IROps Op64, bool VFOp64, IROps Op80,
}
inline void X87StackOptimization::UpdateTopForPop_Slow() {
const auto PopContainer = [](auto& container) {
const auto begin = std::begin(container);
std::rotate(begin, std::next(begin), std::end(container));
};
// Pop the top of the x87 stack
GetOffsetTopWithCache_Slow(1);
std::rotate(TopOffsetCache.begin(), std::next(TopOffsetCache.begin()), TopOffsetCache.end());
std::rotate(TopOffsetAddressCache.begin(), std::next(TopOffsetAddressCache.begin()), TopOffsetAddressCache.end());
std::rotate(TopValueCache.begin(), std::next(TopValueCache.begin()), TopValueCache.end());
std::rotate(FlushValuesPending.begin(), std::next(FlushValuesPending.begin()), FlushValuesPending.end());
std::rotate(TopValidCache.begin(), std::next(TopValidCache.begin()), TopValidCache.end());
PopContainer(TopOffsetCache);
PopContainer(TopOffsetAddressCache);
PopContainer(TopValueCache);
PopContainer(FlushValuesPending);
PopContainer(TopValidCache);
FlushTopPending = true;
}
inline void X87StackOptimization::UpdateTopForPush_Slow() {
// Pop the top of the x87 stack
const auto PushContainer = [](auto& container) {
const auto end = std::end(container);
std::rotate(std::begin(container), std::prev(end), end);
};
// Push the top of the x87 stack
GetOffsetTopWithCache_Slow(1, true);
std::rotate(TopOffsetCache.begin(), std::prev(TopOffsetCache.end()), TopOffsetCache.end());
std::rotate(TopOffsetAddressCache.begin(), std::prev(TopOffsetAddressCache.end()), TopOffsetAddressCache.end());
std::rotate(TopValueCache.begin(), std::prev(TopValueCache.end()), TopValueCache.end());
std::rotate(FlushValuesPending.begin(), std::prev(FlushValuesPending.end()), FlushValuesPending.end());
std::rotate(TopValidCache.begin(), std::prev(TopValidCache.end()), TopValidCache.end());
PushContainer(TopOffsetCache);
PushContainer(TopOffsetAddressCache);
PushContainer(TopValueCache);
PushContainer(FlushValuesPending);
PushContainer(TopValidCache);
FlushTopPending = true;
}
@@ -724,7 +732,6 @@ void X87StackOptimization::Run(IREmitter* Emit) {
// Initialize IREmit member
IREmit = Emit;
IR = &CurrentIR;
// Run optimization proper
for (auto [BlockNode, BlockHeader] : CurrentIR.GetBlocks()) {
@@ -785,6 +792,12 @@ void X87StackOptimization::Run(IREmitter* Emit) {
break;
}
case OP_F80FYL2XP1STACK: {
HandleBinopStack(OP_F64FYL2XP1, false, OP_F80FYL2XP1, 1, 0, 1);
StackPop();
break;
}
case OP_F80ATANSTACK: {
HandleBinopStack(OP_F64ATAN, false, OP_F80ATAN, 1, 1, 0);
StackPop();
@@ -862,25 +875,20 @@ void X87StackOptimization::Run(IREmitter* Emit) {
Ref SinValue {};
Ref CosValue {};
if (ReducedPrecisionMode) {
SinValue = IREmit->_F64SIN(St0);
CosValue = IREmit->_F64COS(St0);
}
#ifdef VIXL_SIMULATOR
if (DisableVixlIndirectCalls() == 0) {
if (ReducedPrecisionMode) {
SinValue = IREmit->_F64SIN(St0);
CosValue = IREmit->_F64COS(St0);
} else {
SinValue = IREmit->_F80SIN(St0);
CosValue = IREmit->_F80COS(St0);
}
} else
else if (DisableVixlIndirectCalls() == 0) {
SinValue = IREmit->_F80SIN(St0);
CosValue = IREmit->_F80COS(St0);
}
#endif
{
else {
SinValue = IREmit->_AllocateFPR(OpSize::i128Bit, OpSize::i128Bit);
CosValue = IREmit->_AllocateFPR(OpSize::i128Bit, OpSize::i128Bit);
if (ReducedPrecisionMode) {
IREmit->_F64SINCOS(St0, SinValue, CosValue);
} else {
IREmit->_F80SINCOS(St0, SinValue, CosValue);
}
IREmit->_F80SINCOS(St0, SinValue, CosValue);
}
// Push values
@@ -932,7 +940,6 @@ void X87StackOptimization::Run(IREmitter* Emit) {
UpdateTopForPush_Slow();
StoreStackValueAtOffset_Slow(SourceNode);
} else {
auto* SourceNode = CurrentIR.GetNode(Op->X80Src);
if (Op->OriginalValue.IsInvalid()) {
// No original value to track - just push the converted data
StackData.push(StackMemberInfo {SourceNode});
@@ -1018,11 +1025,11 @@ void X87StackOptimization::Run(IREmitter* Emit) {
}
if (ReducedPrecisionMode) {
StoreStackMem_Reduced_Helper(Op, StackNode);
StoreStackMem_Reduced_Helper(CurrentIR, Op, StackNode);
break;
}
StoreStackMem_Helper(Op, StackNode);
StoreStackMem_Helper(CurrentIR, Op, StackNode);
break;
}
@@ -1086,7 +1093,7 @@ void X87StackOptimization::Run(IREmitter* Emit) {
ResultNode = IREmit->_VFNeg(OpSize::i64Bit, OpSize::i64Bit, Value);
} else {
Ref HelperNode = IREmit->_LoadNamedVectorConstant(OpSize::i128Bit, IR::NamedVectorConstant::NAMED_VECTOR_F80_SIGN_MASK);
ResultNode = IREmit->_VXor(OpSize::i128Bit, OpSize::i8Bit, Value, HelperNode);
ResultNode = IREmit->_VXor(OpSize::i128Bit, Value, HelperNode);
}
StoreStackValue(ResultNode);
break;
@@ -1101,7 +1108,7 @@ void X87StackOptimization::Run(IREmitter* Emit) {
} else {
// Intermediate insts
Ref HelperNode = IREmit->_LoadNamedVectorConstant(OpSize::i128Bit, IR::NamedVectorConstant::NAMED_VECTOR_F80_SIGN_MASK);
ResultNode = IREmit->_VAndn(OpSize::i128Bit, OpSize::i8Bit, Value, HelperNode);
ResultNode = IREmit->_VAndn(OpSize::i128Bit, Value, HelperNode);
}
StoreStackValue(ResultNode);
break;
@@ -1226,11 +1233,9 @@ void X87StackOptimization::Run(IREmitter* Emit) {
SynchronizeStackValues();
FlushCachedRegs();
}
return;
}
fextl::unique_ptr<Pass> CreateX87StackOptimizationPass(const FEXCore::HostFeatures& Features, OpSize GPROpSize) {
fextl::unique_ptr<Pass> CreateX87StackOptimizationPass(const HostFeatures& Features, OpSize GPROpSize) {
return fextl::make_unique<X87StackOptimization>(Features, GPROpSize);
}
} // namespace FEXCore::IR
+46 -31
View File
@@ -1,5 +1,6 @@
// SPDX-License-Identifier: MIT
#include "Utils/Allocator/HostAllocator.h"
#include "Utils/Allocator.h"
#include <FEXCore/Utils/Allocator.h>
#include <FEXCore/Utils/CompilerDefs.h>
#include <FEXCore/Utils/LogManager.h>
@@ -32,20 +33,18 @@ std::pmr::memory_resource* get_default_resource() {
}
} // namespace fextl::pmr
#ifndef _WIN32
namespace FEXCore::Allocator {
#ifndef _WIN32
MMAP_Hook mmap {::mmap};
MUNMAP_Hook munmap {::munmap};
uint64_t HostVASize {};
using GLIBC_MALLOC_Hook = void* (*)(size_t, const void* caller);
using GLIBC_REALLOC_Hook = void* (*)(void*, size_t, const void* caller);
using GLIBC_FREE_Hook = void (*)(void*, const void* caller);
fextl::unique_ptr<Alloc::HostAllocator> Alloc64 {};
static fextl::unique_ptr<Alloc::HostAllocator> Alloc64 {};
void* FEX_mmap(void* addr, size_t length, int prot, int flags, int fd, off_t offset) {
static void* FEX_mmap(void* addr, size_t length, int prot, int flags, int fd, off_t offset) {
void* Result = Alloc64->Mmap(addr, length, prot, flags, fd, offset);
if (Result >= (void*)-4096) {
errno = -(uint64_t)Result;
@@ -69,7 +68,7 @@ void VirtualName(const char* Name, void* Ptr, size_t Size) {
}
}
int FEX_munmap(void* addr, size_t length) {
static int FEX_munmap(void* addr, size_t length) {
int Result = Alloc64->Munmap(addr, length);
if (Result != 0) {
@@ -103,9 +102,11 @@ void ClearHooks() {
}
#pragma GCC diagnostic pop
FEX_DEFAULT_VISIBILITY size_t DetermineVASize() {
if (HostVASize) {
return HostVASize;
FEX_DEFAULT_VISIBILITY size_t GetHostVABits() {
static uint64_t HostVABits = 0;
if (HostVABits) {
return HostVABits;
}
static constexpr std::array<uintptr_t, 7> TLBSizes = {
@@ -113,27 +114,18 @@ FEX_DEFAULT_VISIBILITY size_t DetermineVASize() {
};
for (auto Bits : TLBSizes) {
uintptr_t Size = 1ULL << Bits;
// Just try allocating
// We can't actually determine VA size on ARM safely
auto Find = [](uintptr_t Size) -> bool {
for (int i = 0; i < 64; ++i) {
// Try grabbing a some of the top pages of the range
// x86 allocates some high pages in the top end
void* Ptr = ::mmap(reinterpret_cast<void*>(Size - FEXCore::Utils::FEX_PAGE_SIZE * i), FEXCore::Utils::FEX_PAGE_SIZE, PROT_NONE,
MAP_FIXED_NOREPLACE | MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
if (Ptr != (void*)~0ULL) {
::munmap(Ptr, FEXCore::Utils::FEX_PAGE_SIZE);
if (Ptr == (void*)(Size - FEXCore::Utils::FEX_PAGE_SIZE * i)) {
return true;
}
}
}
return false;
};
if (Find(Size)) {
HostVASize = Bits;
// We can't actually determine VA size on ARM safely.
// Instead, try allocating the page at the top of the range.
// If this succeeds OR the page is reported as already existing,
// we know we're in valid VA space. Otherwise, we must go lower.
void* Addr = reinterpret_cast<void*>((1ULL << Bits) - FEXCore::Utils::FEX_PAGE_SIZE);
void* Ptr = ::mmap(Addr, FEXCore::Utils::FEX_PAGE_SIZE, PROT_NONE, MAP_FIXED_NOREPLACE | MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
if (Ptr != (void*)~0ULL) {
::munmap(Ptr, FEXCore::Utils::FEX_PAGE_SIZE);
}
if (Ptr != (void*)~0ULL || errno == EEXIST) {
HostVABits = Bits;
return Bits;
}
}
@@ -261,9 +253,19 @@ fextl::vector<MemoryRegion> StealMemoryRegion(uintptr_t Begin, uintptr_t End) {
}
// Block remaining memory gaps
bool SupportsDontDump = true;
for (auto RegionIt = Regions.begin(); RegionIt != Regions.end(); ++RegionIt) {
auto Alloc = ::mmap(RegionIt->Ptr, RegionIt->Size, PROT_NONE, MAP_ANONYMOUS | MAP_NORESERVE | MAP_PRIVATE | MAP_FIXED_NOREPLACE, -1, 0);
if (SupportsDontDump) {
// Mark these regions as don't dump so that coredump doesn't try dumping large unmapped regions.
// Ideally coredump would be smart enough to only dump resident pages, but here we are.
auto Result = madvise(RegionIt->Ptr, RegionIt->Size, MADV_DONTDUMP);
if (Result == -1) {
SupportsDontDump = false;
}
}
LogMan::Throw::AFmt(Alloc != MAP_FAILED, "StealMemoryRegion: mmap({}, {:x}) failed: {}", fmt::ptr(RegionIt->Ptr), RegionIt->Size, errno);
LogMan::Throw::AFmt(Alloc == RegionIt->Ptr, "mmap returned {} instead of {}", Alloc, fmt::ptr(RegionIt->Ptr));
}
@@ -272,7 +274,7 @@ fextl::vector<MemoryRegion> StealMemoryRegion(uintptr_t Begin, uintptr_t End) {
}
fextl::vector<MemoryRegion> Setup48BitAllocatorIfExists(size_t PageSize) {
size_t Bits = FEXCore::Allocator::DetermineVASize();
size_t Bits = FEXCore::Allocator::GetHostVABits();
if (Bits < 48) {
return {};
}
@@ -304,5 +306,18 @@ void UnlockAfterFork(FEXCore::Core::InternalThreadState* Thread, bool Child) {
Alloc64->UnlockAfterFork(Thread, Child);
}
}
} // namespace FEXCore::Allocator
#else
void VirtualNameNOP(const char*, const void*, size_t) {}
void VirtualTHPNOP(const void* Ptr, size_t Size, THPControl Control) {}
VirtualNamePtr VirtualName {VirtualNameNOP};
VirtualTHPPtr VirtualTHPControl {VirtualTHPNOP};
void SetupHooks(size_t PageSize, HookPtrs Ptrs) {
VirtualName = Ptrs.VirtualName;
VirtualTHPControl = Ptrs.VirtualTHPControl;
}
#endif
} // namespace FEXCore::Allocator
+2
View File
@@ -1,11 +1,13 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <cstddef>
namespace FEXCore::Core {
struct InternalThreadState;
}
namespace FEXCore::Allocator {
void InitializeAllocator(size_t PageSize);
void LockBeforeFork(FEXCore::Core::InternalThreadState* Thread);
void UnlockAfterFork(FEXCore::Core::InternalThreadState* Thread, bool Child);
} // namespace FEXCore::Allocator
@@ -7,11 +7,9 @@
#include <FEXCore/Utils/MathUtils.h>
#include <FEXCore/Utils/SignalScopeGuards.h>
#include <FEXCore/Utils/TypeDefines.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/MathUtils.h>
#include <FEXCore/fextl/sstream.h>
#include <FEXHeaderUtils/Syscalls.h>
#include <FEXCore/fextl/memory.h>
#include <FEXCore/fextl/sstream.h>
#include <FEXCore/fextl/vector.h>
#include <algorithm>
@@ -114,24 +112,24 @@ private:
return sizeof(LiveVMARegion) + FEXCore::FlexBitSet<FlexBitElementType>::SizeInBytes(NumElements);
}
static void InitializeVMARegionUsed(LiveVMARegion* Region, size_t AdditionalSize) {
size_t SizeOfLiveRegion =
static void InitializeVMARegionUsed(LiveVMARegion* Region) {
const size_t SizeOfLiveRegion =
FEXCore::AlignUp(LiveVMARegion::GetFEXManagedVMARegionSize(Region->SlabInfo->RegionSize), FEXCore::Utils::FEX_PAGE_SIZE);
size_t SizePlusManagedData = SizeOfLiveRegion + AdditionalSize;
Region->FreeSpace = Region->SlabInfo->RegionSize - SizePlusManagedData;
Region->FreeSpace = Region->SlabInfo->RegionSize - SizeOfLiveRegion;
size_t NumManagedPages = SizePlusManagedData >> FEXCore::Utils::FEX_PAGE_SHIFT;
size_t NumManagedPages = SizeOfLiveRegion >> FEXCore::Utils::FEX_PAGE_SHIFT;
size_t ManagedSize = NumManagedPages << FEXCore::Utils::FEX_PAGE_SHIFT;
// Use madvise to set the full tracking region to zero.
// This ensures unused pages are zero, while not having the backing pages consuming memory.
::madvise(Region->UsedPages.Memory + ManagedSize, (Region->SlabInfo->RegionSize >> FEXCore::Utils::FEX_PAGE_SHIFT) - ManagedSize,
MADV_DONTNEED);
auto* MemoryAsBytes = reinterpret_cast<uint8_t*>(Region->UsedPages.Memory);
const auto TrackingRegionSize = Region->SlabInfo->RegionSize - ManagedSize;
::madvise(MemoryAsBytes + ManagedSize, TrackingRegionSize, MADV_DONTNEED);
// Use madvise to claim WILLNEED on the beginning pages for initial state tracking.
// Improves performance of the following MemClear by not doing a page level fault dance for data necessary to track >170TB of used pages.
::madvise(Region->UsedPages.Memory, ManagedSize, MADV_WILLNEED);
::madvise(MemoryAsBytes, ManagedSize, MADV_WILLNEED);
// Set our reserved pages
Region->UsedPages.MemSet(NumManagedPages);
@@ -154,28 +152,27 @@ private:
FEXCore::ForkableUniqueMutex AllocationMutex;
void DetermineVASize();
LiveVMARegion* MakeRegionActive(ReservedRegionListType::iterator ReservedIterator, uint64_t UsedSize) {
LiveVMARegion* MakeRegionActive(ReservedRegionListType::iterator ReservedIterator) {
ReservedVMARegion* ReservedRegion = *ReservedIterator;
ReservedRegions->erase(ReservedIterator);
// mprotect the new region we've allocated
size_t SizeOfLiveRegion =
const size_t SizeOfLiveRegion =
FEXCore::AlignUp(LiveVMARegion::GetFEXManagedVMARegionSize(ReservedRegion->RegionSize), FEXCore::Utils::FEX_PAGE_SIZE);
size_t SizePlusManagedData = UsedSize + SizeOfLiveRegion;
auto Res = mprotect(reinterpret_cast<void*>(ReservedRegion->Base), SizePlusManagedData, PROT_READ | PROT_WRITE);
auto Res = mprotect(reinterpret_cast<void*>(ReservedRegion->Base), SizeOfLiveRegion, PROT_READ | PROT_WRITE);
LOGMAN_THROW_A_FMT(Res != -1, "Couldn't mprotect region: {} '{}' Likely occurs when running out of memory or Maximum VMAs", errno,
strerror(errno));
FEXCore::Allocator::VirtualName("FEXMem_Misc", reinterpret_cast<void*>(ReservedRegion->Base), SizePlusManagedData);
FEXCore::Allocator::VirtualName("FEXMem_Misc", reinterpret_cast<void*>(ReservedRegion->Base), SizeOfLiveRegion);
LiveVMARegion* LiveRange = new (reinterpret_cast<void*>(ReservedRegion->Base)) LiveVMARegion();
// Copy over the reserved data
LiveRange->SlabInfo = ReservedRegion;
// Initialize VMA
LiveVMARegion::InitializeVMARegionUsed(LiveRange, UsedSize);
LiveVMARegion::InitializeVMARegionUsed(LiveRange);
// Add to our active tracked ranges
auto LiveIter = LiveRegions->emplace_back(LiveRange);
@@ -187,7 +184,7 @@ private:
};
void OSAllocator_64Bit::DetermineVASize() {
size_t Bits = FEXCore::Allocator::DetermineVASize();
size_t Bits = FEXCore::Allocator::GetHostVABits();
uintptr_t Size = 1ULL << Bits;
UPPER_BOUND = Size;
@@ -224,7 +221,7 @@ OSAllocator_64Bit::LiveVMARegion* OSAllocator_64Bit::FindLiveRegionForAddress(ui
uintptr_t RegionEnd = ReservedRegion->Base + ReservedRegion->RegionSize;
if (Addr >= ReservedRegion->Base && AddrEnd < RegionEnd) {
// Found one, let's make it active
LiveRegion = MakeRegionActive(it, 0);
LiveRegion = MakeRegionActive(it);
break;
}
}
@@ -394,7 +391,7 @@ again:
size_t lengthPlusManagedData = length + lengthOfLiveRegion;
for (auto it = ReservedRegions->begin(); it != ReservedRegions->end(); ++it) {
if ((*it)->RegionSize >= lengthPlusManagedData) {
MakeRegionActive(it, 0);
MakeRegionActive(it);
goto again;
}
}
@@ -623,14 +620,14 @@ fextl::unique_ptr<T> make_alloc_unique(FEXCore::Allocator::MemoryRegion& Base, A
fextl::unique_ptr<Alloc::HostAllocator> Create64BitAllocatorWithRegions(fextl::vector<FEXCore::Allocator::MemoryRegion>& Regions) {
// This is a bit tricky as we can't allocate memory safely except from the Regions provided. Otherwise we might overwrite memory pages we
// don't own. Scan the memory regions and find the smallest one.
FEXCore::Allocator::MemoryRegion& Smallest = Regions[0];
for (auto& it : Regions) {
if (it.Size <= Smallest.Size) {
Smallest = it;
FEXCore::Allocator::MemoryRegion* Smallest = &Regions[0];
for (auto& Region : Regions) {
if (Region.Size <= Smallest->Size) {
Smallest = &Region;
}
}
return make_alloc_unique<OSAllocator_64Bit>(Smallest, Regions);
return make_alloc_unique<OSAllocator_64Bit>(*Smallest, Regions);
}
} // namespace Alloc::OSAllocator
+2 -2
View File
@@ -39,10 +39,10 @@ struct FlexBitSet final {
Memory[Element / MinimumSizeBits] &= ~(1ULL << (Element % MinimumSizeBits));
}
void MemClear(size_t Elements) {
memset(Memory, 0, FEXCore::AlignUp(Elements / MinimumSizeBits, MinimumSizeBits));
memset(Memory, 0, SizeInBytes(Elements));
}
void MemSet(size_t Elements) {
memset(Memory, 0xFF, FEXCore::AlignUp(Elements / MinimumSizeBits, MinimumSizeBits));
memset(Memory, 0xFF, SizeInBytes(Elements));
}
// Range scanning results
+3 -1
View File
@@ -2,7 +2,6 @@
#ifdef ENABLE_FEX_ALLOCATOR
#include <rpmalloc/rpmalloc.h>
#ifndef _WIN32
#include <linux/prctl.h>
#include <sys/prctl.h>
#include <sys/mman.h>
#else
@@ -109,6 +108,9 @@ static void* FEX_rp_mmap(size_t size, size_t alignment, size_t* offset, size_t*
#define PR_SET_VMA_ANON_NAME 0
#endif
prctl(PR_SET_VMA, PR_SET_VMA_ANON_NAME, ptr, map_size, global_config.page_name);
// Disable HUGEPAGE on allocation from rpmalloc.
madvise(ptr, map_size, MADV_NOHUGEPAGE);
}
if (ptr == nullptr) {
+55 -19
View File
@@ -2,7 +2,7 @@
#include "Interface/Core/CPUBackend.h"
#include "Interface/Context/Context.h"
#include "Utils/SpinWaitLock.h"
#include <FEXCore/Utils/SpinWaitLock.h>
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXCore/Utils/EnumUtils.h>
@@ -343,11 +343,12 @@ static bool RunCASPAL(uint64_t* GPRs, uint32_t Size, uint32_t DesiredReg1, uint3
// 32bit
uint64_t Addr = GPRs[AddressReg];
// Lower register must be even, so only upper register can be 31.
uint32_t DesiredLower = GPRs[DesiredReg1];
uint32_t DesiredUpper = GPRs[DesiredReg2];
uint32_t DesiredUpper = DesiredReg2 == 31 ? 0 : GPRs[DesiredReg2];
uint32_t ExpectedLower = GPRs[ExpectedReg1];
uint32_t ExpectedUpper = GPRs[ExpectedReg2];
uint32_t ExpectedUpper = ExpectedReg2 == 31 ? 0 : GPRs[ExpectedReg2];
// Cross-cacheline CAS doesn't work on ARM
// It isn't even guaranteed to work on x86
@@ -1352,7 +1353,9 @@ static std::optional<uint64_t> DoCAS(uint32_t Size, uint64_t Desired, uint64_t E
}
static bool RunCASAL(uint64_t* GPRs, uint32_t Size, uint32_t DesiredReg, uint32_t ExpectedReg, uint32_t AddressReg, uint32_t* StrictSplitLockMutex) {
std::optional<uint64_t> Res = DoCAS(Size, GPRs[DesiredReg], GPRs[ExpectedReg], GPRs[AddressReg], StrictSplitLockMutex);
uint64_t Desired = DesiredReg == 31 ? 0 : GPRs[DesiredReg];
uint64_t Expected = ExpectedReg == 31 ? 0 : GPRs[ExpectedReg];
std::optional<uint64_t> Res = DoCAS(Size, Desired, Expected, GPRs[AddressReg], StrictSplitLockMutex);
if (!Res.has_value()) {
return false;
}
@@ -1384,6 +1387,8 @@ static bool HandleAtomicMemOp(uint32_t Instr, uint64_t* GPRs, uint32_t* StrictSp
uint8_t Op = (Instr >> 12) & 0xF;
uint64_t Source = SourceReg == 31 ? 0 : GPRs[SourceReg];
if (Size == 2) {
auto NOPExpected = [](uint16_t SrcVal, uint16_t) -> uint16_t {
return SrcVal;
@@ -1420,7 +1425,7 @@ static bool HandleAtomicMemOp(uint32_t Instr, uint64_t* GPRs, uint32_t* StrictSp
default: LogMan::Msg::EFmt("Unhandled JIT SIGBUS Atomic mem op 0x{:02x}", Op); return false;
}
auto Res = DoCAS16<true>(GPRs[SourceReg],
auto Res = DoCAS16<true>(Source,
0, // Unused
Addr, NOPExpected, DesiredFunction, StrictSplitLockMutex);
// If we passed and our destination register is not zero
@@ -1465,7 +1470,7 @@ static bool HandleAtomicMemOp(uint32_t Instr, uint64_t* GPRs, uint32_t* StrictSp
default: LogMan::Msg::EFmt("Unhandled JIT SIGBUS Atomic mem op 0x{:02x}", Op); return false;
}
auto Res = DoCAS32<true>(GPRs[SourceReg],
auto Res = DoCAS32<true>(Source,
0, // Unused
Addr, NOPExpected, DesiredFunction, StrictSplitLockMutex);
// If we passed and our destination register is not zero
@@ -1510,7 +1515,7 @@ static bool HandleAtomicMemOp(uint32_t Instr, uint64_t* GPRs, uint32_t* StrictSp
default: LogMan::Msg::EFmt("Unhandled JIT SIGBUS Atomic mem op 0x{:02x}", Op); return false;
}
auto Res = DoCAS64<true>(GPRs[SourceReg],
auto Res = DoCAS64<true>(Source,
0, // Unused
Addr, NOPExpected, DesiredFunction, StrictSplitLockMutex);
// If we passed and our destination register is not zero
@@ -1572,9 +1577,11 @@ static bool HandleAtomicStore(uint32_t Instr, uint64_t* GPRs, int64_t Offset, ui
uint64_t Addr = GPRs[AddressReg] + Offset;
constexpr bool DoRetry = false;
uint64_t Data = DataReg == 31 ? 0 : GPRs[DataReg];
if (Size == 2) {
DoCAS16<DoRetry>(
GPRs[DataReg],
Data,
0, // Unused
Addr,
[](uint16_t SrcVal, uint16_t) -> uint16_t {
@@ -1589,7 +1596,7 @@ static bool HandleAtomicStore(uint32_t Instr, uint64_t* GPRs, int64_t Offset, ui
return true;
} else if (Size == 4) {
DoCAS32<DoRetry>(
GPRs[DataReg],
Data,
0, // Unused
Addr,
[](uint32_t SrcVal, uint32_t) -> uint32_t {
@@ -1604,7 +1611,7 @@ static bool HandleAtomicStore(uint32_t Instr, uint64_t* GPRs, int64_t Offset, ui
return true;
} else if (Size == 8) {
DoCAS64<DoRetry>(
GPRs[DataReg],
Data,
0, // Unused
Addr,
[](uint64_t SrcVal, uint64_t) -> uint64_t {
@@ -1834,6 +1841,7 @@ static uint64_t HandleAtomicLoadstoreExclusive(uintptr_t ProgramCounter, uint64_
return Desired;
};
uint64_t Source = DataSourceReg == 31 ? 0 : GPRs[DataSourceReg];
if (Size == 2) {
using AtomicType = uint16_t;
CASDesiredFn<AtomicType> DesiredFunction {};
@@ -1852,7 +1860,7 @@ static uint64_t HandleAtomicLoadstoreExclusive(uintptr_t ProgramCounter, uint64_
default: LogMan::Msg::EFmt("Unhandled JIT SIGBUS Atomic mem op 0x{:02x}", FEXCore::ToUnderlying(AtomicOp)); return false;
}
auto Res = DoCAS16<DoRetry>(GPRs[DataSourceReg],
auto Res = DoCAS16<DoRetry>(Source,
0, // Unused
Addr, NOPExpected, DesiredFunction, StrictSplitLockMutex);
@@ -1879,7 +1887,7 @@ static uint64_t HandleAtomicLoadstoreExclusive(uintptr_t ProgramCounter, uint64_
default: LogMan::Msg::EFmt("Unhandled JIT SIGBUS Atomic mem op 0x{:02x}", FEXCore::ToUnderlying(AtomicOp)); return false;
}
auto Res = DoCAS32<DoRetry>(GPRs[DataSourceReg],
auto Res = DoCAS32<DoRetry>(Source,
0, // Unused
Addr, NOPExpected, DesiredFunction, StrictSplitLockMutex);
@@ -1906,7 +1914,7 @@ static uint64_t HandleAtomicLoadstoreExclusive(uintptr_t ProgramCounter, uint64_
default: LogMan::Msg::EFmt("Unhandled JIT SIGBUS Atomic mem op 0x{:02x}", FEXCore::ToUnderlying(AtomicOp)); return false;
}
auto Res = DoCAS64<DoRetry>(GPRs[DataSourceReg],
auto Res = DoCAS64<DoRetry>(Source,
0, // Unused
Addr, NOPExpected, DesiredFunction, StrictSplitLockMutex);
if (AtomicFetch && ResultReg != 31) {
@@ -1947,8 +1955,24 @@ std::optional<int32_t> HandleUnalignedAccess(FEXCore::Core::InternalThreadState*
uint32_t* StrictSplitLockMutex {CTX->Config.StrictInProcessSplitLocks ? &CTX->StrictSplitLockMutex : nullptr};
if (!IsJIT) [[unlikely]] {
if ((Instr & LDAXR_MASK) == LDAR_INST || // LDAR*
(Instr & LDAXR_MASK) == LDAPR_INST) { // LDAPR*
if ((Instr & ArchHelpers::Arm64::CASPAL_MASK) == ArchHelpers::Arm64::CASPAL_INST) { // CASPAL
if (ArchHelpers::Arm64::HandleCASPAL(Instr, GPRs, StrictSplitLockMutex)) {
// Skip this instruction now
return 4;
} else {
LogMan::Msg::EFmt("Unhandled JIT SIGBUS CASPAL: PC: 0x{:x} Instruction: 0x{:08x}\n", ProgramCounter, PC[0]);
return std::nullopt;
}
} else if ((Instr & ArchHelpers::Arm64::CASAL_MASK) == ArchHelpers::Arm64::CASAL_INST) { // CASAL
if (ArchHelpers::Arm64::HandleCASAL(GPRs, Instr, StrictSplitLockMutex)) {
// Skip this instruction now
return 4;
} else {
LogMan::Msg::EFmt("Unhandled JIT SIGBUS CASAL: PC: 0x{:x} Instruction: 0x{:08x}\n", ProgramCounter, PC[0]);
return std::nullopt;
}
} else if ((Instr & LDAXR_MASK) == LDAR_INST || // LDAR*
(Instr & LDAXR_MASK) == LDAPR_INST) { // LDAPR*
if (ArchHelpers::Arm64::HandleAtomicLoad(Instr, GPRs, 0)) {
// Skip this instruction now
return 4;
@@ -1991,22 +2015,34 @@ std::optional<int32_t> HandleUnalignedAccess(FEXCore::Core::InternalThreadState*
} else if ((Instr & ArchHelpers::Arm64::STLXR_MASK) == ArchHelpers::Arm64::STLXR_INST) { // STLXR*
uint32_t StatusReg = Instr << 11 >> 27;
// // Emulate exclusive store by validating the address and value against the last unaligned LDAXR*.
if (GPRs[AddrReg] != Thread->ExclusiveStore.Addr || Size > Thread->ExclusiveStore.Size) {
uint32_t SizeBytes = 1u << Size;
if (GPRs[AddrReg] != Thread->ExclusiveStore.Addr || SizeBytes > Thread->ExclusiveStore.Size) {
if (StatusReg != 31) {
GPRs[StatusReg] = 1;
}
return 4;
}
if (std::optional<uint64_t> Prev =
DoCAS(Size, DataReg == 31 ? 0 : GPRs[DataReg], Thread->ExclusiveStore.Store, GPRs[AddrReg], StrictSplitLockMutex)) {
DoCAS(SizeBytes, DataReg == 31 ? 0 : GPRs[DataReg], Thread->ExclusiveStore.Store, GPRs[AddrReg], StrictSplitLockMutex)) {
if (StatusReg != 31) {
GPRs[StatusReg] = !!memcmp(&Thread->ExclusiveStore.Store, &*Prev, Size);
GPRs[StatusReg] = !!memcmp(&Thread->ExclusiveStore.Store, &*Prev, SizeBytes);
}
Thread->ExclusiveStore.Size = 0;
return 4;
}
} else if ((Instr & ArchHelpers::Arm64::ATOMIC_MEM_MASK) == ArchHelpers::Arm64::ATOMIC_MEM_INST) { // Atomic memory op
if (ArchHelpers::Arm64::HandleAtomicMemOp(Instr, GPRs, StrictSplitLockMutex)) {
// Skip this instruction now
return 4;
} else {
uint8_t Op = (PC[0] >> 12) & 0xF;
LogMan::Msg::EFmt("Unhandled JIT SIGBUS Atomic mem op 0x{:02x}: PC: 0x{:x} Instruction: 0x{:08x}\n", Op, ProgramCounter, PC[0]);
return std::nullopt;
}
}
return 0;
LogMan::Msg::EFmt("Unhandled non-JIT atomic");
return std::nullopt;
}
const auto Frame = Thread->CurrentFrame;
+2
View File
@@ -1,4 +1,6 @@
// SPDX-License-Identifier: MIT
#include <FEXCore/Utils/CompilerDefs.h>
namespace FEXCore::Assert {
// This function can not be inlined
[[noreturn]]
@@ -32,7 +32,7 @@ public:
// Differs from Itanium specification
LOGMAN_THROW_A_FMT(PMF.adj == 0, "C++ Pointer-To-Member representation didn't have adj == 0. Are you trying to cast a virtual member?");
#else
#error Don't know how to cast Member to function here. Likely just Itanium
#error "Don't know how to cast Member to function here. Likely just Itanium"
#endif
return PMF.ptr;
}
@@ -54,7 +54,7 @@ public:
"members.");
return PMF.ptr;
#else
#error Don't know how to cast Member to function here. Likely just Itanium
#error "Don't know how to cast Member to function here. Likely just Itanium"
#endif
}
+3 -3
View File
@@ -1,11 +1,11 @@
// SPDX-License-Identifier: MIT
#include "Utils/SpinWaitLock.h"
#include <FEXCore/Utils/SpinWaitLock.h>
namespace FEXCore::Utils::SpinWaitLock {
#ifdef ARCHITECTURE_arm64
constexpr uint64_t NanosecondsInSecond = 1'000'000'000ULL;
static uint32_t GetCycleCounterFrequency() {
static uint64_t GetCycleCounterFrequency() {
uint64_t Result {};
__asm("mrs %[Res], CNTFRQ_EL0" : [Res] "=r"(Result));
return Result;
@@ -21,7 +21,7 @@ static uint64_t CalculateCyclesPerNanosecond() {
return NanosecondsInSecond / CounterFrequency;
}
uint32_t CycleCounterFrequency = GetCycleCounterFrequency();
uint64_t CycleCounterFrequency = GetCycleCounterFrequency();
uint64_t CyclesPerNanosecond = CalculateCyclesPerNanosecond();
#endif
} // namespace FEXCore::Utils::SpinWaitLock
+23
View File
@@ -0,0 +1,23 @@
// SPDX-License-Identifier: MIT
#include <FEXCore/Utils/WildcardMatcher.h>
namespace FEXCore::Utils::Wildcard {
static bool matchHelper(std::string_view pattern, std::string_view text, size_t p_idx, size_t t_idx) {
if (p_idx == pattern.size()) {
// Pattern exhausted
return (t_idx == text.size());
} else if (pattern[p_idx] == '*') {
// Wildcard: Try matching zero characters, or one or more characters
return matchHelper(pattern, text, p_idx + 1, t_idx) || (t_idx < text.size() && matchHelper(pattern, text, p_idx, t_idx + 1));
} else {
// Match normally
return (t_idx < text.size() && pattern[p_idx] == text[t_idx] && matchHelper(pattern, text, p_idx + 1, t_idx + 1));
}
}
bool Matches(std::string_view pattern, std::string_view text) {
return matchHelper(pattern, text, 0, 0);
}
} // namespace FEXCore::Utils::Wildcard
Loaded 100 of 636 files, more files were not shown because too many files have changed in this diff. Show more