Compare commits

..
591 Commits
Author SHA1 Message Date
Ryan Houdek 9fbdc00bd6 Docs: Update for release FEX-2609.1 2026-09-21 11:44:20 -07:00
Ryan Houdek 116053cd29 InstcountCI: Update 2026-09-21 11:43:34 -07:00
Ryan Houdek 536722f59d FEXCore/OpcodeDispatcher: Fixes syscall instruction on Linux
I forgot on Linux by default we didn't have the syscall instructions
count as block end. Change this so that it counts as block end now.

This has the additional benefit that now the frontend needs to modify
the RIP manually as well which is fine as it's what arm64ec and wow64
does.

Also add back the UnimplementedOp in RDPID that accidentally got caught
up. Also increment DiskCache version as both changes will change
codegen.

Fixes #5942
2026-09-21 11:43:28 -07:00
Ryan Houdek 395b132f34 Docs: Update for release FEX-2609 2026-09-07 23:42:35 -07:00
Ryan Houdek 566a5e266d Merge pull request #5934 from Plagman/plagman/multi_key_drifting_mr
DiskCache: support multiple entries per lookup key
2026-09-07 23:38:31 -07:00
Pierre-Loup A. Griffais cc5686de37 InstcountCI: Update 2026-09-07 22:51:31 -07:00
Pierre-Loup A. Griffais 049c932675 DiskCache: support multiple entries per lookup key
The LookupKey based on data available at Lookup time gives a list of possible
candidate entries, which may have different guest code sizes/footprints.

If only one candidate, it's stored in-line in the map like before - if we grow
past that, an additional (multi)map is allocated, sorted by footprint.

The footprint sorting lets us reduce the amount of hashes performed at Lookup
to the strict minimum.

Anon entries keys are based on a limited size decoded prefix, so this is an
important part of geting the best hit rate possible out of all the other work
with anon keys, as there's often multiple candidates with the same prefix.

Cap the max amount of entries per bucket to limit growth.
2026-09-07 22:50:22 -07:00
LC 84fabf4ed8 Merge pull request #5933 from Sonicadvance1/235
FEXCore/Allocator: Support naming VMA regions from rpmalloc
2026-09-07 23:36:30 -04:00
Ryan Houdek 310c2504f6 Merge pull request #5930 from cjacek/raise-exception
Windows: Use NtRaiseException to raise guest exception
2026-09-07 15:20:37 -07:00
Ryan Houdek 021c4fa4bf FEXCore/Allocator: Support naming VMA regions from rpmalloc
This allows WTF to catch the allocations just like on Linux. Punch our
unixlib path all the way through to rpmalloc so it gets named and
tracked properly.
2026-09-07 15:13:33 -07:00
Ryan Houdek 319c5bf023 Externals/rpmalloc: Update 2026-09-07 15:13:14 -07:00
Ryan Houdek 023d0cdaa6 Merge pull request #5927 from Plagman/plagman/mono_more_caching
DiskCache: also patch branches for more anon hit rate
2026-09-07 14:09:35 -07:00
Ryan Houdek 5a4f84cad9 Merge pull request #5870 from cjacek/fmt-update
External: Update fmt to upstream master to fix building with LLVM 23
2026-09-07 12:24:52 -07:00
Jacek Caban eb27391954 External: Update fmt to upstream master to fix building with LLVM 23
Includes https://github.com/fmtlib/fmt/pull/4868.
2026-09-07 16:23:16 +02:00
Pierre-Loup A. Griffais ffaa5d6316 InstcountCI: Update 2026-09-07 00:56:39 -07:00
Pierre-Loup A. Griffais 8bf22b6382 DiskCache: add plumbing for guest-patching for branches
To catch many conditional jumps found in Mono JITted code.
2026-09-07 00:47:47 -07:00
Ryan Houdek 42c663269b Merge pull request #5910 from Plagman/plagman/mono_initial_caching
DiskCache: detect inline data, hash around it and patch it on Lookup
2026-09-06 23:45:17 -07:00
Pierre-Loup A. Griffais 983c5defc1 InstcountCI: Update 2026-09-06 23:25:53 -07:00
Pierre-Loup A. Griffais c92805475d DiskCache: detect inline data, hash around it and patch it on Lookup
Only detect certain kinds of mov reg,immediate so far, which was the bulk of
Mono JIT activity.
2026-09-06 23:25:29 -07:00
Ryan Houdek 50e6eee95a Merge pull request #5931 from Plagman/plagman/key_flags_mr
DiskCache: key fixes
2026-09-06 22:32:00 -07:00
Pierre-Loup A. Griffais 31cda95039 DiskCache: also fold FileId into the key for file-backed blocks 2026-09-06 21:47:18 -07:00
Pierre-Loup A. Griffais 970ce4bc13 DiskCache: fold more codegen-affecting environment into key
Fixes a hang during load with Three Point Hospital
2026-09-06 21:40:54 -07:00
Jacek Caban e4c6154be2 WOW64: Use NtRaiseException to raise guest exception
Instead of returning to the currently dispatched exception, which doesn't send a debug event.
2026-09-06 15:22:43 +02:00
Jacek Caban 064a6e96c7 ARM64EC: Use NtRaiseException to raise guest exception
Instead of returning to KiUserExceptionDispatcher, which doesn't send a debug event.
2026-09-06 15:20:30 +02:00
Jacek Caban d704bd3cef Windows: Rise a non-continuable excetpion for STATUS_STACK_BUFFER_OVERRUN 2026-09-06 15:20:30 +02:00
Jacek Caban 9f511d5d78 WOW64: Remove unused variable 2026-09-06 15:20:30 +02:00
Ryan Houdek a77d00acae Merge pull request #5929 from Plagman/plagman/smc_mr
DiskCache: SMC regression fixes
2026-09-06 03:24:15 -07:00
Pierre-Loup A. Griffais 5d11b8b0d5 DiskCache: SMC regression fixes
One from a bad conflict resolution :/
2026-09-06 02:00:05 -07:00
Ryan Houdek bf70dee3ae Merge pull request #5928 from Plagman/plagman/jittail_mr
Frontend: fix nondeterminism by rolling back DecodedMaxAddress on error
2026-09-05 22:39:05 -07:00
Pierre-Loup A. Griffais 4a065ad1a2 InstcountCI: Update 2026-09-05 21:58:07 -07:00
Pierre-Loup A. Griffais 96d955e7b1 Frontend: fix nondeterminism by rolling back DecodedMaxAddress on error
Otherwise GuestSize ends up different across compiles of the same valid code
extents, including the copy of it baked in JITCodeTail of the host code itself.
2026-09-05 21:53:30 -07:00
LC abec314720 Merge pull request #5926 from Sonicadvance1/234
Frontend: Stop asserting on too long of an instruction
2026-09-05 03:29:29 -04:00
LC ec63be4028 Merge pull request #5925 from Sonicadvance1/233
Wine: Change behaviour around SMC and NtReadFile
2026-09-05 03:28:17 -04:00
Ryan Houdek 179b4423a7 Wine: Change behaviour around SMC and NtReadFile
JIT invalidation around syscalls and kernel calls is mighty fickle.
This is why the previous code path was born. In order to properly handle
this case we need strong coordination between Wine and FEX around
invalidating code and memory protections. This doesn't quite exist today
so we're kind of stuck with a kludge solution. Instead of forcing the
Persona 5 code invalidation on to every process, only do it on P5R.

This fixes a hang in msiexec with PhysX trying to do a blocking read and
jitting code or creating threads, while also maintaining the P5R
approach. The full comment is in the file about the reasoning.
2026-09-04 20:08:18 -07:00
Ryan Houdek 9956a528b9 Frontend: Stop asserting on too long of an instruction
Return the same as the NonExecutableRange for now. I have a unittest
that is mostly correct with this change, just need to fixup some of the
supplementary data.

For now push this change to get a bad assert removed that can happen
during code discovery. I'll fix up the remaining details afterwards.
2026-09-04 20:05:18 -07:00
Ryan Houdek e3c6b37f70 Merge pull request #5924 from Sonicadvance1/232
Common/Config: Fixes Cache directory when `STEAM_COMPAT_SHADER_PATH` on Win32
2026-09-04 16:08:44 -07:00
LC b78bde7ec8 Merge pull request #5923 from Sonicadvance1/231
OpcodeDispatcher: Explicitly handle no-nop vector moves
2026-09-04 17:54:36 -04:00
Ryan Houdek 6ff7e82e3b Merge pull request #5918 from Plagman/plagman/validation_mr
DiskCache: optional validation-only mode
2026-09-04 14:34:38 -07:00
Ryan Houdek 37675e0351 Common/Config: Fixes Cache directory when STEAM_COMPAT_SHADER_PATH on Win32
Moves cache data from %APPDATA% to wherever `STEAM_COMPAT_SHADER_PATH`
points. Turns out this was easier than expected.

Basically moving disk cache from:
- compatdata/<AppID>/pfx/drive_c/users/steamuser/AppData/Local/fex-emu/DiskCache
to:
- shadercache/<AppID>/fex-emu/DiskCache/

Fixes #5913
2026-09-04 14:20:57 -07:00
Ryan Houdek 38c9f14dbd OpcodeDispatcher: Explicitly handle no-nop vector moves
Just be very explicit about this rather than questioning why MMX moves
can't be a nop in the regular vector move code.
Doesn't change codegen so binarycacheversion doesn't need to change.
2026-09-04 12:19:21 -07:00
Ryan Houdek 7187bc0e59 Merge pull request #5922 from simon902/MOVVectorUnalignedOp_Selfmove
MOVVectorUnalignedOp: Fix same-register MMX writes being treated as a Nop
2026-09-04 12:07:19 -07:00
Simon Scherer c5700cb0cc InstcountCI: Update 2026-09-04 17:03:46 +02:00
Simon Scherer 8ed4c52d58 OpcodeDispatcher: Don't skip the MMX state transition for same-register writes in MOVVectorUnalignedOp 2026-09-04 17:00:45 +02:00
Simon Scherer bd65c3e568 unittests/ASM: Test PFRCPIT1 self-move 2026-09-04 16:58:10 +02:00
LC 7052f91495 Merge pull request #5919 from simon902/fst_tag_word
Mark tag word valid for FST ST(0) self-store
2026-09-04 06:34:08 -04:00
Simon Scherer 040d3115d2 InstcountCI: Update 2026-09-04 11:53:13 +02:00
Simon Scherer 7236c2bf0a Fix x87 tag word not updating on FST ST(0) self-store 2026-09-04 11:51:12 +02:00
Simon Scherer ed3fa2f1c6 unittests/ASM: Test FST ST(0) tag word update 2026-09-04 11:50:21 +02:00
LC a454845ffd Merge pull request #5909 from simon902/fxch_tag_word
Mark both operands valid in the tag word for FXCH
2026-09-04 02:50:17 -04:00
Pierre-Loup A. Griffais 40bfff985f DiskCache: optional validation-only mode
Diffs guest code in lookup miss hash mismatches and host code after lookup hit.
2026-09-03 23:31:05 -07:00
Simon Scherer 96d65cd6bc InstcountCI: Update 2026-09-04 08:30:03 +02:00
Simon Scherer e4e45bcc59 x87StackOptimizationPass: Set both operands as valid in the tag word for FXCH 2026-09-04 08:26:17 +02:00
Simon Scherer 45fbab212a unittests/ASM: Test FXCH for tag word 2026-09-04 08:26:17 +02:00
Ryan Houdek 8243b93e4d Merge pull request #5917 from lioncash/wfe
System_Tests: Enable WFET/WFIT tests
2026-09-03 23:18:10 -07:00
Ryan Houdek 52cc5689db Merge pull request #5916 from lioncash/bf
ARMEmitter: Fill in bfloat API holes
2026-09-03 23:16:20 -07:00
LC cfef3d2cbf Merge pull request #5892 from simon902/FDECSTP_FDECSTP_C1
Clear C1 for fdecstp and fincstp
2026-09-04 01:57:31 -04:00
Simon Scherer fd64ea54fe InstcountCI: Update 2026-09-04 07:32:09 +02:00
Simon Scherer d1b2333195 OpcodeDispatcher: Clear C1 to 0 for X87ModifySTP (fdecstp and fincstp) 2026-09-04 07:22:05 +02:00
Simon Scherer 97fbd11177 unittests/ASM: Test fdecstp and fincstp for C1 clear 2026-09-04 07:22:05 +02:00
LC e7836f26dd Merge pull request #5914 from Sonicadvance1/230
Core: Adds more validation around pool buffer ownership
2026-09-03 23:13:43 -04:00
Ryan Houdek cf6da6ac26 Merge pull request #5915 from lioncash/vixl
Externals: Update vixl
2026-09-03 19:33:19 -07:00
Ryan Houdek a0cb91daa4 Core: Adds more validation around pool buffer ownership
All buffers should be disowned leaving their respective compilation
sites, and reowning a buffer should never have the flag already be
owned.

Throw an assert in both cases because that would be a programming error
and result in some squirrely buffer handling
2026-09-03 14:39:01 -07:00
Ryan Houdek e010e8adf0 Merge pull request #5897 from Plagman/plagman/anon_key_mr
DiskCache: initial anon caching
2026-09-03 14:38:26 -07:00
LC 7fe7c39a00 Merge pull request #5911 from javelina-pkwy/fix/windows-syscalls-arm64ec
Windows: Fix syscall regression
2026-09-03 14:48:02 -04:00
Justin Becker b03dc847f5 Fix syscall regression 2026-09-03 11:09:19 -07:00
LC ac6be77592 Merge pull request #5900 from simon902/fincstp_tag_word
x87StackOptimizationPass: Generalize ST(0) tag invalidation on pop
2026-09-03 09:43:37 -04:00
Lioncache 35330df728 System_Tests: Enable WFET/WFIT tests
vixl supports these now.
2026-09-03 07:24:28 -04:00
Lioncache bd9897b164 ARMEmitter: Handle BFCVTNT 2026-09-03 06:33:02 -04:00
Lioncache e704275ffd ARMEmitter: Handle BFCVT 2026-09-03 06:32:57 -04:00
Lioncache e33cb3c831 SVE_Tests: Enable BFMLAL tests 2026-09-03 05:52:27 -04:00
Lioncache 91264f05e2 Externals: Update vixl
Keeps it up to date with the latest HEAD

Also amends the recent disassembler changes.
2026-09-03 04:24:35 -04:00
Simon Scherer 30e4650a8b InstcountCI: Update 2026-09-03 08:32:59 +02:00
Simon Scherer add92f681b unittests/ASM: Test fpatan and fyl2x to clear tag 2026-09-03 08:29:15 +02:00
Simon Scherer 21b8f8e20d x87StackOptimizationPass: Generalize ST(0) tag invalidation on pop
UpdateTopForPop_Slow() now invalidates ST(0)'s tag by default,
so every slow-path StackPop() does this consistently instead of
the previous special case only in OP_POPSTACKDESTROY.

FINCSTP is an exception as it only moves the stack pointer without
invalidating the tag.
2026-09-03 08:26:14 +02:00
Simon Scherer ca055cd7d5 unittests/ASM: Test FINCSTP for tag word 2026-09-03 08:05:40 +02:00
Pierre-Loup A. Griffais 42e207932b InstcountCI: Update 2026-09-02 19:55:47 -07:00
Pierre-Loup A. Griffais 3157685f0f DiskCache: initial anon caching
Decode a few bytes in advance to get a hashable prefix to use as key.

Generate touched pages dynamically since they can be misaligned now, as the
cached-hit guest code isn't necessarily in the same spot as the store was.
2026-09-02 19:53:34 -07:00
LC d833f3ae67 Merge pull request #5906 from Sonicadvance1/229
OpcodeDispatcher: Validate instruction encodings with crc
2026-09-02 22:18:02 -04:00
Ryan Houdek fbf5187794 InstcountCI: Update 2026-09-02 18:35:36 -07:00
Ryan Houdek f2c8adf7a3 OpcodeDispatcher: Validate instruction encodings with crc
PR #5902 technically introduced a bug where we would read past the end
of bounds for thunk instructions when full smc was enabled. Luckily this
never occurs in practice as the Mono hacks never are on VDSO boundaries,
and no one is expected to enable full smc detection really.

Switch this path over to using crc32 unconditionally. This raises our
minspec technically to armv8-a+crc, but nothing that matters shipped
without crc so it's fine.

This also is a minor speed and JIT size reduction due less branches
polluting the BTB. But really only for mono/unity games.
Requires revving the DiskCache version again.
2026-09-02 18:35:08 -07:00
Ryan Houdek d472ce190c CMake: Always enable crc32
This effectively raises our minspec to ARMv8.0-a + crc.
2026-09-02 18:25:19 -07:00
LC c4c82f23ee Merge pull request #5905 from Sonicadvance1/228
FEXCore: Support ThreadRemoveCodeEntry as relocatable
2026-09-02 20:32:07 -04:00
Ryan Houdek 20e9e2044e InstcountCI: Update 2026-09-02 17:15:36 -07:00
Ryan Houdek 1a0412969b FEXCore: Support ThreadRemoveCodeEntry as relocatable
Frontend just needs to generate a relocatable entry and pass it to the
backend.
2026-09-02 17:14:49 -07:00
LC 55c80e6c4f Merge pull request #5902 from Sonicadvance1/226
FEXCore/Frontend: Ensure instruction sizes account for sha256 in thunk op
2026-09-02 19:56:53 -04:00
LC 83469208c9 Merge pull request #5904 from javelina-pkwy/opt/pmulhrsw
Faster translation for *PMULHRSW
2026-09-02 19:54:24 -04:00
Justin Becker 3471ca67c4 Bump DiskCache version number 2026-09-02 16:40:17 -07:00
Ryan Houdek 4048faa54b FEXCore/Frontend: Ensure instruction sizes account for sha256 in thunk op
Fixes the hashing mechanism from DiskCache.cpp, allowing it to properly
handle thunk relocations.

Fixes #5896
2026-09-02 14:32:08 -07:00
jubecker 61d94ac117 Potential improvement for *PMULHRSW 2026-09-02 14:31:18 -07:00
LC 511c45c4c6 Merge pull request #5886 from simon902/IE_sticky
Preserve previous value of IE exception flag
2026-09-02 06:49:35 -04:00
Simon Scherer bce8b13366 InstcountCI: Update 2026-09-02 07:47:24 +02:00
Simon Scherer 4a7fda4516 OpcodeDispatcher: Preserve IE exception flag for FCOMI and FTST 2026-09-02 07:42:04 +02:00
Simon Scherer b29561e131 unittests/ASM: Test stickyness of IE exception flag 2026-09-02 07:42:04 +02:00
Ryan Houdek 5a68359f37 Merge pull request #5896 from Plagman/plagman/skipthunkrelocs_mr
DiskCache: disable with thunk relocs for now
2026-09-01 17:54:03 -07:00
Ryan Houdek 00b1249e53 Merge pull request #5895 from lioncash/move
Utils/File: Properly handle move assignment/construction
2026-09-01 17:53:09 -07:00
Pierre-Loup A. Griffais 8be7c1bf8e DiskCache: disable with thunk relocs for now
Block extents currently don't include some of their data.
2026-09-01 17:04:31 -07:00
Lioncache aaca8772bd Utils/File: Properly handle move assignment/construction
Ensures that whenever a file handle is transferred anywhere, that the
moved from instance won't end up closing the file handle when the
destructor runs.
2026-09-01 16:00:11 -04:00
Ryan Houdek ab6e67596f Merge pull request #5891 from Plagman/plagman/guestextents_mr
DiskCache: hash guest code according to its real extents
2026-09-01 10:19:59 -07:00
LC 433091c326 Merge pull request #5871 from mcfi/pr5814
Resolve symlinks continuously until we find the actual binary name.
2026-09-01 10:12:49 -04:00
Pierre-Loup A. Griffais c01bfcd52f InstcountCI: Update 2026-09-01 03:51:12 -07:00
Pierre-Loup A. Griffais 673a387808 DiskCache: hash guest code according to its real extents
Does less hashing, improves hit rate when there's data adjacent to code,
and/or when the SMC check makes us rebuild code that hasn't actually been
changed.
2026-09-01 01:03:24 -07:00
LC a6e74cdbe6 Merge pull request #5889 from Sonicadvance1/224
FEXCore: Removes syscall optimization
2026-08-31 22:47:00 -04:00
LC ff4da8a620 Merge pull request #5890 from Sonicadvance1/225
FEXUnixlib: Removes unused variable/warning.
2026-08-31 22:43:08 -04:00
Ryan Houdek 1558a2c65d InstcountCI: Update 2026-08-31 19:23:20 -07:00
Ryan Houdek 66cad978c3 FEXCore: Removes syscall optimization
The JIT was doing a bunch of additional work where it was saving and
restoring registers and then juggling the arguments back in to a stack
frame. All of this is nonsensical without the optimization where we
could call syscalls inline without a stack frame.

Instead remove this optimization entirely and behave like a "generic"
syscall path always. The Linux syscall handler now pulls the arguments
out of the CPU context directly and stores the result back in to RAX
directly as well.

This has knock-on effects where technically syscalls are
going to be slightly faster because no stack frame setup for the
arguments, but additionally we are going to be able to have syscalls be
proper serialization points where we can interrupt the syscall and
long-jump out without problems.

Bumps the DiskCache version again because it causes codegen to change.
2026-08-31 19:23:18 -07:00
Ryan Houdek c11a0cef68 FEXUnixlib: Removes unused variable/warning. 2026-08-31 19:09:24 -07:00
Ryan Houdek 8cf2bf7ad8 Merge pull request #5884 from Plagman/plagman/lookup_opt_mr
DiskCache: Lookup optimizations
2026-08-31 14:05:34 -07:00
LC ef5439e5ae Merge pull request #5885 from Plagman/plagman/linux_writerthread_signalblock_mr
DiskCache: on Linux, block signals in Writer thread
2026-08-31 10:18:14 -04:00
Pierre-Loup A. Griffais c0b24b2f2e DiskCache: on Linux, block signals in Writer thread
Clean up thread flags while we're at it.
2026-08-30 22:27:05 -07:00
Pierre-Loup A. Griffais 0d3a5f96f0 InstcountCI: Update 2026-08-30 17:17:54 -07:00
LC 73dc3b3edd Merge pull request #5883 from Sonicadvance1/223
Scripts/fit_native: Classify Oryon-1/3
2026-08-30 20:16:07 -04:00
Pierre-Loup A. Griffais 01bc89b82e DiskCache: Lookup optimizations
- quicker misses from storing GuestHash/GuestSize in index
 - use GuestSize to pull less data from disk on hit
2026-08-30 16:50:11 -07:00
Ryan Houdek 2f663db52c Scripts/fit_native: Classify Oryon-1/3
If the compiler is new enough than classify Oryon-1. Clang doesn't
support Oryon-3 so just claim its Oryon-1 in that case.
2026-08-30 15:36:59 -07:00
LC e17fdf9e90 Merge pull request #5882 from Sonicadvance1/222
InstCountCI: Enforce DiskCache version matching
2026-08-30 17:41:35 -04:00
Ryan Houdek b695c5ae11 InstCountCI: Update
Only adds DiskCache version to json to ensure it's tracked. No asm
changes here.
2026-08-30 14:26:45 -07:00
Ryan Houdek f5935b4006 InstCountCI: Enforce DiskCache version matching
Two added failure modes here. If the disk cache version has changed then
the json files must be updated to the new version to ensure correct
tracking.

Additional failure mode is that codegen actually changed but the disk
cache version hasn't. This is the expected common failure mode and we
need to do additional work before updating json results. This just means
incrementing the disk cache version before updating the instcountci
results. Have a fairly lengthy error message to showcase how much of an
impact this might have.
2026-08-30 14:26:43 -07:00
Ryan Houdek eca6569bd3 FEXCore: Expose DiskCache version to the frontend
With a comment about includes currently being incorrectly shared to the
frontend and needs to get resolved.
2026-08-30 14:26:43 -07:00
LC 51c241d1f8 Merge pull request #5881 from Sonicadvance1/222
UnixLib: Removes old non-unixlib handling
2026-08-30 16:58:07 -04:00
Ryan Houdek 83ead40b78 UnixLib: Removes old non-unixlib handling
Everything that we care about supports the unixlib path now. Also turns
out we were doing `svc #0` on Windows when unixlib didn't exist which is
kind of funny.

Remove the legacy hacky path as it is no longer necessary.
2026-08-30 13:08:52 -07:00
LC dea5c62a44 Merge pull request #5880 from Plagman/plagman/lookup_less_allocs_mr
DiskCache: get rid of more extra copies/allocs on Lookup
2026-08-30 07:57:30 -04:00
Pierre-Loup A. Griffais b3f902166b DiskCache: get rid of more extra copies/allocs on Lookup
Reorganize disk format a bit so that entrypoints and guest pages can be used
as is from the original blob allocation, with some in-place relocation.
2026-08-30 01:52:42 -07:00
Ryan Houdek 98964c5527 Merge pull request #5875 from lioncash/config
Config: Fix assertion being hit in config generation
2026-08-28 15:47:11 -07:00
LC 8ccc8dab49 Merge pull request #5865 from Sonicadvance1/221
FEXCore/Config:  Annotate all config options that can affect codegen
2026-08-28 17:59:11 -04:00
LC e1d8881f57 Merge pull request #5866 from Plagman/plagman/packed_applyrelocs_mr
DiskCache: Apply relocations from disk blob
2026-08-28 17:58:41 -04:00
Ben Niu 9742a560f5 Resolve symlinks continuously until we find the actual binary name.
When a guest executable is invoked through two or more levels of symlink,
FEX expands $ORIGIN in its DT_RPATH/DT_RUNPATH to the directory of the
intermediate symlink rather than the directory of the fully-resolved binary.
Libraries referenced relative to $ORIGIN then fail to load.

This PR fixes the issue by continuously chasing the symlinks until we find
the actual executable name.
2026-08-28 14:20:25 -07:00
Lioncache 7a8b9dd341 Config: Fix assertion being hit in config generation 2026-08-28 16:51:11 -04:00
Ryan Houdek 063af711af Merge pull request #5869 from cjacek/wow64-syscall
WOW64: Update frame state ESP and EIP before unlocking context for syscall
2026-08-28 13:30:51 -07:00
Jacek Caban aed71ed5e9 WOW64: Update frame state ESP and EIP before unlocking context for syscall
Fixes reported CPU context when queried while processing the syscall.
2026-08-28 18:52:12 +02:00
Pierre-Loup A. Griffais 8b9237fd52 DiskCache: Apply relocations from disk blob
Removes one allocation in Lookup path.
2026-08-27 22:45:09 -07:00
Ryan Houdek fd180a16d3 Merge pull request #5861 from Plagman/plagman/cache_map_mr
DiskCache: use memory-mapped IO for reads when possible
2026-08-27 22:18:13 -07:00
Pierre-Loup A. Griffais 2a82f58195 DiskCache: use memory-mapped IO for reads when possible
On Wine, use a unixlib call to get a quality mapping that can track a growing
file.
2026-08-27 21:15:22 -07:00
Ryan Houdek 34ffd2d6e2 FEXCore/Config: Annotate all config options that can affect codegen
Instead serializing the world, allow the config to be data driven. A
couple host feature options aren't serialized as they get explained
elsewhere.
2026-08-27 15:08:20 -07:00
LC 62c9c130c4 Merge pull request #5864 from simon902/pf2id_saturation
Fix pf2id overflow saturation
2026-08-27 09:16:04 -04:00
LC fca23d88c3 Merge pull request #5863 from simon902/vfaddp-mmx-width
Fix PFACC producing wrong results due to missing MMX-sized path in VFAddP
2026-08-27 09:15:30 -04:00
LC 3a179a3a40 Merge pull request #5862 from simon902/pf2iw_saturation
Fix pf2iw to saturate on 32-to-16-bit truncation
2026-08-27 08:47:54 -04:00
Simon Scherer fe17c447d9 InstcountCI: Update 2026-08-27 11:50:52 +02:00
Simon Scherer b21a8dafa7 OpcodeDispatcher: Fix pf2id overflow saturation 2026-08-27 11:50:45 +02:00
Simon Scherer 926a871dd4 unittests/ASM: Test pf2id overflow saturation 2026-08-27 11:40:20 +02:00
Simon Scherer 79ef5031eb InstcountCI: Update 2026-08-27 10:36:31 +02:00
Simon Scherer f0135eb332 JIT/VectorOps: Separate 64-bit VFAddP path for MMX-sized operands 2026-08-27 10:36:22 +02:00
Simon Scherer 566b25a35b unittests/ASM: Test pfacc simple addition 2026-08-27 10:33:05 +02:00
Simon Scherer e6748ea1ed InstcountCI: Update 2026-08-27 09:36:28 +02:00
Simon Scherer 5ab2c723ff OpcodeDispatcher: For PF2IWOp use VSQXTN instead of VUnZip to saturate values falling outside the 16-bit range 2026-08-27 09:36:17 +02:00
Simon Scherer cac8785aca unittests/ASM: Test pf2iw with values falling outside of 16-bit range 2026-08-27 09:29:40 +02:00
LC 0814780c2f Merge pull request #5860 from Plagman/plagman/bucket_cleanup_mr
DiskCache: clean up bucket key building a bit
2026-08-26 21:57:02 -04:00
Pierre-Loup A. Griffais 753bce5e73 DiskCache: clean up bucket key building a bit
Less manual math like that.
2026-08-26 17:35:54 -07:00
Ryan Houdek d04ad79eaf Merge pull request #5857 from 06393993/windows-split-cmake-cxxflags
Windows: Split CMAKE_CXX_FLAGS before passing to execute_process
2026-08-26 16:37:35 -07:00
LC baca86bc3f Merge pull request #5859 from Sonicadvance1/220
Config: Filter the `CPUFeatureRegisters` option on serialization
2026-08-26 15:33:59 -04:00
LC adfff26d58 Merge pull request #5858 from Sonicadvance1/219
LookupCache: Add L1 entry on shrink
2026-08-26 15:33:38 -04:00
Ryan Houdek ef3a6963f3 Config: Filter the CPUFeatureRegisters option on serialization
This is unfiltered data that ends up in HostFeatures. If it was
serialized when set then it would effectively never allow the code
serialization to be used.
2026-08-26 12:22:08 -07:00
Ryan Houdek 19dba4a3ec LookupCache: Add L1 entry on shrink
If we're shrinking the L1 cache then we just deleted the entry that we
just looked up. Add it back to ensure we don't get yet another lookup
for this entry.
2026-08-26 12:18:31 -07:00
Ryan Houdek 49c1cda5c0 Merge pull request #5856 from Hojun-Cho/lookupcache_l1_grow_clear
LookupCache: Clear L1 entries on dynamic cache growth
2026-08-26 12:12:44 -07:00
Kaiyi Li 2ea749c7ea Windows: Split CMAKE_CXX_FLAGS before passing to execute_process
CMAKE_CXX_FLAGS is a space-separated string rather than a semicolon-separated
list. Passing ${CMAKE_CXX_FLAGS} directly to execute_process(COMMAND ...)
passes the entire multi-flag string as a single argv argument to the compiler,
causing option parsing to fail when multiple flags are present (such as
flags configured via the CXXFLAGS environment variable).

Use separate_arguments() to convert CMAKE_CXX_FLAGS into a list so each flag
is passed as an individual argument.
2026-08-26 15:52:58 +00:00
LC 9d4b493ec0 Merge pull request #5852 from Plagman/plagman/keybucket_mr
DiskCache: bucket/key bookkeeping
2026-08-26 10:59:48 -04:00
Hojun-Cho 7e3e4f81ac LookupCache: Clear L1 entries on dynamic cache growth
Growing widens L1PointerMask without touching the table, so an entry whose
address has the new mask bit set sits where InvalidateCache no longer looks,
and a later shrink hands it back to the JIT. Wipe [0, old) on grow, the way
the shrink wipes [new, MAX); the still-valid entries go with it.
2026-08-26 20:51:25 +09:00
Pierre-Loup A. Griffais e740af05b0 DiskCache: bucket/key bookkeeping
Compute a bucket hash and use it in the cache path to avoid grouping entries
that will never make sense together. Don't trust the path, though, and also
lace it into the keys themselves, so that eg. a RO cache miss can never turn
into corruption.

Keep a readable metadata entry at the beginning of the cache, with readable
version, bitness and serialized config.

Use printable characters in FOZ key names as intended.
2026-08-25 22:21:09 -07:00
LC 4ed80fd071 Merge pull request #5855 from Sonicadvance1/218
HostFeatures: Fixes crash in mingw build
2026-08-25 20:37:52 -04:00
Ryan Houdek 59116a06b9 HostFeatures: Fixes crash in mingw build
Turns out packed enum classes without specifying an underlying type
causes problems. Declare its underlying type as uint32_t to match
everything else here.
2026-08-25 17:19:35 -07:00
LC a409220e54 Merge pull request #5854 from Sonicadvance1/217
OpcodeDispatcher: Fixes SHLD by 16 behaviour
2026-08-25 18:33:39 -04:00
Ryan Houdek 485f1c6acf InstcountCI: Update 2026-08-25 15:18:17 -07:00
Ryan Houdek 6646a5cc72 OpcodeDispatcher: Fixes SHLD by 16 behaviour
We were assuming that SHLD undefined behaviour matches SHL, but the
specification actually changes a `ge` comparison to `gt`, which means a
shift of 16 isn't UB!

Thanks to the impeccable @OFFTKP in #5842 for bringing this up as it took a bit
for me to figure out what was actually wrong here. I modified their
unittest to cover more just to ensure we don't break it.
2026-08-25 15:15:36 -07:00
Ryan Houdek 2a76d3b153 Merge pull request #5809 from simon902/fcomi_flag_zero
FCOMI fails to clear OF, SF, AF
2026-08-25 13:22:28 -07:00
LC 7614e48322 Merge pull request #5851 from Sonicadvance1/215
FEXCore/HostFeatures: Adds HostFeatures hashing support
2026-08-25 06:13:21 -04:00
Ryan Houdek 631b8c3f58 FEXCore/HostFeatures: Adds HostFeatures hashing support
As long as the hash is smaller than 64-bits we can just return the bits
encoded directly. Codegen slightly changes with this packed
representation, but doesn't really matter.

Also removes ICacheLineSize as that doesn't actually affect codegen for
us. Once we add 27 more HostFeatures we can switch the hash over to
XXH3.
2026-08-24 20:50:35 -07:00
LC b4b38a92ef Merge pull request #5850 from Sonicadvance1/214
unittests/FEXLinuxTests: Fixes race condition in signal_sra_state
2026-08-24 22:14:15 -04:00
LC c3769a6937 Merge pull request #5849 from Sonicadvance1/213
FEXCore/Context: Removes InitialRIP/RSP from CreateThread
2026-08-24 22:11:34 -04:00
Ryan Houdek 8d69852785 unittests/FEXLinuxTests: Fixes race condition in signal_sra_state
This was using a one second alarm which could catch the JIT while it is
still busy. Instead wait for it to let an alarm thread that it is ready,
then tgkill the thread. This removes the race that this was hitting.
2026-08-24 19:02:39 -07:00
Ryan Houdek b83dd97762 FEXCore/Context: Removes InitialRIP/RSP from CreateThread
We actually never use this anymore, we instead always pass zero for
both, and then rely on the thread inheritance model or setting the
values manually. Now that we expose visibility of the
InternalThreadState to the frontend they just access it directly.

Just a smidge of cleanup, NFC.
2026-08-24 18:32:35 -07:00
LC 22498dc871 Merge pull request #5846 from Sonicadvance1/211
FEXCore/Config: Support Serializing
2026-08-24 20:50:30 -04:00
LC 96172840f1 Merge pull request #5848 from Sonicadvance1/212
code-format-helper: More dependabot upgrades
2026-08-24 20:47:42 -04:00
Ryan Houdek e8880e4e98 code-format-helper: More dependabot upgrades
Only one this time.
2026-08-24 17:07:50 -07:00
Ryan Houdek 26f6cda589 FEXCore/Config: Support Serializing
Serializes every option, even ones that are set to default to ensure
validation that if any config option value is added or changed that they
are captured.

Skips a handful of options that are either meta options, environment
options that don't matter, or HostFeatures which is handled elsewhere.
(HostFeatures will be controlling bucketing rather than the remaining
options).

This serialization is currently 1413 bytes and generates in 17580ns on
my A1A. So it's not the fastest, definitely don't want to be generating
it constantly per process.
2026-08-24 16:10:42 -07:00
LC f0b45010ec Merge pull request #5840 from Sonicadvance1/209
Avoid double offset relocations
2026-08-24 15:44:37 -04:00
Ryan Houdek ad3939a44f Avoid double offset relocations
FEX Relocations now live at an offset from the `CodeData.BlockBegin` of the
code. Regardless of where the relocation moves to, it should always be
relative to that address. This is what makes it PIC compatible.

We were preemptively offsetting the relocation location to be relative
to the memory base in the buffer, which is unnecessary and causes code
caching to basically relocate twice to get the real location.

So in JIT.cpp, stop relocating the offsets, they're already relative to
`BlockBegin`, which is offset 0.

Then when storing the relocation, stop relocating offsets AGAIN because it's
already relative to the code being serialized.

Then when loading the relocations in `CodeCache::ApplyCodeRelocations`
stop relocating offsets YET ANOTHER TIME.

All this is to say that relocation offsets are already PIC and relative
to offset 0, so we don't need to do it three times.
2026-08-24 12:17:11 -07:00
LC c9b23eb0e7 Merge pull request #5845 from Plagman/plagman/thread_priority_mr
DiskCache: make Writer thread low-priority
2026-08-23 18:09:20 -04:00
LC 44f69f4dd7 Merge pull request #5844 from Plagman/plagman/stats_mr
DiskCache: add some SHM stats
2026-08-23 18:08:36 -04:00
Pierre-Loup A. Griffais c3fb6ccaaa DiskCache: make Writer thread low-priority 2026-08-23 14:01:32 -07:00
Pierre-Loup A. Griffais b47a36a47b DiskCache: add some SHM stats 2026-08-23 13:51:01 -07:00
LC af9b438eec Merge pull request #5841 from OFFTKP/blsmsk
unittests/ASM: Test BLSR/BLSMSK CF flag
2026-08-22 23:13:27 -04:00
LC fa9bbf081b Merge pull request #5843 from Sonicadvance1/210
FEXCore: Fixes ever shrinking JIT code buffer
2026-08-22 23:05:34 -04:00
Ryan Houdek db3817a260 FEXCore: Fixes ever shrinking JIT code buffer
I accidentally replaced a couple usages of `AllocatedSize` with
`GetAllocatedSize()`. This resulted in a JIT buffer that ran out of
space would actually allocate a slightly smaller buffer each time, and
then it cascades downwards resulting in catastrophic performance.

Fix the use in `SharedCodeBufferManager.cpp` and `Core.cpp` which were
incorrect and renames the function to be more explicit.
2026-08-22 18:31:22 -07:00
Paris Oplopoios 635befb4c8 FEXCore: Fix CF calculation for BLSMSK and BLSR for 32-bit operands 2026-08-22 15:47:50 +03:00
Paris Oplopoios a69daa2524 unittests/ASM: Test BLSR/BLSMSK CF flag 2026-08-22 15:34:13 +03:00
LC b563703701 Merge pull request #5839 from Sonicadvance1/208
FEX: Only fsync on assert
2026-08-21 19:56:07 -04:00
Ryan Houdek 0717f4689b FEX: Only fsync on assert
The rest of the messages should be buffered, no reason to sync those.
2026-08-21 16:26:14 -07:00
LC bf30f6af4c Merge pull request #5838 from Sonicadvance1/207
JIT: Fixes some alignment logic
2026-08-21 18:07:39 -04:00
Ryan Houdek dfcbdac347 JIT: Fixes some alignment logic
Noticed while taking a look at the relocations that we were technically
not doing alignment before writing down code size.

- Make sure Align16B isn't used with unaligned code with assert
- Switch an `Align` over to `Align(16)` to force 16-byte alignment
  - Without NOP insertion, as this is data at this point, so just zeros.
- Record data size after that alignment
- Remove the `Align16B` that occurred afterwards
  - Previous query between alignments would leave us with up to 12 bytes
    unaccounted for.
- Ensure everything is using the correct sizes by not querying again
- Ensure that emission buffer abuse can't happen by zeroing the buffer.
2026-08-21 14:55:03 -07:00
LC 8d12f3d5aa Merge pull request #5837 from Sonicadvance1/206
SharedCodeBufferManager: Leak less internal details about implementation
2026-08-21 17:31:30 -04:00
Ryan Houdek 02f8ab5f87 SharedCodeBufferManager: Leak less internal details about implementation
The various places that were using the CodeBuffer object were using
internal implementation details that are changing as we move over to a
bitmap allocator.

Preempt this by hiding some of the implementation details early without
changing behaviour. `GetBufferBase` is still technically leaking some of
the internal details, but it needs changes around how relocations are
being handled and how the disk cache validation works in order to handle
that right now.

Should be no functional change.
2026-08-21 14:06:12 -07:00
LC b2b6b263bb Merge pull request #5835 from Plagman/plagman/cache_thread_mr
DiskCache: offload Store to a WorkQueueThread
2026-08-21 14:36:39 -04:00
Pierre-Loup A. Griffais e3f208f61f DiskCache: offload Store to a WorkQueueThread
With all Stores happening on the same thread now, we can also make locking
more granular for extra perf. Move to positioned IO for everything, as we
can't reliably track the cursor with that faster locking model.

Add some bounds checking to index population to protect against corruption.
2026-08-21 11:16:07 -07:00
Pierre-Loup A. Griffais 008d990f20 Utils: add WorkQueueThread
Straightforward queue for arbitrary work.
2026-08-21 11:16:07 -07:00
Pierre-Loup A. Griffais c038bb5794 Windows: add Thread implementation
So we can make a worker thread that the guest (hopefully) won't see.
2026-08-21 11:16:07 -07:00
Pierre-Loup A. Griffais efc5eaf9c7 Utils/File: add PRead(), PWrite() and Size()
Because PRead/PWrite don't have the same side effects on the file cursor
between Linux and Windows (:/), make them all-or-nothing.
2026-08-21 11:16:07 -07:00
Ryan Houdek c1e9d19809 Merge pull request #5836 from cjacek/int-3
Windows: Handle interrupt 3 in HandleGuestException
2026-08-21 09:30:01 -07:00
Jacek Caban 85efb47a3b Windows: Handle interrupt 3 in HandleGuestException 2026-08-21 13:16:43 +02:00
LC bc69693b6c Merge pull request #5834 from Sonicadvance1/205
FEXCore/SharedCodeBufferManager: Pivot what tracks memory allocations
2026-08-20 19:47:47 -04:00
Ryan Houdek c9c5a75b76 FEXCore/SharedCodeBufferManager: Pivot what tracks memory allocations
It's soon going to change how these buffers are managed, where the
CodeBuffer is going to manage its own allocations soon once it changes
over to the bitmap allocator. Additionally the Manager class is actually
going to do proper management, pooling, and invalidation handling.

Split the task preemptively before we switch to the bitmap allocator to
reduce churn. A little change in the CodeCache where it needs to query
the codebuffer directly rather than the context, but fairly safe.

Shouldn't be any real behaviour change.
2026-08-20 16:35:28 -07:00
Ryan Houdek f50279a2e7 Merge pull request #5832 from Plagman/plagman/cache_mr
Disk Cache initial implementation
2026-08-20 12:06:59 -07:00
Pierre-Loup A. Griffais 561c32b45d Disk Cache initial implementation
Serializes code blocks to disk - only blocks coming from known regions, for now

Disabled by default, key and versioning still needs work, but works for testing
2026-08-19 18:20:49 -07:00
Ryan Houdek f42dc71972 Merge pull request #5831 from cjacek/int-assert
Windows: Handle 0x2c interrupt in HandleGuestException
2026-08-18 16:39:39 -07:00
Jacek Caban ea67665bbf Windows: Handle 0x2c interrupt in HandleGuestException 2026-08-19 00:41:07 +02:00
LC c3b4d4b7bb Merge pull request #5829 from Sonicadvance1/204
unittests: Disable siglongjmp_branch_invalid on 32-bit
2026-08-17 23:21:41 -04:00
Ryan Houdek 492dac719f unittests: Disable siglongjmp_branch_invalid on 32-bit
This test is trying to execute "invalid" code from the last two pages of
the address space to the first two pages of the address space. But
failed to noticed that the last two pages of the 32-bit x86 address
space are actually valid, usually containing VDSO things.
2026-08-17 18:22:03 -07:00
Ryan Houdek 9377bac5e7 Merge pull request #5825 from javelina-pkwy/fix/siglongjmp-deadlock
FEXCore: prevent deadlock when branching to MAX_UINT64
2026-08-17 17:55:39 -07:00
Justin Becker e0bbfdda84 Cleaner change 2026-08-17 12:05:24 -07:00
Ryan Houdek 9618b5adef Merge pull request #5828 from Sonicadvance1/203
Steam/CompatTool: Become more picky about configs
2026-08-17 10:22:19 -07:00
Ryan Houdek a3d609ebb1 Merge pull request #5820 from javelina-pkwy/lea-reg-reg
Decoder: fix illegal LEA encoding
2026-08-17 09:58:47 -07:00
Ryan Houdek 2b0d94536f Steam/CompatTool: Become more picky about configs
If `STEAM_COMPAT_FEX_CONFIG` is missing options, then instead of having
an opinion about what those options should be, just leave them unset.
This allows FEX's regular default option handling to kick in for missing
configuration options.

Where previously if an option was missing from the config, it would
default to boolean false, which may or may not be the default depending
on option.
2026-08-17 09:34:25 -07:00
LC 73ab3bc56d Merge pull request #5826 from cjacek/crt-printf
Windows/CRT: Add printf and puts stubs
2026-08-16 10:07:21 -04:00
Jacek Caban 412c49a8ce Windows/CRT: Add printf and puts stubs
Fixes PE builds with VIXL disassembler enabled.
2026-08-16 15:14:36 +02:00
Justin Becker 6f29dfcbb8 Probe before taking lock in Compile*() 2026-08-13 16:53:05 -07:00
LC f3ab82a73f Merge pull request #5823 from Sonicadvance1/202
SharedCodeBufferManager: Allocate JIT space atomically.
2026-08-13 18:15:59 -04:00
Ryan Houdek 6734c9ed3e SharedCodeBufferManager: Allocate JIT space atomically.
This removes the fairly long lived lock that the buffer allocator held
while doing significantly more work than intended while holding that
lock.

As the first step towards moving over to the atomic bitmap allocator,
change this to be atomic to closer match what the new allocator is
doing. Since we are just doing linear allocations, this is an easy
convert and should give a good stutter improvement.
2026-08-13 14:58:01 -07:00
Ryan Houdek 71afe47675 Merge pull request #5821 from Sonicadvance1/200
Thunks: Adds some new PV paths
2026-08-12 22:47:56 -07:00
LC 7075377a63 Merge pull request #5822 from Sonicadvance1/201
clang-format: Slight whitespace difference
2026-08-12 20:13:16 -04:00
Ryan Houdek 7c1036df09 clang-format: Slight whitespace difference 2026-08-12 16:55:38 -07:00
Ryan Houdek a312347589 Thunks: Adds some new PV paths
Slight PV behaviour changes meant we missed this.
2026-08-12 16:54:01 -07:00
LC f386c62dba Merge pull request #5815 from Sonicadvance1/198
FEXCore/Utils: Implements a new atomic segmented bitmap allocator
2026-08-12 17:58:19 -04:00
Ryan Houdek 30a81484d7 FEXCore/unittests: Adds tests for atomic bitmap allocator 2026-08-12 14:32:31 -07:00
Ryan Houdek 00a6b046a8 FEXCore/Utils: Implements a new atomic segmented bitmap allocator
This is tailored towards our needs for our JIT and eventually replacing
the linear allocator. Allowing us to reallocate memory for code blocks
that have been invalidated, letting us keep a single code buffer around
for longer and using less memory overall.

In particular, high-invalidation games that ship anti-tamper tend to
emit millions of ~128-byte blocks in just a handful of minutes which
causes our current linear allocator to consume gigabytes very quickly.
This will allow us to more aggressively reuse the allocation space and
reduce the memory load in those situations.

There's some additional resize tuning that needs some work whence it is
in situ which doesn't need to be done now.
2026-08-12 14:32:31 -07:00
Justin Becker c156498c5c Add 16 bit and 32 bit variants 2026-08-11 16:19:49 -07:00
LC adea3e410f Merge pull request #5793 from Sonicadvance1/193
FEXCore: Adds a lock-free atomic bitset that supports contiguous range allocations
2026-08-10 17:24:37 -04:00
Ryan Houdek 40940ae0b3 FEXCore/unittests: Adds atomic bitset unittest
Hammers the API in a couple of ways to make sure it works.
2026-08-10 14:12:10 -07:00
Ryan Houdek 1bab28dad4 FEXCore: Adds a lock-free atomic bitset that supports contiguous range allocations
This thing is a bit intense, so some requirements from the start:
- It needs to be lock-free and thread-safe
- It needs to support contiguous range allocations
- It needs to support allocations larger than a single atomic word

These requirements kind of fly in the face of most bitset allocators
where they will support some parts of these requirements, or just throw
a mutex in front of the whole thing.

Some implementation details:
- If allocating only 1-bit, trivial and always succeeds if there is space
- If allocating <= 64-bit, then always succeeds if there is at least
  those many contiguous bits within a single atomic word
  - Allocation can fail if there are cross-word contiguous bits of the
    size available
  - Introduces some sparsity
- If allocating > 64-bits then it falls down the longer scan path.
  - Searches for contiguous bits of free space between multiple atomic
    words.
  - If found, will attempt to allocate tracking which bits were allocated
  - If allocation fails, unwind bits already acquired and continue
    scanning

Some downsides to this implementation:
- Allocations can fail if sparsity builds up
- Heavily contended allocations can be worse than a lock
  - If larger than atomic word allocations are in flight.
- Unwinding larger than word allocations and continuing scanning adds
  overhead, a lock would have won at that point.
- A small bit of false sharing where an atomic word is read without
  acquire semantics for scanning can technically overlook some
  allocations that no longer exist.
- Slower than a linear allocator, but that's not unexpected.

Most of these downsides are okay for our use case, which is code buffer
allocations with the ability to do partial invalidation. If the atomic
bitset fails to fit an allocation, we can throw away the code buffer
like we currently do.

The bitmap allocator that uses this lock-free atomic bitset is still
in-flight but this is one complex container that can land independently.
2026-08-10 14:12:09 -07:00
Ryan Houdek 4838265589 FEXCore/MathUtils: Adds helper for alignment by power of 2 size
Useful for removing integer division instructions when we know the
source value is aligned to be power of two. As integer division is quite
slow, we want to use this when possible.
2026-08-10 14:08:07 -07:00
Ryan Houdek f6d20a1a88 Merge pull request #5819 from cjacek/clang-warnings
Fix warnings in llvm-mingw builds
2026-08-10 10:37:40 -07:00
Jacek Caban 11be444d45 toolchain_mingw: Don't use -static-libgcc -static-libstdc++
Those are not supported by Clang and cause -Wunused-command-line-argument warnings.
They are also redundant when -static is used.
2026-08-10 15:05:26 +02:00
Jacek Caban 363bf85b5b LongJump: Silence -Winline-asm warnings on ARM64EC
Fixes disallowed registers warnings.
2026-08-10 15:05:26 +02:00
Jacek Caban 720039ca6d CRT: Avoid implicit char to unsigned char casts in strtoumax
Fixes -Wpointer-sign warnings.
2026-08-10 15:05:26 +02:00
Jacek Caban 4e9399e029 CRT: Silence -Wunused-but-set-variable warnings in libm.h 2026-08-10 15:05:26 +02:00
Jacek Caban 466ee3cbcf CRT: Use DLLEXPORT_FUNC for more functions
Fixes locally defined symbol imported linker warnings.
2026-08-10 15:05:26 +02:00
Jacek Caban 390d5d2fb4 OfflineCompiler: Avoid unused function in win32 builds
Fixes -Wunused-function warning.
2026-08-10 15:05:26 +02:00
Jacek Caban b543df8299 HostFeatures: Avoid unused function in win32 builds
Fixes -Wunused-function warning.
2026-08-10 15:05:26 +02:00
Jacek Caban de16c18961 SignalScopeGuards: Mark template function helpers as inline
Fixes -Wunused-template warnings.
2026-08-10 15:05:26 +02:00
Jacek Caban 2158309c51 InterpreterFallbacks: Remove unused template
Fixes -Wunused-template warning.
2026-08-10 15:05:26 +02:00
Ryan Houdek 430846d7f2 Merge pull request #5818 from OFFTKP/ffreep
unittests/ASM: Ensure ffreep always has an operand
2026-08-09 10:46:43 -07:00
Paris Oplopoios f4e362c477 unittests/ASM: Ensure ffreep always has an operand 2026-08-09 16:46:13 +03:00
LC 7966bfb077 Merge pull request #5816 from Sonicadvance1/199
Misc: Adds some missing headers
2026-08-09 06:19:59 -04:00
LC a4e04dd7b4 Merge pull request #5808 from Sonicadvance1/197
FEXCore: Fixes SourceOutline description
2026-08-09 06:19:41 -04:00
Ryan Houdek b161a74365 Misc: Adds some missing headers
Newer compiler and libraries got angry that these were missing.
2026-08-07 14:52:00 -07:00
Ryan Houdek fd141ed6d7 Merge pull request #5811 from Claudemirovsky/fix/compilation/archlinux-mingw-llvm
CMake: Fix MINGW compilation under ArchLinux
2026-08-06 22:00:10 -07:00
Claudemirovsky e2ff0dd1cd CMake: Fix MINGW compilation under ArchLinux 2026-08-07 01:15:42 -03:00
Simon Scherer 6284c39eb3 InstcountCI: Update 2026-08-06 16:31:49 +02:00
Simon Scherer 7f0bdf8d63 OpcodeDispatcher: Zero OF, SF and AF for FCOMI and FCOMIF64 2026-08-06 16:29:29 +02:00
Simon Scherer 015beff4cc unittests/ASM: Test OF,SF,AF flags for fcomi 2026-08-06 16:14:27 +02:00
Ryan Houdek 92e43c25d4 Merge pull request #5807 from Sonicadvance1/196
CPUID: Adds a few new bits
2026-08-05 19:56:21 -07:00
Ryan Houdek 69fe85274f Merge pull request #5801 from Sonicadvance1/94
Scripts: Fixes failure in doc_outline_generator
2026-08-05 19:56:10 -07:00
Ryan Houdek 0122ef9e83 FEXCore: Fixes SourceOutline description 2026-08-05 19:53:33 -07:00
Ryan Houdek c772c0e4e7 Merge pull request #5794 from Sonicadvance1/194
Win32: Actually set app config path and name
2026-08-05 19:50:37 -07:00
Ryan Houdek c171c06192 Merge pull request #5806 from simon902/F64ToI32Precision
Fix precision loss for F64 to i32 conversion
2026-08-05 14:40:41 -07:00
Ryan Houdek 9365e6240b CPUID: Adds a few new bits
The two page-size extensions are a nop so might as well as enable them.
For the debug flag, we already set the duplicated flag in 8000_0001.edx, but missed this one.
Doesn't add anything new for the FEX side, but Burnout Paradise (and
remastered) is incorrectly checking for SSE2 support by checking if this is set.

Closes #5805 although their (ML?) write-up was incorrect.
2026-08-05 14:20:56 -07:00
Simon Scherer f8e3571c66 JIT: Skip redundant rounding in Vector_F64ToI32 2026-08-05 13:20:28 +02:00
Simon Scherer 6ebcf65451 JIT: Fix f64->i32 precision loss in non-SVE path in Vector_F64ToI32 2026-08-05 13:17:58 +02:00
Simon Scherer 4844729e93 unittests/ASM: Test lossy precision non-SVE path in Vector_F64ToI32 2026-08-05 13:13:04 +02:00
Ryan Houdek 9686454161 Merge pull request #5804 from sunshineinabox/PR_thunkgen
thunkgen: satisfy the clang 22 ComplierInstance VFS invariant.
2026-08-04 23:51:02 -07:00
sunshineinabox 167d79f0fa hunkgen: satisfy the clang 22 ComplierInstance VFS invariant. 2026-08-04 23:37:11 -07:00
Ryan Houdek b1275edb63 Scripts: Fixes failure in doc_outline_generator
While it would be better to fix the error in the source, it shouldn't be
a case of blocking release. Print the line that was failed to parse and
then continue onwards.

In particular hit by `ERROR:root:Failure to parse Thread shared code buffer management`
2026-08-04 17:11:10 -07:00
Ryan Houdek e869aa644a Docs: Update for release FEX-2608 2026-08-04 16:55:43 -07:00
Ryan Houdek 68740b3c65 Merge pull request #5792 from Sonicadvance1/192
CI: Disable ranges-v3 from trying to build native
2026-08-03 15:30:52 -07:00
Ryan Houdek 4caad9bf55 Merge pull request #5800 from Sonicadvance1/195
#5795 but with clang_format
2026-08-03 15:30:14 -07:00
Iaying 7629323548 Fix an XMM register bug in SpillSRA, along with adding a test that reproduces the bug 2026-08-03 15:15:41 -07:00
Ryan Houdek 681636cd68 Win32: Actually set app config path and name
Apparently we never set this and it happened to not be a problem. I
needed it to gather some data so fix it.
2026-07-30 20:51:29 -07:00
Ryan Houdek 2fdbff3d1c Merge pull request #5790 from mstorsjo/libc++23
Fix building for Windows with libc++ 23
2026-07-30 14:29:13 -07:00
Ryan Houdek 5c7df98768 CI: Disable ranges-v3 from trying to build native
We don't want this.
2026-07-30 13:59:44 -07:00
Martin Storsjö efe38f2ee2 Add more function stubs for libc++ 23 on Windows
These are needed by libc++ when targeting Windows since
https://github.com/llvm/llvm-project/commit/8a531c3608c722ad529be448d6ecef06ba107228,
which is included in libc++ 23.

In a very brief test, it seems like we don't need to actually
implement them.
2026-07-28 23:37:53 +03:00
Martin Storsjö 08031a2767 Add missing includes
This fixes compilation with libc++ 23, which has removed a number
of unnecessary transitive includes in its headers.

Include <cstdlib> in StringConv.h for std::strtoll and std::strtoull.

Include <cstdlib> for the declarations of malloc/free/realloc/calloc
in Alloc.cpp. (Without this, the functions we define end up with
C++ name mangling.)

Include <stdarg.h> in IO.cpp for va_start/va_end.
2026-07-28 23:17:43 +03:00
Justin Becker 34a87cc94c Decoder: fix illegal LEA encoding 2026-07-27 15:48:35 -07:00
Ryan Houdek d295d9f08e Merge pull request #5789 from lioncash/sig
SignalDelegator: Remove unused Required parameter in handler setting
2026-07-27 13:14:14 -07:00
Ryan Houdek 61d033f21e Merge pull request #5788 from lioncash/config
Config: Minor cleanup
2026-07-27 13:09:21 -07:00
Ryan Houdek 27315e33ab Merge pull request #5784 from FrontMage/fix/instruction-fetch-fault-priority
Frontend: Prioritize instruction fetch faults
2026-07-27 12:54:53 -07:00
LC e8b5cd18bd SignalDelegator: Remove unused Required parameter in handler setting
The required flag is set by the subsequent frontend functions that
follow the calls to these functions.
2026-07-27 15:06:24 -04:00
LC 1a77f41846 Config: Add missing override specifier 2026-07-27 14:01:11 -04:00
LC d1947f715c Config: Move strings in constructor where applicable
Same thing, just a little less memory churn.
2026-07-27 14:00:06 -04:00
LC 3f7971b7d2 Config: Mark internally linked where applicable
Makes it obvious these aren't supposed to be exposed.
2026-07-27 13:56:53 -04:00
LC 2cbd5cb3e6 Config: Amend prototypes where applicable
Previously, these didn't match up with the implementation (luckily it's
only used internally at the moment).
2026-07-27 13:54:25 -04:00
FrontMage 151b4d4c2d Frontend: Prioritize instruction fetch faults 2026-07-25 09:01:21 +08:00
Ryan Houdek 7a3fdefafb Merge pull request #5786 from OFFTKP/fist
Extend FIST tests to check for indefinite value
2026-07-24 09:27:22 -07:00
Paris Oplopoios 86c20d0519 Extend FIST tests to check for indefinite value 2026-07-24 17:04:10 +03:00
Ryan Houdek 464ec9d0bc Merge pull request #5783 from FrontMage/fix/inactive-jit-guard-range
FEXCore: Ignore inactive JIT guard ranges
2026-07-23 17:50:51 -07:00
FrontMage 35518a5fa0 FEXCore: Ignore inactive JIT guard ranges 2026-07-24 07:57:48 +08:00
Ryan Houdek d028c7942b Merge pull request #5782 from lioncash/validation
IRValidation: Minor cleanups
2026-07-23 15:44:48 -07:00
LC 585286a617 IRValidation: Remove unused members from BlockInfo
HasExit is assigned to but never used, but we check this condition a
different way right after leaving the main loop anyway.
2026-07-24 16:00:20 -04:00
LC 4904fd43e9 IRValidation: Make BlockInfo private
This isn't used outside the context of the pass.
2026-07-24 16:00:20 -04:00
LC 9e3f287c1f IRValidation: Use C instead of CW
This op isn't mutated anywhere in the pass.
2026-07-24 16:00:20 -04:00
LC f5e4e26e08 IRValidation: Move var closer to usage
Same behavior, just more compact.
2026-07-24 16:00:20 -04:00
LC 6c5a39e164 IRValidation: Turn ORs with true into assignment
These are just unconditional setting to true anyway.
2026-07-24 16:00:17 -04:00
Ryan Houdek 7469fdb0d6 Merge pull request #5781 from FrontMage/fix/multiblock-block-local-errors
FEXCore: Isolate multiblock error state per block
2026-07-23 15:11:13 -07:00
FrontMage fcf9fd77d7 FEXCore: Isolate multiblock error state per block 2026-07-23 16:52:19 +08:00
Ryan Houdek 0589d9b872 Merge pull request #5779 from lioncash/x87
x87StackOptimizationPass: Minor cleanup
2026-07-22 18:58:17 -07:00
LC 6d4c80adff x87StackOptimizationPass: Remove IR member
This is only used in the store helpers, so we can just pass it in
directly
2026-07-23 02:26:46 -04:00
LC 56dd470528 x87StackOptimizationPass: Remove unnecesary return in Run()
It's a void function, so we don't need this at the end
2026-07-23 02:21:26 -04:00
LC 6741f53d87 x87StackOptimizationPass: Mark getValidMask()/getInvalidMask() as const
These don't modify instance state.
2026-07-23 02:18:33 -04:00
LC 19550c5417 x87StackOptimizationPass: Pass by const reference in setTop()
Avoids redundant copies. Just a minor codegen saving.
2026-07-23 02:17:02 -04:00
Ryan Houdek cfa3dfaac7 Merge pull request #5780 from mrpippy/unicode
Windows: Fixes around using Unicode functions
2026-07-22 18:48:52 -07:00
Brendan Shanks 0ed0bc1dd5 CMake: Define UNICODE when building for Windows 2026-07-22 15:17:17 -07:00
Brendan Shanks 074743d6e8 Windows: Explicitly use *A/*W Win32 functions 2026-07-22 15:16:44 -07:00
Brendan Shanks 8b446de059 Windows: Use GetModuleHandleW() to avoid unnecessary string conversions 2026-07-22 15:16:44 -07:00
Ryan Houdek f2e35f336f Merge pull request #5778 from lioncash/buf
SharedCodeBufferManager: Minor header tidying
2026-07-21 09:18:25 -07:00
LC 10a0fe2e71 SharedCodeBufferManager: Make AllocateNew() signature consistent with declaration 2026-07-23 01:42:41 -04:00
LC c9add0d292 SharedCodeBufferManager: Hoist prctl define into util header
Same behavior, but just moves the potential define to be alongside all
of the others in the wrapper header.
2026-07-23 01:40:49 -04:00
LC c239d09ea0 SharedCodeBufferManager: Add missing header
Ensures the page size define is always visible.
2026-07-23 00:16:56 -04:00
LC a652a5811b Merge pull request #5777 from Sonicadvance1/191
FEXCore: Split out CodeBuffer management to its own file
2026-07-21 07:31:13 -04:00
Ryan Houdek d2c92808f5 FEXCore: Split out CodeBuffer management to its own file
NFC

- Renames CodeBufferManager to SharedCodeBufferManager to be more
  explicit about it being shared between threads
- Renames `CodeBuffers` to `SharedCodeBuffers` to make it more explicit
  about sharing these buffers between threads.
- Separates the Manager to its own file so it is distinct from the rest
  of the CPUBackend code

Makes it easier to parse ownership and lifetime semantics of these
buffers.
2026-07-20 18:09:29 -07:00
LC 2464633431 Merge pull request #5776 from Sonicadvance1/190
JIT: Remove JIT detection string
2026-07-20 21:07:23 -04:00
LC 15e76e88b4 Merge pull request #5775 from Sonicadvance1/189
JIT: Rename temporary CPU buffer allocator
2026-07-20 20:50:56 -04:00
Ryan Houdek fe1ac1bc1d JIT: Remove JIT detection string
Now that we have VMA region naming enabled on JIT buffers, this is no
longer used. Confirming a region is a JIT buffer is now just a case of
comparing the name that shows up in `/procfs/maps` rather than dumping
the first bytes of an unknown region.
2026-07-20 17:44:15 -07:00
Ryan Houdek 9edd27b214 JIT: Rename temporary CPU buffer allocator
`TempAllocator` was a bit too opaque as to what the allocator was for,
so I kept needing to lookup its usage every couple of months. Rename it
to `TempCodeBufferAllocator` so I can remember that it is a temporary
allocator for the staging JIT code buffer more easily.

NFC
2026-07-20 17:32:03 -07:00
Ryan Houdek eb7e02ea1d Merge pull request #5772 from lioncash/pass
PassManager: Simplify initialization interface
2026-07-19 16:28:07 -07:00
LC 53befc68c9 PassManager: Ensure GetPass() only queries the underlying pass mappings
Previously this would create an entry in the map if it didn't exist.
2026-07-21 12:10:26 -04:00
LC 19f95d89ec PassManager: Add basic documentation 2026-07-21 12:10:26 -04:00
LC ecb9b3b7f8 PassManager: Constrain GetPass() template to Pass-derived objects
Makes the particular conversion types constrained to catch any trivial
misuses.
2026-07-21 12:10:26 -04:00
LC d619e36523 PassManager: Pass string by const reference where applicable
Gets rid of potential extraneous copies. We also add handling for cases
where two passes with the same name are unintentionally added.
Previously we'd blindly overwrite the mapping.
2026-07-21 12:09:16 -04:00
LC fa80d11960 PassManager: Remove SyscallHandler member
This isn't used anymore, so we can get rid of it to further simplify
initialization.
2026-07-21 11:28:12 -04:00
LC 2893d2b64f PassManager: Simplify pass initialization
We don't conditionally add any passes, so we can simplify the interface
so that we just add all existing passes at once. Makes the core
initialization process a little more straightforward.
2026-07-21 11:28:09 -04:00
Ryan Houdek 99b8df4e6f Merge pull request #5773 from lioncash/fdres
ThreadManager: Fix error return values in FrontendAllocateSlots()
2026-07-19 16:14:50 -07:00
LC 331655182d ThreadManager: Fix error return values in FrontendAllocateSlots()
If ftruncate or the mmap ever fail for whatever reason, then we need to
return the current size, rather than the new size.
2026-07-21 15:12:41 -04:00
Ryan Houdek ec95330dcd Merge pull request #5765 from lioncash/signal
SignalDelegator: Group members together
2026-07-19 16:13:26 -07:00
LC 3959863462 SignalDelegator: Group members together
Hides all public members and situates all of them together to make for
an easier overview.
2026-07-19 21:10:35 -04:00
LC c00b2c9584 Merge pull request #5774 from Sonicadvance1/93
gitlab: Fixes CI
2026-07-19 18:43:37 -04:00
Ryan Houdek 6c354f3987 gitlab: Fixes CI 2026-07-19 12:23:46 -07:00
Ryan Houdek 3bd4d244a4 Merge pull request #5771 from lioncash/fmt
Externals: Update fmt to 12.2.0
2026-07-19 11:50:58 -07:00
LC 9888de25fe Externals: Update fmt to 12.2.0
Keeps fmt up to date.
2026-07-21 09:32:23 -04:00
Ryan Houdek 58c247b30c Merge pull request #5770 from lioncash/cast
CPUBackend: Remove unnecessary reinterpret_casts
2026-07-19 00:57:29 -07:00
LC 22bd10f3b1 CPUBackend: Remove unnecessary reinterpret_casts
This both take a void*, so the casting is unnecessary to begin with,
since this would occur anyway without it. We can also avoid a
duplication to reduce line noise.
2026-07-21 08:22:17 -04:00
Ryan Houdek f374b4775a Merge pull request #5769 from lioncash/bound
Core: Remove unnecessary bounds check in GenerateIR()
2026-07-18 21:35:34 -07:00
LC ff213bbc5e Core: Move vars closer to usage scope in GenerateIR()
Makes it so their purpose is more easily seen
2026-07-21 04:44:47 -04:00
LC 56a4ca6e6a Core: Remove unnecessary bounds check in GenerateIR()
We already check the bounds in the loop prior to calling at().
2026-07-21 04:40:22 -04:00
Ryan Houdek d6b38b6b1c Merge pull request #5768 from lioncash/stream
IRDumper: stringstream -> ostringstream
2026-07-18 21:33:42 -07:00
LC 04d06d386f IRDumper: stringstream -> ostringstream
These are purely output operations, so we don't need to use the more
heavyweight class.
2026-07-21 04:29:03 -04:00
LC aa26a780ed Merge pull request #5767 from Sonicadvance1/188
AVX128: Optimize 256-bit vmovmaskpd as well
2026-07-17 16:56:23 -04:00
Ryan Houdek f73b93dbc2 InstcountCI: Update 2026-07-17 13:17:39 -07:00
Ryan Houdek c4a5ac892f AVX128: Optimize 256-bit vmovmaskpd as well
Similar to #5757, but once the elements have been zipped together, we
can treat it identically to the 128-bit 32-bit element path.

Closes #3782
2026-07-17 13:15:47 -07:00
Ryan Houdek f129ca0c61 unittests/vmovmskpd: Extend test to have different lower and upper results between 128-bit lanes. 2026-07-17 13:12:27 -07:00
Ryan Houdek 941f0fbf8d Merge pull request #5764 from lioncash/alloc
LinuxAllocator: Reduce MemAllocator32Bit size by 16 bytes
2026-07-17 08:00:25 -07:00
LC 6e37dca566 LinuxAllocator: Reduce MemAllocator32Bit size by 16 bytes
These constants don't need to be member vars.
2026-07-19 18:37:38 -04:00
Ryan Houdek 1cffa009fe Merge pull request #5763 from lioncash/const
IREmitter: Mark some helpers as const
2026-07-17 07:59:47 -07:00
Tony Wasserka 83f4c9d101 Merge pull request #5760 from neobrain/fix_emitter_constants
Arm64Emitter: Fix incorrect condition for constant NOP padding
2026-07-17 13:05:50 +02:00
Tony Wasserka c0c95da796 Arm64Emitter: Fix incorrect condition for constant NOP padding
This needs to be enabled when *generating* caches, not at runtime when we're
loading them (unless we're compiling for validation).

Previous code would incorrectly disable NOP padding in FEXOfflineCompiler and
instead enable it at runtime when it wasn't needed.
2026-07-17 12:47:51 +02:00
LC 8b612a87c6 IREmitter: Mark some helpers as const
These don't modify internal state.
2026-07-17 03:20:36 -04:00
Ryan Houdek 2ac95f446b Merge pull request #5762 from lioncash/core
FEXCore: Resolve missing prototype warnings
2026-07-17 00:04:47 -07:00
LC c5eddd922d FEXCore: Resolve missing prototype warnings
Makes sure we mark everything internally linked as necessary, or make
declarations visible to their implementation.
2026-07-17 02:48:45 -04:00
Ryan Houdek b58be2c073 Merge pull request #5761 from lioncash/sys
LinuxEmulation: Resolve missing prototype warnings
2026-07-16 22:00:10 -07:00
LC a3d0b67777 LinuxEmulation: Resolve missing prototype warnings
Ensures all functions are marked whether they're intended to be
internally linked or not.
2026-07-17 00:12:28 -04:00
Ryan Houdek 0467d523c0 Merge pull request #5759 from neobrain/fix_codebuffer_max_size
CodeCache: Use maximal code buffer size when generating code caches, too
2026-07-16 13:53:42 -07:00
Ryan Houdek b0af054a95 Merge pull request #5757 from MoonFlowww/avx128-vmovmsk-256
AVX_128: Optimize VMOVMSKPS from 11 to 7 instructions
2026-07-16 13:53:00 -07:00
Ryan Houdek 6846f10510 Merge pull request #5758 from lioncash/x87
x87StackOptimizationPass: Make use of std::array for FixedSizeStack
2026-07-16 12:39:08 -07:00
Ryan Houdek e24f232f52 Merge pull request #5756 from lioncash/fill
Arm64Emitter: Pull FillSpecialRegs bools into a struct
2026-07-16 12:36:31 -07:00
Ryan Houdek a1cc5d034e Merge pull request #5755 from lioncash/host
HostRunner: Tidy up interface
2026-07-16 12:35:18 -07:00
Ryan Houdek b478f54aea Merge pull request #5754 from lioncash/vdso
VDSO_Emulation: Mark relevant members as internally linked
2026-07-16 12:34:27 -07:00
Ryan Houdek 7bb380a086 Merge pull request #5753 from lioncash/pipe
FEXServer: Fix some missing declaration warnings
2026-07-16 12:33:54 -07:00
Tony Wasserka 228c351396 CodeCache: Use maximal code buffer size when generating code caches, too
This is less likely to happen, but will still be required for very large libraries.
2026-07-16 16:37:18 +02:00
LC 4bb675a530 x87StackOptimizationPass: Reduce noise in slow push/pop paths
Deduplicates the repeated rotate behavior.
2026-07-16 10:31:40 -04:00
LC 8d62773570 x87StackOptimizationPass: Make helpers internally linked
Makes it obvious they're only used in this TU and allows the compiler to
warn if they ever become unused.
2026-07-16 09:56:14 -04:00
LC 9e26c57643 x87StackOptimizationPass: Fix isValid()
Previously this wouldn't have worked, since .first isn't a valid member.
The only reason it wasn't caught is because the function is never
instantiated.
2026-07-16 09:56:14 -04:00
LC 35bc502062 x87StackOptimizationPass: Make use of std::array for FixedSizeStack
Reduces the overall generated code for state management.

Drops the overall text size from 11447294 to 11441918
2026-07-16 09:56:05 -04:00
moonfloww 7189e1e280 InstcountCI: Update 2026-07-16 14:44:04 +02:00
LC 4b4aa1cdbe Arm64Emitter: Pull FillSpecialRegs bools into a struct
Makes this easily expandable over time without modifying the prototype,
and lets us be a little more informative at call sites.
2026-07-16 08:07:57 -04:00
moonfloww 3680282b30 new vmovmsk from 11 to 7 ins. 2026-07-16 13:57:12 +02:00
LC a4d5fbae63 HostRunner: Tidy up interface
We've accumulated a bunch of forward declarations that are no longer
necessary. We also don't need to pass the signal delegator as a
reference, since we're not modifying the pointer itself, it's just
passed in to register a signal handler.
2026-07-16 07:39:03 -04:00
LC 2943b87c83 VDSO_Emulation: Mark relevant members as internally linked
Silences missing prototype warnings and makes it obvious they're only
used within the translation unit.
2026-07-16 07:28:50 -04:00
LC 9575506391 FEXServer: Fix some missing declaration warnings
Makes sure prototypes are visible to their implementation. Also marks
functions internally linked where applicable.

Also makes it a little more visibly obvious which bits are exposed for
use elsewhere.
2026-07-16 07:10:32 -04:00
Ryan Houdek 31c2449d6e Merge pull request #5752 from lioncash/config
FEXGetConfig: Add convenience option for dumping system/tso info
2026-07-15 09:48:49 -07:00
LC 70eafc4f6f FEXGetConfig: Mark helpers as static where applicable
Makes them internally linked, and also lets them be caught by the
compiler when they're unused.
2026-07-15 12:22:12 -04:00
LC 38cbd2aeb2 FEXGetConfig: Add convenience option for dumping system/tso info
Just lets you get a broad overview all at once instead of needing to
type out every long command.

Now it's easier to be lazy and just pass "-e", or "--all-emu-info".
2026-07-15 12:22:10 -04:00
LC ba762a9326 Merge pull request #5747 from Sonicadvance1/187
HostFeatures: Pull MMFR3 identification register
2026-07-15 12:20:52 -04:00
Ryan Houdek 921ce59054 Merge pull request #5750 from lioncash/pred
VectorOps: Make use of unpredicated shifts
2026-07-15 09:15:04 -07:00
Ryan Houdek a7627ba39a Merge pull request #5751 from lioncash/sq
VectorOps: Add trivial case handling in VSQXTN2
2026-07-15 09:06:59 -07:00
Ryan Houdek 34ef28dea5 Merge pull request #5749 from lioncash/calc
RedundantFlagCalculationElimination: Minor tidying
2026-07-15 09:05:45 -07:00
Ryan Houdek 0f2463dda2 Merge pull request #5748 from lioncash/invariant
RegisterAllocationPass: Ensure pair reg invariant
2026-07-15 09:04:38 -07:00
Ryan Houdek 5ec852637c Merge pull request #5745 from lioncash/addv
VectorOps: Simplify 256-bit VAddV
2026-07-15 09:03:57 -07:00
LC ba8b0afe7a VectorOps: Use unpredicated shifts where applicable for 256-bit scalar shifts
Lets us trim some output
2026-07-15 09:10:57 -04:00
LC 98d45a6a9b VectorOps: Make use of unpredicated immediate shifts
Same behavior, just without introducing a predicate register dependency.
2026-07-15 08:00:58 -04:00
LC 6bc808cb06 VectorOps: Add trivial case handling in VSQXTN2
Lets us generate much more optimal code in the event the destination and
lower source are the same.
2026-07-15 07:49:50 -04:00
LC 1325fef703 RFCE: Prefer accessing ops with C instead of CW
CW is only intended when the op needs to be writable, but most of these
are only reading data.
2026-07-15 07:06:23 -04:00
LC d2d0ef1803 RFCE: Remove unnecessary std::invoke()
We can just call this normally (and also make the constituent helper
function internally linked).
2026-07-15 07:03:09 -04:00
LC d6fb60d512 RegisterAllocationPass: Ensure pair reg invariant
Allows us to actually catch if this requirement ever gets broken in
the future.
2026-07-15 06:46:24 -04:00
LC ef35474f88 VectorOps: Simplify 256-bit VAddV
Didn't read the manual close enough on the first read award.
2026-07-15 04:43:47 -04:00
Ryan Houdek 372891361c HostFeatures: Pull MMFR3 identification register
This has the S1POE flag that we will want to use in the future.
2026-07-14 20:09:31 -07:00
Ryan Houdek 50c75d1f43 Move Linux version calculation to common code 2026-07-14 20:06:37 -07:00
Ryan Houdek 30f2a7b23b Merge pull request #5746 from lioncash/shadow
x87StackOptimizationPass: Remove shadowing variable in PUSHSTACK case
2026-07-14 13:32:13 -07:00
LC f2212a497b x87StackOptimizationPass: Remove shadowing variable in PUSHSTACK case
No behavioral change, since the one in the outer scope does the same thing.
2026-07-14 08:01:30 -04:00
Ryan Houdek 12e8cf008a Merge pull request #5744 from lioncash/telem
AtomicOps: Avoid constrained unpredictable case in TelemetrySetValue()
2026-07-13 16:37:51 -07:00
LC 9e8e87bbb2 AtomicOps: Avoid constrained unpredictable case in TelemetrySetValue()
STLXR cannot use the same register as both the status register and the
value register, otherwise it's architecturally unpredictable
behavior.

Only applies to hardware without FEAT_LSE, so this only meaningfully
affects hardware using the v8.0 spec, since FEAT_LSE becomes mandatory
in v8.1 and newer.
2026-07-13 19:08:17 -04:00
Ryan Houdek 76c4ebb36f Merge pull request #5743 from lioncash/str
StringUtils: Handle strings entirely composed of whitespace in trims
2026-07-13 15:40:50 -07:00
Ryan Houdek 9ff322eeed Merge pull request #5742 from lioncash/sema
x32/Semaphore: Fix storing of message type in msgrcv
2026-07-13 15:40:07 -07:00
LC 4afa49824e StringUtils: Handle strings entirely composed of whitespace in trims
Previously this wouldn't handle fully whitespaced strings.
2026-07-13 17:37:09 -04:00
LC 23402bf31b x32/Semaphore: Fix storing of message type in msgrcv
This was previously storing into the local compat handler, not the
actual managed message.
2026-07-13 17:30:27 -04:00
LC 50be718b72 x32/Semaphore: Mark _ipc as static
This isn't used outside of the translation unit.
2026-07-13 17:30:24 -04:00
Ryan Houdek 903e7db427 Merge pull request #5741 from lioncash/file 2026-07-13 14:13:13 -07:00
LC 5e5e9e0803 Utils/File: Fix handle releasing
ShouldClose was never being set in the event we opened a regular file.
The only time it was set (to false) is when it's used to encapsulate
stderr and stdout.

So anything opened by a File instance was essentially held open.
2026-07-13 16:01:28 -04:00
Ryan Houdek ce27754b9d Merge pull request #5740 from lioncash/ra
RegisterAllocationPass: Function cleanup
2026-07-13 12:42:45 -07:00
LC 4254c0f5a9 RegisterAllocationPass: Function cleanup
Marks a few functions const or static to clarify usage a little more.
2026-07-13 15:26:51 -04:00
Ryan Houdek 7efc3ecaba Merge pull request #5739 from lioncash/zero
Vector: Indicate 128-bit zero vector in DefaultX87State()
2026-07-13 11:59:22 -07:00
LC 2934b01d58 Vector: Indicate 128-bit zero vector in DefaultX87State()
Same functional behavior, just makes it visually match the store size
below. Technically also avoids delegating off to the 64-bit element
path if a 128-bit constant zero is already loaded.
2026-07-13 14:07:12 -04:00
LC 192e363701 Merge pull request #5738 from Sonicadvance1/186
64BitAllocator: Removes unused additional size argument
2026-07-13 13:44:11 -04:00
Ryan Houdek 1287365616 64BitAllocator: Removes unused additional size argument
This used to be used for the intrusively allocated `LiveVMARegion` but
that is all handled internally to the object now, making this
unnecessary. It was always receiving zero and doing nothing so just
remove it.
2026-07-13 10:18:46 -07:00
Ryan Houdek 4fa539fbb2 Merge pull request #5733 from lioncash/alloc
64BitAllocator: Avoid madvising more than necessary in InitializeVMARegionsUsed()
2026-07-13 10:16:57 -07:00
Ryan Houdek 28cdae4687 Merge pull request #5737 from lioncash/bsl
VectorOps: Simplify SVE 256-bit VOrn with BSL2N
2026-07-13 10:05:08 -07:00
LC 842e22915c VectorOps: Simplify SVE 256-bit VOrn with BSL2N
Lets us shave off an instruction and also avoid using a temporary
register in some cases. We can also tweak our worst case that requires a
predicate to eliminate the temporary as well.

We can also expand our cmpps cases, so that we can reflect the
BSL2N usages in instcountci.
2026-07-13 12:27:54 -04:00
Ryan Houdek bd150233ce Merge pull request #5736 from lioncash/ushrni
VectorOps: Make SVE shift==0 case symmetric with ASIMD
2026-07-13 07:56:25 -07:00
LC 24720b67da VectorOps: Make SVE shift==0 case symmetric with ASIMD
Ensures that we have consistent behavior.
2026-07-13 10:26:42 -04:00
Ryan Houdek 356d461123 Merge pull request #5734 from lioncash/bytes
Common/BitSet: Amend byte size retrieval
2026-07-13 07:07:00 -07:00
Ryan Houdek 3cbcc7b9f8 Merge pull request #5735 from lioncash/ir
IR: Enclose straggler Desc comments in brackets
2026-07-13 07:06:06 -07:00
LC 329f12a888 json_ir_generator: Join successive write calls together for allocator helpers
We can just write these out as cohesive units. Also makes adding to them
less annoying.
2026-07-13 09:22:55 -04:00
LC c0b2eec5de IR: Enclose straggler Desc comments in brackets
Ensures the comments get rendered properly in output. We can also
make sure that the IR generation script catches this in the future.
2026-07-13 08:59:45 -04:00
LC 9b8ae25491 Common/BitSet: Amend byte size retrieval
This needs to divide by 8 to get a proper byte size for all type sizes.
The only usage of this is currently a uint64_t, so it worked by
coincidence, since sizeof(uint64_t) == 8.
2026-07-13 08:17:27 -04:00
LC c0ee865e3e 64BitAllocator: Avoid madvising more than necessary in InitializeVMARegionsUsed
Because our bitset type is uint64_t, then that means Memory + ManagedSize
is more like: Memory + (ManagedSize * 8), which is way larger of a base
than we need.
2026-07-12 20:38:19 -04:00
Ryan Houdek 46ec2797ff Merge pull request #5732 from lioncash/vec
Crypto: Clarify zero vector size in SHA1RNDS4Op()
2026-07-12 16:14:26 -07:00
LC fb2cdc8541 Crypto: Clarify zero vector size in SHA1RNDS4Op()
This ends up zeroing out the whole 128-bit vector.
2026-07-12 18:49:27 -04:00
Ryan Houdek 850ef70496 Merge pull request #5731 from lioncash/xar
Crypto: Make use of XAR in SHA1NEXTE when available
2026-07-12 14:26:14 -07:00
LC 9d3c388664 Crypto: Make use of XAR in SHA1NEXTE when available
Lets us shave an instruction off on hardware that supports XAR.

Closes #5730
2026-07-12 16:02:12 -04:00
Ryan Houdek f2b679f602 Merge pull request #5728 from lioncash/halves
x32/FD: Combine offset halves directly
2026-07-12 11:15:10 -07:00
Ryan Houdek b9d97dffe7 Merge pull request #5727 from lioncash/vmsplice
x32/FD: Make use of SanitizeIOCount for vector construction in vmsplice
2026-07-12 11:14:31 -07:00
Ryan Houdek 7ff0466c2c Merge pull request #5726 from lioncash/file
Utils/File: Handle dual read/write case
2026-07-12 11:09:00 -07:00
Ryan Houdek a135325185 Merge pull request #5725 from lioncash/dead
Signals: Preprocessor disable intentional dead code
2026-07-12 11:06:30 -07:00
LC a6c8f0d300 x32/FD: Combine offset halves directly
Shortens these up a little.
2026-07-12 13:31:48 -04:00
LC 046750354e x32/FD: Make use of SanitizeIOCount for vector construction in vmsplice
Makes this consistent with the other fd syscalls that make temporary
buffers.
2026-07-12 13:01:58 -04:00
LC e23d703873 Utils/File: Handle dual read/write case
According to POSIX open docs, this is a completely separate flag that
isn't a combination of O_RDONLY and O_WRONLY, so we need to handle this
separately.

Makes the codepath behaviorally symmetric with the Windows one.
2026-07-12 12:41:02 -04:00
LC ca2d2520d2 Signals: Preprocessor disable intentional dead code in userfaultfd
Noticed this when going through the syscalls. Avoids potential warnings.
2026-07-12 12:16:07 -04:00
Ryan Houdek bc16f902d1 Merge pull request #5724 from lioncash/file
WinAPI/IO: Fix handling of end of file offset in SetFilePointerEx
2026-07-11 22:44:55 -07:00
LC f5ae888597 WinAPI/IO: Fix handling of end of file offset in SetFilePointerEx
This just means the end of the file is being used as the base offset.

Also note that according to the documentation for SetFilePositionEx,
that setting the position beyond the current file size is not considered
an error as far as the API is concerned.
2026-07-12 01:26:23 -04:00
Ryan Houdek 376e3af058 Merge pull request #5723 from lioncash/gdb 2026-07-11 21:54:57 -07:00
Ryan Houdek af63c0a9e1 Merge pull request #5722 from lioncash/container 2026-07-11 21:54:20 -07:00
LC 7d149ebec4 GdbServer: Add missing log format argument 2026-07-12 00:28:29 -04:00
LC 37c809471b ElfContainer: Amend entry iteration in GetDynamicLibs()
These were using i in the termination condition, which is for section
headers, not entries.
2026-07-12 00:19:28 -04:00
Ryan Houdek 12baceb859 Merge pull request #5721 from lioncash/win
AllocatorHooks: Amend VirtualProtect for Windows
2026-07-11 20:49:16 -07:00
Ryan Houdek 3e6b3c8c40 Merge pull request #5720 from lioncash/hdr
64BitAllocator: Remove duplicate headers
2026-07-11 20:33:09 -07:00
LC 25f2021711 AllocatorHooks: Amend VirtualProtect for Windows
VirtualProtect returns non-zero on success, also the old protection flag
parameter isn't allowed to be null.
2026-07-11 23:24:30 -04:00
LC 8b30f7dbb5 64BitAllocator: Remove duplicate headers
These are already included.
2026-07-11 23:08:29 -04:00
Ryan Houdek 76b35dfeb0 Merge pull request #5719 from lioncash/small
64BitAllocator: Avoid overwriting Region[0] in Create64BitAllocatorWithRegions
2026-07-11 19:30:08 -07:00
Ryan Houdek 4518d5831b Merge pull request #5718 from lioncash/absolute
Filesystem: Fix Absolute() on Windows
2026-07-11 17:11:18 -07:00
Ryan Houdek 066250851f Merge pull request #5717 from lioncash/reg
RegisterAllocationPass: Amend type cast in DecodeSRANode()
2026-07-11 17:10:27 -07:00
Ryan Houdek 341a5196ba Merge pull request #5701 from lioncash/thread
Thread: Amend new thread handling in HandleNewClone()
2026-07-11 17:09:59 -07:00
LC c878a89e95 Filesystem: Fix Absolute() on Windows
sizeof(*Fill) will only ever be 1, so we wouldn't actually copy much of
anything.
2026-07-11 17:25:26 -04:00
LC c72d59a3f1 64BitAllocator: Avoid overwriting Region[0] in Create64BitAllocatorWithRegions
Since this was a reference, this would end up overwriting Region[0] with
whatever the smallest region was instead of just being a running pointer
to what happened to be the current smallest region.

We can switch over to a pointer to avoid obliterating the first memory
region.
2026-07-11 16:52:01 -04:00
LC 995e2657bb RegisterAllocationPass: Amend type cast in DecodeSRANode()
Uses the proper type for StoreRegister. Same behavior though, due to
layout.
2026-07-11 16:39:35 -04:00
Ryan Houdek fe6d6397d6 Merge pull request #5716 from lioncash/vec 2026-07-11 13:29:47 -07:00
LC c7f52bcee4 Vector: Centralize masking in InsertScalarFCMPOp
Ensures that even if someone threw bogus constants in the upper bits of
the immediate, that the special-cased comparison types would still be
handled properly.

We can move the masking in the AVX variant too, just to be consistent.
2026-07-11 16:07:19 -04:00
Ryan Houdek fb4d8a6d14 Merge pull request #5715 from lioncash/singlestep
Dispatcher: Avoid loading unnecessary reg in vixl single step
2026-07-11 12:25:13 -07:00
LC 929f9a7ad2 Dispatcher: Avoid loading unnecessary reg in vixl single step
CompileSingleStep only takes one uint64_t, not two.
2026-07-11 15:03:29 -04:00
Ryan Houdek 48ce5bf6e6 Merge pull request #5714 from lioncash/readahead
x32/FD: Fix readahead upper offset type
2026-07-11 11:20:49 -07:00
Ryan Houdek 617a518714 Merge pull request #5713 from lioncash/select
x32/FD: Correct total word calculation in select() variants
2026-07-11 11:19:29 -07:00
Ryan Houdek 480f45f2f3 Merge pull request #5712 from lioncash/bpf
BPFEmitter: Amend instruction class checking in HandleStore()
2026-07-11 11:08:58 -07:00
Ryan Houdek 0e8b01c1b5 Merge pull request #5711 from lioncash/fault
x64/Thread: Amend faulting copy handling related to LDTs
2026-07-11 11:04:47 -07:00
Ryan Houdek b750d6772f Merge pull request #5710 from lioncash/close
Common/Async: Handle fd closing a little better
2026-07-11 11:03:10 -07:00
Ryan Houdek 7ce172a497 Merge pull request #5709 from lioncash/size
FlexBitSet: Simplify MemClear/MemSet
2026-07-11 10:57:58 -07:00
Ryan Houdek 821dfe5b98 Merge pull request #5708 from lioncash/mrs
MiscOps: Fix round mode clearing for RP/RM modes in PushRoundingMode
2026-07-11 10:56:27 -07:00
Ryan Houdek 6c57b2f8f9 Merge pull request #5707 from lioncash/offset
MemoryOps: Avoid double application of base offset in {Load,Store}ContextIndexed case
2026-07-11 10:49:27 -07:00
Ryan Houdek 6e0c9d159d Merge pull request #5706 from lioncash/odd
SignalDelegator: Remove odd double negation in GuestSigProcMask
2026-07-11 10:48:42 -07:00
LC a90a7dbec7 x32/FD: Fix readahead upper offset type
This should be a uint32_t
2026-07-11 13:33:11 -04:00
LC 2250bd58a2 x32/FD: Deduplicate guest and host fd set management
Same behavior, but less copy pastey
2026-07-11 13:23:19 -04:00
LC 126be45ed1 x32/FD: Correct total word calculation in select() variants
Previously this would result in a larger amount of words specified than
necessary.

e.g. Given nfds = 1:

With AlignUp(1, 32) / 4, we'd end up with supposedly eight words, when it
should only be one word.

On the other extreme, given a full fd set of 1024 fds, then we'd end up
with 256 words, when it should only be 32.
2026-07-11 12:53:08 -04:00
LC e021a55abd BPFEmitter: Amend instruction class checking in HandleStore()
This was previously checking for a load class, which would result in ST
clobbering the index register.
2026-07-11 12:26:11 -04:00
LC 2ad6254894 x64/Thread: Amend faulting copy handling related to LDTs
CopyToUser doesn't return the number of bytes copied, but rather returns
0 to indicate success, otherwise a fault has occurred (and the SIGSEGV
handler has set X0 to EFAULT)
2026-07-11 11:57:38 -04:00
LC 5bd97c2f63 Common/Async: Handle fd closing a little better
We should be checking against -1, rather than just anything non-zero.
2026-07-11 10:49:08 -04:00
LC 570d1f2271 FlexBitSet: Simplify MemClear/MemSet
We can just make use of the helpers already in the interface.
2026-07-11 10:33:47 -04:00
LC 613e9ef701 MiscOps: Fix round mode clearing for RP/RM modes in PushRoundingMode
Previously this had the potential to not clear rounding bits properly
depending on incoming FPCR state.
2026-07-11 10:24:07 -04:00
LC 5080c6ffc5 MemoryOps: Avoid double application of base offset in {Load,Store}ContextIndexed unaligned case 2026-07-11 09:56:56 -04:00
LC f16bc12b7d SignalDelegator: Remove odd double negation in GuestSigProcMask
We can just reduce it to normal null comparisons.
2026-07-11 05:42:48 -04:00
Ryan Houdek d2f096187c Merge pull request #5705 from lioncash/thread3 2026-07-10 22:34:04 -07:00
Ryan Houdek bd1e61befd Merge pull request #5704 from lioncash/host 2026-07-10 22:33:30 -07:00
Ryan Houdek b97165bec9 Merge pull request #5703 from lioncash/size 2026-07-10 22:32:45 -07:00
LC 8b0f07b2dc Syscalls: Remove unnecessary usages of namespace FEXCore::IR
These aren't necessary.
2026-07-10 21:02:12 -04:00
LC c0ba45f6de HostFeatures: Shrink feature setting in FillFeatureFlags
Allows us to unify most of the flag setting, so the flag name only needs
to be stated once, reducing likelihood of typos.
2026-07-10 20:34:46 -04:00
LC 017c898ed0 x64/Signals: Amend set size in rt_sigtimedwait
We should be checking the size passed in, not the sizeof of it.
2026-07-10 19:40:05 -04:00
Ryan Houdek b193a0c9fd Merge pull request #5702 from lioncash/sbss
HostFeatures: Amend SSBS2 signifying
2026-07-10 16:38:07 -07:00
LC db0c7b0562 HostFeatures: Amend SSBS2 signifying 2026-07-10 19:13:29 -04:00
LC 85b8e91b57 Thread: Amend new thread handling in HandleNewClone()
Ensures that newly cloned threads get tracked properly.
2026-07-10 18:35:58 -04:00
Ryan Houdek 87301ca154 Merge pull request #5700 from lioncash/bitset
Common/Bitset: Minor API changes
2026-07-10 10:35:45 -07:00
Ryan Houdek 650f5b784d Merge pull request #5699 from lioncash/spillops
Arm64Emitter: Wire up conditional FPR spilling in SpillForPreserveAllABICall
2026-07-10 10:35:19 -07:00
Ryan Houdek c8dd9eefa6 Merge pull request #5698 from lioncash/dead
ConversionOps: Remove redundant code in Vector_FToS
2026-07-10 10:35:06 -07:00
Ryan Houdek 0bdf15977f Merge pull request #5697 from lioncash/thread
x32/Thread: Move writability check around in waitpid
2026-07-10 10:34:56 -07:00
Ryan Houdek 8a844f9cca Merge pull request #5695 from lioncash/socket
Socket: Pass size by reference in getsockopt
2026-07-10 10:34:33 -07:00
Ryan Houdek 8fbde84380 Merge pull request #5696 from lioncash/time
x32/Time: Correct sizeof expression in utimensat
2026-07-10 10:31:47 -07:00
Ryan Houdek 00826d3327 Merge pull request #5694 from lioncash/rlimit
x32/Info: Only modify output in getrlimit/ugetrlimit if successful
2026-07-10 10:27:51 -07:00
Ryan Houdek d5572e322f Merge pull request #5693 from lioncash/ir
IR: Use begin block type in operator--
2026-07-10 10:27:13 -07:00
Ryan Houdek 37814111de Merge pull request #5692 from lioncash/unary
Vector: Use unary handler for scalar unary insertions
2026-07-10 10:26:44 -07:00
Ryan Houdek 1417888a89 Merge pull request #5691 from lioncash/fpr
Vector: LoadSourceGPR -> LoadSourceFPR for MASKMOVOp
2026-07-10 10:25:26 -07:00
Ryan Houdek 5d846c3ca9 Merge pull request #5690 from lioncash/cache
MemoryOps: Make use of current working reg for cache operations
2026-07-10 10:23:55 -07:00
Ryan Houdek 5dd0477440 Merge pull request #5689 from lioncash/msg
x32/Msg: Fix result comparison in mq_getsetattr
2026-07-10 10:15:15 -07:00
Ryan Houdek c6f823f855 Merge pull request #5688 from lioncash/pidfd
Thread: Amend pidfd_open check
2026-07-10 10:14:49 -07:00
Ryan Houdek 2af7c24e79 Merge pull request #5687 from lioncash/limit
x32/Thread: Fix off-by-one in get_thread_area
2026-07-10 10:14:29 -07:00
Ryan Houdek 94ccfefa84 Merge pull request #5686 from lioncash/fd
x32/FD: Ensure sendfile updates offset if set
2026-07-10 10:14:18 -07:00
LC 46f3bec37e Common/BitSet: Ensure internal pointer is always initialized
Provides deterministic state.
2026-07-10 10:31:45 -04:00
LC f5d2e0db29 Common/BitSet: Mark getters as const
These don't modify internal state.
2026-07-10 10:31:05 -04:00
LC 88ee56f471 Common/BitSet: Amend Clear() behavior
Ensures the bits are actually being unset.
2026-07-10 10:26:07 -04:00
LC 8436154276 Arm64Emitter: Wire up conditional FPR spilling in SpillForPreserveAllABICall
Technically, this parameter wasn't being used at all. It was wired up
for filling, but not spilling.
2026-07-10 10:10:25 -04:00
LC 150b25f29e ConversionOps: Remove redundant code in Vector_FToS
These are already defined in an outer scope.
2026-07-10 09:16:35 -04:00
LC f4c50105ba x32/Thread: Move writability check around in waitpid
Same behavior, but catches the write before it actually occurs.
2026-07-10 08:08:45 -04:00
LC 3395bedc2b x32/Time: Correct sizeof expression in utimensat
Ensures we check the proper type.
2026-07-10 08:04:56 -04:00
LC 7a14210e2b Socket: Pass size by reference in getsockopt
Previously this was passing by value.
2026-07-10 08:01:29 -04:00
LC d206b67ca8 x32/Info: Only modify output in getrlimit/ugetrlimit if successful
Avoids trampling over input data.
2026-07-10 07:49:44 -04:00
LC 0bfef9008b IR: Use begin block type in operator--
Same behavior, just more correct from a descriptive PoV
2026-07-10 07:31:25 -04:00
LC 401e542fcd Vector: Use unary handler for scalar unary insertions
Same behavior, but just uses a more proper handler.
2026-07-10 07:27:54 -04:00
LC 210514f74b Vector: LoadSourceGPR -> LoadSourceFPR for MASKMOVOp 2026-07-10 07:24:39 -04:00
LC dd838d4ad3 MemoryOps: Make use of current working reg for cache operations
TMP1 technically isn't initialized properly here until after the first
iteration.
2026-07-10 07:13:17 -04:00
Ryan Houdek 5f2d19c7aa Merge pull request #5685 from lioncash/faddv 2026-07-10 03:33:43 -07:00
LC d5ae87b5ca Thread: Amend pidfd_open check
Checks for success.
2026-07-10 06:25:43 -04:00
Ryan Houdek 95e7c866cf Merge pull request #5682 from lioncash/pid 2026-07-10 03:20:23 -07:00
LC 02028eb1ad x32/Msg: Fix result comparison in mq_getsetattr
Checks against failure instead of 1.
2026-07-10 06:19:11 -04:00
Ryan Houdek 2effdb04bd Merge pull request #5684 from lioncash/midr 2026-07-10 03:18:25 -07:00
Ryan Houdek ab31e3e3bb Merge pull request #5683 from lioncash/mul 2026-07-10 03:18:09 -07:00
LC 37b1432514 x32/Thread: Fix off-by-one in get_thread_area
12, 13, and 14 are the only valid TLS areas.
2026-07-10 06:06:30 -04:00
LC 6e8bc337aa x32/FD: Ensure sendfile updates offset if set 2026-07-10 05:58:49 -04:00
Ryan Houdek 8347566815 Merge pull request #5681 from lioncash/timer 2026-07-10 02:51:20 -07:00
LC ede09a03db VectorOps: Fix 256-bit FADDV path
Avoids falling down to the SVE-128 path.
2026-07-10 05:50:22 -04:00
LC 849d60253c CPUID: Fix MIDR walking in SetupHostHybridFlag() 2026-07-10 05:46:16 -04:00
LC b5660c8a92 MiscOps: Avoid stack misalignment in ProcessorID
This needs to be an add.
2026-07-10 05:41:54 -04:00
LC 54263bb5a7 ALUOps: Fix 32-bit MulH case
These need to be 64-bit multiply and ubfx. Thankfully this case wasn't
actually hit in practice.
2026-07-10 05:39:14 -04:00
LC 4069a9f7d5 Timer: Fix typo in timer_gettime
This should be passed by reference rather than by value.
2026-07-10 05:33:06 -04:00
Ryan Houdek 77467eaf0c Merge pull request #5680 from lioncash/usrai
IR: Remove unused VUShraI
2026-07-10 01:50:51 -07:00
LC caa030714b IR: Remove unused VUShraI
Given that this is currently unused and that we don't have the signed
equivalent implemented, we can just remove this for now.
2026-07-10 04:15:52 -04:00
Ryan Houdek facbc78e0a Merge pull request #5679 from lioncash/macro
ALUOps: Remove unused macros
2026-07-10 00:50:52 -07:00
LC 6488dcbb01 ALUOps: Remove unused macros
These are now unused.
2026-07-10 03:30:03 -04:00
LC 37265b109a Merge pull request #5676 from Sonicadvance1/185
FEX: Remove FEXInterpreter binary
2026-07-09 18:39:36 -04:00
Ryan Houdek ebe7342d10 Merge pull request #5677 from mrpippy/protontso
Windows/UnixLib: Fix enabling TSO through legacy Proton codepath
2026-07-09 15:33:27 -07:00
Ryan Houdek 523bbef034 FEX: Remove FEXInterpreter binary
It's been ten months, a bit longer than than I was expecting to keep
this around. Go ahead and remove it now.
2026-07-09 15:11:30 -07:00
Brendan Shanks 84e127a637 Windows/UnixLib: Fix enabling TSO through legacy Proton codepath 2026-07-09 15:00:50 -07:00
Ryan Houdek 4a091df8cc Merge pull request #5672 from neobrain/feature_woa_code_cache_bitness
CodeCache/WoA: Support mixed WoW64/ARM64EC processing
2026-07-09 14:55:21 -07:00
Ryan Houdek 92a171ce53 Merge pull request #5675 from lioncash/branch
VectorOps: Join identical branches in VFMLS/VFNMLS
2026-07-09 13:58:54 -07:00
LC 1b1e46ff6c VectorOps: Join identical branches in VFMLS/VFNMLS
Same thing, just a little less redundant.
2026-07-09 16:33:26 -04:00
Ryan Houdek 3370d9af15 Merge pull request #5670 from simon902/MOVDoverride
Fix movd when prefixed with 0x66
2026-07-09 13:02:08 -07:00
Ryan Houdek 9306de79ad Merge pull request #5667 from simon902/CVTTSS2SIOverride
Fix cvttss2si when prefixed with 0x66
2026-07-09 12:52:00 -07:00
Ryan Houdek ff7a54add8 Merge pull request #5668 from OFFTKP/inf
Fix element getting overwritten in 66_5B test
2026-07-09 12:26:46 -07:00
Ryan Houdek c3d1157696 Merge pull request #5669 from OFFTKP/lzcnt
Fix LZCNT tests reading out of bounds
2026-07-09 12:24:43 -07:00
Ryan Houdek d0f03cb148 Merge pull request #5674 from lioncash/insertq
Vector: Trim one instruction off insertq
2026-07-09 12:24:06 -07:00
LC 9b7c9f0fb6 Vector: Trim one instruction off insertq
We can fold a bitwise not and and pair into a bic
2026-07-09 14:58:04 -04:00
Ryan Houdek e508b6df0d Merge pull request #5673 from lioncash/vbitwise
IR: Remove need to specify element size for vector bitwise ops
2026-07-09 11:42:37 -07:00
LC 710b85be70 IR: Remove need to specify element size for vector bitwise ops
Element size doesn't really mean anything here, considering all bits are
acted upon independently of segmentation.

Makes using these ops a little bit less noisy.
2026-07-09 13:26:51 -04:00
Tony Wasserka 66455b708a CodeCache: Switch between 32-/64-bit compilers during cache generation 2026-07-09 17:17:19 +02:00
Tony Wasserka 8b8000b98a CodeCache/WoA: Run cache generation in a subprocess to improve robustness 2026-07-09 17:16:16 +02:00
Tony Wasserka 54236df6e0 CodeCache: Record main executable bitness in code map
Code maps already contain the main executable they were recorded from, so
it's convenient to capture the executable's bitness along the way.
2026-07-09 17:12:45 +02:00
LC 6cd2a48910 Merge pull request #5666 from Sonicadvance1/184
FEXCore: Fixes a crash with multiblock if `ProcessorID` IR op is encountered
2026-07-09 10:16:06 -04:00
Simon Scherer fe08b96844 OpcodeDispatcher: Fix cvttss2si when prefixed with 0x66 2026-07-09 11:58:18 +02:00
Simon Scherer 4d78901420 OpcodeDispatcher: Fix movd when prefixed with 0x66 2026-07-09 11:46:48 +02:00
Simon Scherer c06468025c unittests/ASM: Test movd prefixed with 0x66 2026-07-09 11:46:10 +02:00
Paris Oplopoios 148e539025 Fix LZCNT tests reading out of bounds 2026-07-09 12:36:18 +03:00
Paris Oplopoios 92b96ff30d Fix element getting overwritten in 66_5B test 2026-07-09 11:54:38 +03:00
Simon Scherer 778df0c93b unittests/ASM: Test cvttss2si prefixed with 0x66 2026-07-09 09:24:59 +02:00
Ryan Houdek 9d18ecc5cb FEXCore: Fixes a crash with multiblock if ProcessorID IR op is encountered
If during multiblock code discovery a RDTSCP/RDPID instruction was
encountered then ProcessorID has an assert at JIT compile time. Make
sure to early exit with an illegal instruction encoding early instead.
Also make sure to correctly report RDPID support in CPUID, it's
technically a different bit than RDTSCP.

Fixes a crash in Crusader Kings 3's Paradox Launcher installer. Although
the installer seems to fail otherwise for some reason.
2026-07-08 17:41:03 -07:00
Ryan Houdek 5f2455c502 Merge pull request #5665 from lioncash/blendop
[SVE256] Handle 256-bit blend operations much more efficiently
2026-07-08 16:43:20 -07:00
LC 6bc67609a3 [SVE256] Handle 256-bit blend operations much more efficiently
We can massage a given selector into a valid predicate register bitmask
and then simply perform a merging move, which eliminates most busywork
around optimizing 256-bit blends.

In the future, once we drop SVE2.1 support in, we can use PMOV to
eliminate the load from memory and related constant management.
2026-07-08 17:35:12 -04:00
LC 1bd3945dd9 Merge pull request #5577 from Sonicadvance1/168
Context: Add support for single-step RIP ranges
2026-07-08 15:47:28 -04:00
Ryan Houdek 168f4b1e6b Context: Add support for single-step RIP ranges
Useful when debugging a range.
2026-07-08 12:21:16 -07:00
LC 8a8827c980 Merge pull request #5664 from simon902/CMPXCHGZeroing
OpcodeDispatcher: Fix 32bit cmpxchg zero extension with eax as first operand
2026-07-08 15:14:12 -04:00
Ryan Houdek 21a968b84d Merge pull request #5663 from simon902/PDEPoverlap
JIT/ALUOps: Fix operand overlapping bug for pdep
2026-07-08 11:08:12 -07:00
Simon Scherer 84fab84b3f InstcountCI: Update 2026-07-08 15:00:18 +02:00
Simon Scherer 3d65c030a8 OpcodeDispatcher: Fix 32bit cmpxchg zero extension with eax as destination operand and remove incorrect comment. 2026-07-08 14:58:32 +02:00
Simon Scherer 59097bab20 unittests/ASM: Test cmpxchg with eax as destination 2026-07-08 14:52:19 +02:00
Simon Scherer 4cbacd9261 InstcountCI: Update 2026-07-08 10:10:30 +02:00
Simon Scherer 655102fc7d JIT/ALUOps: Fix operand overlapping bug for pdep 2026-07-08 10:09:47 +02:00
Simon Scherer f718f46545 unittests/ASM: Test overlapping operands for pdep 2026-07-08 09:47:12 +02:00
LC 71d4e2c320 Merge pull request #5660 from Sonicadvance1/183
Wow64: Spin loop on atomic with WFE
2026-07-07 22:53:50 -04:00
Ryan Houdek 41241d7500 Wow64: Spin loop on atomic with WFE
Instead of burning roughly a million watts, put this spinloop on a WFE.
This tends to occur on a crash during shutdown that isn't fully able to
be avoided. The least we can do is not consume all the power in the
world.
2026-07-07 16:09:00 -07:00
Ryan Houdek b90c9836cb Merge pull request #5662 from lioncash/alias
OpcodeDispatcher: Remove asterisk from BMI source args
2026-07-07 11:22:59 -07:00
Ryan Houdek aff3fcf76d Merge pull request #5658 from neobrain/fix_woa_code_cache_ec
CodeCache: Mark executable memory as EC code on ARM64EC
2026-07-07 11:22:15 -07:00
Ryan Houdek ec2aa4063a Merge pull request #5661 from lioncash/blend
unittests: Add stress tests for VBLEND{PD, PS}
2026-07-07 11:13:56 -07:00
LC 718f2e01f7 OpcodeDispatcher: Remove asterisk from BMI source args
Keeps it consistent with the rest of the code and prevents breakages
whenever the Ref alias gets turned into its own value type.
2026-07-07 14:07:38 -04:00
LC ba9f7fb1b5 unittests: Add stress tests for VBLEND{PD, PS}
Forgot about these two
2026-07-07 13:49:09 -04:00
Tony Wasserka 0ac6b3e8f3 CodeCache: Mark executable memory as EC code on ARM64EC
See bd5b817c3a.
2026-07-07 15:55:27 +02:00
Ryan Houdek dddad1c2ca Merge pull request #5659 from lioncash/shuf
unittests: Add stress tests for VSHUF{PD, PS}
2026-07-06 15:50:08 -07:00
LC 95bfff20a4 unittests: Add stress tests for VSHUF{PD, PS}
Covers the remaining shuffle paths
2026-07-06 18:01:52 -04:00
LC 7a6f0def85 Merge pull request #5653 from Sonicadvance1/182
Config: Fixes AppOverrides with FEX_APP_CONFIG
2026-07-06 17:01:55 -04:00
Ryan Houdek 103d4d76be Config: Fixes AppOverrides with FEX_APP_CONFIG 2026-07-06 12:43:14 -07:00
Ryan Houdek db9414a756 Merge pull request #5655 from neobrain/feature_woa_cache_loading
Windows/ImageTracker: Adapt code cache loading logic to FEXOfflineCompiler
2026-07-06 12:40:40 -07:00
Ryan Houdek 5a0bf1bb5f Merge pull request #5657 from neobrain/fix_foc_syscall_abi_woa
FEXOfflineCompiler: Fix improper syscall ABI on WoA
2026-07-06 12:37:52 -07:00
LC 444c37fe2e Merge pull request #5656 from neobrain/fix_invalid_iterator_deref
Core: Fix dereference of invalid iterator
2026-07-06 10:13:04 -04:00
Tony Wasserka d3d735370f FEXOfflineCompiler: Fix improper syscall ABI on WoA 2026-07-06 15:16:30 +02:00
Tony Wasserka 852e93aa74 Core: Fix dereference of invalid iterator 2026-07-06 15:07:37 +02:00
Tony Wasserka dc1be2efe8 Windows/ImageTracker: Unindent refactored code 2026-07-06 14:58:16 +02:00
Tony Wasserka f2734ac608 Windows/ImageTracker: Adapt code cache loading logic to FEXOfflineCompiler 2026-07-06 14:58:16 +02:00
Tony Wasserka 57ca49dc5c Merge pull request #5501 from bylaws/finishloadwin
Windows/ImageTracker: Wire up LoadCache and EnableLoadedSection
2026-07-06 14:58:01 +02:00
Billy Laws 5ef3134a4d Windows/ImageTracker: Wire up LoadCache and EnableLoadedSection
The lazy code loading refactor replaced LoadData with the new
LoadCache/EnableLoadedSection API but left the Windows path as TODOs.
Implement the wiring: LoadAOTImages now calls LoadCache +
RegisterMappedCodeBuffer for each mapped cache file, and HandleImageMap
calls EnableLoadedSection (with nullptr thread since lazy mapping is not
yet implemented on Windows).
2026-07-06 14:44:59 +02:00
Ryan Houdek 5b91642883 Merge pull request #5654 from ShadowCurse/cache_va_size
Allocator: fix the caching of host va size
2026-07-05 18:32:10 -07:00
Egor Lazarchuk 1eb5abb8db Allocator: rename DetermineVASize to GetHostVABits
`DetermineVASize` does not return the size of VA, but the number of bits
it can use. Change the naming to make it more self explanatory.
In the mean time also move `HostVASize` global into `GetHostVABits`
since it is not and should not be used directly.
2026-07-05 12:46:41 +01:00
Egor Lazarchuk a65e1bf7e5 Allocator: fix the caching of host va size
Commit abf9724 ("Allocator: Fix and optimize VA range detection")
removed assignment to the `HostVASize` global thus making each call to
`DetermineVASize` redo all the work with potential to produce incorrect
results. Set the global again to fix this.
2026-07-05 12:46:29 +01:00
Ryan Houdek 3f98202b0f Merge pull request #5651 from lioncash/perm
unittests: Add stress tests for VPERMIL{PD, PS}
2026-07-04 12:13:24 -07:00
LC e7727c39f3 unittests: Add stress tests for VPERMIL{PD, PS}
While unlikely to be used in practice over other kind of
shuffling and blending, these should also have stress tests
to make sure they do the right thing.
2026-07-04 15:00:30 -04:00
Ryan Houdek 17e5637664 Merge pull request #5650 from lioncash/aes256
[SVE256] Handle 256-bit AES operations
2026-07-03 19:06:36 -07:00
LC 7fd9b897c2 [SVE256] Handle 256-bit AES operations
Currently we split these into two 128-bit operations since VIXL doesn't
have support for the unified SVE operations yet.

Now we fully support VAES on SVE256.
2026-07-03 21:52:59 -04:00
Ryan Houdek f8967aa207 Merge pull request #5649 from lioncash/pclmul256
[SVE256] EncryptionOps: Handle 256-bit VPCLMULQDQ
2026-07-03 18:33:53 -07:00
LC 5145324806 [SVE256] EncryptionOps: Handle 256-bit VPCLMULQDQ
Since vixl now handles this, we can drop this support right in.
2026-07-03 21:22:05 -04:00
Ryan Houdek 4233fb6270 Merge pull request #5648 from lioncash/aes
[SVE256] Ensure SSE insertion behavior for AES/SHA/PCLMUL operations
2026-07-03 17:33:19 -07:00
LC d11b19fd2b [SVE256] Ensure insertion behavior for PCLMUL SSE operations
Also includes accompanying test to ensure it never breaks.
2026-07-03 20:15:34 -04:00
LC 684c568033 [SVE256] Ensure insertion behavior for SHA SSE operations
These slipped through, so now we can add tests for them to prevent that
from happening again.
2026-07-03 20:09:50 -04:00
LC ab4fb7b3ad [SVE256] Ensure insertion behavior for AES operations on SSE
These slipped through, so now we can add tests for them to prevent that
from happening again.
2026-07-03 19:35:31 -04:00
Ryan Houdek 91017dbedb Merge pull request #5647 from lioncash/perm128
unittests: Add stress test for vperm2f128/vperm2i128
2026-07-03 15:32:34 -07:00
LC f5e8a051a8 unittests: Add stress test for vperm2f128/vperm2i128
Lets us test all possible immediate encodings for proper behavior.
2026-07-03 16:17:12 -04:00
LC 1d695f6db4 Merge pull request #5646 from Sonicadvance1/181
Github: More dependabot things
2026-07-02 22:22:45 -04:00
Ryan Houdek d4bdfd0592 Github: More dependabot things
They just never stop.
2026-07-02 19:05:12 -07:00
387 changed files with 19656 additions and 4881 deletions

No files matched your search

+1 -1
View File
@@ -43,7 +43,7 @@ jobs:
distrobox upgrade steamrt4
distrobox enter --name steamrt4 -- sudo apt-get install -y \
git cmake ninja-build ccache \
lld clang \
lld clang clang-tools \
libclang-dev llvm-dev \
libstdc++-14-dev-i386-cross libgcc-14-dev-i386-cross \
libstdc++-14-dev-amd64-cross libgcc-14-dev-amd64-cross
+1 -1
View File
@@ -24,7 +24,7 @@ runs:
cmake -S . -B build_${{ inputs.target }} -DCMAKE_BUILD_TYPE=$BUILD_TYPE -DCMAKE_TOOLCHAIN_FILE=Data/CMake/toolchain_mingw.cmake \
-DMINGW_TRIPLE=${_cc}-w64-mingw32 -DCMAKE_INSTALL_LIBDIR=/usr/lib/wine/aarch64-windows -G Ninja \
-DENABLE_LTO=False -DENABLE_ASSERTIONS=False -DENABLE_JEMALLOC_GLIBC_ALLOC=False \
-DBUILD_TESTING=False -DCMAKE_INSTALL_PREFIX=/usr -DTUNE_ARCH=generic -DTUNE_CPU=none
-DBUILD_TESTING=False -DCMAKE_INSTALL_PREFIX=/usr -DTUNE_ARCH=generic -DTUNE_CPU=none -DRANGES_NATIVE=OFF
- name: Build
shell: bash
+2 -2
View File
@@ -31,12 +31,12 @@ build:
- apt-get -y update
- apt-get install -y
git cmake ninja-build ccache
lld clang
lld clang clang-tools
libclang-dev llvm-dev
libstdc++-14-dev-i386-cross libgcc-14-dev-i386-cross
libstdc++-14-dev-amd64-cross libgcc-14-dev-amd64-cross
- cmake -E make_directory build/
- cmake -DCMAKE_BUILD_TYPE=Release -G Ninja -DBUILD_STEAM_SUPPORT=True -DENABLE_LTO=True -DENABLE_ASSERTIONS=False -DBUILD_THUNKS=True -DBUILD_FEXCONFIG=False -DBUILD_TESTING=False -DENABLE_CLANG_THUNKS=True -DUSE_LINKER=lld -DCMAKE_INSTALL_PREFIX=/usr -DTUNE_ARCH=armv8.2-a -DTUNE_CPU=none . -B build/
- cmake -DCMAKE_BUILD_TYPE=Release -G Ninja -DBUILD_STEAM_SUPPORT=True -DENABLE_LTO=True -DENABLE_ASSERTIONS=False -DBUILD_THUNKS=True -DBUILD_FEXCONFIG=False -DBUILD_TESTING=False -DENABLE_CLANG_THUNKS=True -DUSE_LINKER=lld -DCMAKE_INSTALL_PREFIX=/usr -DTUNE_ARCH=armv8.2-a -DTUNE_CPU=none -DRANGES_NATIVE=OFF . -B build/
- cmake --build build/ --config Release
- DESTDIR=$(pwd)/install/ cmake --build build/ --config Release -t install
+24 -8
View File
@@ -195,6 +195,9 @@ if (ENABLE_GDB_SYMBOLS)
endif()
add_compile_definitions(_LARGEFILE64_SOURCE)
if (WIN32)
add_compile_definitions(UNICODE _UNICODE)
endif()
set(CMAKE_CXX_STANDARD 20)
set(CMAKE_EXPORT_COMPILE_COMMANDS ON)
@@ -425,11 +428,16 @@ else ()
file(GENERATE OUTPUT CTestTestfile.cmake CONTENT "# No tests since BUILD_TESTING is disabled")
endif()
find_package(fmt QUIET)
if (NOT fmt_FOUND)
# Disable fmt install
if (MINGW)
set(FMT_INSTALL OFF)
add_subdirectory(External/fmt/)
else()
find_package(fmt QUIET)
if (NOT fmt_FOUND)
# Disable fmt install
set(FMT_INSTALL OFF)
add_subdirectory(External/fmt/)
endif()
endif()
find_package(range-v3 QUIET)
@@ -477,12 +485,20 @@ endif()
set(FEX_TUNE_COMPILE_FLAGS)
if (NOT TUNE_ARCH STREQUAL "generic")
check_cxx_compiler_flag("-march=${TUNE_ARCH}" COMPILER_SUPPORTS_ARCH_TYPE)
if(COMPILER_SUPPORTS_ARCH_TYPE)
list(APPEND FEX_TUNE_COMPILE_FLAGS "-march=${TUNE_ARCH}")
else()
message(FATAL_ERROR "Trying to compile arch type '${TUNE_ARCH}' but the compiler doesn't support this")
set(TUNE_ARCH_STRING "${TUNE_ARCH}")
if(ARCHITECTURE_arm64)
set(TUNE_ARCH_STRING "${TUNE_ARCH}+crc")
endif()
check_cxx_compiler_flag("-march=${TUNE_ARCH_STRING}" COMPILER_SUPPORTS_ARCH_TYPE)
if(COMPILER_SUPPORTS_ARCH_TYPE)
list(APPEND FEX_TUNE_COMPILE_FLAGS "-march=${TUNE_ARCH_STRING}")
else()
message(FATAL_ERROR "Trying to compile arch type '${TUNE_ARCH_STRING}' but the compiler doesn't support this")
endif()
elseif(ARCHITECTURE_arm64)
# Need to always append crc
check_cxx_compiler_flag("-march=armv8-a+crc" COMPILER_SUPPORTS_ARCH_TYPE)
list(APPEND FEX_TUNE_COMPILE_FLAGS "-march=armv8-a+crc")
endif()
if (TUNE_CPU STREQUAL "native")
+18 -11
View File
@@ -270,6 +270,9 @@ public:
void fcvtxnt(ZRegister zd, PRegisterMerge pg, ZRegister zn) {
SVEFloatConvertOdd(0b00, 0b10, pg, zn, zd);
}
void bfcvtnt(ZRegister zd, PRegisterMerge pg, ZRegister zn) {
SVEFloatConvertOdd(0b10, 0b10, pg, zn, zd);
}
///< Size is destination size
void fcvtnt(SubRegSize size, ZRegister zd, PRegisterMerge pg, ZRegister zn) {
LOGMAN_THROW_A_FMT(size == SubRegSize::i32Bit || size == SubRegSize::i16Bit, "Unsupported size in {}", __func__);
@@ -292,8 +295,6 @@ public:
SVEFloatConvertOdd(ConvertedSrcSize, ConvertedDestSize, pg, zn, zd);
}
// XXX: BFCVTNT
// SVE2 floating-point pairwise operations
void faddp(SubRegSize size, ZRegister zd, PRegisterMerge pg, ZRegister zn, ZRegister zm) {
SVEFloatPairwiseArithmetic(0b000, size, pg, zd, zn, zm);
@@ -2312,15 +2313,15 @@ public:
// SVE floating-point convert precision
void fcvt(SubRegSize to, SubRegSize from, ZRegister zd, PRegisterMerge pg, ZRegister zn) {
LOGMAN_THROW_A_FMT(to != from, "to and from sizes cannot be the same.");
LOGMAN_THROW_A_FMT(to != SubRegSize::i8Bit && from != SubRegSize::i8Bit, "Can't use 8-bit element size");
SVEFPConvertPrecision(to, from, zd, pg, zn);
}
void fcvtx(ZRegister zd, PRegisterMerge pg, ZRegister zn) {
LOGMAN_THROW_A_FMT(pg <= PReg::p7, "Can only use p0-p7 as a governing predicate");
uint32_t Instr = 0b0110'0101'0000'1010'1010'0000'0000'0000;
Instr |= pg.Idx() << 10;
Instr |= zn.Idx() << 5;
Instr |= zd.Idx();
dc32(Instr);
SVEFPConvertPrecision(SubRegSize::i32Bit, SubRegSize::i8Bit, zd, pg, zn);
}
void bfcvt(ZRegister zd, PRegisterMerge pg, ZRegister zn) {
SVEFPConvertPrecision(SubRegSize::i32Bit, SubRegSize::i32Bit, zd, pg, zn);
}
// SVE floating-point unary operations
@@ -3847,14 +3848,19 @@ private:
void SVEFPConvertPrecision(SubRegSize to, SubRegSize from, ZRegister zd, PRegister pg, ZRegister zn) {
LOGMAN_THROW_A_FMT(pg <= PReg::p7, "Can only use p0-p7 as a governing predicate");
LOGMAN_THROW_A_FMT(to != from, "to and from sizes cannot be the same.");
LOGMAN_THROW_A_FMT(to != SubRegSize::i8Bit && to != SubRegSize::i128Bit && from != SubRegSize::i8Bit && from != SubRegSize::i128Bit,
"Can't use 8-bit or 128-bit element size");
LOGMAN_THROW_A_FMT(to != SubRegSize::i128Bit && from != SubRegSize::i128Bit, "Can't use 128-bit element size");
// Encodings for the to and from sizes can get a little funky
// depending on what is being converted to/from.
const uint32_t op = [&] {
switch (from) {
case SubRegSize::i8Bit: {
switch (to) {
case SubRegSize::i32Bit: return 0x00020000U;
default: return UINT32_MAX;
}
}
case SubRegSize::i16Bit: {
switch (to) {
case SubRegSize::i32Bit: return 0x00810000U;
@@ -3866,6 +3872,7 @@ private:
case SubRegSize::i32Bit: {
switch (to) {
case SubRegSize::i16Bit: return 0x00800000U;
case SubRegSize::i32Bit: return 0x00820000U;
case SubRegSize::i64Bit: return 0x00C30000U;
default: return UINT32_MAX;
}
+2 -2
View File
@@ -9,8 +9,8 @@ set(CMAKE_AR ${MINGW_TRIPLE}-ar)
# Compile everything as static to avoid requiring the MinGW runtime libraries, force page aligned sections so that
# debug symbols work correctly, and disable loop alignment to workaround an LLVM bug
# (https://github.com/llvm/llvm-project/issues/47432)
set(CMAKE_SHARED_LINKER_FLAGS_INIT "-static -static-libgcc -static-libstdc++ -Wl,--file-alignment=4096,/mllvm:-align-loops=1")
set(CMAKE_EXE_LINKER_FLAGS_INIT "-static -static-libgcc -static-libstdc++ -Wl,--file-alignment=4096,/mllvm:-align-loops=1")
set(CMAKE_SHARED_LINKER_FLAGS_INIT "-static -Wl,--file-alignment=4096,/mllvm:-align-loops=1")
set(CMAKE_EXE_LINKER_FLAGS_INIT "-static -Wl,--file-alignment=4096,/mllvm:-align-loops=1")
set(CMAKE_C_STANDARD_LIBRARIES "" CACHE STRING "" FORCE)
set(CMAKE_CXX_STANDARD_LIBRARIES "" CACHE STRING "" FORCE)
set(CMAKE_STANDARD_LIBRARIES "" CACHE STRING "" FORCE)
+50 -53
View File
@@ -210,56 +210,53 @@ click==8.1.7 \
--hash=sha256:ae74fb96c20a0277a1d615f1e4d73c8414f5a98db8b799a7931d1582f3390c28 \
--hash=sha256:ca9853ad459e787e2192211578cc907e7594e294c7ccc834310722b41b9ca6de
# via black
cryptography==48.0.0 \
--hash=sha256:0890f502ddf7d9c6426129c3f49f5c0a39278ed7cd6322c8755ffca6ee675a13 \
--hash=sha256:0c558d2cdffd8f4bbb30fc7134c74d2ca9a476f830bb053074498fbc86f41ed6 \
--hash=sha256:16cd65b9330583e4619939b3a3843eec1e6e789744bb01e7c7e2e62e33c239c8 \
--hash=sha256:18349bbc56f4743c8b12dc32e2bccb2cf83ee8b69a3bba74ef8ae857e26b3d25 \
--hash=sha256:1e2d54c8be6152856a36f0882ab231e70f8ec7f14e93cf87db8a2ed056bf160c \
--hash=sha256:22a5cb272895dce158b2cacdfdc3debd299019659f42947dbdac6f32d68fe832 \
--hash=sha256:27241b1dc9962e056062a8eef1991d02c3a24569c95975bd2322a8a52c6e5e12 \
--hash=sha256:2b4d59804e8408e2fea7d1fbaf218e5ec984325221db76e6a241a9abd6cdd95c \
--hash=sha256:2eb992bbd4661238c5a397594c83f5b4dc2bc5b848c365c8f991b6780efcc5c7 \
--hash=sha256:369a6348999f94bbd53435c894377b20ab95f25a9065c283570e70150d8abc3c \
--hash=sha256:3cb07a3ed6431663cd321ea8a000a1314c74211f823e4177fefa2255e057d1ec \
--hash=sha256:40ba1f85eaa6959837b1d51c9767e230e14612eea4ef110ee8854ada22da1bf5 \
--hash=sha256:4defde8685ae324a9eb9d818717e93b4638ef67070ac9bc15b8ca85f63048355 \
--hash=sha256:55b7718303bf06a5753dcdccf2f3945cf18ad7bffde41b61226e4db31ab89a9c \
--hash=sha256:561215ea3879cb1cbbf272867e2efda62476f240fb58c64de6b393ae19246741 \
--hash=sha256:58d00498e8933e4a194f3076aee1b4a97dfec1a6da444535755822fe5d8b0b86 \
--hash=sha256:59baa2cb386c4f0b9905bd6eb4c2a79a69a128408fd31d32ca4d7102d4156321 \
--hash=sha256:5a5ed8fde7a1d09376ca0b40e68cd59c69fe23b1f9768bd5824f54681626032a \
--hash=sha256:5b012212e08b8dd5edc78ef54da83dd9892fd9105323b3993eff6bea65dc21d7 \
--hash=sha256:5c3932f4436d1cccb036cb0eaef46e6e2db91035166f1ad6505c3c9d5a635920 \
--hash=sha256:614d0949f4790582d2cc25553abd09dd723025f0c0e7c67376a1d77196743d6e \
--hash=sha256:76341972e1eff8b4bea859f09c0d3e64b96ce931b084f9b9b7db8ef364c30eff \
--hash=sha256:77a2ccbbe917f6710e05ba9adaa25fb5075620bf3ea6fb751997875aff4ae4bd \
--hash=sha256:7995ef305d7165c3f11ae07f2517e5a4f1d5c18da1376a0a9ed496336b69e5f3 \
--hash=sha256:7ce4bfae76319a532a2dc68f82cc32f5676ee792a983187dac07183690e5c66f \
--hash=sha256:7e8eac43dfca5c4cccc6dad9a80504436fca53bb9bc3100a2386d730fbe6b602 \
--hash=sha256:84cf79f0dc8b36ac5da873481716e87aef31fcfa0444f9e1d8b4b2cece142855 \
--hash=sha256:8c7378637d7d88016fa6791c159f698b3d3eed28ebf844ac36b9dc04a14dae18 \
--hash=sha256:8cd666227ef7af430aa5914a9910e0ddd703e75f039cef0825cd0da71b6b711a \
--hash=sha256:906cbf0670286c6e0044156bc7d4af9cbb0ef6db9f73e52c3ec56ba6bdde5336 \
--hash=sha256:9071196d81abc88b3516ac8cdfad32e2b66dd4a5393a8e68a961e9161ddc6239 \
--hash=sha256:9249e3cd978541d665967ac2cb2787fd6a62bddf1e75b3e347a594d7dacf4f74 \
--hash=sha256:984a20b0f62a26f48a3396c72e4bc34c66e356d356bf370053066b3b6d54634a \
--hash=sha256:9be5aafa5736574f8f15f262adc81b2a9869e2cfe9014d52a44633905b40d52c \
--hash=sha256:9c459db21422be75e2809370b829a87eb37f74cd785fc4aa9ea1e5f43b47cda4 \
--hash=sha256:9ccdac7d40688ecb5a3b4a604b8a88c8002e3442d6c60aead1db2a89a041560c \
--hash=sha256:a0e692c683f4df67815a2d258b324e66f4738bd7a96a218c826dce4f4bd05d8f \
--hash=sha256:a5da777e32ffed6f85a7b2b3f7c5cbc88c146bfcd0a1d7baf5fcc6c52ee35dd4 \
--hash=sha256:a64697c641c7b1b2178e573cbc31c7c6684cd56883a478d75143dbb7118036db \
--hash=sha256:ad64688338ed4bc1a6618076ba75fd7194a5f1797ac60b47afe926285adb3166 \
--hash=sha256:bd72e68b06bb1e96913f97dd4901119bc17f39d4586a5adf2d3e47bc2b9d58b5 \
--hash=sha256:c17dfe85494deaeddc5ce251aebd1d60bbe6afc8b62071bb0b469431a000124f \
--hash=sha256:c18684a7f0cc9a3cb60328f496b8e3372def7c5d2df39ac267878b05565aaaae \
--hash=sha256:cc90c0b39b2e3c65ef52c804b72e3c58f8a04ab2a1871272798e5f9572c17d20 \
--hash=sha256:db63bf618e5dea46c07de12e900fe1cdd2541e6dc9dbae772a70b7d4d4765f6a \
--hash=sha256:ea8990436d914540a40ab24b6a77c0969695ed52f4a4874c5137ccf7045a7057 \
--hash=sha256:ecde28a596bead48b0cfd2a1b4416c3d43074c2d785e3a398d7ec1fc4d0f7fbb \
--hash=sha256:f5333311663ea94f75dd408665686aaf426563556bb5283554a3539177e03b8c \
--hash=sha256:fdfef35d751d510fcef5252703621574364fec16418c4a1e5e1055248401054b
cryptography==50.0.0 \
--hash=sha256:031e2d5dd4bb9caa3ca9c82e5a197fd8ae680232cee62603d1a813f3f07e3d03 \
--hash=sha256:06a32a980526a6ab9a4b9bf8f7385800791e2bb960903cb6b530e4817509a3b7 \
--hash=sha256:07479a1cb08219ab719147e742e76090c9c773321959bb94946fffdd397a6437 \
--hash=sha256:07949c449a1abcf60d1ee6e88956d89404c7df3c8258f46589e912988e551987 \
--hash=sha256:105110f43a471dbd0060b9c9516cb8a6a79233631a04cc2ba16f28323ac6e025 \
--hash=sha256:11b74db56cdbe3cdee6e3f6982ecb70334fa10dce99ed58bf7894aaaa3b2a037 \
--hash=sha256:12b9c6996425c76ea6c457ace4f3073e715b8c545add07cd1a8f3a4f90691269 \
--hash=sha256:1489e263a8048bb8b6a8bac662eb2d402ea5d2b7b4699b72f385f1e2772db105 \
--hash=sha256:19736989797678c6af1e55cd49055cdbcb55d8f6b5583ac5335f933aba9101dc \
--hash=sha256:1b4a266766514614f8aa60416e71f2fc6e575d36e7bdc90f644fadb2f4b75b95 \
--hash=sha256:2a8183b489dc1f7f80f135780fadc1108f14b31b8a40411c7a5b17425f65f28b \
--hash=sha256:37fdb0d0111f1e2ff07139dfb79f1b49531f8e213c46f1163dd7642979b58c47 \
--hash=sha256:3f5735ffe4996d28b809371756219f5354864902a3b9e7c0b9ee87041209fc9c \
--hash=sha256:49e7d93abdbd2990caced757e5fade25302f719c3c8fb6e6fff2dde98999fc41 \
--hash=sha256:5e34edd123674534acd70147f0ca331eaa2c74e6325fb2028c886aa26ba0b68c \
--hash=sha256:62598a8a57f815db4c6259a4e97d857dab56697e7de8e8ab02352ab74da1995d \
--hash=sha256:65c2c3add92b45fd0709db8594536aea39c2a67af0e27ffcf049c498501140b7 \
--hash=sha256:6ba6a53445bd3cfa809ef3ef5f1589aa6ba08784a1d962bf47d0940e871dab1c \
--hash=sha256:6e7d61120573a7f2cd94cc095f9e81f6967c61ccdf194285aa143ecec8e0b708 \
--hash=sha256:7cec5b856506da6defb290f30c9ee687d5f5e8cb0bd3f6459dde43b0b4fa40ef \
--hash=sha256:80b63928fa35083b33966ce1efb70e5b9607181e49dcd1c22c8c005e319f667f \
--hash=sha256:82148ec5bddac30b51a5b3c1945075f896fa022cb93f8e4a01e9f6ee95292c5f \
--hash=sha256:828743d939e9629bc267b8e2d08d8bb67cd4319c771a33d4b18b22dd8fb7440a \
--hash=sha256:8d89f3976b10b4ce31118de72329025f70d2c6ead14a8217c5514dd2c6d5a78f \
--hash=sha256:8eb5e1172eb569ea8a872796576e6a67c276351728b6455d5beb01242b027c6a \
--hash=sha256:900131fafd8aead39ac7dd3a7e833be754c17a95cfd91221636949fe4eb0aa8a \
--hash=sha256:910d11e1a385c654bf738bf3e6b8e6ed5de0f5610fcae2be9e5b398d8081d20e \
--hash=sha256:910e1d2668e7de9648f2bcee30e180db2a6b15c30f887d7c4c93ddf96e3992e3 \
--hash=sha256:9aa87839c383bdbab6ef865787a1fb877af8dd03464c4400322726feaaadfc6d \
--hash=sha256:a1b30560f2acc95aa8b2e06e716a13dbfc97314747b80d9707e307f77b40d6b3 \
--hash=sha256:a91296cb61e8df6f86d0c19cc4068228da256bf59bf86049fbd821084565327f \
--hash=sha256:b42a28c1844fd9de8f3f7d540e36b66f3a9c83fceac7170ebc7a6a19edd9dcae \
--hash=sha256:bd1c592e4d5974f0d08d4888e432157adba757c66da0246918e43677fafa2d30 \
--hash=sha256:c87f62a3d3b9888ed0fdde100ec06aa61ca9cd44bad9057d1dff9a516b5f5bb9 \
--hash=sha256:c99c003e088647b8a5b7c145d6f78c335f6348332b62e142d411c4b63d1460b9 \
--hash=sha256:ccdc4a71a4dabae05de219404f9f4abc38e3b58422177ff93d0da05967dafa07 \
--hash=sha256:d24fead1d4d076e1bfb006dcec392074a3cd8d7b4fc8a595aa64073b2b7a96ba \
--hash=sha256:d58c3db7cd6eed54e6c06744db55456b65ebd7492ddeae9c1e93cfca7aa857d3 \
--hash=sha256:d764dcf130c428ef66786f866dd750f53182bc608813489915e9fc106bb0c82f \
--hash=sha256:df2a58a472f332225671c35b0a830208b86d004f82baa8530fa3782c85646533 \
--hash=sha256:e722f16708d854fe924790e051061f6704a472c3bac347b6fd88033ea8dd0dc5 \
--hash=sha256:ecfed7367f965a0328cfbdd70da860f15441f002f613185668c6e6ebf5a0ac11 \
--hash=sha256:eeac2acb5a20ed25e0ad6d1df9891a520b78b404266b6d11778f25d5d691a6c9 \
--hash=sha256:f59e38625469987d7ef6d495323c55e7db6c212eaf6112267e0d3b565a2e9c9f \
--hash=sha256:f89831ef99dd7dd169ab06d63a831adb9e20a87aac6d380266bbda5823349169 \
--hash=sha256:fd9192b7b70c573d7f214eb1ae35e00d359f6f5e4b27c7e21e30de1fc6204645
# via
# -r requirements_formatting.txt.in
# pyjwt
@@ -311,9 +308,9 @@ pygithub==2.6.1 \
--hash=sha256:6f2fa6d076ccae475f9fc392cc6cdbd54db985d4f69b8833a28397de75ed6ca3 \
--hash=sha256:b5c035392991cca63959e9453286b41b54d83bf2de2daa7d7ff7e4312cebf3bf
# via -r requirements_formatting.txt.in
pyjwt==2.12.1 \
--hash=sha256:28ca37c070cad8ba8cd9790cd940535d40274d22f80ab87f3ac6a713e6e8454c \
--hash=sha256:c74a7a2adf861c04d002db713dd85f84beb242228e671280bf709d765b03672b
pyjwt==2.13.0 \
--hash=sha256:41571c89ca91598c79e8ef18a2d07367d4810fbbd6f637794879baf1b7703423 \
--hash=sha256:66adcc2aff09b3f1bbd95fc1e1577df8ac8723c978552fd43304c8a290ac5728
# via
# -r requirements_formatting.txt.in
# pygithub
+2 -2
View File
@@ -1,10 +1,10 @@
black>=26.3.1
darker==2.1.1
PyGithub==2.6.1
cryptography>=46.0.7
cryptography>=50.0.0
urllib3>=2.7.0
requests>=2.33.0
idna>=3.15
certifi>=2024.7.4
PyNaCl>=1.6.2
PyJWT>=2.12.1
PyJWT>=2.13.0
+1 -1
+1 -1
+28
View File
@@ -407,6 +407,32 @@ def print_parse_enum_options(options):
output_argloader.write("#endif\n")
def print_affects_codegen_options(options, unnamed_options):
output_argloader.write("#ifdef CONFIG_AFFECTSCODEGEN\n")
output_argloader.write("#undef CONFIG_AFFECTSCODEGEN\n")
TotalConfigOptions = 0
for op_group, group_vals in options.items():
for op_key, op_vals in group_vals.items():
TotalConfigOptions += 1
for op_group, group_vals in unnamed_options.items():
for op_key, op_vals in group_vals.items():
TotalConfigOptions += 1
output_argloader.write("constexpr static std::array<bool, {}> Config_AffectsCodeGen = {{{{\n".format(TotalConfigOptions))
for op_group, group_vals in options.items():
for op_key, op_vals in group_vals.items():
assert "AffectsCodeGen" in op_vals, "All config options must be marked if they affect codegen."
output_argloader.write("\t{}, // {}\n".format(op_vals["AffectsCodeGen"], op_key))
for op_group, group_vals in unnamed_options.items():
for op_key, op_vals in group_vals.items():
assert "AffectsCodeGen" in op_vals, "All config options must be marked if they affect codegen."
output_argloader.write("\t{}, // {}\n".format(op_vals["AffectsCodeGen"], op_key))
output_argloader.write("}};\n")
output_argloader.write("#endif\n")
if (len(sys.argv) < 5):
sys.exit()
@@ -451,4 +477,6 @@ print_parse_jsonloader_options(options);
# Generate enum variable options
print_parse_enum_options(options);
print_affects_codegen_options(options, unnamed_options);
output_argloader.close()
+64 -60
View File
@@ -251,6 +251,10 @@ def parse_ops(ops):
if "Desc" in op_val:
OpDef.Desc = op_val["Desc"]
if not isinstance(OpDef.Desc, list):
ExitError(f"Desc field for op {OpDef.Name} must be an array of strings")
if not all(isinstance(item, str) for item in OpDef.Desc):
ExitError(f"Desc field for op {OpDef.Name} must only contain strings")
if "DynamicDispatch" in op_val:
OpDef.DynamicDispatch = bool(op_val["DynamicDispatch"])
@@ -603,77 +607,77 @@ def print_validation(op):
def print_ir_allocator_helpers():
output_file.write("#ifdef IROP_ALLOCATE_HELPERS\n")
output_file.write("\ttemplate <class T>\n")
output_file.write("\tstruct Wrapper final {\n")
output_file.write("\t\tT *first;\n")
output_file.write("\t\tOrderedNode *Node; ///< Actual offset of this IR in ths list\n")
output_file.write("\n")
output_file.write("\t\toperator Wrapper<IROp_Header>() const { return Wrapper<IROp_Header> {reinterpret_cast<IROp_Header*>(first), Node}; }\n")
output_file.write("\t\toperator OrderedNode *() { return Node; }\n")
output_file.write("\t\toperator const OrderedNode *() const { return Node; }\n")
output_file.write("\t\toperator OpNodeWrapper () const { return Node->Header.Value; }\n")
output_file.write("\t};\n")
output_file.write("\ttemplate <class T>\n"
"\tstruct Wrapper final {\n"
"\t\tT *first;\n"
"\t\tOrderedNode *Node; ///< Actual offset of this IR in ths list\n"
"\n"
"\t\toperator Wrapper<IROp_Header>() const { return Wrapper<IROp_Header> {reinterpret_cast<IROp_Header*>(first), Node}; }\n"
"\t\toperator OrderedNode *() { return Node; }\n"
"\t\toperator const OrderedNode *() const { return Node; }\n"
"\t\toperator OpNodeWrapper () const { return Node->Header.Value; }\n"
"\t};\n")
output_file.write("\ttemplate <class T>\n")
output_file.write("\tusing IRPair = Wrapper<T>;\n\n")
output_file.write("\ttemplate <class T>\n"
"\tusing IRPair = Wrapper<T>;\n\n")
output_file.write("\tIRPair<IROp_Header> AllocateRawOp(size_t HeaderSize) {\n")
output_file.write("\t\tauto Op = reinterpret_cast<IROp_Header*>(DualListData.DataAllocate(HeaderSize));\n")
output_file.write("\t\tmemset(Op, 0, HeaderSize);\n")
output_file.write("\t\tOp->Op = IROps::OP_DUMMY;\n")
output_file.write("\t\treturn IRPair<IROp_Header>{Op, CreateNode(Op)};\n")
output_file.write("\t}\n\n")
output_file.write("\tIRPair<IROp_Header> AllocateRawOp(size_t HeaderSize) {\n"
"\t\tauto Op = reinterpret_cast<IROp_Header*>(DualListData.DataAllocate(HeaderSize));\n"
"\t\tmemset(Op, 0, HeaderSize);\n"
"\t\tOp->Op = IROps::OP_DUMMY;\n"
"\t\treturn IRPair<IROp_Header>{Op, CreateNode(Op)};\n"
"\t}\n\n")
output_file.write("\ttemplate<class T, IROps T2>\n")
output_file.write("\tT *AllocateOrphanOp() {\n")
output_file.write("\t\tsize_t Size = FEXCore::IR::GetSize(T2);\n")
output_file.write("\t\tauto Op = reinterpret_cast<T*>(DualListData.DataAllocate(Size));\n")
output_file.write("\t\tmemset(Op, 0, Size);\n")
output_file.write("\t\tOp->Header.Op = T2;\n")
output_file.write("\t\treturn Op;\n")
output_file.write("\t}\n\n")
output_file.write("\ttemplate<class T, IROps T2>\n"
"\tT *AllocateOrphanOp() {\n"
"\t\tsize_t Size = FEXCore::IR::GetSize(T2);\n"
"\t\tauto Op = reinterpret_cast<T*>(DualListData.DataAllocate(Size));\n"
"\t\tmemset(Op, 0, Size);\n"
"\t\tOp->Header.Op = T2;\n"
"\t\treturn Op;\n"
"\t}\n\n")
output_file.write("\ttemplate<class T, IROps T2>\n")
output_file.write("\tIRPair<T> AllocateOp() {\n")
output_file.write("\t\tsize_t Size = FEXCore::IR::GetSize(T2);\n")
output_file.write("\t\tauto Op = reinterpret_cast<T*>(DualListData.DataAllocate(Size));\n")
output_file.write("\t\tmemset(Op, 0, Size);\n")
output_file.write("\t\tOp->Header.Op = T2;\n")
output_file.write("\t\treturn IRPair<T>{Op, CreateNode(&Op->Header)};\n")
output_file.write("\t}\n\n")
output_file.write("\ttemplate<class T, IROps T2>\n"
"\tIRPair<T> AllocateOp() {\n"
"\t\tsize_t Size = FEXCore::IR::GetSize(T2);\n"
"\t\tauto Op = reinterpret_cast<T*>(DualListData.DataAllocate(Size));\n"
"\t\tmemset(Op, 0, Size);\n"
"\t\tOp->Header.Op = T2;\n"
"\t\treturn IRPair<T>{Op, CreateNode(&Op->Header)};\n"
"\t}\n\n")
output_file.write("\tIR::OpSize GetOpSize(const OrderedNode *Op) const {\n")
output_file.write("\t\tauto HeaderOp = Op->Header.Value.GetNode(DualListData.DataBegin());\n")
output_file.write("\t\treturn HeaderOp->Size;\n")
output_file.write("\t}\n\n")
output_file.write("\tIR::OpSize GetOpSize(const OrderedNode *Op) const {\n"
"\t\tauto HeaderOp = Op->Header.Value.GetNode(DualListData.DataBegin());\n"
"\t\treturn HeaderOp->Size;\n"
"\t}\n\n")
output_file.write("\tIR::OpSize GetOpElementSize(const OrderedNode *Op) const {\n")
output_file.write("\t\tauto HeaderOp = Op->Header.Value.GetNode(DualListData.DataBegin());\n")
output_file.write("\t\treturn HeaderOp->ElementSize;\n")
output_file.write("\t}\n\n")
output_file.write("\tIR::OpSize GetOpElementSize(const OrderedNode *Op) const {\n"
"\t\tauto HeaderOp = Op->Header.Value.GetNode(DualListData.DataBegin());\n"
"\t\treturn HeaderOp->ElementSize;\n"
"\t}\n\n")
output_file.write("\tuint8_t GetOpElements(const OrderedNode *Op) const {\n")
output_file.write("\t\tLOGMAN_THROW_A_FMT(OpHasDest(Op), \"Op {} has no dest\\n\", GetOpName(Op));\n")
output_file.write("\t\treturn IR::OpSizeToSize(GetOpSize(Op)) / IR::OpSizeToSize(GetOpElementSize(Op));\n")
output_file.write("\t}\n\n")
output_file.write("\tuint8_t GetOpElements(const OrderedNode *Op) const {\n"
"\t\tLOGMAN_THROW_A_FMT(OpHasDest(Op), \"Op {} has no dest\\n\", GetOpName(Op));\n"
"\t\treturn IR::OpSizeToSize(GetOpSize(Op)) / IR::OpSizeToSize(GetOpElementSize(Op));\n"
"\t}\n\n")
output_file.write("\tbool OpHasDest(const OrderedNode *Op) const {\n")
output_file.write("\t\tauto HeaderOp = Op->Header.Value.GetNode(DualListData.DataBegin());\n")
output_file.write("\t\treturn GetHasDest(HeaderOp->Op);\n")
output_file.write("\t}\n\n")
output_file.write("\tbool OpHasDest(const OrderedNode *Op) const {\n"
"\t\tauto HeaderOp = Op->Header.Value.GetNode(DualListData.DataBegin());\n"
"\t\treturn GetHasDest(HeaderOp->Op);\n"
"\t}\n\n")
output_file.write("\tIROps GetOpType(const OrderedNode *Op) const {\n")
output_file.write("\t\tauto HeaderOp = Op->Header.Value.GetNode(DualListData.DataBegin());\n")
output_file.write("\t\treturn HeaderOp->Op;\n")
output_file.write("\t}\n\n")
output_file.write("\tIROps GetOpType(const OrderedNode *Op) const {\n"
"\t\tauto HeaderOp = Op->Header.Value.GetNode(DualListData.DataBegin());\n"
"\t\treturn HeaderOp->Op;\n"
"\t}\n\n")
output_file.write("\tFEXCore::IR::RegClass GetOpRegClass(const OrderedNode *Op) const {\n")
output_file.write("\t\treturn GetRegClass(GetOpType(Op));\n")
output_file.write("\t}\n\n")
output_file.write("\tFEXCore::IR::RegClass GetOpRegClass(const OrderedNode *Op) const {\n"
"\t\treturn GetRegClass(GetOpType(Op));\n"
"\t}\n\n")
output_file.write("\tstd::string_view const& GetOpName(const OrderedNode *Op) const {\n")
output_file.write("\t\treturn IR::GetName(GetOpType(Op));\n")
output_file.write("\t}\n\n")
output_file.write("\tstd::string_view const& GetOpName(const OrderedNode *Op) const {\n"
"\t\treturn IR::GetName(GetOpType(Op));\n"
"\t}\n\n")
# Generate helpers with operands
for op in IROps:
+4
View File
@@ -18,12 +18,14 @@ set(SRCS
Common/JitSymbols.cpp
Interface/Context/Context.cpp
Interface/Core/LookupCache.cpp
Interface/Core/DiskCache.cpp
Interface/Core/CodeCache.cpp
Interface/Core/Core.cpp
Interface/Core/CPUBackend.cpp
Interface/Core/Addressing.cpp
Interface/Core/CPUID.cpp
Interface/Core/Frontend.cpp
Interface/Core/SharedCodeBufferManager.cpp
Interface/Core/OpcodeDispatcher/AVX_128.cpp
Interface/Core/OpcodeDispatcher/Crypto.cpp
Interface/Core/OpcodeDispatcher/Flags.cpp
@@ -68,6 +70,7 @@ set(SRCS
Utils/LongJump.cpp
Utils/Telemetry.cpp
Utils/Threads.cpp
Utils/WorkQueueThread.cpp
Utils/Profiler.cpp)
if (ARCHITECTURE_arm64)
@@ -300,6 +303,7 @@ add_library(JemallocLibs STATIC Utils/AllocatorHooks.cpp)
if (ENABLE_FEX_ALLOCATOR)
target_compile_definitions(JemallocLibs PRIVATE ENABLE_FEX_ALLOCATOR=1)
target_link_libraries(JemallocLibs PUBLIC rpmalloc)
target_include_directories(JemallocLibs PRIVATE "${PROJECT_SOURCE_DIR}/include/")
endif()
if (ENABLE_JEMALLOC_GLIBC_ALLOC)
set_source_files_properties(Interface/HLE/Thunks/Thunks.cpp PROPERTIES COMPILE_DEFINITIONS ENABLE_JEMALLOC_GLIBC=1)
+19 -14
View File
@@ -18,7 +18,7 @@ struct BitSet final {
constexpr static size_t MinimumSize = sizeof(ElementType);
constexpr static size_t MinimumSizeBits = sizeof(ElementType) * 8;
ElementType* Memory;
ElementType* Memory {};
void Allocate(size_t Elements) {
size_t AllocateSize = ToBytes(Elements);
LOGMAN_THROW_A_FMT((AllocateSize * MinimumSize) >= Elements, "Fail");
@@ -33,14 +33,15 @@ struct BitSet final {
FEXCore::Allocator::free(Memory);
Memory = nullptr;
}
bool Get(T Element) {
[[nodiscard]]
bool Get(T Element) const {
return (Memory[Element / MinimumSizeBits] & (1ULL << (Element % MinimumSizeBits))) != 0;
}
void Set(T Element) {
Memory[Element / MinimumSizeBits] |= (1ULL << (Element % MinimumSizeBits));
}
void Clear(T Element) {
Memory[Element / MinimumSizeBits] &= (1ULL << (Element % MinimumSizeBits));
Memory[Element / MinimumSizeBits] &= ~(1ULL << (Element % MinimumSizeBits));
}
void MemClear(size_t Elements) {
memset(Memory, 0, ToBytes(Elements));
@@ -48,13 +49,15 @@ struct BitSet final {
void MemSet(size_t Elements) {
memset(Memory, 0xFF, ToBytes(Elements));
}
uint32_t ToBytes(size_t Elements) {
return AlignUp(Elements, MinimumSizeBits) / MinimumSize;
[[nodiscard]]
static size_t ToBytes(size_t Elements) {
return AlignUp(Elements, MinimumSizeBits) / 8;
}
// This very explicitly doesn't let you take an address
// Is only a getter
bool operator[](T Element) {
[[nodiscard]]
bool operator[](T Element) const {
return Get(Element);
}
};
@@ -62,35 +65,37 @@ struct BitSet final {
template<typename T>
struct BitSetView final {
using ElementType = T;
constexpr static size_t MinimumSize = sizeof(ElementType);
constexpr static size_t MinimumSizeBits = sizeof(ElementType) * 8;
constexpr static size_t MinimumSize = BitSet<T>::MinimumSize;
constexpr static size_t MinimumSizeBits = BitSet<T>::MinimumSizeBits;
ElementType* Memory;
ElementType* Memory {};
void GetView(BitSet<T>& Set, uint64_t ElementOffset) {
LOGMAN_THROW_A_FMT((ElementOffset % MinimumSize) == 0, "Bitset view offset needs to be aligned to size of backing element");
Memory = &Set.Memory[ElementOffset / MinimumSizeBits];
}
bool Get(T Element) {
[[nodiscard]]
bool Get(T Element) const {
return (Memory[Element / MinimumSizeBits] & (1ULL << (Element % MinimumSizeBits))) != 0;
}
void Set(T Element) {
Memory[Element / MinimumSizeBits] |= (1ULL << (Element % MinimumSizeBits));
}
void Clear(T Element) {
Memory[Element / MinimumSizeBits] &= (1ULL << (Element % MinimumSizeBits));
Memory[Element / MinimumSizeBits] &= ~(1ULL << (Element % MinimumSizeBits));
}
void MemClear(size_t Elements) {
memset(Memory, 0, AlignUp(Elements / MinimumSizeBits, MinimumSizeBits));
memset(Memory, 0, BitSet<T>::ToBytes(Elements));
}
void MemSet(size_t Elements) {
memset(Memory, 0xFF, AlignUp(Elements / MinimumSizeBits, MinimumSizeBits));
memset(Memory, 0xFF, BitSet<T>::ToBytes(Elements));
}
// This very explicitly doesn't let you take an address
// Is only a getter
bool operator[](T Element) {
[[nodiscard]]
bool operator[](T Element) const {
return Get(Element);
}
};
+1
View File
@@ -4,6 +4,7 @@
#include <concepts>
#include <string_view>
#include <cstdlib>
namespace FEXCore::StrConv {
template<std::integral T>
+59 -2
View File
@@ -1,9 +1,10 @@
// SPDX-License-Identifier: MIT
#include "Common/StringConv.h"
#include "FEXCore/Utils/EnumUtils.h"
#include "Utils/Config.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/Utils/Allocator.h>
#include <FEXCore/Utils/EnumUtils.h>
#include <FEXCore/Utils/FileLoading.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/StringUtils.h>
@@ -37,8 +38,16 @@ namespace detail {
#define OPT_STRARRAY(group, enum, json, default) OPT_STR(group, enum, json, default)
#define OPT_STRENUM(group, enum, json, default) const uint64_t P(enum) = FEXCore::ToUnderlying(P(default));
#include <FEXCore/Config/ConfigValues.inl>
constexpr static std::array<std::string_view, FEXCore::Config::ConfigOption::CONFIG_MAX> option_names = {
#define OPT_BASE(type, group, enum, json, default) #json,
#include <FEXCore/Config/ConfigValues.inl>
};
} // namespace detail
std::string_view GetConfigJSONName(FEXCore::Config::ConfigOption option) {
return FEXCore::Config::detail::option_names[option];
}
enum Paths {
PATH_DATA_DIR_LOCAL = 0,
PATH_DATA_DIR_GLOBAL,
@@ -47,6 +56,7 @@ enum Paths {
PATH_CONFIG_FILE_LOCAL,
PATH_CONFIG_FILE_GLOBAL,
PATH_CONFIG_TELEMETRY_FOLDER,
PATH_CACHE_DIR,
PATH_LAST,
};
static std::array<fextl::string, Paths::PATH_LAST> Paths;
@@ -63,6 +73,10 @@ void SetConfigFileLocation(const std::string_view Path, bool Global) {
Paths[PATH_CONFIG_FILE_LOCAL + Global] = Path;
}
void SetCacheDirectory(const std::string_view Path) {
Paths[PATH_CACHE_DIR] = Path;
}
const fextl::string& GetTelemetryDirectory() {
auto& Path = Paths[PATH_CONFIG_TELEMETRY_FOLDER];
if (Path.empty()) {
@@ -90,6 +104,10 @@ const fextl::string& GetConfigFileLocation(bool Global) {
return Paths[PATH_CONFIG_FILE_LOCAL + Global];
}
const fextl::string& GetCacheDirectory() {
return Paths[PATH_CACHE_DIR];
}
fextl::string GetApplicationConfig(const std::string_view Program, bool Global) {
fextl::string ConfigFile = GetConfigDirectory(Global);
@@ -252,7 +270,7 @@ void Load() {
}
}
fextl::string ExpandPath(const fextl::string& ContainerPrefix, const fextl::string& PathName) {
static fextl::string ExpandPath(const fextl::string& ContainerPrefix, const fextl::string& PathName) {
if (PathName.empty()) {
return {};
}
@@ -501,4 +519,43 @@ void Value<T>::GetListIfExists(FEXCore::Config::ConfigOption Option, StringArray
}
}
template void Value<StringArrayType>::GetListIfExists(FEXCore::Config::ConfigOption Option, StringArrayType* List);
#define CONFIG_AFFECTSCODEGEN
#include <FEXCore/Config/ConfigOptions.inl>
fextl::string SerializeForCache() {
fextl::string Config {};
auto append_string_triple = [](fextl::string& Config, std::string_view Key, ConfigOption Option, auto Value) {
Config.append(Key);
Config.append(1, '\0');
Config.append(fextl::fmt::format("{}", FEXCore::ToUnderlying(Option)));
Config.append(1, '\0');
Config.append(fextl::fmt::format("{}", Value));
Config.append(1, '\0');
};
const auto SerializeValue = [&Config, append_string_triple]<typename T, ConfigOption Option>(auto ConfigVal, const auto Default) {
if (!Config_AffectsCodeGen[FEXCore::ToUnderlying(Option)]) {
// Skip everything that the config says doesn't affect codegen.
return;
}
append_string_triple(Config, FEXCore::Config::GetConfigJSONName(Option), Option, ConfigVal());
};
#define OPT_BASE(type, group, enum, json, default) \
SerializeValue.template operator()<type, CONFIG_##enum>(FEXCore::Config::Get_##enum(), default);
#define OPT_STR(group, enum, json, default) \
SerializeValue.template operator()<fextl::string, CONFIG_##enum>(FEXCore::Config::Get_##enum(), default);
#define OPT_STRARRAY(group, enum, json, default) // Unsupported.
#define OPT_STRENUM(group, enum, json, default) // Unsupported.
#include <FEXCore/Config/ConfigValues.inl>
return Config;
}
FEX_DEFAULT_VISIBILITY bool CheckConfigMatches(std::string_view Config) {
// Serialize current config and just check if it matches.
return SerializeForCache() == Config;
}
} // namespace FEXCore::Config
+134 -3
View File
@@ -4,6 +4,7 @@
"Multiblock": {
"Type": "bool",
"Default": "true",
"AffectsCodeGen": "true",
"Desc": [
"Controls multiblock code compilation",
"Can cause long JIT compilation times and stutter"
@@ -12,6 +13,7 @@
"MaxInst": {
"Type": "int32",
"Default": "5000",
"AffectsCodeGen": "true",
"Desc": [
"Maximum number of instruction to store in a block"
]
@@ -19,6 +21,7 @@
"EnableCodeCachingWIP": {
"Type": "bool",
"Default": "false",
"AffectsCodeGen": "true",
"Desc": [
"Enable the code caching subsystem"
]
@@ -26,6 +29,7 @@
"EnableLazyCodeCachingWIP": {
"Type": "bool",
"Default": "false",
"AffectsCodeGen": "false",
"Desc": [
"Enable lazy loading of chunks in code caches"
]
@@ -33,6 +37,7 @@
"EnableCodeCacheValidation": {
"Type": "bool",
"Default": "false",
"AffectsCodeGen": "false",
"Desc": [
"Enable expensive validation when loading code caches"
]
@@ -40,6 +45,8 @@
"HostFeatures": {
"Type": "strenum",
"Default": "FEXCore::Config::HostFeatures::OFF",
"AffectsCodeGen": "true",
"Comment": "Technically affects codegen, but this is serialized elsewhere.",
"Enums": {
"ENABLESVE": "enablesve",
"DISABLESVE": "disablesve",
@@ -115,6 +122,7 @@
"SmallTSCScale": {
"Type": "bool",
"Default": "true",
"AffectsCodeGen": "true",
"Desc": [
"Scales the cycle counter on systems that have low frequencies."
]
@@ -122,6 +130,7 @@
"HideHybrid": {
"Type": "bool",
"Default": "true",
"AffectsCodeGen": "false",
"Desc": [
"Hides hybrid CPU core arrangement."
]
@@ -129,15 +138,74 @@
"CPUFeatureRegisters": {
"Type": "str",
"Default": "",
"AffectsCodeGen": "false",
"Comment": "Technically affects codegen, but this is serialized in to HostFeatures.",
"Desc": [
"Allows overriding cpu feature flags for manual testing"
]
},
"DiskCache": {
"Type": "bool",
"Default": "false",
"AffectsCodeGen": "false",
"Desc": [
"Enables disk caching for code blocks"
]
},
"DiskCacheFileMapping": {
"Type": "bool",
"Default": "true",
"AffectsCodeGen": "false",
"Desc": [
"Maps cache files for faster reading"
]
},
"DiskCacheValidation": {
"Type": "bool",
"Default": "false",
"AffectsCodeGen": "false",
"Desc": [
"Debug mode that does nothing but validate code hits"
]
},
"DiskCacheRelocationFilter": {
"Type": "bool",
"Default": "true",
"AffectsCodeGen": "false",
"Desc": [
"Don't cache blocks with relocations pointing outside of any known region"
]
},
"DiskCacheAnonCaching": {
"Type": "bool",
"Default": "true",
"AffectsCodeGen": "false",
"Desc": [
"Attempt to cache anonymous code"
]
},
"DiskCachePath": {
"Type": "str",
"Default": "",
"AffectsCodeGen": "false",
"Desc": [
"Optional base directory override for disk cache"
]
},
"DiskCacheRODBNames": {
"Type": "str",
"Default": "",
"AffectsCodeGen": "false",
"Desc": [
"Optional list of extra read-only disk cache DBs to consider"
]
}
},
"Emulation": {
"RootFS": {
"Type": "str",
"Default": "",
"AffectsCodeGen": "false",
"Desc": [
"Which Root filesystem prefix to use",
"This can be a filesystem path",
@@ -152,6 +220,7 @@
"ThunkHostLibs": {
"Type": "str",
"Default": "@CMAKE_INSTALL_FULL_LIBDIR@/fex-emu/HostThunks",
"AffectsCodeGen": "false",
"Desc": [
"Folder to find the host-side thunking libraries."
]
@@ -159,6 +228,7 @@
"ThunkGuestLibs": {
"Type": "str",
"Default": "@CMAKE_INSTALL_PREFIX@/share/fex-emu/GuestThunks",
"AffectsCodeGen": "false",
"Desc": [
"Folder to find the guest-side thunking libraries."
]
@@ -166,6 +236,7 @@
"ThunkConfig": {
"Type": "str",
"Default": "",
"AffectsCodeGen": "false",
"Desc": [
"A json file specifying where to overlay the thunks.",
"This can be a filesystem path",
@@ -180,6 +251,7 @@
"Env": {
"Type": "strarray",
"Default": "",
"AffectsCodeGen": "false",
"Desc": [
"Adds an environment variable to the emulated environment."
]
@@ -187,6 +259,7 @@
"HostEnv": {
"Type": "strarray",
"Default": "",
"AffectsCodeGen": "false",
"Desc": [
"Adds an environment variable to the host environment.",
"This can be useful for setting environment variables that thunks can pick up.",
@@ -196,6 +269,7 @@
"AdditionalArguments": {
"Type": "strarray",
"Default": "",
"AffectsCodeGen": "false",
"Desc": [
"Allows the user to pass additional arguments to the application"
]
@@ -203,6 +277,7 @@
"DisableL2Cache": {
"Type": "bool",
"Default": "true",
"AffectsCodeGen": "false",
"Desc": [
"Disables FEXCore's JIT L2 cache lookup. Saving memory.",
"Can potentially introduce more stutters."
@@ -211,6 +286,7 @@
"DynamicL1Cache": {
"Type": "bool",
"Default": "true",
"AffectsCodeGen": "false",
"Desc": [
"Switches FEXCore's JIT L1 cache to be dynamically sized. Saving memory.",
"Can potentially introduce more stutters."
@@ -219,6 +295,7 @@
"DynamicL1CacheIncreaseCountHeuristic": {
"Type": "uint64",
"Default": "250",
"AffectsCodeGen": "false",
"Desc": [
"Threshold of lookups per second that the L1 dynamic cache should increase its size.",
"Lower numbers means more aggressive scaling upward to the maximum size.",
@@ -230,6 +307,7 @@
"DynamicL1CacheDecreaseCountHeuristic": {
"Type": "uint64",
"Default": "50",
"AffectsCodeGen": "false",
"Desc": [
"Threshold of lookups per second that the L1 dynamic cache should decrease its size.",
"The higher the number, the more aggressively it reduces the L1 cache size.",
@@ -243,6 +321,7 @@
"SingleStep": {
"Type": "bool",
"Default": "false",
"AffectsCodeGen": "true",
"Desc": [
"Single stepping configuration."
]
@@ -250,6 +329,7 @@
"GdbServer": {
"Type": "bool",
"Default": "false",
"AffectsCodeGen": "true",
"Desc": [
"Enables the GDB server."
]
@@ -257,6 +337,7 @@
"DumpIR": {
"Type": "str",
"Default": "no",
"AffectsCodeGen": "false",
"Desc": [
"Folder to dump the IR in to.",
"[no, stdout, stderr, server, <Folder>]"
@@ -265,6 +346,7 @@
"PassManagerDumpIR": {
"Type": "strenum",
"Default": "FEXCore::Config::PassManagerDumpIR::OFF",
"AffectsCodeGen": "false",
"Enums": {
"BEFOREOPT": "beforeopt",
"AFTEROPT": "afteropt",
@@ -283,6 +365,7 @@
"DumpGPRs": {
"Type": "bool",
"Default": "false",
"AffectsCodeGen": "false",
"Desc": [
"When the test harness ends, print the GPR state."
]
@@ -290,6 +373,7 @@
"O0": {
"Type": "bool",
"Default": "false",
"AffectsCodeGen": "true",
"Desc": [
"Disables optimizations passes for debugging."
]
@@ -297,6 +381,7 @@
"GlobalJITNaming": {
"Type": "bool",
"Default": "false",
"AffectsCodeGen": "false",
"Desc": [
"Uses JITSymbols to name all JIT state as one symbol",
"Useful for querying how much time is spent inside of the JIT",
@@ -306,6 +391,7 @@
"LibraryJITNaming": {
"Type": "bool",
"Default": "false",
"AffectsCodeGen": "false",
"Desc": [
"Uses JITSymbols to name JIT symbols grouped by library",
"Useful for querying how much time is spent in each guest library",
@@ -315,6 +401,7 @@
"BlockJITNaming": {
"Type": "bool",
"Default": "false",
"AffectsCodeGen": "false",
"Desc": [
"Uses JITSymbols to name JIT symbols",
"Useful for determining hot blocks of code",
@@ -324,6 +411,7 @@
"GDBSymbols": {
"Type": "bool",
"Default": "false",
"AffectsCodeGen": "false",
"Desc": [
"Integrates with GDB using the JIT interface.",
"Needs the fex jit loader in GDB, which can be loaded via `jit-reader-load libFEXGDBReader.so.`",
@@ -334,6 +422,7 @@
"InjectLibSegFault": {
"Type": "bool",
"Default": "false",
"AffectsCodeGen": "false",
"Desc": [
"Sets the environment variable LD_PRELOAD=libSegFault.so",
"This allows the user to very easily enable libSegFault without dealing with environment variables",
@@ -345,6 +434,7 @@
"Disassemble": {
"Type": "strenum",
"Default": "FEXCore::Config::Disassemble::OFF",
"AffectsCodeGen": "false",
"Enums": {
"DISPATCHER": "dispatcher",
"BLOCKS": "blocks",
@@ -361,6 +451,7 @@
"X86Disassemble": {
"Type": "bool",
"Default": "false",
"AffectsCodeGen": "false",
"Desc": [
"Enables x86/x86-64 guest disassembly output for compiled blocks.",
"Requires FEX to be built with -DENABLE_ZYDIS=TRUE"
@@ -369,6 +460,7 @@
"ForceSVEWidth": {
"Type": "uint32",
"Default": "0",
"AffectsCodeGen": "true",
"Desc": [
"Allows overriding the SVE width in the vixl simulator.",
"Useful as a debugging feature."
@@ -377,6 +469,7 @@
"DisableTelemetry": {
"Type": "bool",
"Default": "false",
"AffectsCodeGen": "true",
"Desc": [
"Disables telemetry at runtime.",
"Useful for CI instcountCI mostly"
@@ -387,6 +480,7 @@
"SilentLog": {
"Type": "bool",
"Default": "true",
"AffectsCodeGen": "false",
"Desc": [
"Disables logging"
]
@@ -394,6 +488,7 @@
"OutputLog": {
"Type": "str",
"Default": "server",
"AffectsCodeGen": "false",
"Desc": [
"File to write FEX output to.",
"[stderr, server, <Filename>]"
@@ -402,6 +497,7 @@
"TelemetryDirectory": {
"Type": "str",
"Default": "",
"AffectsCodeGen": "false",
"Desc": [
"Redirects the telemetry folder that FEX usually writes to.",
"By default telemetry data is stored in {$FEX_APP_DATA_LOCATION,{$XDG_DATA_HOME,$HOME}/fex-emu/Telemetry/}"
@@ -410,6 +506,7 @@
"ProfileStats": {
"Type": "bool",
"Default": "false",
"AffectsCodeGen": "false",
"Desc": [
"Enables FEX's low-overhead sampling profile statistics.",
"Requires a supported version of Mangohud to see the results"
@@ -418,6 +515,7 @@
"EnableGpuvisProfiling": {
"Type": "bool",
"Default": "false",
"AffectsCodeGen": "false",
"Desc": [
"Enables profiling when FEX was built with the gpuvis profiler backend."
]
@@ -427,6 +525,7 @@
"SMCChecks": {
"Type": "uint8",
"Default": "FEXCore::Config::CONFIG_SMC_MTRACK",
"AffectsCodeGen": "true",
"TextDefault": "mtrack",
"ArgumentHandler": "SMCCheckHandler",
"Desc": [
@@ -439,6 +538,7 @@
"TSOEnabled": {
"Type": "bool",
"Default": "true",
"AffectsCodeGen": "true",
"Desc": [
"Controls TSO IR ops.",
"Highly likely to break any multithreaded application if disabled."
@@ -447,6 +547,7 @@
"VectorTSOEnabled": {
"Type": "bool",
"Default": "false",
"AffectsCodeGen": "true",
"Desc": [
"When TSO emulation is enabled, controls if vector loadstores should also be atomic."
]
@@ -454,6 +555,7 @@
"MemcpySetTSOEnabled": {
"Type": "bool",
"Default": "false",
"AffectsCodeGen": "true",
"Desc": [
"When TSO emulation is enabled, controls if memcpy and memset should also be atomic.",
"Only affects REP MOVS and REP STOS instructions"
@@ -462,6 +564,7 @@
"HalfBarrierTSOEnabled": {
"Type": "bool",
"Default": "true",
"AffectsCodeGen": "true",
"Desc": [
"When TSO emulation is enabled, controls if unaligned loads and stores should be backpatched to half-barrier atomics.",
"Can be dangerous due to aligned loadstores through the same code now become non-atomic."
@@ -470,6 +573,7 @@
"StrictInProcessSplitLocks": {
"Type": "bool",
"Default": "false",
"AffectsCodeGen": "false",
"Desc": [
"Strict global lock when handling an unaligned atomic that crosses a 16-byte or cacheline granularity",
"This is required to ensure a split-lock doesn't tear inside the process"
@@ -478,6 +582,7 @@
"KernelUnalignedAtomicBackpatching": {
"Type": "bool",
"Default": "true",
"AffectsCodeGen": "false",
"Desc": [
"When the kernel unaligned atomic handler is enabled, use backpatching to reduce kernel context switches."
]
@@ -485,6 +590,7 @@
"VolatileMetadata": {
"Type": "bool",
"Default": "true",
"AffectsCodeGen": "true",
"Desc": [
"Use volatile metadata in PE files to inform TSO instructions when available.",
"When metadata is unavailable falls back to the currently enabled TSO options."
@@ -493,6 +599,7 @@
"X87ReducedPrecision": {
"Type": "bool",
"Default": "false",
"AffectsCodeGen": "true",
"Desc": [
"Emulates X87 floating point using 64-bit precision. This reduces emulation accuracy and may result in rendering bugs."
]
@@ -500,6 +607,7 @@
"StallProcess": {
"Type": "bool",
"Default": "false",
"AffectsCodeGen": "false",
"Desc": [
"Forces a process to stall out on initialization",
"Useful for a process that keeps restarting and doesn't work"
@@ -508,6 +616,7 @@
"HideHypervisorBit": {
"Type": "bool",
"Default": "false",
"AffectsCodeGen": "false",
"Desc": [
"Hides the hypervisor CPUID bit when set.",
"Should only be used for applications that have issues with this set."
@@ -516,6 +625,7 @@
"StartupSleep": {
"Type": "uint32",
"Default": "0",
"AffectsCodeGen": "false",
"Desc": [
"Sleeps the process at startup for a duration of seconds.",
"Useful if an application crashes too quickly to attach a debugger."
@@ -524,6 +634,7 @@
"StartupSleepProcName": {
"Type": "str",
"Default": "",
"AffectsCodeGen": "false",
"Desc": [
"Contrains the startup sleep to only apply to processes that match this name."
]
@@ -531,6 +642,7 @@
"MonoHacks": {
"Type": "bool",
"Default": "true",
"AffectsCodeGen": "true",
"Desc": [
"Permits a hook-based SMC approach and smaller JIT blocks when mono is detected."
]
@@ -540,6 +652,7 @@
"ServerSocketPath": {
"Type": "str",
"Default": "",
"AffectsCodeGen": "false",
"Desc": [
"Override for a FEXServer socket path. Only useful for chroots."
]
@@ -547,6 +660,7 @@
"NeedsSeccomp": {
"Type": "bool",
"Default": "false",
"AffectsCodeGen": "false",
"Desc": [
"Disables inline syscalls in order to support seccomp handling"
]
@@ -554,6 +668,7 @@
"ExtendedVolatileMetadata": {
"Type": "str",
"Default": "",
"AffectsCodeGen": "true",
"Desc": [
"Configuration provided volatile metadata. Only implemented for WoW64/arm64ec.",
"Limited in its use but can be handy.",
@@ -578,15 +693,18 @@
"Misc": {
"INTERPRETER_INSTALLED": {
"Type": "bool",
"Default": "false"
"Default": "false",
"AffectsCodeGen": "false"
},
"APP_FILENAME": {
"Type": "str",
"Default": ""
"Default": "",
"AffectsCodeGen": "false"
},
"APP_CONFIG_NAME": {
"Type": "str",
"Default": "",
"AffectsCodeGen": "false",
"Desc": [
"This is the application config name that has been loaded.",
"This differs from APP_FILENAME in two ways",
@@ -597,16 +715,29 @@
},
"IS64BIT_MODE": {
"Type": "bool",
"Default": "false"
"Default": "false",
"AffectsCodeGen": "false",
"Comment": "Technically affects codegen, but this is serialized elsewhere."
},
"DISABLE_VIXL_INDIRECT_RUNTIME_CALLS": {
"Type": "bool",
"Default": "true",
"AffectsCodeGen": "false",
"Comment": "Technically affects codegen, but only shows up in the test harness.",
"Desc": [
"This option is used for the InstructionCountCI so it can generate the same codegen between Arm64 hosts and vixl simulator hosts.",
"Vixl simulator indirect runtime calls are a special hlt instruction with metadata after it. Effectively making a custom call instruction.",
"With visual simulator calls disabled, the code generation would be the same as on a native Arm64 host, but running the code is broken."
]
},
"CONFIG_VERSION": {
"Type": "uint32",
"Default": "0",
"AffectsCodeGen": "true",
"Comment": [
"Meta option that if config has ever changed definitions dramatically enough that we can rev the version.",
"Be mindful that this will invalidate all caches!"
]
}
}
}
@@ -55,4 +55,10 @@ FEXCore::CPUID::FunctionResults FEXCore::Context::ContextImpl::RunCPUIDFunctionN
bool FEXCore::Context::ContextImpl::IsAddressInCodeBuffer(FEXCore::Core::InternalThreadState* Thread, uintptr_t Address) const {
return Thread->CPUBackend->IsAddressInCodeBuffer(Address) || CodeCache.IsAddressInMappedCodeBuffer(Address);
}
bool FEXCore::Context::ContextImpl::RequiresRelocatableConstants() const {
// Support relocation when generating a cache or when generating reference code for validation
return CodeCache.IsGeneratingCache || FEXCore::Config::Get_ENABLECODECACHEVALIDATION() || DiskCache.IsWritingDiskCache();
}
} // namespace FEXCore::Context
+39 -13
View File
@@ -4,11 +4,13 @@
#include "Common/JitSymbols.h"
#include "Interface/Core/CPUBackend.h"
#include "Interface/Core/CPUID.h"
#include "Interface/Core/SharedCodeBufferManager.h"
#include <Interface/IR/IntrusiveIRList.h>
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/Context.h>
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Core/HostFeatures.h>
#include <FEXCore/Core/DiskCache.h>
#include <FEXCore/IR/IR.h>
#include <FEXCore/Utils/CompilerDefs.h>
#include <FEXCore/Utils/SignalScopeGuards.h>
@@ -120,17 +122,20 @@ public:
* Note that FEX relocations are unrelated to ELF/PE relocations.
*
* @param GuestDelta Guest address offset to apply to RIP-relative data
* @param RelocationOffset Offset to subtract from relocation target offsets
* @param ForStorage True for serializing data (producing deterministic output); false for de-serializing it (resolving dynamic symbols)
*
* @return Returns true on success
*/
[[nodiscard]]
bool ApplyCodeRelocations(uint64_t GuestDelta, std::span<std::byte> Code, std::span<const CPU::Relocation> Relocations,
uint32_t RelocationOffset, bool ForStorage);
bool ApplyCodeRelocations(uint64_t GuestDelta, std::span<std::byte> Code, std::span<const CPU::Relocation> Relocations, bool ForStorage);
// Same but on disk cache packed relocations
[[nodiscard]]
bool ApplyPackedCodeRelocations(uint64_t GuestDelta, std::span<std::byte> Code, std::span<const DiskCache::BlobSmallRelocation> SmallRelocs,
std::span<const DiskCache::BlobThunkRelocation> ThunkRelocs);
};
class ContextImpl final : public FEXCore::Context::Context, public CPU::CodeBufferManager {
class ContextImpl final : public FEXCore::Context::Context, public CPU::SharedCodeBufferManager {
public:
// Context base class implementation.
bool InitCore() override;
@@ -155,32 +160,32 @@ public:
void SetXMMRegistersFromState(FEXCore::Core::InternalThreadState* Thread, const __uint128_t* XMM_Low, const __uint128_t* YMM_High) override;
/**
* @brief Used to create FEX thread objects in preparation for creating a true OS thread. Does set a TID or PID.
* @brief Used to create FEX thread objects in preparation for creating a true OS thread.
*
* @param InitialRIP The starting RIP of this thread
* @param StackPointer The starting RSP of this thread
* @param NewThreadState The initial thread state to setup for our state, if inheriting.
*
* @return The InternalThreadState object that tracks all of the emulated thread's state
*
* Usecases:
* Parent thread Creation:
* - Thread = CreateThread(InitialRIP, InitialStack, nullptr, 0);
* - Thread = CreateThread();
* - Thread->CurrentFrame->State.rip = InitialRIP;
* - Thread->CurrentFrame->State.gregs[FEXCore::X86State::REG_RSP] = InitialStack;
* - CTX->ExecuteThread(Thread);
* OS thread Creation:
* - Thread = CreateThread(0, 0, NewState, PPID);
* - Thread = CreateThread(NewState);
* - Thread->ExecutionThread = FEXCore::Threads::Thread::Create(ThreadHandler, Arg);
* - ThreadHandler calls `CTX->ExecuteThread(Thread)`
* OS fork (New thread created with a clone of thread state):
* - clone{2, 3}
* - Thread = CreateThread(0, 0, CopyOfThreadState, PPID);
* - Thread = CreateThread(CopyOfThreadState);
* - ExecuteThread(Thread); // Starts executing without creating another host thread
* Thunk callback executing guest code from native host thread
* - Thread = CreateThread(0, 0, NewState, PPID);
* - Thread = CreateThread(NewState);
* - HandleCallback(Thread, RIP);
*/
FEXCore::Core::InternalThreadState* CreateThread(uint64_t InitialRIP, uint64_t StackPointer, const FEXCore::Core::CPUState* NewThreadState) override;
FEXCore::Core::InternalThreadState* CreateThread(const FEXCore::Core::CPUState* NewThreadState) override;
/**
* @brief Destroys this FEX thread object and stops tracking it internally
@@ -201,6 +206,8 @@ public:
FEXCore::CPUID::XCRResults RunXCRFunction(uint32_t Function) override;
FEXCore::CPUID::FunctionResults RunCPUIDFunctionName(uint32_t Function, uint32_t Leaf, uint32_t CPU) override;
virtual void InitDiskCache() override {}
CodeCache& GetCodeCache() override {
return CodeCache;
}
@@ -242,11 +249,15 @@ public:
}
void MarkMonoBackpatcherBlock(uint64_t BlockEntry) override;
std::atomic<uint64_t>& GetMonoBackPatcherBlock() {
return MonoBackpatcherBlock;
}
// Manual debugging tooling which is useful for developers.
struct TrackingEmpty {
// RIP stepping handling
virtual void AddSingleStepTarget(uint64_t GuestRIP) {}
virtual void AddSingleStepTargetRange(uint64_t RIPBegin, uint64_t RipEnd) {}
virtual void AllTargetSingleStep() {}
virtual void RemoveSingleStepTarget(uint64_t GuestRIP) {}
virtual bool IsSingleStepTarget(uint64_t GuestRIP) {
@@ -269,6 +280,10 @@ public:
SingleStepTargets.emplace(GuestRIP);
}
virtual void AddSingleStepTargetRange(uint64_t RIPBegin, uint64_t RIPEnd) override {
SingleStepRanges.emplace_back(Range {RIPBegin, RIPEnd});
}
void RemoveSingleStepTarget(uint64_t GuestRIP) override {
SingleStepTargets.erase(GuestRIP);
}
@@ -278,7 +293,7 @@ public:
}
bool IsSingleStepTarget(uint64_t GuestRIP) override {
return SingleStepEverything || SingleStepTargets.contains(GuestRIP);
return SingleStepEverything || SingleStepTargets.contains(GuestRIP) || IsInRange(GuestRIP);
}
void AddWriteWatchPoint(uint64_t Ptr) override {
@@ -302,6 +317,14 @@ public:
fextl::set<uint64_t> SingleStepTargets {};
fextl::set<uint64_t> WatchWriteTargets {};
fextl::set<uint64_t> WatchReadTargets {};
struct Range {
uint64_t Begin, End;
};
fextl::vector<Range> SingleStepRanges {};
bool IsInRange(uint64_t RIP) const {
return std::ranges::any_of(SingleStepRanges, [RIP](const auto& range) { return RIP >= range.Begin && RIP <= range.End; });
}
static bool ContainsRange(const fextl::set<uint64_t>& Set, uint64_t Ptr, size_t Size) {
for (auto it = Set.lower_bound(Ptr); it != Set.end(); --it) {
@@ -361,6 +384,7 @@ public:
FEXCore::HLE::SourcecodeResolver* SourcecodeResolver {};
FEXCore::ThunkHandler* ThunkHandler {};
fextl::unique_ptr<FEXCore::CPU::Dispatcher> Dispatcher;
DiskCache::DiskCache DiskCache;
CodeCache CodeCache;
fextl::unique_ptr<CodeMapWriter> CodeMapWriter;
@@ -438,6 +462,8 @@ public:
return Config.MonoHacks && MonoDetected;
}
bool RequiresRelocatableConstants() const;
protected:
void UpdateAtomicTSOEmulationConfig() {
if (SupportsHardwareTSO) {
@@ -360,6 +360,7 @@ namespace x32 {
Arm64Emitter::Arm64Emitter(FEXCore::Context::ContextImpl* ctx, void* EmissionPtr, size_t size)
: Emitter(static_cast<uint8_t*>(EmissionPtr), size)
, EmitterCTX {ctx}
, SupportCodeRelocations {ctx->RequiresRelocatableConstants()}
#ifdef VIXL_SIMULATOR
, Simulator {&SimDecoder, stdout, vixl::aarch64::SimStack(SimulatorStackSize).Allocate()}
#endif
@@ -425,7 +426,7 @@ void Arm64Emitter::LoadConstant(ARMEmitter::Size s, ARMEmitter::Register Reg, ui
NOPPad = false;
} else if (Pad == PadType::AUTOPAD) {
// Force NOP padding to ensure relocated constants always have enough encoding space available
NOPPad = EnableCodeCaching;
NOPPad = SupportCodeRelocations;
}
bool Is64Bit = s == ARMEmitter::Size::i64Bit;
@@ -633,7 +634,7 @@ void Arm64Emitter::PopCalleeSavedRegisters() {
}
}
void Arm64Emitter::FillSpecialRegs(ARMEmitter::Register TmpReg, ARMEmitter::Register TmpReg2, bool SetFIZ, bool SetPredRegs) {
void Arm64Emitter::FillSpecialRegs(ARMEmitter::Register TmpReg, ARMEmitter::Register TmpReg2, const FillSpecialRegsOptions& Options) {
#ifndef VIXL_SIMULATOR
if (EmitterCTX->HostFeatures.SupportsAFP) {
// Enable AFP features when filling JIT state.
@@ -649,7 +650,7 @@ void Arm64Emitter::FillSpecialRegs(ARMEmitter::Register TmpReg, ARMEmitter::Regi
(1U << 2) | // NEP
(1U << 1)); // AH
if (SetFIZ) {
if (Options.SetFIZ) {
// Insert MXCSR.DAZ in to FIZ
ldr(TmpReg2.W(), STATE.R(), offsetof(FEXCore::Core::CPUState, mxcsr));
bfxil(ARMEmitter::Size::i64Bit, TmpReg, TmpReg2, 6, 1);
@@ -659,7 +660,7 @@ void Arm64Emitter::FillSpecialRegs(ARMEmitter::Register TmpReg, ARMEmitter::Regi
}
#endif
if (SetPredRegs && EmitterCTX->HostFeatures.SupportsSVE()) {
if (Options.SetPredRegs && EmitterCTX->HostFeatures.SupportsSVE()) {
// Set up predicate registers.
// We don't bother spilling these in SpillStaticRegs,
// since all that matters is we restore them on a fill.
@@ -822,7 +823,7 @@ void Arm64Emitter::FillStaticRegs(FillStaticRegOptions Options) {
msr(ARMEmitter::SystemRegister::NZCV, TmpReg);
}
FillSpecialRegs(TmpReg, TmpReg2, true, Options.FPRs);
FillSpecialRegs(TmpReg, TmpReg2, {.SetFIZ = true, .SetPredRegs = Options.FPRs});
if (Options.FPRs) {
if (EmitterCTX->HostFeatures.SupportsAVX && EmitterCTX->HostFeatures.SupportsSVE256) {
@@ -1059,6 +1060,7 @@ size_t Arm64Emitter::SpillForPreserveAllABICall(ARMEmitter::Register TmpReg, boo
SpillStaticRegs(TmpReg, {
.GPRSpillMask = PreserveSRAMask,
.FPRSpillMask = PreserveSRAFPRMask,
.FPRs = FPRs,
});
sub(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, ARMEmitter::Reg::rsp, SPOffset);
@@ -1121,6 +1123,7 @@ void Arm64Emitter::FillForPreserveAllABICall(bool FPRs) {
void Arm64Emitter::Align16B() {
uint64_t CurrentOffset = GetCursorAddress<uint64_t>();
LOGMAN_THROW_A_FMT((CurrentOffset & 3) == 0, "Can't Align16B code that isn't 4-byte aligned!");
for (uint64_t i = (-CurrentOffset & 0xF); i != 0; i -= 4) {
nop();
}
@@ -129,7 +129,21 @@ protected:
std::span<const ARMEmitter::VRegister> GeneralFPRegisters {};
uint32_t PairRegisters = 0;
void FillSpecialRegs(ARMEmitter::Register TmpReg, ARMEmitter::Register TmpReg2, bool SetFIZ, bool SetPredRegs);
bool SupportCodeRelocations;
struct FillSpecialRegsOptions {
// Whether or not to set the FPCR.FIZ (flush inputs to zero) bit in the FPCR to
// the current value of the emulated MXCSR.DAZ bit.
// Will only attempt to do so, even when set to true, if and only if the host system
// supports FEAT_AFP.
bool SetFIZ {};
// Whether or not FillSpecialRegs should load our SVE predicate temporaries
// with certain canned values that accelerate some operations. Will (obviously)
// not load predicates, even if set to true, on host systems that do not support SVE.
bool SetPredRegs {};
};
void FillSpecialRegs(ARMEmitter::Register TmpReg, ARMEmitter::Register TmpReg2, const FillSpecialRegsOptions& Options);
// Correlate an ARM register back to an x86 register index.
// Returning REG_INVALID if there was no mapping.
@@ -308,8 +322,6 @@ protected:
FEX_CONFIG_OPT(Disassemble, DISASSEMBLE);
#endif
FEX_CONFIG_OPT(EnableCodeCaching, ENABLECODECACHINGWIP);
};
} // namespace FEXCore::CPU
+10 -109
View File
@@ -11,17 +11,8 @@
#include <cstdint>
#ifndef _WIN32
#include <sys/prctl.h>
#endif
namespace FEXCore {
namespace CPU {
static constexpr size_t INITIAL_CODE_SIZE = 1024 * 1024 * 16;
// We don't want to move above 128MB atm because that means we will have to encode longer jumps
static constexpr size_t MAX_CODE_SIZE = 1024 * 1024 * 128;
constexpr static uint64_t NamedVectorConstants[FEXCore::IR::NamedVectorConstant::NAMED_VECTOR_CONST_POOL_MAX][2] = {
{0x0003'0002'0001'0000ULL, 0x0007'0006'0005'0004ULL}, // NAMED_VECTOR_INCREMENTAL_U16_INDEX
{0x000B'000A'0009'0008ULL, 0x000F'000E'000D'000CULL}, // NAMED_VECTOR_INCREMENTAL_U16_INDEX_UPPER
@@ -275,9 +266,9 @@ namespace CPU {
return TotalLUT;
}()};
CPUBackend::CPUBackend(CodeBufferManager& CodeBuffers, FEXCore::Core::InternalThreadState* ThreadState)
CPUBackend::CPUBackend(SharedCodeBufferManager& SharedCodeBuffers, FEXCore::Core::InternalThreadState* ThreadState)
: ThreadState(ThreadState)
, CodeBuffers(CodeBuffers) {
, SharedCodeBuffers(SharedCodeBuffers) {
auto& Ptrs = ThreadState->CurrentFrame->Pointers;
@@ -316,11 +307,11 @@ namespace CPU {
CPUBackend::~CPUBackend() = default;
auto CPUBackend::GetEmptyCodeBuffer() -> CodeBuffer* {
auto CPUBackend::AcquireNewSharedCodeBuffer() -> CodeBuffer* {
auto PrevCodeBuffer = CurrentCodeBuffer;
// Resize the code buffer and reallocate our code size
CurrentCodeBuffer = CodeBuffers.StartLargerCodeBuffer();
CurrentCodeBuffer = SharedCodeBuffers.StartLargerCodeBuffer();
RegisterForSignalHandler(std::move(PrevCodeBuffer));
return CurrentCodeBuffer.get();
@@ -338,7 +329,7 @@ namespace CPU {
}
fextl::shared_ptr<CodeBuffer> CPUBackend::CheckCodeBufferUpdate() {
auto NewCodeBuffer = CodeBuffers.GetLatest();
auto NewCodeBuffer = SharedCodeBuffers.GetLatest();
if (CurrentCodeBuffer != NewCodeBuffer) {
RegisterForSignalHandler(CurrentCodeBuffer);
return std::exchange(CurrentCodeBuffer, NewCodeBuffer);
@@ -346,107 +337,17 @@ namespace CPU {
return nullptr;
}
GuestToHostMap& GetLookupCache(const CodeBuffer& Buffer) {
return *Buffer.LookupCache;
}
CodeBuffer::CodeBuffer(size_t Size)
: AllocatedSize(Size) {
Ptr = static_cast<uint8_t*>(FEXCore::Allocator::VirtualAlloc(Size, true));
LOGMAN_THROW_A_FMT(!!Ptr, "Couldn't allocate code buffer");
// Protect the last page of the allocated buffer to trigger SIGSEGV on write access
uintptr_t LastPageAddr = AlignDown(reinterpret_cast<uintptr_t>(Ptr) + Size - 1, FEXCore::Utils::FEX_PAGE_SIZE);
if (!FEXCore::Allocator::VirtualProtect(reinterpret_cast<void*>(LastPageAddr), FEXCore::Utils::FEX_PAGE_SIZE,
FEXCore::Allocator::ProtectOptions::None)) {
LogMan::Msg::EFmt("Failed to mprotect last page of code buffer.");
}
FEXCore::Allocator::VirtualName("FEXMemJIT", reinterpret_cast<void*>(Ptr), Size);
// Huge-pages reduce the amount of iTLB misses dramatically when it works.
FEXCore::Allocator::VirtualTHPControl(reinterpret_cast<void*>(Ptr), Size, FEXCore::Allocator::THPControl::Enable);
LookupCache = fextl::make_unique<GuestToHostMap>();
}
CodeBuffer::~CodeBuffer() {
FEXCore::Allocator::VirtualFree(Ptr, AllocatedSize);
}
auto CodeBufferManager::AllocateNew(size_t Size) -> fextl::shared_ptr<CodeBuffer> {
#ifndef _WIN32
// MDWE (Memory-Deny-Write-Execute) is a new Linux 6.3 feature.
// It's equivalent to systemd's `MemoryDenyWriteExecute` but implemented entirely in the kernel.
//
// MDWE prevents applications from creating RWX memory mappings.
// This prevents FEX from doing anything JIT related, as FEX uses RWX for JIT memory mappings.
//
// A potential workaround to make FEX work with MDWE is to call mprotect every time we need to write or modify code.
// Alternatively, FEX could use a memory mirror where one half is mapped as RW and the other is RX.
//
// Once MDWE is enabled with the prctl, the feature is sealed and it can /NOT/ be turned off.
//
// Status of MDWE is queried through prctl using `PR_GET_MDWE`:
// -1: The kernel doesn't support MDWE
// 0: MDWE is supported but disabled
// >0: MDWE is enabled, hence prohibiting RWX mappings
#ifndef PR_GET_MDWE
#define PR_GET_MDWE 66
#endif
int MDWE = ::prctl(PR_GET_MDWE, 0, 0, 0, 0);
if (MDWE != -1 && MDWE != 0) {
LogMan::Msg::EFmt("MDWE was set to 0x{:x} which means FEX can't allocate executable memory", MDWE);
}
#endif
auto Buffer = fextl::make_shared<CodeBuffer>(Size);
Latest = Buffer;
LatestOffset = 0;
OnCodeBufferAllocated(Buffer);
return Buffer;
}
fextl::shared_ptr<CodeBuffer> CodeBufferManager::GetLatest() {
if (!Latest) {
if (FEXCore::Config::Get_ENABLECODECACHINGWIP()) {
// Start with a larger code buffer to avoid resizes that would discard
// code loaded from caches
AllocateNew(MAX_CODE_SIZE);
} else {
AllocateNew(INITIAL_CODE_SIZE);
}
}
return Latest;
}
fextl::shared_ptr<CodeBuffer> CodeBufferManager::StartLargerCodeBuffer() {
if (!Latest) {
// Allocate initial CodeBuffer and return it
return GetLatest();
}
auto NewCodeBufferSize = GetLatest()->AllocatedSize;
NewCodeBufferSize = std::min<size_t>(NewCodeBufferSize * 2, MAX_CODE_SIZE);
return AllocateNew(NewCodeBufferSize);
}
bool CPUBackend::IsAddressInCodeBuffer(uintptr_t Address) const {
auto CheckCodeBuffer = [](CodeBuffer& Buffer, uintptr_t Address) {
// The last page of the code buffer is protected, so we need to exclude it from the valid range
// when checking if the address is in the code buffer.
uintptr_t LastPageAddr = AlignDown(reinterpret_cast<uintptr_t>(Buffer.Ptr) + Buffer.AllocatedSize - 1, FEXCore::Utils::FEX_PAGE_SIZE);
return (Address >= reinterpret_cast<uintptr_t>(Buffer.Ptr) && Address < LastPageAddr);
const auto CheckCodeBuffer = [](const CodeBuffer& Buffer, uintptr_t Address) {
const auto BufferPtr = reinterpret_cast<uintptr_t>(Buffer.GetBufferBase());
const uintptr_t LastPageAddr = BufferPtr + Buffer.UsableSize();
return (Address >= BufferPtr && Address < LastPageAddr);
};
if (CheckCodeBuffer(*CurrentCodeBuffer, Address)) {
return true;
}
for (auto& Buffer : SignalHandlerCodeBuffers) {
for (const auto& Buffer : SignalHandlerCodeBuffers) {
if (CheckCodeBuffer(*Buffer, Address)) {
return true;
}
+13 -56
View File
@@ -8,6 +8,8 @@ $end_info$
#pragma once
#include "Interface/Core/SharedCodeBufferManager.h"
#include <FEXCore/Utils/CompilerDefs.h>
#include <FEXCore/Utils/SignalScopeGuards.h>
#include <FEXCore/fextl/memory.h>
@@ -16,6 +18,7 @@ $end_info$
#include <FEXCore/fextl/map.h>
#include <cstdint>
#include <span>
namespace FEXCore::CPU {
union Relocation;
@@ -41,63 +44,10 @@ namespace CodeSerialize {
struct GuestToHostMap;
namespace CPU {
struct CodeBuffer {
uint8_t* Ptr;
size_t AllocatedSize; // including guard page; see UsableSize()
fextl::unique_ptr<GuestToHostMap> LookupCache;
CodeBuffer(size_t Size);
CodeBuffer(const CodeBuffer&) = delete;
CodeBuffer& operator=(const CodeBuffer&) = delete;
CodeBuffer(CodeBuffer&& oth) = delete;
CodeBuffer& operator=(CodeBuffer&&) = delete;
~CodeBuffer();
/// Returns the number of bytes available for storing code
size_t UsableSize() const {
return AllocatedSize - FEXCore::Utils::FEX_PAGE_SIZE;
}
};
/**
* A manager that coordinates access to the CodeBuffer used for compiling new code across threads.
*
* The CodeBuffer is managed as a partially persistent data structure:
* - Exactly one CodeBuffer is now designated as "active", which means data can be appended to it
* - Lossy modifications to the active CodeBuffer will not invalidate any data in use by other threads (which is what enables save CodeBuffer sharing across threads)
* - Instead, such lossy modifications trigger a new "version" of the data in the modifying thread. Old versions of the CodeBuffer persist as read-only data for use by the other threads.
* - The other threads can update their version of the CodeBuffer. This will decrease the reference count and eventually trigger deallocation of the old version
*/
class CodeBufferManager {
public:
// Get the CodeBuffer that was most recently allocated.
// This is the only CodeBuffer that data may be written to.
fextl::shared_ptr<CodeBuffer> GetLatest();
// Allocate a new CodeBuffer with geometric growth up to an internal maximum.
// Subsequent calls to GetLatest will point to the returned buffer.
fextl::shared_ptr<CodeBuffer> StartLargerCodeBuffer();
// Write offset into the latest CodeBuffer
std::size_t LatestOffset {};
// Protects writes to the latest CodeBuffer and changes to LatestOffset
FEXCore::ForkableUniqueMutex CodeBufferWriteMutex;
virtual void OnCodeBufferAllocated(const std::shared_ptr<CodeBuffer>&) {};
private:
fextl::shared_ptr<CodeBuffer> Latest;
fextl::shared_ptr<CodeBuffer> AllocateNew(size_t Size);
};
class CPUBackend {
public:
CPUBackend(CodeBufferManager&, FEXCore::Core::InternalThreadState*);
CPUBackend(SharedCodeBufferManager&, FEXCore::Core::InternalThreadState*);
virtual ~CPUBackend();
@@ -107,6 +57,8 @@ namespace CPU {
fextl::map<uint64_t, uint8_t*> EntryPoints;
// The total size of the codeblock from [BlockBegin, BlockBegin+Size).
size_t Size;
// Offset of BlockBegin from the start of the CodeBuffer it lives in
uint64_t HostCodeOffset;
};
// Header that can live at the start of a JIT block.
@@ -166,6 +118,10 @@ namespace CPU {
virtual CompiledCode CompileCode(uint64_t Entry, uint64_t Size, bool SingleInst, const FEXCore::IR::IRListView* IR,
FEXCore::Core::DebugData* DebugData, bool CheckTF) = 0;
virtual CompiledCode LoadCachedCode(std::span<const uint8_t> HostBytes) {
return {};
}
virtual fextl::vector<FEXCore::CPU::Relocation> TakeRelocations(uint64_t GuestBaseAddress) = 0;
virtual void ClearCache() {}
@@ -189,8 +145,9 @@ namespace CPU {
FEXCore::Core::InternalThreadState* ThreadState;
// Acquires a new shared code buffer, setting `CurrentCodeBuffer` and returning a pointer to it.
[[nodiscard]]
CodeBuffer* GetEmptyCodeBuffer();
CodeBuffer* AcquireNewSharedCodeBuffer();
// This is the code buffer containing the main code under execution by this thread.
// CheckCodeBufferUpdate must be used before compiling new code.
@@ -199,7 +156,7 @@ namespace CPU {
// Old CodeBuffer generations required to be valid until returning from signal handlers
fextl::vector<fextl::shared_ptr<CodeBuffer>> SignalHandlerCodeBuffers;
CodeBufferManager& CodeBuffers;
SharedCodeBufferManager& SharedCodeBuffers;
private:
void RegisterForSignalHandler(fextl::shared_ptr<CodeBuffer>);
+15 -9
View File
@@ -100,7 +100,7 @@ namespace ProductNames {
#endif
} // namespace ProductNames
uint32_t GetCPUID_Syscall() {
static uint32_t GetCPUID_Syscall() {
uint32_t CPU {};
FHU::Syscalls::getcpu(&CPU, nullptr);
return CPU;
@@ -148,7 +148,7 @@ uint64_t GetCycleCounterFrequency() {
return Result;
}
uint32_t GetCPUID_TPIDRRO() {
static uint32_t GetCPUID_TPIDRRO() {
uint64_t Result {};
__asm("mrs %[Res], TPIDRRO_EL0" : [Res] "=r"(Result));
return Result;
@@ -316,9 +316,8 @@ void CPUIDEmu::SetupHostHybridFlag() {
// Walk our list of CPUMIDRs to find the most little core
for (size_t j = LowestMIDRIdx; j < CPUMIDRs.size(); ++j) {
auto& MIDROption = CPUMIDRs[i];
const auto& MIDROption = CPUMIDRs[j];
if ((MIDROption.Implementer == Implementer && MIDROption.Part == Part) || (MIDROption.Implementer == 0 && MIDROption.Part == 0)) {
LowestMIDRIdx = j;
LowestMIDR = MIDR;
break;
@@ -494,8 +493,8 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_01h(uint32_t Leaf) const {
Res.edx = (1 << 0) | // FPU
(1 << 1) | // Virtual 8086 mode enhancements
(0 << 2) | // Debugging extensions
(0 << 3) | // Page size extension
(1 << 2) | // Debugging extensions
(1 << 3) | // Page size extension
(1 << 4) | // RDTSC supported
(1 << 5) | // MSR supported
(1 << 6) | // PAE
@@ -650,6 +649,13 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_06h(uint32_t Leaf) const {
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_07h(uint32_t Leaf) const {
FEXCore::CPUID::FunctionResults Res {};
if (Leaf == 0) {
#ifndef _WIN32
constexpr uint32_t SUPPORTS_RDPID = 1;
#else
// RDPID under WIN32 is only supported if CPUIndex is available in TPIDRRO.
const uint32_t SUPPORTS_RDPID = SupportsCPUIndexInTPIDRRO;
#endif
// Disable Enhanced REP MOVS when TSO is enabled.
// vcruntime140 memmove will use `rep movsb` in this case which completely destroys perf in Hades(appId 1145360)
// This is due to LRCPC performance on Cortex being abysmal.
@@ -715,7 +721,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_07h(uint32_t Leaf) const {
(0 << 19) | // MPX MAWAU
(0 << 20) | // MPX MAWAU
(0 << 21) | // MPX MAWAU
(1 << 22) | // RDPID Read Processor ID
(SUPPORTS_RDPID << 22) | // RDPID Read Processor ID
(0 << 23) | // AES Key Locker
(1 << 24) | // bus-lock-detect
(0 << 25) | // CLDEMOTE
@@ -1091,7 +1097,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0001h(uint32_t Leaf) con
(1 << 23) | // MMX
(1 << 24) | // FXSAVE/FXRSTOR
(1 << 25) | // FXSAVE/FXRSTOR Optimizations
(0 << 26) | // 1 gigabit pages
(1 << 26) | // 1 gigabit pages
(SUPPORTS_RDTSCP << 27) | // RDTSCP
(0 << 28) | // Reserved
(1 << 29) | // Long Mode
@@ -1341,7 +1347,7 @@ FEXCore::CPUID::XCRResults CPUIDEmu::XCRFunction_0h() const {
CPUIDEmu::CPUIDEmu(const FEXCore::Context::ContextImpl* ctx)
: CTX {ctx}
, SupportsCPUIndexInTPIDRRO {CTX->HostFeatures.SupportsCPUIndexInTPIDRRO}
, SupportsCPUIndexInTPIDRRO {CTX->HostFeatures.SupportsCPUIndexInTPIDRRO != 0}
, GetCPUID {GetCPUID_Syscall} {
Cores = CTX->HostFeatures.CPUMIDRs.size();
+166 -48
View File
@@ -50,6 +50,10 @@ MappedCodeCacheFile::~MappedCodeCacheFile() {
if (!CodeBuffer.empty()) {
FEXCore::Allocator::munmap(CodeBuffer.data(), CodeBuffer.size_bytes());
}
#elif defined(_M_ARM64EC)
if (!CodeBuffer.empty()) {
FEXCore::Allocator::VirtualFree(CodeBuffer.data(), CodeBuffer.size_bytes());
}
#endif
}
@@ -107,16 +111,22 @@ fextl::map<CodeMapFileId, CodeMap::ParsedContents> CodeMap::ParseCodeMap(std::if
break;
}
Ret[Info.ExternalFileId].Filename = std::move(Filename);
} else if (Entry.FileId == SetExecutableFileId {}.Marker.FileId && Entry.BlockOffset == SetExecutableFileId {}.Marker.BlockOffset) {
} else if ((Entry.FileId == SetExecutableFileId::Marker32.FileId && Entry.BlockOffset == SetExecutableFileId::Marker32.BlockOffset) ||
(Entry.FileId == SetExecutableFileId::Marker64.FileId && Entry.BlockOffset == SetExecutableFileId::Marker64.BlockOffset)) {
CodeMapFileId ExecutableFileId;
File.read(reinterpret_cast<char*>(&ExecutableFileId), sizeof(ExecutableFileId));
if (!File) {
break;
}
Ret[ExecutableFileId].IsExecutable = true;
Ret[ExecutableFileId].ExecutableBitness =
(Entry.FileId == SetExecutableFileId::Marker32.FileId && Entry.BlockOffset == SetExecutableFileId::Marker32.BlockOffset) ? 32 : 64;
} else {
if (!Ret.contains(Entry.FileId)) {
LogMan::Msg::EFmt("Code map referenced unknown file id {:016x}", Entry.FileId);
if (Entry.FileId == 0xffff'ffff'ffff'ffff) {
ERROR_AND_DIE_FMT("Malformed code map");
} else {
LogMan::Msg::EFmt("Code map referenced unknown file id {:016x}", Entry.FileId);
}
} else {
Ret[Entry.FileId].Blocks.insert(Entry.BlockOffset);
}
@@ -222,8 +232,8 @@ void CodeMapWriter::AppendLibraryLoad(const FEXCore::ExecutableFileInfo& FileInf
AppendData(std::as_bytes(std::span {Data, TotalSize}));
}
void CodeMapWriter::AppendSetMainExecutable(const FEXCore::ExecutableFileInfo& FileInfo) {
CodeMap::SetExecutableFileId Data {.ExecutableFileId = FileInfo.FileId};
void CodeMapWriter::AppendSetMainExecutable(const FEXCore::ExecutableFileInfo& FileInfo, bool Is64Bit) {
CodeMap::SetExecutableFileId Data {Is64Bit ? CodeMap::SetExecutableFileId::Marker64 : CodeMap::SetExecutableFileId::Marker32, FileInfo.FileId};
AppendData(std::span {reinterpret_cast<const std::byte*>(&Data), sizeof(Data)});
}
@@ -304,7 +314,7 @@ bool CodeCache::SaveData(Core::InternalThreadState& Thread, int fd, const Execut
std::ranges::copy(GIT_HASH, header.FEXVersion);
header.NumBlocks = LookupCache.BlockList.size();
header.NumCodePages = LookupCache.CodePages.size();
header.CodeBufferSize = FEXCore::AlignUp(CTX.LatestOffset, Utils::FEX_PAGE_SIZE);
header.CodeBufferSize = FEXCore::AlignUp(CodeBuffer->AllocatedSpaceUsed(), Utils::FEX_PAGE_SIZE);
header.NumRelocations = Relocations.size();
header.SerializedBaseAddress = SerializedBaseAddress;
::write(fd, &header, sizeof(header));
@@ -327,7 +337,7 @@ bool CodeCache::SaveData(Core::InternalThreadState& Thread, int fd, const Execut
Guest -= SourceBinary.FileStartVA;
::write(fd, &Guest, sizeof(Guest));
uint64_t HostCode = Host->HostCode - reinterpret_cast<uintptr_t>(CodeBuffer->Ptr);
uint64_t HostCode = Host->HostCode - reinterpret_cast<uintptr_t>(CodeBuffer->GetBufferBase());
::write(fd, &HostCode, sizeof(HostCode));
uint64_t NumCodePages = Host->CodePages.size();
::write(fd, &NumCodePages, sizeof(NumCodePages));
@@ -351,8 +361,9 @@ bool CodeCache::SaveData(Core::InternalThreadState& Thread, int fd, const Execut
}
// Dump the host code (relocated for position-independent serialization)
std::span CodeBufferData(reinterpret_cast<std::byte*>(CodeBuffer->Ptr), reinterpret_cast<std::byte*>(CodeBuffer->Ptr) + CTX.LatestOffset);
if (!ApplyCodeRelocations(SerializedBaseAddress, CodeBufferData, Relocations, 0, true)) {
std::span CodeBufferData(reinterpret_cast<std::byte*>(CodeBuffer->GetBufferBase()),
reinterpret_cast<std::byte*>(CodeBuffer->GetBufferBase()) + CodeBuffer->AllocatedSpaceUsed());
if (!ApplyCodeRelocations(SerializedBaseAddress, CodeBufferData, Relocations, true)) {
LOGMAN_THROW_A_FMT(false, "Failed to apply code relocations");
return false;
}
@@ -395,7 +406,7 @@ void CodeCache::Validate(const ExecutableFileSectionInfo& Section, fextl::set<ui
ERROR_AND_DIE_FMT("Failed to create cache load validation context");
}
ValidationThread.reset(ValidationCTX->CreateThread(0, 0, nullptr));
ValidationThread.reset(ValidationCTX->CreateThread(nullptr));
auto Frame = ValidationThread->CurrentFrame;
Frame->State.segment_arrays[FEXCore::Core::CPUState::SEGMENT_ARRAY_INDEX_GDT] = &ValidationGDT[0];
@@ -416,11 +427,12 @@ void CodeCache::Validate(const ExecutableFileSectionInfo& Section, fextl::set<ui
while (CachedCode.size_bytes() > NewCodeBuffer->UsableSize()) {
ValidationCTX->ClearCodeCache(ValidationThread.get());
NewCodeBuffer = ValidationCTX->GetLatest();
LogMan::Msg::IFmt("Increased cache validation code buffer size to {} MiB", NewCodeBuffer->AllocatedSize / 1024 / 1024);
LogMan::Msg::IFmt("Increased cache validation code buffer size to {} MiB", NewCodeBuffer->TotalAllocationSize() / 1024 / 1024);
}
std::span<std::byte> CodeBufferRangeRef =
std::as_writable_bytes(std::span {NewCodeBuffer->Ptr, NewCodeBuffer->Ptr + NewCodeBuffer->UsableSize()}).subspan(0, CachedCode.size_bytes());
std::as_writable_bytes(std::span {NewCodeBuffer->GetBufferBase(), NewCodeBuffer->GetBufferBase() + NewCodeBuffer->UsableSize()})
.subspan(0, CachedCode.size_bytes());
while (!GuestBlocks.empty()) {
auto [CompiledBlocks, _, _2, _3, _4] = ValidationCTX->CompileCode(ValidationThread.get(), *GuestBlocks.begin(), 0 /* TODO: Set MaxInst? */);
@@ -434,12 +446,12 @@ void CodeCache::Validate(const ExecutableFileSectionInfo& Section, fextl::set<ui
NewRelocations.erase(std::remove_if(NewRelocations.begin(), NewRelocations.end(), [](const CPU::Relocation& Reloc) {
return Reloc.Header.Type != CPU::RelocationTypes::RELOC_NAMED_SYMBOL_LITERAL && Reloc.Header.Type != CPU::RelocationTypes::RELOC_NAMED_THUNK_MOVE;
}));
(void)ApplyCodeRelocations(Section.FileStartVA, CodeBufferRangeRef, NewRelocations, 0, false);
(void)ApplyCodeRelocations(Section.FileStartVA, CodeBufferRangeRef, NewRelocations, false);
if (ValidationCTX->LatestOffset <= CodeBufferRangeRef.size()) {
if (NewCodeBuffer->AllocatedSpaceUsed() <= CodeBufferRangeRef.size()) {
// Reference compilation produced fewer bytes than our cache, so validation is going to fail.
// Make sure we don't output any garbage bytes though.
CodeBufferRangeRef = CodeBufferRangeRef.subspan(0, ValidationCTX->LatestOffset);
CodeBufferRangeRef = CodeBufferRangeRef.subspan(0, NewCodeBuffer->AllocatedSpaceUsed());
}
auto [Mismatch, _] = std::mismatch(CodeBufferRangeRef.begin(), CodeBufferRangeRef.end(), CachedCode.begin());
@@ -466,7 +478,7 @@ void CodeCache::Validate(const ExecutableFileSectionInfo& Section, fextl::set<ui
if (tail->RIP >= Section.BeginVA && tail->RIP < Section.EndVA) {
auto [IRView, TotalInstructions, TotalInstructionsLength, StartAddr, Length, _] =
ValidationCTX->GenerateIR(ValidationThread.get(), tail->RIP, false, FEXCore::Config::Get_MAXINST());
fextl::stringstream ss;
fextl::ostringstream ss;
FEXCore::IR::Dump(&ss, &*IRView);
LogMan::Msg::EFmt("IR:\n{}", ss.str());
} else {
@@ -490,47 +502,141 @@ void CodeCache::Validate(const ExecutableFileSectionInfo& Section, fextl::set<ui
// Reset Context state for next validation
ValidationThread->LookupCache->ClearCache(ValidationThread->LookupCache->AcquireWriteLock());
ValidationCTX->LatestOffset = 0;
NewCodeBuffer->Reset();
LogMan::Msg::IFmt(" successfully validated cache");
}
bool CodeCache::ApplyCodeRelocations(uint64_t GuestEntry, std::span<std::byte> Code,
std::span<const FEXCore::CPU::Relocation> EntryRelocations, uint32_t RelocationOffset, bool ForStorage) {
CPU::Arm64Emitter Emitter(&CTX, Code.data(), Code.size_bytes());
for (size_t j = 0; j < EntryRelocations.size(); ++j) {
const FEXCore::CPU::Relocation& Reloc = EntryRelocations[j];
LOGMAN_THROW_A_FMT(Reloc.Header.Offset >= RelocationOffset, "Invalid relocation offset");
LOGMAN_THROW_A_FMT(Reloc.Header.Offset - RelocationOffset < Code.size_bytes(), "Invalid relocation offset");
Emitter.SetCursorOffset(Reloc.Header.Offset - RelocationOffset);
static inline void ApplySymbolLiteralRelocation(ContextImpl& CTX, const CPU::RelocNamedSymbolLiteral::NamedSymbol Symbol,
uint64_t GuestEntry, CPU::Arm64Emitter& Emitter, bool ForStorage) {
// Generate a literal so we can place it
uint64_t Pointer = ForStorage ? 0 : GetNamedSymbolLiteral(CTX, Symbol);
Emitter.dc64(Pointer);
}
switch (Reloc.Header.Type) {
static inline bool
ApplyThunkMoveRelocation(ContextImpl& CTX, const IR::SHA256Sum* Symbol, uint32_t RegisterIndex, CPU::Arm64Emitter& Emitter, bool ForStorage) {
uint64_t Pointer = ForStorage ? 0 : reinterpret_cast<uint64_t>(CTX.ThunkHandler->LookupThunk(*Symbol));
if (Pointer == ~0ULL) {
return false;
}
// TODO: Pointers are required to fit within 48-bit VA space.
// But forcing 6-byte broke relocations.
Emitter.LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Register(RegisterIndex), Pointer, CPU::Arm64Emitter::PadType::DOPAD);
return true;
}
static inline void ApplyRIPLiteralRelocation(ContextImpl& CTX, uint64_t GuestRIP, uint64_t GuestEntry, CPU::Arm64Emitter& Emitter) {
Emitter.dc64(GuestEntry + GuestRIP);
}
static inline void
ApplyRIPMoveRelocation(ContextImpl& CTX, uint64_t GuestRIP, uint8_t RegisterIndex, uint64_t GuestEntry, CPU::Arm64Emitter& Emitter) {
uint64_t Pointer = GuestRIP + GuestEntry;
// TODO: Pointers are required to fit within 48-bit VA space.
// But forcing 6-byte broke relocations.
Emitter.LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Register(RegisterIndex), Pointer, CPU::Arm64Emitter::PadType::DOPAD);
}
static inline void ApplyPatchableDataRelocation(uint64_t SiteAddress, uint8_t ValueSize, uint8_t RegisterIndex, CPU::Arm64Emitter& Emitter) {
uint64_t Value = 0;
memcpy(&Value, reinterpret_cast<const void*>(SiteAddress), ValueSize);
Emitter.LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Register(RegisterIndex), Value, CPU::Arm64Emitter::PadType::DOPAD);
}
static inline int64_t ReadLiveGuestDisplacement(uint64_t SiteAddress, uint8_t ValueSize) {
uint64_t Raw = 0;
memcpy(&Raw, reinterpret_cast<const void*>(SiteAddress), ValueSize);
// manual sign-extension from guest live bytes
// 1/2 sizes not permitted in DetectDataMasks currently
if (ValueSize == 4) {
return (int32_t)Raw;
} else {
return (int64_t)Raw;
}
}
static inline void ApplyPatchableRIPLiteralRelocation(uint64_t SiteAddress, uint8_t ValueSize, CPU::Arm64Emitter& Emitter) {
Emitter.dc64(SiteAddress + ValueSize + ReadLiveGuestDisplacement(SiteAddress, ValueSize));
}
static inline void ApplyPatchableRIPMoveRelocation(uint64_t SiteAddress, uint8_t ValueSize, uint8_t RegisterIndex, CPU::Arm64Emitter& Emitter) {
const uint64_t Target = SiteAddress + ValueSize + ReadLiveGuestDisplacement(SiteAddress, ValueSize);
Emitter.LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Register(RegisterIndex), Target, CPU::Arm64Emitter::PadType::DOPAD);
}
bool CodeCache::ApplyPackedCodeRelocations(uint64_t GuestEntry, std::span<std::byte> Code,
std::span<const DiskCache::BlobSmallRelocation> SmallRelocs,
std::span<const DiskCache::BlobThunkRelocation> ThunkRelocs) {
CPU::Arm64Emitter Emitter(&CTX, Code.data(), Code.size_bytes());
for (auto& Reloc : SmallRelocs) {
LOGMAN_THROW_A_FMT(Reloc.Offset < Code.size_bytes(), "Invalid relocation offset");
Emitter.SetCursorOffset(Reloc.Offset);
switch ((CPU::RelocationTypes)Reloc.Type) {
case FEXCore::CPU::RelocationTypes::RELOC_NAMED_SYMBOL_LITERAL: {
// Generate a literal so we can place it
uint64_t Pointer = ForStorage ? 0 : GetNamedSymbolLiteral(CTX, Reloc.NamedSymbolLiteral.Symbol);
Emitter.dc64(Pointer);
break;
}
case FEXCore::CPU::RelocationTypes::RELOC_NAMED_THUNK_MOVE: {
uint64_t Pointer = ForStorage ? 0 : reinterpret_cast<uint64_t>(CTX.ThunkHandler->LookupThunk(Reloc.NamedThunkMove.Symbol));
if (Pointer == ~0ULL) {
return false;
}
// TODO: Pointers are required to fit within 48-bit VA space.
// But forcing 6-byte broke relocations.
Emitter.LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Register(Reloc.NamedThunkMove.RegisterIndex), Pointer,
CPU::Arm64Emitter::PadType::DOPAD);
ApplySymbolLiteralRelocation(CTX, (CPU::RelocNamedSymbolLiteral::NamedSymbol)Reloc.Named.Symbol, GuestEntry, Emitter, false);
break;
}
case FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_LITERAL: {
Emitter.dc64(GuestEntry + Reloc.GuestRIP.GuestRIP);
ApplyRIPLiteralRelocation(CTX, Reloc.RIPLiteral.GuestRIP, GuestEntry, Emitter);
break;
}
case FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_MOVE: {
uint64_t Pointer = Reloc.GuestRIP.GuestRIP + GuestEntry;
// TODO: Pointers are required to fit within 48-bit VA space.
// But forcing 6-byte broke relocations.
Emitter.LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Register(Reloc.GuestRIP.RegisterIndex), Pointer, CPU::Arm64Emitter::PadType::DOPAD);
ApplyRIPMoveRelocation(CTX, Reloc.RIPMove.GuestRIP, Reloc.RIPMove.RegisterIndex, GuestEntry, Emitter);
break;
}
case FEXCore::CPU::RelocationTypes::RELOC_GUEST_PATCHABLE_DATA_MOVE: {
ApplyPatchableDataRelocation(GuestEntry + Reloc.PatchableData.SiteOffset, Reloc.PatchableData.ValueSize,
Reloc.PatchableData.RegisterIndex, Emitter);
break;
}
case FEXCore::CPU::RelocationTypes::RELOC_GUEST_PATCHABLE_RIP_LITERAL: {
ApplyPatchableRIPLiteralRelocation(GuestEntry + Reloc.PatchableData.SiteOffset, Reloc.PatchableData.ValueSize, Emitter);
break;
}
case FEXCore::CPU::RelocationTypes::RELOC_GUEST_PATCHABLE_RIP_MOVE: {
ApplyPatchableRIPMoveRelocation(GuestEntry + Reloc.PatchableData.SiteOffset, Reloc.PatchableData.ValueSize,
Reloc.PatchableData.RegisterIndex, Emitter);
break;
}
default: ERROR_AND_DIE_FMT("Unknown packed relocation type {}", ToUnderlying((CPU::RelocationTypes)Reloc.Type));
}
}
for (auto& Reloc : ThunkRelocs) {
LOGMAN_THROW_A_FMT(Reloc.Offset < Code.size_bytes(), "Invalid relocation offset");
Emitter.SetCursorOffset(Reloc.Offset);
if (!ApplyThunkMoveRelocation(CTX, (const IR::SHA256Sum*)Reloc.SymbolHash, Reloc.RegisterIndex, Emitter, false)) {
return false;
}
}
return true;
}
bool CodeCache::ApplyCodeRelocations(uint64_t GuestEntry, std::span<std::byte> Code,
std::span<const FEXCore::CPU::Relocation> EntryRelocations, bool ForStorage) {
CPU::Arm64Emitter Emitter(&CTX, Code.data(), Code.size_bytes());
for (size_t j = 0; j < EntryRelocations.size(); ++j) {
const FEXCore::CPU::Relocation& Reloc = EntryRelocations[j];
LOGMAN_THROW_A_FMT(Reloc.Header.Offset < Code.size_bytes(), "Invalid relocation offset");
Emitter.SetCursorOffset(Reloc.Header.Offset);
switch (Reloc.Header.Type) {
case FEXCore::CPU::RelocationTypes::RELOC_NAMED_SYMBOL_LITERAL: {
ApplySymbolLiteralRelocation(CTX, Reloc.NamedSymbolLiteral.Symbol, GuestEntry, Emitter, ForStorage);
break;
}
case FEXCore::CPU::RelocationTypes::RELOC_NAMED_THUNK_MOVE: {
if (!ApplyThunkMoveRelocation(CTX, &Reloc.NamedThunkMove.Symbol, Reloc.NamedThunkMove.RegisterIndex, Emitter, ForStorage)) {
return false;
}
break;
}
case FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_LITERAL: {
ApplyRIPLiteralRelocation(CTX, Reloc.GuestRIP.GuestRIP, GuestEntry, Emitter);
break;
}
case FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_MOVE: {
ApplyRIPMoveRelocation(CTX, Reloc.GuestRIP.GuestRIP, Reloc.GuestRIP.RegisterIndex, GuestEntry, Emitter);
break;
}
@@ -600,7 +706,16 @@ CodeCache::LoadCache(std::span<std::byte> CacheFile, const ExecutableFileInfo& F
return nullptr;
}
auto CodeBuffer = std::span {static_cast<std::byte*>(CodeBufferAllocation), header.CodeBufferSize};
#else
#elif defined(_M_ARM64EC)
// TODO: Implement lazy mapping on Windows
// NOTE: The executed code must have MEM_EXTENDED_PARAMETER_EC_CODE set, so we can't operate on the mapped cache file directly
void* CodeBufferAllocation = Allocator::VirtualAlloc(header.CodeBufferSize, true);
if (!CodeBufferAllocation) {
LogMan::Msg::EFmt("Failed to allocate code cache memory");
return nullptr;
}
auto CodeBuffer = std::span {reinterpret_cast<std::byte*>(CodeBufferAllocation), header.CodeBufferSize};
#else // WoW64
// TODO: Implement lazy mapping on Windows
auto CodeBuffer = CodeDataInFile;
#endif
@@ -852,7 +967,7 @@ void CodeCache::FinalizeCodePages(MappedCodeCacheFile& Code, std::span<std::byte
auto StagingSpan = std::span {Staging, Size};
for (size_t i = StartPage; i < EndPage; ++i) {
auto PageRelocations = SpanPageRelocations(Code, i);
(void)ApplyCodeRelocations(Code.GuestBase, StagingSpan, PageRelocations, static_cast<uint32_t>(StartOffset), false);
(void)ApplyCodeRelocations(Code.GuestBase, StagingSpan, PageRelocations, false);
Code.LoadedPages[i] = true;
}
@@ -870,9 +985,12 @@ void CodeCache::FinalizeCodePages(MappedCodeCacheFile& Code, std::span<std::byte
Allocator::VirtualDontNeed(Code.CodeBufferInFile.data() + StartOffset, Size);
#else
// TODO: Implement lazy mapping on Windows
#ifdef _M_ARM64EC
memcpy(Code.CodeBuffer.data() + StartOffset, Code.CodeBufferInFile.data() + StartOffset, Size);
#endif
for (size_t i = StartPage; i < EndPage; ++i) {
auto PageRelocations = SpanPageRelocations(Code, i);
(void)ApplyCodeRelocations(Code.GuestBase, Code.CodeBuffer, PageRelocations, 0, false);
(void)ApplyCodeRelocations(Code.GuestBase, Code.CodeBuffer, PageRelocations, false);
Code.LoadedPages[i] = true;
}
#endif
+141 -47
View File
@@ -76,6 +76,9 @@ $end_info$
#include <unordered_map>
#include <utility>
#include <xxhash.h>
#if defined(ARCHITECTURE_arm64)
#include <arm_acle.h>
#endif
namespace FEXCore::Context {
ContextImpl::ContextImpl(const FEXCore::HostFeatures& Features)
@@ -103,6 +106,8 @@ ContextImpl::ContextImpl(const FEXCore::HostFeatures& Features)
// Track atomic TSO emulation configuration.
UpdateAtomicTSOEmulationConfig();
DiskCache.Init(this);
}
struct GetFrameBlockInfoResult {
@@ -342,6 +347,11 @@ void ContextImpl::SetFlagsFromCompactedEFLAGS(FEXCore::Core::InternalThreadState
}
bool ContextImpl::InitCore() {
if (CodeCache.IsGeneratingCache || FEXCore::Config::Get_ENABLECODECACHINGWIP()) {
// Start with a larger code buffer to avoid resizes that would discard code
StartMaximalCodeBuffer();
}
// Initialize the CPU core signal handlers & DispatcherConfig
Dispatcher = FEXCore::CPU::Dispatcher::Create(this);
@@ -390,7 +400,7 @@ void ContextImpl::InitializeCompiler(FEXCore::Core::InternalThreadState* Thread)
Thread->OpDispatcher->SetMultiblock(Config.Multiblock);
Thread->LookupCache = fextl::make_unique<FEXCore::LookupCache>(this);
Thread->FrontendDecoder = fextl::make_unique<FEXCore::Frontend::Decoder>(Thread);
Thread->PassManager = fextl::make_unique<FEXCore::IR::PassManager>();
Thread->PassManager = fextl::make_unique<FEXCore::IR::PassManager>(this);
Thread->CurrentFrame->State.L1Pointer = Thread->LookupCache->GetL1Pointer();
Thread->CurrentFrame->State.L1Mask = Thread->LookupCache->GetScaledL1PointerMask();
@@ -399,28 +409,20 @@ void ContextImpl::InitializeCompiler(FEXCore::Core::InternalThreadState* Thread)
Dispatcher->InitThreadPointers(Thread);
Thread->PassManager->AddDefaultPasses(this);
Thread->PassManager->AddDefaultValidationPasses();
Thread->PassManager->RegisterSyscallHandler(SyscallHandler);
// Create CPU backend
Thread->PassManager->InsertRegisterAllocationPass(this);
Thread->CPUBackend = FEXCore::CPU::CreateArm64JITCore(this, Thread);
// We finalize *after* the CPU backend is initialized, as the CPU backend will
// provide necessary register information to the register allocation pass.
Thread->PassManager->Finalize();
}
FEXCore::Core::InternalThreadState*
ContextImpl::CreateThread(uint64_t InitialRIP, uint64_t StackPointer, const FEXCore::Core::CPUState* NewThreadState) {
FEXCore::Core::InternalThreadState* ContextImpl::CreateThread(const FEXCore::Core::CPUState* NewThreadState) {
FEXCore::Core::InternalThreadState* Thread = new FEXCore::Core::InternalThreadState {
.CTX = this,
};
FEXCore::Allocator::VirtualName("FEXMem_ThreadState", Thread, sizeof(*Thread));
Thread->CurrentFrame->State.gregs[X86State::REG_RSP] = StackPointer;
Thread->CurrentFrame->State.rip = InitialRIP;
// Copy over the new thread state to the new object
if (NewThreadState) {
memcpy(&Thread->CurrentFrame->State, NewThreadState, sizeof(FEXCore::Core::CPUState));
@@ -480,7 +482,7 @@ void ContextImpl::LockBeforeFork(FEXCore::Core::InternalThreadState* Thread) {
void ContextImpl::OnCodeBufferAllocated(const fextl::shared_ptr<CPU::CodeBuffer>& Buffer) {
if (Config.GlobalJITNaming()) {
Symbols.RegisterJITSpace(Buffer->Ptr, Buffer->AllocatedSize);
Symbols.RegisterJITSpace(Buffer->GetBufferBase(), Buffer->TotalAllocationSize());
}
{
@@ -505,11 +507,11 @@ void ContextImpl::ClearCodeCache(FEXCore::Core::InternalThreadState* Thread, boo
static void IRDumper(FEXCore::Core::InternalThreadState* Thread, IR::IREmitter* IREmitter, uint64_t GuestRIP) {
FEXCore::File::File FD = FEXCore::File::File::GetStdERR();
fextl::stringstream out;
fextl::ostringstream out;
auto NewIR = IREmitter->ViewIR();
FEXCore::IR::Dump(&out, &NewIR);
fextl::fmt::print(FD, "IR-ShouldDump-{} 0x{:x}:\n{}\n@@@@@\n", NewIR.PostRA() ? "post" : "pre", GuestRIP, out.str());
};
}
bool ContextImpl::CheckIfBlockIsCacheable(FEXCore::Core::InternalThreadState& Thread, uint64_t GuestRIP, uint64_t MaxInst) {
return Thread.FrontendDecoder->CheckIfCacheable(Thread, reinterpret_cast<const uint8_t*>(GuestRIP), GuestRIP, MaxInst);
@@ -538,18 +540,14 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
}
if (!HasCustomIR) {
const uint8_t* GuestCode {};
GuestCode = reinterpret_cast<const uint8_t*>(GuestRIP);
const auto* GuestCode = reinterpret_cast<const uint8_t*>(GuestRIP);
bool HadDispatchError {false};
bool HadInvalidInst {false};
Thread->FrontendDecoder->DecodeLoop(GuestCode);
Thread->FrontendDecoder->DecodeInstructionsAtEntry(Thread, GuestCode, GuestRIP, MaxInst);
const auto* BlockInfo = Thread->FrontendDecoder->GetDecodedBlockInfo();
const auto& CodeBlocks = BlockInfo->Blocks;
auto BlockInfo = Thread->FrontendDecoder->GetDecodedBlockInfo();
auto CodeBlocks = &BlockInfo->Blocks;
Thread->OpDispatcher->BeginFunction(GuestRIP, CodeBlocks, BlockInfo->TotalInstructionCount, BlockInfo->Is64BitMode,
Thread->OpDispatcher->BeginFunction(GuestRIP, &CodeBlocks, BlockInfo->TotalInstructionCount, BlockInfo->Is64BitMode,
AreMonoHacksActive() && MonoBackpatcherBlock.load(std::memory_order_relaxed) == GuestRIP);
const auto GPRSize = Thread->OpDispatcher->GetGPROpSize();
@@ -563,11 +561,17 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
}
#endif
for (size_t j = 0; j < CodeBlocks->size(); ++j) {
const FEXCore::Frontend::Decoder::DecodedBlocks& Block = CodeBlocks->at(j);
for (size_t j = 0; j < CodeBlocks.size(); ++j) {
const auto& Block = CodeBlocks[j];
// Dispatch failures and invalid instructions terminate only the decoded
// block that contains them. Other block targets in the same multiblock
// compilation unit are independent entry paths.
bool HadDispatchError {false};
bool HadInvalidInst {false};
#ifdef ZYDIS_DISASSEMBLER
if (FEXCore::Config::Get_X86DISASSEMBLE() && CodeBlocks->size() > 1) {
if (FEXCore::Config::Get_X86DISASSEMBLE() && CodeBlocks.size() > 1) {
LogMan::Msg::IFmt(" Block {} Entry={:#x} NumInsts={}", j, Block.Entry, Block.NumInstructions);
}
#endif
@@ -575,7 +579,7 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
bool BlockInForceTSOValidRange = false;
auto InstForceTSOIt = ForceTSOInstructions.end();
if (ForceTSOValidRanges.Contains({Block.Entry, Block.Entry + Block.Size})) {
if (auto It = ForceTSOInstructions.lower_bound(Block.Entry); *It < Block.Entry + Block.Size) {
if (auto It = ForceTSOInstructions.lower_bound(Block.Entry); It != ForceTSOInstructions.end() && *It < Block.Entry + Block.Size) {
InstForceTSOIt = It;
BlockInForceTSOValidRange = true;
}
@@ -584,18 +588,16 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
// Set the block entry point
Thread->OpDispatcher->SetNewBlockIfChanged(Block.Entry);
uint64_t BlockInstructionsLength {};
// Reset any block-specific state
Thread->OpDispatcher->StartNewBlock();
uint64_t InstsInBlock = Block.NumInstructions;
const uint64_t InstsInBlock = Block.NumInstructions;
if (InstsInBlock == 0) {
// Special case for an empty instruction block.
Thread->OpDispatcher->ExitFunction(Thread->OpDispatcher->_InlineEntrypointOffset(GPRSize, Block.Entry - GuestRIP));
}
uint64_t BlockInstructionsLength {};
for (size_t i = 0; i < InstsInBlock; ++i) {
uint64_t InstAddress = Block.Entry + BlockInstructionsLength;
const FEXCore::X86Tables::X86InstInfo* TableInfo {nullptr};
@@ -639,9 +641,28 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
if (Config.SMCChecks == FEXCore::Config::CONFIG_SMC_FULL || Block.ForceFullSMCDetection) {
auto ExistingCodePtr = reinterpret_cast<uint8_t*>(Block.Entry + BlockInstructionsLength);
auto InstAddressReg = Thread->OpDispatcher->_EntrypointOffset(GPRSize, InstAddress - GuestRIP);
std::array<uint8_t, 0x10> CodeOriginal;
memcpy(CodeOriginal.data(), ExistingCodePtr, DecodedInfo->InstSize);
auto CodeChanged = Thread->OpDispatcher->_ValidateCode(CodeOriginal, InstAddressReg, DecodedInfo->InstSize);
auto crc32 = [](const uint8_t* Ptr, size_t Size) -> uint32_t {
#if defined(ARCHITECTURE_arm64)
uint32_t Result {};
#define do_crc(type, suffix) \
while (Size >= sizeof(type)) { \
Result = __crc32##suffix(Result, *reinterpret_cast<const type*>(Ptr)); \
Ptr += sizeof(type); \
Size -= sizeof(type); \
}
do_crc(uint64_t, d);
do_crc(uint32_t, w);
do_crc(uint16_t, h);
do_crc(uint8_t, b);
return Result;
#else
// Unsupported on non-arm.
return 0;
#endif
};
auto CodeChanged = Thread->OpDispatcher->_ValidateCode(
Thread->OpDispatcher->Constant(crc32(ExistingCodePtr, DecodedInfo->InstSize)), InstAddressReg, DecodedInfo->InstSize);
auto InvalidateCodeCond = Thread->OpDispatcher->CondJump(CodeChanged);
@@ -651,7 +672,12 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
Thread->OpDispatcher->SetCurrentCodeBlock(CodeWasChangedBlock);
Thread->OpDispatcher->StartNewBlock();
Thread->OpDispatcher->_ThreadRemoveCodeEntry();
// Generate a relocatable entry for invalidation purposes.
auto EntryReg = Thread->OpDispatcher->_EntrypointOffset(GPRSize, 0);
Thread->OpDispatcher->_ThreadRemoveCodeEntry(EntryReg);
// Exit the function at this instruction after invalidation.
Thread->OpDispatcher->ExitFunction(Thread->OpDispatcher->_InlineEntrypointOffset(GPRSize, InstAddress - GuestRIP));
auto NextOpBlock = Thread->OpDispatcher->CreateNewCodeBlockAfter(CurrentBlock);
@@ -790,6 +816,8 @@ ContextImpl::CompileCodeResult ContextImpl::CompileCode(FEXCore::Core::InternalT
auto [IRView, TotalInstructions, TotalInstructionsLength, StartAddr, Length, NeedsAddGuestCodeRanges] =
GenerateIR(Thread, GuestRIP, Config.GDBSymbols(), MaxInst);
if (!IRView) {
Thread->FrontendDecoder->ValidateDisownedOrFree();
Thread->OpDispatcher->ValidateDisownedOrFree();
// OpDispatcher IR already released in this case.
return {{}, nullptr, 0, 0, false};
}
@@ -803,6 +831,8 @@ ContextImpl::CompileCodeResult ContextImpl::CompileCode(FEXCore::Core::InternalT
if (auto Block = Thread->LookupCache->FindBlock(Thread, GuestRIP)) {
// Raced to compile, release the OpDispatcher IR.
Thread->OpDispatcher->DelayedDisownBuffer();
Thread->FrontendDecoder->ValidateDisownedOrFree();
Thread->OpDispatcher->ValidateDisownedOrFree();
return {.CompiledCode = {.BlockBegin = reinterpret_cast<uint8_t*>(Block), .EntryPoints = {{GuestRIP, reinterpret_cast<uint8_t*>(Block)}}},
.DebugData = nullptr,
.StartAddr = 0,
@@ -821,6 +851,8 @@ ContextImpl::CompileCodeResult ContextImpl::CompileCode(FEXCore::Core::InternalT
// Release the IR
Thread->OpDispatcher->DelayedDisownBuffer();
Thread->FrontendDecoder->ValidateDisownedOrFree();
Thread->OpDispatcher->ValidateDisownedOrFree();
return {
.CompiledCode = std::move(CompiledCode),
.DebugData = std::move(DebugData),
@@ -857,6 +889,47 @@ uintptr_t ContextImpl::CompileBlock(FEXCore::Core::CpuStateFrame* Frame, uint64_
return HostCode;
}
Thread->FrontendDecoder->SetupDecodeInstructionsAtEntry(Thread, GuestRIP, MaxInst);
std::optional<ExecutableFileSectionInfo> Region = SyscallHandler->LookupExecutableFileSection(Thread, GuestRIP);
std::optional<DiskCache::CodeHitData> Hit;
std::optional<uint64_t> DiskCacheGuestCodeKey;
{
FEXCORE_PROFILE_ACCUMULATION(Thread, AccumulatedDiskCacheLookupTime);
Hit = DiskCache.Lookup(Thread, Region, GuestRIP, DiskCacheGuestCodeKey);
if (Hit && !DiskCache.IsValidating()) {
auto LoadedCode = Thread->CPUBackend->LoadCachedCode(Hit->HostCode);
if (LoadedCode.BlockBegin) {
for (auto& CodePage : Hit->GuestPages) {
if (Thread->LookupCache->AddBlockExecutableRange(Thread, Hit->EntryPointRIPs, CodePage, FEXCore::Utils::FEX_PAGE_SIZE)) {
SyscallHandler->MarkGuestExecutableRange(Thread, CodePage, FEXCore::Utils::FEX_PAGE_SIZE);
}
}
LOGMAN_THROW_A_FMT(Hit->EntryPointRIPs.size() == Hit->EntryPointHostOffsets.size(), "Mismatched Disk Cache entrypoint pairs!");
uintptr_t CachedHostCode = 0;
for (size_t i = 0; i < Hit->EntryPointRIPs.size(); i++) {
void* HostAddr = LoadedCode.BlockBegin + Hit->EntryPointHostOffsets[i];
Thread->LookupCache->AddBlockMapping(Thread, Hit->EntryPointRIPs[i], Hit->GuestPages, HostAddr);
if (Hit->EntryPointRIPs[i] == GuestRIP) {
CachedHostCode = reinterpret_cast<uintptr_t>(HostAddr);
}
}
LOGMAN_THROW_A_FMT(CachedHostCode != 0, "Couldn't find GuestRIP in Disk Cache entrypoints!");
FEXCORE_PROFILE_INSTANT_INCREMENT(Thread, AccumulatedDiskCacheHitCount, 1);
Thread->FrontendDecoder->DelayedDisownBuffer();
Thread->FrontendDecoder->ValidateDisownedOrFree();
Thread->OpDispatcher->ValidateDisownedOrFree();
return CachedHostCode;
}
}
FEXCORE_PROFILE_INSTANT_INCREMENT(Thread, AccumulatedDiskCacheMissCount, 1);
}
// Accumulate a JIT count now, as even if another thread raced us, it should count as a compile.
FEXCORE_PROFILE_INSTANT_INCREMENT(Thread, AccumulatedJITCount, 1);
@@ -869,6 +942,13 @@ uintptr_t ContextImpl::CompileBlock(FEXCore::Core::CpuStateFrame* Frame, uint64_
return reinterpret_cast<uintptr_t>(CodePtr);
}
if (DiskCacheGuestCodeKey && Hit && DiskCache.IsValidating()) {
DiskCache.Validate(*DiskCacheGuestCodeKey, *Hit, CompiledCode, Region);
}
// if this ever fires, we need to serialize the offset into disk cache
LOGMAN_THROW_A_FMT(StartAddr == GuestRIP, "StartAddr offset from GuestRIP");
// The core managed to compile the code.
if (Config.BlockJITNaming()) {
auto FragmentBasePtr = CompiledCode.BlockBegin;
@@ -908,11 +988,6 @@ uintptr_t ContextImpl::CompileBlock(FEXCore::Core::CpuStateFrame* Frame, uint64_
}
}
// Clear any relocations that might have been generated
if (!CodeCache.IsGeneratingCache) {
Thread->CPUBackend->ClearRelocations();
}
fextl::vector<uint64_t> CodePages;
if (NeedsAddGuestCodeRanges) {
@@ -928,19 +1003,37 @@ uintptr_t ContextImpl::CompileBlock(FEXCore::Core::CpuStateFrame* Frame, uint64_
}
}
// Insert to lookup cache
// Disk Cache
if (!CodeCache.IsGeneratingCache) {
if (DiskCacheGuestCodeKey) {
std::span<const FEXCore::CPU::Relocation> Relocations;
if (DebugData && DebugData->Relocations) {
Relocations = *DebugData->Relocations;
}
std::span<const uint8_t> GuestCode = {reinterpret_cast<const uint8_t*>(StartAddr), Length};
const Frontend::Decoder::DecodedBlockInformation* BlockInfo =
NeedsAddGuestCodeRanges ? Thread->FrontendDecoder->GetDecodedBlockInfo() : nullptr;
DiskCache.Store(Thread, Region, GuestRIP, *DiskCacheGuestCodeKey, GuestCode, CompiledCode, Relocations, BlockInfo);
}
if (CodeMapWriter && Region && Region->FileStartVA != 0) {
CodeMapWriter->AppendBlock(*Region, GuestRIP);
}
}
// Insert to lookup cache
for (auto [GuestAddr, HostAddr] : CompiledCode.EntryPoints) {
Thread->LookupCache->AddBlockMapping(Thread, GuestAddr, CodePages, HostAddr);
}
if (CodeMapWriter) {
auto Region = SyscallHandler->LookupExecutableFileSection(Thread, GuestRIP);
if (Region && Region->FileStartVA != 0) {
CodeMapWriter->AppendBlock(*Region, GuestRIP);
}
// Clear any relocations that might have been generated
if (!CodeCache.IsGeneratingCache) {
Thread->CPUBackend->ClearRelocations();
}
Thread->FrontendDecoder->ValidateDisownedOrFree();
Thread->OpDispatcher->ValidateDisownedOrFree();
return (uintptr_t)CodePtr;
}
@@ -953,6 +1046,7 @@ uintptr_t ContextImpl::CompileSingleStep(FEXCore::Core::CpuStateFrame* Frame, ui
// Invalidate might take a unique lock on this, to guarantee that during invalidation no code gets compiled
auto lk = GuardSignalDeferringSection<std::shared_lock>(CodeInvalidationMutex, Thread);
Thread->FrontendDecoder->SetupDecodeInstructionsAtEntry(Thread, GuestRIP, 1);
auto [CompiledCode, DebugData, StartAddr, Length, _] = CompileCode(Thread, GuestRIP, 1);
auto CodePtr = CompiledCode.EntryPoints[GuestRIP];
if (CodePtr == nullptr) {
File diff suppressed because it is too large. Load diff
@@ -121,7 +121,7 @@ void Dispatcher::EmitDispatcher() {
ldr(REG_CALLRET_SP, STATE_PTR(CpuStateFrame, State.callret_sp));
FillSpecialRegs(TMP1, TMP2, false, true);
FillSpecialRegs(TMP1, TMP2, {.SetFIZ = false, .SetPredRegs = true});
// As ARM64EC uses this as an entrypoint for both guest calls and host returns, opportunistically try to return
// using the call-ret stack to avoid unbalancing it.
@@ -357,7 +357,7 @@ void Dispatcher::EmitDispatcher() {
ldr(ARMEmitter::XReg::x4, &l_CompileSingleStep);
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<uintptr_t, void*, void*, uint64_t, uint64_t>(ARMEmitter::Reg::r4);
GenerateIndirectRuntimeCall<uintptr_t, void*, void*, uint64_t>(ARMEmitter::Reg::r4);
} else {
blr(ARMEmitter::Reg::r4); // { CTX, Frame, RIP }
}
+247 -87
View File
@@ -69,7 +69,6 @@ static uint32_t MapVEXToReg(uint8_t vvvv, bool HasXMM) {
Decoder::Decoder(FEXCore::Core::InternalThreadState* Thread)
: Thread {Thread}
, CTX {static_cast<FEXCore::Context::ContextImpl*>(Thread->CTX)}
, OSABI {CTX->SyscallHandler ? CTX->SyscallHandler->GetOSABI() : FEXCore::HLE::SyscallOSABI::OS_UNKNOWN}
, PoolObject {CTX->FrontendAllocator, sizeof(FEXCore::X86Tables::DecodedInst) * DefaultDecodedBufferSize} {
FEX_CONFIG_OPT(ReducedPrecision, X87REDUCEDPRECISION);
@@ -89,6 +88,11 @@ Decoder::Decoder(FEXCore::Core::InternalThreadState* Thread)
}
bool Decoder::CheckRangeExecutable(uint64_t Address, uint64_t Size) {
// Check for wraparound
if (Address + Size < Address) {
return false;
}
while (Address < ExecutableRangeBase || Address + Size > ExecutableRangeEnd) {
auto RangeInfo = CTX->SyscallHandler->QueryGuestExecutableRange(Thread, Address);
ExecutableRangeBase = RangeInfo.Base;
@@ -110,9 +114,8 @@ bool Decoder::CheckRangeExecutable(uint64_t Address, uint64_t Size) {
}
uint8_t Decoder::ReadByte() {
LOGMAN_THROW_A_FMT(InstructionSize < MAX_INST_SIZE, "Max instruction size exceeded!");
std::optional<uint8_t> Byte = PeekByte(0);
if (!Byte) {
if (!Byte || InstructionSize == MAX_INST_SIZE) {
HitNonExecutableRange = true;
// Pretend we read 0, the main decode loop will see HitNonExecutableRange and rollback the instruction.
return 0;
@@ -137,6 +140,8 @@ std::pair<uint64_t, bool> Decoder::ReadData(uint8_t Size) {
uint64_t Res = 0;
uint64_t Address = reinterpret_cast<uint64_t>(InstStream.InstStream + InstructionSize);
LastFieldReadOffset = (uint8_t)InstructionSize;
LastFieldReadSize = Size;
if (CheckRangeExecutable(Address, Size)) {
std::memcpy(&Res, &InstStream.AdjustedInstStream[InstructionSize], Size);
} else {
@@ -634,10 +639,7 @@ Decoder::DecodedBlockStatus Decoder::NormalOp(const FEXCore::X86Tables::X86InstI
CurrentDest->Data.GPR.GPR = MapVEXToReg(Options.vvvv, HasXMMDst);
}
if (Bytes != 0) {
LOGMAN_THROW_A_FMT(Bytes <= 8, "Number of bytes should be <= 8 for literal src");
if (Bytes <= 8 && Bytes > 0) {
auto [Literal, IsRelocation] = ReadData(Bytes);
if (IsRelocation) {
DecodeInst->Src[CurrentSrc].Type = DecodedOperand::OpType::LiteralRelocation;
@@ -662,6 +664,11 @@ Decoder::DecodedBlockStatus Decoder::NormalOp(const FEXCore::X86Tables::X86InstI
DecodeInst->Src[CurrentSrc].Data.Literal.Value = Literal;
}
Bytes = 0;
} else {
// All real x86 instructions have byte sizes that are 8-bytes or less.
// Thunk instruction has an additional 32-byte SHA256 payload that needs to be accounted for.
InstructionSize += Bytes;
Bytes = 0;
}
@@ -1091,13 +1098,21 @@ Decoder::DecodedBlockStatus Decoder::DecodeInstruction(uint64_t PC) {
if (ErrorDuringDecoding != DecodedBlockStatus::SUCCESS || HitNonExecutableRange || HitBadRelocation) [[unlikely]] {
// Put an invalid instruction in the stream so the core can raise SIGILL if hit
// Error while decoding instruction. We don't know the table or instruction size
const auto InstSize = DecodeInst->InstSize;
DecodeInst->TableInfo = nullptr;
auto Result = ErrorDuringDecoding != DecodedBlockStatus::SUCCESS ? ErrorDuringDecoding :
DecodeInst->InstSize ? DecodedBlockStatus::PARTIAL_DECODE_INST :
HitNonExecutableRange ? DecodedBlockStatus::NOEXEC_INST :
DecodedBlockStatus::BAD_RELOCATION;
DecodeInst->InstSize = 0;
return Result;
// A decode error can be caused by substituting zero for an inaccessible
// instruction byte, so the instruction fetch fault takes priority.
if (HitNonExecutableRange) {
return InstSize ? DecodedBlockStatus::PARTIAL_DECODE_INST : DecodedBlockStatus::NOEXEC_INST;
}
if (HitBadRelocation) {
return DecodedBlockStatus::BAD_RELOCATION;
}
return ErrorDuringDecoding;
} else if (!DecodeInst->TableInfo || (DecodeInst->TableInfo->Type == TYPE_INST && !DecodeInst->TableInfo->OpcodeDispatcher.OpDispatch)) {
// If there wasn't an error during decoding but we have no dispatcher for the instruction then claim invalid instruction.
return DecodedBlockStatus::INVALID_INST;
@@ -1321,6 +1336,13 @@ void Decoder::AddBranchTarget(uint64_t Target) {
.BlockStatus = BlockIt->BlockStatus,
};
if (BlockIt->DataMasks.size()) {
auto MaskIt = std::lower_bound(BlockIt->DataMasks.begin(), BlockIt->DataMasks.end(), SplitAddr,
[](const DataMask& Mask, uint64_t Addr) { return Mask.FieldAddress < Addr; });
SplitBlock.DataMasks.assign(MaskIt, BlockIt->DataMasks.end());
BlockIt->DataMasks.erase(MaskIt, BlockIt->DataMasks.end());
}
BlockIt->Size = SplitOffset;
BlockIt->NumInstructions = SplitIdx;
@@ -1345,7 +1367,8 @@ const Decoder::DecodeStream Decoder::AdjustAddrForSpecialRegion(const uint8_t* _
constexpr uint64_t VSyscall_Base = 0xFFFF'FFFF'FF60'0000ULL;
constexpr uint64_t VSyscall_End = VSyscall_Base + 0x1000;
if (OSABI == FEXCore::HLE::SyscallOSABI::OS_LINUX64 && RIP >= VSyscall_Base && RIP < VSyscall_End) {
if (BlockInfo.Is64BitMode && CTX->HostFeatures.HostType == FEXCore::HostFeatures::HostTypeEnum::Linux && RIP >= VSyscall_Base &&
RIP < VSyscall_End) {
// VSyscall
// This doesn't exist on AArch64 and on x86_64 hosts this is emulated with faults to a region mapped with --xp permissions
// Offset 0: vgettimeofday
@@ -1365,106 +1388,140 @@ const Decoder::DecodeStream Decoder::AdjustAddrForSpecialRegion(const uint8_t* _
}
bool Decoder::CheckIfCacheable(FEXCore::Core::InternalThreadState& Thread, const uint8_t* InstStream, uint64_t PC, uint64_t MaxInst) {
DecodeInstructionsAtEntry(&Thread, InstStream, PC, MaxInst);
SetupDecodeInstructionsAtEntry(&Thread, PC, MaxInst);
DecodeLoop(InstStream);
bool Uncacheable = HitBadRelocation;
DelayedDisownBuffer();
return !Uncacheable;
}
void Decoder::DecodeInstructionsAtEntry(FEXCore::Core::InternalThreadState* Thread, const uint8_t* _InstStream, uint64_t PC, uint64_t MaxInst) {
FEXCORE_PROFILE_SCOPED("DecodeInstructions");
BlockInfo.TotalInstructionCount = 0;
BlockInfo.Blocks.clear();
VisitedBlocks.clear();
// Reset internal state management
DecodedSize = 0;
MaxCondBranchForward = 0;
MaxCondBranchBackwards = ~0ULL;
DecodedBuffer = PoolObject.ReownOrClaimBuffer();
// Decode operating mode from thread's CS segment.
const auto CSSegment = Core::CPUState::GetSegmentFromIndex(Thread->CurrentFrame->State, Thread->CurrentFrame->State.cs_idx);
BlockInfo.Is64BitMode = CSSegment->L == 1;
LOGMAN_THROW_A_FMT(BlockInfo.Is64BitMode == CTX->Config.Is64BitMode, "Expected operating mode to not change at runtime!");
EntryPoint = PC;
BlockInfo.EntryPoints = {PC};
uint64_t TotalInstructions {};
SectionMinAddress = 0;
SectionMaxAddress = ~0ULL;
Relocations = nullptr;
if (CTX->GetCodeCache().IsGeneratingCache || EnableCodeCacheValidation) {
// If generating cache, attempt to load section bounds and relocations
if (auto SectionInfo = CTX->SyscallHandler->LookupExecutableFileSection(Thread, EntryPoint)) {
SectionMinAddress = SectionInfo->FileStartVA;
SectionMaxAddress = SectionInfo->EndVA;
Relocations = &SectionInfo->FileInfo.Relocations;
}
void Decoder::DetectDataMasks(uint64_t OpAddress, DecodedBlocks& Block) {
if (LastFieldReadSize < 4) {
return;
}
DecodedMinAddress = EntryPoint;
DecodedMaxAddress = EntryPoint;
FEXCore::X86Tables::DecodedOperand* LiteralToPatch = nullptr;
DataMaskType Type;
// Entry is a jump target
BlocksToDecode = {PC};
// mov reg,imm
if (DecodeInst->OP >= 0xB8 && DecodeInst->OP <= 0xBF) {
for (auto& Src : DecodeInst->Src) {
if (Src.IsLiteral()) {
LiteralToPatch = &Src;
break;
}
}
uint64_t CurrentCodePage = PC & FEXCore::Utils::FEX_PAGE_MASK;
BlockInfo.CodePages = {CurrentCodePage};
if (MaxInst == 0) {
MaxInst = CTX->Config.MaxInstPerBlock;
// we could filter to certain high values that are more likely to be pointers/etc?
// const uint64_t Value = Lit->Data.Literal.Value;
// if (LiteralToPatch && Value < 0x1000000ULL) {
// LiteralToPatch = nullptr;
// }
Type = DataMaskType::MOV;
}
bool EntryBlock {true};
bool FinalInstruction {false};
// jmp/call branches that use a literal rip-relative offset
// some of those may be inlined by multiblock and will be cleaned up at decode end
if (DecodeInst->TableInfo->Flags & X86Tables::InstFlags::FLAGS_SETS_RIP && DecodeInst->Src[0].IsLiteral()) {
LiteralToPatch = &DecodeInst->Src[0];
Type = DataMaskType::BRANCH;
}
while (!FinalInstruction && !BlocksToDecode.empty()) {
auto BlockDecodeIt = BlocksToDecode.begin();
uint64_t RIPToDecode = *BlockDecodeIt;
BlocksToDecode.erase(BlockDecodeIt);
VisitedBlocks.emplace(RIPToDecode);
// todo add a bunch more
auto BlockSuccIt = std::lower_bound(BlockInfo.Blocks.begin(), BlockInfo.Blocks.end(), RIPToDecode,
[](const auto& a, uint64_t Address) { return a.Entry < Address; });
if (LiteralToPatch) {
Block.DataMasks.push_back({OpAddress + LastFieldReadOffset, Type, LastFieldReadSize});
LOGMAN_THROW_A_FMT(BlockSuccIt == BlockInfo.Blocks.end() || BlockSuccIt->Entry != RIPToDecode, "unexpected");
LiteralToPatch->Type = X86Tables::DecodedOperand::OpType::LiteralPatchable;
LiteralToPatch->Data.LiteralPatchable.FieldOffset = LastFieldReadOffset;
LiteralToPatch->Data.LiteralPatchable.Width = LastFieldReadSize;
}
}
NextBlockStartAddress = ~0ULL;
if (!BlocksToDecode.empty()) {
// We just erased the lowest, the front is then the second lowest
NextBlockStartAddress = *BlocksToDecode.begin();
void Decoder::PruneInlinedBranchDataMasks() {
for (auto& Block : BlockInfo.Blocks) {
if (!Block.DataMasks.size()) {
continue;
}
if (BlockSuccIt != BlockInfo.Blocks.end() && BlockSuccIt->Entry < NextBlockStartAddress) {
NextBlockStartAddress = BlockSuccIt->Entry;
const auto& LastInst = Block.DecodedInstructions[Block.NumInstructions - 1];
const auto& LastMask = Block.DataMasks.back();
if (LastMask.Type != DataMaskType::BRANCH) {
continue;
}
LOGMAN_THROW_A_FMT(NextBlockStartAddress > RIPToDecode, "unexpected");
// Insert the block now so it can be looked up and split if necessary on a backward edge
auto BlockIt = BlockInfo.Blocks.emplace(BlockSuccIt);
const uint64_t NextInst = LastInst.PC + LastInst.InstSize;
if (LastMask.FieldAddress < LastInst.PC || LastMask.FieldAddress + LastMask.ValueSize > NextInst) {
continue;
}
BlockIt->Entry = RIPToDecode;
BlockIt->Size = 0;
BlockIt->IsEntryPoint = EntryBlock;
if (std::ranges::binary_search(BlockInfo.Blocks, NextInst + LastInst.Src[0].Data.LiteralPatchable.Value, std::less {}, &DecodedBlocks::Entry)) {
Block.DataMasks.pop_back();
}
}
}
uint64_t PCOffset = 0;
uint64_t BlockStartOffset = DecodedSize;
bool EraseBlock = true; // Unset once the block contains an instruction
void Decoder::DecodeLoop(const uint8_t* _InstStream, uint64_t GuestSizePause) {
// counter-intuitively, the masks are also needed for lookup on anon prefix decodes, not just stores
bool WantsDataMasks = CTX->DiskCache.IsReadingDiskCache() || CTX->DiskCache.IsWritingDiskCache();
// remove this if we ever fixup ValidateCode crc constant after relocations
if (CTX->Config.SMCChecks == FEXCore::Config::CONFIG_SMC_FULL) {
WantsDataMasks = false;
}
BlockIt->DecodedInstructions = &DecodedBuffer[BlockStartOffset];
BlockIt->NumInstructions = 0;
while (!FinalInstruction && (Paused || !BlocksToDecode.empty())) {
bool Pausing = false;
fextl::vector<DecodedBlocks>::iterator BlockIt;
if (!Paused || BlockResume == -1) {
auto BlockDecodeIt = BlocksToDecode.begin();
uint64_t RIPToDecode = *BlockDecodeIt;
BlocksToDecode.erase(BlockDecodeIt);
VisitedBlocks.emplace(RIPToDecode);
// Do a bit of pointer math to figure out where we are in code
InstStream = AdjustAddrForSpecialRegion(_InstStream, EntryPoint, RIPToDecode);
auto BlockSuccIt = std::lower_bound(BlockInfo.Blocks.begin(), BlockInfo.Blocks.end(), RIPToDecode,
[](const auto& a, uint64_t Address) { return a.Entry < Address; });
LOGMAN_THROW_A_FMT(BlockSuccIt == BlockInfo.Blocks.end() || BlockSuccIt->Entry != RIPToDecode, "unexpected");
NextBlockStartAddress = ~0ULL;
if (!BlocksToDecode.empty()) {
// We just erased the lowest, the front is then the second lowest
NextBlockStartAddress = *BlocksToDecode.begin();
}
if (BlockSuccIt != BlockInfo.Blocks.end() && BlockSuccIt->Entry < NextBlockStartAddress) {
NextBlockStartAddress = BlockSuccIt->Entry;
}
LOGMAN_THROW_A_FMT(NextBlockStartAddress == ~0ULL || NextBlockStartAddress > RIPToDecode, "unexpected");
// Insert the block now so it can be looked up and split if necessary on a backward edge
BlockIt = BlockInfo.Blocks.emplace(BlockSuccIt);
BlockIt->Entry = RIPToDecode;
BlockIt->Size = 0;
BlockIt->IsEntryPoint = EntryBlock;
PCOffset = 0;
BlockStartOffset = DecodedSize;
EraseBlock = true; // Unset once the block contains an instruction
BlockIt->DecodedInstructions = &DecodedBuffer[BlockStartOffset];
BlockIt->NumInstructions = 0;
// Do a bit of pointer math to figure out where we are in code
InstStream = AdjustAddrForSpecialRegion(_InstStream, EntryPoint, RIPToDecode);
} else if (BlockResume != -1) {
BlockIt = BlockInfo.Blocks.begin() + BlockResume;
BlockResume = -1;
}
Paused = false;
while (1) {
InstructionSize = 0;
// MAX_INST_SIZE assumes worst case
auto OpAddress = RIPToDecode + PCOffset;
auto OpAddress = BlockIt->Entry + PCOffset;
auto OpMaxAddress = OpAddress + MAX_INST_SIZE;
auto OpMinPage = OpAddress & FEXCore::Utils::FEX_PAGE_MASK;
@@ -1486,6 +1543,7 @@ void Decoder::DecodeInstructionsAtEntry(FEXCore::Core::InternalThreadState* Thre
BlockInfo.CodePages.insert(CurrentCodePage);
}
LastFieldReadSize = 0;
BlockIt->BlockStatus = DecodeInstruction(OpAddress);
if (HitBadRelocation) {
BlockInfo.TotalInstructionCount = 0;
@@ -1510,6 +1568,11 @@ void Decoder::DecodeInstructionsAtEntry(FEXCore::Core::InternalThreadState* Thre
++BlockIt->NumInstructions;
BlockIt->Size += DecodeInst->InstSize;
// if we weren't provided relocations (guest JIT), try to detect what we can
if (WantsDataMasks && BlockIt->BlockStatus == DecodedBlockStatus::SUCCESS && BlockInfo.Is64BitMode && !Relocations) {
DetectDataMasks(OpAddress, *BlockIt);
}
// Can not continue this block at all on invalid instruction
if (BlockIt->BlockStatus != DecodedBlockStatus::SUCCESS) [[unlikely]] {
if (!EntryBlock && BlockIt->BlockStatus != DecodedBlockStatus::BAD_RELOCATION) {
@@ -1519,6 +1582,9 @@ void Decoder::DecodeInstructionsAtEntry(FEXCore::Core::InternalThreadState* Thre
TotalInstructions -= BlockIt->NumInstructions;
DecodedSize = BlockStartOffset;
InstStream -= PCOffset;
if (DecodedMaxAddress == OpEndAddress) {
DecodedMaxAddress -= PCOffset;
}
EraseBlock = true;
} else {
LogMan::Msg::EFmt("{} instruction in entry block: {:X}",
@@ -1532,6 +1598,15 @@ void Decoder::DecodeInstructionsAtEntry(FEXCore::Core::InternalThreadState* Thre
break;
}
if (GuestSizePause) {
if (GuestSizePause > DecodeInst->InstSize) {
GuestSizePause -= DecodeInst->InstSize;
} else {
GuestSizePause = 0;
Pausing = true;
}
}
// Check if we need to end the entire multiblock
FinalInstruction = DecodedSize >= MaxInst || DecodedSize >= DefaultDecodedBufferSize || TotalInstructions >= MaxInst;
if (FinalInstruction) {
@@ -1544,7 +1619,12 @@ void Decoder::DecodeInstructionsAtEntry(FEXCore::Core::InternalThreadState* Thre
// If the branch target is within our multiblock range then we can keep going on
// We don't want to short circuit this since we want to calculate our ranges still
// NOTE: This will invalidate BlockIt, this is fine as we immediately break from the loop and EraseBlock cannot be true
BlockIt->ForceFullSMCDetection = CTX->AreMonoHacksActive() && IsBranchMonoTailcall(BlockIt->NumInstructions);
if (CTX->AreMonoHacksActive() && IsBranchMonoTailcall(BlockIt->NumInstructions)) {
BlockIt->ForceFullSMCDetection = true;
// todo abandon patching this for now, as the crc will fail and it will lock up redoing it over and over
// we should fix the crc at relocation if this is important
BlockIt->DataMasks.clear();
}
BranchTargetInMultiblockRange();
}
@@ -1553,6 +1633,17 @@ void Decoder::DecodeInstructionsAtEntry(FEXCore::Core::InternalThreadState* Thre
PCOffset += DecodeInst->InstSize;
InstStream += DecodeInst->InstSize;
if (Pausing) {
Pausing = false;
Paused = true;
BlockResume = BlockIt - BlockInfo.Blocks.begin();
break;
}
}
if (Paused) {
break;
}
// NOTE: BlockIt is only valid here in the EraseBlock case
@@ -1564,6 +1655,16 @@ void Decoder::DecodeInstructionsAtEntry(FEXCore::Core::InternalThreadState* Thre
CurrentBlockTargets.clear();
EntryBlock = false;
if (Pausing && !BlocksToDecode.empty() && !FinalInstruction) {
Paused = true;
BlockResume = -1;
break;
}
}
if (Paused) {
return;
}
BlockInfo.TotalInstructionCount = TotalInstructions;
@@ -1571,6 +1672,65 @@ void Decoder::DecodeInstructionsAtEntry(FEXCore::Core::InternalThreadState* Thre
for (auto& Block : BlockInfo.Blocks) {
Block.IsEntryPoint = BlockInfo.EntryPoints.contains(Block.Entry);
}
// now that multiblock has settled down, remove any branch masks we put down that didn't end the block
if (WantsDataMasks) {
PruneInlinedBranchDataMasks();
}
}
void Decoder::SetupDecodeInstructionsAtEntry(FEXCore::Core::InternalThreadState* Thread, uint64_t PC, uint64_t MaxInst) {
FEXCORE_PROFILE_SCOPED("DecodeInstructions");
BlockInfo.TotalInstructionCount = 0;
BlockInfo.Blocks.clear();
VisitedBlocks.clear();
// Reset internal state management
Paused = false;
BlockResume = -1;
DecodedSize = 0;
if (MaxInst == 0) {
MaxInst = CTX->Config.MaxInstPerBlock;
}
this->MaxInst = MaxInst;
MaxCondBranchForward = 0;
MaxCondBranchBackwards = ~0ULL;
DecodedBuffer = PoolObject.ReownOrClaimBuffer();
// Decode operating mode from thread's CS segment.
const auto CSSegment = Core::CPUState::GetSegmentFromIndex(Thread->CurrentFrame->State, Thread->CurrentFrame->State.cs_idx);
BlockInfo.Is64BitMode = CSSegment->L == 1;
LOGMAN_THROW_A_FMT(BlockInfo.Is64BitMode == CTX->Config.Is64BitMode, "Expected operating mode to not change at runtime!");
EntryPoint = PC;
BlockInfo.EntryPoints = {PC};
TotalInstructions = 0;
SectionMinAddress = 0;
SectionMaxAddress = ~0ULL;
Relocations = nullptr;
if (CTX->GetCodeCache().IsGeneratingCache || EnableCodeCacheValidation) {
// If generating cache, attempt to load section bounds and relocations
if (auto SectionInfo = CTX->SyscallHandler->LookupExecutableFileSection(Thread, EntryPoint)) {
SectionMinAddress = SectionInfo->FileStartVA;
SectionMaxAddress = SectionInfo->EndVA;
Relocations = &SectionInfo->FileInfo.Relocations;
}
}
DecodedMinAddress = EntryPoint;
DecodedMaxAddress = EntryPoint;
// Entry is a jump target
BlocksToDecode = {PC};
CurrentCodePage = PC & FEXCore::Utils::FEX_PAGE_MASK;
BlockInfo.CodePages = {CurrentCodePage};
EntryBlock = true;
FinalInstruction = false;
}
} // namespace FEXCore::Frontend
+32 -5
View File
@@ -19,9 +19,6 @@
namespace FEXCore::Context {
class ContextImpl;
}
namespace FEXCore::HLE {
enum class SyscallOSABI;
}
namespace FEXCore::Frontend {
class Decoder final {
@@ -35,6 +32,14 @@ public:
UNIMPLEMENTED_INST,
};
enum class DataMaskType : uint8_t { MOV, BRANCH };
struct DataMask final {
uint64_t FieldAddress;
DataMaskType Type;
uint8_t ValueSize;
};
// New Frontend decoding
struct DecodedBlocks final {
uint64_t Entry {};
@@ -44,6 +49,7 @@ public:
DecodedBlockStatus BlockStatus;
bool IsEntryPoint {};
bool ForceFullSMCDetection {};
fextl::vector<DataMask> DataMasks;
};
struct DecodedBlockInformation final {
@@ -57,7 +63,8 @@ public:
Decoder(FEXCore::Core::InternalThreadState* Thread);
bool CheckIfCacheable(FEXCore::Core::InternalThreadState&, const uint8_t* InstStream, uint64_t PC, uint64_t MaxInst);
void DecodeInstructionsAtEntry(FEXCore::Core::InternalThreadState* Thread, const uint8_t* InstStream, uint64_t PC, uint64_t MaxInst);
void SetupDecodeInstructionsAtEntry(FEXCore::Core::InternalThreadState* Thread, uint64_t PC, uint64_t MaxInst);
void DecodeLoop(const uint8_t* InstStream, uint64_t GuestPause = 0);
const DecodedBlockInformation* GetDecodedBlockInfo() const {
return &BlockInfo;
@@ -74,6 +81,10 @@ public:
PoolObject.DelayedDisownBuffer();
}
void ValidateDisownedOrFree() const {
PoolObject.ValidateDisownedOrFree();
}
void ResetExecutableRangeCache() {
ExecutableRangeBase = ExecutableRangeEnd = 0;
}
@@ -89,7 +100,6 @@ private:
FEXCore::Core::InternalThreadState* Thread;
FEXCore::Context::ContextImpl* CTX;
const FEXCore::HLE::SyscallOSABI OSABI {};
FEX_CONFIG_OPT(EnableCodeCacheValidation, ENABLECODECACHEVALIDATION);
@@ -102,6 +112,9 @@ private:
void AddBranchTarget(uint64_t Target);
void DetectDataMasks(uint64_t OpAddress, DecodedBlocks& Block);
void PruneInlinedBranchDataMasks();
bool CheckRangeExecutable(uint64_t Address, uint64_t Size);
uint8_t ReadByte();
@@ -121,6 +134,19 @@ private:
FEXCore::X86Tables::DecodedInst* DecodedBuffer {};
Utils::PoolBufferWithTimedRetirement<FEXCore::X86Tables::DecodedInst*, 5000, 500> PoolObject;
size_t DecodedSize {};
uint64_t TotalInstructions {};
uint64_t CurrentCodePage {};
bool EntryBlock {};
bool FinalInstruction {};
uint64_t MaxInst {};
bool Paused {};
int64_t BlockResume = -1;
uint64_t PCOffset {};
uint64_t BlockStartOffset {};
bool EraseBlock {};
uint8_t LastFieldReadOffset;
uint8_t LastFieldReadSize;
uint64_t ExecutableRangeBase {};
uint64_t ExecutableRangeEnd {};
@@ -155,6 +181,7 @@ private:
static constexpr size_t MAX_INST_SIZE = 15;
uint8_t InstructionSize {};
// Contains the full decoded instruction, unless it is a `Thunk` instruction.
std::array<uint8_t, MAX_INST_SIZE> Instruction;
uint8_t LastEscapePrefix {};
FEXCore::X86Tables::DecodedInst* DecodeInst;
@@ -10,11 +10,6 @@
namespace FEXCore::CPU {
template<typename R, typename... Args>
static FallbackInfo GetFallbackInfo(R (*fn)(Args...), FEXCore::Core::FallbackHandlerIndex HandlerIndex) {
return {FABI_UNKNOWN, HandlerIndex};
}
void InterpreterOps::FillFallbackIndexPointers(Core::FallbackABIInfo* Info, uint64_t* ABIHandlers) {
Info[Core::OPINDEX_F80CVTTO_4] = {ABIHandlers[FABI_F80_I16_F32_PTR],
reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTO>::handle4)};
@@ -216,12 +211,6 @@ bool InterpreterOps::GetFallbackHandler(const IR::IROp_Header* IROp, FallbackInf
return true; \
}
#define COMMON_F64_OP(OP) \
case IR::OP_F64##OP: { \
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F64##OP>::handle, Core::OPINDEX_F64##OP); \
return true; \
}
#define COMMON_UNARY_F64_OP(OP) \
case IR::OP_F64##OP: { \
*Info = {FABI_F64_F64_PTR, Core::OPINDEX_F64##OP}; \
+8 -6
View File
@@ -13,9 +13,6 @@ $end_info$
namespace FEXCore::CPU {
#define GRD(Node) (IROp->Size <= 4 ? GetDst<RA_32>(Node) : GetDst<RA_64>(Node))
#define GRS(Node) (IROp->Size <= 4 ? GetReg<RA_32>(Node) : GetReg<RA_64>(Node))
#define DEF_BINOP_WITH_CONSTANT(FEXOp, VarOp, ConstOp) \
DEF_OP(FEXOp) { \
auto Op = IROp->C<IR::IROp_##FEXOp>(); \
@@ -67,6 +64,11 @@ DEF_OP(EntrypointOffset) {
InsertGuestRIPMove(GetReg(Node), Constant & Mask);
}
DEF_OP(PatchableGuestData) {
auto Op = IROp->C<IR::IROp_PatchableGuestData>();
InsertGuestPatchableDataMove(GetReg(Node), Op->Value, Op->SiteAddress, (uint8_t)Op->SiteSize);
}
DEF_OP(InlineConstant) {
// nop
}
@@ -421,8 +423,8 @@ DEF_OP(MulH) {
if (OpSize == IR::OpSize::i32Bit) {
sxtw(TMP1, Src1.W());
sxtw(TMP2, Src2.W());
mul(ARMEmitter::Size::i32Bit, Dst, TMP1, TMP2);
ubfx(ARMEmitter::Size::i32Bit, Dst, Dst, 32, 32);
mul(ARMEmitter::Size::i64Bit, Dst, TMP1, TMP2);
ubfx(ARMEmitter::Size::i64Bit, Dst, Dst, 32, 32);
} else {
smulh(Dst.X(), Src1.X(), Src2.X());
}
@@ -774,7 +776,7 @@ DEF_OP(PDep) {
// Now, they're copied, so we can start setting Dest (even if it overlaps with
// one of them). Handle early exit case
mov(EmitSize, Dest, 0);
(void)cbz(EmitSize, OrigMask, &Done);
(void)cbz(EmitSize, Mask, &Done);
// Setup for first iteration
neg(EmitSize, T0, Mask);
@@ -58,7 +58,8 @@ Arm64JITCore::NamedSymbolLiteralPair Arm64JITCore::InsertNamedSymbolLiteral(FEXC
void Arm64JITCore::PlaceNamedSymbolLiteral(NamedSymbolLiteralPair Lit) {
switch (Lit.MoveABI.Header.Type) {
case RelocationTypes::RELOC_NAMED_SYMBOL_LITERAL:
case RelocationTypes::RELOC_GUEST_RIP_LITERAL: {
case RelocationTypes::RELOC_GUEST_RIP_LITERAL:
case RelocationTypes::RELOC_GUEST_PATCHABLE_RIP_LITERAL: {
Lit.MoveABI.Header.Offset = GetCursorOffset();
break;
}
@@ -102,6 +103,48 @@ void Arm64JITCore::InsertGuestRIPMove(ARMEmitter::Register Reg, uint64_t Constan
Relocations.emplace_back(MoveABI);
}
auto Arm64JITCore::InsertGuestPatchableRIPLiteral(uint64_t GuestRIP, uint64_t SiteAddress, uint8_t ValueSize) -> NamedSymbolLiteralPair {
return {
.Lit = GuestRIP,
.MoveABI =
{
.GuestPatchableData = {.Header =
{
.Offset = 0, // Set by PlaceNamedSymbolLiteral
.Type = FEXCore::CPU::RelocationTypes::RELOC_GUEST_PATCHABLE_RIP_LITERAL,
},
.RegisterIndex = 0, // unused
.ValueSize = ValueSize,
// NOTE: Cache serialization will subtract the unit entry address later
.SiteAddress = SiteAddress},
},
};
}
void Arm64JITCore::InsertGuestPatchableDataMove(ARMEmitter::Register Reg, uint64_t Value, uint64_t SiteAddress, uint8_t ValueSize) {
Relocation MoveABI = Relocation::Default();
MoveABI.GuestPatchableData.Header = {.Offset = GetCursorOffset(), .Type = FEXCore::CPU::RelocationTypes::RELOC_GUEST_PATCHABLE_DATA_MOVE};
MoveABI.GuestPatchableData.RegisterIndex = Reg.Idx();
MoveABI.GuestPatchableData.ValueSize = ValueSize;
MoveABI.GuestPatchableData.SiteAddress = SiteAddress;
// this might get patched on disk cache load
LoadConstant(ARMEmitter::Size::i64Bit, Reg, Value, FEXCore::CPU::Arm64Emitter::PadType::DOPAD);
Relocations.emplace_back(MoveABI);
}
void Arm64JITCore::InsertGuestPatchableRIPMove(ARMEmitter::Register Reg, uint64_t Value, uint64_t SiteAddress, uint8_t ValueSize) {
Relocation MoveABI = Relocation::Default();
MoveABI.GuestPatchableData.Header = {.Offset = GetCursorOffset(), .Type = FEXCore::CPU::RelocationTypes::RELOC_GUEST_PATCHABLE_RIP_MOVE};
MoveABI.GuestPatchableData.RegisterIndex = Reg.Idx();
MoveABI.GuestPatchableData.ValueSize = ValueSize;
MoveABI.GuestPatchableData.SiteAddress = SiteAddress;
// this might get patched on disk cache load
LoadConstant(ARMEmitter::Size::i64Bit, Reg, Value, FEXCore::CPU::Arm64Emitter::PadType::DOPAD);
Relocations.emplace_back(MoveABI);
}
fextl::vector<FEXCore::CPU::Relocation> Arm64JITCore::TakeRelocations(uint64_t GuestBaseAddress) {
// Rebase relocations to library base address
for (auto& Relocation : Relocations) {
@@ -342,8 +342,8 @@ DEF_OP(TelemetrySetValue) {
(void)Bind(&LoopTop);
ldaxr(ARMEmitter::SubRegSize::i64Bit, TMP3, TMP2);
orr(ARMEmitter::Size::i32Bit, TMP3, TMP3, Src);
stlxr(ARMEmitter::SubRegSize::i64Bit, TMP3, TMP3, TMP2);
(void)cbnz(ARMEmitter::Size::i32Bit, TMP3, &LoopTop);
stlxr(ARMEmitter::SubRegSize::i64Bit, TMP4, TMP3, TMP2);
(void)cbnz(ARMEmitter::Size::i32Bit, TMP4, &LoopTop);
}
#endif
}
+48 -61
View File
@@ -83,7 +83,11 @@ DEF_OP(ExitFunction) {
if (NewRIP < EC_CODE_BITMAP_MAX_ADDRESS && RtlIsEcCode(NewRIP)) {
str(REG_CALLRET_SP, STATE_PTR(CpuStateFrame, State.callret_sp));
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, StaticRegisters[X86State::REG_RSP], 0);
InsertGuestRIPMove(EC_CALL_CHECKER_PC_REG, NewRIP);
if (Op->PatchSiteAddress) {
InsertGuestPatchableRIPMove(EC_CALL_CHECKER_PC_REG, NewRIP, Op->PatchSiteAddress, Op->PatchSiteSize);
} else {
InsertGuestRIPMove(EC_CALL_CHECKER_PC_REG, NewRIP);
}
ldr(TMP2, STATE_PTR(CpuStateFrame, Pointers.ExitFunctionEC));
br(TMP2);
} else {
@@ -173,6 +177,7 @@ DEF_OP(ExitFunction) {
ARMEmitter::ForwardLabel TFUnset;
ldrb(TMP1, STATE_PTR(CpuStateFrame, State.flags[X86State::RFLAG_TF_RAW_LOC]));
(void)cbz(ARMEmitter::Size::i32Bit, TMP1, &TFUnset);
// todo do we need to account for cache patching here?
InsertGuestRIPMove(TMP1, NewRIP);
str(TMP1, STATE, offsetof(FEXCore::Core::CpuStateFrame, State.rip));
ldr(TMP2, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.DispatcherLoopTop));
@@ -180,7 +185,7 @@ DEF_OP(ExitFunction) {
(void)Bind(&TFUnset);
}
EmitLinkedBranch(NewRIP, Op->Hint == IR::BranchHint::Call);
EmitLinkedBranch(NewRIP, Op->Hint == IR::BranchHint::Call, Op->PatchSiteAddress, Op->PatchSiteSize);
(void)Bind(&l_CallReturn);
#ifdef ARCHITECTURE_arm64ec
}
@@ -277,11 +282,9 @@ DEF_OP(CondJump) {
}
DEF_OP(Syscall) {
auto Op = IROp->C<IR::IROp_Syscall>();
// Arguments are passed as follows:
// X0: SyscallHandler
// X1: ThreadState
// X2: Pointer to SyscallArguments
PushDynamicRegs(TMP1);
@@ -300,31 +303,18 @@ DEF_OP(Syscall) {
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, GPRSpillMask & 0xFFFF);
str(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CpuStateFrame, InSyscallInfo));
uint64_t SPOffset = AlignUp(FEXCore::HLE::SyscallArguments::MAX_ARGS * 8, 16);
sub(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, ARMEmitter::Reg::rsp, SPOffset);
for (uint32_t i = 0; i < FEXCore::HLE::SyscallArguments::MAX_ARGS; ++i) {
if (Op->Header.Args[i].IsInvalid()) {
continue;
}
str(GetReg(Op->Header.Args[i]).X(), ARMEmitter::Reg::rsp, i * 8);
}
ldr(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.SyscallHandlerObj));
ldr(ARMEmitter::XReg::x3, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.SyscallHandlerFunc));
mov(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r1, STATE.R());
// SP supporting move
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r2, ARMEmitter::Reg::rsp, 0);
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<uint64_t, void*, void*, void*>(ARMEmitter::Reg::r3);
} else {
blr(ARMEmitter::Reg::r3);
}
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, ARMEmitter::Reg::rsp, SPOffset);
// Result is now in x0
// Fix the stack and any values that were stepped on
// Syscall result is in any static register that the frontend desired.
FillStaticRegs({
.OptionalReg = ARMEmitter::Reg::r1,
.OptionalReg2 = ARMEmitter::Reg::r2,
@@ -337,14 +327,6 @@ DEF_OP(Syscall) {
str(ARMEmitter::XReg::zr, STATE, offsetof(FEXCore::Core::CpuStateFrame, InSyscallInfo));
PopDynamicRegs();
const auto OSABI = CTX->SyscallHandler->GetOSABI();
if (OSABI != FEXCore::HLE::SyscallOSABI::OS_GENERIC) {
// Move result to its destination register.
// Only if `NORETURNEDRESULT` wasn't set, otherwise we might overwrite the CPUState refilled with `FillStaticRegs`
mov(ARMEmitter::Size::i64Bit, GetReg(Node), ARMEmitter::Reg::r0);
}
}
DEF_OP(Thunk) {
@@ -379,53 +361,61 @@ DEF_OP(Thunk) {
DEF_OP(ValidateCode) {
auto Op = IROp->C<IR::IROp_ValidateCode>();
auto OldCode = Op->CodeOriginal.data();
auto Base = GetReg(Op->Header.Args[0]).X();
auto Base = GetReg(Op->Address).X();
int len = Op->CodeLength;
int Offset = 0;
ARMEmitter::ForwardLabel Fail;
const auto Dst = GetReg(Node);
const auto CRC32Reg = GetReg(Op->crc);
auto EmitCheck = [&](size_t Size, auto&& LoadData) {
while (len >= Size) {
LoadData();
sub(ARMEmitter::Size::i64Bit, TMP1, TMP1, TMP2);
cbnz_OrRestart(ARMEmitter::Size::i64Bit, TMP1, &Fail);
len -= Size;
Offset += Size;
}
};
// Changes to TMP1
auto WorkingReg = ARMEmitter::XReg::zr;
auto BaseReg = TMP2;
auto TmpDataReg = TMP3;
mov(ARMEmitter::Size::i64Bit, BaseReg, Base);
EmitCheck(8, [&]() {
ldr(TMP1, Base, Offset);
LoadConstant(ARMEmitter::Size::i64Bit, TMP2, *(const uint64_t*)(OldCode + Offset));
});
while (len >= 8) {
ldr<ARMEmitter::IndexType::POST>(TmpDataReg, BaseReg, 8);
crc32x(TMP1, WorkingReg, TmpDataReg);
len -= 8;
WorkingReg = TMP1;
}
EmitCheck(4, [&]() {
ldr(TMP1.W(), Base, Offset);
LoadConstant(ARMEmitter::Size::i32Bit, TMP2, *(const uint32_t*)(OldCode + Offset));
});
while (len >= 4) {
ldr<ARMEmitter::IndexType::POST>(TmpDataReg.W(), BaseReg, 4);
crc32w(TMP1.W(), WorkingReg.W(), TmpDataReg.W());
len -= 4;
WorkingReg = TMP1;
}
EmitCheck(2, [&]() {
ldrh(TMP1.W(), Base, Offset);
LoadConstant(ARMEmitter::Size::i32Bit, TMP2, *(const uint16_t*)(OldCode + Offset));
});
while (len >= 2) {
ldrh<ARMEmitter::IndexType::POST>(TmpDataReg.W(), BaseReg, 2);
crc32h(TMP1.W(), WorkingReg.W(), TmpDataReg.W());
len -= 2;
WorkingReg = TMP1;
}
EmitCheck(1, [&]() {
ldrb(TMP1.W(), Base, Offset);
LoadConstant(ARMEmitter::Size::i32Bit, TMP2, *(const uint8_t*)(OldCode + Offset));
});
while (len >= 1) {
ldrb<ARMEmitter::IndexType::POST>(TmpDataReg.W(), BaseReg, 1);
crc32b(TMP1.W(), WorkingReg.W(), TmpDataReg.W());
len -= 1;
WorkingReg = TMP1;
}
sub(ARMEmitter::Size::i32Bit, Dst, TMP1, CRC32Reg);
ARMEmitter::ForwardLabel End;
LoadConstant(ARMEmitter::Size::i32Bit, Dst, 0);
b_OrRestart(&End);
BindOrRestart(&Fail);
cbz_OrRestart(ARMEmitter::Size::i32Bit, Dst, &End);
LoadConstant(ARMEmitter::Size::i32Bit, Dst, 1);
BindOrRestart(&End);
}
DEF_OP(ThreadRemoveCodeEntry) {
auto Op = IROp->C<IR::IROp_ThreadRemoveCodeEntry>();
// Move the entry to ABI before saving state.
mov(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r1, GetReg(Op->Entry));
PushDynamicRegs(TMP4);
SpillStaticRegs(TMP4);
@@ -434,9 +424,6 @@ DEF_OP(ThreadRemoveCodeEntry) {
// X1: RIP
mov(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, STATE.R());
// TODO: Relocations don't seem to be wired up to this...?
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r1, Entry, CPU::Arm64Emitter::PadType::AUTOPAD);
ldr(ARMEmitter::XReg::x2, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.ThreadRemoveCodeEntryFromJIT));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<void, void*, void*>(ARMEmitter::Reg::r2);
@@ -292,8 +292,6 @@ DEF_OP(Vector_FToS) {
frinti(SubEmitSize, Dst.Z(), Mask.Merging(), Vector.Z());
fcvtzs(Dst.Z(), SubEmitSize, Mask.Merging(), Dst.Z(), SubEmitSize);
} else {
const auto Dst = GetVReg(Node);
const auto Vector = GetVReg(Op->Vector);
if (OpSize == IR::OpSize::i64Bit) {
frinti(SubEmitSize, Dst.D(), Vector.D());
fcvtzs(SubEmitSize, Dst.D(), Dst.D());
@@ -559,26 +557,30 @@ DEF_OP(Vector_F64ToI32) {
}
}
} else {
// This has a known precision issue that isn't easily resolvable without throwing away performance.
// Doing the conversion in multi-stage steps has an issue that you can lose precision in the f32->i32 step if your source was f64.
// To get around this with ASIMD FEX needs to use fcvtzs (Scalar, Integer, to GPR) for each F64 to be directly converted to i32.
// This is a very costly transform that the SVE path doesn't need to do since it supports f64->i32 directly.
// If this precision issue is necessary then we can add an option for it in the future.
///< Round float to integral depending on rounding mode.
///< skip TowardsZero as fcvtzs below already truncates toward zero on its own
auto CVTReg = Dst.Q();
switch (Round) {
case IR::RoundMode::Nearest: frintn(ARMEmitter::SubRegSize::i64Bit, Dst.Q(), Vector.Q()); break;
case IR::RoundMode::NegInfinity: frintm(ARMEmitter::SubRegSize::i64Bit, Dst.Q(), Vector.Q()); break;
case IR::RoundMode::PosInfinity: frintp(ARMEmitter::SubRegSize::i64Bit, Dst.Q(), Vector.Q()); break;
case IR::RoundMode::TowardsZero: frintz(ARMEmitter::SubRegSize::i64Bit, Dst.Q(), Vector.Q()); break;
case IR::RoundMode::TowardsZero: CVTReg = Vector.Q(); break;
case IR::RoundMode::Host: frinti(ARMEmitter::SubRegSize::i64Bit, Dst.Q(), Vector.Q()); break;
}
// Now narrow from f64 to f32.
fcvtn(ARMEmitter::SubRegSize::i32Bit, Dst.Q(), Dst.Q());
///< Convert f64 directly to i64
fcvtzs(ARMEmitter::SubRegSize::i64Bit, Dst.Q(), CVTReg);
///< Convert the two F32 integrals to real integers.
fcvtzs(ARMEmitter::SubRegSize::i32Bit, Dst.D(), Dst.D());
///< Saturating narrow i64 -> i32
///
///< The caller(Vector_CVT_Float_To_Int32Impl) only fixes up positive overflow:
///< it tests MaxF > Src (MaxF = 2^31) and swaps in CVTMAX_I32 (0x80000000) where
///< the test fails.
///
///< Sources below INT32_MIN are handled by sqxtn:
///< ARM saturates to INT32_MIN, which is 0x80000000 the same value as
///< x86's integer-indefinite value.
sqxtn(ARMEmitter::SubRegSize::i32Bit, Dst.D(), Dst.D());
}
}
@@ -324,24 +324,62 @@ DEF_OP(PCLMUL) {
const auto Op = IROp->C<IR::IROp_PCLMUL>();
const auto OpSize = IROp->Size;
const auto Is256Bit = OpSize == IR::OpSize::i256Bit;
LOGMAN_THROW_A_FMT(!Is256Bit || HostSupportsSVE256, "Need SVE256 support in order to use {} with 256-bit operation", __func__);
const auto Dst = GetVReg(Node);
const auto Src1 = GetVReg(Op->Src1);
const auto Src2 = GetVReg(Op->Src2);
LOGMAN_THROW_A_FMT(OpSize == IR::OpSize::i128Bit, "Currently only supports 128-bit operations.");
switch (Op->Selector) {
case 0b00000000: pmull(ARMEmitter::SubRegSize::i128Bit, Dst.D(), Src1.D(), Src2.D()); break;
case 0b00000001:
dup(ARMEmitter::SubRegSize::i64Bit, VTMP1.Q(), Src1.Q(), 1);
pmull(ARMEmitter::SubRegSize::i128Bit, Dst.D(), VTMP1.D(), Src2.D());
break;
case 0b00010000:
dup(ARMEmitter::SubRegSize::i64Bit, VTMP1.Q(), Src2.Q(), 1);
pmull(ARMEmitter::SubRegSize::i128Bit, Dst.D(), VTMP1.D(), Src1.D());
break;
case 0b00010001: pmull2(ARMEmitter::SubRegSize::i128Bit, Dst.Q(), Src1.Q(), Src2.Q()); break;
default: LOGMAN_MSG_A_FMT("Unknown PCLMUL selector: {}", Op->Selector); break;
if (HostSupportsSVE256 && Is256Bit) {
switch (Op->Selector) {
case 0b00000000: {
pmullb(ARMEmitter::SubRegSize::i128Bit, Dst.Z(), Src1.Z(), Src2.Z());
break;
}
case 0b00000001: {
trn2(ARMEmitter::SubRegSize::i64Bit, VTMP1.Z(), Src1.Z(), Src1.Z());
pmullb(ARMEmitter::SubRegSize::i128Bit, Dst.Z(), VTMP1.Z(), Src2.Z());
break;
}
case 0b00010000:
trn2(ARMEmitter::SubRegSize::i64Bit, VTMP1.Z(), Src2.Z(), Src2.Z());
pmullb(ARMEmitter::SubRegSize::i128Bit, Dst.Z(), Src1.Z(), VTMP1.Z());
break;
case 0b00010001: {
pmullt(ARMEmitter::SubRegSize::i128Bit, Dst.Z(), Src1.Z(), Src2.Z());
break;
}
default: {
LOGMAN_MSG_A_FMT("Unknown PCLMUL selector: {}", Op->Selector);
break;
}
}
} else {
switch (Op->Selector) {
case 0b00000000: {
pmull(ARMEmitter::SubRegSize::i128Bit, Dst.D(), Src1.D(), Src2.D());
break;
}
case 0b00000001: {
dup(ARMEmitter::SubRegSize::i64Bit, VTMP1.Q(), Src1.Q(), 1);
pmull(ARMEmitter::SubRegSize::i128Bit, Dst.D(), VTMP1.D(), Src2.D());
break;
}
case 0b00010000: {
dup(ARMEmitter::SubRegSize::i64Bit, VTMP1.Q(), Src2.Q(), 1);
pmull(ARMEmitter::SubRegSize::i128Bit, Dst.D(), VTMP1.D(), Src1.D());
break;
}
case 0b00010001: {
pmull2(ARMEmitter::SubRegSize::i128Bit, Dst.Q(), Src1.Q(), Src2.Q());
break;
}
default: {
LOGMAN_MSG_A_FMT("Unknown PCLMUL selector: {}", Op->Selector);
break;
}
}
}
}
+81 -63
View File
@@ -616,13 +616,13 @@ void Arm64JITCore::Op_NoOp(const IR::IROp_Header* IROp, IR::Ref Node) {}
Arm64JITCore::Arm64JITCore(FEXCore::Context::ContextImpl* ctx, FEXCore::Core::InternalThreadState* Thread)
: CPUBackend(*ctx, Thread)
, Arm64Emitter(ctx)
, HostSupportsSVE128 {ctx->HostFeatures.SupportsSVE128}
, HostSupportsSVE256 {ctx->HostFeatures.SupportsSVE256}
, HostSupportsSVE128 {ctx->HostFeatures.SupportsSVE128 != 0}
, HostSupportsSVE256 {ctx->HostFeatures.SupportsSVE256 != 0}
, HostSupportsAVX256 {ctx->HostFeatures.SupportsAVX && ctx->HostFeatures.SupportsSVE256}
, HostSupportsRPRES {ctx->HostFeatures.SupportsRPRES}
, HostSupportsAFP {ctx->HostFeatures.SupportsAFP}
, HostSupportsRPRES {ctx->HostFeatures.SupportsRPRES != 0}
, HostSupportsAFP {ctx->HostFeatures.SupportsAFP != 0}
, CTX {ctx}
, TempAllocator(ctx->CPUBackendAllocator, 0) {
, TempCodeBufferAllocator(ctx->CPUBackendAllocator, 0) {
RAPass = Thread->PassManager->GetPass<IR::RegisterAllocationPass>("RA");
@@ -630,7 +630,7 @@ Arm64JITCore::Arm64JITCore(FEXCore::Context::ContextImpl* ctx, FEXCore::Core::In
RAPass->AddRegisters(IR::RegClass::GPRFixed, StaticRegisters.size());
RAPass->AddRegisters(IR::RegClass::FPR, GeneralFPRegisters.size());
RAPass->AddRegisters(IR::RegClass::FPRFixed, StaticFPRegisters.size());
RAPass->PairRegs = PairRegisters;
RAPass->SetNumPairRegs(PairRegisters);
{
// Set up pointers that the JIT needs to load
@@ -666,26 +666,17 @@ Arm64JITCore::Arm64JITCore(FEXCore::Context::ContextImpl* ctx, FEXCore::Core::In
Ptrs.LDIV = reinterpret_cast<uint64_t>(LDIV);
}
CurrentCodeBuffer = CodeBuffers.GetLatest();
CurrentCodeBuffer = SharedCodeBuffers.GetLatest();
ThreadState->LookupCache->Shared = CurrentCodeBuffer->LookupCache.get();
}
void Arm64JITCore::EmitDetectionString() {
const char JITString[] = "FEXJIT::Arm64JITCore::";
EmitString(JITString);
Align();
}
void Arm64JITCore::ClearCache() {
// NOTE: Holding on to the reference here is required to ensure validity of the WriteLock mutex
auto PrevCodeBuffer = CurrentCodeBuffer;
auto lk = PrevCodeBuffer->LookupCache->AcquireWriteLock();
auto CodeBuffer = GetEmptyCodeBuffer();
SetBuffer(CodeBuffer->Ptr, CodeBuffer->AllocatedSize);
EmitDetectionString();
ThreadState->LookupCache->ChangeGuestToHostMapping(*PrevCodeBuffer, *CurrentCodeBuffer->LookupCache, lk);
auto CodeBuffer = AcquireNewSharedCodeBuffer();
ThreadState->LookupCache->ChangeGuestToHostMapping(*PrevCodeBuffer, *CodeBuffer->LookupCache, lk);
}
Arm64JITCore::~Arm64JITCore() {}
@@ -822,6 +813,33 @@ void Arm64JITCore::EmitEntryPoint(ARMEmitter::BackwardLabel& HeaderLabel, bool C
EmitSuspendInterruptCheck();
}
CodeBuffer::CodeBufferAllocation Arm64JITCore::AllocateCodeBufferInSharedCache(size_t Size) {
CodeBuffer::CodeBufferAllocation AllocatedInfo {};
LOGMAN_THROW_A_FMT(CurrentCodeBuffer->LookupCache.get() == ThreadState->LookupCache->Shared, "INVARIANT VIOLATED: SharedLookupCache "
"doesn't match up!\n");
// Bring CodeBuffer up to date
if (auto Prev = CheckCodeBufferUpdate()) {
Allocator::VirtualDontNeed(ThreadState->CallRetStackBase, FEXCore::Core::InternalThreadState::CALLRET_STACK_SIZE);
auto lk = ThreadState->LookupCache->AcquireWriteLock();
ThreadState->LookupCache->ChangeGuestToHostMapping(*Prev, *CurrentCodeBuffer->LookupCache, lk);
}
// Attempt to allocate a buffer from the SharedCodeBuffers.
while (AllocatedInfo.BufferAllocationOffset == nullptr) {
AllocatedInfo = CurrentCodeBuffer->AtomicAllocateBuffer(Size);
if (AllocatedInfo.BufferAllocationOffset == nullptr) {
// If it didn't fit then clear the buffer and try again.
// This has the possibility of migrating the SharedCodeBuffer. See above in `Arm64JITCore::ClearCache()`
CTX->ClearCodeCache(ThreadState);
continue;
}
}
return AllocatedInfo;
}
CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size, bool SingleInst, const FEXCore::IR::IRListView* IR,
FEXCore::Core::DebugData* DebugData, bool CheckTF) {
FEXCORE_PROFILE_SCOPED("Arm64::CompileCode");
@@ -843,7 +861,7 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
case RestartOptions::Control::EnableFarARM64Jumps: RequiresFarARM64Jumps = true; break;
case RestartOptions::Control::NeedsLargerJITSpace:
// Get rid of the claimed buffer immediately, we can't fit in it at all.
TempAllocator.UnclaimBuffer();
TempCodeBufferAllocator.UnclaimBuffer();
SSANodeMultiplier *= 2;
break;
default: LOGMAN_MSG_A_FMT("Unhandled Arm64 restart condition!");
@@ -864,7 +882,7 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
// JIT output is first written to a temporary buffer and later relocated to the CodeBuffer.
// This minimizes lock contention of CodeBufferWriteMutex.
auto TempCodeBufferInfo = TempAllocator.ReownOrClaimBufferWithSize(DesiredBufferRange);
auto TempCodeBufferInfo = TempCodeBufferAllocator.ReownOrClaimBufferWithSize(DesiredBufferRange);
auto TempCodeBuffer = TempCodeBufferInfo.Ptr;
const uint32_t UsableBufferRange = TempCodeBufferInfo.Size - FEXCore::Utils::FEX_PAGE_SIZE;
@@ -874,6 +892,7 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
ThreadState->JITGuardOverflowArgument = FEXCore::ToUnderlying(RestartOptions::Control::NeedsLargerJITSpace);
CodeData.BlockBegin = GetCursorAddress<uint8_t*>();
LOGMAN_THROW_A_FMT(GetCursorOffset() == 0, "Needs to be zero");
// Put the code header at the start of the data block.
ARMEmitter::BackwardLabel JITCodeHeaderLabel {};
@@ -909,7 +928,6 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
PendingCallReturnTargetLabel = nullptr;
for (auto [BlockNode, BlockHeader] : IR->GetBlocks()) {
using namespace FEXCore::IR;
auto BlockIROp = BlockHeader->CW<FEXCore::IR::IROp_CodeBlock>();
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
LOGMAN_THROW_A_FMT(BlockIROp->Header.Op == IR::OP_CODEBLOCK, "IR type failed to be a code block");
@@ -1001,9 +1019,15 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
// This is a ExitFunctionLinkData struct
BindOrRestart(&l_ExitLink);
dc64(0); // HostCode
PlaceNamedSymbolLiteral(InsertGuestRIPLiteral(PendingJumpThunk.GuestRIP)); // GuestRIP
dc64(PendingJumpThunk.CallerAddress - ThunkAddress); // CallerOffset
dc64(0); // HostCode
if (PendingJumpThunk.PatchSiteAddress) {
// GuestRIP with an extra step
PlaceNamedSymbolLiteral(
InsertGuestPatchableRIPLiteral(PendingJumpThunk.GuestRIP, PendingJumpThunk.PatchSiteAddress, PendingJumpThunk.PatchSiteSize));
} else {
PlaceNamedSymbolLiteral(InsertGuestRIPLiteral(PendingJumpThunk.GuestRIP)); // GuestRIP
}
dc64(PendingJumpThunk.CallerAddress - ThunkAddress); // CallerOffset
}
BindOrRestart(&l_ExitLink);
@@ -1063,69 +1087,47 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
}
SetCursorOffset(JITRIPEntriesLocation - CodeData.BlockBegin);
Align();
// Make sure code is 16B aligned on the tail.
// Can't use Align16B here as vl64pair can cause non-4byte alignment.
Align(16);
CodeData.Size = GetCursorAddress<uint8_t*>() - CodeData.BlockBegin;
// Beginning of emission is guaranteed to be offset zero. So the code data size is just the current cursor offset.
CodeData.Size = GetCursorOffset();
// Finalize and write block tail data
JITBlockTail.Size = CodeData.Size;
{
auto PrevCur = GetCursorOffset();
memcpy(JITBlockTailLocation, &JITBlockTail, sizeof(JITBlockTail));
SetCursorOffset(JITBlockTailLocation - CodeData.BlockBegin + offsetof(JITCodeTail, RIP));
PlaceNamedSymbolLiteral(InsertGuestRIPLiteral(JITBlockTail.RIP));
SetCursorOffset(PrevCur);
// Emitter buffer is no longer used, guard against misuse by setting to nullptr.
SetBuffer(nullptr, 0);
}
// Migrate the compile output from temporary storage to the actual CodeBuffer.
// This can block progress in other compiling threads, so the duration of the lock should be as small as possible.
{
auto CodeBufferLock = std::unique_lock {CodeBuffers.CodeBufferWriteMutex};
LOGMAN_THROW_A_FMT(CodeData.Size % 16 == 0, "Needs to be 16B aligned!");
// Query size of generated code
const auto TempSize = GetCursorOffset();
// Bring CodeBuffer up to date
{
LOGMAN_THROW_A_FMT(CurrentCodeBuffer->LookupCache.get() == ThreadState->LookupCache->Shared, "INVARIANT VIOLATED: SharedLookupCache "
"doesn't match up!\n");
if (auto Prev = CheckCodeBufferUpdate()) {
Allocator::VirtualDontNeed(ThreadState->CallRetStackBase, FEXCore::Core::InternalThreadState::CALLRET_STACK_SIZE);
auto lk = ThreadState->LookupCache->AcquireWriteLock();
ThreadState->LookupCache->ChangeGuestToHostMapping(*Prev, *CurrentCodeBuffer->LookupCache, lk);
}
// NOTE: 16-byte alignment of the new cursor offset must be preserved for block linking records
SetBuffer(CurrentCodeBuffer->Ptr, CurrentCodeBuffer->AllocatedSize);
SetCursorOffset(CodeBuffers.LatestOffset);
Align16B();
if ((GetCursorOffset() + TempSize) > CurrentCodeBuffer->UsableSize()) {
CTX->ClearCodeCache(ThreadState);
}
CodeBuffers.LatestOffset = GetCursorOffset();
}
auto AllocatedInfo = AllocateCodeBufferInSharedCache(CodeData.Size);
// NOTE: 16-byte alignment of the new cursor offset must be preserved for block linking records
LOGMAN_THROW_A_FMT((reinterpret_cast<uintptr_t>(AllocatedInfo.BufferAllocationOffset) % 16) == 0, "Allocated buffer wasn't 16B "
"aligned?");
// Adjust host addresses
const auto Delta = GetCursorAddress<uint8_t*>() - CodeData.BlockBegin;
const auto Delta = AllocatedInfo.BufferAllocationOffset - CodeData.BlockBegin;
CodeData.BlockBegin += Delta;
for (auto& EntryPoint : CodeData.EntryPoints) {
EntryPoint.second += Delta;
}
CodeBegin += Delta;
for (std::size_t Idx = PrevNumAllocations; Idx != Relocations.size(); ++Idx) {
Relocations[Idx].Header.Offset += CodeBuffers.LatestOffset;
}
CodeData.HostCodeOffset = CodeData.BlockBegin - CurrentCodeBuffer->GetBufferBase();
// Copy over CodeBuffer contents
memcpy(GetCursorAddress<uint8_t*>(), TempCodeBuffer, TempSize);
SetCursorOffset(CodeBuffers.LatestOffset + TempSize);
CodeBuffers.LatestOffset = GetCursorOffset();
memcpy(AllocatedInfo.BufferAllocationOffset, TempCodeBuffer, CodeData.Size);
}
TempAllocator.DelayedDisownBuffer();
TempCodeBufferAllocator.DelayedDisownBuffer();
ClearICache(CodeBegin, CodeOnlySize);
@@ -1161,6 +1163,22 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
return std::move(CodeData);
}
CPUBackend::CompiledCode Arm64JITCore::LoadCachedCode(std::span<const uint8_t> HostBytes) {
// we stored it aligned, better still be?
LOGMAN_THROW_A_FMT(HostBytes.size() % 16 == 0, "Needs to be 16B aligned!");
auto AllocatedInfo = AllocateCodeBufferInSharedCache(HostBytes.size());
uint8_t* Dest = AllocatedInfo.BufferAllocationOffset;
memcpy(Dest, HostBytes.data(), HostBytes.size());
ClearICache(Dest, HostBytes.size());
CPUBackend::CompiledCode Result;
Result.BlockBegin = Dest;
Result.Size = HostBytes.size();
Result.HostCodeOffset = Dest - CurrentCodeBuffer->GetBufferBase();
return Result;
}
void Arm64JITCore::ResetStack() {
if (SpillSlots == 0) {
return;
+18 -5
View File
@@ -54,6 +54,9 @@ public:
CPUBackend::CompiledCode CompileCode(uint64_t Entry, uint64_t Size, bool SingleInst, const FEXCore::IR::IRListView* IR,
FEXCore::Core::DebugData* DebugData, bool CheckTF) override;
[[nodiscard]]
CPUBackend::CompiledCode LoadCachedCode(std::span<const uint8_t> HostBytes) override;
void ClearCache() override;
void ClearRelocations() override {
@@ -102,10 +105,12 @@ private:
uint64_t CallerAddress;
uint64_t GuestRIP;
ARMEmitter::ForwardLabel Label;
uint64_t PatchSiteAddress = 0;
uint8_t PatchSiteSize = 0;
};
fextl::vector<PendingJumpThunk> PendingJumpThunks;
Utils::PoolBufferWithTimedRetirement<uint8_t*, 5000, 500> TempAllocator;
Utils::PoolBufferWithTimedRetirement<uint8_t*, 5000, 500> TempCodeBufferAllocator;
static uint64_t ExitFunctionLink(FEXCore::Core::CpuStateFrame* Frame, FEXCore::Context::ExitFunctionLinkData* Record);
@@ -342,8 +347,8 @@ private:
uint32_t End;
};
void EmitLinkedBranch(uint64_t GuestRIP, bool Call) {
PendingJumpThunks.push_back({GetCursorAddress<uint64_t>(), GuestRIP, {}});
void EmitLinkedBranch(uint64_t GuestRIP, bool Call, uint64_t PatchSiteAddress = 0, uint8_t PatchSiteSize = 0) {
PendingJumpThunks.push_back({GetCursorAddress<uint64_t>(), GuestRIP, {}, PatchSiteAddress, PatchSiteSize});
auto& Thunk = PendingJumpThunks.back();
BindOrRestart(&Thunk.Label);
if (Call) {
@@ -526,8 +531,6 @@ private:
FEXCore::UncheckedLongJump::LongJump(ThreadState->RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
// This is purely a debugging aid for developers to see if they are in JIT code space when inspecting raw memory
void EmitDetectionString();
IR::RegisterAllocationPass* RAPass {};
FEXCore::Core::DebugData* DebugData {};
@@ -561,6 +564,9 @@ private:
*/
void InsertGuestRIPMove(ARMEmitter::Register Reg, uint64_t Constant);
void InsertGuestPatchableDataMove(ARMEmitter::Register Reg, uint64_t Value, uint64_t SiteAddress, uint8_t ValueSize);
void InsertGuestPatchableRIPMove(ARMEmitter::Register Reg, uint64_t Value, uint64_t SiteAddress, uint8_t ValueSize);
/**
* @brief Inserts a named symbol as a literal in memory
*
@@ -580,6 +586,11 @@ private:
*/
NamedSymbolLiteralPair InsertGuestRIPLiteral(uint64_t GuestRIP);
/**
* @brief Like InsertGuestRIPLiteral, but with patch information to recompute value from live guest bytes at cache load time
*/
NamedSymbolLiteralPair InsertGuestPatchableRIPLiteral(uint64_t GuestRIP, uint64_t SiteAddress, uint8_t ValueSize);
/**
* @brief Place the named symbol literal relocation in memory
*
@@ -626,6 +637,8 @@ private:
void EmitEntryPoint(ARMEmitter::BackwardLabel& HeaderLabel, bool CheckTF);
[[nodiscard]] CodeBuffer::CodeBufferAllocation AllocateCodeBufferInSharedCache(size_t Size);
#define DEF_OP(x) void Op_##x(IR::IROp_Header const* IROp, IR::Ref Node)
///< Unhandled handler
+10 -10
View File
@@ -267,7 +267,7 @@ DEF_OP(LoadContextIndexed) {
ldr(Dst.Q(), TMP1, Op->BaseOffset);
} else {
add(ARMEmitter::Size::i64Bit, TMP1, TMP1, Op->BaseOffset);
ldur(Dst.Q(), TMP1, Op->BaseOffset);
ldur(Dst.Q(), TMP1);
}
break;
case IR::OpSize::i256Bit:
@@ -333,7 +333,7 @@ DEF_OP(StoreContextIndexed) {
str(Value.Q(), TMP1, Op->BaseOffset);
} else {
add(ARMEmitter::Size::i64Bit, TMP1, TMP1, Op->BaseOffset);
stur(Value.Q(), TMP1, Op->BaseOffset);
stur(Value.Q(), TMP1);
}
break;
case IR::OpSize::i256Bit:
@@ -2401,13 +2401,13 @@ DEF_OP(CacheLineClear) {
// Clear dcache only
// icache doesn't matter here since the guest application shouldn't be calling clflush on JIT code.
// check host cacheline size again x86_64 size to ensure at least 64 bytes are cleaned
if (CTX->HostFeatures.DCacheLineSize >= 64U) {
if (CTX->HostFeatures.DCacheSize() >= 64U) {
dc(ARMEmitter::DataCacheOperation::CIVAC, MemReg);
} else {
auto CurrentWorkingReg = MemReg.X();
for (size_t i = 0; i < std::max(1U, 64U / CTX->HostFeatures.DCacheLineSize); ++i) {
dc(ARMEmitter::DataCacheOperation::CIVAC, TMP1);
add(ARMEmitter::Size::i64Bit, TMP1, CurrentWorkingReg, CTX->HostFeatures.DCacheLineSize);
for (size_t i = 0; i < std::max(1U, 64U / CTX->HostFeatures.DCacheSize()); ++i) {
dc(ARMEmitter::DataCacheOperation::CIVAC, CurrentWorkingReg);
add(ARMEmitter::Size::i64Bit, TMP1, CurrentWorkingReg, CTX->HostFeatures.DCacheSize());
CurrentWorkingReg = TMP1;
}
}
@@ -2430,13 +2430,13 @@ DEF_OP(CacheLineClean) {
// Clean dcache only
// check host cacheline size again x86_64 size to ensure at least 64 bytes are cleaned
if (CTX->HostFeatures.DCacheLineSize >= 64U) {
if (CTX->HostFeatures.DCacheSize() >= 64U) {
dc(ARMEmitter::DataCacheOperation::CVAC, MemReg);
} else {
auto CurrentWorkingReg = MemReg.X();
for (size_t i = 0; i < std::max(1U, 64U / CTX->HostFeatures.DCacheLineSize); ++i) {
dc(ARMEmitter::DataCacheOperation::CVAC, TMP1);
add(ARMEmitter::Size::i64Bit, TMP1, CurrentWorkingReg, CTX->HostFeatures.DCacheLineSize);
for (size_t i = 0; i < std::max(1U, 64U / CTX->HostFeatures.DCacheSize()); ++i) {
dc(ARMEmitter::DataCacheOperation::CVAC, CurrentWorkingReg);
add(ARMEmitter::Size::i64Bit, TMP1, CurrentWorkingReg, CTX->HostFeatures.DCacheSize());
CurrentWorkingReg = TMP1;
}
}
@@ -168,7 +168,7 @@ DEF_OP(PushRoundingMode) {
} else {
LOGMAN_THROW_A_FMT(Op->RoundMode == 1 || Op->RoundMode == 2, "expect a valid round mode");
and_(ARMEmitter::Size::i64Bit, TMP1, Dest, ~(Op->RoundMode << 22));
and_(ARMEmitter::Size::i64Bit, TMP1, Dest, ~(3 << 22));
orr(ARMEmitter::Size::i64Bit, TMP1, TMP1, (Op->RoundMode == 2 ? 1 : 2) << 22);
}
@@ -282,7 +282,7 @@ DEF_OP(ProcessorID) {
// Load the values returned by the kernel
ldp<ARMEmitter::IndexType::OFFSET>(ARMEmitter::WReg::w0, ARMEmitter::WReg::w1, ARMEmitter::Reg::rsp);
// Deallocate stack space
sub(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, ARMEmitter::Reg::rsp, 16);
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, ARMEmitter::Reg::rsp, 16);
// Now that we are done in the syscall we need to carefully peel back the state
// First unspill the registers from before
@@ -25,6 +25,18 @@ enum class RelocationTypes : uint32_t {
// 4 instruction constant generation
// Aligned to struct RelocGuestRIP
RELOC_GUEST_RIP_MOVE,
// The frontend flagged those regions as patchable by the disk cache
// Aligned to struct RelocGuestPatchableData
RELOC_GUEST_PATCHABLE_DATA_MOVE,
// Same as GuestRipLiteral but patchable
// Aligned to struct RelocGuestPatchableData
RELOC_GUEST_PATCHABLE_RIP_LITERAL,
// Like PATCHABLE_RIP_LITERAL but puts it in a register
// Aligned to struct RelocGuestPatchableData
RELOC_GUEST_PATCHABLE_RIP_MOVE,
};
struct FEX_PACKED RelocationHeader final {
@@ -73,6 +85,20 @@ struct RelocGuestRIP final {
uint32_t pad2[6] {};
};
struct RelocGuestPatchableData final {
RelocationHeader Header {};
uint8_t RegisterIndex;
uint8_t ValueSize;
char Pad[2];
uint64_t SiteAddress;
uint32_t pad2[6] {};
};
union Relocation {
// Clang 16 Can't default-initialize this union
static Relocation Default() {
@@ -93,6 +119,8 @@ union Relocation {
RelocNamedThunkMove NamedThunkMove;
RelocGuestRIP GuestRIP;
RelocGuestPatchableData GuestPatchableData;
};
uint64_t GetNamedSymbolLiteral(FEXCore::Context::ContextImpl&, RelocNamedSymbolLiteral::NamedSymbol);
+213 -121
View File
@@ -1139,9 +1139,16 @@ DEF_OP(VOrn) {
const auto Vector2 = GetVReg(Op->Vector2);
if (HostSupportsSVE256 && Is256Bit) {
const auto Pred = PRED_TMP_32B.Merging();
not_(ARMEmitter::SubRegSize::i8Bit, VTMP1.Z(), Pred, Vector2.Z());
orr(Dst.Z(), Vector1.Z(), VTMP1.Z());
if (Dst == Vector1) {
bsl2n(Dst.Z(), Dst.Z(), Vector2.Z(), Dst.Z());
} else if (Dst == Vector2) {
const auto Pred = PRED_TMP_32B.Merging();
not_(ARMEmitter::SubRegSize::i8Bit, Dst.Z(), Pred, Dst.Z());
orr(Dst.Z(), Vector1.Z(), Dst.Z());
} else {
movprfx(Dst.Z(), Vector1.Z());
bsl2n(Dst.Z(), Dst.Z(), Vector2.Z(), Vector1.Z());
}
} else if (Is128Bit) {
orn(Dst.Q(), Vector1.Q(), Vector2.Q());
} else {
@@ -1165,8 +1172,7 @@ DEF_OP(VFAddV) {
if (HostSupportsSVE256 && Is256Bit) {
const auto Pred = PRED_TMP_32B.Merging();
faddv(SubRegSize.Vector, Dst, Pred, Vector.Z());
}
if (HostSupportsSVE128) {
} else if (HostSupportsSVE128) {
const auto Pred = PRED_TMP_16B.Merging();
faddv(SubRegSize.Vector, Dst, Pred, Vector.Z());
} else {
@@ -1193,20 +1199,16 @@ DEF_OP(VAddV) {
const auto Vector = GetVReg(Op->Vector);
if (HostSupportsSVE256 && Is256Bit) {
// SVE doesn't have an equivalent ADDV instruction, so we make do
// by performing two Adv. SIMD ADDV operations on the high and low
// 128-bit lanes and then sum them up.
const auto Mask = PRED_TMP_32B.Zeroing();
const auto CompactPred = ARMEmitter::PReg::p0;
// Select all our upper elements to run ADDV over them.
not_(CompactPred, Mask, PRED_TMP_16B);
compact(ARMEmitter::SubRegSize::i64Bit, VTMP1.Z(), CompactPred, Vector.Z());
addv(SubRegSize.Vector, VTMP2.Q(), Vector.Q());
addv(SubRegSize.Vector, VTMP1.Q(), VTMP1.Q());
add(SubRegSize.Vector, Dst.Q(), VTMP1.Q(), VTMP2.Q());
if (ElementSize == IR::OpSize::i64Bit) {
const auto Mask = PRED_TMP_32B.Zeroing();
uaddv(SubRegSize.Vector, Dst.D(), Mask, Vector.Z());
} else {
const auto Mask = ARMEmitter::PReg::p0;
uaddv(SubRegSize.Vector, VTMP1.D(), Mask, Vector.Z());
mov_imm(ARMEmitter::SubRegSize::i64Bit, Dst.Z(), 0);
ptrue(SubRegSize.Vector, Mask, ARMEmitter::PredicatePattern::SVE_VL1);
mov(SubRegSize.Vector, Dst.Z(), Mask.Merging(), VTMP1.Z());
}
} else {
if (ElementSize == IR::OpSize::i64Bit) {
addp(SubRegSize.Scalar, Dst, Vector);
@@ -1295,6 +1297,7 @@ DEF_OP(VFAddP) {
const auto Op = IROp->C<IR::IROp_VFAddP>();
const auto OpSize = IROp->Size;
const auto IsScalar = OpSize == IR::OpSize::i64Bit;
const auto Is256Bit = OpSize == IR::OpSize::i256Bit;
LOGMAN_THROW_A_FMT(!Is256Bit || HostSupportsSVE256, "Need SVE256 support in order to use {} with 256-bit operation", __func__);
@@ -1325,6 +1328,8 @@ DEF_OP(VFAddP) {
// Merge upper half with lower half.
splice<ARMEmitter::OpType::Destructive>(ARMEmitter::SubRegSize::i64Bit, Dst.Z(), PRED_TMP_16B, Dst.Z(), VTMP2.Z());
} else if (IsScalar) {
faddp(SubRegSize, Dst.D(), VectorLower.D(), VectorUpper.D());
} else {
faddp(SubRegSize, Dst.Q(), VectorLower.Q(), VectorUpper.Q());
}
@@ -2785,17 +2790,17 @@ DEF_OP(VUShrSWide) {
const auto Vector = GetVReg(Op->Vector);
if (HostSupportsSVE256 && Is256Bit) {
const auto Mask = PRED_TMP_32B.Merging();
dup(ARMEmitter::SubRegSize::i64Bit, VTMP1.Z(), ShiftScalar.Z(), 0);
if (Dst != Vector) {
// NOTE: SVE LSR is a destructive operation.
movprfx(Dst.Z(), Vector.Z());
}
if (ElementSize == IR::OpSize::i64Bit) {
const auto Mask = PRED_TMP_32B.Merging();
if (Dst != Vector) {
// NOTE: SVE LSR is a destructive operation.
movprfx(Dst.Z(), Vector.Z());
}
lsr(SubRegSize, Dst.Z(), Mask, Dst.Z(), VTMP1.Z());
} else {
lsr_wide(SubRegSize, Dst.Z(), Mask, Dst.Z(), VTMP1.Z());
lsr_wide(SubRegSize, Dst.Z(), Vector.Z(), VTMP1.Z());
}
} else if (HostSupportsSVE128) {
const auto Mask = PRED_TMP_16B.Merging();
@@ -2851,17 +2856,17 @@ DEF_OP(VSShrSWide) {
const auto Vector = GetVReg(Op->Vector);
if (HostSupportsSVE256 && Is256Bit) {
const auto Mask = PRED_TMP_32B.Merging();
dup(ARMEmitter::SubRegSize::i64Bit, VTMP1.Z(), ShiftScalar.Z(), 0);
if (Dst != Vector) {
// NOTE: SVE LSR is a destructive operation.
movprfx(Dst.Z(), Vector.Z());
}
if (ElementSize == IR::OpSize::i64Bit) {
const auto Mask = PRED_TMP_32B.Merging();
if (Dst != Vector) {
// NOTE: SVE LSR is a destructive operation.
movprfx(Dst.Z(), Vector.Z());
}
asr(SubRegSize, Dst.Z(), Mask, Dst.Z(), VTMP1.Z());
} else {
asr_wide(SubRegSize, Dst.Z(), Mask, Dst.Z(), VTMP1.Z());
asr_wide(SubRegSize, Dst.Z(), Vector.Z(), VTMP1.Z());
}
} else if (HostSupportsSVE128) {
const auto Mask = PRED_TMP_16B.Merging();
@@ -2917,17 +2922,17 @@ DEF_OP(VUShlSWide) {
const auto Vector = GetVReg(Op->Vector);
if (HostSupportsSVE256 && Is256Bit) {
const auto Mask = PRED_TMP_32B.Merging();
dup(ARMEmitter::SubRegSize::i64Bit, VTMP1.Z(), ShiftScalar.Z(), 0);
if (Dst != Vector) {
// NOTE: SVE LSR is a destructive operation.
movprfx(Dst.Z(), Vector.Z());
}
if (ElementSize == IR::OpSize::i64Bit) {
const auto Mask = PRED_TMP_32B.Merging();
if (Dst != Vector) {
// NOTE: SVE LSR is a destructive operation.
movprfx(Dst.Z(), Vector.Z());
}
lsl(SubRegSize, Dst.Z(), Mask, Dst.Z(), VTMP1.Z());
} else {
lsl_wide(SubRegSize, Dst.Z(), Mask, Dst.Z(), VTMP1.Z());
lsl_wide(SubRegSize, Dst.Z(), Vector.Z(), VTMP1.Z());
}
} else if (HostSupportsSVE128) {
const auto Mask = PRED_TMP_16B.Merging();
@@ -3175,19 +3180,12 @@ DEF_OP(VUShrI) {
movi(ARMEmitter::SubRegSize::i64Bit, Dst.Q(), 0);
} else {
if (HostSupportsSVE256 && Is256Bit) {
const auto Mask = PRED_TMP_32B.Merging();
if (BitShift == 0) {
if (Dst != Vector) {
mov(Dst.Z(), Vector.Z());
}
} else {
// SVE LSR is destructive, so lets set up the destination if
// Vector doesn't already alias it.
if (Dst != Vector) {
movprfx(Dst.Z(), Vector.Z());
}
lsr(SubRegSize, Dst.Z(), Mask, Dst.Z(), BitShift);
lsr(SubRegSize, Dst.Z(), Vector.Z(), BitShift);
}
} else {
if (BitShift == 0) {
@@ -3201,48 +3199,6 @@ DEF_OP(VUShrI) {
}
}
DEF_OP(VUShraI) {
const auto Op = IROp->C<IR::IROp_VUShraI>();
const auto OpSize = IROp->Size;
const auto BitShift = Op->BitShift;
const auto SubRegSize = ConvertSubRegSize8(IROp);
const auto Is256Bit = OpSize == IR::OpSize::i256Bit;
LOGMAN_THROW_A_FMT(!Is256Bit || HostSupportsSVE256, "Need SVE256 support in order to use {} with 256-bit operation", __func__);
const auto Dst = GetVReg(Node);
const auto DestVector = GetVReg(Op->DestVector);
const auto Vector = GetVReg(Op->Vector);
if (HostSupportsSVE256 && Is256Bit) {
if (Dst == DestVector) {
usra(SubRegSize, Dst.Z(), Vector.Z(), BitShift);
} else {
if (Dst != Vector) {
mov(Dst.Z(), DestVector.Z());
usra(SubRegSize, Dst.Z(), Vector.Z(), BitShift);
} else {
mov(VTMP1.Z(), DestVector.Z());
usra(SubRegSize, Dst.Z(), Vector.Z(), BitShift);
mov(Dst.Z(), VTMP1.Z());
}
}
} else {
if (Dst == DestVector) {
usra(SubRegSize, Dst.Q(), Vector.Q(), BitShift);
} else {
if (Dst != Vector) {
mov(Dst.Q(), DestVector.Q());
usra(SubRegSize, Dst.Q(), Vector.Q(), BitShift);
} else {
mov(VTMP1.Q(), DestVector.Q());
usra(SubRegSize, VTMP1.Q(), Vector.Q(), BitShift);
mov(Dst.Q(), VTMP1.Q());
}
}
}
}
DEF_OP(VSShrI) {
const auto Op = IROp->C<IR::IROp_VSShrI>();
const auto OpSize = IROp->Size;
@@ -3258,19 +3214,12 @@ DEF_OP(VSShrI) {
const auto Vector = GetVReg(Op->Vector);
if (HostSupportsSVE256 && Is256Bit) {
const auto Mask = PRED_TMP_32B.Merging();
if (Shift == 0) {
if (Dst != Vector) {
mov(Dst.Z(), Vector.Z());
}
} else {
// SVE ASR is destructive, so lets set up the destination if
// Vector doesn't already alias it.
if (Dst != Vector) {
movprfx(Dst.Z(), Vector.Z());
}
asr(SubRegSize, Dst.Z(), Mask, Dst.Z(), Shift);
asr(SubRegSize, Dst.Z(), Vector.Z(), Shift);
}
} else {
if (Shift == 0) {
@@ -3300,19 +3249,12 @@ DEF_OP(VShlI) {
movi(ARMEmitter::SubRegSize::i64Bit, Dst.Q(), 0);
} else {
if (HostSupportsSVE256 && Is256Bit) {
const auto Mask = PRED_TMP_32B.Merging();
if (BitShift == 0) {
if (Dst != Vector) {
mov(Dst.Z(), Vector.Z());
}
} else {
// SVE LSL is destructive, so lets set up the destination if
// Vector doesn't already alias it.
if (Dst != Vector) {
movprfx(Dst.Z(), Vector.Z());
}
lsl(SubRegSize, Dst.Z(), Mask, Dst.Z(), BitShift);
lsl(SubRegSize, Dst.Z(), Vector.Z(), BitShift);
}
} else {
if (BitShift == 0) {
@@ -3339,8 +3281,13 @@ DEF_OP(VUShrNI) {
const auto Vector = GetVReg(Op->Vector);
if (HostSupportsSVE256 && Is256Bit) {
shrnb(SubRegSize, Dst.Z(), Vector.Z(), BitShift);
uzp1(SubRegSize, Dst.Z(), Dst.Z(), Dst.Z());
if (BitShift == 0) {
mov_imm(ARMEmitter::SubRegSize::i64Bit, VTMP1.Z(), 0);
uzp1(SubRegSize, Dst.Z(), Dst.Z(), VTMP1.Z());
} else {
shrnb(SubRegSize, Dst.Z(), Vector.Z(), BitShift);
uzp1(SubRegSize, Dst.Z(), Dst.Z(), Dst.Z());
}
} else {
if (BitShift == 0) {
xtn(SubRegSize, Dst.D(), Vector.D());
@@ -3388,6 +3335,55 @@ DEF_OP(VUShrNI2) {
}
}
DEF_OP(VRSHRN) {
const auto Op = IROp->C<IR::IROp_VRSHRN>();
const auto OpSize = IROp->Size;
const auto BitShift = Op->BitShift;
const auto SubRegSize = ConvertSubRegSize4(IROp);
const auto Is256Bit = OpSize == IR::OpSize::i256Bit;
LOGMAN_THROW_A_FMT(!Is256Bit || HostSupportsSVE256, "Need SVE256 support in order to use {} with 256-bit operation", __func__);
const auto Dst = GetVReg(Node);
const auto Vector = GetVReg(Op->Vector);
if (HostSupportsSVE256 && Is256Bit) {
rshrnb(SubRegSize, Dst.Z(), Vector.Z(), BitShift);
uzp1(SubRegSize, Dst.Z(), Dst.Z(), Dst.Z());
} else {
rshrn(SubRegSize, Dst.D(), Vector.D(), BitShift);
}
}
DEF_OP(VRSHRNPair) {
const auto Op = IROp->C<IR::IROp_VRSHRNPair>();
const auto OpSize = IROp->Size;
const auto BitShift = Op->BitShift;
const auto SubRegSize = ConvertSubRegSize4(IROp);
const auto Is256Bit = OpSize == IR::OpSize::i256Bit;
LOGMAN_THROW_A_FMT(!Is256Bit || HostSupportsSVE256, "Need SVE256 support in order to use {} with 256-bit operation", __func__);
const auto Dst = GetVReg(Node);
const auto VectorLower = GetVReg(Op->VectorLower);
auto VectorUpper = GetVReg(Op->VectorUpper);
if (HostSupportsSVE256 && Is256Bit) {
rshrnb(SubRegSize, VTMP1.Z(), VectorLower.Z(), BitShift);
rshrnb(SubRegSize, VTMP2.Z(), VectorUpper.Z(), BitShift);
uzp1(SubRegSize, Dst.Z(), VTMP1.Z(), VTMP2.Z());
} else {
if (Dst == VectorUpper) {
// RSHRN writes the lower half and would destroy the upper input.
mov(VTMP1.Q(), VectorUpper.Q());
VectorUpper = VTMP1;
}
rshrn(SubRegSize, Dst.D(), VectorLower.D(), BitShift);
rshrn2(SubRegSize, Dst.Q(), VectorUpper.Q(), BitShift);
}
}
DEF_OP(VSXTL) {
const auto Op = IROp->C<IR::IROp_VSXTL>();
const auto OpSize = IROp->Size;
@@ -3593,9 +3589,13 @@ DEF_OP(VSQXTN2) {
mov(Dst.Q(), VectorLower.Q());
ins(ARMEmitter::SubRegSize::i32Bit, Dst, 1, VTMP2, 0);
} else {
mov(VTMP1.Q(), VectorLower.Q());
sqxtn2(SubRegSize, VTMP1, VectorUpper);
mov(Dst.Q(), VTMP1.Q());
if (Dst == VectorLower) {
sqxtn2(SubRegSize, VectorLower, VectorUpper);
} else {
mov(VTMP1.Q(), VectorLower.Q());
sqxtn2(SubRegSize, VTMP1, VectorUpper);
mov(Dst.Q(), VTMP1.Q());
}
}
}
}
@@ -4451,13 +4451,9 @@ DEF_OP(VFMLS) {
if (Is128Bit) {
fneg(SubRegSize, DestTmp.Q(), VectorAddend.Q());
} else {
fneg(SubRegSize, DestTmp.D(), VectorAddend.D());
}
if (Is128Bit) {
fmla(SubRegSize, DestTmp.Q(), Vector1.Q(), Vector2.Q());
} else {
fneg(SubRegSize, DestTmp.D(), VectorAddend.D());
fmla(SubRegSize, DestTmp.D(), Vector1.D(), Vector2.D());
}
@@ -4611,13 +4607,9 @@ DEF_OP(VFNMLS) {
if (Is128Bit) {
fneg(SubRegSize, DestTmp.Q(), VectorAddend.Q());
} else {
fneg(SubRegSize, DestTmp.D(), VectorAddend.D());
}
if (Is128Bit) {
fmls(SubRegSize, DestTmp.Q(), Vector1.Q(), Vector2.Q());
} else {
fneg(SubRegSize, DestTmp.D(), VectorAddend.D());
fmls(SubRegSize, DestTmp.D(), Vector1.D(), Vector2.D());
}
@@ -4631,6 +4623,106 @@ DEF_OP(VFNMLS) {
}
}
DEF_OP(VBlendImm) {
LOGMAN_THROW_A_FMT(HostSupportsSVE128 || HostSupportsSVE256, "Host must support SVE to use {}", __func__);
auto Op = IROp->C<IR::IROp_VBlendImm>();
const auto OpSize = IROp->Size;
const auto Is256Bit = OpSize == IR::OpSize::i256Bit;
const auto SubRegSize = ConvertSubRegSize8(IROp);
const auto ElementSize = IROp->ElementSize;
const auto Selector = Op->Selector;
const auto GoverningPredicate = Is256Bit ? PRED_TMP_32B : PRED_TMP_16B;
const auto Dst = GetVReg(Node);
const auto LHS = GetVReg(Op->LHS);
const auto RHS = GetVReg(Op->RHS);
const auto DstIsNonAliasing = Dst != LHS && Dst != RHS;
// Silly case where two blending sources are the same.
if (LHS == RHS) {
if (DstIsNonAliasing) {
mov(SubRegSize, Dst.Z(), GoverningPredicate.Merging(), LHS.Z());
}
return;
}
// We'll need to expand our selector to match its predicate equivalent.
// The lowest bit of each predicate element being set to 1 signifies
// that it's enabled.
const auto MakePredicateMask = [ElementSize, Is256Bit, OpSize](uint16_t Imm) {
if (ElementSize == IR::OpSize::i8Bit) {
// Since we use a u16 selector, we have enough bits for every byte in a
// 128-bit lane, so we don't need to do anything here except replicate the
// bits in the event of 256-bit.
return Is256Bit ? uint32_t(Imm) << 16 | Imm : Imm;
}
uint32_t Mask = 0;
const auto DataSize = IR::OpSizeToSize(ElementSize);
const auto NumElements = IR::NumElements(OpSize, ElementSize);
for (uint32_t i = 0; i < NumElements; i++) {
if (((Imm >> i) & 1) != 0) {
Mask |= 1U << (DataSize * i);
}
}
return Mask;
};
// Our predicate that we'll be firing our constructed bitmask into.
constexpr auto Predicate = ARMEmitter::PReg::p0.Merging();
// TODO: We can completely eliminate this via PMOV in SVE2.1
ARMEmitter::ForwardLabel AfterLabel;
ARMEmitter::BackwardLabel ConstantLabel;
(void)b(&AfterLabel);
(void)Bind(&ConstantLabel);
const auto PredicateMask = MakePredicateMask(Selector);
if (Dst == RHS) {
dc32(~PredicateMask);
} else {
dc32(PredicateMask);
}
(void)Bind(&AfterLabel);
(void)adr(TMP1, &ConstantLabel);
ldr(Predicate, TMP1);
if (Dst == LHS) {
mov(SubRegSize, LHS.Z(), Predicate, RHS.Z());
} else if (Dst == RHS) {
mov(SubRegSize, RHS.Z(), Predicate, LHS.Z());
} else {
mov(SubRegSize, Dst.Z(), GoverningPredicate.Merging(), LHS.Z());
mov(SubRegSize, Dst.Z(), Predicate, RHS.Z());
}
}
DEF_OP(VXar) {
LOGMAN_THROW_A_FMT(HostSupportsSVE128 || HostSupportsSVE256, "Host must support SVE to use {}", __func__);
auto Op = IROp->C<IR::IROp_VXar>();
const auto SubRegSize = ConvertSubRegSize8(IROp);
const auto ElementSizeBits = IR::OpSizeAsBits(IROp->ElementSize);
const auto Dst = GetVReg(Node);
const auto LHS = GetVReg(Op->LHS);
const auto RHS = GetVReg(Op->RHS);
const auto Rotate = Op->Rotate;
LOGMAN_THROW_A_FMT(Rotate >= 1 && Rotate <= ElementSizeBits, "Rotate immediate must be within [1, {}]", ElementSizeBits);
if (Dst == LHS) {
xar(SubRegSize, Dst.Z(), RHS.Z(), Rotate);
} else if (Dst == RHS) {
movprfx(VTMP1.Z(), LHS.Z());
xar(SubRegSize, VTMP1.Z(), RHS.Z(), Rotate);
mov(Dst.Z(), VTMP1.Z());
} else {
movprfx(Dst.Z(), LHS.Z());
xar(SubRegSize, Dst.Z(), RHS.Z(), Rotate);
}
}
DEF_OP(VFCopySign) {
auto Op = IROp->C<IR::IROp_VFCopySign>();
const auto OpSize = IROp->Size;
+22 -9
View File
@@ -13,6 +13,7 @@
#include <FEXCore/fextl/memory_resource.h>
#include <cstdint>
#include <span>
#include <stddef.h>
#include <utility>
#include <mutex>
@@ -93,13 +94,15 @@ struct GuestToHostMap {
GuestToHostMap();
// Adds to Guest -> Host code mapping
const BlockEntry& AddBlockMapping(uint64_t Address, const fextl::vector<uint64_t>& CodePages, void* HostCode, const LookupCacheWriteLockToken&) {
const BlockEntry& AddBlockMapping(uint64_t Address, std::span<const uint64_t> CodePages, void* HostCode, const LookupCacheWriteLockToken&) {
// This may replace an existing mapping
// NOTE: Generally no previous entry should exist, however there is one exception:
// If the backend updates the active thread's CodeBuffer, the new associated LookupCache
// may already contain the block address. Since is comparatively rare, we'll just leak
// one of the two blocks in this case.
return BlockList.insert_or_assign(Address, BlockEntry {(uintptr_t)HostCode, CodePages}).first->second;
return BlockList
.insert_or_assign(Address, BlockEntry {(uintptr_t)HostCode, fextl::vector<uint64_t>(CodePages.begin(), CodePages.end())})
.first->second;
}
const BlockEntry* FindBlock(uint64_t Address, const LookupCacheReadLockToken&) {
@@ -220,7 +223,7 @@ public:
}
if (HostPtr && DynamicL1Cache()) {
UpdateDynamicL1Stats(Thread);
UpdateDynamicL1Stats(Thread, Address, HostPtr);
}
FEXCORE_PROFILE_INSTANT_INCREMENT(Thread, AccumulatedCacheMissCount, 1);
@@ -228,7 +231,7 @@ public:
return HostPtr;
}
void UpdateDynamicL1Stats(FEXCore::Core::InternalThreadState* Thread) {
void UpdateDynamicL1Stats(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestAddress, uint64_t HostCode) {
// If host pointer was found in L2 or L3, then add it to the counter.
// Keeping track not L1 misses, but specifically L2/L3 hits.
++L2L3CacheHits;
@@ -242,12 +245,18 @@ public:
if (AveragePerSecond >= DynamicL1CacheIncreaseCountHeuristic()) {
if (CurrentL1Entries < MAX_L1_ENTRIES) {
// Entries whose address has the new mask bit set would be unreachable by InvalidateCache
FEXCore::Allocator::VirtualDontNeed(reinterpret_cast<void*>(L1Pointer), CurrentL1Entries * sizeof(LookupCacheEntry), false);
CurrentL1Entries <<= 1;
L1PointerMask = CurrentL1Entries - 1;
// Update the thread's L1 pointer mask to increase how much cache it uses.
// Since we're in C-code, this is safe to update here.
Thread->CurrentFrame->State.L1Mask = GetScaledL1PointerMask();
// If L1 was just shrunk, then we just removed our cached entry. Add it back.
AddL1Entry(GuestAddress, HostCode);
}
} else if (AveragePerSecond < DynamicL1CacheDecreaseCountHeuristic()) {
if (CurrentL1Entries > MIN_L1_ENTRIES) {
@@ -275,7 +284,7 @@ public:
// Appends a list of Block {Address} to CodePages [Start, Start + Length)
// Returns true if new pages are marked as containing code
bool AddBlockExecutableRange(FEXCore::Core::InternalThreadState* Thread, const fextl::set<uint64_t>& Addresses, uint64_t Start, uint64_t Length) {
bool AddBlockExecutableRange(FEXCore::Core::InternalThreadState* Thread, auto& Addresses, uint64_t Start, uint64_t Length) {
std::optional<FEXCore::SHMStats::AccumulationBlock<uint64_t>> LockTime(
Thread->ThreadStats ? &Thread->ThreadStats->AccumulatedCacheWriteLockTime : nullptr);
auto lk = Shared->AcquireWriteLock();
@@ -285,7 +294,7 @@ public:
}
// Adds to Guest -> Host code mapping
void AddBlockMapping(FEXCore::Core::InternalThreadState* Thread, uint64_t Address, const fextl::vector<uint64_t>& CodePages, void* HostCode) {
void AddBlockMapping(FEXCore::Core::InternalThreadState* Thread, uint64_t Address, std::span<const uint64_t> CodePages, void* HostCode) {
std::optional<FEXCore::SHMStats::AccumulationBlock<uint64_t>> LockTime(
Thread->ThreadStats ? &Thread->ThreadStats->AccumulatedCacheWriteLockTime : nullptr);
auto lk = Shared->AcquireWriteLock();
@@ -380,15 +389,19 @@ public:
}
private:
void AddL1Entry(uint64_t GuestAddress, uint64_t HostCode) {
auto& L1Entry = reinterpret_cast<LookupCacheEntry*>(L1Pointer)[GuestAddress & L1PointerMask];
L1Entry.GuestCode = GuestAddress;
L1Entry.HostCode = HostCode;
}
void CacheBlockMapping(uint64_t Address, const GuestToHostMap::BlockEntry& Entry, bool L1Only, const LookupCacheBaseLockToken& lk) {
for (const auto& CodePage : Entry.CodePages) {
CachedCodePages[CodePage >> 12].insert(Address);
}
// Do L1
auto& L1Entry = reinterpret_cast<LookupCacheEntry*>(L1Pointer)[Address & L1PointerMask];
L1Entry.GuestCode = Address;
L1Entry.HostCode = Entry.HostCode;
AddL1Entry(Address, Entry.HostCode);
if (!DisableL2Cache() && !L1Only) {
// Do ful map
@@ -36,35 +36,6 @@ using X86Tables::OpToIndex;
#define OpcodeArgs [[maybe_unused]] FEXCore::X86Tables::DecodedOp Op
void OpDispatchBuilder::SyscallOp(OpcodeArgs, bool IsSyscallInst) {
constexpr size_t SyscallArgs = 7;
using SyscallArray = std::array<uint64_t, SyscallArgs>;
size_t NumArguments {};
const SyscallArray* GPRIndexes {};
static constexpr SyscallArray GPRIndexes_64 = {
FEXCore::X86State::REG_RAX, FEXCore::X86State::REG_RDI, FEXCore::X86State::REG_RSI, FEXCore::X86State::REG_RDX,
FEXCore::X86State::REG_R10, FEXCore::X86State::REG_R8, FEXCore::X86State::REG_R9,
};
static constexpr SyscallArray GPRIndexes_32 = {
FEXCore::X86State::REG_RAX, FEXCore::X86State::REG_RBX, FEXCore::X86State::REG_RCX, FEXCore::X86State::REG_RDX,
FEXCore::X86State::REG_RSI, FEXCore::X86State::REG_RDI, FEXCore::X86State::REG_RBP,
};
const auto OSABI = CTX->SyscallHandler->GetOSABI();
if (OSABI == FEXCore::HLE::SyscallOSABI::OS_LINUX64) {
NumArguments = GPRIndexes_64.size();
GPRIndexes = &GPRIndexes_64;
} else if (OSABI == FEXCore::HLE::SyscallOSABI::OS_LINUX32) {
NumArguments = GPRIndexes_32.size();
GPRIndexes = &GPRIndexes_32;
} else if (OSABI == FEXCore::HLE::SyscallOSABI::OS_GENERIC) {
// All registers will be spilled before the syscall and filled afterwards so no JIT-side argument handling is necessary.
NumArguments = 0;
GPRIndexes = nullptr;
} else {
ERROR_AND_DIE_FMT("Unhandled OSABI syscall");
}
// Calculate flags early.
CalculateDeferredFlags();
@@ -72,13 +43,6 @@ void OpDispatchBuilder::SyscallOp(OpcodeArgs, bool IsSyscallInst) {
auto NewRIP = GetRelocatedPC(Op, -Op->InstSize);
_StoreContextGPR(GPRSize, NewRIP, offsetof(FEXCore::Core::CPUState, rip));
Ref Arguments[SyscallArgs] {
InvalidNode, InvalidNode, InvalidNode, InvalidNode, InvalidNode, InvalidNode, InvalidNode,
};
for (size_t i = 0; i < NumArguments; ++i) {
Arguments[i] = LoadGPRRegister(GPRIndexes->at(i));
}
if (IsSyscallInst) {
// If this is the `Syscall` instruction rather than `int 0x80` then we need to do some additional work.
// RCX = RIP after this instruction
@@ -94,12 +58,7 @@ void OpDispatchBuilder::SyscallOp(OpcodeArgs, bool IsSyscallInst) {
}
FlushRegisterCache();
auto SyscallOp = _Syscall(Arguments[0], Arguments[1], Arguments[2], Arguments[3], Arguments[4], Arguments[5], Arguments[6]);
// Generic ABI doesn't store result in RAX.
if (OSABI != FEXCore::HLE::SyscallOSABI::OS_GENERIC) {
StoreGPRRegister(X86State::REG_RAX, SyscallOp);
}
_Syscall();
if (Op->TableInfo->Flags & X86Tables::InstFlags::FLAGS_BLOCK_END) {
// RIP could have been updated after coming back from the Syscall.
@@ -1526,7 +1485,7 @@ void OpDispatchBuilder::SHLDImmediateOp(OpcodeArgs) {
Res = _Extr(OpSizeFromSrc(Op), Dest, Src, Size - Shift);
}
CalculateFlags_ShiftLeftImmediate(OpSizeFromSrc(Op), Res, Dest, Shift);
CalculateFlags_ShiftLeftImmediate(OpSizeFromSrc(Op), Res, Dest, Shift, true);
CalculateDeferredFlags();
StoreResultGPR(Op, Res);
} else if (Shift == 0 && Size == 32) {
@@ -1689,8 +1648,8 @@ void OpDispatchBuilder::RotateOp(OpcodeArgs, bool Left, bool IsImmediate, bool I
}
void OpDispatchBuilder::ANDNBMIOp(OpcodeArgs) {
auto* Src1 = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.AllowUpperGarbage = true});
auto* Src2 = LoadSourceGPR(Op, Op->Src[1], Op->Flags, {.AllowUpperGarbage = true});
auto Src1 = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.AllowUpperGarbage = true});
auto Src2 = LoadSourceGPR(Op, Op->Src[1], Op->Flags, {.AllowUpperGarbage = true});
auto Dest = _Andn(OpSizeFromSrc(Op), Src2, Src1);
@@ -1703,8 +1662,8 @@ void OpDispatchBuilder::BEXTRBMIOp(OpcodeArgs) {
// along with some edge-case handling and flag setting.
LOGMAN_THROW_A_FMT(Op->InstSize >= 4, "No masking needed");
auto* Src1 = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.AllowUpperGarbage = true});
auto* Src2 = LoadSourceGPR(Op, Op->Src[1], Op->Flags, {.AllowUpperGarbage = true});
auto Src1 = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.AllowUpperGarbage = true});
auto Src2 = LoadSourceGPR(Op, Op->Src[1], Op->Flags, {.AllowUpperGarbage = true});
const auto Size = OpSizeFromSrc(Op);
const auto SrcSize = IR::OpSizeAsBits(Size);
@@ -1746,7 +1705,7 @@ void OpDispatchBuilder::BLSIBMIOp(OpcodeArgs) {
LOGMAN_THROW_A_FMT(Op->InstSize >= 4, "No masking needed");
const auto Size = OpSizeFromSrc(Op);
auto* Src = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.AllowUpperGarbage = true});
auto Src = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.AllowUpperGarbage = true});
auto NegatedSrc = _Neg(Size, Src);
auto Result = _And(Size, Src, NegatedSrc);
@@ -1767,14 +1726,14 @@ void OpDispatchBuilder::BLSMSKBMIOp(OpcodeArgs) {
LOGMAN_THROW_A_FMT(Op->InstSize >= 4, "No masking needed");
const auto Size = OpSizeFromSrc(Op);
auto* Src = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.AllowUpperGarbage = true});
auto Src = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.AllowUpperGarbage = true});
auto Result = _Xor(Size, Sub(Size, Src, 1), Src);
StoreResultGPR(Op, Result);
InvalidatePF_AF();
// CF set according to the Src
auto CFInv = To01(OpSize::i64Bit, Src);
auto CFInv = To01(Size, Src);
// The output of BLSMSK is always nonzero, so TST will clear Z (along with C
// and O) while setting S.
@@ -1787,12 +1746,12 @@ void OpDispatchBuilder::BLSRBMIOp(OpcodeArgs) {
LOGMAN_THROW_A_FMT(Op->InstSize >= 4, "No masking needed");
const auto Size = OpSizeFromSrc(Op);
auto* Src = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.AllowUpperGarbage = true});
auto Src = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.AllowUpperGarbage = true});
auto Result = _And(Size, Sub(Size, Src, 1), Src);
StoreResultGPR(Op, Result);
auto CFInv = To01(OpSize::i64Bit, Src);
auto CFInv = To01(Size, Src);
SetNZ_ZeroCV(Size, Result);
SetCFInverted(CFInv);
@@ -1807,8 +1766,8 @@ void OpDispatchBuilder::BMI2Shift(OpcodeArgs) {
const auto Size = OpSizeFromSrc(Op);
const auto SrcSize = Op->Src[0].IsGPR() ? GPRSize : Size;
auto* Src = LoadSourceGPR_WithOpSize(Op, Op->Src[0], SrcSize, Op->Flags);
auto* Shift = LoadSourceGPR_WithOpSize(Op, Op->Src[1], GPRSize, Op->Flags, {.AllowUpperGarbage = true});
auto Src = LoadSourceGPR_WithOpSize(Op, Op->Src[0], SrcSize, Op->Flags);
auto Shift = LoadSourceGPR_WithOpSize(Op, Op->Src[1], GPRSize, Op->Flags, {.AllowUpperGarbage = true});
Ref Result;
if (Op->OP == 0x6F7) {
@@ -1831,9 +1790,9 @@ void OpDispatchBuilder::BZHI(OpcodeArgs) {
// In 32-bit mode we only look at bottom 32-bit, no 8 or 16-bit BZHI so no
// need to zero-extend sources
auto* Src = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.AllowUpperGarbage = true});
auto Src = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.AllowUpperGarbage = true});
auto* Index = LoadSourceGPR(Op, Op->Src[1], Op->Flags, {.AllowUpperGarbage = true});
auto Index = LoadSourceGPR(Op, Op->Src[1], Op->Flags, {.AllowUpperGarbage = true});
// Clear the high bits specified by the index. A64 only considers bottom bits
// of the shift, so we don't need to mask bottom 8-bits ourselves.
@@ -1878,8 +1837,8 @@ void OpDispatchBuilder::RORX(OpcodeArgs) {
return;
}
auto* Src = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.AllowUpperGarbage = true});
auto* Result = Src;
auto Src = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.AllowUpperGarbage = true});
auto Result = Src;
if (DoRotation) [[likely]] {
Result = _Ror(OpSizeFromSrc(Op), Src, _InlineConstant(Amount));
}
@@ -1916,8 +1875,8 @@ void OpDispatchBuilder::MULX(OpcodeArgs) {
void OpDispatchBuilder::PDEP(OpcodeArgs) {
LOGMAN_THROW_A_FMT(Op->InstSize >= 4, "No masking needed");
auto* Input = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.AllowUpperGarbage = true});
auto* Mask = LoadSourceGPR(Op, Op->Src[1], Op->Flags, {.AllowUpperGarbage = true});
auto Input = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.AllowUpperGarbage = true});
auto Mask = LoadSourceGPR(Op, Op->Src[1], Op->Flags, {.AllowUpperGarbage = true});
auto Result = _PDep(OpSizeFromSrc(Op), Input, Mask);
StoreResultGPR(Op, Op->Dest, Result);
@@ -1925,8 +1884,8 @@ void OpDispatchBuilder::PDEP(OpcodeArgs) {
void OpDispatchBuilder::PEXT(OpcodeArgs) {
LOGMAN_THROW_A_FMT(Op->InstSize >= 4, "No masking needed");
auto* Input = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.AllowUpperGarbage = true});
auto* Mask = LoadSourceGPR(Op, Op->Src[1], Op->Flags, {.AllowUpperGarbage = true});
auto Input = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.AllowUpperGarbage = true});
auto Mask = LoadSourceGPR(Op, Op->Src[1], Op->Flags, {.AllowUpperGarbage = true});
auto Result = _PExt(OpSizeFromSrc(Op), Input, Mask);
StoreResultGPR(Op, Op->Dest, Result);
@@ -1936,8 +1895,8 @@ void OpDispatchBuilder::ADXOp(OpcodeArgs) {
const auto OpSize = OpSizeFromSrc(Op);
// Only 32/64-bit anyway so allow garbage, we use 32-bit ops.
auto* Src = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.AllowUpperGarbage = true});
auto* Before = LoadSourceGPR(Op, Op->Dest, Op->Flags, {.AllowUpperGarbage = true});
auto Src = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.AllowUpperGarbage = true});
auto Before = LoadSourceGPR(Op, Op->Dest, Op->Flags, {.AllowUpperGarbage = true});
// Handles ADCX and ADOX
const bool IsADCX = Op->OP == 0x1F6;
@@ -3877,7 +3836,6 @@ void OpDispatchBuilder::CMPXCHGOp(OpcodeArgs) {
// This allows us to only hit the ZEXT case on failure
Ref RAXResult = NZCVSelect(OpSize::i64Bit, CondClass::EQ, Src3, Src1Lower);
// When the size is 4 we need to make sure not zext the GPR when the comparison fails
StoreGPRRegister(X86State::REG_RAX, RAXResult);
} else {
StoreGPRRegister(X86State::REG_RAX, Src1Lower, Size);
@@ -3891,7 +3849,7 @@ void OpDispatchBuilder::CMPXCHGOp(OpcodeArgs) {
if (GPRSize == OpSize::i64Bit && Size == OpSize::i32Bit) {
Src2Lower = _Bfe(GPRSize, IR::OpSizeAsBits(Size), 0, Src2);
}
Ref DestResult = Trivial ? Src2 : NZCVSelect(OpSize::i64Bit, CondClass::EQ, Src2Lower, Src1);
Ref DestResult = Trivial ? Src2Lower : NZCVSelect(OpSize::i64Bit, CondClass::EQ, Src2Lower, Src1);
// Store in to GPR Dest
if (GPRSize == OpSize::i64Bit && Size == OpSize::i32Bit) {
@@ -4399,6 +4357,9 @@ AddressMode OpDispatchBuilder::DecodeAddress(const X86Tables::DecodedOp& Op, con
A.NonTSO |= IsNonTSOReg(AccessType, Operand.Data.SIB.Base) || IsNonTSOReg(AccessType, Operand.Data.SIB.Index);
} else if (Operand.IsLiteralRelocation()) {
A.Base = _EntrypointOffset(GPRSize, Operand.Data.LiteralRelocation.EntrypointOffset);
} else if (Operand.IsLiteralPatchable()) {
A.Base = _PatchableGuestData(OpSize::i64Bit, Operand.Data.LiteralPatchable.Value, Op->PC + Operand.Data.LiteralPatchable.FieldOffset,
static_cast<uint64_t>(Operand.Data.LiteralPatchable.Width));
} else {
LOGMAN_MSG_A_FMT("Unknown Src Type: {}\n", Operand.Type);
}
@@ -4618,7 +4579,7 @@ void OpDispatchBuilder::StoreResult(RegClass Class, X86Tables::DecodedOp Op, Ref
}
OpDispatchBuilder::OpDispatchBuilder(FEXCore::Context::ContextImpl* ctx, FEXCore::Core::InternalThreadState* Thread)
: IREmitter {ctx->OpDispatcherAllocator, ctx->HostFeatures.SupportsTSOImm9}
: IREmitter {ctx->OpDispatcherAllocator, ctx->HostFeatures.SupportsTSOImm9 != 0}
, CTX {ctx}
, Thread {Thread} {
if (CTX->HostFeatures.SupportsAVX && CTX->HostFeatures.SupportsSVE256) {
@@ -5038,7 +4999,11 @@ void OpDispatchBuilder::RDTSCPOp(OpcodeArgs) {
// - Explicitly use an MFENCE before this instruction if you want this behaviour
// This instruction is not an execution fence, so subsequent instructions can execute after this
// - Explicitly use an LFENCE after RDTSCP if you want to block this behaviour
if (CTX->HostFeatures.HostType != FEXCore::HostFeatures::HostTypeEnum::Linux && !CTX->HostFeatures.SupportsCPUIndexInTPIDRRO) {
// RDTSCP is unsupported on Win32 platforms if TPIDRRO isn't supported.
UnimplementedOp(Op);
return;
}
auto Counter = CycleCounter(true);
auto ID = _ProcessorID();
@@ -5048,6 +5013,11 @@ void OpDispatchBuilder::RDTSCPOp(OpcodeArgs) {
}
void OpDispatchBuilder::RDPIDOp(OpcodeArgs) {
if (CTX->HostFeatures.HostType != FEXCore::HostFeatures::HostTypeEnum::Linux && !CTX->HostFeatures.SupportsCPUIndexInTPIDRRO) {
// RDTSCP is unsupported on Win32 platforms if TPIDRRO isn't supported.
UnimplementedOp(Op);
return;
}
StoreResultGPR(Op, _ProcessorID());
}
@@ -202,9 +202,10 @@ public:
FlushRegisterCache();
return _ExitFunction(GetOpSize(NewRIP), NewRIP, Hint, InvalidNode, InvalidNode);
}
IRPair<IROp_ExitFunction> ExitFunction(Ref NewRIP, BranchHint Hint, Ref CallReturnAddress, Ref CallReturnBlock) {
IRPair<IROp_ExitFunction> ExitFunction(Ref NewRIP, BranchHint Hint, Ref CallReturnAddress, Ref CallReturnBlock,
uint64_t PatchSiteAddress = 0, uint64_t PatchSiteSize = 0) {
FlushRegisterCache();
return _ExitFunction(GetOpSize(NewRIP), NewRIP, Hint, CallReturnAddress, CallReturnBlock);
return _ExitFunction(GetOpSize(NewRIP), NewRIP, Hint, CallReturnAddress, CallReturnBlock, PatchSiteAddress, PatchSiteSize);
}
IRPair<IROp_Break> Break(BreakDefinition Reason) {
FlushRegisterCache();
@@ -360,6 +361,7 @@ public:
void MOVGPRNTOp(OpcodeArgs);
void MOVVectorAlignedOp(OpcodeArgs);
void MOVVectorUnalignedOp(OpcodeArgs);
void MOVVectorUnalignedNoNopOp(OpcodeArgs);
void MOVVectorNTOp(OpcodeArgs, bool IsAVX);
void ALUOp(OpcodeArgs, FEXCore::IR::IROps ALUIROp, FEXCore::IR::IROps AtomicFetchOp, unsigned SrcIdx);
void LSLOp(OpcodeArgs);
@@ -799,6 +801,7 @@ public:
void VPFCMPOp(OpcodeArgs, uint8_t CompType);
void PI2FWOp(OpcodeArgs);
void PF2IWOp(OpcodeArgs);
void PF2IDOp(OpcodeArgs);
void PMULHRWOp(OpcodeArgs);
@@ -839,12 +842,12 @@ public:
void SHA256MSG2Op(OpcodeArgs);
void SHA256RNDS2Op(OpcodeArgs);
void AESImcOp(OpcodeArgs);
void AESImcOp(OpcodeArgs, bool IsAVX);
void AESEncOp(OpcodeArgs);
void AESEncLastOp(OpcodeArgs);
void AESDecOp(OpcodeArgs);
void AESDecLastOp(OpcodeArgs);
void AESKeyGenAssist(OpcodeArgs);
void AESKeyGenAssist(OpcodeArgs, bool IsAVX);
void VFMAImpl(OpcodeArgs, IROps IROp, bool Scalar, uint8_t Src1Idx, uint8_t Src2Idx, uint8_t AddendIdx);
void VFMAddSubImpl(OpcodeArgs, bool AddSub, uint8_t Src1Idx, uint8_t Src2Idx, uint8_t AddendIdx);
@@ -1508,18 +1511,24 @@ private:
return _GetRelocatedPC(Op, Offset, false);
}
void ExitRelocatedPC(const FEXCore::X86Tables::DecodedOp& Op, int64_t Offset = 0) {
ExitFunction(_GetRelocatedPC(Op, Offset, true /* Inline */));
void ExitRelocatedPC(const FEXCore::X86Tables::DecodedOp& Op, int64_t Offset, BranchHint Hint, Ref CallReturnAddress, Ref CallReturnBlock) {
uint64_t PatchOffset = 0;
uint64_t PatchSize = 0;
if (Op->Src[0].IsLiteralPatchable() && Offset && Offset == (int64_t)Op->Src[0].Literal()) {
PatchOffset = Op->PC + Op->Src[0].Data.LiteralPatchable.FieldOffset;
PatchSize = Op->Src[0].Data.LiteralPatchable.Width;
}
ExitFunction(_GetRelocatedPC(Op, Offset, true /* Inline */), Hint, CallReturnAddress, CallReturnBlock, PatchOffset, PatchSize);
}
void ExitRelocatedPC(const FEXCore::X86Tables::DecodedOp& Op, int64_t Offset, BranchHint Hint, Ref CallReturnAddress, Ref CallReturnBlock) {
ExitFunction(_GetRelocatedPC(Op, Offset, true /* Inline */), Hint, CallReturnAddress, CallReturnBlock);
void ExitRelocatedPC(const FEXCore::X86Tables::DecodedOp& Op, int64_t Offset = 0) {
ExitRelocatedPC(Op, Offset, BranchHint::None, InvalidNode, InvalidNode);
}
[[nodiscard]]
static bool IsOperandMem(const X86Tables::DecodedOperand& Operand, bool Load) {
// Literals are immediates as sources but memory addresses as destinations.
return !(Load && (Operand.IsLiteral() || Operand.IsLiteralRelocation())) && !Operand.IsGPR();
return !(Load && (Operand.IsLiteral() || Operand.IsLiteralRelocation() || Operand.IsLiteralPatchable())) && !Operand.IsGPR();
}
[[nodiscard]]
@@ -2131,16 +2140,16 @@ private:
}
// Compares two floats and sets flags for a COMISS instruction
void Comiss(IR::OpSize ElementSize, Ref Src1, Ref Src2, bool InvalidateAF = false) {
void Comiss(IR::OpSize ElementSize, Ref Src1, Ref Src2) {
// First, set flags according to Arm FCMP.
HandleNZCVWrite();
_FCmp(ElementSize, Src1, Src2);
CFInverted = false;
ComissFlags(InvalidateAF);
ComissFlags();
}
// Sets flags for a COMISS instruction
void ComissFlags(bool InvalidateAF = false) {
void ComissFlags() {
LOGMAN_THROW_A_FMT(!NZCVDirty, "only expected after fcmp");
// We need to set PF according to the unordered flag. We'd rather do this
@@ -2155,12 +2164,15 @@ private:
Ref V_inv = GetRFLAG(FEXCore::X86State::RFLAG_OF_RAW_LOC, true);
SetRFLAG<FEXCore::X86State::RFLAG_PF_RAW_LOC>(V_inv);
if (!InvalidateAF) {
// Zero AF. Note that the comparison sets the raw PF to 0/1 above, so
// PF[4] is 0 so the XOR with PF will have no effect, so setting the AF
// byte to zero will indeed zero AF as intended.
SetRFLAG<FEXCore::X86State::RFLAG_AF_RAW_LOC>(Constant(0));
}
// Intel: OF, SF, and AF set to zero
// AMD: no mention of OF, SF and AF but actual hardware seems to always zero
//
// Zero AF. Note that the comparison sets the raw PF to 0/1 above, so
// PF[4] is 0 so the XOR with PF will have no effect, so setting the AF
// byte to zero will indeed zero AF as intended.
// OF and SF are zeroed:
// _AXFLAG always produces N=0 (SF), V=0 (OF)
SetRFLAG<FEXCore::X86State::RFLAG_AF_RAW_LOC>(Constant(0));
// Convert NZCV from the Arm representation to an eXternal representation
// that's totally not a euphemism for x86, nuh-uh. But maps to exactly we
@@ -2372,7 +2384,7 @@ private:
void CalculateFlags_MUL(IR::OpSize SrcSize, Ref Res, Ref High);
void CalculateFlags_UMUL(Ref High);
void CalculateFlags_Logical(IR::OpSize SrcSize, Ref Res);
void CalculateFlags_ShiftLeftImmediate(IR::OpSize SrcSize, Ref Res, Ref Src1, uint64_t Shift);
void CalculateFlags_ShiftLeftImmediate(IR::OpSize SrcSize, Ref Res, Ref Src1, uint64_t Shift, bool DoubleWide = false);
void CalculateFlags_ShiftRightImmediate(IR::OpSize SrcSize, Ref Res, Ref Src1, uint64_t Shift);
void CalculateFlags_ShiftRightDoubleImmediate(IR::OpSize SrcSize, Ref Res, Ref Src1, uint64_t Shift);
void CalculateFlags_ShiftRightImmediateCommon(IR::OpSize SrcSize, Ref Res, Ref Src1, uint64_t Shift);
@@ -603,7 +603,7 @@ void OpDispatchBuilder::AVX128_CVTFPR_To_GPR(OpcodeArgs, IR::OpSize SrcElementSi
void OpDispatchBuilder::AVX128_VANDN(OpcodeArgs) {
AVX128_VectorBinaryImpl(Op, OpSizeFromSrc(Op), OpSize::i128Bit,
[this](IR::OpSize _ElementSize, Ref Src1, Ref Src2) { return _VAndn(OpSize::i128Bit, _ElementSize, Src2, Src1); });
[this](IR::OpSize, Ref Src1, Ref Src2) { return _VAndn(OpSize::i128Bit, Src2, Src1); });
}
void OpDispatchBuilder::AVX128_VPACKSS(OpcodeArgs, IR::OpSize ElementSize) {
@@ -865,13 +865,14 @@ void OpDispatchBuilder::AVX128_MOVMSK(OpcodeArgs, IR::OpSize ElementSize) {
GPR = Mask4Byte(Src.Low);
}
} else if (ElementSize == OpSize::i32Bit) {
auto GPRLow = Mask4Byte(Src.Low);
auto GPRHigh = Mask4Byte(Src.High);
GPR = _Orlshl(OpSize::i64Bit, GPRLow, GPRHigh, 4);
Ref Fused = _VUnZip2(OpSize::i128Bit, OpSize::i16Bit, Src.Low, Src.High);
Fused = _VUShrI(OpSize::i128Bit, OpSize::i16Bit, Fused, 15);
auto ConstantUSHL = LoadAndCacheNamedVectorConstant(OpSize::i128Bit, NAMED_VECTOR_INCREMENTAL_U16_INDEX);
Fused = _VUShl(OpSize::i128Bit, OpSize::i16Bit, Fused, ConstantUSHL, false);
Fused = _VAddV(OpSize::i128Bit, OpSize::i16Bit, Fused);
GPR = _VExtractToGPR(OpSize::i128Bit, OpSize::i16Bit, Fused, 0);
} else {
auto GPRLow = Mask8Byte(Src.Low);
auto GPRHigh = Mask8Byte(Src.High);
GPR = _Orlshl(OpSize::i64Bit, GPRLow, GPRHigh, 2);
GPR = Mask4Byte(_VUnZip2(OpSize::i128Bit, OpSize::i32Bit, Src.Low, Src.High));
}
StoreResultGPR_WithOpSize(Op, Op->Dest, GPR, GetGPROpSize());
}
@@ -885,7 +886,7 @@ void OpDispatchBuilder::AVX128_MOVMSKB(OpcodeArgs) {
auto Mask1Byte = [this](Ref Src, Ref VMask) {
auto VCMP = _VCMPLTZ(OpSize::i128Bit, OpSize::i8Bit, Src);
auto VAnd = _VAnd(OpSize::i128Bit, OpSize::i8Bit, VCMP, VMask);
auto VAnd = _VAnd(OpSize::i128Bit, VCMP, VMask);
auto VAdd1 = _VAddP(OpSize::i128Bit, OpSize::i8Bit, VAnd, VAnd);
auto VAdd2 = _VAddP(OpSize::i128Bit, OpSize::i8Bit, VAdd1, VAdd1);
@@ -1729,8 +1730,8 @@ void OpDispatchBuilder::AVX128_VTESTP(OpcodeArgs, IR::OpSize ElementSize) {
{
// Calculate ZF first.
auto AndLow = _VAnd(OpSize::i128Bit, OpSize::i8Bit, Src2.Low, Src1.Low);
auto AndHigh = _VAnd(OpSize::i128Bit, OpSize::i8Bit, Src2.High, Src1.High);
auto AndLow = _VAnd(OpSize::i128Bit, Src2.Low, Src1.Low);
auto AndHigh = _VAnd(OpSize::i128Bit, Src2.High, Src1.High);
auto ShiftLow = _VUShrI(OpSize::i128Bit, ElementSize, AndLow, ElementSizeInBits - 1);
auto ShiftHigh = _VUShrI(OpSize::i128Bit, ElementSize, AndHigh, ElementSizeInBits - 1);
@@ -1749,8 +1750,8 @@ void OpDispatchBuilder::AVX128_VTESTP(OpcodeArgs, IR::OpSize ElementSize) {
{
// Calculate CF Second
auto AndLow = _VAndn(OpSize::i128Bit, OpSize::i8Bit, Src2.Low, Src1.Low);
auto AndHigh = _VAndn(OpSize::i128Bit, OpSize::i8Bit, Src2.High, Src1.High);
auto AndLow = _VAndn(OpSize::i128Bit, Src2.Low, Src1.Low);
auto AndHigh = _VAndn(OpSize::i128Bit, Src2.High, Src1.High);
auto ShiftLow = _VUShrI(OpSize::i128Bit, ElementSize, AndLow, ElementSizeInBits - 1);
auto ShiftHigh = _VUShrI(OpSize::i128Bit, ElementSize, AndHigh, ElementSizeInBits - 1);
@@ -1788,11 +1789,11 @@ void OpDispatchBuilder::AVX128_PTest(OpcodeArgs) {
}
// For 256-bit, we need to unroll. This is nontrivial.
Ref Test1Low = _VAnd(OpSize::i128Bit, OpSize::i8Bit, Src1.Low, Src2.Low);
Ref Test2Low = _VAndn(OpSize::i128Bit, OpSize::i8Bit, Src2.Low, Src1.Low);
Ref Test1Low = _VAnd(OpSize::i128Bit, Src1.Low, Src2.Low);
Ref Test2Low = _VAndn(OpSize::i128Bit, Src2.Low, Src1.Low);
Ref Test1High = _VAnd(OpSize::i128Bit, OpSize::i8Bit, Src1.High, Src2.High);
Ref Test2High = _VAndn(OpSize::i128Bit, OpSize::i8Bit, Src2.High, Src1.High);
Ref Test1High = _VAnd(OpSize::i128Bit, Src1.High, Src2.High);
Ref Test2High = _VAndn(OpSize::i128Bit, Src2.High, Src1.High);
// Element size must be less than 32-bit for the sign bit tricks.
Ref Test1Max = _VUMax(OpSize::i128Bit, OpSize::i16Bit, Test1Low, Test1High);
@@ -2009,13 +2010,13 @@ void OpDispatchBuilder::AVX128_VFMAddSubImpl(OpcodeArgs, bool AddSub, uint8_t Sr
ConstantEOR = LoadAndCacheNamedVectorConstant(
OpSize::i128Bit, ElementSize == OpSize::i32Bit ? NAMED_VECTOR_PSUBADDPS_INVERT : NAMED_VECTOR_PSUBADDPD_INVERT);
}
auto InvertedSourceLow = _VXor(OpSize::i128Bit, ElementSize, Sources[AddendIdx - 1].Low, ConstantEOR);
auto InvertedSourceLow = _VXor(OpSize::i128Bit, Sources[AddendIdx - 1].Low, ConstantEOR);
Result.Low = _VFMLA(OpSize::i128Bit, ElementSize, Sources[Src1Idx - 1].Low, Sources[Src2Idx - 1].Low, InvertedSourceLow);
if (Is128Bit) {
Result.High = LoadZeroVector(OpSize::i128Bit);
} else {
auto InvertedSourceHigh = _VXor(OpSize::i128Bit, ElementSize, Sources[AddendIdx - 1].High, ConstantEOR);
auto InvertedSourceHigh = _VXor(OpSize::i128Bit, Sources[AddendIdx - 1].High, ConstantEOR);
Result.High = _VFMLA(OpSize::i128Bit, ElementSize, Sources[Src1Idx - 1].High, Sources[Src2Idx - 1].High, InvertedSourceHigh);
}
AVX128_StoreResult_WithOpSize(Op, Op->Dest, Result);
@@ -26,17 +26,25 @@ void OpDispatchBuilder::SHA1NEXTEOp(OpcodeArgs) {
Ref Dest = LoadSourceFPR(Op, Op->Dest, Op->Flags);
Ref Src = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
// ARMv8 SHA1 extension provides a `SHA1H` instruction which does a fixed rotate by 30.
// This only operates on element 0 rather than element 3. We don't have the luxury of rewriting the x86 SHA algorithm to take advantage of this.
// Move the element to zero, rotate, and then move back (Using duplicates).
// Saves one instruction versus that path that doesn't support SHA extension.
auto Duplicated = _VDupElement(OpSize::i128Bit, OpSize::i32Bit, Dest, 3);
auto Sha1HRotated = _VSha1H(Duplicated);
auto RotatedNode = _VDupElement(OpSize::i128Bit, OpSize::i32Bit, Sha1HRotated, 0);
auto Tmp = _VAdd(OpSize::i128Bit, OpSize::i32Bit, Src, RotatedNode);
auto Result = _VInsElement(OpSize::i128Bit, OpSize::i32Bit, 3, 3, Src, Tmp);
Ref Result {};
if (CTX->HostFeatures.SupportsSVE128) {
auto ZeroVec = LoadZeroVector(OpSize::i128Bit);
auto Tmp = _VInsElement(OpSize::i128Bit, OpSize::i32Bit, 3, 3, ZeroVec, Dest);
auto Xar = _VXar(OpSize::i128Bit, OpSize::i32Bit, ZeroVec, Tmp, 2);
Result = _VAdd(OpSize::i128Bit, OpSize::i32Bit, Src, Xar);
} else {
// ARMv8 SHA1 extension provides a `SHA1H` instruction which does a fixed rotate by 30.
// This only operates on element 0 rather than element 3. We don't have the luxury of rewriting the x86 SHA algorithm to take advantage of this.
// Move the element to zero, rotate, and then move back (Using duplicates).
// Saves one instruction versus that path that doesn't support SHA extension.
auto Duplicated = _VDupElement(OpSize::i128Bit, OpSize::i32Bit, Dest, 3);
auto Sha1HRotated = _VSha1H(Duplicated);
auto RotatedNode = _VDupElement(OpSize::i128Bit, OpSize::i32Bit, Sha1HRotated, 0);
auto Tmp = _VAdd(OpSize::i128Bit, OpSize::i32Bit, Src, RotatedNode);
Result = _VInsElement(OpSize::i128Bit, OpSize::i32Bit, 3, 3, Src, Tmp);
}
StoreResultFPR(Op, Result);
StoreResult_WithAVXInsert(VectorOpType::SSE, RegClass::FPR, Op, Result);
}
void OpDispatchBuilder::SHA1MSG1Op(OpcodeArgs) {
@@ -50,9 +58,9 @@ void OpDispatchBuilder::SHA1MSG1Op(OpcodeArgs) {
Ref NewVec = _VExtr(OpSize::i128Bit, OpSize::i64Bit, Dest, Src, 1);
// [W0, W1, W2, W3] ^ [W2, W3, W4, W5]
Ref Result = _VXor(OpSize::i128Bit, OpSize::i8Bit, Dest, NewVec);
Ref Result = _VXor(OpSize::i128Bit, Dest, NewVec);
StoreResultFPR(Op, Result);
StoreResult_WithAVXInsert(VectorOpType::SSE, RegClass::FPR, Op, Result);
}
void OpDispatchBuilder::SHA1MSG2Op(OpcodeArgs) {
@@ -70,7 +78,7 @@ void OpDispatchBuilder::SHA1MSG2Op(OpcodeArgs) {
// The result is swizzled differently than expected
auto Result = SHADataShuffle(_VSha1SU1(Src1, Src2));
StoreResultFPR(Op, Result);
StoreResult_WithAVXInsert(VectorOpType::SSE, RegClass::FPR, Op, Result);
}
void OpDispatchBuilder::SHA1RNDS4Op(OpcodeArgs) {
@@ -99,7 +107,7 @@ void OpDispatchBuilder::SHA1RNDS4Op(OpcodeArgs) {
break;
}
const auto ZeroRegister = LoadZeroVector(OpSize::i32Bit);
const auto ZeroRegister = LoadZeroVector(OpSize::i128Bit);
Ref Src1 = SHADataShuffle(Dest);
Ref Src2 = SHADataShuffle(Src);
@@ -112,7 +120,7 @@ void OpDispatchBuilder::SHA1RNDS4Op(OpcodeArgs) {
case 3: Result = SHADataShuffle(_VSha1P(Src1, ZeroRegister, Src2)); break;
}
StoreResultFPR(Op, Result);
StoreResult_WithAVXInsert(VectorOpType::SSE, RegClass::FPR, Op, Result);
}
void OpDispatchBuilder::SHA256MSG1Op(OpcodeArgs) {
@@ -125,7 +133,7 @@ void OpDispatchBuilder::SHA256MSG1Op(OpcodeArgs) {
auto Result = _VSha256U0(Dest, Src);
StoreResultFPR(Op, Result);
StoreResult_WithAVXInsert(VectorOpType::SSE, RegClass::FPR, Op, Result);
}
void OpDispatchBuilder::SHA256MSG2Op(OpcodeArgs) {
@@ -142,7 +150,7 @@ void OpDispatchBuilder::SHA256MSG2Op(OpcodeArgs) {
auto Result = _VSha256U1(Src1, Src2);
StoreResultFPR(Op, Result);
StoreResult_WithAVXInsert(VectorOpType::SSE, RegClass::FPR, Op, Result);
}
void OpDispatchBuilder::SHA256RNDS2Op(OpcodeArgs) {
@@ -177,17 +185,22 @@ void OpDispatchBuilder::SHA256RNDS2Op(OpcodeArgs) {
auto B = _VSha256H2(EFGH, ABCD, Key);
auto Result = shuffle_abcd(A, B);
StoreResultFPR(Op, Result);
StoreResult_WithAVXInsert(VectorOpType::SSE, RegClass::FPR, Op, Result);
}
void OpDispatchBuilder::AESImcOp(OpcodeArgs) {
void OpDispatchBuilder::AESImcOp(OpcodeArgs, bool IsAVX) {
if (!CTX->HostFeatures.SupportsAES) {
UnimplementedOp(Op);
return;
}
Ref Src = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
Ref Result = _VAESImc(Src);
StoreResultFPR(Op, Result);
if (IsAVX) {
StoreResultFPR(Op, Result);
} else {
StoreResult_WithAVXInsert(VectorOpType::SSE, RegClass::FPR, Op, Result);
}
}
void OpDispatchBuilder::AESEncOp(OpcodeArgs) {
@@ -198,19 +211,30 @@ void OpDispatchBuilder::AESEncOp(OpcodeArgs) {
Ref Dest = LoadSourceFPR(Op, Op->Dest, Op->Flags);
Ref Src = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
Ref Result = _VAESEnc(OpSize::i128Bit, Dest, Src, LoadZeroVector(OpSize::i128Bit));
StoreResultFPR(Op, Result);
StoreResult_WithAVXInsert(VectorOpType::SSE, RegClass::FPR, Op, Result);
}
void OpDispatchBuilder::VAESEncOp(OpcodeArgs) {
const auto DstSize = OpSizeFromDst(Op);
const auto Is128Bit = DstSize == OpSize::i128Bit;
// TODO: Handle 256-bit VAESENC.
LOGMAN_THROW_A_FMT(Is128Bit, "256-bit VAESENC unimplemented");
const auto Is256Bit = DstSize == OpSize::i256Bit;
Ref State = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
Ref Key = LoadSourceFPR(Op, Op->Src[1], Op->Flags);
Ref Result = _VAESEnc(DstSize, State, Key, LoadZeroVector(DstSize));
Ref ZeroVec = LoadZeroVector(DstSize);
Ref Result {};
if (Is256Bit) {
// TODO: Handle as one operation once vixl supports it.
auto UpperState = _VDupElement(DstSize, OpSize::i128Bit, State, 1);
auto UpperKey = _VDupElement(DstSize, OpSize::i128Bit, Key, 1);
auto Lower = _VAESEnc(OpSize::i128Bit, State, Key, ZeroVec);
auto Upper = _VAESEnc(OpSize::i128Bit, UpperState, UpperKey, ZeroVec);
Result = _VInsElement(DstSize, OpSize::i128Bit, 1, 0, Lower, Upper);
} else {
Result = _VAESEnc(DstSize, State, Key, ZeroVec);
}
StoreResultFPR(Op, Result);
}
@@ -223,19 +247,30 @@ void OpDispatchBuilder::AESEncLastOp(OpcodeArgs) {
Ref Dest = LoadSourceFPR(Op, Op->Dest, Op->Flags);
Ref Src = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
Ref Result = _VAESEncLast(OpSize::i128Bit, Dest, Src, LoadZeroVector(OpSize::i128Bit));
StoreResultFPR(Op, Result);
StoreResult_WithAVXInsert(VectorOpType::SSE, RegClass::FPR, Op, Result);
}
void OpDispatchBuilder::VAESEncLastOp(OpcodeArgs) {
const auto DstSize = OpSizeFromDst(Op);
const auto Is128Bit = DstSize == OpSize::i128Bit;
// TODO: Handle 256-bit VAESENCLAST.
LOGMAN_THROW_A_FMT(Is128Bit, "256-bit VAESENCLAST unimplemented");
const auto Is256Bit = DstSize == OpSize::i256Bit;
Ref State = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
Ref Key = LoadSourceFPR(Op, Op->Src[1], Op->Flags);
Ref Result = _VAESEncLast(DstSize, State, Key, LoadZeroVector(DstSize));
Ref ZeroVec = LoadZeroVector(DstSize);
Ref Result {};
if (Is256Bit) {
// TODO: Handle as one operation once vixl supports it.
auto UpperState = _VDupElement(DstSize, OpSize::i128Bit, State, 1);
auto UpperKey = _VDupElement(DstSize, OpSize::i128Bit, Key, 1);
auto Lower = _VAESEncLast(OpSize::i128Bit, State, Key, ZeroVec);
auto Upper = _VAESEncLast(OpSize::i128Bit, UpperState, UpperKey, ZeroVec);
Result = _VInsElement(DstSize, OpSize::i128Bit, 1, 0, Lower, Upper);
} else {
Result = _VAESEncLast(DstSize, State, Key, ZeroVec);
}
StoreResultFPR(Op, Result);
}
@@ -248,19 +283,30 @@ void OpDispatchBuilder::AESDecOp(OpcodeArgs) {
Ref Dest = LoadSourceFPR(Op, Op->Dest, Op->Flags);
Ref Src = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
Ref Result = _VAESDec(OpSize::i128Bit, Dest, Src, LoadZeroVector(OpSize::i128Bit));
StoreResultFPR(Op, Result);
StoreResult_WithAVXInsert(VectorOpType::SSE, RegClass::FPR, Op, Result);
}
void OpDispatchBuilder::VAESDecOp(OpcodeArgs) {
const auto DstSize = OpSizeFromDst(Op);
const auto Is128Bit = DstSize == OpSize::i128Bit;
// TODO: Handle 256-bit VAESDEC.
LOGMAN_THROW_A_FMT(Is128Bit, "256-bit VAESDEC unimplemented");
const auto Is256Bit = DstSize == OpSize::i256Bit;
Ref State = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
Ref Key = LoadSourceFPR(Op, Op->Src[1], Op->Flags);
Ref Result = _VAESDec(DstSize, State, Key, LoadZeroVector(DstSize));
Ref ZeroVec = LoadZeroVector(DstSize);
Ref Result {};
if (Is256Bit) {
// TODO: Handle as one operation once vixl supports it.
auto UpperState = _VDupElement(DstSize, OpSize::i128Bit, State, 1);
auto UpperKey = _VDupElement(DstSize, OpSize::i128Bit, Key, 1);
auto Lower = _VAESDec(OpSize::i128Bit, State, Key, ZeroVec);
auto Upper = _VAESDec(OpSize::i128Bit, UpperState, UpperKey, ZeroVec);
Result = _VInsElement(DstSize, OpSize::i128Bit, 1, 0, Lower, Upper);
} else {
Result = _VAESDec(DstSize, State, Key, ZeroVec);
}
StoreResultFPR(Op, Result);
}
@@ -273,19 +319,30 @@ void OpDispatchBuilder::AESDecLastOp(OpcodeArgs) {
Ref Dest = LoadSourceFPR(Op, Op->Dest, Op->Flags);
Ref Src = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
Ref Result = _VAESDecLast(OpSize::i128Bit, Dest, Src, LoadZeroVector(OpSize::i128Bit));
StoreResultFPR(Op, Result);
StoreResult_WithAVXInsert(VectorOpType::SSE, RegClass::FPR, Op, Result);
}
void OpDispatchBuilder::VAESDecLastOp(OpcodeArgs) {
const auto DstSize = OpSizeFromDst(Op);
const auto Is128Bit = DstSize == OpSize::i128Bit;
// TODO: Handle 256-bit VAESDECLAST.
LOGMAN_THROW_A_FMT(Is128Bit, "256-bit VAESDECLAST unimplemented");
const auto Is256Bit = DstSize == OpSize::i256Bit;
Ref State = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
Ref Key = LoadSourceFPR(Op, Op->Src[1], Op->Flags);
Ref Result = _VAESDecLast(DstSize, State, Key, LoadZeroVector(DstSize));
Ref ZeroVec = LoadZeroVector(DstSize);
Ref Result {};
if (Is256Bit) {
// TODO: Handle as one operation once vixl supports it.
auto UpperState = _VDupElement(DstSize, OpSize::i128Bit, State, 1);
auto UpperKey = _VDupElement(DstSize, OpSize::i128Bit, Key, 1);
auto Lower = _VAESDecLast(OpSize::i128Bit, State, Key, ZeroVec);
auto Upper = _VAESDecLast(OpSize::i128Bit, UpperState, UpperKey, ZeroVec);
Result = _VInsElement(DstSize, OpSize::i128Bit, 1, 0, Lower, Upper);
} else {
Result = _VAESDecLast(DstSize, State, Key, ZeroVec);
}
StoreResultFPR(Op, Result);
}
@@ -298,14 +355,19 @@ Ref OpDispatchBuilder::AESKeyGenAssistImpl(OpcodeArgs) {
return _VAESKeyGenAssist(Src, KeyGenSwizzle, LoadZeroVector(OpSize::i128Bit), RCON);
}
void OpDispatchBuilder::AESKeyGenAssist(OpcodeArgs) {
void OpDispatchBuilder::AESKeyGenAssist(OpcodeArgs, bool IsAVX) {
if (!CTX->HostFeatures.SupportsAES) {
UnimplementedOp(Op);
return;
}
Ref Result = AESKeyGenAssistImpl(Op);
StoreResultFPR(Op, Result);
if (IsAVX) {
StoreResultFPR(Op, Result);
} else {
StoreResult_WithAVXInsert(VectorOpType::SSE, RegClass::FPR, Op, Result);
}
}
void OpDispatchBuilder::PCLMULQDQOp(OpcodeArgs) {
@@ -317,8 +379,8 @@ void OpDispatchBuilder::PCLMULQDQOp(OpcodeArgs) {
Ref Src = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
const auto Selector = static_cast<uint8_t>(Op->Src[1].Literal());
auto Res = _PCLMUL(OpSize::i128Bit, Dest, Src, Selector & 0b1'0001);
StoreResultFPR(Op, Res);
auto Result = _PCLMUL(OpSize::i128Bit, Dest, Src, Selector & 0b1'0001);
StoreResult_WithAVXInsert(VectorOpType::SSE, RegClass::FPR, Op, Result);
}
void OpDispatchBuilder::VPCLMULQDQOp(OpcodeArgs) {
@@ -7,7 +7,7 @@ constexpr DispatchTableEntry OpDispatch_DDDTable[] = {
{0x0C, 1, &OpDispatchBuilder::PI2FWOp},
{0x0D, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Int_To_Float, OpSize::i32Bit, false, false>},
{0x1C, 1, &OpDispatchBuilder::PF2IWOp},
{0x1D, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Vector_CVT_Float_To_Int, OpSize::i32Bit, false, false>},
{0x1D, 1, &OpDispatchBuilder::PF2IDOp},
{0x86, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorUnaryOp, IR::OP_VFRECPPRECISION, OpSize::i32Bit>},
{0x87, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::RSqrt3DNowOp, false>},
@@ -26,8 +26,8 @@ constexpr DispatchTableEntry OpDispatch_DDDTable[] = {
{0xA0, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPFCMPOp, 2>},
{0xA4, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFMAX, OpSize::i32Bit>},
// Can be treated as a move
{0xA6, 1, &OpDispatchBuilder::MOVVectorUnalignedOp},
{0xA7, 1, &OpDispatchBuilder::MOVVectorUnalignedOp},
{0xA6, 1, &OpDispatchBuilder::MOVVectorUnalignedNoNopOp},
{0xA7, 1, &OpDispatchBuilder::MOVVectorUnalignedNoNopOp},
{0xAA, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUROp, IR::OP_VFSUB, OpSize::i32Bit>},
{0xAE, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFADDP, OpSize::i32Bit>},
@@ -35,7 +35,7 @@ constexpr DispatchTableEntry OpDispatch_DDDTable[] = {
{0xB0, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPFCMPOp, 0>},
{0xB4, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFMUL, OpSize::i32Bit>},
// Can be treated as a move
{0xB6, 1, &OpDispatchBuilder::MOVVectorUnalignedOp},
{0xB6, 1, &OpDispatchBuilder::MOVVectorUnalignedNoNopOp},
{0xB7, 1, &OpDispatchBuilder::PMULHRWOp},
{0xBB, 1, &OpDispatchBuilder::PSWAPDOp},
@@ -432,7 +432,7 @@ void OpDispatchBuilder::CalculateFlags_Logical(IR::OpSize SrcSize, Ref Res) {
SetNZP_ZeroCV(SrcSize, Res);
}
void OpDispatchBuilder::CalculateFlags_ShiftLeftImmediate(IR::OpSize SrcSize, Ref UnmaskedRes, Ref Src1, uint64_t Shift) {
void OpDispatchBuilder::CalculateFlags_ShiftLeftImmediate(IR::OpSize SrcSize, Ref UnmaskedRes, Ref Src1, uint64_t Shift, bool DoubleWide) {
// No flags changed if shift is zero
if (Shift == 0) {
return;
@@ -447,8 +447,12 @@ void OpDispatchBuilder::CalculateFlags_ShiftLeftImmediate(IR::OpSize SrcSize, Re
// Extract the last bit shifted in to CF. Shift is already masked, but for
// 8/16-bit it might be >= SrcSizeBits, in which case CF is cleared. There's
// nothing to do in that case since we already cleared CF above.
//
// - Double-wide shift has UB when shift is GREATER-THAN operand.
// - Single-wide shift has UB when shift is GREATER-THAN-EQUAL operand.
const auto SrcSizeBits = IR::OpSizeAsBits(SrcSize);
if (Shift < SrcSizeBits) {
const bool ShouldSetCF = DoubleWide ? (Shift <= SrcSizeBits) : (Shift < SrcSizeBits);
if (ShouldSetCF) {
SetCFDirect(Src1, SrcSizeBits - Shift, true);
}
}
@@ -79,7 +79,7 @@ constexpr DispatchTableEntry OpDispatch_H0F38Table[] = {
{OPD(PF_38_NONE, 0xCC), 1, &OpDispatchBuilder::SHA256MSG1Op},
{OPD(PF_38_NONE, 0xCD), 1, &OpDispatchBuilder::SHA256MSG2Op},
{OPD(PF_38_66, 0xDB), 1, &OpDispatchBuilder::AESImcOp},
{OPD(PF_38_66, 0xDB), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AESImcOp, false>},
{OPD(PF_38_66, 0xDC), 1, &OpDispatchBuilder::AESEncOp},
{OPD(PF_38_66, 0xDD), 1, &OpDispatchBuilder::AESEncLastOp},
{OPD(PF_38_66, 0xDE), 1, &OpDispatchBuilder::AESDecOp},
@@ -37,7 +37,7 @@ constexpr auto OpDispatchTableGenH0F3A = []() consteval {
{OPD(REX, PF_3A_66, 0x63), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPCMPISTRIOp, false>},
{OPD(REX, PF_3A_NONE, 0xCC), 1, &OpDispatchBuilder::SHA1RNDS4Op},
{OPD(REX, PF_3A_66, 0xDF), 1, &OpDispatchBuilder::AESKeyGenAssist},
{OPD(REX, PF_3A_66, 0xDF), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AESKeyGenAssist, false>},
};
return std::to_array(Table);
@@ -42,6 +42,12 @@ void OpDispatchBuilder::MOVVectorUnalignedOp(OpcodeArgs) {
StoreResult_WithAVXInsert(VectorOpType::SSE, RegClass::FPR, Op, Src);
}
void OpDispatchBuilder::MOVVectorUnalignedNoNopOp(OpcodeArgs) {
// Moves to same register might have secondary-effects and can't convert to a nop.
Ref Src = LoadSourceFPR(Op, Op->Src[0], Op->Flags, {.Align = OpSize::i8Bit});
StoreResult_WithAVXInsert(VectorOpType::SSE, RegClass::FPR, Op, Src);
}
void OpDispatchBuilder::MOVVectorNTOp(OpcodeArgs, bool IsAVX) {
const auto Size = OpSizeFromDst(Op);
@@ -366,13 +372,13 @@ Ref OpDispatchBuilder::VectorScalarUnaryInsertALUOpImpl(OpcodeArgs, IROps IROp,
void OpDispatchBuilder::VectorScalarUnaryInsertALUOp(OpcodeArgs, IROps IROp, IR::OpSize ElementSize) {
const auto DstSize = GetGuestVectorLength();
auto Result = VectorScalarInsertALUOpImpl(Op, IROp, DstSize, ElementSize, Op->Dest, Op->Src[0], false);
auto Result = VectorScalarUnaryInsertALUOpImpl(Op, IROp, DstSize, ElementSize, Op->Dest, Op->Src[0], false);
StoreResultFPR_WithOpSize(Op, Op->Dest, Result, DstSize);
}
void OpDispatchBuilder::AVXVectorScalarUnaryInsertALUOp(OpcodeArgs, IROps IROp, IR::OpSize ElementSize) {
const auto DstSize = GetGuestVectorLength();
auto Result = VectorScalarInsertALUOpImpl(Op, IROp, DstSize, ElementSize, Op->Src[0], Op->Src[1], true);
auto Result = VectorScalarUnaryInsertALUOpImpl(Op, IROp, DstSize, ElementSize, Op->Src[0], Op->Src[1], true);
StoreResultFPR_WithOpSize(Op, Op->Dest, Result, DstSize);
}
@@ -523,14 +529,14 @@ Ref OpDispatchBuilder::InsertScalarFCMPOpImpl(OpSize Size, IR::OpSize OpDstSize,
case VectorCompareType::NLT_US: // NGT(Swapped operand)
case VectorCompareType::NLT_UQ: {
Ref Result = _VFCMPLT(ElementSize, ElementSize, Src1, Src2);
Result = _VNot(ElementSize, ElementSize, Result);
Result = _VNot(ElementSize, Result);
// Insert the lower bits
return _VInsElement(OpDstSize, ElementSize, 0, 0, Src1, Result);
}
case VectorCompareType::NLE_US: // NGE(Swapped operand)
case VectorCompareType::NLE_UQ: {
Ref Result = _VFCMPLE(ElementSize, ElementSize, Src1, Src2);
Result = _VNot(ElementSize, ElementSize, Result);
Result = _VNot(ElementSize, Result);
// Insert the lower bits
return _VInsElement(OpDstSize, ElementSize, 0, 0, Src1, Result);
}
@@ -539,14 +545,14 @@ Ref OpDispatchBuilder::InsertScalarFCMPOpImpl(OpSize Size, IR::OpSize OpDstSize,
case VectorCompareType::NGT_UQ:
case VectorCompareType::NGT_US: {
Ref Result = _VFCMPLT(ElementSize, ElementSize, Src2, Src1);
Result = _VNot(ElementSize, ElementSize, Result);
Result = _VNot(ElementSize, Result);
// Insert the lower bits
return _VInsElement(OpDstSize, ElementSize, 0, 0, Src1, Result);
}
case VectorCompareType::NGE_UQ:
case VectorCompareType::NGE_US: {
Ref Result = _VFCMPLE(ElementSize, ElementSize, Src2, Src1);
Result = _VNot(ElementSize, ElementSize, Result);
Result = _VNot(ElementSize, Result);
// Insert the lower bits
return _VInsElement(OpDstSize, ElementSize, 0, 0, Src1, Result);
}
@@ -567,10 +573,10 @@ Ref OpDispatchBuilder::InsertScalarFCMPOpImpl(OpSize Size, IR::OpSize OpDstSize,
// If either of the sources are unordered, then returns true.
Ref Src1_U = _VFCMPEQ(Size, ElementSize, Src1, Src1);
Ref Src2_U = _VFCMPEQ(Size, ElementSize, Src2, Src2);
auto Ordered = _VAnd(Size, ElementSize, Src1_U, Src2_U);
auto Ordered = _VAnd(Size, Src1_U, Src2_U);
Ref Compare_Ordered = _VFCMPEQ(Size, ElementSize, Src1, Src2);
Ref Result = _VOrn(Size, ElementSize, Compare_Ordered, Ordered);
Ref Result = _VOrn(Size, Compare_Ordered, Ordered);
// Insert the lower bits
return _VInsElement(OpDstSize, ElementSize, 0, 0, Src1, Result);
@@ -582,8 +588,8 @@ Ref OpDispatchBuilder::InsertScalarFCMPOpImpl(OpSize Size, IR::OpSize OpDstSize,
Ref Src2_U = _VFCMPEQ(Size, ElementSize, Src2, Src2);
Ref Compare_Ordered = _VFCMPEQ(Size, ElementSize, Src1, Src2);
Ref Result = _VAndn(Size, ElementSize, Src1_U, Compare_Ordered);
Result = _VAnd(Size, ElementSize, Result, Src2_U);
Ref Result = _VAndn(Size, Src1_U, Compare_Ordered);
Result = _VAnd(Size, Result, Src2_U);
// Insert the lower bits
return _VInsElement(OpDstSize, ElementSize, 0, 0, Src1, Result);
@@ -598,14 +604,14 @@ Ref OpDispatchBuilder::InsertScalarFCMPOpImpl(OpSize Size, IR::OpSize OpDstSize,
}
void OpDispatchBuilder::InsertScalarFCMPOp(OpcodeArgs, IR::OpSize ElementSize) {
const uint8_t CompType = Op->Src[1].Literal();
const uint8_t CompType = Op->Src[1].Literal() & 0b111;
const auto DstSize = GetGuestVectorLength();
const auto SrcSize = OpSizeFromSrc(Op);
Ref Src1 = LoadSourceFPR_WithOpSize(Op, Op->Dest, DstSize, Op->Flags);
Ref Src2 = LoadSourceFPR_WithOpSize(Op, Op->Src[0], SrcSize, Op->Flags, {.AllowUpperGarbage = true});
Ref Result = InsertScalarFCMPOpImpl(DstSize, OpSizeFromDst(Op), ElementSize, Src1, Src2, CompType & 0b111, false);
Ref Result = InsertScalarFCMPOpImpl(DstSize, OpSizeFromDst(Op), ElementSize, Src1, Src2, CompType, false);
// ARM doesn't have any instructions that handle the semantics of NLT and NLE directly.
// In fact, these are the two SSE compatison types where we cannot use VFCMPScalarInsert
@@ -623,7 +629,7 @@ void OpDispatchBuilder::InsertScalarFCMPOp(OpcodeArgs, IR::OpSize ElementSize) {
}
void OpDispatchBuilder::AVXInsertScalarFCMPOp(OpcodeArgs, IR::OpSize ElementSize) {
const uint8_t CompType = Op->Src[2].Literal();
const uint8_t CompType = Op->Src[2].Literal() & 0b11111;
const auto DstSize = GetGuestVectorLength();
const auto SrcSize = OpSizeFromSrc(Op);
@@ -633,7 +639,7 @@ void OpDispatchBuilder::AVXInsertScalarFCMPOp(OpcodeArgs, IR::OpSize ElementSize
Ref Src1 = LoadSourceFPR_WithOpSize(Op, Op->Src[0], DstSize, Op->Flags);
Ref Src2 = LoadSourceFPR_WithOpSize(Op, Op->Src[1], SrcSize, Op->Flags, {.AllowUpperGarbage = true});
Ref Result = InsertScalarFCMPOpImpl(DstSize, OpSizeFromDst(Op), ElementSize, Src1, Src2, CompType & 0b11111, true);
Ref Result = InsertScalarFCMPOpImpl(DstSize, OpSizeFromDst(Op), ElementSize, Src1, Src2, CompType, true);
StoreResultFPR_WithOpSize(Op, Op->Dest, Result, DstSize);
}
@@ -789,7 +795,7 @@ void OpDispatchBuilder::MOVMSKOpOne(OpcodeArgs) {
Ref VMask = LoadAndCacheNamedVectorConstant(SrcSize, NAMED_VECTOR_MOVMASKB);
auto VCMP = _VCMPLTZ(SrcSize, OpSize::i8Bit, Src);
auto VAnd = _VAnd(SrcSize, OpSize::i8Bit, VCMP, VMask);
auto VAnd = _VAnd(SrcSize, VCMP, VMask);
// Since we also handle the MM MOVMSKB here too,
// we need to clamp the lower bound.
@@ -884,7 +890,7 @@ Ref OpDispatchBuilder::PSHUFBOpImpl(IR::OpSize SrcSize, Ref Src1, Ref Src2, Ref
// the lane splitting behavior, so cap the maximum size at 16.
const auto SanitizedSrcSize = std::min(SrcSize, OpSize::i128Bit);
Ref MaskedIndices = _VAnd(SrcSize, SrcSize, Src2, MaskVector);
Ref MaskedIndices = _VAnd(SrcSize, Src2, MaskVector);
Ref Low = _VTBL1(SanitizedSrcSize, Src1, MaskedIndices);
if (!Is256Bit) {
@@ -1999,7 +2005,7 @@ void OpDispatchBuilder::VANDNOp(OpcodeArgs) {
Ref Src1 = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
Ref Src2 = LoadSourceFPR(Op, Op->Src[1], Op->Flags);
Ref Dest = _VAndn(SrcSize, SrcSize, Src2, Src1);
Ref Dest = _VAndn(SrcSize, Src2, Src1);
StoreResultFPR(Op, Dest);
}
@@ -2581,9 +2587,10 @@ Ref OpDispatchBuilder::CVTGPR_To_FPRImpl(OpcodeArgs, IR::OpSize DstElementSize,
}
Ref OpDispatchBuilder::CVTFPR_To_GPRImpl(OpcodeArgs, Ref Src, IR::OpSize SrcElementSize, bool HostRoundingMode) {
// GPR size is determined by REX.W
// Source Element size is determined by instruction
const auto GPRSize = OpSizeFromDst(Op);
// GPR size is determined by REX.W
// But instruction does not support 16bit register operands
const auto GPRSize = std::max(OpSize::i32Bit, OpSizeFromDst(Op));
if (CTX->HostFeatures.SupportsFRINTTS) {
// When we have FRINTTS, this is a two-step process. First, we round to the
@@ -2616,7 +2623,9 @@ void OpDispatchBuilder::CVTFPR_To_GPR(OpcodeArgs, IR::OpSize SrcElementSize, boo
const auto SrcSize = Op->Src[0].IsGPR() ? OpSize::i128Bit : SrcElementSize;
Ref Src = LoadSourceFPR_WithOpSize(Op, Op->Src[0], SrcSize, Op->Flags);
Ref Result = CVTFPR_To_GPRImpl(Op, Src, SrcElementSize, HostRoundingMode);
StoreResultGPR(Op, Result);
const auto DestSize = std::max(OpSize::i32Bit, OpSizeFromDst(Op));
StoreResultGPR_WithOpSize(Op, Op->Dest, Result, DestSize);
}
Ref OpDispatchBuilder::Vector_CVT_Int_To_FloatImpl(OpcodeArgs, IR::OpSize SrcElementSize, bool Widen) {
@@ -2782,12 +2791,12 @@ void OpDispatchBuilder::XMM_To_MMX_Vector_CVT_Float_To_Int(OpcodeArgs, IR::OpSiz
void OpDispatchBuilder::MASKMOVOp(OpcodeArgs) {
const auto Size = OpSizeFromSrc(Op);
Ref MaskSrc = LoadSourceGPR(Op, Op->Src[0], Op->Flags);
Ref MaskSrc = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
// Mask only cares about the top bit of each byte
MaskSrc = _VCMPLTZ(Size, OpSize::i8Bit, MaskSrc);
// Vector that will overwrite byte elements.
Ref VectorSrc = LoadSourceGPR(Op, Op->Dest, Op->Flags);
Ref VectorSrc = LoadSourceFPR(Op, Op->Dest, Op->Flags);
// RDI source (DS prefix by default)
auto MemDest = MakeSegmentAddress(X86State::REG_RDI, Op->Flags, X86Tables::DecodeFlags::FLAG_DS_PREFIX);
@@ -2839,11 +2848,15 @@ void OpDispatchBuilder::MOVBetweenGPR_FPR(OpcodeArgs, VectorOpType VectorType) {
if (Op->Src[0].IsGPR()) {
// Loading from GPR and moving to Vector.
Ref Src = LoadSourceFPR_WithOpSize(Op, Op->Src[0], GetGPROpSize(), Op->Flags);
const auto SrcSize = std::max(OpSize::i32Bit, OpSizeFromSrc(Op));
// zext to 128bit
Result = _VCastFromGPR(OpSize::i128Bit, OpSizeFromSrc(Op), Src);
Result = _VCastFromGPR(OpSize::i128Bit, SrcSize, Src);
} else {
// Loading from Memory as a scalar. Zero extend
Result = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
const auto SrcSize = std::max(OpSize::i32Bit, OpSizeFromSrc(Op));
Result = LoadSourceFPR_WithOpSize(Op, Op->Src[0], SrcSize, Op->Flags);
}
StoreResult_WithAVXInsert(VectorType, RegClass::FPR, Op, Result);
@@ -2851,14 +2864,19 @@ void OpDispatchBuilder::MOVBetweenGPR_FPR(OpcodeArgs, VectorOpType VectorType) {
Ref Src = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
if (Op->Dest.IsGPR()) {
const auto ElementSize = OpSizeFromDst(Op);
const auto DstSize = std::max(OpSize::i32Bit, OpSizeFromDst(Op));
// Extract element from GPR. Zero extending in the process.
Src = _VExtractToGPR(OpSizeFromSrc(Op), ElementSize, Src, 0);
Src = _VExtractToGPR(OpSizeFromSrc(Op), DstSize, Src, 0);
StoreResultGPR(Op, Op->Dest, Src);
StoreResultGPR_WithOpSize(Op, Op->Dest, Src, DstSize);
} else {
const auto DstSize = std::max(OpSize::i32Bit, OpSizeFromDst(Op));
// Storing first element to memory.
Ref Dest = LoadSourceGPR(Op, Op->Dest, Op->Flags, {.LoadData = false});
_StoreMemFPR(OpSizeFromDst(Op), Dest, Src, OpSize::i8Bit);
_StoreMemFPR(DstSize, Dest, Src, OpSize::i8Bit);
}
}
}
@@ -2878,24 +2896,24 @@ Ref OpDispatchBuilder::VFCMPOpImpl(OpSize Size, IR::OpSize ElementSize, Ref Src1
case VectorCompareType::NLT_US: // NGT(Swapped operand)
case VectorCompareType::NLT_UQ: {
Ref Result = _VFCMPLT(Size, ElementSize, Src1, Src2);
return _VNot(Size, ElementSize, Result);
return _VNot(Size, Result);
}
case VectorCompareType::NLE_US: // NGE(Swapped operand)
case VectorCompareType::NLE_UQ: {
Ref Result = _VFCMPLE(Size, ElementSize, Src1, Src2);
return _VNot(Size, ElementSize, Result);
return _VNot(Size, Result);
}
case VectorCompareType::ORD_Q:
case VectorCompareType::ORD_S: return _VFCMPORD(Size, ElementSize, Src1, Src2);
case VectorCompareType::NGT_UQ:
case VectorCompareType::NGT_US: {
Ref Result = _VFCMPLT(Size, ElementSize, Src2, Src1);
return _VNot(Size, ElementSize, Result);
return _VNot(Size, Result);
}
case VectorCompareType::NGE_UQ:
case VectorCompareType::NGE_US: {
Ref Result = _VFCMPLE(Size, ElementSize, Src2, Src1);
return _VNot(Size, ElementSize, Result);
return _VNot(Size, Result);
}
case VectorCompareType::GT_OQ:
case VectorCompareType::GT_OS: return _VFCMPLT(Size, ElementSize, Src2, Src1);
@@ -2906,10 +2924,10 @@ Ref OpDispatchBuilder::VFCMPOpImpl(OpSize Size, IR::OpSize ElementSize, Ref Src1
// If either of the sources are unordered, then returns true.
Ref Src1_U = _VFCMPEQ(Size, ElementSize, Src1, Src1);
Ref Src2_U = _VFCMPEQ(Size, ElementSize, Src2, Src2);
auto Ordered = _VAnd(Size, ElementSize, Src1_U, Src2_U);
auto Ordered = _VAnd(Size, Src1_U, Src2_U);
Ref Compare_Ordered = _VFCMPEQ(Size, ElementSize, Src1, Src2);
return _VOrn(Size, ElementSize, Compare_Ordered, Ordered);
return _VOrn(Size, Compare_Ordered, Ordered);
}
case VectorCompareType::NEQ_OQ:
case VectorCompareType::NEQ_OS: {
@@ -2918,8 +2936,8 @@ Ref OpDispatchBuilder::VFCMPOpImpl(OpSize Size, IR::OpSize ElementSize, Ref Src1
Ref Src2_U = _VFCMPEQ(Size, ElementSize, Src2, Src2);
Ref Compare_Ordered = _VFCMPEQ(Size, ElementSize, Src1, Src2);
Ref Result = _VAndn(Size, ElementSize, Src1_U, Compare_Ordered);
return _VAnd(Size, ElementSize, Result, Src2_U);
Ref Result = _VAndn(Size, Src1_U, Compare_Ordered);
return _VAnd(Size, Result, Src2_U);
}
case VectorCompareType::FALSE_OQ:
case VectorCompareType::FALSE_OS: return LoadZeroVector(Size);
@@ -3289,7 +3307,7 @@ void OpDispatchBuilder::DefaultX87State(OpcodeArgs) {
// On top of resetting the flags to a default state, we also need to clear
// all of the ST0-7/MM0-7 registers to zero.
Ref ZeroVector = LoadZeroVector(OpSize::i64Bit);
Ref ZeroVector = LoadZeroVector(OpSize::i128Bit);
for (uint32_t i = 0; i < Core::CPUState::NUM_MMS; ++i) {
_StoreContextFPR(OpSize::i128Bit, ZeroVector, MMBaseOffset() + i * 16);
}
@@ -3485,7 +3503,7 @@ Ref OpDispatchBuilder::ADDSUBPOpImpl(OpSize Size, IR::OpSize ElementSize, Ref Sr
} else {
auto ConstantEOR =
LoadAndCacheNamedVectorConstant(Size, ElementSize == OpSize::i32Bit ? NAMED_VECTOR_PADDSUBPS_INVERT : NAMED_VECTOR_PADDSUBPD_INVERT);
auto InvertedSource = _VXor(Size, ElementSize, Src2, ConstantEOR);
auto InvertedSource = _VXor(Size, Src2, ConstantEOR);
return _VFAdd(Size, ElementSize, Src1, InvertedSource);
}
}
@@ -3571,15 +3589,25 @@ void OpDispatchBuilder::PF2IWOp(OpcodeArgs) {
// Float to int32_t
Src = _Vector_FToZS(Size, OpSize::i32Bit, Src);
// We now need to transpose the lower 16-bits of each element together
// Only needing to move the upper element down in this case
Src = _VUnZip(Size, OpSize::i16Bit, Src, Src);
// Truncate the 32-bit integers to 16-bit
// Saturate values outside the 16-bit range to smallest and largest 16-bit values
Src = _VSQXTN(Size, OpSize::i32Bit, Src);
// Now we need to sign extend the 16bit value to 32-bit
Src = _VSXTL(Size, OpSize::i16Bit, Src);
StoreResultFPR_WithOpSize(Op, Op->Dest, Src, Size);
}
void OpDispatchBuilder::PF2IDOp(OpcodeArgs) {
Ref Src = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
const auto Size = OpSizeFromDst(Op);
Src = _Vector_FToZS(Size, OpSize::i32Bit, Src);
StoreResultFPR_WithOpSize(Op, Op->Dest, Src, Size);
}
void OpDispatchBuilder::PMULHRWOp(OpcodeArgs) {
const auto Size = OpSizeFromSrc(Op);
@@ -3764,32 +3792,14 @@ void OpDispatchBuilder::VPMULHWOp(OpcodeArgs, bool Signed) {
}
Ref OpDispatchBuilder::PMULHRSWOpImpl(OpSize Size, Ref Src1, Ref Src2) {
Ref Res {};
if (Size == OpSize::i64Bit) {
// Implementation is more efficient for 8byte registers
Res = _VSMull(Size << 1, OpSize::i16Bit, Src1, Src2);
Res = _VSShrI(Size << 1, OpSize::i32Bit, Res, 14);
auto OneVector = _VectorImm(Size << 1, OpSize::i32Bit, 1);
Res = _VAdd(Size << 1, OpSize::i32Bit, Res, OneVector);
return _VUShrNI(Size << 1, OpSize::i32Bit, Res, 1);
Ref Res = _VSMull(Size << 1, OpSize::i16Bit, Src1, Src2);
return _VRSHRN(Size << 1, OpSize::i32Bit, Res, 15);
} else {
// 128-bit and 256-bit are less efficient
Ref ResultLow;
Ref ResultHigh;
ResultLow = _VSMull(Size, OpSize::i16Bit, Src1, Src2);
ResultHigh = _VSMull2(Size, OpSize::i16Bit, Src1, Src2);
ResultLow = _VSShrI(Size, OpSize::i32Bit, ResultLow, 14);
ResultHigh = _VSShrI(Size, OpSize::i32Bit, ResultHigh, 14);
auto OneVector = _VectorImm(Size, OpSize::i32Bit, 1);
ResultLow = _VAdd(Size, OpSize::i32Bit, ResultLow, OneVector);
ResultHigh = _VAdd(Size, OpSize::i32Bit, ResultHigh, OneVector);
// Combine the results
Res = _VUShrNI(Size, OpSize::i32Bit, ResultLow, 1);
return _VUShrNI2(Size, OpSize::i32Bit, Res, ResultHigh, 1);
Ref ResultLow = _VSMull(Size, OpSize::i16Bit, Src1, Src2);
Ref ResultHigh = _VSMull2(Size, OpSize::i16Bit, Src1, Src2);
return _VRSHRNPair(Size, OpSize::i32Bit, ResultLow, ResultHigh, 15);
}
}
@@ -4349,8 +4359,8 @@ void OpDispatchBuilder::AVXVectorVariableBlend(OpcodeArgs, IR::OpSize ElementSiz
}
void OpDispatchBuilder::PTestOpImpl(OpSize Size, Ref Dest, Ref Src) {
Ref Test1 = _VAnd(Size, OpSize::i8Bit, Dest, Src);
Ref Test2 = _VAndn(Size, OpSize::i8Bit, Src, Dest);
Ref Test1 = _VAnd(Size, Dest, Src);
Ref Test2 = _VAndn(Size, Src, Dest);
// Element size must be less than 32-bit for the sign bit tricks.
Test1 = _VUMaxV(Size, OpSize::i16Bit, Test1);
@@ -4384,11 +4394,11 @@ void OpDispatchBuilder::VTESTOpImpl(OpSize SrcSize, IR::OpSize ElementSize, Ref
Ref Mask = _VDupFromGPR(SrcSize, ElementSize, Constant(MaskConstant));
Ref AndTest = _VAnd(SrcSize, OpSize::i8Bit, Src2, Src1);
Ref AndNotTest = _VAndn(SrcSize, OpSize::i8Bit, Src2, Src1);
Ref AndTest = _VAnd(SrcSize, Src2, Src1);
Ref AndNotTest = _VAndn(SrcSize, Src2, Src1);
Ref MaskedAnd = _VAnd(SrcSize, OpSize::i8Bit, AndTest, Mask);
Ref MaskedAndNot = _VAnd(SrcSize, OpSize::i8Bit, AndNotTest, Mask);
Ref MaskedAnd = _VAnd(SrcSize, AndTest, Mask);
Ref MaskedAndNot = _VAnd(SrcSize, AndNotTest, Mask);
Ref MaxAnd = _VUMaxV(SrcSize, OpSize::i16Bit, MaskedAnd);
Ref MaxAndNot = _VUMaxV(SrcSize, OpSize::i16Bit, MaskedAndNot);
@@ -4491,7 +4501,7 @@ Ref OpDispatchBuilder::DPPOpImpl(IR::OpSize DstSize, Ref Src1, Ref Src2, uint8_t
// Now mask results based on IndexMask.
if (SrcMask != SizeMask) {
auto InputMask = LoadAndCacheIndexedNamedVectorConstant(DstSize, NamedIndexMask, SrcMask * 16);
Temp = _VAnd(DstSize, ElementSize, Temp, InputMask);
Temp = _VAnd(DstSize, Temp, InputMask);
}
// Now due a float reduction
@@ -4913,7 +4923,7 @@ void OpDispatchBuilder::VPERM2Op(OpcodeArgs) {
Ref OpDispatchBuilder::VPERMDIndices(OpSize DstSize, Ref Indices, Ref IndexMask, Ref Repeating3210) {
// Get rid of any junk unrelated to the relevant selector index bits (bits [2:0])
Ref SanitizedIndices = _VAnd(DstSize, OpSize::i8Bit, Indices, IndexMask);
Ref SanitizedIndices = _VAnd(DstSize, Indices, IndexMask);
// Build up the broadcasted index mask. e.g. On x86-64, the selector index
// is always in the lower 3 bits of a 32-bit element. However, in order to
@@ -5108,15 +5118,12 @@ void OpDispatchBuilder::VPERMQOp(OpcodeArgs) {
}
Ref OpDispatchBuilder::VBLENDOpImpl(IR::OpSize VecSize, IR::OpSize ElementSize, Ref Src1, Ref Src2, uint64_t Selector) {
const auto IsWordElements = ElementSize == OpSize::i16Bit;
const auto Is256Bit = VecSize == OpSize::i256Bit;
if (VecSize == OpSize::i256Bit) {
return _VBlendImm(VecSize, ElementSize, Src1, Src2, Selector);
}
const auto ElementsPerLane = uint32_t(IR::NumElements(OpSize::i128Bit, ElementSize));
// PBLENDW uses the same immediate size for 128-bit and 256-bit
// while all the others double in size.
const auto MaskSize = Is256Bit && !IsWordElements ? ElementsPerLane * 2 : ElementsPerLane;
const auto Mask = (1U << MaskSize) - 1;
const auto Mask = (1U << ElementsPerLane) - 1;
// Now, we determine which mask portion has the higher population count.
// we use this to determine which source we use as the base to insert into.
@@ -5133,7 +5140,7 @@ Ref OpDispatchBuilder::VBLENDOpImpl(IR::OpSize VecSize, IR::OpSize ElementSize,
// In the event we tie, then we can just use Src1 and only perform incoming insertions
// that come from Src2.
const auto NumSrc2Bits = uint32_t(std::popcount(Selector & Mask));
const auto NumSrc1Bits = MaskSize - NumSrc2Bits;
const auto NumSrc1Bits = ElementsPerLane - NumSrc2Bits;
const auto IsUsingSrc1 = NumSrc1Bits >= NumSrc2Bits;
Ref Result = IsUsingSrc1 ? Src1 : Src2;
Ref Source = IsUsingSrc1 ? Src2 : Src1;
@@ -5316,7 +5323,7 @@ Ref OpDispatchBuilder::VPERMILRegOpImpl(OpSize DstSize, IR::OpSize ElementSize,
// Sanitize indices first
const auto ShiftAmount = 0b11 >> static_cast<uint32_t>(IsPD);
Ref IndexMask = _VectorImm(DstSize, ElementSize, ShiftAmount);
Ref SanitizedIndices = _VAnd(DstSize, OpSize::i8Bit, Indices, IndexMask);
Ref SanitizedIndices = _VAnd(DstSize, Indices, IndexMask);
Ref IndexTrn1 = _VTrn(DstSize, OpSize::i8Bit, SanitizedIndices, SanitizedIndices);
Ref IndexTrn2 = _VTrn(DstSize, OpSize::i16Bit, IndexTrn1, IndexTrn1);
@@ -5518,7 +5525,7 @@ void OpDispatchBuilder::VFMAddSubImpl(OpcodeArgs, bool AddSub, uint8_t Src1Idx,
LoadAndCacheNamedVectorConstant(Size, ElementSize == OpSize::i32Bit ? NAMED_VECTOR_PSUBADDPS_INVERT : NAMED_VECTOR_PSUBADDPD_INVERT);
}
auto InvertedSourc = _VXor(Size, ElementSize, Sources[AddendIdx - 1], ConstantEOR);
auto InvertedSourc = _VXor(Size, Sources[AddendIdx - 1], ConstantEOR);
Ref Result = _VFMLA(Size, ElementSize, Sources[Src1Idx - 1], Sources[Src2Idx - 1], InvertedSourc);
if (!Is256Bit) {
@@ -5665,7 +5672,7 @@ void OpDispatchBuilder::Extrq_imm(OpcodeArgs) {
const uint64_t Mask = ~0ULL >> (MaskWidth == 0 ? 0 : (64 - MaskWidth));
const Ref MaskVector = _VCastFromGPR(OpSize::i128Bit, OpSize::i64Bit, _Constant(Mask));
Result = _VAnd(OpSize::i128Bit, OpSize::i64Bit, Result, MaskVector);
Result = _VAnd(OpSize::i128Bit, Result, MaskVector);
StoreResult_WithAVXInsert(VectorOpType::SSE, RegClass::FPR, Op, Result);
}
@@ -5681,7 +5688,7 @@ void OpDispatchBuilder::Insertq_imm(OpcodeArgs) {
Ref MaskVector = _VCastFromGPR(OpSize::i128Bit, OpSize::i64Bit, _Constant(Mask));
// Mask incoming source.
Src = _VAnd(OpSize::i64Bit, OpSize::i64Bit, Src, MaskVector);
Src = _VAnd(OpSize::i64Bit, Src, MaskVector);
// If shifting then shift source and mask in to the correct location.
if (Shift) {
@@ -5689,11 +5696,8 @@ void OpDispatchBuilder::Insertq_imm(OpcodeArgs) {
MaskVector = _VShlI(OpSize::i128Bit, OpSize::i64Bit, MaskVector, Shift);
}
// Negate the mask.
MaskVector = _VNot(OpSize::i64Bit, OpSize::i64Bit, MaskVector);
Dest = _VAnd(OpSize::i64Bit, OpSize::i64Bit, Dest, MaskVector);
const Ref Result = _VOr(OpSize::i64Bit, OpSize::i64Bit, Dest, Src);
Dest = _VAndn(OpSize::i64Bit, Dest, MaskVector);
const Ref Result = _VOr(OpSize::i64Bit, Dest, Src);
StoreResult_WithAVXInsert(VectorOpType::SSE, RegClass::FPR, Op, Result);
}
@@ -5710,15 +5714,15 @@ void OpDispatchBuilder::Extrq(OpcodeArgs) {
};
// Bits[5:0] = Mask width in bits
const Ref MaskWidthBits = _VAnd(OpSize::i64Bit, OpSize::i64Bit, Src, ElementMask);
const Ref MaskWidthBits = _VAnd(OpSize::i64Bit, Src, ElementMask);
// Bits[13:8] = Shift right in bits
const Ref ShiftBits = _VAnd(OpSize::i64Bit, OpSize::i64Bit, _VUShrI(OpSize::i64Bit, OpSize::i64Bit, Src, 8), ElementMask);
const Ref ShiftBits = _VAnd(OpSize::i64Bit, _VUShrI(OpSize::i64Bit, OpSize::i64Bit, Src, 8), ElementMask);
// First shift in to the correct position.
Ref Result = _VUShr(OpSize::i64Bit, OpSize::i64Bit, Dest, ShiftBits, false);
Result = _VAnd(OpSize::i128Bit, OpSize::i64Bit, Result, GenerateMask(MaskWidthBits));
Result = _VAnd(OpSize::i128Bit, Result, GenerateMask(MaskWidthBits));
StoreResult_WithAVXInsert(VectorOpType::SSE, RegClass::FPR, Op, Result);
}
@@ -5737,21 +5741,19 @@ void OpDispatchBuilder::Insertq(OpcodeArgs) {
};
// Bits[5:0] = Mask width in bits
const Ref MaskWidthBits = _VAnd(OpSize::i64Bit, OpSize::i64Bit, SelectorBits, ElementMask);
const Ref MaskWidthBits = _VAnd(OpSize::i64Bit, SelectorBits, ElementMask);
// Bits[13:8] = Shift right in bits
const Ref ShiftBits = _VAnd(OpSize::i64Bit, OpSize::i64Bit, _VUShrI(OpSize::i64Bit, OpSize::i64Bit, SelectorBits, 8), ElementMask);
const Ref ShiftBits = _VAnd(OpSize::i64Bit, _VUShrI(OpSize::i64Bit, OpSize::i64Bit, SelectorBits, 8), ElementMask);
// Extract the source data and put in to the correct location
const Ref SrcMask = GenerateMask(MaskWidthBits);
Ref SrcData = _VAnd(OpSize::i128Bit, OpSize::i64Bit, Src, SrcMask);
Ref SrcData = _VAnd(OpSize::i128Bit, Src, SrcMask);
SrcData = _VUShl(OpSize::i128Bit, OpSize::i64Bit, SrcData, ShiftBits, false);
// Generate a destination mask
const Ref DstMask = _VNot(OpSize::i64Bit, OpSize::i64Bit, _VUShl(OpSize::i128Bit, OpSize::i64Bit, SrcMask, ShiftBits, false));
Ref Result = _VAnd(OpSize::i64Bit, OpSize::i64Bit, Dest, DstMask);
Result = _VOr(OpSize::i64Bit, OpSize::i64Bit, Result, SrcData);
Ref Result = _VAndn(OpSize::i64Bit, Dest, _VUShl(OpSize::i128Bit, OpSize::i64Bit, SrcMask, ShiftBits, false));
Result = _VOr(OpSize::i64Bit, Result, SrcData);
StoreResult_WithAVXInsert(VectorOpType::SSE, RegClass::FPR, Op, Result);
}
@@ -139,9 +139,7 @@ void OpDispatchBuilder::FST(OpcodeArgs, IR::OpSize Width) {
void OpDispatchBuilder::FSTToStack(OpcodeArgs) {
const uint8_t Offset = Op->OP & 7;
if (Offset != 0) {
_StoreStackToStack(Offset);
}
_StoreStackToStack(Offset);
if (Op->TableInfo->Flags & X86Tables::InstFlags::FLAGS_POP) {
_PopStackDestroy();
@@ -172,8 +170,9 @@ void OpDispatchBuilder::FIST(OpcodeArgs, bool Truncate) {
Ref IsOverflow = _NZCVSelect01(CondClass::UGE);
// Set Invalid Operation flag if overflow or special value
// The x87 exception flags are sticky. Preserve earlier result
Ref InvalidFlag = _Or(OpSize::i64Bit, IsSpecial, IsOverflow);
SetRFLAG<FEXCore::X86State::X87FLAG_IE_LOC>(InvalidFlag);
SetRFLAG<FEXCore::X86State::X87FLAG_IE_LOC>(_Or(OpSize::i32Bit, GetRFLAG(FEXCore::X86State::X87FLAG_IE_LOC), InvalidFlag));
}
Data = _F80CVTInt(Size, Data, Truncate);
@@ -575,7 +574,7 @@ void OpDispatchBuilder::X87FRSTOR(OpcodeArgs) {
for (int i = 0; i < 7; ++i) {
Ref Reg = _LoadMemFPR(OpSize::i128Bit, Mem, Constant((IR::OpSizeToSize(Size) * 7) + (10 * i)), OpSize::i8Bit, MemOffsetType::SXTX, 1);
// Mask off the top bits
Reg = _VAnd(OpSize::i128Bit, OpSize::i128Bit, Reg, Mask);
Reg = _VAnd(OpSize::i128Bit, Reg, Mask);
if (ReducedPrecisionMode) {
// Convert to double precision
Reg = _F80CVT(OpSize::i64Bit, Reg);
@@ -667,7 +666,6 @@ void OpDispatchBuilder::FCOMI(OpcodeArgs, IR::OpSize Width, bool Integer, OpDisp
SetRFLAG<FEXCore::X86State::X87FLAG_C2_LOC>(HostFlag_Unordered);
SetRFLAG<FEXCore::X86State::X87FLAG_C3_LOC>(HostFlag_ZF);
} else {
// OF, SF, AF, PF all undefined
SetCFDirect(HostFlag_CF);
SetRFLAG<FEXCore::X86State::RFLAG_ZF_RAW_LOC>(HostFlag_ZF);
@@ -675,10 +673,17 @@ void OpDispatchBuilder::FCOMI(OpcodeArgs, IR::OpSize Width, bool Integer, OpDisp
// TODO: This could perhaps be optimized?
auto PF = _Xor(OpSize::i32Bit, HostFlag_Unordered, Constant(1));
SetRFLAG<FEXCore::X86State::RFLAG_PF_RAW_LOC>(PF);
// Intel: OF, SF, and AF set to zero
// AMD: no mention of OF, SF and AF but actual hardware seems to always zero
SetRFLAG<FEXCore::X86State::RFLAG_OF_RAW_LOC>(Constant(0));
SetRFLAG<FEXCore::X86State::RFLAG_SF_RAW_LOC>(Constant(0));
SetRFLAG<FEXCore::X86State::RFLAG_AF_RAW_LOC>(Constant(0));
}
// Set Invalid Operation flag when unordered (NaN comparison)
SetRFLAG<FEXCore::X86State::X87FLAG_IE_LOC>(HostFlag_Unordered);
// The x87 exception flags are sticky. Preserve earlier result
SetRFLAG<FEXCore::X86State::X87FLAG_IE_LOC>(_Or(OpSize::i32Bit, GetRFLAG(FEXCore::X86State::X87FLAG_IE_LOC), HostFlag_Unordered));
if (PopTwice) {
_PopStackDestroy();
@@ -703,7 +708,8 @@ void OpDispatchBuilder::FTST(OpcodeArgs) {
SetRFLAG<FEXCore::X86State::X87FLAG_C3_LOC>(HostFlag_ZF);
// Set Invalid Operation flag when unordered (NaN comparison)
SetRFLAG<FEXCore::X86State::X87FLAG_IE_LOC>(HostFlag_Unordered);
// The x87 exception flags are sticky. Preserve earlier result
SetRFLAG<FEXCore::X86State::X87FLAG_IE_LOC>(_Or(OpSize::i32Bit, GetRFLAG(FEXCore::X86State::X87FLAG_IE_LOC), HostFlag_Unordered));
}
void OpDispatchBuilder::X87OpHelper(OpcodeArgs, FEXCore::IR::IROps IROp, bool ZeroC2) {
@@ -719,6 +725,9 @@ void OpDispatchBuilder::X87ModifySTP(OpcodeArgs, bool Inc) {
} else {
_DecStackTop();
}
// C1 set to 0
SetRFLAG<FEXCore::X86State::X87FLAG_C1_LOC>(Constant(0));
}
// Operations dealing with loading and storing environment pieces
@@ -367,7 +367,7 @@ void OpDispatchBuilder::FCOMIF64(OpcodeArgs, IR::OpSize Width, bool Integer, OpD
} else {
HandleNZCVWrite();
_F80CmpValue(b);
ComissFlags(true /* InvalidateAF */);
ComissFlags();
}
if (PopTwice) {
@@ -0,0 +1,99 @@
// SPDX-License-Identifier: MIT
#include "Interface/Core/LookupCache.h"
#include "Interface/Core/SharedCodeBufferManager.h"
#include <FEXCore/fextl/memory.h>
#include <FEXCore/Utils/AllocatorHooks.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/MathUtils.h>
#ifndef _WIN32
#include <FEXCore/Utils/PrctlUtils.h>
#endif
namespace FEXCore::CPU {
static constexpr size_t INITIAL_CODE_SIZE = 1024 * 1024 * 16;
// We don't want to move above 128MB atm because that means we will have to encode longer jumps
static constexpr size_t MAX_CODE_SIZE = 1024 * 1024 * 128;
CodeBuffer::CodeBuffer(size_t Size)
: AllocatedSize(Size) {
Ptr = static_cast<uint8_t*>(FEXCore::Allocator::VirtualAlloc(Size, true));
LOGMAN_THROW_A_FMT(!!Ptr, "Couldn't allocate code buffer");
// Protect the last page of the allocated buffer to trigger SIGSEGV on write access
uintptr_t LastPageAddr = AlignDown(reinterpret_cast<uintptr_t>(Ptr) + Size - 1, FEXCore::Utils::FEX_PAGE_SIZE);
if (!FEXCore::Allocator::VirtualProtect(reinterpret_cast<void*>(LastPageAddr), FEXCore::Utils::FEX_PAGE_SIZE,
FEXCore::Allocator::ProtectOptions::None)) {
LogMan::Msg::EFmt("Failed to mprotect last page of code buffer.");
}
FEXCore::Allocator::VirtualName("FEXMemJIT", Ptr, Size);
// Huge-pages reduce the amount of iTLB misses dramatically when it works.
FEXCore::Allocator::VirtualTHPControl(Ptr, Size, FEXCore::Allocator::THPControl::Enable);
LookupCache = fextl::make_unique<GuestToHostMap>();
CodeBufferEnd = Ptr + UsableSize();
CodeBufferOffset = Ptr;
}
CodeBuffer::~CodeBuffer() {
FEXCore::Allocator::VirtualFree(Ptr, AllocatedSize);
}
fextl::shared_ptr<CodeBuffer> SharedCodeBufferManager::AllocateNew(size_t Size) {
#ifndef _WIN32
// MDWE (Memory-Deny-Write-Execute) is a new Linux 6.3 feature.
// It's equivalent to systemd's `MemoryDenyWriteExecute` but implemented entirely in the kernel.
//
// MDWE prevents applications from creating RWX memory mappings.
// This prevents FEX from doing anything JIT related, as FEX uses RWX for JIT memory mappings.
//
// A potential workaround to make FEX work with MDWE is to call mprotect every time we need to write or modify code.
// Alternatively, FEX could use a memory mirror where one half is mapped as RW and the other is RX.
//
// Once MDWE is enabled with the prctl, the feature is sealed and it can /NOT/ be turned off.
//
// Status of MDWE is queried through prctl using `PR_GET_MDWE`:
// -1: The kernel doesn't support MDWE
// 0: MDWE is supported but disabled
// >0: MDWE is enabled, hence prohibiting RWX mappings
int MDWE = ::prctl(PR_GET_MDWE, 0, 0, 0, 0);
if (MDWE != -1 && MDWE != 0) {
LogMan::Msg::EFmt("MDWE was set to 0x{:x} which means FEX can't allocate executable memory", MDWE);
}
#endif
auto Buffer = fextl::make_shared<CodeBuffer>(Size);
Latest = Buffer;
OnCodeBufferAllocated(Buffer);
return Buffer;
}
fextl::shared_ptr<CodeBuffer> SharedCodeBufferManager::GetLatest() {
if (!Latest) {
AllocateNew(INITIAL_CODE_SIZE);
}
return Latest;
}
fextl::shared_ptr<CodeBuffer> SharedCodeBufferManager::StartLargerCodeBuffer() {
if (!Latest) {
// Allocate initial CodeBuffer and return it
return GetLatest();
}
auto NewCodeBufferSize = GetLatest()->TotalAllocationSize();
NewCodeBufferSize = std::min<size_t>(NewCodeBufferSize * 2, MAX_CODE_SIZE);
return AllocateNew(NewCodeBufferSize);
}
fextl::shared_ptr<CodeBuffer> SharedCodeBufferManager::StartMaximalCodeBuffer() {
return AllocateNew(MAX_CODE_SIZE);
}
} // namespace FEXCore::CPU
@@ -0,0 +1,135 @@
// SPDX-License-Identifier: MIT
/*
$info$
category: code buffer ~ Thread shared code buffer management
tags: backend|shared
$end_info$
*/
#pragma once
#include <FEXCore/fextl/memory.h>
#include <FEXCore/Utils/MathUtils.h>
#include <FEXCore/Utils/SignalScopeGuards.h>
#include <FEXCore/Utils/TypeDefines.h>
#include <cstddef>
#include <cstdint>
namespace FEXCore {
struct GuestToHostMap;
}
namespace FEXCore::CPU {
struct CodeBuffer {
fextl::unique_ptr<GuestToHostMap> LookupCache;
CodeBuffer(size_t Size);
CodeBuffer(const CodeBuffer&) = delete;
CodeBuffer& operator=(const CodeBuffer&) = delete;
CodeBuffer(CodeBuffer&& oth) = delete;
CodeBuffer& operator=(CodeBuffer&&) = delete;
~CodeBuffer();
// Atomically allocate a fixed size buffer out of the current allocated codebuffer.
// Lockless because it's just a linear allocator.
struct CodeBufferAllocation {
const uint8_t* BufferBase;
uint8_t* BufferAllocationOffset;
};
CodeBufferAllocation AtomicAllocateBuffer(size_t Size) {
Size = FEXCore::AlignUp(Size, 16);
LOGMAN_THROW_A_FMT(reinterpret_cast<uintptr_t>(CodeBufferOffset.load()) % 16 == 0, "Buffer needs to always be 16B aligned!");
auto ExpectedOffset = CodeBufferOffset.load(std::memory_order_relaxed);
auto DesiredOffset = ExpectedOffset + Size;
if (DesiredOffset > CodeBufferEnd) {
// Couldn't fit.
return {};
}
while (!CodeBufferOffset.compare_exchange_strong(ExpectedOffset, DesiredOffset)) {
DesiredOffset = ExpectedOffset + Size;
if (DesiredOffset > CodeBufferEnd) {
// Couldn't fit.
return {};
}
}
// Managed to fit.
return {
.BufferBase = Ptr,
.BufferAllocationOffset = ExpectedOffset,
};
}
// Returns the total number of bytes available for storing code
size_t UsableSize() const {
return AllocatedSize - FEXCore::Utils::FEX_PAGE_SIZE;
}
// Returns the full size of the buffer, including the guard page.
size_t TotalAllocationSize() const {
return AllocatedSize;
}
// Returns the num of bytes currently allocated from the allocator.
size_t AllocatedSpaceUsed() const {
return CodeBufferOffset - Ptr;
}
// Trivially reset the allocator.
void Reset() {
CodeBufferOffset = Ptr;
}
// Returns the base of the buffer.
uint8_t* GetBufferBase() const {
return Ptr;
}
private:
uint8_t* Ptr;
uint8_t* CodeBufferEnd;
size_t AllocatedSize; // including guard page; see UsableSize()
// Code buffer allocation information.
std::atomic<uint8_t*> CodeBufferOffset {};
};
/**
* A manager that coordinates access to the CodeBuffer used for compiling new code across threads.
*
* The CodeBuffer is managed as a partially persistent data structure:
* - Exactly one CodeBuffer is now designated as "active", which means data can be appended to it
* - Lossy modifications to the active CodeBuffer will not invalidate any data in use by other threads (which is what enables save CodeBuffer sharing across threads)
* - Instead, such lossy modifications trigger a new "version" of the data in the modifying thread. Old versions of the CodeBuffer persist as read-only data for use by the other threads.
* - The other threads can update their version of the CodeBuffer. This will decrease the reference count and eventually trigger deallocation of the old version
*/
class SharedCodeBufferManager {
public:
virtual ~SharedCodeBufferManager() = default;
// Get the CodeBuffer that was most recently allocated.
// This is the only CodeBuffer that data may be written to.
fextl::shared_ptr<CodeBuffer> GetLatest();
// Allocate a new CodeBuffer with geometric growth up to an internal maximum.
// Subsequent calls to GetLatest will point to the returned buffer.
fextl::shared_ptr<CodeBuffer> StartLargerCodeBuffer();
// Allocate a new CodeBuffer with maximum internal size.
// Subsequent calls to GetLatest will point to the returned buffer.
fextl::shared_ptr<CodeBuffer> StartMaximalCodeBuffer();
virtual void OnCodeBufferAllocated(const std::shared_ptr<CodeBuffer>&) {};
private:
fextl::shared_ptr<CodeBuffer> Latest;
fextl::shared_ptr<CodeBuffer> AllocateNew(size_t Size);
};
} // namespace FEXCore::CPU
@@ -318,7 +318,7 @@ const std::array<X86InstInfo, MAX_PRIMARY_TABLE_SIZE> BaseOps = []() consteval {
{0x8A, 1, X86InstInfo{"MOV", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0}},
{0x8B, 1, X86InstInfo{"MOV", TYPE_INST, FLAGS_MODRM, 0}},
{0x8C, 1, X86InstInfo{"MOV", TYPE_INST, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x8D, 1, X86InstInfo{"LEA", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_MODRM, 0}},
{0x8D, 1, X86InstInfo{"LEA", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{0x8E, 1, X86InstInfo{"MOV", TYPE_INST, GenFlagsSameSize(SIZE_16BIT) | FLAGS_MODRM, 0}},
{0x8F, 1, X86InstInfo{"POP", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_ZERO_REG | FLAGS_DEBUG_MEM_ACCESS, 0}},
{0x90, 8, X86InstInfo{"XCHG", TYPE_INST, FLAGS_SF_REX_IN_BYTE | FLAGS_SF_SRC_RAX, 0}},
@@ -362,7 +362,7 @@ const std::array<X86InstInfo, MAX_PRIMARY_TABLE_SIZE> BaseOps = []() consteval {
{0xCA, 1, X86InstInfo{"RETF", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_BLOCK_END, 2}},
{0xCB, 1, X86InstInfo{"RETF", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_BLOCK_END, 0}},
{0xCC, 1, X86InstInfo{"INT3", TYPE_INST, FLAGS_BLOCK_END, 0}},
{0xCD, 1, X86InstInfo{"INT", TYPE_INST, DEFAULT_SYSCALL_FLAGS, 1}},
{0xCD, 1, X86InstInfo{"INT", TYPE_INST, FLAGS_NO_OVERLAY | FLAGS_BLOCK_END, 1}},
{0xCE, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_CE] }}},
{0xCF, 1, X86InstInfo{"IRET", TYPE_INST, FLAGS_SETS_RIP | FLAGS_BLOCK_END, 0}},
@@ -26,8 +26,8 @@ enum Secondary_LUT {
constexpr std::array<X86InstInfo[2], ENTRY_MAX> Secondary_ArchSelect_LUT = {{
{
{"SYSCALL", TYPE_INST, DEFAULT_SYSCALL_FLAGS, 0, { .OpDispatch = &IR::OpDispatchBuilder::NOPOp } },
{"SYSCALL", TYPE_INST, DEFAULT_SYSCALL_FLAGS, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::SyscallOp, true> } },
{"SYSCALL", TYPE_INST, FLAGS_NO_OVERLAY | FLAGS_BLOCK_END, 0, { .OpDispatch = &IR::OpDispatchBuilder::NOPOp } },
{"SYSCALL", TYPE_INST, FLAGS_NO_OVERLAY | FLAGS_BLOCK_END, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::SyscallOp, true> } },
},
{
{"PUSH FS", TYPE_INST, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_FS_PREFIX> } },
@@ -303,7 +303,7 @@ constexpr std::array<X86InstInfo, MAX_SECOND_TABLE_SIZE> SecondBaseOps = []() co
{0x3E, 1, X86InstInfo{"CALLBACKRET", TYPE_INST, FLAGS_BLOCK_END | FLAGS_NO_OVERLAY | FLAGS_SETS_RIP, 0}},
// This was originally used by VIA to jump to its alternative instruction set. Used for OP_THUNK
{0x3F, 1, X86InstInfo{"ALTINST", TYPE_INST, FLAGS_BLOCK_END | FLAGS_NO_OVERLAY | FLAGS_SETS_RIP, 0}},
{0x3F, 1, X86InstInfo{"ALTINST", TYPE_INST, FLAGS_BLOCK_END | FLAGS_NO_OVERLAY | FLAGS_SETS_RIP, sizeof(IR::SHA256Sum)}},
#endif
};
@@ -808,7 +808,7 @@ namespace AVX256 {
{OPD(2, 0b01, 0xB6), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VFMAddSubImpl, true, 2, 3, 1>}, // VFMADDSUB
{OPD(2, 0b01, 0xB7), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VFMAddSubImpl, false, 2, 3, 1>}, // VFMSUBADD
{OPD(2, 0b01, 0xDB), 1, &OpDispatchBuilder::AESImcOp},
{OPD(2, 0b01, 0xDB), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AESImcOp, true>},
{OPD(2, 0b01, 0xDC), 1, &OpDispatchBuilder::VAESEncOp},
{OPD(2, 0b01, 0xDD), 1, &OpDispatchBuilder::VAESEncLastOp},
{OPD(2, 0b01, 0xDE), 1, &OpDispatchBuilder::VAESDecOp},
@@ -860,7 +860,7 @@ namespace AVX256 {
{OPD(3, 0b01, 0x62), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPCMPISTRMOp, true>},
{OPD(3, 0b01, 0x63), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VPCMPISTRIOp, true>},
{OPD(3, 0b01, 0xDF), 1, &OpDispatchBuilder::AESKeyGenAssist},
{OPD(3, 0b01, 0xDF), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::AESKeyGenAssist, true>},
};
#undef OPD
@@ -128,6 +128,7 @@ struct DecodedOperand {
RIPRelativeRelocation,
Literal,
LiteralRelocation,
LiteralPatchable,
SIB,
SIBRelocation
};
@@ -159,6 +160,9 @@ struct DecodedOperand {
bool IsLiteralRelocation() const {
return Type == OpType::LiteralRelocation;
}
bool IsLiteralPatchable() const {
return Type == OpType::LiteralPatchable;
}
bool IsSIB() const {
return Type == OpType::SIB;
}
@@ -167,7 +171,7 @@ struct DecodedOperand {
}
uint64_t Literal() const {
LOGMAN_THROW_A_FMT(IsLiteral(), "Precondition: must be a literal");
LOGMAN_THROW_A_FMT(IsLiteral() || IsLiteralPatchable(), "Precondition: must be a literal");
return Data.Literal.Value;
}
@@ -196,6 +200,12 @@ struct DecodedOperand {
int64_t EntrypointOffset;
} LiteralRelocation;
struct {
uint64_t Value;
uint8_t Size;
uint8_t FieldOffset;
uint8_t Width;
} LiteralPatchable;
struct {
int64_t Offset;
uint8_t Scale;
@@ -409,13 +419,6 @@ namespace InstFlags {
constexpr InstFlagType SIZE_256BIT = 0b110;
constexpr InstFlagType SIZE_64BITDEF = 0b111; // Default mode is 64bit instead of typical 32bit
#ifndef _WIN32
constexpr uint32_t DEFAULT_SYSCALL_FLAGS = FLAGS_NO_OVERLAY;
#else
// Syscall ends a block on WIN32 because the instruction can update the CPU's RIP.
constexpr uint32_t DEFAULT_SYSCALL_FLAGS = FLAGS_NO_OVERLAY | FLAGS_BLOCK_END;
#endif
constexpr InstFlagType GetSizeDstFlags(InstFlagType Flags) {
return (Flags >> FLAGS_SIZE_DST_OFF) & SIZE_MASK;
}
+2 -2
View File
@@ -643,7 +643,7 @@ public:
auto IROp = Node.GetNode(BaseList)->Op(IRList);
if (IROp->Op == OP_BEGINBLOCK) {
auto BeginBlock = IROp->C<IROp_EndBlock>();
auto BeginBlock = IROp->C<IROp_BeginBlock>();
Node = BeginBlock->BlockHeader;
} else if (IROp->Op == OP_CODEBLOCK) {
@@ -675,7 +675,7 @@ inline NodeID NodeWrapperBase<Type>::ID() const {
[[nodiscard]]
bool IsBlockExit(FEXCore::IR::IROps Op);
void Dump(fextl::stringstream* out, const IRListView* IR);
void Dump(fextl::ostringstream* out, const IRListView* IR);
constexpr auto format_as(FEXCore::IR::NodeID ID) {
return ID.Value;
+109 -61
View File
@@ -195,13 +195,13 @@
"HasSideEffects": true
},
"GPR = ValidateCode Array16:$CodeOriginal, GPR:$Address, u8:$CodeLength": {
"GPR = ValidateCode GPR:$crc, GPR:$Address, u8:$CodeLength": {
"HasSideEffects": true,
"HasDest": true,
"DestSize": "OpSize::i64Bit"
},
"ThreadRemoveCodeEntry": {
"ThreadRemoveCodeEntry GPR:$Entry": {
"HasSideEffects": true
},
@@ -312,8 +312,8 @@
"HasSideEffects": true,
"RAOverride": "2"
},
"ExitFunction OpSize:#Size, GPR:$NewRIP, BranchHint:$Hint, GPR:$CallReturnAddress, SSA:$CallReturnBlock": {
"Desc": ["Exits the current JIT function with a target RIP"
"ExitFunction OpSize:#Size, GPR:$NewRIP, BranchHint:$Hint, GPR:$CallReturnAddress, SSA:$CallReturnBlock, i64:$PatchSiteAddress{0}, i64:$PatchSiteSize{0}": {
"Desc": ["Exits the current JIT function with a target RIP - optionally patchable from guest bytes for caching"
],
"Inline": ["Any"],
"HasSideEffects": true,
@@ -326,11 +326,10 @@
"CallbackReturn": {
"HasSideEffects": true
},
"GPR = Syscall GPR:$SyscallID, GPR:$Arg0, GPR:$Arg1, GPR:$Arg2, GPR:$Arg3, GPR:$Arg4, GPR:$Arg5": {
"Syscall": {
"HasSideEffects": true,
"Desc": ["Dispatches a guest syscall through to the SyscallHandler class"
],
"DestSize": "OpSize::i64Bit"
]
},
"Thunk GPR:$ArgPtr, SHA256Sum:$ThunkNameHash": {
@@ -954,6 +953,13 @@
]
},
"GPR = PatchableGuestData OpSize:#Size, i64:$Value, i64:$SiteAddress, i64:$SiteSize": {
"Desc": ["Loads Value in a patchable way",
"On disk cache load the value is patched from live guest bytes at SiteAddress"
],
"DestSize": "Size"
},
"GPR = Constant i64:$Constant, ConstPad:$Pad{IR::ConstPad::NoPad}, i32:$MaxBytes{0}": {
"Desc": ["Generates a 64bit constant inside of a GPR",
"Unsupported to create a constant in FPR"
@@ -1864,9 +1870,9 @@
"DestSize": "RegisterSize",
"ElementSize": "ElementSize"
},
"FPR = VNot OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$Vector": {
"FPR = VNot OpSize:#RegisterSize, FPR:$Vector": {
"DestSize": "RegisterSize",
"ElementSize": "ElementSize"
"ElementSize": "OpSize::i8Bit"
},
"FPR = VAbs OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$Vector": {
@@ -2004,15 +2010,6 @@
"BitShift > 0"
]
},
"FPR = VUShraI OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$DestVector, FPR:$Vector, u8:$BitShift": {
"TiedSource": 0,
"DestSize": "RegisterSize",
"ElementSize": "ElementSize",
"EmitValidation": [
"ElementSize >= FEXCore::IR::OpSize::i8Bit && ElementSize <= FEXCore::IR::OpSize::i64Bit",
"BitShift > 0 && BitShift <= IR::OpSizeAsBits(ElementSize)"
]
},
"FPR = VSShrI OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$Vector, u8:$BitShift": {
"TiedSource": 0,
"DestSize": "RegisterSize",
@@ -2025,7 +2022,7 @@
"FPR = VUShrNI OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$Vector, u8:$BitShift": {
"TiedSource": 0,
"Desc": "Unsigned shifts right each element and then narrows to the next lower element size",
"Desc": ["Unsigned shifts right each element and then narrows to the next lower element size"],
"DestSize": "RegisterSize",
"ElementSize": "ElementSize >> 1",
"EmitValidation": [
@@ -2046,8 +2043,32 @@
"BitShift > 0 && BitShift <= IR::OpSizeAsBits(ElementSize)"
]
},
"FPR = VRSHRN OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$Vector, u8:$BitShift": {
"TiedSource": 0,
"Desc": ["Rounding shift right each element and then narrows to the next lower element size",
"Writes result to the bottom half of the destination register, upper half is zeroed"
],
"DestSize": "RegisterSize",
"ElementSize": "ElementSize >> 1",
"EmitValidation": [
"ElementSize >= FEXCore::IR::OpSize::i16Bit && ElementSize <= FEXCore::IR::OpSize::i64Bit",
"BitShift > 0 && BitShift <= IR::OpSizeAsBits(ElementSize)"
]
},
"FPR = VRSHRNPair OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$VectorLower, FPR:$VectorUpper, u8:$BitShift": {
"Desc": ["Rounding shift right and narrow a pair of vectors into one result"],
"DestSize": "RegisterSize",
"ElementSize": "ElementSize >> 1",
"EmitValidation": [
"RegisterSize == FEXCore::IR::OpSize::i128Bit || RegisterSize == FEXCore::IR::OpSize::i256Bit",
"ElementSize >= FEXCore::IR::OpSize::i16Bit && ElementSize <= FEXCore::IR::OpSize::i64Bit",
"BitShift > 0 && BitShift <= IR::OpSizeAsBits(ElementSize)"
]
},
"FPR = VSXTL OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$Vector": {
"Desc": "Sign extends elements from the source element size to the next size up",
"Desc": ["Sign extends elements from the source element size to the next size up"],
"DestSize": "RegisterSize",
"ElementSize": "ElementSize << 1"
},
@@ -2059,7 +2080,7 @@
"ElementSize": "ElementSize << 1"
},
"FPR = VSSHLL OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$Vector, u8:$BitShift{0}": {
"Desc": "Sign extends elements from the source element size to the next size up",
"Desc": ["Sign extends elements from the source element size to the next size up"],
"DestSize": "RegisterSize",
"ElementSize": "ElementSize << 1"
},
@@ -2071,7 +2092,7 @@
"ElementSize": "ElementSize << 1"
},
"FPR = VUXTL OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$Vector": {
"Desc": "Zero extends elements from the source element size to the next size up",
"Desc": ["Zero extends elements from the source element size to the next size up"],
"DestSize": "RegisterSize",
"ElementSize": "ElementSize << 1"
},
@@ -2153,43 +2174,55 @@
"ElementSize": "ElementSize"
},
"FPR = VAnd OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$Vector1, FPR:$Vector2": {
"FPR = VAnd OpSize:#RegisterSize, FPR:$Vector1, FPR:$Vector2": {
"DestSize": "RegisterSize",
"ElementSize": "ElementSize",
"ElementSize": "OpSize::i8Bit",
"EmitValidation": [
"RegisterSize == FEXCore::IR::OpSize::i256Bit || RegisterSize == FEXCore::IR::OpSize::i128Bit || RegisterSize == FEXCore::IR::OpSize::i64Bit"
]
},
"FPR = VAndn OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$Vector1, FPR:$Vector2": {
"FPR = VAndn OpSize:#RegisterSize, FPR:$Vector1, FPR:$Vector2": {
"DestSize": "RegisterSize",
"ElementSize": "ElementSize",
"ElementSize": "OpSize::i8Bit",
"EmitValidation": [
"RegisterSize == FEXCore::IR::OpSize::i256Bit || RegisterSize == FEXCore::IR::OpSize::i128Bit || RegisterSize == FEXCore::IR::OpSize::i64Bit"
]
},
"FPR = VOrn OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$Vector1, FPR:$Vector2": {
"FPR = VOrn OpSize:#RegisterSize, FPR:$Vector1, FPR:$Vector2": {
"DestSize": "RegisterSize",
"ElementSize": "ElementSize",
"ElementSize": "OpSize::i8Bit",
"EmitValidation": [
"RegisterSize == FEXCore::IR::OpSize::i256Bit || RegisterSize == FEXCore::IR::OpSize::i128Bit || RegisterSize == FEXCore::IR::OpSize::i64Bit"
]
},
"FPR = VOr OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$Vector1, FPR:$Vector2": {
"FPR = VOr OpSize:#RegisterSize, FPR:$Vector1, FPR:$Vector2": {
"DestSize": "RegisterSize",
"ElementSize": "ElementSize",
"ElementSize": "OpSize::i8Bit",
"EmitValidation": [
"RegisterSize == FEXCore::IR::OpSize::i256Bit || RegisterSize == FEXCore::IR::OpSize::i128Bit || RegisterSize == FEXCore::IR::OpSize::i64Bit"
]
},
"FPR = VXor OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$Vector1, FPR:$Vector2": {
"FPR = VXor OpSize:#RegisterSize, FPR:$Vector1, FPR:$Vector2": {
"DestSize": "RegisterSize",
"ElementSize": "OpSize::i8Bit",
"EmitValidation": [
"RegisterSize == FEXCore::IR::OpSize::i256Bit || RegisterSize == FEXCore::IR::OpSize::i128Bit || RegisterSize == FEXCore::IR::OpSize::i64Bit"
]
},
"FPR = VXar OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$LHS, FPR:$RHS, u8:$Rotate": {
"Desc": [
"Performs an XOR of corresponding elements and then rotates them right by",
"an amount between [1, ElementSize]"
],
"DestSize": "RegisterSize",
"ElementSize": "ElementSize",
"EmitValidation": [
"RegisterSize == FEXCore::IR::OpSize::i256Bit || RegisterSize == FEXCore::IR::OpSize::i128Bit || RegisterSize == FEXCore::IR::OpSize::i64Bit"
"RegisterSize == IR::OpSize::i256Bit || RegisterSize == IR::OpSize::i128Bit"
]
},
@@ -2214,7 +2247,7 @@
},
"FPR = VAddP OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$VectorLower, FPR:$VectorUpper": {
"Desc": "Does a horizontal pairwise add of elements across the two source vectors",
"Desc": ["Does a horizontal pairwise add of elements across the two source vectors"],
"DestSize": "RegisterSize",
"ElementSize": "ElementSize"
},
@@ -2271,7 +2304,7 @@
"ElementSize": "ElementSize"
},
"FPR = VFAddP OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$VectorLower, FPR:$VectorUpper": {
"Desc": "Does a horizontal pairwise add of elements across the two source vectors with float element types",
"Desc": ["Does a horizontal pairwise add of elements across the two source vectors with float element types"],
"DestSize": "RegisterSize",
"ElementSize": "ElementSize"
},
@@ -2321,29 +2354,27 @@
"ElementSize": "ElementSize << 1"
},
"FPR = VUMull2 OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$Vector1, FPR:$Vector2": {
"Desc": "Multiplies the high elements with size extension",
"Desc": ["Multiplies the high elements with size extension"],
"DestSize": "RegisterSize",
"ElementSize": "ElementSize << 1"
},
"FPR = VSMull2 OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$Vector1, FPR:$Vector2": {
"Desc": "Multiplies the high elements with size extension",
"Desc": ["Multiplies the high elements with size extension"],
"DestSize": "RegisterSize",
"ElementSize": "ElementSize << 1"
},
"FPR = VUMulH OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$Vector1, FPR:$Vector2": {
"Desc": "Wide unsigned multiply returning the high results",
"Desc": ["Wide unsigned multiply returning the high results"],
"DestSize": "RegisterSize",
"ElementSize": "ElementSize"
},
"FPR = VSMulH OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$Vector1, FPR:$Vector2": {
"Desc": "Wide signed multiply returning the high results",
"Desc": ["Wide signed multiply returning the high results"],
"DestSize": "RegisterSize",
"ElementSize": "ElementSize"
},
"FPR = VUABDL OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$Vector1, FPR:$Vector2": {
"Desc": ["Unsigned Absolute Difference Long"
],
"Desc": ["Unsigned Absolute Difference Long"],
"DestSize": "RegisterSize",
"ElementSize": "ElementSize << 1"
},
@@ -2564,6 +2595,24 @@
"DestSize": "RegisterSize",
"ElementSize": "ElementSize",
"TiedSource": 2
},
"FPR = VBlendImm OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$LHS, FPR:$RHS, u16:$Selector": {
"Desc": [
"Functions the same way an immediate blend operation on x86 would.",
"That is: (e.g. using 16-bit elements)",
" if (Selector[0] == 1)",
" Dst[15:0] = RHS[15:0]",
" else",
" Dst[15:0] = LHS[15:0]",
" <etc for the rest of the elements along the vector>",
"",
"Note that like x86, due to the selector size, the operation of this IR op",
"uses a 128-bit lane granularity, so each blending selector independently operates",
"on each 128-bit element that composes the vector."
],
"DestSize": "RegisterSize",
"ElementSize": "ElementSize",
"TiedSource": 0
}
},
"Conv": {
@@ -2601,7 +2650,7 @@
},
"FPR = Vector_SToF OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$Vector": {
"Desc": "Vector op: Converts signed integer to same size float",
"Desc": ["Vector op: Converts signed integer to same size float"],
"DestSize": "RegisterSize",
"ElementSize": "ElementSize"
},
@@ -2613,12 +2662,12 @@
"ElementSize": "ElementSize"
},
"FPR = Vector_FToZS OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$Vector": {
"Desc": "Vector op: Converts float to signed integer, rounding towards zero",
"Desc": ["Vector op: Converts float to signed integer, rounding towards zero"],
"DestSize": "RegisterSize",
"ElementSize": "ElementSize"
},
"FPR = Vector_FToF OpSize:#RegisterSize, OpSize:#DestElementSize, FPR:$Vector, OpSize:$SrcElementSize": {
"Desc": "Vector op: Converts float from source element size to destination size (fp32<->fp64)",
"Desc": ["Vector op: Converts float from source element size to destination size (fp32<->fp64)"],
"DestSize": "RegisterSize",
"ElementSize": "DestElementSize"
},
@@ -2673,75 +2722,74 @@
},
"Crypto": {
"FPR = VAESImc FPR:$Vector": {
"Desc": "Does a stage of the inverse mix column transformation",
"Desc": ["Does a stage of the inverse mix column transformation"],
"DestSize": "OpSize::i128Bit"
},
"FPR = VAESEnc OpSize:#RegisterSize, FPR:$State, FPR:$Key, FPR:$ZeroReg": {
"Desc": "Does a step of AES encryption",
"Desc": ["Does a step of AES encryption"],
"DestSize": "RegisterSize"
},
"FPR = VAESEncLast OpSize:#RegisterSize, FPR:$State, FPR:$Key, FPR:$ZeroReg": {
"Desc": "Does the last step of AES encryption",
"Desc": ["Does the last step of AES encryption"],
"DestSize": "RegisterSize"
},
"FPR = VAESDec OpSize:#RegisterSize, FPR:$State, FPR:$Key, FPR:$ZeroReg": {
"Desc": "Does a step of AES decryption",
"Desc": ["Does a step of AES decryption"],
"DestSize": "RegisterSize"
},
"FPR = VAESDecLast OpSize:#RegisterSize, FPR:$State, FPR:$Key, FPR:$ZeroReg": {
"Desc": "Does the last step of AES decryption",
"Desc": ["Does the last step of AES decryption"],
"DestSize": "RegisterSize"
},
"FPR = VAESKeyGenAssist FPR:$Src, FPR:$KeyGenTBLSwizzle, FPR:$ZeroReg, u8:$RCON": {
"Desc": "Assists in key generation",
"Desc": ["Assists in key generation"],
"DestSize": "OpSize::i128Bit"
},
"FPR = VSha1H FPR:$Src": {
"Desc": "Does vector scalar SHA1H instruction",
"Desc": ["Does vector scalar SHA1H instruction"],
"DestSize": "FEXCore::IR::OpSize::i32Bit"
},
"FPR = VSha1C FPR:$Src1, FPR:$Src2, FPR:$Src3": {
"Desc": "Does vector SHA1C instruction",
"Desc": ["Does vector SHA1C instruction"],
"DestSize": "FEXCore::IR::OpSize::i128Bit",
"TiedSource": 0
},
"FPR = VSha1M FPR:$Src1, FPR:$Src2, FPR:$Src3": {
"Desc": "Does vector SHA1M instruction",
"Desc": ["Does vector SHA1M instruction"],
"DestSize": "FEXCore::IR::OpSize::i128Bit",
"TiedSource": 0
},
"FPR = VSha1P FPR:$Src1, FPR:$Src2, FPR:$Src3": {
"Desc": "Does vector SHA1P instruction",
"Desc": ["Does vector SHA1P instruction"],
"DestSize": "FEXCore::IR::OpSize::i128Bit",
"TiedSource": 0
},
"FPR = VSha1SU1 FPR:$Src1, FPR:$Src2": {
"Desc": "Does vector scalar SHA1H instruction",
"Desc": ["Does vector scalar SHA1H instruction"],
"DestSize": "FEXCore::IR::OpSize::i128Bit",
"TiedSource": 0
},
"FPR = VSha256U0 FPR:$Src1, FPR:$Src2": {
"Desc": "Does vector scalar VSha256U0 instruction",
"Desc": ["Does vector scalar VSha256U0 instruction"],
"DestSize": "FEXCore::IR::OpSize::i128Bit",
"TiedSource": 0
},
"FPR = VSha256U1 FPR:$Src1, FPR:$Src2": {
"Desc": "Does vector scalar VSha256U1 instruction",
"Desc": ["Does vector scalar VSha256U1 instruction"],
"DestSize": "FEXCore::IR::OpSize::i128Bit"
},
"FPR = VSha256H FPR:$Src1, FPR:$Src2, FPR:$Src3": {
"Desc": "Does vector scalar VSha256H instruction",
"Desc": ["Does vector scalar VSha256H instruction"],
"DestSize": "FEXCore::IR::OpSize::i128Bit",
"TiedSource": 0
},
"FPR = VSha256H2 FPR:$Src1, FPR:$Src2, FPR:$Src3": {
"Desc": "Does vector scalar VSha256H2 instruction",
"Desc": ["Does vector scalar VSha256H2 instruction"],
"DestSize": "FEXCore::IR::OpSize::i128Bit",
"TiedSource": 0
},
"GPR = CRC32 GPR:$Src1, GPR:$Src2, OpSize:$SrcSize": {
"Desc": ["CRC32 using polynomial 0x1EDC6F41"
],
"Desc": ["CRC32 using polynomial 0x1EDC6F41"],
"DestSize": "OpSize::i32Bit"
},
"FPR = PCLMUL OpSize:#RegisterSize, FPR:$Src1, FPR:$Src2, u8:$Selector": {
+18 -22
View File
@@ -30,19 +30,19 @@ namespace FEXCore::IR {
#include <FEXCore/IR/IRDefines.inc>
static void PrintArg(fextl::stringstream* out, const IRListView*, const SHA256Sum& Arg) {
static void PrintArg(fextl::ostringstream* out, const IRListView*, const SHA256Sum& Arg) {
*out << fextl::fmt::format("sha256:{:02x}", fmt::join(Arg.data, ""));
}
static void PrintArg(fextl::stringstream* out, const IRListView*, uint64_t Arg) {
static void PrintArg(fextl::ostringstream* out, const IRListView*, uint64_t Arg) {
*out << fextl::fmt::format("#{:#x}", Arg);
}
static void PrintArg(fextl::stringstream* out, const IRListView*, const char* const Arg) {
static void PrintArg(fextl::ostringstream* out, const IRListView*, const char* const Arg) {
*out << fextl::fmt::format("'{}'", Arg);
}
static void PrintArg(fextl::stringstream* out, const IRListView*, CondClass Arg) {
static void PrintArg(fextl::ostringstream* out, const IRListView*, CondClass Arg) {
if (Arg == CondClass::AL) {
*out << "ALWAYS";
return;
@@ -55,7 +55,7 @@ static void PrintArg(fextl::stringstream* out, const IRListView*, CondClass Arg)
*out << CondNames[FEXCore::ToUnderlying(Arg)];
}
static void PrintArg(fextl::stringstream* out, const IRListView*, MemOffsetType Arg) {
static void PrintArg(fextl::ostringstream* out, const IRListView*, MemOffsetType Arg) {
static constexpr std::array<std::string_view, 3> Names = {
"SXTX",
"UXTW",
@@ -65,7 +65,7 @@ static void PrintArg(fextl::stringstream* out, const IRListView*, MemOffsetType
*out << Names[FEXCore::ToUnderlying(Arg)];
}
static void PrintArg(fextl::stringstream* out, const IRListView*, RegClass Arg) {
static void PrintArg(fextl::ostringstream* out, const IRListView*, RegClass Arg) {
*out << [Arg] {
switch (Arg) {
case RegClass::Invalid: return "Invalid";
@@ -79,7 +79,7 @@ static void PrintArg(fextl::stringstream* out, const IRListView*, RegClass Arg)
}();
}
static void PrintArg(fextl::stringstream* out, const IRListView* IR, OrderedNodeWrapper Arg) {
static void PrintArg(fextl::ostringstream* out, const IRListView* IR, OrderedNodeWrapper Arg) {
if (Arg.IsImmediate()) {
auto PhyReg = PhysicalRegister(Arg);
@@ -128,7 +128,7 @@ static void PrintArg(fextl::stringstream* out, const IRListView* IR, OrderedNode
}
}
static void PrintArg(fextl::stringstream* out, const IRListView*, FenceType Arg) {
static void PrintArg(fextl::ostringstream* out, const IRListView*, FenceType Arg) {
*out << [Arg] {
switch (Arg) {
case FenceType::Load: return "Loads";
@@ -140,7 +140,7 @@ static void PrintArg(fextl::stringstream* out, const IRListView*, FenceType Arg)
}();
}
static void PrintArg(fextl::stringstream* out, const IRListView*, RoundMode Arg) {
static void PrintArg(fextl::ostringstream* out, const IRListView*, RoundMode Arg) {
*out << [Arg] {
switch (Arg) {
case RoundMode::Nearest: return "Nearest";
@@ -153,7 +153,7 @@ static void PrintArg(fextl::stringstream* out, const IRListView*, RoundMode Arg)
}();
}
static void PrintArg(fextl::stringstream* out, const IRListView*, ConstPad Arg) {
static void PrintArg(fextl::ostringstream* out, const IRListView*, ConstPad Arg) {
*out << [Arg] {
switch (Arg) {
case ConstPad::NoPad: return "NoPad";
@@ -164,7 +164,7 @@ static void PrintArg(fextl::stringstream* out, const IRListView*, ConstPad Arg)
}();
}
static void PrintArg(fextl::stringstream* out, const IRListView*, NamedVectorConstant Arg) {
static void PrintArg(fextl::ostringstream* out, const IRListView*, NamedVectorConstant Arg) {
*out << [Arg] {
// clang-format off
switch (Arg) {
@@ -260,7 +260,7 @@ static void PrintArg(fextl::stringstream* out, const IRListView*, NamedVectorCon
}();
}
static void PrintArg(fextl::stringstream* out, const IRListView*, IndexNamedVectorConstant Arg) {
static void PrintArg(fextl::ostringstream* out, const IRListView*, IndexNamedVectorConstant Arg) {
*out << [Arg] {
// clang-format off
switch (Arg) {
@@ -286,7 +286,7 @@ static void PrintArg(fextl::stringstream* out, const IRListView*, IndexNamedVect
}();
}
static void PrintArg(fextl::stringstream* out, const IRListView*, OpSize Arg) {
static void PrintArg(fextl::ostringstream* out, const IRListView*, OpSize Arg) {
*out << [Arg] {
switch (Arg) {
case OpSize::iUnsized: return "Unsized";
@@ -303,7 +303,7 @@ static void PrintArg(fextl::stringstream* out, const IRListView*, OpSize Arg) {
}();
}
static void PrintArg(fextl::stringstream* out, const IRListView*, FloatCompareOp Arg) {
static void PrintArg(fextl::ostringstream* out, const IRListView*, FloatCompareOp Arg) {
*out << [Arg] {
switch (Arg) {
case FloatCompareOp::EQ: return "FEQ";
@@ -317,14 +317,14 @@ static void PrintArg(fextl::stringstream* out, const IRListView*, FloatCompareOp
}();
}
static void PrintArg(fextl::stringstream* out, const IRListView*, FEXCore::IR::BreakDefinition Arg) {
static void PrintArg(fextl::ostringstream* out, const IRListView*, FEXCore::IR::BreakDefinition Arg) {
*out << "{" << Arg.ErrorRegister << ".";
*out << static_cast<uint32_t>(Arg.Signal) << ".";
*out << static_cast<uint32_t>(Arg.TrapNumber) << ".";
*out << static_cast<uint32_t>(Arg.si_code) << "}";
}
static void PrintArg(fextl::stringstream* out, const IRListView*, ShiftType Arg) {
static void PrintArg(fextl::ostringstream* out, const IRListView*, ShiftType Arg) {
*out << [Arg] {
switch (Arg) {
case ShiftType::LSL: return "LSL";
@@ -336,7 +336,7 @@ static void PrintArg(fextl::stringstream* out, const IRListView*, ShiftType Arg)
}();
}
static void PrintArg(fextl::stringstream* out, const IRListView*, BranchHint Arg) {
static void PrintArg(fextl::ostringstream* out, const IRListView*, BranchHint Arg) {
*out << [Arg] {
switch (Arg) {
case BranchHint::None: return "None";
@@ -348,11 +348,7 @@ static void PrintArg(fextl::stringstream* out, const IRListView*, BranchHint Arg
}();
}
static void PrintArg(fextl::stringstream* out, const IRListView*, const std::array<uint8_t, 0x10>& Arg) {
*out << fextl::fmt::format("{:02x}", fmt::join(Arg, ""));
}
void Dump(fextl::stringstream* out, const IRListView* IR) {
void Dump(fextl::ostringstream* out, const IRListView* IR) {
auto HeaderOp = IR->GetHeader();
int8_t CurrentIndent = 0;
+3 -4
View File
@@ -33,7 +33,7 @@ bool IsBlockExit(FEXCore::IR::IROps Op) {
}
}
RegClass IREmitter::WalkFindRegClass(Ref Node) {
RegClass IREmitter::WalkFindRegClass(Ref Node) const {
auto Class = GetOpRegClass(Node);
switch (Class) {
case RegClass::GPR:
@@ -45,9 +45,8 @@ RegClass IREmitter::WalkFindRegClass(Ref Node) {
}
// Complex case, needs to be handled on an op by op basis
uintptr_t DataBegin = DualListData.DataBegin();
FEXCore::IR::IROp_Header* IROp = Node->Op(DataBegin);
const uintptr_t DataBegin = DualListData.DataBegin();
const auto* IROp = Node->Op(DataBegin);
switch (IROp->Op) {
case IROps::OP_LOADREGISTER: {
+13 -9
View File
@@ -36,6 +36,10 @@ public:
DualListData.DelayedDisownBuffer();
}
void ValidateDisownedOrFree() const {
DualListData.ValidateDisownedOrFree();
}
IRListView ViewIR() {
return IRListView(&DualListData);
}
@@ -45,7 +49,7 @@ public:
*
* @{ */
RegClass WalkFindRegClass(Ref Node);
RegClass WalkFindRegClass(Ref Node) const;
// These inlining helpers are used by IRDefines.inc so define first.
Ref InlineMem(OpSize Size, Ref Offset, MemOffsetType OffsetType, uint8_t& OffsetScale, bool TSO = false) {
@@ -310,14 +314,14 @@ public:
}
/** @} */
RegClass WalkFindRegClass(OrderedNodeWrapper ssa) {
Ref RealNode = ssa.GetNode(DualListData.ListBegin());
RegClass WalkFindRegClass(OrderedNodeWrapper ssa) const {
auto RealNode = ssa.GetNode(DualListData.ListBegin());
return WalkFindRegClass(RealNode);
}
bool IsValueConstant(OrderedNodeWrapper ssa, uint64_t* Constant = nullptr) {
Ref RealNode = ssa.GetNode(DualListData.ListBegin());
FEXCore::IR::IROp_Header* IROp = RealNode->Op(DualListData.DataBegin());
bool IsValueConstant(OrderedNodeWrapper ssa, uint64_t* Constant = nullptr) const {
auto RealNode = ssa.GetNode(DualListData.ListBegin());
const auto* IROp = RealNode->Op(DualListData.DataBegin());
if (IROp->Op == OP_CONSTANT) {
auto Op = IROp->C<IR::IROp_Constant>();
if (Constant) {
@@ -328,9 +332,9 @@ public:
return false;
}
bool IsValueInlineConstant(OrderedNodeWrapper ssa) {
Ref RealNode = ssa.GetNode(DualListData.ListBegin());
FEXCore::IR::IROp_Header* IROp = RealNode->Op(DualListData.DataBegin());
bool IsValueInlineConstant(OrderedNodeWrapper ssa) const {
auto RealNode = ssa.GetNode(DualListData.ListBegin());
const auto* IROp = RealNode->Op(DualListData.DataBegin());
if (IROp->Op == OP_INLINECONSTANT) {
return true;
}
@@ -129,6 +129,10 @@ public:
PoolObject.DelayedDisownBuffer();
}
void ValidateDisownedOrFree() const {
PoolObject.ValidateDisownedOrFree();
}
private:
Utils::PoolBufferWithTimedRetirement<uintptr_t, 5000, 500> PoolObject;
};
+34 -5
View File
@@ -13,6 +13,7 @@ $end_info$
#include "Interface/IR/Passes/RegisterAllocationPass.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/Profiler.h>
namespace FEXCore::IR {
@@ -66,25 +67,42 @@ void PassManager::Finalize() {
}
}
void PassManager::AddDefaultPasses(FEXCore::Context::ContextImpl* ctx) {
void PassManager::AddDefaultPasses(Context::ContextImpl* ctx) {
FEX_CONFIG_OPT(DisablePasses, O0);
// We only specifically disable optimization passes if desired, as IR output should
// still be well-formed regardless of the modifications made to it.
if (!DisablePasses()) {
InsertPass(CreateX87StackOptimizationPass(ctx->HostFeatures, ctx->Config.Is64BitMode ? IR::OpSize::i64Bit : IR::OpSize::i32Bit));
InsertPass(CreateDeadFlagCalculationEliminination());
}
}
void PassManager::AddDefaultValidationPasses() {
InsertPass(IR::CreateRegisterAllocationPass(&ctx->CPUID), "RA");
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
InsertValidationPass(Validation::CreateIRValidation(), "IRValidation");
#endif
}
void PassManager::InsertRegisterAllocationPass(FEXCore::Context::ContextImpl* ctx) {
InsertPass(IR::CreateRegisterAllocationPass(&ctx->CPUID), "RA");
Pass* PassManager::InsertPass(fextl::unique_ptr<Pass> Pass, const fextl::string& Name) {
auto* PassPtr = InsertAt(Passes.end(), std::move(Pass))->get();
AttemptNameMapping(Name, PassPtr);
return PassPtr;
}
PassManager::PassArrayType::iterator PassManager::InsertAt(PassArrayType::iterator pos, fextl::unique_ptr<Pass> Pass) {
Pass->RegisterPassManager(this);
return Passes.insert(pos, std::move(Pass));
}
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
void PassManager::InsertValidationPass(fextl::unique_ptr<Pass> Pass, const fextl::string& Name) {
Pass->RegisterPassManager(this);
auto* PassPtr = ValidationPasses.emplace_back(std::move(Pass)).get();
AttemptNameMapping(Name, PassPtr);
}
#endif
void PassManager::Run(IREmitter* IREmit) {
FEXCORE_PROFILE_SCOPED("PassManager::Run");
@@ -98,4 +116,15 @@ void PassManager::Run(IREmitter* IREmit) {
}
#endif
}
void PassManager::AttemptNameMapping(const fextl::string& Name, Pass* NewPass) {
if (Name.empty()) {
// Empty name is a 'don't care' case. e.g. Passes that just need to run,
// but don't need to be actively looked up.
return;
}
const auto Result = NameToPassMaping.emplace(Name, NewPass);
LOGMAN_THROW_A_FMT(Result.second, "Tried to insert pass with name '{}'. But name is already used", Name);
}
} // namespace FEXCore::IR
+31 -43
View File
@@ -8,23 +8,18 @@ $end_info$
#pragma once
#include <FEXCore/Config/Config.h>
#include <FEXCore/Utils/ThreadPoolAllocator.h>
#include <FEXCore/fextl/memory.h>
#include <FEXCore/fextl/string.h>
#include <FEXCore/fextl/unordered_map.h>
#include <FEXCore/fextl/vector.h>
#include <functional>
#include <concepts>
#include <utility>
namespace FEXCore::Context {
class ContextImpl;
}
namespace FEXCore::HLE {
class SyscallHandler;
}
namespace FEXCore::IR {
class PassManager;
class IREmitter;
@@ -44,64 +39,57 @@ protected:
class PassManager final {
public:
void AddDefaultPasses(FEXCore::Context::ContextImpl* ctx);
void AddDefaultValidationPasses();
Pass* InsertPass(fextl::unique_ptr<Pass> Pass, fextl::string Name = "") {
auto PassPtr = InsertAt(Passes.end(), std::move(Pass))->get();
if (!Name.empty()) {
NameToPassMaping[Name] = PassPtr;
}
return PassPtr;
explicit PassManager(Context::ContextImpl* CTX) {
AddDefaultPasses(CTX);
}
void InsertRegisterAllocationPass(FEXCore::Context::ContextImpl* ctx);
// Executes all of the passes added to the manager.
// If assertions are enabled, this will also run all validation passes.
void Run(IREmitter* IREmit);
bool HasPass(fextl::string Name) const {
// Inserts a new pass into the manager, optionally also assigning a name to it
// for use in the lookup functions,
Pass* InsertPass(fextl::unique_ptr<Pass> Pass, const fextl::string& Name = "");
// Whether or not a pass with the given name is within the manager.
bool HasPass(const fextl::string& Name) const {
return NameToPassMaping.contains(Name);
}
template<typename T>
T* GetPass(fextl::string Name) {
return dynamic_cast<T*>(NameToPassMaping[Name]);
// Retrieves a pass from the manager that has the given name assigned to it.
// Will return nullptr if the pass doesn't exist.
template<std::derived_from<Pass> T>
T* GetPass(const fextl::string& Name) {
return dynamic_cast<T*>(GetPass(Name));
}
Pass* GetPass(fextl::string Name) {
return NameToPassMaping[Name];
}
void RegisterSyscallHandler(FEXCore::HLE::SyscallHandler* Handler) {
SyscallHandler = Handler;
Pass* GetPass(const fextl::string& Name) {
const auto Iter = NameToPassMaping.find(Name);
if (Iter == NameToPassMaping.end()) {
return nullptr;
}
return Iter->second;
}
// Finalizes the pass manager state and assumes no other passes will be added after called.
// This will reorganize the execution order of the passes if necessary.
void Finalize();
protected:
FEXCore::HLE::SyscallHandler* SyscallHandler {};
private:
void AddDefaultPasses(Context::ContextImpl* ctx);
using PassArrayType = fextl::vector<fextl::unique_ptr<Pass>>;
PassArrayType::iterator InsertAt(PassArrayType::iterator pos, fextl::unique_ptr<Pass> Pass) {
Pass->RegisterPassManager(this);
return Passes.insert(pos, std::move(Pass));
}
PassArrayType::iterator InsertAt(PassArrayType::iterator pos, fextl::unique_ptr<Pass> Pass);
PassArrayType Passes;
fextl::unordered_map<fextl::string, Pass*> NameToPassMaping;
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
fextl::vector<fextl::unique_ptr<Pass>> ValidationPasses;
void InsertValidationPass(fextl::unique_ptr<Pass> Pass, fextl::string Name = "") {
Pass->RegisterPassManager(this);
auto PassPtr = ValidationPasses.emplace_back(std::move(Pass)).get();
if (!Name.empty()) {
NameToPassMaping[Name] = PassPtr;
}
}
void InsertValidationPass(fextl::unique_ptr<Pass> Pass, const fextl::string& Name = "");
#endif
void AttemptNameMapping(const fextl::string& Name, Pass* NewPass);
FEX_CONFIG_OPT(Is64BitMode, IS64BIT_MODE);
FEX_CONFIG_OPT(PassManagerDumpIR, PASSMANAGERDUMPIR);
};
@@ -57,7 +57,7 @@ void IRDumper::Run(IREmitter* IREmit) {
}
if (FD.IsValid() || DumpToLog) {
fextl::stringstream out;
fextl::ostringstream out;
FEXCore::IR::Dump(&out, &IR);
if (FD.IsValid()) {
fextl::fmt::print(FD, "IR-{} 0x{:x}:\n{}\n@@@@@\n", IR.PostRA() ? "post" : "pre", +HeaderOp->OriginalRIP, out.str());
@@ -10,6 +10,7 @@ $end_info$
#include "Interface/IR/IREmitter.h"
#include "Interface/IR/PassManager.h"
#include "Interface/IR/RegisterAllocationData.h"
#include "Interface/IR/Passes.h"
#include "Interface/IR/Passes/IRValidation.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
@@ -46,7 +47,7 @@ void IRValidation::Run(IREmitter* IREmit) {
OffsetToBlockMap.clear();
EntryBlock = nullptr;
uint32_t Count = CurrentIR.GetSSACount();
const auto Count = CurrentIR.GetSSACount();
if (Count > MaxNodes) {
NodeIsLive.Realloc(Count);
}
@@ -59,7 +60,7 @@ void IRValidation::Run(IREmitter* IREmit) {
#endif
for (auto [BlockNode, BlockHeader] : CurrentIR.GetBlocks()) {
auto BlockIROp = BlockHeader->CW<FEXCore::IR::IROp_CodeBlock>();
auto BlockIROp = BlockHeader->C<FEXCore::IR::IROp_CodeBlock>();
LOGMAN_THROW_A_FMT(BlockIROp->Header.Op == OP_CODEBLOCK, "IR type failed to be a code block");
if (!EntryBlock) {
@@ -77,15 +78,15 @@ void IRValidation::Run(IREmitter* IREmit) {
const auto OpSize = IROp->Size;
if (GetHasDest(IROp->Op)) {
HadError |= OpSize == IR::OpSize::iInvalid;
// Does the op have a destination of size 0?
// Does the op have an unsized destination?
if (OpSize == IR::OpSize::iInvalid) {
HadError = true;
Errors << "%" << ID << ": Had destination but with no size" << std::endl;
}
// Does the node have zero uses? Should have been DCE'd
if (CodeNode->GetUses() == 0) {
HadWarning |= true;
HadWarning = true;
Warnings << "%" << ID << ": Destination created but had no uses" << std::endl;
}
@@ -98,27 +99,26 @@ void IRValidation::Run(IREmitter* IREmit) {
// If no register class was assigned
if (AssignedClass == IR::RegClass::Invalid) {
HadError |= true;
HadError = true;
Errors << "%" << ID << ": Had destination but with no register class assigned" << std::endl;
}
// If no physical register was assigned
if (PhyReg.IsInvalid()) {
HadError |= true;
HadError = true;
Errors << "%" << ID << ": Had destination but with no register assigned" << std::endl;
}
// Assigned class wasn't the expected class and it is a non-complex op
if (AssignedClass != ExpectedClass && ExpectedClass != IR::RegClass::Complex) {
HadWarning |= true;
HadWarning = true;
Warnings << "%" << ID << ": Destination had register class " << uint32_t(AssignedClass) << " When register class "
<< uint32_t(ExpectedClass) << " Was expected" << std::endl;
}
}
}
uint8_t NumArgs = IR::GetRAArgs(IROp->Op);
const uint8_t NumArgs = IR::GetRAArgs(IROp->Op);
for (uint32_t i = 0; i < NumArgs; ++i) {
OrderedNodeWrapper Arg = IROp->Args[i];
const auto ArgID = Arg.ID();
@@ -126,8 +126,6 @@ void IRValidation::Run(IREmitter* IREmit) {
continue;
}
IROps Op = CurrentIR.GetOp<IROp_Header>(Arg)->Op;
if (ArgID.IsValid()) {
Uses[ArgID.Value]++;
}
@@ -135,10 +133,11 @@ void IRValidation::Run(IREmitter* IREmit) {
// We do not validate the location of inline constants because it's
// irrelevant, they're ignored by RA and always inlined to where they
// need to be. This lets us pool inline constants globally.
bool Ignore = (Op == OP_IRHEADER || Op == OP_INLINECONSTANT);
const IROps Op = CurrentIR.GetOp<IROp_Header>(Arg)->Op;
const bool Ignore = (Op == OP_IRHEADER || Op == OP_INLINECONSTANT);
if (!Ignore && ArgID.IsValid() && !NodeIsLive.Get(ArgID.Value)) {
HadError |= true;
HadError = true;
Errors << "%" << ID << ": Arg[" << i << "] references invalid %" << ArgID << std::endl;
}
}
@@ -147,7 +146,6 @@ void IRValidation::Run(IREmitter* IREmit) {
switch (IROp->Op) {
case IR::OP_EXITFUNCTION: {
CurrentBlock->HasExit = true;
break;
}
case IR::OP_CONDJUMP: {
@@ -163,7 +161,7 @@ void IRValidation::Run(IREmitter* IREmit) {
const FEXCore::IR::IROp_Header* FalseTargetOp = CurrentIR.GetOp<IROp_Header>(FalseTargetNode);
if (TrueTargetOp->Op != OP_CODEBLOCK) {
HadError |= true;
HadError = true;
Errors << "CondJump %" << ID << ": True Target Jumps to Op that isn't the begining of a block" << std::endl;
} else {
auto Block = OffsetToBlockMap.try_emplace(Op->TrueBlock.ID()).first;
@@ -171,7 +169,7 @@ void IRValidation::Run(IREmitter* IREmit) {
}
if (FalseTargetOp->Op != OP_CODEBLOCK) {
HadError |= true;
HadError = true;
Errors << "CondJump %" << ID << ": False Target Jumps to Op that isn't the begining of a block" << std::endl;
} else {
auto Block = OffsetToBlockMap.try_emplace(Op->FalseBlock.ID()).first;
@@ -187,7 +185,7 @@ void IRValidation::Run(IREmitter* IREmit) {
const FEXCore::IR::IROp_Header* TargetOp = CurrentIR.GetOp<IROp_Header>(TargetNode);
if (TargetOp->Op != OP_CODEBLOCK) {
HadError |= true;
HadError = true;
Errors << "Jump %" << ID << ": Jump to Op that isn't the begining of a block" << std::endl;
} else {
auto Block = OffsetToBlockMap.try_emplace(Op->Header.Args[0].ID()).first;
@@ -204,7 +202,7 @@ void IRValidation::Run(IREmitter* IREmit) {
// Blocks can only have zero (Exit), 1 (Unconditional branch) or 2 (Conditional) successors
size_t NumSuccessors = CurrentBlock->Successors.size();
if (NumSuccessors > 2) {
HadError |= true;
HadError = true;
Errors << "%" << BlockID << " Has " << NumSuccessors << " successors which is too many" << std::endl;
}
@@ -220,7 +218,7 @@ void IRValidation::Run(IREmitter* IREmit) {
{
auto Op = GetOp(CodeCurrent);
if (Op != IR::OP_ENDBLOCK) {
HadError |= true;
HadError = true;
Errors << "%" << BlockID << " Failed to end block with EndBlock" << std::endl;
}
}
@@ -231,7 +229,7 @@ void IRValidation::Run(IREmitter* IREmit) {
{
auto Op = GetOp(CodeCurrent);
if (!IsBlockExit(Op)) {
HadError |= true;
HadError = true;
Errors << "%" << BlockID << " Didn't have a block exit IR op as its last instruction" << std::endl;
}
}
@@ -243,7 +241,7 @@ void IRValidation::Run(IREmitter* IREmit) {
for (uint32_t i = 0; i < CurrentIR.GetSSACount(); i++) {
auto [Node, IROp] = CurrentIR.at(IR::NodeID {i})();
if (Node->NumUses != Uses[i] && IROp->Op != OP_CODEBLOCK && IROp->Op != OP_IRHEADER) {
HadError |= true;
HadError = true;
Errors << "%" << i << " Has " << Uses[i] << " Uses, but reports " << Node->NumUses << std::endl;
}
}
@@ -251,7 +249,7 @@ void IRValidation::Run(IREmitter* IREmit) {
HadWarning = false;
if (HadError || HadWarning) {
fextl::stringstream Out;
fextl::ostringstream Out;
FEXCore::IR::Dump(&Out, &CurrentIR);
if (HadError) {
@@ -8,20 +8,16 @@
namespace FEXCore::IR::Validation {
struct BlockInfo {
bool HasExit;
const OrderedNode* BlockNode;
fextl::vector<OrderedNode*> Predecessors;
fextl::vector<OrderedNode*> Successors;
};
class IRValidation final : public FEXCore::IR::Pass {
public:
~IRValidation();
void Run(IREmitter* IREmit) override;
private:
struct BlockInfo {
fextl::vector<OrderedNode*> Predecessors;
fextl::vector<OrderedNode*> Successors;
};
BitSet<uint64_t> NodeIsLive {};
OrderedNode* EntryBlock {};
@@ -7,6 +7,7 @@ $end_info$
#include "Interface/IR/IR.h"
#include "Interface/IR/IREmitter.h"
#include "Interface/IR/Passes.h"
#include "Interface/IR/PassManager.h"
#include <FEXCore/Core/X86Enums.h>
@@ -196,7 +197,7 @@ unsigned DeadFlagCalculationEliminination::FlagsForCondClassType(CondClass Cond)
}
}
constexpr FlagInfo ClassifyConst(IROps Op) {
static constexpr FlagInfo ClassifyConst(IROps Op) {
switch (Op) {
case OP_ANDWITHFLAGS:
return FlagInfo::Pack({
@@ -332,15 +333,15 @@ constexpr FlagInfo ClassifyConst(IROps Op) {
}
}
constexpr auto FlagInfos = std::invoke([] {
constexpr auto FlagInfos = [] {
std::array<FlagInfo, OP_LAST> ret = {};
for (unsigned i = 0; i < OP_LAST; ++i) {
ret[i] = ClassifyConst((IROps)i);
ret[i] = ClassifyConst(IROps(i));
}
return ret;
});
}();
FlagInfo DeadFlagCalculationEliminination::Classify(IROp_Header* IROp) {
FlagInfo Info = FlagInfos[IROp->Op];
@@ -351,22 +352,22 @@ FlagInfo DeadFlagCalculationEliminination::Classify(IROp_Header* IROp) {
switch (IROp->Op) {
case OP_NZCVSELECT:
case OP_NZCVSELECTINCREMENT: {
auto Op = IROp->CW<IR::IROp_NZCVSelect>();
auto Op = IROp->C<IR::IROp_NZCVSelect>();
return FlagInfo::Pack({.Read = FlagsForCondClassType(Op->Cond)});
}
case OP_NZCVSELECTV: {
auto Op = IROp->CW<IR::IROp_NZCVSelectV>();
auto Op = IROp->C<IR::IROp_NZCVSelectV>();
return FlagInfo::Pack({.Read = FlagsForCondClassType(Op->Cond)});
}
case OP_NEG: {
auto Op = IROp->CW<IR::IROp_Neg>();
auto Op = IROp->C<IR::IROp_Neg>();
return FlagInfo::Pack({.Read = FlagsForCondClassType(Op->Cond)});
}
case OP_CONDJUMP: {
auto Op = IROp->CW<IR::IROp_CondJump>();
auto Op = IROp->C<IR::IROp_CondJump>();
if (!Op->FromNZCV) {
return FlagInfo::Pack({});
}
@@ -376,7 +377,7 @@ FlagInfo DeadFlagCalculationEliminination::Classify(IROp_Header* IROp) {
case OP_CONDSUBNZCV:
case OP_CONDADDNZCV: {
auto Op = IROp->CW<IR::IROp_CondAddNZCV>();
auto Op = IROp->C<IR::IROp_CondAddNZCV>();
return FlagInfo::Pack({
.Read = FlagsForCondClassType(Op->Cond),
.Write = FLAG_NZCV,
@@ -385,7 +386,7 @@ FlagInfo DeadFlagCalculationEliminination::Classify(IROp_Header* IROp) {
}
case OP_RMIFNZCV: {
auto Op = IROp->CW<IR::IROp_RmifNZCV>();
auto Op = IROp->C<IR::IROp_RmifNZCV>();
static_assert(FLAG_N == (1 << 3), "rmif mask lines up with our bits");
static_assert(FLAG_Z == (1 << 2), "rmif mask lines up with our bits");
@@ -399,7 +400,7 @@ FlagInfo DeadFlagCalculationEliminination::Classify(IROp_Header* IROp) {
}
case OP_INVALIDATEFLAGS: {
auto Op = IROp->CW<IR::IROp_InvalidateFlags>();
auto Op = IROp->C<IR::IROp_InvalidateFlags>();
unsigned Flags = 0;
// TODO: Make this translation less silly
@@ -536,7 +537,7 @@ bool DeadFlagCalculationEliminination::ProcessBlock(IREmitter* IREmit, IRListVie
// Initialize the FlagsRead mask according to the exit instruction.
auto [ExitNode, ExitOp] = CodeLast();
if (ExitOp->Op == IR::OP_CONDJUMP) {
auto Op = ExitOp->CW<IR::IROp_CondJump>();
auto Op = ExitOp->C<IR::IROp_CondJump>();
FlagsRead = CFG.Get(Op->TrueBlock)->Flags | CFG.Get(Op->FalseBlock)->Flags;
} else if (ExitOp->Op == IR::OP_JUMP) {
FlagsRead = CFG.Get(ExitOp->Args[0])->Flags;
@@ -643,7 +644,7 @@ void DeadFlagCalculationEliminination::OptimizeParity(IREmitter* IREmit, IRListV
for (auto [CodeNode, IROp] : CurrentIR.GetCode(Block)) {
if (IROp->Op == OP_STOREPF) {
auto Op = IROp->CW<IR::IROp_StorePF>();
auto Op = IROp->C<IR::IROp_StorePF>();
auto Generator = CurrentIR.GetOp<IR::IROp_Header>(Op->Value);
// Determine if we only write 0/1 to the parity flag.
@@ -696,7 +697,7 @@ void DeadFlagCalculationEliminination::Run(IREmitter* IREmit) {
--CodeLast;
auto [ExitNode, ExitOp] = CodeLast();
if (ExitOp->Op == IR::OP_CONDJUMP) {
auto Op = ExitOp->CW<IR::IROp_CondJump>();
auto Op = ExitOp->C<IR::IROp_CondJump>();
CFG.RecordEdge(Block->ID, Op->TrueBlock);
CFG.RecordEdge(Block->ID, Op->FalseBlock);
@@ -33,7 +33,7 @@ namespace {
Ref RegToSSA[32];
};
IR::RegClass GetRegClassFromNode(IR::IRListView* IR, IR::IROp_Header* IROp) {
IR::RegClass GetRegClassFromNode(const IR::IROp_Header* IROp) {
const auto Class = IR::GetRegClass(IROp->Op);
if (Class != IR::RegClass::Complex) {
return Class;
@@ -49,9 +49,14 @@ namespace {
case IR::OP_FILLREGISTER: return IROp->C<IR::IROp_FillRegister>()->Class;
default: return IR::RegClass::Invalid;
}
};
}
} // Anonymous namespace
void RegisterAllocationPass::SetNumPairRegs(uint32_t NumRegs) {
LOGMAN_THROW_A_FMT((NumRegs % 2) == 0, "Number of pair regs must be even. (Given: {})", NumRegs);
PairRegs = NumRegs;
}
class ConstrainedRAPass final : public RegisterAllocationPass {
public:
explicit ConstrainedRAPass(const FEXCore::CPUIDEmu* CPUID)
@@ -85,27 +90,27 @@ private:
// SourcesNextUses is read backwards, this tracks the index
int64_t SourceIndex {};
bool Rematerializable(IROp_Header* IROp) {
static bool Rematerializable(const IROp_Header* IROp) {
return IROp->Op == OP_CONSTANT;
}
Ref InsertFill(Ref Node) {
IROp_Header* IROp = IR->GetOp<IROp_Header>(Node);
const auto* IROp = IR->GetOp<IROp_Header>(Node);
// Remat if we can
if (Rematerializable(IROp)) {
const auto Op = IROp->C<IR::IROp_Constant>();
uint64_t Const = Op->Constant;
const auto* Op = IROp->C<IR::IROp_Constant>();
const uint64_t Const = Op->Constant;
return IREmit->_Constant(Const, Op->Pad, Op->MaxBytes);
}
// Otherwise fill from stack
uint32_t SlotPlusOne = SpillSlots[IR->GetID(Node).Value];
const uint32_t SlotPlusOne = SpillSlots[IR->GetID(Node).Value];
LOGMAN_THROW_A_FMT(SlotPlusOne >= 1, "Node must have been spilled");
const auto RegClass = GetRegClassFromNode(IR, IROp);
const auto RegClass = GetRegClassFromNode(IROp);
return IREmit->_FillRegister(IROp->Size, IROp->ElementSize, SlotPlusOne - 1, RegClass);
};
}
// IP of next-use of each source. IPs are measured from the end of the
// block, so we don't need to size the block up-front.
@@ -113,32 +118,35 @@ private:
bool AnySpilled {};
bool IsValidArg(OrderedNodeWrapper Arg) {
bool IsValidArg(OrderedNodeWrapper Arg) const {
if (Arg.IsInvalid()) {
return false;
}
auto Op = IR->GetOp<IROp_Header>(Arg)->Op;
return Op != OP_INLINECONSTANT && Op != OP_INLINEENTRYPOINTOFFSET;
};
}
RegisterClassData* GetClass(PhysicalRegister Reg) {
return &Classes[Reg.Class];
};
}
const RegisterClassData* GetClass(PhysicalRegister Reg) const {
return &Classes[Reg.Class];
}
uint32_t GetRegBits(PhysicalRegister Reg) {
return 1 << Reg.Reg;
};
static uint32_t GetRegBits(PhysicalRegister Reg) {
return 1U << Reg.Reg;
}
bool IsInRegisterFile(Ref Node) {
bool IsInRegisterFile(Ref Node) const {
auto ID = IR->GetID(Node).Value;
LOGMAN_THROW_A_FMT(ID < SSAToReg.size(), "Only old nodes looked up");
PhysicalRegister Reg = SSAToReg[ID];
RegisterClassData* Class = GetClass(Reg);
const PhysicalRegister Reg = SSAToReg[ID];
const RegisterClassData* Class = GetClass(Reg);
return (Class->Available & GetRegBits(Reg)) == 0 && Class->RegToSSA[Reg.Reg] == Node;
};
}
void FreeReg(PhysicalRegister Reg) {
RegisterClassData* Class = GetClass(Reg);
@@ -147,7 +155,7 @@ private:
LOGMAN_THROW_A_FMT(!(Class->Available & RegBits), "Register double-free");
Class->Available |= RegBits;
};
}
bool HasSource(IROp_Header* I, PhysicalRegister Reg) {
int NumArgs = IR::GetRAArgs(I->Op);
@@ -170,13 +178,13 @@ private:
}
return false;
};
}
Ref DecodeSRANode(const IROp_Header* IROp, Ref Node) {
if (IROp->Op == OP_LOADREGISTER || IROp->Op == OP_LOADPF || IROp->Op == OP_LOADAF) {
return Node;
} else if (IROp->Op == OP_STOREREGISTER) {
auto V = IROp->C<IR::IROp_StorePF>()->Value;
auto V = IROp->C<IR::IROp_StoreRegister>()->Value;
V.ClearKill();
return IR->GetNode(V);
} else if (IROp->Op == OP_STOREPF || IROp->Op == OP_STOREAF) {
@@ -186,9 +194,9 @@ private:
}
return nullptr;
};
}
PhysicalRegister DecodeSRAReg(const IROp_Header* IROp, Ref Node) {
PhysicalRegister DecodeSRAReg(const IROp_Header* IROp, Ref Node) const {
uint8_t FlagOffset = Classes[FEXCore::ToUnderlying(RegClass::GPRFixed)].Count - 2;
if (IROp->Op == OP_STOREREGISTER) {
@@ -207,9 +215,9 @@ private:
return PhysicalRegister {RegClass::GPRFixed, uint8_t(Op->Reg)};
}
}
};
}
bool IsTrivial(Ref Node, const IROp_Header* Header) {
bool IsTrivial(Ref Node, const IROp_Header* Header) const {
switch (Header->Op) {
case OP_ALLOCATEGPR: return true;
case OP_ALLOCATEGPRAFTER: return true;
@@ -320,7 +328,7 @@ private:
// If we already spilled the Candidate, we don't need to spill again.
// Similarly, if we can rematerialize the instruction, we don't spill it.
if (!Spilled && Header->Op != OP_CONSTANT) {
LOGMAN_THROW_A_FMT(Reg.AsRegClass() == GetRegClassFromNode(IR, Header), "Consistent");
LOGMAN_THROW_A_FMT(Reg.AsRegClass() == GetRegClassFromNode(Header), "Consistent");
// SpillSlots allocation is deferred.
if (SpillSlots.empty()) {
@@ -340,7 +348,7 @@ private:
// Now that we've spilled the value, take it out of the register file
FreeReg(Reg);
AnySpilled = true;
};
}
void RemapReg(Ref Node, PhysicalRegister Reg) {
RegisterClassData* Class = GetClass(Reg);
@@ -350,7 +358,7 @@ private:
if (Index < SSAToReg.size()) {
SSAToReg[Index] = Reg;
}
};
}
// Record a given assignment of register Reg to Node.
void SetReg(Ref Node, PhysicalRegister Reg) {
@@ -363,7 +371,7 @@ private:
RemapReg(Node, Reg);
Node->Reg = Reg.Raw;
};
}
// Assign a register for a given Node, spilling if necessary.
void AssignReg(IROp_Header* IROp, IROp_CodeBlock* Block, Ref CodeNode, IROp_Header* Pivot) {
@@ -419,7 +427,7 @@ private:
}
}
RegClass ClassType = GetRegClassFromNode(IR, IROp);
RegClass ClassType = GetRegClassFromNode(IROp);
RegisterClassData* Class = &Classes[FEXCore::ToUnderlying(ClassType)];
// Spill to make room in the register file.
@@ -432,7 +440,7 @@ private:
LOGMAN_THROW_A_FMT(Class->Available != 0, "Post-condition of spilling");
unsigned Reg = std::countr_zero(Class->Available);
SetReg(CodeNode, PhysicalRegister(ClassType, Reg));
};
}
};
void ConstrainedRAPass::AddRegisters(IR::RegClass Class, uint32_t RegisterCount) {
@@ -441,7 +449,7 @@ void ConstrainedRAPass::AddRegisters(IR::RegClass Class, uint32_t RegisterCount)
Classes[FEXCore::ToUnderlying(Class)].Count = RegisterCount;
}
inline bool KillMove(IROp_Header* LastOp, IROp_Header* IROp, Ref LastNode, Ref CodeNode) {
static bool KillMove(const IROp_Header* LastOp, IROp_Header* IROp, Ref LastNode, Ref CodeNode) {
// 32-bit moves in x86_64 are represented as a Bfe, detect them.
if (LastOp->Op == OP_BFE && LastOp->C<IR::IROp_Bfe>()->lsb == 0 && LastOp->C<IR::IROp_Bfe>()->Width == 32) {
auto Op = IROp->Op;
@@ -459,7 +467,7 @@ inline bool KillMove(IROp_Header* LastOp, IROp_Header* IROp, Ref LastNode, Ref C
return LastOp->Op == OP_STOREREGISTER;
}
inline bool IsSignext(const IROp_Header* IROp, OrderedNodeWrapper Src, OpSize Size) {
static bool IsSignext(const IROp_Header* IROp, OrderedNodeWrapper Src, OpSize Size) {
if (IROp->Op == OP_SBFE) {
auto Sbfe = IROp->C<IR::IROp_Sbfe>();
return Sbfe->Width == 1 && Sbfe->lsb == (IR::OpSizeAsBits(Size) - 1) && Sbfe->Src == Src;
@@ -468,7 +476,7 @@ inline bool IsSignext(const IROp_Header* IROp, OrderedNodeWrapper Src, OpSize Si
}
}
inline bool IsZero(const IROp_Header* IROp) {
static bool IsZero(const IROp_Header* IROp) {
return IROp->Op == OP_CONSTANT && IROp->C<IROp_Constant>()->Constant == 0;
}
@@ -6,10 +6,11 @@ $end_info$
*/
#pragma once
#include "Interface/IR/PassManager.h"
#include <cstdint>
#include <memory>
#include <stdint.h>
namespace FEXCore::IR {
enum class RegClass : uint32_t;
@@ -18,6 +19,9 @@ class RegisterAllocationPass : public FEXCore::IR::Pass {
public:
virtual void AddRegisters(RegClass Class, uint32_t RegisterCount) = 0;
void SetNumPairRegs(uint32_t NumRegs);
protected:
// Number of GPRs usable for pairs at start of GPR set. Must be even.
uint32_t PairRegs {};
};
@@ -3,6 +3,7 @@
#include "Interface/Core/Interpreter/Fallbacks/FallbackOpHandler.h"
#include "Interface/IR/IR.h"
#include "Interface/IR/IREmitter.h"
#include "Interface/IR/Passes.h"
#include "Interface/IR/PassManager.h"
#include "FEXCore/IR/IR.h"
#include "FEXCore/Utils/Profiler.h"
@@ -32,14 +33,14 @@
namespace FEXCore::IR {
// FIXME(pmatos): copy from OpcodeDispatcher.h
inline uint32_t MMBaseOffset() {
static uint32_t MMBaseOffset() {
return static_cast<uint32_t>(offsetof(Core::CPUState, mm[0][0]));
}
// Similar helper to the one in OpcodeDispatcher.h except we do not
// need to handle flags, etc.
template<typename T>
void DeriveOp(Ref& RefV, IROps NewOp, IREmitter::IRPair<T> Expr) {
static void DeriveOp(Ref& RefV, IROps NewOp, IREmitter::IRPair<T> Expr) {
Expr.first->Header.Op = NewOp;
RefV = Expr;
}
@@ -52,8 +53,8 @@ template<typename T>
class FixedSizeStack {
public:
struct StackSlotEntry final {
StackSlot Type;
T Value;
StackSlot Type = StackSlot::UNUSED;
T Value = T::Invalid;
};
static constexpr uint8_t size = 8;
@@ -64,8 +65,7 @@ public:
// If SlowPath is true, then TopOffset is always zero.
int8_t TopOffset = 0;
FixedSizeStack()
: buffer(FixedSizeStack::size, {StackSlot::UNUSED, T::Invalid}) {}
FixedSizeStack() = default;
void push(const T& Value) {
rotate();
@@ -92,25 +92,23 @@ public:
return buffer[Offset];
}
void setTop(T Value, size_t Offset = 0) {
void setTop(const T& Value, size_t Offset = 0) {
buffer[Offset] = {StackSlot::VALID, Value};
}
bool isValid(size_t Offset) const {
return buffer[Offset].first;
return buffer[Offset].Type == StackSlot::VALID;
}
void clear() {
for (auto& Elem : buffer) {
Elem = {StackSlot::UNUSED, T::Invalid};
}
buffer.fill({StackSlot::UNUSED, T::Invalid});
TopOffset = 0;
}
void dump() const {
LogMan::Msg::DFmt("-- Stack");
for (size_t i = 0; i < 8; i++) {
for (size_t i = 0; i < buffer.size(); i++) {
const auto& [Valid, Element] = buffer[i];
if (Valid == StackSlot::VALID) {
LogMan::Msg::DFmt("| ST{}: 0x{:x}", i, (uintptr_t)(Element.StackDataNode));
@@ -126,7 +124,7 @@ public:
}
// Returns a mask to set in AbridgedTagWord
uint8_t getValidMask() {
uint8_t getValidMask() const {
uint8_t Mask = 0;
for (size_t i = 0; i < buffer.size(); i++) {
if (buffer[i].Type == StackSlot::VALID) {
@@ -137,7 +135,7 @@ public:
}
// Returns a mask to set in AbridgedTagWord
uint8_t getInvalidMask() {
uint8_t getInvalidMask() const {
uint8_t Mask = 0;
for (size_t i = 0; i < buffer.size(); i++) {
if (buffer[i].Type == StackSlot::INVALID) {
@@ -148,7 +146,7 @@ public:
}
private:
fextl::vector<StackSlotEntry> buffer;
std::array<StackSlotEntry, size> buffer {};
};
class X87StackOptimization final : public Pass {
@@ -201,11 +199,11 @@ private:
}
}
void StoreStackMem_Helper(const IROp_StoreStackMem* Op, Ref StackNode) {
void StoreStackMem_Helper(const IRListView& IR, const IROp_StoreStackMem* Op, Ref StackNode) {
LOGMAN_THROW_A_FMT(!ReducedPrecisionMode, "Full precision mode expected.");
Ref AddrNode = IR->GetNode(Op->Addr);
Ref Offset = IR->GetNode(Op->Offset);
Ref AddrNode = IR.GetNode(Op->Addr);
Ref Offset = IR.GetNode(Op->Offset);
OpSize Align = Op->Align;
MemOffsetType OffsetType = Op->OffsetType;
uint8_t OffsetScale = Op->OffsetScale;
@@ -229,11 +227,11 @@ private:
// Performs a store to memory from a value the stack passed in as StackNode.
// This is the version dealing with the reduced precision case.
void StoreStackMem_Reduced_Helper(const IROp_StoreStackMem* Op, Ref StackNode) {
void StoreStackMem_Reduced_Helper(const IRListView& IR, const IROp_StoreStackMem* Op, Ref StackNode) {
LOGMAN_THROW_A_FMT(ReducedPrecisionMode, "Reduced precision mode expected.");
Ref AddrNode = IR->GetNode(Op->Addr);
Ref Offset = IR->GetNode(Op->Offset);
Ref AddrNode = IR.GetNode(Op->Addr);
Ref Offset = IR.GetNode(Op->Offset);
OpSize Align = Op->Align;
MemOffsetType OffsetType = Op->OffsetType;
uint8_t OffsetScale = Op->OffsetScale;
@@ -274,7 +272,7 @@ private:
Ref LoadStackValueAtOffset_Slow(uint8_t Offset = 0);
void StoreStackValueAtOffset_Slow(Ref Value, uint8_t Offset = 0, bool SetValid = true);
// Update Top value in slow path for a pop
void UpdateTopForPop_Slow();
void UpdateTopForPop_Slow(bool InvalidateTag = true);
void UpdateTopForPush_Slow();
// Synchronizes the current simulated stack with the actual values.
// Returns a new value for Top, that's synchronized between the simulated stack
@@ -292,10 +290,10 @@ private:
void Reset();
struct StackMemberInfo {
StackMemberInfo() = delete;
StackMemberInfo(Ref Data)
constexpr StackMemberInfo() = default;
constexpr StackMemberInfo(Ref Data)
: StackDataNode(Data) {}
StackMemberInfo(Ref Data, Ref Source, OpSize Size)
constexpr StackMemberInfo(Ref Data, Ref Source, OpSize Size)
: StackDataNode(Data)
, Source({Size, Source}) {}
Ref StackDataNode {}; // Reference to the data in the Stack.
@@ -357,9 +355,9 @@ private:
// On the slow path TopCache is always the last obtained version of top.
// TopOffset is ignored
bool SlowPath = false;
// Keeping IREmitter not to pass arguments around
IREmitter* IREmit = nullptr;
IRListView* IR = nullptr;
};
inline const X87StackOptimization::StackMemberInfo X87StackOptimization::StackMemberInfo::Invalid {nullptr};
@@ -575,25 +573,39 @@ void X87StackOptimization::HandleBinopStack(IROps Op64, bool VFOp64, IROps Op80,
HandleBinopValue(Op64, VFOp64, Op80, DestStackOffset, StackOffset2 != DestStackOffset, StackOffset1, StackNode, Reverse);
}
inline void X87StackOptimization::UpdateTopForPop_Slow() {
inline void X87StackOptimization::UpdateTopForPop_Slow(bool InvalidateTag) {
const auto PopContainer = [](auto& container) {
const auto begin = std::begin(container);
std::rotate(begin, std::next(begin), std::end(container));
};
if (InvalidateTag) {
SetX87ValidTag(0, false);
}
// Pop the top of the x87 stack
GetOffsetTopWithCache_Slow(1);
std::rotate(TopOffsetCache.begin(), std::next(TopOffsetCache.begin()), TopOffsetCache.end());
std::rotate(TopOffsetAddressCache.begin(), std::next(TopOffsetAddressCache.begin()), TopOffsetAddressCache.end());
std::rotate(TopValueCache.begin(), std::next(TopValueCache.begin()), TopValueCache.end());
std::rotate(FlushValuesPending.begin(), std::next(FlushValuesPending.begin()), FlushValuesPending.end());
std::rotate(TopValidCache.begin(), std::next(TopValidCache.begin()), TopValidCache.end());
PopContainer(TopOffsetCache);
PopContainer(TopOffsetAddressCache);
PopContainer(TopValueCache);
PopContainer(FlushValuesPending);
PopContainer(TopValidCache);
FlushTopPending = true;
}
inline void X87StackOptimization::UpdateTopForPush_Slow() {
// Pop the top of the x87 stack
const auto PushContainer = [](auto& container) {
const auto end = std::end(container);
std::rotate(std::begin(container), std::prev(end), end);
};
// Push the top of the x87 stack
GetOffsetTopWithCache_Slow(1, true);
std::rotate(TopOffsetCache.begin(), std::prev(TopOffsetCache.end()), TopOffsetCache.end());
std::rotate(TopOffsetAddressCache.begin(), std::prev(TopOffsetAddressCache.end()), TopOffsetAddressCache.end());
std::rotate(TopValueCache.begin(), std::prev(TopValueCache.end()), TopValueCache.end());
std::rotate(FlushValuesPending.begin(), std::prev(FlushValuesPending.end()), FlushValuesPending.end());
std::rotate(TopValidCache.begin(), std::prev(TopValidCache.end()), TopValidCache.end());
PushContainer(TopOffsetCache);
PushContainer(TopOffsetAddressCache);
PushContainer(TopValueCache);
PushContainer(FlushValuesPending);
PushContainer(TopValidCache);
FlushTopPending = true;
}
@@ -724,7 +736,6 @@ void X87StackOptimization::Run(IREmitter* Emit) {
// Initialize IREmit member
IREmit = Emit;
IR = &CurrentIR;
// Run optimization proper
for (auto [BlockNode, BlockHeader] : CurrentIR.GetBlocks()) {
@@ -933,7 +944,6 @@ void X87StackOptimization::Run(IREmitter* Emit) {
UpdateTopForPush_Slow();
StoreStackValueAtOffset_Slow(SourceNode);
} else {
auto* SourceNode = CurrentIR.GetNode(Op->X80Src);
if (Op->OriginalValue.IsInvalid()) {
// No original value to track - just push the converted data
StackData.push(StackMemberInfo {SourceNode});
@@ -1019,11 +1029,11 @@ void X87StackOptimization::Run(IREmitter* Emit) {
}
if (ReducedPrecisionMode) {
StoreStackMem_Reduced_Helper(Op, StackNode);
StoreStackMem_Reduced_Helper(CurrentIR, Op, StackNode);
break;
}
StoreStackMem_Helper(Op, StackNode);
StoreStackMem_Helper(CurrentIR, Op, StackNode);
break;
}
@@ -1031,22 +1041,17 @@ void X87StackOptimization::Run(IREmitter* Emit) {
const auto* Op = IROp->C<IROp_StoreStackToStack>();
auto Offset = Op->StackLocation;
if (Offset != 0) {
auto Value = MigrateToSlowPath_IfInvalid();
auto Value = MigrateToSlowPath_IfInvalid();
// Need to store st0 to stack location - basically a copy.
if (SlowPath) {
StoreStackValueAtOffset_Slow(LoadStackValueAtOffset_Slow(), Offset);
} else {
StackData.setTop(*Value, Offset);
}
// Need to store st0 to stack location - basically a copy.
if (SlowPath) {
StoreStackValueAtOffset_Slow(LoadStackValueAtOffset_Slow(), Offset);
} else {
StackData.setTop(*Value, Offset);
}
break;
}
case OP_POPSTACKDESTROY: {
if (SlowPath) {
SetX87ValidTag(0, false);
}
StackPop();
break;
}
@@ -1067,8 +1072,8 @@ void X87StackOptimization::Run(IREmitter* Emit) {
// Slow path: do actual memory operations
Ref ValueTop = LoadStackValue();
Ref ValueOffset = LoadStackValue(Offset);
StoreStackValue(ValueOffset);
StoreStackValue(ValueTop, Offset);
StoreStackValue(ValueOffset, 0, true);
StoreStackValue(ValueTop, Offset, true);
} else {
// Fast path: swap complete StackMemberInfo preserving Source metadata
StackData.setTop(StackMemberOffset, 0);
@@ -1087,7 +1092,7 @@ void X87StackOptimization::Run(IREmitter* Emit) {
ResultNode = IREmit->_VFNeg(OpSize::i64Bit, OpSize::i64Bit, Value);
} else {
Ref HelperNode = IREmit->_LoadNamedVectorConstant(OpSize::i128Bit, IR::NamedVectorConstant::NAMED_VECTOR_F80_SIGN_MASK);
ResultNode = IREmit->_VXor(OpSize::i128Bit, OpSize::i8Bit, Value, HelperNode);
ResultNode = IREmit->_VXor(OpSize::i128Bit, Value, HelperNode);
}
StoreStackValue(ResultNode);
break;
@@ -1102,7 +1107,7 @@ void X87StackOptimization::Run(IREmitter* Emit) {
} else {
// Intermediate insts
Ref HelperNode = IREmit->_LoadNamedVectorConstant(OpSize::i128Bit, IR::NamedVectorConstant::NAMED_VECTOR_F80_SIGN_MASK);
ResultNode = IREmit->_VAndn(OpSize::i128Bit, OpSize::i8Bit, Value, HelperNode);
ResultNode = IREmit->_VAndn(OpSize::i128Bit, Value, HelperNode);
}
StoreStackValue(ResultNode);
break;
@@ -1172,7 +1177,7 @@ void X87StackOptimization::Run(IREmitter* Emit) {
case OP_INCSTACKTOP: {
if (SlowPath) {
UpdateTopForPop_Slow();
UpdateTopForPop_Slow(false);
} else {
StackData.rotate(false);
}
@@ -1227,8 +1232,6 @@ void X87StackOptimization::Run(IREmitter* Emit) {
SynchronizeStackValues();
FlushCachedRegs();
}
return;
}
fextl::unique_ptr<Pass> CreateX87StackOptimizationPass(const HostFeatures& Features, OpSize GPROpSize) {
+11 -9
View File
@@ -38,15 +38,13 @@ namespace FEXCore::Allocator {
MMAP_Hook mmap {::mmap};
MUNMAP_Hook munmap {::munmap};
uint64_t HostVASize {};
using GLIBC_MALLOC_Hook = void* (*)(size_t, const void* caller);
using GLIBC_REALLOC_Hook = void* (*)(void*, size_t, const void* caller);
using GLIBC_FREE_Hook = void (*)(void*, const void* caller);
fextl::unique_ptr<Alloc::HostAllocator> Alloc64 {};
static fextl::unique_ptr<Alloc::HostAllocator> Alloc64 {};
void* FEX_mmap(void* addr, size_t length, int prot, int flags, int fd, off_t offset) {
static void* FEX_mmap(void* addr, size_t length, int prot, int flags, int fd, off_t offset) {
void* Result = Alloc64->Mmap(addr, length, prot, flags, fd, offset);
if (Result >= (void*)-4096) {
errno = -(uint64_t)Result;
@@ -70,7 +68,7 @@ void VirtualName(const char* Name, void* Ptr, size_t Size) {
}
}
int FEX_munmap(void* addr, size_t length) {
static int FEX_munmap(void* addr, size_t length) {
int Result = Alloc64->Munmap(addr, length);
if (Result != 0) {
@@ -104,9 +102,11 @@ void ClearHooks() {
}
#pragma GCC diagnostic pop
FEX_DEFAULT_VISIBILITY size_t DetermineVASize() {
if (HostVASize) {
return HostVASize;
FEX_DEFAULT_VISIBILITY size_t GetHostVABits() {
static uint64_t HostVABits = 0;
if (HostVABits) {
return HostVABits;
}
static constexpr std::array<uintptr_t, 7> TLBSizes = {
@@ -125,6 +125,7 @@ FEX_DEFAULT_VISIBILITY size_t DetermineVASize() {
::munmap(Ptr, FEXCore::Utils::FEX_PAGE_SIZE);
}
if (Ptr != (void*)~0ULL || errno == EEXIST) {
HostVABits = Bits;
return Bits;
}
}
@@ -273,7 +274,7 @@ fextl::vector<MemoryRegion> StealMemoryRegion(uintptr_t Begin, uintptr_t End) {
}
fextl::vector<MemoryRegion> Setup48BitAllocatorIfExists(size_t PageSize) {
size_t Bits = FEXCore::Allocator::DetermineVASize();
size_t Bits = FEXCore::Allocator::GetHostVABits();
if (Bits < 48) {
return {};
}
@@ -316,6 +317,7 @@ VirtualTHPPtr VirtualTHPControl {VirtualTHPNOP};
void SetupHooks(size_t PageSize, HookPtrs Ptrs) {
VirtualName = Ptrs.VirtualName;
VirtualTHPControl = Ptrs.VirtualTHPControl;
SetupAllocatorHooks(VirtualName);
}
#endif
@@ -7,11 +7,9 @@
#include <FEXCore/Utils/MathUtils.h>
#include <FEXCore/Utils/SignalScopeGuards.h>
#include <FEXCore/Utils/TypeDefines.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/MathUtils.h>
#include <FEXCore/fextl/sstream.h>
#include <FEXHeaderUtils/Syscalls.h>
#include <FEXCore/fextl/memory.h>
#include <FEXCore/fextl/sstream.h>
#include <FEXCore/fextl/vector.h>
#include <algorithm>
@@ -114,24 +112,24 @@ private:
return sizeof(LiveVMARegion) + FEXCore::FlexBitSet<FlexBitElementType>::SizeInBytes(NumElements);
}
static void InitializeVMARegionUsed(LiveVMARegion* Region, size_t AdditionalSize) {
size_t SizeOfLiveRegion =
static void InitializeVMARegionUsed(LiveVMARegion* Region) {
const size_t SizeOfLiveRegion =
FEXCore::AlignUp(LiveVMARegion::GetFEXManagedVMARegionSize(Region->SlabInfo->RegionSize), FEXCore::Utils::FEX_PAGE_SIZE);
size_t SizePlusManagedData = SizeOfLiveRegion + AdditionalSize;
Region->FreeSpace = Region->SlabInfo->RegionSize - SizePlusManagedData;
Region->FreeSpace = Region->SlabInfo->RegionSize - SizeOfLiveRegion;
size_t NumManagedPages = SizePlusManagedData >> FEXCore::Utils::FEX_PAGE_SHIFT;
size_t NumManagedPages = SizeOfLiveRegion >> FEXCore::Utils::FEX_PAGE_SHIFT;
size_t ManagedSize = NumManagedPages << FEXCore::Utils::FEX_PAGE_SHIFT;
// Use madvise to set the full tracking region to zero.
// This ensures unused pages are zero, while not having the backing pages consuming memory.
::madvise(Region->UsedPages.Memory + ManagedSize, (Region->SlabInfo->RegionSize >> FEXCore::Utils::FEX_PAGE_SHIFT) - ManagedSize,
MADV_DONTNEED);
auto* MemoryAsBytes = reinterpret_cast<uint8_t*>(Region->UsedPages.Memory);
const auto TrackingRegionSize = Region->SlabInfo->RegionSize - ManagedSize;
::madvise(MemoryAsBytes + ManagedSize, TrackingRegionSize, MADV_DONTNEED);
// Use madvise to claim WILLNEED on the beginning pages for initial state tracking.
// Improves performance of the following MemClear by not doing a page level fault dance for data necessary to track >170TB of used pages.
::madvise(Region->UsedPages.Memory, ManagedSize, MADV_WILLNEED);
::madvise(MemoryAsBytes, ManagedSize, MADV_WILLNEED);
// Set our reserved pages
Region->UsedPages.MemSet(NumManagedPages);
@@ -154,28 +152,27 @@ private:
FEXCore::ForkableUniqueMutex AllocationMutex;
void DetermineVASize();
LiveVMARegion* MakeRegionActive(ReservedRegionListType::iterator ReservedIterator, uint64_t UsedSize) {
LiveVMARegion* MakeRegionActive(ReservedRegionListType::iterator ReservedIterator) {
ReservedVMARegion* ReservedRegion = *ReservedIterator;
ReservedRegions->erase(ReservedIterator);
// mprotect the new region we've allocated
size_t SizeOfLiveRegion =
const size_t SizeOfLiveRegion =
FEXCore::AlignUp(LiveVMARegion::GetFEXManagedVMARegionSize(ReservedRegion->RegionSize), FEXCore::Utils::FEX_PAGE_SIZE);
size_t SizePlusManagedData = UsedSize + SizeOfLiveRegion;
auto Res = mprotect(reinterpret_cast<void*>(ReservedRegion->Base), SizePlusManagedData, PROT_READ | PROT_WRITE);
auto Res = mprotect(reinterpret_cast<void*>(ReservedRegion->Base), SizeOfLiveRegion, PROT_READ | PROT_WRITE);
LOGMAN_THROW_A_FMT(Res != -1, "Couldn't mprotect region: {} '{}' Likely occurs when running out of memory or Maximum VMAs", errno,
strerror(errno));
FEXCore::Allocator::VirtualName("FEXMem_Misc", reinterpret_cast<void*>(ReservedRegion->Base), SizePlusManagedData);
FEXCore::Allocator::VirtualName("FEXMem_Misc", reinterpret_cast<void*>(ReservedRegion->Base), SizeOfLiveRegion);
LiveVMARegion* LiveRange = new (reinterpret_cast<void*>(ReservedRegion->Base)) LiveVMARegion();
// Copy over the reserved data
LiveRange->SlabInfo = ReservedRegion;
// Initialize VMA
LiveVMARegion::InitializeVMARegionUsed(LiveRange, UsedSize);
LiveVMARegion::InitializeVMARegionUsed(LiveRange);
// Add to our active tracked ranges
auto LiveIter = LiveRegions->emplace_back(LiveRange);
@@ -187,7 +184,7 @@ private:
};
void OSAllocator_64Bit::DetermineVASize() {
size_t Bits = FEXCore::Allocator::DetermineVASize();
size_t Bits = FEXCore::Allocator::GetHostVABits();
uintptr_t Size = 1ULL << Bits;
UPPER_BOUND = Size;
@@ -224,7 +221,7 @@ OSAllocator_64Bit::LiveVMARegion* OSAllocator_64Bit::FindLiveRegionForAddress(ui
uintptr_t RegionEnd = ReservedRegion->Base + ReservedRegion->RegionSize;
if (Addr >= ReservedRegion->Base && AddrEnd < RegionEnd) {
// Found one, let's make it active
LiveRegion = MakeRegionActive(it, 0);
LiveRegion = MakeRegionActive(it);
break;
}
}
@@ -394,7 +391,7 @@ again:
size_t lengthPlusManagedData = length + lengthOfLiveRegion;
for (auto it = ReservedRegions->begin(); it != ReservedRegions->end(); ++it) {
if ((*it)->RegionSize >= lengthPlusManagedData) {
MakeRegionActive(it, 0);
MakeRegionActive(it);
goto again;
}
}
@@ -623,14 +620,14 @@ fextl::unique_ptr<T> make_alloc_unique(FEXCore::Allocator::MemoryRegion& Base, A
fextl::unique_ptr<Alloc::HostAllocator> Create64BitAllocatorWithRegions(fextl::vector<FEXCore::Allocator::MemoryRegion>& Regions) {
// This is a bit tricky as we can't allocate memory safely except from the Regions provided. Otherwise we might overwrite memory pages we
// don't own. Scan the memory regions and find the smallest one.
FEXCore::Allocator::MemoryRegion& Smallest = Regions[0];
for (auto& it : Regions) {
if (it.Size <= Smallest.Size) {
Smallest = it;
FEXCore::Allocator::MemoryRegion* Smallest = &Regions[0];
for (auto& Region : Regions) {
if (Region.Size <= Smallest->Size) {
Smallest = &Region;
}
}
return make_alloc_unique<OSAllocator_64Bit>(Smallest, Regions);
return make_alloc_unique<OSAllocator_64Bit>(*Smallest, Regions);
}
} // namespace Alloc::OSAllocator
+2 -2
View File
@@ -39,10 +39,10 @@ struct FlexBitSet final {
Memory[Element / MinimumSizeBits] &= ~(1ULL << (Element % MinimumSizeBits));
}
void MemClear(size_t Elements) {
memset(Memory, 0, FEXCore::AlignUp(Elements / MinimumSizeBits, MinimumSizeBits));
memset(Memory, 0, SizeInBytes(Elements));
}
void MemSet(size_t Elements) {
memset(Memory, 0xFF, FEXCore::AlignUp(Elements / MinimumSizeBits, MinimumSizeBits));
memset(Memory, 0xFF, SizeInBytes(Elements));
}
// Range scanning results
+23
View File
@@ -1,4 +1,6 @@
// SPDX-License-Identifier: MIT
#include <FEXCore/Utils/PrctlUtils.h>
#ifdef ENABLE_FEX_ALLOCATOR
#include <rpmalloc/rpmalloc.h>
#ifndef _WIN32
@@ -20,16 +22,32 @@
namespace FEXCore::Allocator {
using mmap_hook_type = void* (*)(void* addr, size_t length, int prot, int flags, int fd, off_t offset);
using munmap_hook_type = int (*)(void* addr, size_t length);
using vma_name_hook_type = void (*)(const char* name, const void* address, size_t size);
#ifdef ENABLE_FEX_ALLOCATOR
typedef void* (*rp_mmap_hook_type)(size_t size, size_t alignment, size_t* offset, size_t* mapped_size);
typedef void (*rp_munmap_hook_type)(void* address, size_t offset, size_t mapped_size);
typedef void (*vma_name_hook_type)(const char* name, const void* address, size_t size);
extern "C" rp_mmap_hook_type rp_mmap_hook;
extern "C" rp_munmap_hook_type rp_munmap_hook;
extern "C" vma_name_hook_type rp_name_hook;
#ifndef _WIN32
mmap_hook_type fex_mmap_hook = ::mmap;
munmap_hook_type fex_munmap_hook = ::munmap;
static inline void LocalVirtualName(const char* Name, const void* Ptr, size_t Size) {
#ifndef _WIN32
static bool Supports {true};
if (Supports) {
auto Result = prctl(PR_SET_VMA, PR_SET_VMA_ANON_NAME, Ptr, Size, Name);
if (Result == -1) {
// Disable any additional attempts.
Supports = false;
}
}
#endif
}
#endif
// Assume a 64KB page size until told otherwise.
@@ -172,6 +190,11 @@ void InitializeAllocator(size_t PageSize) {
rpmalloc_initialize_config(&global_interface, &global_config);
rp_mmap_hook = FEX_rp_mmap;
rp_munmap_hook = FEX_rp_memory_unmap;
rp_name_hook = LocalVirtualName;
}
#else
void SetupAllocatorHooks(vma_name_hook_type NameHook) {
rp_name_hook = NameHook;
}
#endif
+2
View File
@@ -1,4 +1,6 @@
// SPDX-License-Identifier: MIT
#include <FEXCore/Utils/CompilerDefs.h>
namespace FEXCore::Assert {
// This function can not be inlined
[[noreturn]]
+3
View File
@@ -6,6 +6,9 @@
namespace FEXCore::UncheckedLongJump {
#if defined(ARCHITECTURE_arm64)
#ifdef __arm64ec__
#pragma clang diagnostic ignored "-Winline-asm"
#endif
[[nodiscard]]
FEX_DEFAULT_VISIBILITY FEX_NAKED uint64_t SetJump(JumpBuf& Buffer) {
__asm volatile(R"(
+3 -3
View File
@@ -7,7 +7,7 @@
#include <unistd.h>
namespace FEXCore::Threads {
static fextl::unique_ptr<FEXCore::Threads::Thread> CreateThread_Default(ThreadFunc Func, void* Arg) {
static fextl::unique_ptr<FEXCore::Threads::Thread> CreateThread_Default(ThreadFunc Func, void* Arg, FEXCore::Threads::Flags Flags) {
ERROR_AND_DIE_FMT("Frontend didn't setup thread creation!");
}
@@ -20,8 +20,8 @@ static FEXCore::Threads::Pointers Ptrs = {
.CleanupAfterFork = CleanupAfterFork_Default,
};
fextl::unique_ptr<FEXCore::Threads::Thread> FEXCore::Threads::Thread::Create(ThreadFunc Func, void* Arg) {
return Ptrs.CreateThread(Func, Arg);
fextl::unique_ptr<FEXCore::Threads::Thread> FEXCore::Threads::Thread::Create(ThreadFunc Func, void* Arg, FEXCore::Threads::Flags Flags) {
return Ptrs.CreateThread(Func, Arg, Flags);
}
void FEXCore::Threads::Thread::CleanupAfterFork() {
+54
View File
@@ -0,0 +1,54 @@
// SPDX-License-Identifier: MIT
#include <FEXCore/Utils/WorkQueueThread.h>
#include <FEXCore/Utils/LogManager.h>
namespace FEXCore {
WorkQueueThread::WorkQueueThread(FEXCore::Threads::Flags ThreadFlags) {
Thread = FEXCore::Threads::Thread::Create(ThreadEntry, this, ThreadFlags);
}
WorkQueueThread::~WorkQueueThread() {
{
std::unique_lock lk {Mutex};
Stop = true;
}
CV.notify_one();
if (Thread && Thread->joinable()) {
Thread->join(nullptr);
}
}
void WorkQueueThread::QueueWork(fextl::unique_ptr<WorkItem> Work) {
{
std::unique_lock lk {Mutex};
Queue.push_back(std::move(Work));
}
CV.notify_one();
}
void WorkQueueThread::ThreadProc() {
while (true) {
fextl::unique_ptr<WorkItem> Work;
{
std::unique_lock lk {Mutex};
while (!(Stop || !Queue.empty())) {
CV.wait(lk);
}
if (Queue.empty()) {
// nothing to do? must be stopping
LOGMAN_THROW_A_FMT(Stop, "WorkQueueThread wakes up empty but no Stop?");
return;
}
Work = std::move(Queue.front());
Queue.pop_front();
}
Work->Run();
// Work is destroyed here
}
}
} // namespace FEXCore
+584
View File
@@ -0,0 +1,584 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <FEXCore/Utils/Allocator.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/MathUtils.h>
#include <atomic>
#include <bit>
#include <cstddef>
#include <cstdint>
#include <cstring>
namespace FEXCore::Utils {
/**
* A bitset that supports allocating contiguous ranges atomically.
* - Lock-free, with the caveat that on contention for > 64-bit it would be faster to acquire a lock.
* - If low-contention then atomic-behaviour wins.
* - Three modes of allocation:
* - 1-bit, single-atomic.
* - <= 64-bit, single-atomic, contained within single word (introduces sparsity).
* - > 64-bit, multiple-atomic, contiguous, roll-back on contiguous allocation failure.
* - Can return failure to allocate even if there is space in certain circumstances.
* - If the allocated size crosses multiple words.
* - Race to allocation caused contention.
* - Remembers last allocation/free for inner-word allocations to improve performance.
* - Large greater than atomic-word scans always scan from the start.
* - Resetting the bitset with clear() is lower cost when page_size=true.
* - MADV_DONTNEED replaces pages with zero-page
* - When page_size=false, basic memset is also fairly quick.
*/
template<bool track_last_allocation = false, bool page_sized = true>
class atomic_bitset final {
public:
void init(void* ptr, size_t bits) {
LOGMAN_THROW_A_FMT(bits != 0, "Can't init zero");
base = reinterpret_cast<uint64_t*>(ptr);
bits_to_track = bits;
words_to_track = bits / WORD_SIZE_BITS;
LOGMAN_THROW_A_FMT(bits % WORD_SIZE_BITS == 0, "Bits to track must match uint64_t");
if constexpr (page_sized) {
LOGMAN_THROW_A_FMT(bits % (4096 * 8) == 0, "Bits to track must match bit count in page");
}
last_allocation_track.set_last_allocation(0);
}
// Allocate a contiguous buffer of bits.
// Returns initial bit offset on success, ~0ULL on failure.
size_t allocate(size_t count) {
LOGMAN_THROW_A_FMT(count != 0, "Can't allocate zero");
LOGMAN_THROW_A_FMT(count <= bits_to_track, "Can't allocate larger than size");
if (count == 1) [[likely]] {
// Common and trivial case.
return allocate_one(last_allocation_track.get_last_allocation(), words_to_track);
} else if (count <= WORD_SIZE_BITS) [[likely]] {
// Allocate up to a single word. Don't allow cross-word allocations
// Could cause some sparsity
return allocate_inside_word(count, last_allocation_track.get_last_allocation(), words_to_track);
}
// TODO: Always scans from beginning to end.
// Support iterative scanning.
return allocate_large_amount(count, 0, words_to_track);
}
// Frees a contiguous set of bits.
void free(size_t index, size_t count) {
LOGMAN_THROW_A_FMT(count != 0, "Can't free zero");
LOGMAN_THROW_A_FMT(index < bits_to_track, "Can't free beyond end");
if (count == 1) [[likely]] {
free_one(index);
return;
} else if ((index % WORD_SIZE_BITS + count) <= WORD_SIZE_BITS) [[likely]] {
free_inside_word(index, count);
return;
}
free_large_amount(index, count);
}
// Clears the entire bitset.
// Not thread safe!
void clear() {
const size_t bytes = words_to_track * sizeof(uint64_t);
if constexpr (page_sized) {
// VirtualDontNeed replaces pages with zero page.
FEXCore::Allocator::VirtualDontNeed(base, bytes);
} else {
memset(base, 0, bytes);
}
last_allocation_track.set_last_allocation(0);
}
// Checks if a single bit is set.
bool is_set(size_t index) const {
const size_t word_index = index / WORD_SIZE_BITS;
const size_t word_offset = index % WORD_SIZE_BITS;
auto word_atomic = std::atomic_ref<uint64_t>(base[word_index]);
const uint64_t bit_mask = 1ULL << word_offset;
return (word_atomic.load() & bit_mask) != 0;
}
size_t size_in_bits() const {
return bits_to_track;
}
constexpr static size_t invalid() {
return ~0ULL;
}
// Debug interface
// non-atomically returns the number of set bits in the bitset.
size_t popcount() const {
size_t count {};
// Just ensure all store are visible.
std::atomic_thread_fence(std::memory_order_release);
for (size_t word_index = 0; word_index < words_to_track; ++word_index) {
auto word_atomic = std::atomic_ref<uint64_t>(base[word_index]);
count += std::popcount(word_atomic.load(std::memory_order_relaxed));
}
return count;
}
private:
uint64_t* base {};
size_t bits_to_track {};
size_t words_to_track {};
struct data_to_track_nop {
constexpr static size_t get_last_allocation() {
return 0;
}
constexpr static void set_last_allocation(size_t) {}
};
struct data_to_track {
std::atomic<size_t> last_allocation_word {};
constexpr size_t get_last_allocation() const {
// It's okay if this isn't up to date, full scan of the region still occurs.
return last_allocation_word.load(std::memory_order_relaxed);
}
constexpr void set_last_allocation(size_t word) {
last_allocation_word = word;
}
};
using data_type = typename std::conditional<track_last_allocation, data_to_track, data_to_track_nop>::type;
data_type last_allocation_track {};
constexpr static size_t WORD_SIZE_BITS = sizeof(uint64_t) * 8;
size_t allocate_one(size_t beginning_word_index, size_t ending_word_index) {
// Trivial spin.
for (size_t i = beginning_word_index; i < ending_word_index; ++i) {
auto word_atomic = std::atomic_ref<uint64_t>(base[i]);
auto expected_word = word_atomic.load(std::memory_order_relaxed);
if (expected_word == ~0ULL) {
// Won't pass.
continue;
}
// Spin on the word trying to acquire a bit.
// Uncontended case should immediately succeed.
// Contended case can spin the whole word and lose every acquire.
do {
const auto zero_bit = std::countr_one(expected_word);
const auto bit_mask = 1ULL << zero_bit;
// If mask was already set, then we raced to acquire (returned value will be 1).
// If mask not set, then we will have acquired (returned value will be 0).
expected_word = word_atomic.fetch_or(bit_mask);
if ((expected_word & bit_mask) == 0) {
// Acquired the bit, return the offset.
last_allocation_track.set_last_allocation(i);
return i * WORD_SIZE_BITS + zero_bit;
}
// Bit was already acquired.
expected_word |= bit_mask;
} while (expected_word != ~0ULL);
}
if constexpr (track_last_allocation) {
if (beginning_word_index) {
// One more chance to get an allocation.
// Scan before the previous allocation to see if any free slots have appeared.
return allocate_one(0, beginning_word_index);
}
}
// Failure to acquire here.
return invalid();
}
size_t allocate_inside_word(size_t count, size_t beginning_word_index, size_t ending_word_index) {
for (size_t i = beginning_word_index; i < ending_word_index; ++i) {
auto word_atomic = std::atomic_ref<uint64_t>(base[i]);
auto expected_word = word_atomic.load(std::memory_order_relaxed);
// Spin on the word trying to acquire a bit.
// Uncontended case should immediately succeed.
// Contended case can spin the whole word and lose every acquire.
while (expected_word != ~0ULL) {
auto zero_bit = std::countr_one(expected_word);
bool fits = false;
uint64_t bit_mask = count == WORD_SIZE_BITS ? ~0ULL : ((1ULL << count) - 1);
for (; (zero_bit + count) <= WORD_SIZE_BITS; ++zero_bit) {
// Check if the bits fit.
uint64_t tmp_bit_mask = bit_mask << zero_bit;
if ((expected_word & tmp_bit_mask) == 0) {
fits = true;
break;
}
}
// Could never fit, early exit.
if (!fits) {
break;
}
// Shift bit_mask to the desired location.
bit_mask <<= zero_bit;
uint64_t desired_word {};
bool acquired = true;
do {
if (expected_word & bit_mask) {
// Couldn't acquire this field, move to the next.
acquired = false;
break;
}
// We desire setting a single bit.
desired_word = expected_word | bit_mask;
} while (!word_atomic.compare_exchange_strong(expected_word, desired_word));
if (acquired) {
// Acquired the bit, return the offset.
last_allocation_track.set_last_allocation(i);
return i * WORD_SIZE_BITS + zero_bit;
}
}
}
if constexpr (track_last_allocation) {
if (beginning_word_index) {
// One more chance to get an allocation.
// Scan before the previous allocation to see if any free slots have appeared.
return allocate_inside_word(count, 0, beginning_word_index);
}
}
// Failure to acquire here.
return invalid();
}
size_t allocate_large_amount(size_t count, size_t beginning_word_index, size_t ending_word_index) {
// This version of the code needs to explicitly deal with large allocations that can't fit in a word.
// So we are always scanning minimum 2 words.
const size_t num_words_to_scan = FEXCore::AlignUpPowerOf2(count, WORD_SIZE_BITS) / WORD_SIZE_BITS;
const size_t last_word_to_scan = ending_word_index - num_words_to_scan - 1;
// Scan forward to find the first word.
for (size_t base_index = beginning_word_index; base_index < last_word_to_scan;) {
uint64_t leading_zeros {};
size_t center_word_index = 1;
bool has_center {};
size_t tail_bits {};
size_t remaining_bits = count;
auto check_head_fitment = [&]() -> bool {
auto base_word_atomic = std::atomic_ref<uint64_t>(base[base_index]);
auto base_expected_word = base_word_atomic.load(std::memory_order_relaxed);
// Count the leading zeros, if it is above zero then we can start here.
leading_zeros = std::countl_zero(base_expected_word);
if (leading_zeros == 0) {
// Nope.
++base_index;
return false;
}
return true;
};
auto check_center_fitment = [&]() -> bool {
// Subtract the number of zeros.
remaining_bits -= leading_zeros;
has_center = remaining_bits >= WORD_SIZE_BITS;
while (remaining_bits >= WORD_SIZE_BITS) {
// All words in-between head and tail must be zero.
auto center_word_atomic = std::atomic_ref<uint64_t>(base[base_index + center_word_index]);
auto center_expected_word = center_word_atomic.load(std::memory_order_relaxed);
if (center_expected_word != 0) {
// Couldn't fit, won't ever fit in this range, so jump ahead to the last scanned item.
// Might still be able to start on the tail of this word.
base_index += center_word_index;
return false;
}
++center_word_index;
remaining_bits -= WORD_SIZE_BITS;
}
return true;
};
auto check_tail_fitment = [&]() -> bool {
tail_bits = remaining_bits;
if (tail_bits) {
// Now for the tail (if it is necessary).
size_t tail_index = center_word_index;
// Count the trailing zeros, if it fits out remaining bits then we can try and allocate.
auto tail_word_atomic = std::atomic_ref<uint64_t>(base[base_index + tail_index]);
auto tail_expected_word = tail_word_atomic.load(std::memory_order_relaxed);
const auto trailing_zeros = std::countr_zero(tail_expected_word);
if (trailing_zeros < remaining_bits) {
// Couldn't fit, but also won't ever fit in this range. Jump ahead to this tail item.
// Might still be able to start on the tail of this word.
base_index += tail_index;
return false;
}
}
return true;
};
if (!check_head_fitment()) {
continue;
}
if (!check_center_fitment()) {
continue;
}
if (!check_tail_fitment()) {
continue;
}
// We can try fitting!
if (attempt_allocate_range_from_base(base_index, count, leading_zeros, tail_bits, has_center)) {
const size_t head_leading_offset = (WORD_SIZE_BITS - leading_zeros);
return base_index * WORD_SIZE_BITS + head_leading_offset;
}
// Failure to fit here means we could never fit in this full range. Jump past all the bits
// Rescanning the tail to avoid fragmented sparsity on the tails.
base_index += num_words_to_scan - 1;
}
return invalid();
}
bool attempt_allocate_range_from_base(size_t base_index, size_t count, size_t head_bits, size_t tail_bits, bool has_center) {
const bool has_tail = tail_bits != 0;
const uint64_t head_bit_mask = head_bits == WORD_SIZE_BITS ? ~0ULL : (((1ULL << head_bits) - 1) << (WORD_SIZE_BITS - head_bits));
const uint64_t tail_bit_mask = (1ULL << tail_bits) - 1;
const uint64_t center_words_count = (count - head_bits - tail_bits) / WORD_SIZE_BITS;
const uint64_t center_words_base_index = base_index + 1;
const uint64_t tail_word_base_index = center_words_base_index + center_words_count;
// Three distinct sections, all of which need to support rewinding.
// - Head: Setting the leading zeros to one
// - Always exists. Can be a full word, or partial.
// - Center: Setting all in-between words to ~0ULL
// - Might not exist if tail is smaller than a word
// - Always full words if it does exist.
// - Tail: Set all trailing zeros up to the size to 1
// - Might not exist if Center perfectly aligned to word edge.
// - Always partial words, otherwise it would be considered "Center".
bool set_head {true};
bool set_center {true};
bool set_tail {true};
size_t num_center_set {};
// Head first.
auto set_head_word = [&]() -> bool {
auto head_word_atomic = std::atomic_ref<uint64_t>(base[base_index]);
auto head_expected_word = head_word_atomic.load(std::memory_order_relaxed);
uint64_t desired_word {};
do {
if (head_expected_word & head_bit_mask) {
// Another thread raced and allocated.
return false;
}
// Set the whole mask in one atomic operation.
desired_word = head_expected_word | head_bit_mask;
} while (!head_word_atomic.compare_exchange_strong(head_expected_word, desired_word));
return true;
};
auto clear_head_word = [&]() {
auto head_word_atomic = std::atomic_ref<uint64_t>(base[base_index]);
head_word_atomic.fetch_and(~head_bit_mask);
};
auto set_center_words = [&]() -> bool {
const size_t end_center_word_index = center_words_base_index + center_words_count;
for (size_t center_index = center_words_base_index; center_index < end_center_word_index; ++center_index) {
auto center_word_atomic = std::atomic_ref<uint64_t>(base[center_index]);
auto center_expected_word = center_word_atomic.load(std::memory_order_relaxed);
// Center words set a full ~0ULL mask.
const uint64_t desired_word {~0ULL};
do {
if (center_expected_word) {
// Another thread raced and allocated.
return false;
}
} while (!center_word_atomic.compare_exchange_strong(center_expected_word, desired_word));
++num_center_set;
}
return true;
};
auto clear_center_words = [&]() {
const size_t end_center_word_index = center_words_base_index + num_center_set;
for (size_t center_index = center_words_base_index; center_index < end_center_word_index; ++center_index) {
auto center_word_atomic = std::atomic_ref<uint64_t>(base[center_index]);
center_word_atomic.store(0);
}
};
auto set_tail_word = [&]() -> bool {
auto tail_word_atomic = std::atomic_ref<uint64_t>(base[tail_word_base_index]);
auto tail_expected_word = tail_word_atomic.load(std::memory_order_relaxed);
uint64_t desired_word {};
do {
if (tail_expected_word & tail_bit_mask) {
// Another thread raced and allocated.
return false;
}
// Set the whole mask in one atomic operation.
desired_word = tail_expected_word | tail_bit_mask;
} while (!tail_word_atomic.compare_exchange_strong(tail_expected_word, desired_word));
return true;
};
set_head = set_head_word();
// Do the center if it exists.
if (set_head && has_center) {
set_center = set_center_words();
}
// Do the tail if it exists.
if (set_head && set_center && has_tail) {
set_tail = set_tail_word();
}
if (set_head && set_center && set_tail) {
return true;
}
// Some stage failed to set, rewind everything.
if (set_head) {
// Clear the head if it was set
clear_head_word();
}
if (has_center && num_center_set) {
// `set_center` might not be set, but it still managed to set some of the words.
clear_center_words();
}
// Tail doesn't need to rewind as it will never have been set if we got here.
return false;
}
#if defined(__ARM_FEATURE_ATOMICS) && __ARM_FEATURE_ATOMICS == 1
// Might violate memory-ordering requirements?
// TODO: Verify and enable or delete depending.
// Provides an 11% (Cortex-X4) to 25% (AmpereOneA) performance improvement.
constexpr static bool use_stclr {};
static inline void stclr(uint64_t value, uint64_t* addr) {
asm volatile("stclrl %[Val], [%[addr]];" ::[Val] "r"(value), [addr] "r"(addr) : "memory");
}
#endif
void free_one(size_t index) {
const size_t word_index = index / WORD_SIZE_BITS;
const size_t word_offset = index % WORD_SIZE_BITS;
last_allocation_track.set_last_allocation(word_index);
const uint64_t bic_bit_mask = 1ULL << word_offset;
#if defined(__ARM_FEATURE_ATOMICS) && __ARM_FEATURE_ATOMICS == 1
if constexpr (use_stclr) {
stclr(bic_bit_mask, &base[word_index]);
return;
}
#endif
auto word_atomic = std::atomic_ref<uint64_t>(base[word_index]);
word_atomic.fetch_and(~bic_bit_mask);
}
void free_inside_word(size_t index, size_t count) {
const size_t word_index = index / WORD_SIZE_BITS;
const size_t word_offset = index % WORD_SIZE_BITS;
auto word_atomic = std::atomic_ref<uint64_t>(base[word_index]);
last_allocation_track.set_last_allocation(word_index);
if (count == WORD_SIZE_BITS) {
word_atomic.store(0);
return;
}
const uint64_t bic_bit_mask = ((1ULL << count) - 1) << word_offset;
#if defined(__ARM_FEATURE_ATOMICS) && __ARM_FEATURE_ATOMICS == 1
if constexpr (use_stclr) {
stclr(bic_bit_mask, &base[word_index]);
return;
}
#endif
word_atomic.fetch_and(~bic_bit_mask);
}
void free_large_amount(size_t index, size_t count) {
// Incoming count can be less than WORD_SIZE_BITS if it is unaligned and crossing multiple words.
// Needs to always handle at minimum a head plus center and/or tail arrangement.
const size_t base_index = index / WORD_SIZE_BITS;
const size_t last_index = FEXCore::AlignUpPowerOf2(index + count, WORD_SIZE_BITS) / WORD_SIZE_BITS;
const size_t num_words_to_scan = last_index - base_index;
LOGMAN_THROW_A_FMT(num_words_to_scan > 1, "Needs to be larger than 1 ({}, {})", index, count);
size_t remaining_bits = count;
const uint64_t head_offset_start = index % WORD_SIZE_BITS;
const uint64_t head_bits = WORD_SIZE_BITS - head_offset_start;
uint64_t head_mask = head_bits == WORD_SIZE_BITS ? ~0ULL : (((1ULL << head_bits) - 1) << head_offset_start);
remaining_bits -= head_bits;
auto head_word_atomic = std::atomic_ref<uint64_t>(base[base_index]);
head_word_atomic.fetch_and(~head_mask);
const size_t remaining_center_words = remaining_bits / WORD_SIZE_BITS;
for (size_t i = 0; i < remaining_center_words; ++i) {
// Handle centers if they exist.
auto center_word_atomic = std::atomic_ref<uint64_t>(base[base_index + i + 1]);
center_word_atomic.store(0);
remaining_bits -= WORD_SIZE_BITS;
}
if (remaining_bits) {
// Handle tail if they exist, must always be less than WORD_SIZE_BITS.
LOGMAN_THROW_A_FMT(remaining_bits < WORD_SIZE_BITS, "Too large");
const uint64_t tail_mask = (1ULL << remaining_bits) - 1;
auto tail_word_atomic = std::atomic_ref<uint64_t>(base[base_index + remaining_center_words + 1]);
tail_word_atomic.fetch_and(~tail_mask);
}
}
};
} // namespace FEXCore::Utils
@@ -0,0 +1,474 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <FEXCore/Utils/TypeDefines.h>
#include "Utils/atomic_bitset.h"
#include <array>
#include <optional>
namespace FEXCore::Utils {
struct bitmap_allocator_settings {
std::array<uint32_t, 3> bucket_sizes;
std::array<uint32_t, 3> bucket_allocation_percentages;
};
// Default allocation settings passed through template for easy tinkering.
constexpr static bitmap_allocator_settings fex_default_allocator_settings = {
.bucket_sizes =
{
16, // 16B - 1KB
512, // 512B - 32KB
2048, // 2KB - 128KB
},
.bucket_allocation_percentages =
{
// 16B bucket special cased to allocate remaining space left over from other buckets.
100,
// 40% in the 512B bucket.
40,
// 10% in the 2KB bucket.
10,
},
};
/**
* A bitmap allocator that is segmented in to three partitioned buckets, with atomic memory allocation.
*
* This class is strongly coupled with FEXCore::Utils::atomic_bitset to have fast and relatively efficient bitmap allocation in a lock-free
* fashion. The bucket granule sizes are roughly calculated to match FEX's needs for typical allocation ranges in a single atomic word. It
* supports allocating larger than an atomic word worth of granules with the expectation that those are relatively uncommon. Once a bucket
* is full, the allocation has a chance to be allocated in to another bucket at a slightly less efficient space usage This is with the
* expectation that we want to keep a buffer around as long as possible, without allocating a new one as buffer migrating is expensive.
*
* TODO: In the future this will support scaling bucket percentages based on which ones filled first. This is currently disabled.
*/
template<bitmap_allocator_settings settings = fex_default_allocator_settings>
class atomic_segmented_bitmap_allocator final {
public:
~atomic_segmented_bitmap_allocator() {
deinit();
}
// Initialize segmented allocator asking for a specific arena size.
// Bitset tracking arena allocations will allocate an independent buffer independent of arena size.
void init(size_t arena_size) {
deinit();
// Align up to page.
arena_size = FEXCore::AlignUpPowerOf2(arena_size, FEXCore::Utils::FEX_PAGE_SIZE);
const auto total_bitset_tracking_memory = calculate_bucket_granules_and_bitset_sizes(arena_size);
// RWX for JIT buffer
arena_base = reinterpret_cast<uint8_t*>(FEXCore::Allocator::VirtualAlloc(arena_size, true));
arena_base_size = arena_size;
// RW for bitset buffer
const auto bitset_size_aligned_to_page = FEXCore::AlignUpPowerOf2(total_bitset_tracking_memory, FEXCore::Utils::FEX_PAGE_SIZE);
bitset_base = reinterpret_cast<uint8_t*>(FEXCore::Allocator::VirtualAlloc(bitset_size_aligned_to_page, false));
bitset_base_size = bitset_size_aligned_to_page;
// Name the buffer. This is only going to be used for the JIT right now.
FEXCore::Allocator::VirtualName("FEXMemJIT", arena_base, arena_size);
FEXCore::Allocator::VirtualName("FEXMem_Misc", bitset_base, bitset_base_size);
// THP enabled for both buffers.
FEXCore::Allocator::VirtualTHPControl(arena_base, arena_size, FEXCore::Allocator::THPControl::Enable);
FEXCore::Allocator::VirtualTHPControl(bitset_base, bitset_base_size, FEXCore::Allocator::THPControl::Enable);
initialize_bitsets_for_buckets();
}
void deinit() {
if (arena_base) {
FEXCore::Allocator::VirtualFree(arena_base, arena_base_size);
}
if (bitset_base) {
FEXCore::Allocator::VirtualFree(bitset_base, bitset_base_size);
}
arena_base = nullptr;
bitset_base = nullptr;
}
// Allocate memory.
// Returns nullptr on failure to allocate.
void* allocate(size_t size) {
const auto allocation_order = find_allocation_order(size);
for (size_t i = 0; i < num_bitset_buckets; ++i) {
const auto bitset_index = get_bitset_index_from_allocation_order(allocation_order, i);
auto& bucket = bitset_buckets[bitset_index];
if (!bucket.bucket_allocation_size) {
continue;
}
const auto bucket_granule_size = bitset_buckets[bitset_index].bucket_granule_size;
const auto aligned_size = FEXCore::AlignUpPowerOf2(size, bucket_granule_size);
const auto bits_to_allocate = FEXCore::DividePow2(aligned_size, bucket_granule_size);
auto allocated_bitset_index = bucket.atomic_bitset.allocate(bits_to_allocate);
if (allocated_bitset_index == bucket.atomic_bitset.invalid()) {
// Mark that the bucket might be full and continue going.
mark_bucket_potentially_full(bucket, bitset_index);
continue;
}
// Bitset range allocated, get the pointer.
return get_ptr_from_bitset(bucket, allocated_bitset_index);
}
return nullptr;
}
// Free memory.
void free(void* ptr, size_t size) {
for (auto& bucket : bitset_buckets) {
if (ptr >= bucket.bucket_allocation_base && ptr < (bucket.bucket_allocation_base + bucket.bucket_allocation_size)) {
// Pointer base must be aligned to granule size, but size doesn't need to be.
// User could have asked for a smaller size and we rounded up to the larger granule.
LOGMAN_THROW_A_FMT(reinterpret_cast<uint64_t>(ptr) % bucket.bucket_granule_size == 0, "Pointer was not aligned to bucket granule "
"size!");
const auto bitset_index = FEXCore::DividePow2(
(reinterpret_cast<uint64_t>(ptr) - reinterpret_cast<uint64_t>(bucket.bucket_allocation_base)), bucket.bucket_granule_size);
const auto bitset_count = FEXCore::DividePow2(FEXCore::AlignUpPowerOf2(size, bucket.bucket_granule_size), bucket.bucket_granule_size);
bucket.atomic_bitset.free(bitset_index, bitset_count);
return;
}
}
LogMan::Msg::AFmt("Attempted to free a pointer that isn't tracked in the bitmap?");
}
// Completely clear the allocator.
// Not thread safe!
void clear() {
// VirtualDontNeed replaces pages with zero page.
FEXCore::Allocator::VirtualDontNeed(arena_base, arena_base_size);
FEXCore::Allocator::VirtualDontNeed(bitset_base, bitset_base_size);
// Reinitialize the bitsets.
initialize_bitsets_for_buckets();
bucket_filled_order = ~0ULL;
}
// Reevaluate bucket allocation capacity percentages.
// Not thread safe!
void reevaluate_bucket_allocations_percentages() {
if (bucket_filled_order.load() == ~0ULL) {
// Buckets never claimed that they could be full.
return;
}
// TODO: Adjust `bitset_allocation_percentages` based on fill order.
// TODO: Reallocate bitsets with `calculate_bucket_granules_and_bitset_sizes`
// Statistically for a Denuvo game, likely bucket[0] will be the first to fill, which will consume additional resources from the larger buckets.
// Statistically for any other game, likely bucket[1] will be the first to fill, which will consume additional resources from bucket[0].
// TODO: Gather growth statistics from a variety of titles to determine trends.
// Reset bucket fill order.
bucket_filled_order = ~0ULL;
}
// Debugger routines.
// Returns a bucket index if the pointer exists in it.
ssize_t find_bucket_index(void* ptr) const {
for (size_t i = 0; i < bitset_buckets.size(); ++i) {
const auto& bucket = bitset_buckets[i];
if (ptr >= bucket.bucket_allocation_base && ptr < (bucket.bucket_allocation_base + bucket.bucket_allocation_size)) {
return i;
}
}
return -1;
}
// Returns the number of buckets.
// Although they may be empty without the ability to allocate.
size_t num_buckets() const {
return bitset_buckets.size();
}
struct bucket_alloc_information {
size_t granule_size {};
size_t allocated {};
size_t free {};
};
// Returns information about a bucket's granule allocations.
std::optional<bucket_alloc_information> get_granule_information(size_t bucket_index) const {
if (bucket_index >= bitset_buckets.size()) {
return std::nullopt;
}
const auto& bucket = bitset_buckets[bucket_index];
const auto count = bucket.atomic_bitset.popcount();
return bucket_alloc_information {
.granule_size = bucket.bucket_granule_size,
.allocated = count,
.free = bucket.atomic_bitset.size_in_bits() - count,
};
}
private:
// Arena base.
uint8_t* arena_base {};
size_t arena_base_size {};
// Bitset tracking.
uint8_t* bitset_base {};
size_t bitset_base_size {};
// Bucket sizes are important for how much can be allocated in a single 64-bit atomic operation.
// Ensure these two buffers are sorted by size.
constexpr static std::array<uint32_t, 3> bucket_granule_sizes = settings.bucket_sizes;
// These must be power of two.
static_assert(std::has_single_bit(bucket_granule_sizes[0]));
static_assert(std::has_single_bit(bucket_granule_sizes[1]));
static_assert(std::has_single_bit(bucket_granule_sizes[2]));
// Ordered by size.
static_assert(bucket_granule_sizes[0] < bucket_granule_sizes[1]);
static_assert(bucket_granule_sizes[1] < bucket_granule_sizes[2]);
constexpr static size_t num_bitset_buckets = bucket_granule_sizes.size();
constexpr static size_t bitset_index_shift = 2;
constexpr static size_t bitset_index_mask = (1U << bitset_index_shift) - 1;
struct bucket_information {
// Memory that the bitset is controlling
uint8_t* bucket_allocation_base {};
size_t bucket_allocation_size {};
// Memory backing the bitset itself.
uint8_t* bitset_memory_base {};
size_t bitset_memory_size {};
// The allocation granule of the bitset.
size_t bucket_granule_size {};
// std::atomic<bool>::fetch_or isn't required to be implemented, so use uint8_t instead.
std::atomic<uint8_t> marked_potentially_full {};
FEXCore::Utils::atomic_bitset<true, false> atomic_bitset {};
};
// Order in which the buckets filled for heuristic scaling of bucket sizes.
std::atomic<uint64_t> bucket_filled_order {~0ULL};
std::array<bucket_information, num_bitset_buckets> bitset_buckets {};
// TODO: Support scaling bitset bucket percentages by order in which they filled.
// Currently this is const and can't change when the buckets run out of space.
constexpr static std::array<uint64_t, num_bitset_buckets> bitset_allocation_percentages {
settings.bucket_allocation_percentages[0],
settings.bucket_allocation_percentages[1],
settings.bucket_allocation_percentages[2],
};
// bucket[0] is special cased to catch everything remaining.
static_assert(settings.bucket_allocation_percentages[0] == 100);
// Calculate bitset sizes based on percentages of the arena size.
// eg - split at 50+40+10:
// - 16MB: 8MB + 6.4MB + 1.6MB
// - 128MB: 64MB + 51.2MB + 12.8MB
size_t calculate_bucket_granules_and_bitset_sizes(size_t arena_size) {
size_t total_bitset_range {};
size_t total_bitset_tracking_memory {};
for (size_t inverse_i = bucket_granule_sizes.size(); inverse_i > 0; --inverse_i) {
const auto i = inverse_i - 1;
const size_t bucket_granule_size = bucket_granule_sizes[i];
const auto allocation_percentage = bitset_allocation_percentages[i];
auto& bucket = bitset_buckets[i];
// Set the bucket granule size.
bucket.bucket_granule_size = bucket_granule_size;
// Calculate the amount of space this bitset tracks.
// If each bucket is tracking data at all it must track at least one 64-bit atomic word of bits.
// 16B bucket: 1KB minimum
// 512B bucket: 32KB minimum
// 2KB bucket: 128KB minimum
// Total: 161KB to reach minimums.
//
// If minimum of each bucket isn't reached, that bucket allocates **ZERO**
size_t bucket_track_size {};
if (allocation_percentage == 100) {
// Use the remaining size to allocate in special case.
bucket_track_size = FEXCore::AlignDown(arena_size - total_bitset_range, bucket_granule_size * WORD_SIZE_BITS);
} else {
bucket_track_size = FEXCore::AlignDown(arena_size * allocation_percentage / 100, bucket_granule_size * WORD_SIZE_BITS);
}
const auto tracked_bits = FEXCore::DividePow2(bucket_track_size, bucket_granule_size);
const auto bitset_memory_size = tracked_bits / 8;
bucket.bucket_allocation_size = bucket_track_size;
bucket.bitset_memory_size = bitset_memory_size;
// Mark the bucket as empty if it has no size.
bucket.marked_potentially_full.store(bucket_track_size == 0, std::memory_order_relaxed);
total_bitset_range += bucket_track_size;
total_bitset_tracking_memory += bitset_memory_size;
LOGMAN_THROW_A_FMT(tracked_bits % WORD_SIZE_BITS == 0, "Wasn't aligned to 64-bits!");
}
LOGMAN_THROW_A_FMT((arena_size - total_bitset_range) == 0, "Still had {} bytes remaining in bitmap allocator init",
(arena_size - total_bitset_range));
return total_bitset_tracking_memory;
}
void initialize_bitsets_for_buckets() {
uint8_t* current_base_arena = arena_base;
uint8_t* current_base_bitset = bitset_base;
for (auto& bucket : bitset_buckets) {
if (!bucket.bucket_allocation_size) {
continue;
}
const auto tracked_bits = FEXCore::DividePow2(bucket.bucket_allocation_size, bucket.bucket_granule_size);
bucket.bucket_allocation_base = current_base_arena;
bucket.atomic_bitset.init(current_base_bitset, tracked_bits);
current_base_arena += bucket.bucket_allocation_size;
current_base_bitset += bucket.bitset_memory_size;
}
current_base_bitset =
reinterpret_cast<uint8_t*>(FEXCore::AlignUpPowerOf2(reinterpret_cast<uint64_t>(current_base_bitset), FEXCore::Utils::FEX_PAGE_SIZE));
LogMan::Throw::AFmt(current_base_arena == (arena_base + arena_base_size), "Failed arena math: 0x{:x} != expected 0x{:x}",
(uint64_t)current_base_arena, (uint64_t)arena_base + arena_base_size);
LogMan::Throw::AFmt(current_base_bitset == (bitset_base + bitset_base_size), "Failed bitset math: 0x{:x} != expected 0x{:x}",
(uint64_t)current_base_bitset, (uint64_t)bitset_base + bitset_base_size);
}
// Marking a bucket as potentially being full.
// Doesn't necessarily mean it is actually full, just allocations have started failing.
// Which this could mean high fragmentation or actually full.
// Regardless mark the buffer based on order of allocations failing.
void mark_bucket_potentially_full(bucket_information& bucket, size_t bitset_index) {
if (bucket.marked_potentially_full.load(std::memory_order_relaxed)) {
return;
}
// fetch_or is faster than CAS here.
bool previous_full = bucket.marked_potentially_full.fetch_or(true);
if (previous_full) {
// Another thread already marking it as full. Minor race.
return;
}
uint64_t expected = bucket_filled_order.load(std::memory_order_relaxed);
uint64_t desired;
do {
desired = (expected << bitset_index_shift) | bitset_index;
} while (!bucket_filled_order.compare_exchange_strong(expected, desired));
}
// Return a pointer to the data backing the bitset based on allocation index.
void* get_ptr_from_bitset(const bucket_information& bucket, size_t allocation_index) const {
return bucket.bucket_allocation_base + bucket.bucket_granule_size * allocation_index;
}
// Returns an ordered list of bitsets for which bitsets we should allocate from.
// Always returns all the bitsets but fitment of the heuristic may change the order.
// LSB is highest priority, MSB is lowest priority.
uint64_t find_allocation_order(size_t size) const {
// Align up by the size of the smallest bucket size
size = FEXCore::AlignUpPowerOf2(size, bitset_buckets[0].bucket_granule_size);
struct tracking {
int32_t index {};
};
// This is effectively setup to be a linear scan std::deque, but without the slow overhead of std::deque.
std::array<tracking, num_bitset_buckets> remaining_bitsets = {{
{.index = 0},
{.index = 1},
{.index = 2},
}};
// Check exact size match.
int32_t exact_match_index = -1;
int32_t smallest_single_atomic_index = -1;
for (auto it = remaining_bitsets.begin(); it != remaining_bitsets.end(); ++it) {
auto& bitset_tracking = *it;
const auto bitset_index = bitset_tracking.index;
const auto bucket_granule_size = bitset_buckets[bitset_index].bucket_granule_size;
if (bucket_granule_size == size) {
exact_match_index = bitset_index;
bitset_tracking.index = -1;
break;
}
}
// Check for smallest match within a 64-bit atomic.
for (auto it = remaining_bitsets.begin(); it != remaining_bitsets.end(); ++it) {
auto& bitset_tracking = *it;
if (bitset_tracking.index == -1) {
continue;
}
const auto bitset_index = bitset_tracking.index;
const auto bucket_granule_size = bitset_buckets[bitset_index].bucket_granule_size;
if (size < bucket_granule_size) {
// If the allocation size is smaller than a single bit, then skip this.
// Ensures we don't burn 2KB buckets with 128B allocations unnecessarily.
continue;
}
const auto bucket_granule_size_per_atomic_word = bucket_granule_size * WORD_SIZE_BITS;
if (size <= bucket_granule_size_per_atomic_word) {
// Fits within a single 64-bit atomic.
smallest_single_atomic_index = bitset_index;
bitset_tracking.index = -1;
break;
}
}
uint64_t allocation_order {};
uint64_t current_shift {};
// Exact match has highest priority.
if (exact_match_index != -1) {
allocation_order = exact_match_index;
current_shift += bitset_index_shift;
}
// Smallest fit within a single atomic word has second priority.
if (smallest_single_atomic_index != -1) {
allocation_order |= (smallest_single_atomic_index << current_shift);
current_shift += bitset_index_shift;
}
// Remaining is sorted by smallest->largest bucket size;
// TODO: Might lead to the bitset allocation heuristic to favor scaling the lower bucket sizes to a larger percentage.
for (auto it : remaining_bitsets) {
if (it.index == -1) {
continue;
}
allocation_order |= (it.index << current_shift);
current_shift += bitset_index_shift;
}
return allocation_order;
}
constexpr static size_t get_bitset_index_from_allocation_order(uint64_t allocation_order, size_t index) {
return (allocation_order >> (index * bitset_index_shift)) & bitset_index_mask;
}
constexpr static size_t WORD_SIZE_BITS = sizeof(uint64_t) * 8;
};
} // namespace FEXCore::Utils
+5
View File
@@ -36,6 +36,7 @@ namespace Handler {
enum ConfigOption {
#define OPT_BASE(type, group, enum, json, default) CONFIG_##enum,
#include <FEXCore/Config/ConfigValues.inl>
CONFIG_MAX,
};
#define ENUMDEFINES
@@ -116,11 +117,13 @@ namespace detail {
FEX_DEFAULT_VISIBILITY void SetDataDirectory(std::string_view Path, bool Global);
FEX_DEFAULT_VISIBILITY void SetConfigDirectory(const std::string_view Path, bool Global);
FEX_DEFAULT_VISIBILITY void SetConfigFileLocation(std::string_view Path, bool Global);
FEX_DEFAULT_VISIBILITY void SetCacheDirectory(const std::string_view Path);
FEX_DEFAULT_VISIBILITY const fextl::string& GetDataDirectory(bool Global = false);
FEX_DEFAULT_VISIBILITY const fextl::string& GetConfigDirectory(bool Global);
FEX_DEFAULT_VISIBILITY const fextl::string& GetConfigFileLocation(bool Global = false);
FEX_DEFAULT_VISIBILITY fextl::string GetApplicationConfig(const std::string_view Program, bool Global);
FEX_DEFAULT_VISIBILITY const fextl::string& GetCacheDirectory();
using LayerValue = std::variant< fextl::string, StringArrayType, uint8_t, int8_t, uint16_t, int16_t, uint32_t, int32_t, uint64_t, int64_t, bool >;
@@ -228,6 +231,8 @@ FEX_DEFAULT_VISIBILITY std::optional<T> GetConv(ConfigOption Option);
FEX_DEFAULT_VISIBILITY std::optional<fextl::string*> Get(ConfigOption Option);
FEX_DEFAULT_VISIBILITY void Set(ConfigOption Option, std::string_view Data);
FEX_DEFAULT_VISIBILITY void Erase(ConfigOption Option);
FEX_DEFAULT_VISIBILITY fextl::string SerializeForCache();
FEX_DEFAULT_VISIBILITY bool CheckConfigMatches(std::string_view Config);
template<typename T>
class FEX_DEFAULT_VISIBILITY Value {
+6 -3
View File
@@ -103,14 +103,17 @@ struct CodeMap {
static constexpr Entry LoadExternalLibrary = {0xffff'ffff'ffff'ffff, 0xffff'ffff};
struct FEX_PACKED SetExecutableFileId {
Entry Marker = {0xffff'ffff'ffff'ffff, 0xffff'fffe};
static constexpr Entry Marker32 = {0xffff'ffff'ffff'ffff, 0xffff'ffee};
static constexpr Entry Marker64 = {0xffff'ffff'ffff'ffff, 0xffff'ffef};
Entry Marker;
CodeMapFileId ExecutableFileId;
};
struct ParsedContents {
fextl::string Filename;
fextl::set<uint64_t> Blocks;
bool IsExecutable = false;
// 32/64 for executables, nullopt for non-executables (libraries)
std::optional<int> ExecutableBitness;
};
// Follows scheme fileid[-nomb]
@@ -152,7 +155,7 @@ public:
void AppendBlock(const FEXCore::ExecutableFileSectionInfo&, uint64_t Entry);
void AppendLibraryLoad(const FEXCore::ExecutableFileInfo&);
void AppendSetMainExecutable(const FEXCore::ExecutableFileInfo&);
void AppendSetMainExecutable(const FEXCore::ExecutableFileInfo&, bool Is64Bit);
// Thread-safely commit any pending data to disk
void Flush(size_t Offset);
+3 -4
View File
@@ -113,15 +113,12 @@ public:
* @brief Create a new thread object that doesn't inherit any state.
* Used to create FEX thread objects in preparation for creating a true OS thread.
*
* @param InitialRIP The starting RIP of this thread
* @param StackPointer The starting RSP of this thread
* @param NewThreadState The thread state to inherit from if not nullptr.
*
* @return A new InternalThreadState object for using with a new guest thread.
*/
FEX_DEFAULT_VISIBILITY virtual FEXCore::Core::InternalThreadState*
CreateThread(uint64_t InitialRIP, uint64_t StackPointer, const FEXCore::Core::CPUState* NewThreadState = nullptr) = 0;
FEX_DEFAULT_VISIBILITY virtual FEXCore::Core::InternalThreadState* CreateThread(const FEXCore::Core::CPUState* NewThreadState = nullptr) = 0;
FEX_DEFAULT_VISIBILITY virtual void DestroyThread(FEXCore::Core::InternalThreadState* Thread) = 0;
#ifndef _WIN32
@@ -136,6 +133,8 @@ public:
FEX_DEFAULT_VISIBILITY virtual FEXCore::CPUID::XCRResults RunXCRFunction(uint32_t Function) = 0;
FEX_DEFAULT_VISIBILITY virtual FEXCore::CPUID::FunctionResults RunCPUIDFunctionName(uint32_t Function, uint32_t Leaf, uint32_t CPU) = 0;
virtual void InitDiskCache() = 0;
virtual AbstractCodeCache& GetCodeCache() = 0;
virtual void SetCodeMapWriter(fextl::unique_ptr<CodeMapWriter>) = 0;
virtual void FlushAndCloseCodeMap() = 0;
+231
View File
@@ -0,0 +1,231 @@
// SPDX-License-Identifier: MIT
#pragma once
#include "FEXCore/Core/CodeCache.h"
#include "FEXCore/Core/Context.h"
#include "Interface/Core/JIT/Relocations.h"
#include "Interface/Core/Frontend.h"
#include "Interface/Core/CPUBackend.h"
#include "FEXCore/Config/Config.h"
#include "FEXCore/Utils/File.h"
#include "FEXCore/Utils/WorkQueueThread.h"
#include "FEXCore/fextl/memory.h"
#include <FEXCore/fextl/string.h>
#include <FEXCore/fextl/unordered_set.h>
#include <FEXCore/fextl/robin_map.h>
#include <FEXCore/fextl/map.h>
#include <FEXCore/fextl/vector.h>
#include <stdint.h>
#include <mutex>
#include <optional>
#include <span>
#include <xxhash.h>
namespace FEXCore {
namespace Context {
class ContextImpl;
}
namespace DiskCache {
namespace MesaFOZ {
#define FOSSILIZE_BLOB_HASH_LENGTH 40 /* SHA1 hexadecimal string length */
struct __attribute__((packed)) foz_payload_key {
uint8_t bytes[FOSSILIZE_BLOB_HASH_LENGTH];
};
struct __attribute__((packed)) foz_payload_header {
uint32_t payload_size;
uint32_t format;
uint32_t crc;
uint32_t uncompressed_size;
};
} // namespace MesaFOZ
class IndexedDB;
struct IndexEntry {
IndexedDB* DB;
uint64_t Offset;
uint32_t Size;
uint32_t GuestSize;
XXH128_hash_t GuestHash;
fextl::vector<uint32_t> GuestExtents;
};
struct IndexCacheHead {
struct IndexEntry MainEntry;
uint64_t MainEntryFootprint;
fextl::unique_ptr<fextl::multimap<uint64_t, IndexEntry>> MoreEntries; // sorted by guest footprint
};
struct __attribute__((packed)) BlobFixedHeader {
uint32_t GuestSize;
uint32_t HostSize;
uint32_t EntryPointCount;
uint32_t SmallRelocCount;
uint32_t ThunkRelocCount;
XXH128_hash_t GuestHash;
};
// packed struct for types 0, 2 and 3. type 1 is bigger and separate below
struct __attribute__((packed)) BlobSmallRelocation {
uint32_t Offset;
uint8_t Type;
union {
struct __attribute__((packed)) {
uint32_t Symbol;
} Named;
struct __attribute__((packed)) {
uint64_t GuestRIP;
} RIPLiteral;
struct __attribute__((packed)) {
uint8_t RegisterIndex;
uint64_t GuestRIP;
} RIPMove;
struct __attribute__((packed)) {
uint8_t RegisterIndex;
uint8_t ValueSize;
uint32_t SiteOffset;
} PatchableData;
};
};
// type 1, implicit
struct __attribute__((packed)) BlobThunkRelocation {
uint32_t Offset;
uint8_t RegisterIndex;
uint8_t SymbolHash[32]; // sha256sum in the real RelocNamedThunkMove
};
struct CodeHitData {
fextl::vector<uint8_t> Blob;
std::span<uint8_t> HostCode;
std::span<const uint64_t> GuestPages;
std::span<uint64_t> EntryPointRIPs;
std::span<const uint32_t> EntryPointHostOffsets;
// the spans above point to memory owned by the Blob vec, so it's important this can't be copied
CodeHitData() = default;
CodeHitData(CodeHitData&&) = default;
CodeHitData& operator=(CodeHitData&&) = default;
CodeHitData(const CodeHitData&) = delete;
CodeHitData& operator=(const CodeHitData&) = delete;
};
using Index = fextl::robin_map<uint64_t, IndexCacheHead>;
class FOZFile {
public:
bool Open(const fextl::string& CacheFileName, bool ReadOnly);
bool Lock(uint32_t TimeoutMS) {
if (!FD) {
return false;
}
return FD->Lock(TimeoutMS);
}
bool Unlock() {
if (!FD) {
return false;
}
return FD->Unlock();
}
File::File::FileHandleType GetHandle() {
return FD ? FD->GetHandle() : (File::File::FileHandleType)-1;
}
ssize_t Size();
bool ReadAll(fextl::vector<uint8_t>& Out); // from first blob
bool ReadBlob(uint64_t Offset, std::span<uint8_t> OutBlob);
bool WriteBlob(const MesaFOZ::foz_payload_key& Key, std::span<const std::span<const uint8_t>> BlobChunks, uint64_t& OutBlobOffset);
private:
static constexpr uint32_t OPEN_LOCK_TIMEOUT_MS = 100;
fextl::string FileName;
fextl::unique_ptr<File::File> FD;
bool ReadOnly = false;
};
class IndexedDB {
public:
bool Open(const fextl::string& CacheDBName, bool ReadOnly);
void PopulateIndex(Index& CacheIndex, bool& FoundMetadata);
bool ReadCacheBlob(uint64_t Offset, std::span<uint8_t> OutBlob);
bool StoreCacheBlob(const MesaFOZ::foz_payload_key& UniqueKey, uint64_t LookupKey, std::span<const uint8_t> Blob, Index& CacheIndex,
std::mutex& IndexMutex, std::span<const uint8_t> IndexBlob);
private:
// stores run on the Writer, so returning quick isn't as important
static constexpr uint32_t STORE_LOCK_TIMEOUT_MS = 1000;
static constexpr uint64_t BIG_MAPPING_SIZE = 1ULL << 33;
static constexpr uint32_t LOOKUP_KEY_MAX_BUCKET_DEPTH = 20;
FOZFile CacheFOZ;
uint8_t* CacheFileMapping = nullptr;
std::atomic<uint64_t> CacheFileSize;
FOZFile IndexFOZ;
bool ReadOnly = false;
};
class DiskCache {
public:
void Init(FEXCore::Context::ContextImpl* CTX);
std::optional<CodeHitData> Lookup(Core::InternalThreadState* Thread, std::optional<ExecutableFileSectionInfo> Region, uint64_t GuestRIP,
std::optional<uint64_t>& GuestCodeKey);
void Validate(uint64_t GuestCodeKey, const CodeHitData& Hit, const CPU::CPUBackend::CompiledCode& CompiledCode,
std::optional<ExecutableFileSectionInfo> Region);
bool Store(Core::InternalThreadState* Thread, std::optional<ExecutableFileSectionInfo> Region, uint64_t GuestRIP, uint64_t GuestCodeKey,
std::span<const uint8_t> GuestCode, const CPU::CPUBackend::CompiledCode& CompiledCode,
std::span<const FEXCore::CPU::Relocation> Relocations, const Frontend::Decoder::DecodedBlockInformation* DecodedBlockInfo);
bool IsWritingDiskCache() const {
return WritingDiskCache;
}
bool IsReadingDiskCache() const {
return ReadingDiskCache;
}
bool IsValidating() const {
return Validation;
}
private:
bool OpenCacheDB(const fextl::string& CacheDBName, bool ReadOnly);
uint64_t MakeLookupKey(Core::InternalThreadState* Thread, const uint64_t ModuleOffset, bool Writable, bool MonoBackpatcher);
bool ReadingDiskCache {};
bool WritingDiskCache {};
FEXCore::Context::ContextImpl* CTX;
XXH128_hash_t BucketHash;
fextl::vector<fextl::unique_ptr<IndexedDB>> ROCacheDBs;
fextl::unique_ptr<IndexedDB> RWCacheDB;
Index Index;
std::mutex IndexLock;
bool FoundMetadata = false;
struct CacheStoreWorkItem;
// the Writer holds references to all this stuff above and needs to be last
fextl::unique_ptr<WorkQueueThread> Writer;
FEX_CONFIG_OPT(EnableDiskCache, DISKCACHE);
FEX_CONFIG_OPT(Validation, DISKCACHEVALIDATION);
FEX_CONFIG_OPT(MapDiskCacheFiles, DISKCACHEFILEMAPPING);
FEX_CONFIG_OPT(RelocationFilter, DISKCACHERELOCATIONFILTER);
FEX_CONFIG_OPT(AnonCaching, DISKCACHEANONCACHING);
FEX_CONFIG_OPT(BasePathOverride, DISKCACHEPATH);
FEX_CONFIG_OPT(RODBNames, DISKCACHERODBNAMES);
};
static constexpr uint16_t AnonPrefixGuestBytes = 64;
// TODO: This header is in global installed header path, but uses internal headers.
// Migrate this once that is fixed.
static constexpr uint16_t FormatVersion = 20;
FEX_DEFAULT_VISIBILITY uint16_t GetFormatVersion();
} // namespace DiskCache
} // namespace FEXCore
@@ -0,0 +1,11 @@
// SPDX-License-Identifier: MIT
#pragma once
#include "FEXCore/Utils/CompilerDefs.h"
#include "FEXCore/Utils/File.h"
namespace FEXCore::DiskCache {
using FileMapperFunc = void* (*)(FEXCore::File::File::FileHandleType Handle, uint64_t MapSize);
FEX_DEFAULT_VISIBILITY void SetFileMapper(FileMapperFunc Func);
} // namespace FEXCore::DiskCache
+56 -44
View File
@@ -13,59 +13,71 @@ namespace FEXCore {
* Not the x86->IR process
*/
struct HostFeatures {
// Changes code generation slightly.
enum class HostTypeEnum : uint32_t {
Unknown,
Linux,
Wow64,
Arm64ec,
};
// Whether or not the host supports any kind of SVE implementation.
[[nodiscard]]
bool SupportsSVE() const {
return SupportsSVE128 || SupportsSVE256;
}
uint32_t DCacheLineSize {};
uint32_t ICacheLineSize {};
bool SupportsCacheMaintenanceOps {};
bool SupportsAES {};
bool SupportsCRC {};
bool SupportsCLZERO {};
bool SupportsAtomics {};
bool SupportsRCPC {};
bool SupportsTSOImm9 {};
bool SupportsRAND {};
bool SupportsAVX {};
bool SupportsSVE128 {};
bool SupportsSVE256 {};
bool SupportsSHA {};
bool SupportsPMULL_128Bit {};
bool SupportsCSSC {};
bool SupportsFCMA {};
bool SupportsFlagM {};
bool SupportsFlagM2 {};
bool SupportsRPRES {};
bool SupportsPreserveAllABI {};
bool SupportsAES256 {};
bool SupportsSVEBitPerm {};
bool SupportsCPUIndexInTPIDRRO {};
bool SupportsFRINTTS {};
bool SupportsECV {};
bool SupportsWFXT {};
bool Supports3DNow {};
bool SupportsSSE4a {};
bool SupportsMOPS {};
bool PreferZVAForVZero {};
[[nodiscard]]
uint32_t DCacheSize() const {
return 4 << DCacheLineLog2;
}
// Float exception behaviour
bool SupportsAFP {};
bool SupportsFloatExceptions {};
// Changes code generation slightly.
enum class HostTypeEnum {
Unknown,
Linux,
Wow64,
Arm64ec,
};
HostTypeEnum HostType {};
[[nodiscard]]
uint64_t HashForCaching() const {
// As long as the number of options is 64-bit or below, we can just return it.
// Skip CPUMIDRs as it doesn't affect codegen.
static_assert(offsetof(HostFeatures, CPUMIDRs) == 8);
uint64_t Result {};
memcpy(&Result, this, sizeof(Result));
return Result;
}
uint32_t DCacheLineLog2 : 4 {};
uint32_t SupportsCacheMaintenanceOps : 1 {};
uint32_t SupportsAES : 1 {};
uint32_t SupportsCRC : 1 {};
uint32_t SupportsCLZERO : 1 {};
uint32_t SupportsAtomics : 1 {};
uint32_t SupportsRCPC : 1 {};
uint32_t SupportsTSOImm9 : 1 {};
uint32_t SupportsRAND : 1 {};
uint32_t SupportsAVX : 1 {};
uint32_t SupportsSVE128 : 1 {};
uint32_t SupportsSVE256 : 1 {};
uint32_t SupportsSHA : 1 {};
uint32_t SupportsPMULL_128Bit : 1 {};
uint32_t SupportsCSSC : 1 {};
uint32_t SupportsFCMA : 1 {};
uint32_t SupportsFlagM : 1 {};
uint32_t SupportsFlagM2 : 1 {};
uint32_t SupportsRPRES : 1 {};
uint32_t SupportsPreserveAllABI : 1 {};
uint32_t SupportsAES256 : 1 {};
uint32_t SupportsSVEBitPerm : 1 {};
uint32_t SupportsCPUIndexInTPIDRRO : 1 {};
uint32_t SupportsFRINTTS : 1 {};
uint32_t SupportsECV : 1 {};
uint32_t SupportsWFXT : 1 {};
uint32_t Supports3DNow : 1 {};
uint32_t SupportsSSE4a : 1 {};
uint32_t SupportsMOPS : 1 {};
uint32_t PreferZVAForVZero : 1 {};
uint32_t SupportsAFP : 1 {};
uint32_t SupportsFloatExceptions : 1 {};
// Flag if this is InstCountCI
bool IsInstCountCI {};
uint32_t IsInstCountCI : 1 {};
HostTypeEnum HostType : 2 {};
uint32_t pad : 26 {};
// MIDR information
// Also used for determining number of CPU cores for CPUID
+1 -30
View File
@@ -17,29 +17,6 @@ struct CpuStateFrame;
} // namespace FEXCore::Core
namespace FEXCore::HLE {
struct SyscallArguments {
static constexpr std::size_t MAX_ARGS = 7;
uint64_t Argument[MAX_ARGS];
};
struct SyscallABI {
// Expectation is that the backend will be aware of how to modify the arguments based on numbering
// Only GPRs expected
uint8_t NumArgs;
// If the syscall has a return then it should be stored in the ABI specific syscall register
// Linux = RAX
bool HasReturn;
int32_t HostSyscallNumber;
};
enum class SyscallOSABI {
OS_UNKNOWN,
OS_LINUX64,
OS_LINUX32,
OS_GENERIC, // No JIT-side argument handling, spill/fill all regs.
};
struct ExecutableRangeInfo {
uint64_t Base;
uint64_t Size;
@@ -53,11 +30,8 @@ class SyscallHandler {
public:
virtual ~SyscallHandler() = default;
virtual uint64_t HandleSyscall(FEXCore::Core::CpuStateFrame* Frame, FEXCore::HLE::SyscallArguments* Args) = 0;
virtual void HandleSyscall(FEXCore::Core::CpuStateFrame* Frame) = 0;
SyscallOSABI GetOSABI() const {
return OSABI;
}
virtual void MarkGuestExecutableRange(FEXCore::Core::InternalThreadState* Thread, uint64_t Start, uint64_t Length) {}
virtual void InvalidateGuestCodeRange(FEXCore::Core::InternalThreadState* Thread, uint64_t Start, uint64_t Length) {}
virtual void MarkOvercommitRange(uint64_t Start, uint64_t Length) {}
@@ -72,8 +46,5 @@ public:
}
virtual void SleepThread(FEXCore::Context::Context* CTX, FEXCore::Core::CpuStateFrame* Frame) {}
protected:
SyscallOSABI OSABI;
};
} // namespace FEXCore::HLE
+1 -1
View File
@@ -12,7 +12,7 @@ struct InternalThreadState;
}
namespace FEXCore::Allocator {
FEX_DEFAULT_VISIBILITY size_t DetermineVASize();
FEX_DEFAULT_VISIBILITY size_t GetHostVABits();
#ifdef GLIBC_ALLOCATOR_FAULT
// Glibc hooks should only fault once we are in main.
@@ -97,7 +97,8 @@ inline bool VirtualProtect(void* Ptr, size_t Size, ProtectOptions options) {
LOGMAN_MSG_A_FMT("Unknown VirtualProtect options combination");
}
return ::VirtualProtect(Ptr, Size, prot, nullptr) == 0;
DWORD OldProt {};
return ::VirtualProtect(Ptr, Size, prot, &OldProt) != 0;
}
FEX_DEFAULT_VISIBILITY extern VirtualNamePtr VirtualName;
@@ -165,6 +166,8 @@ FEX_DEFAULT_VISIBILITY extern void InitializeThread();
#ifndef _WIN32
void SetupAllocatorHooks(void* (*)(void* addr, size_t length, int prot, int flags, int fd, off_t offset), int (*)(void* addr, size_t length));
#else
void SetupAllocatorHooks(void (*)(const char* name, const void* address, size_t size));
#endif
struct FEXAllocOperators {
+189 -9
View File
@@ -3,10 +3,17 @@
#include <FEXCore/fextl/allocator.h>
#include <FEXCore/fextl/string.h>
#include <FEXCore/Utils/EnumOperators.h>
#include "FEXCore/Utils/LogManager.h"
#include <chrono>
#include <thread>
#include <utility>
#ifndef _WIN32
#include <fcntl.h>
#include <unistd.h>
#include <sys/file.h>
#include <sys/stat.h>
#else
#define WIN32_LEAN_AND_MEAN
#include <windows.h>
@@ -39,7 +46,8 @@ public:
File() = default;
File(const char* Filepath, FileModes Modes) {
File(const char* Filepath, FileModes Modes, bool Seekable = true)
: Seekable {Seekable} {
#ifndef _WIN32
auto Disp = TranslateModes(Modes);
Handle = open(Filepath, Disp, DEFAULT_USER_PERMS);
@@ -58,6 +66,40 @@ public:
}
IsValidHandle = Handle != INVALID_HANDLE_VALUE;
#endif
ShouldClose = IsValidHandle;
}
File(File&& Other) noexcept
: ShouldClose(std::exchange(Other.ShouldClose, false))
, IsValidHandle(std::exchange(Other.IsValidHandle, false))
#ifdef _WIN32
, Handle(std::exchange(Other.Handle, INVALID_HANDLE_VALUE))
#else
, Handle(std::exchange(Other.Handle, -1))
#endif
, Seekable(std::exchange(Other.Seekable, false))
, Locked(std::exchange(Other.Locked, false)) {
}
File& operator=(File&& Other) noexcept {
if (this == &Other) {
return *this;
}
std::swap(ShouldClose, Other.ShouldClose);
std::swap(IsValidHandle, Other.IsValidHandle);
std::swap(Handle, Other.Handle);
std::swap(Seekable, Other.Seekable);
std::swap(Locked, Other.Locked);
return *this;
}
File(const File&) = delete;
File& operator=(const File&) = delete;
FileHandleType GetHandle() {
return Handle;
}
/**
@@ -69,6 +111,10 @@ public:
* @return The number of bytes actually written or -1 on error.
*/
ssize_t Write(const void* Buffer, size_t Bytes) {
if (!Seekable) {
LOGMAN_THROW_A_FMT(false, "Can't use non-positioned ops on a non-seekable file!");
return -1;
}
#ifndef _WIN32
return write(Handle, Buffer, Bytes);
#else
@@ -95,6 +141,10 @@ public:
* @return The number of bytes read or -1 on error.
*/
ssize_t Read(void* Buffer, size_t Bytes) {
if (!Seekable) {
LOGMAN_THROW_A_FMT(false, "Can't use non-positioned ops on a non-seekable file!");
return -1;
}
#ifndef _WIN32
return read(Handle, Buffer, Bytes);
#else
@@ -108,6 +158,80 @@ public:
#endif
}
ssize_t PRead(void* Buffer, size_t Bytes, uint64_t Offset) {
if (Seekable) {
LOGMAN_THROW_A_FMT(false, "Can't use positioned ops on a seekable file!");
return -1;
}
#ifndef _WIN32
return pread(Handle, Buffer, Bytes, Offset);
#else
DWORD BytesRead {};
OVERLAPPED Overlapped {};
Overlapped.Offset = static_cast<DWORD>(Offset);
Overlapped.OffsetHigh = static_cast<DWORD>(Offset >> 32);
auto Result = ReadFile(Handle, Buffer, Bytes, &BytesRead, &Overlapped);
if (Result) {
return BytesRead;
}
// Some error, match Linux side.
return -1;
#endif
}
ssize_t PWrite(const void* Buffer, size_t Bytes, uint64_t Offset) {
if (Seekable) {
LOGMAN_THROW_A_FMT(false, "Can't use positioned ops on a seekable file!");
return -1;
}
#ifndef _WIN32
return pwrite(Handle, Buffer, Bytes, Offset);
#else
DWORD BytesWritten {};
OVERLAPPED Overlapped {};
Overlapped.Offset = static_cast<DWORD>(Offset);
Overlapped.OffsetHigh = static_cast<DWORD>(Offset >> 32);
auto Result = WriteFile(Handle, Buffer, Bytes, &BytesWritten, &Overlapped);
if (Result) {
return BytesWritten;
}
// Some error, match Linux side.
return -1;
#endif
}
bool Lock(uint32_t TimeoutMS) {
for (uint32_t i = 0;; ++i) {
if (TryLock()) {
return true;
}
if (i >= TimeoutMS) {
return false;
}
std::this_thread::sleep_for(std::chrono::milliseconds(1));
}
}
bool Unlock() {
if (!Locked) {
return false; // we could return true here :thonk:
}
#ifndef _WIN32
if (flock(Handle, LOCK_UN) == -1) {
return false;
}
#else
OVERLAPPED Overlapped {};
Overlapped.Offset = static_cast<DWORD>(LOCK_SENTINEL_OFFSET);
Overlapped.OffsetHigh = static_cast<DWORD>(LOCK_SENTINEL_OFFSET >> 32);
if (!UnlockFileEx(Handle, 0, 1, 0, &Overlapped)) {
return false;
}
#endif
Locked = false;
return true;
}
~File() {
if (!IsValidHandle) {
return;
@@ -164,6 +288,22 @@ public:
#endif
}
ssize_t Size() {
#ifndef _WIN32
struct stat st;
if (fstat(Handle, &st) != 0) {
return -1;
}
return st.st_size;
#else
LARGE_INTEGER FileSize;
if (!GetFileSizeEx(Handle, &FileSize)) {
return -1;
}
return FileSize.QuadPart;
#endif
}
/**
* @brief Seek the file pointer location.
*
@@ -173,6 +313,10 @@ public:
* @return The current file pointer location or -1.
*/
ssize_t Seek(ssize_t Distance, SeekOp Op) {
if (!Seekable) {
LOGMAN_THROW_A_FMT(false, "Can't use non-positioned ops on a non-seekable file!");
return -1;
}
#ifndef _WIN32
return lseek(Handle, Distance, TranslateSeek(Op));
#else
@@ -189,25 +333,55 @@ public:
protected:
File(FileHandleType Handle, bool ShouldClose)
File(FileHandleType Handle, bool ShouldClose, bool Seekable = true)
: ShouldClose {ShouldClose}
, IsValidHandle {true}
, Handle {Handle} {}
, Handle {Handle}
, Seekable {Seekable} {}
private:
bool TryLock() {
if (Locked) {
return true;
}
#ifndef _WIN32
if (flock(Handle, LOCK_EX | LOCK_NB) == -1) {
return false;
}
#else
// mimic posix advisory-only lock by locking some unattainably-high bit
OVERLAPPED Overlapped {};
Overlapped.Offset = static_cast<DWORD>(LOCK_SENTINEL_OFFSET);
Overlapped.OffsetHigh = static_cast<DWORD>(LOCK_SENTINEL_OFFSET >> 32);
if (!LockFileEx(Handle, LOCKFILE_EXCLUSIVE_LOCK | LOCKFILE_FAIL_IMMEDIATELY, 0, 1, 0, &Overlapped)) {
return false;
}
#endif
Locked = true;
return true;
}
bool ShouldClose {};
bool IsValidHandle {};
FileHandleType Handle {};
bool Seekable = true;
bool Locked = false;
#ifndef _WIN32
static constexpr int DEFAULT_USER_PERMS = S_IRWXU | S_IRWXG | S_IRWXO;
static uint32_t TranslateModes(FileModes Modes) {
const auto IsRead = (Modes & FileModes::READ) == FileModes::READ;
const auto IsWrite = (Modes & FileModes::WRITE) == FileModes::WRITE;
uint32_t Mode {};
if ((Modes & FileModes::READ) == FileModes::READ) {
Mode |= O_RDONLY;
}
if ((Modes & FileModes::WRITE) == FileModes::WRITE) {
Mode |= O_WRONLY;
if (IsRead && IsWrite) {
Mode |= O_RDWR;
} else {
if (IsRead) {
Mode |= O_RDONLY;
} else if (IsWrite) {
Mode |= O_WRONLY;
}
}
if ((Modes & FileModes::CREATE) == FileModes::CREATE) {
Mode |= O_CREAT;
@@ -232,6 +406,8 @@ private:
}
#else
static constexpr int DEFAULT_SHARE_MODE = FILE_SHARE_READ | FILE_SHARE_WRITE | FILE_SHARE_DELETE;
static constexpr uint64_t LOCK_SENTINEL_OFFSET = 1ULL << 62;
struct Disposition {
uint32_t CreationFlag;
uint32_t Access;
@@ -246,7 +422,11 @@ private:
Disp.Access |= GENERIC_WRITE;
}
if ((Modes & FileModes::CREATE) == FileModes::CREATE) {
Disp.CreationFlag = CREATE_ALWAYS;
if ((Modes & FileModes::TRUNCATE) == FileModes::TRUNCATE) {
Disp.CreationFlag = CREATE_ALWAYS;
} else {
Disp.CreationFlag = OPEN_ALWAYS;
}
} else {
Disp.CreationFlag = OPEN_ALWAYS;
}
+12
View File
@@ -18,6 +18,18 @@ constexpr uint64_t AlignDown(uint64_t value, uint64_t size) {
return value - value % size;
}
[[nodiscard]]
constexpr uint64_t AlignUpPowerOf2(uint64_t value, uint64_t size) {
LOGMAN_THROW_A_FMT(std::popcount(size) == 1, "Alignment needs to be power of 2");
return (value + size - 1) & ~(size - 1);
}
[[nodiscard]]
constexpr uint64_t AlignDownPowerOf2(uint64_t value, uint64_t size) {
LOGMAN_THROW_A_FMT(std::popcount(size) == 1, "Alignment needs to be power of 2");
return value & ~(size - 1);
}
// Returns the ilog2 of a power-of-2 integer.
// Asserts in the case that the passed in integer is not a power-of-2.
template<typename T>
+5 -2
View File
@@ -2,8 +2,7 @@
#pragma once
#ifndef _WIN32
#include <sys/mman.h>
#include <sys/user.h>
#include <sys/prctl.h>
#ifndef PR_SET_VMA
@@ -50,4 +49,8 @@
#define PR_SHADOW_STACK_ENABLE (1ULL << 0)
#endif
#ifndef PR_GET_MDWE
#define PR_GET_MDWE 66
#endif
#endif // ifndef _WIN32
+5
View File
@@ -70,6 +70,11 @@ struct ThreadStats {
uint64_t AccumulatedCacheWriteLockTime;
uint64_t AccumulatedJITCount;
uint64_t AccumulatedDiskCacheHitCount;
uint64_t AccumulatedDiskCacheMissCount;
uint64_t AccumulatedDiskCacheLookupTime;
uint64_t Padding;
};
// Ensure 16-byte alignment to take advantage of ARM single-copy atomicity.
Loaded 100 of 387 files, more files were not shown because too many files have changed in this diff. Show more