Compare commits

...
226 Commits
Author SHA1 Message Date
Ryan Houdek a04b0241c2 Docs: Update for release FEX-2605 2026-05-08 19:28:30 -07:00
Ryan Houdek 670fd19d33 Merge pull request #5486 from Sonicadvance1/151
Allocator: Mark large unmapped regions as DONTDUMP
2026-05-08 19:28:13 -07:00
Ryan Houdek a66544f3f4 Allocator: Mark large unmapped regions as DONTDUMP
coredump applications aren't smart enough to only dump resident pages,
so explicitly mark our 128TB and other mapped VA ranges as DONTDUMP.

This will speed up coredumps.
2026-05-08 16:27:04 -07:00
Ryan Houdek 1bfb3aefcc Merge pull request #5485 from Sonicadvance1/150
Windows: Setup `tu_override_uncached_as_cache_coherent` inside of dlls
2026-05-08 15:58:43 -07:00
Ryan Houdek b7bfbc3fcd Windows: Setup tu_override_uncached_as_cache_coherent inside of dlls
To not have this environment variable accidently be enabled on arm64
native Wine games, we need to set it from inside of FEX.

Requires the FEX dlls to set them directly rather than launch scripts.
2026-05-08 13:39:51 -07:00
Ryan Houdek e517f3259c Merge pull request #5484 from neobrain/fix_code_cache_no_guest_wrappers
CodeCache: Fix crash when guest library wrappers aren't installed
2026-05-07 10:59:32 -07:00
Tony Wasserka 8afda92a64 CodeCache: Fix crash when guest library wrappers aren't installed 2026-05-07 18:56:35 +02:00
Ryan Houdek ed216c8d4d Merge pull request #5449 from neobrain/opt_code_cache_writing
CodeCache: Slightly optimize cache file writing
2026-05-06 18:06:51 -07:00
Ryan Houdek 7506cb4ea1 Merge pull request #5483 from neobrain/fix_guest_wrapper_code_cache
CodeCache: Delay cache loading for guest library wrappers until after LoadLib
2026-05-06 18:04:55 -07:00
Tony Wasserka 60bc5944db CodeCache: Slightly optimize cache file writing
ftruncate only requires one call (and one extra seek) instead up to 64 manual
zero writes.
2026-05-06 17:08:15 +02:00
Tony Wasserka 8f0572283a Windows/CRT: Implement ftruncate and _chsize 2026-05-06 17:08:04 +02:00
Tony Wasserka b13b46eefe CodeCache: Delay cache loading for guest library wrappers until after LoadLib
These libraries need to be initialized before relocating their caches,
since the guest function hashes won't be registered before.
2026-05-05 16:29:39 +02:00
LC 4db2a98d7f Merge pull request #5481 from Sonicadvance1/148
win32: Query DCZID_EL0 so clzero works
2026-05-05 08:42:31 -04:00
LC 05ebb07753 Merge pull request #5482 from Sonicadvance1/149
OpcodeDispatcher: Optimize MMX pshufw
2026-05-05 08:41:22 -04:00
Ryan Houdek 694e68b838 Merge pull request #5468 from peppergrayxyz/proc_self_stat
read /proc/self/stat using %lu
2026-05-04 20:21:16 -07:00
Pepper Gray 78320e1433 read /proc/self/stat using %lu
building on clang/musl causes warnings:

```
FEX/Source/Tools/FEXInterpreter/ELFCodeLoader.h:782:29: warning: format specifies type 'unsigned long long *' but the argument has type 'uint64_t *' (aka 'unsigned long *') [-Wformat]
  776 |                             "%llu %llu %llu %*u %*u "   // 26 to 30
      |                              ~~~~
      |                              %lu
  777 |                             "%*u %*u %*u %*u %*u "      // 31 to 35
  778 |                             "%*u %*u %*d %*d %*u "      // 36 to 40
  779 |                             "%*u %*u %*u %*d %llu "     // 40 to 45
  780 |                             "%llu %llu %llu %llu %llu " // 46 to 50
  781 |                             "%llu",                     // 51
  782 |                             &map.start_code, &map.end_code, &map.start_stack, &map.start_data, &map.end_data, &map.start_brk,
      |                             ^~~~~~~~~~~~~~~
```

according to the [man page](https://man7.org/linux/man-pages/man5/proc_pid_stat.5.html)
`/proc/self/stat` uses `%lu`:

read the values as unsigned long (%lu) and then write them to
prctl_mm_map (platform specific format).

fixes: #5467
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-05 04:24:31 +02:00
Ryan Houdek 082e7b2695 InstcountCI: Update 2026-05-04 17:52:54 -07:00
Ryan Houdek cdbffb80c7 unittests: Adds a full coverage pshufw test 2026-05-04 17:52:53 -07:00
Ryan Houdek cb6c8cce55 OpcodeDispatcher: Optimize MMX pshufw
Found through writing a shuffle solver rather than an LLM.

Fixes #3785
2026-05-04 17:52:53 -07:00
Ryan Houdek 47e173e549 Merge pull request #5472 from peppergrayxyz/format
use portable format specifiers
2026-05-04 14:18:28 -07:00
Ryan Houdek 162bd4be97 win32: Query DCZID_EL0 so clzero works
EL0 registers are readable without going through the registry, but we
were failing to populate this register, which was causing clzero to not
be supported.
2026-05-04 12:42:18 -07:00
Ryan Houdek f0764aeafe Merge pull request #5480 from neobrain/refactor_musl_sigmask
SignalDelegator: Simplify support for musl's sigset_t
2026-05-04 10:39:42 -07:00
Ryan Houdek 1efed71696 Merge pull request #5479 from neobrain/refactor_drop_compile_service
FEXCore: Drop unused CompileService
2026-05-04 10:36:44 -07:00
Tony Wasserka 06d77c1c19 SignalDelegator: Simplify support for musl's sigset_t 2026-05-04 16:28:47 +02:00
Tony Wasserka 942d0c631a FEXCore: Drop unused CompileService 2026-05-04 15:52:34 +02:00
Pepper Gray abae5dd93b use portable format specifiers
building with clang/musl causes these warnings:

```
FEX/Source/Tools/FEXServer/ProcessPipe.cpp:100:96: warning: format specifies type 'ssize_t' (aka 'long') but the argument has type 'rlim_t' (aka 'unsigned long long') [-Wformat]
```

- cast platform specific MaxFDs members to uintmax_t and print as PRIuMAX
- use %zu for GetNumFilesOpen (size_t)

fix: #5471
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-04 11:20:55 +02:00
Ryan Houdek 93015a0266 Merge pull request #5458 from peppergrayxyz/largefile64
make 64bit symbols visible to enhance portability (musl)
2026-05-03 23:36:59 -07:00
Ryan Houdek d238db69d3 Merge pull request #5477 from peppergrayxyz/unistd
include missing header unistd.h
2026-05-03 15:18:41 -07:00
Ryan Houdek e0ead236b6 Merge pull request #5473 from peppergrayxyz/ObjectCacheRefCounter
remove dead code (ObjectCacheRefCounter)
2026-05-03 15:16:57 -07:00
Pepper Gray c548262664 include missing header unistd.h
build on clang/musl fails with:

```
FEX/unittests/APITests/Allocator.cpp:17:5: error: use of undeclared identifier 'close'
FEX/unittests/APITests/Allocator.cpp:23:5: error: use of undeclared identifier 'lseek'; did you mean 'fseek'?
FEX/unittests/APITests/Allocator.cpp:23:11: error: cannot initialize a parameter of type 'FILE *' (aka 'struct _IO_FILE *') with an lvalue of type 'int'
FEX/unittests/APITests/Allocator.cpp:24:5: error: use of undeclared identifier 'write'; did you mean '_IO_cookie_io_functions_t::write'?
FEX/unittests/APITests/Allocator.cpp:24:5: error: invalid use of non-static data member 'write'
```

include header to provide defintions

fix: #5476
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-03 15:49:40 +02:00
Pepper Gray bfc51e577f remove dead code (ObjectCacheRefCounter)
building using libc++ failes due to shared_mutex ObjectCacheRefCounter
inflating InternalThreadState beyond FEX_PAGE_SIZE, thus triggering
the static assert:

```
FEXCore/Debug/InternalThreadState.h:133:15: error: static assertion failed
FEX/FEXCore/include/FEXCore/Debug/InternalThreadState.h:133:145: note: expression evaluates to '7680 < 4096'
FEX/FEXCore/include/FEXCore/Debug/InternalThreadState.h:136:58: note: expression evaluates to '12288 == 8192'
```

remove `ObjectCacheRefCounter` as it is not used anywhere.

fix: #5456
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-03 15:13:47 +02:00
Ryan Houdek 85773995e1 Merge pull request #5470 from peppergrayxyz/header_redirect
fix include redirect for <poll.h> and <signal.h>
2026-05-03 04:41:18 -07:00
Pepper Gray 00b4777290 fix include redirect for <poll.h> and <signal.h>
building on musl/clang causes redirecting incorrect #includes warnings:

```
warning: redirecting incorrect #include <sys/poll.h> to <poll.h> [-W#warnings]
warning: redirecting incorrect #include <sys/signal.h> to <signal.h> [-W#warnings]
```

include headers instead of sys/headers.

fix: #5469
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-03 13:22:32 +02:00
Ryan Houdek 197e6de194 Merge pull request #5462 from peppergrayxyz/uc_sigmask
determine sigset_t fieldname to enhance portability (musl)
2026-05-03 03:48:53 -07:00
Ryan Houdek 9908ea4c2f Merge pull request #5466 from peppergrayxyz/libgen_h
include <libgen.h> for basename
2026-05-03 03:46:43 -07:00
Ryan Houdek c402b15bd3 Merge pull request #5464 from peppergrayxyz/tgkill
add header and classpath for tgkill
2026-05-03 03:46:07 -07:00
Ryan Houdek a3f3118ac5 Merge pull request #5460 from peppergrayxyz/sigset_t
use <signal.h> instead of glibc header to enhance portability (musl)
2026-05-03 03:18:20 -07:00
Pepper Gray b1aab0e498 determine sigset_t fieldname to enhance portability (musl)
musl build fails due to access to internal glibc member:

```
FEX/Source/Tools/LinuxEmulation/LinuxSyscalls/SignalDelegator.cpp:655:39: error: no member named '__val' in '__sigset_t'
  655 |       .SigMask = _context->uc_sigmask.__val[0],
      |                  ~~~~~~~~~~~~~~~~~~~~ ^
1 error generated.
```

add check to determine private glibc or musl member name or throw an error.

fixes: #5461
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-03 12:17:09 +02:00
Pepper Gray 927c0ce54a include <libgen.h> for basename
building on clang/musl build fails due to missing symbol:

```
/run/host/home/pepper/dev/ghrepos/FEX/Source/Tools/FEXOfflineCompiler/Main.cpp:321:35: error: use of undeclared identifier 'basename'
  321 |   auto CommandName = std::string {basename(argv[0])} + " " + (argc > 1 ? argv[1] : "");
      |                                   ^~~~~~~~
/run/host/home/pepper/dev/ghrepos/FEX/Source/Tools/FEXOfflineCompiler/Main.cpp:327:43: error: use of undeclared identifier 'basename'
  327 |     fmt::print("Usage: {} <command>\n\n", basename(argv[0]));
      |                                           ^~~~~~~~
```

include missing header.

fixes: #5465
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-03 12:10:07 +02:00
Pepper Gray 57d9dc037d add header and classpath for tgkill
building on clang/musl fails with:

```
FEX/Source/Tools/LinuxEmulation/LinuxSyscalls/GdbServer.cpp:1174:5: error: use of undeclared identifier 'tgkill'
 1174 |     tgkill(::getpid(), ::getpid(), SIGKILL);
      |     ^~~~~~
1 error generated.
```

include and use `FHU::Syscalls::tgkill`.

fixes: #5463
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-03 11:50:23 +02:00
Pepper Gray e3cfe28848 use <signal.h> instead of glibc header to enhance portability (musl)
building using musl fails with:

```
FEX/Source/Tools/LinuxEmulation/LinuxSyscalls/ThreadManager.h:33:10: fatal error: 'bits/types/sigset_t.h' file not found
   33 | #include <bits/types/sigset_t.h>
      |          ^~~~~~~~~~~~~~~~~~~~~~~
1 error generated.
```

`<bits/types/sigset_t.h>` is a glibc internal header, use <signal.h> instead.

fixes: #5459
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-03 11:03:38 +02:00
Pepper Gray ac91f583b8 make 64bit symbols visible to enhance portability (musl)
building using clang/musl fails with:

```
FEX/Source/Tools/LinuxEmulation/LinuxSyscalls/x32/Types.h:548:5: error: member access into incomplete type 'const struct statfs64'
  548 |     COPY(f_bsize);
      |     ^
```

add `_LARGEFILE64_SOURCE` to define large-file feature macros

fix: #5457
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-03 10:52:09 +02:00
Ryan Houdek 215658bf29 Merge pull request #5455 from peppergrayxyz/sys_prctl
use <sys/prctl.h> to enhance portability (clang)
2026-05-02 15:11:33 -07:00
Pepper Gray 92b1a6ea8a use <sys/prctl.h> to enhance portability (clang)
using <sys/prctl.h> and <linux/prctl.h> simultaneously causes clang to
fail:

```
In file included from FEX/FEXCore/Source/Utils/AllocatorHooks.cpp:6:
/usr/include/sys/prctl.h:88:8: error: redefinition of 'prctl_mm_map'
   88 | struct prctl_mm_map {
      |        ^
/usr/include/linux/prctl.h:134:8: note: previous definition is here
  134 | struct prctl_mm_map {
      |        ^
1 error generated.
```

prefer <sys/prctl.h> and do not include <linux/prctl.h>

fix: #5454
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-02 15:36:40 +02:00
Ryan Houdek 8ab00758be Merge pull request #5448 from bylaws/claudefix6
X87: Fix FXTRACT with Inf and NaN inputs
2026-04-30 14:42:08 -07:00
Ryan Houdek e91bda7765 Merge pull request #5447 from bylaws/claudefix5
SoftFloat: Fix FSCALE(0, +Inf) to raise IE and return a quiet NaN
2026-04-30 14:41:24 -07:00
Ryan Houdek 9db211ac97 Merge pull request #5446 from bylaws/claudefix4
OpcodeDispatcher: Fix FIST 16-bit IE for denormal inputs
2026-04-30 14:40:41 -07:00
Ryan Houdek b1381fd3b7 Merge pull request #5444 from bylaws/claudefix2
VectorOps: Fix MAXPS/MAXPD NaN and signed-zero tie selection
2026-04-30 14:39:56 -07:00
Ryan Houdek e5f6a7d85e Merge pull request #5450 from neobrain/fix_code_cache_portable
CodeCache: Make FEXServer call FEXOfflineCompiler by absolute path
2026-04-30 14:17:16 -07:00
Ryan Houdek feae76fc4f Merge pull request #5451 from Sonicadvance1/146
arm64ec: Fixes crash in many games with SDL+Dualsense
2026-04-30 14:16:06 -07:00
Ryan Houdek 015f3cffb9 arm64ec: Fixes crash in many games with SDL+Dualsense
We were pointing to an incorrect function pointer and exploding when a
pending suspend doorbell had occured.
2026-04-30 12:59:53 -07:00
Tony Wasserka 3e5c17ae80 CodeCache: Make FEXServer call FEXOfflineCompiler by absolute path
In portable mode, FEXOfflineCompiler may not be in PATH (and if it is, it's
most likely not a compatible version). Instead, use the executable next to
the FEXServer binary.
2026-04-30 17:14:22 +02:00
Billy Laws 3e278b42f8 InstcountCI: Update 2026-04-29 02:54:34 +00:00
Billy Laws fa953445e9 InstcountCI: Update 2026-04-29 02:53:07 +00:00
Billy Laws a412b1d3b7 InstcountCI: Update 2026-04-29 02:43:19 +00:00
Billy Laws fc680388e3 X87: Fix FXTRACT with Inf and NaN inputs
Both the default F80 softfloat wrappers (FXTRACT_SIG/FXTRACT_EXP) and
the reduced-precision F64 dispatcher fell through to the generic
exponent/significand extraction for Inf and NaN inputs, producing
finite garbage. The F80 wrapper returns input unchanged in the
significand slot and +Inf (or NaN) in the exponent slot; the F64
dispatcher detects the exponent-all-ones case and selects the proper
Inf/NaN result before the existing zero-case fold.
2026-04-29 02:34:41 +00:00
Billy Laws 7c260b45e1 unittests/ASM: Adds tests for FXTRACT Inf/NaN 2026-04-29 02:29:23 +00:00
Ryan Houdek 098c4c57b4 Merge pull request #5443 from bylaws/claudefix1
X86Features: Fix AVX/RAND/PCMLULQDQ feature detection on intel
2026-04-28 19:25:15 -07:00
Billy Laws 7c826e35b4 SoftFloat: Fix FSCALE(0, +Inf) to raise IE
The lhs==0 short-circuit in X80SoftFloat::FSCALE returned lhs
unchanged without calling extF80_mul, so the 0*Inf invalid-operation
case never set softfloat_flag_invalid. Detect +Inf rhs explicitly
in the zero-lhs path and raise the flag, returning QNaN to match
hardware.
2026-04-29 02:17:03 +00:00
Billy Laws 1bd2ff3fc3 unittests/ASM: Adds test for FSCALE(0, +Inf) raising IE 2026-04-29 02:16:57 +00:00
Billy Laws e2fc91fc15 OpcodeDispatcher: Fix FIST 16-bit IE for denormal inputs
The special-value check masked the biased exponent with 0x7fff and
then used TestNZ, so it triggered on Exp==0 (denormals) instead of
Exp==0x7fff (NaN/Inf). Replace the TestNZ with SubWithFlags against
0x7fff so EQ only fires for genuine NaN/Inf.
2026-04-29 02:10:35 +00:00
Billy Laws cf20647b25 unittests/ASM: Adds test for 16-bit FIST with denormal input not setting IE 2026-04-29 02:09:54 +00:00
Billy Laws 8d7071e549 VectorOps: Fix MAXPS/MAXPD NaN and signed-zero tie selection
fcmgt returns false on NaN, so the existing polarity in the non-SVE
fcmgt+bit sequences and in the SVE predicate-merge picked the wrong
source on NaN/tie. Swap the compare operands and flip bit<->bif / add
a predicate not to match x86 second source wins behaviour.
2026-04-29 02:00:58 +00:00
Billy Laws fb006b2c6d unittests/ASM: Adds test for MAXPS/MAXPD NaN and signed-zero tie 2026-04-29 01:57:41 +00:00
Billy Laws 9039eeb3cd X86Features: Fix AVX/RAND/PCMLULQDQ feature detection on intel 2026-04-29 01:57:18 +00:00
Billy Laws dd0702d30f unittests/ASM: Mark SSE4a/CLZERO as required for tests using them 2026-04-29 01:56:35 +00:00
LC 886faf0bd4 Merge pull request #5442 from neobrain/refactor_code_cache_check
CodeCache: Move bounds check to FEXOfflineCompiler
2026-04-28 19:47:52 -04:00
Tony Wasserka 86e28c6d34 CodeCache: Move bounds check to FEXOfflineCompiler
The previous check site would easily fail when loading caches for binaries
with multiple executable sections.

It makes much more sense to refuse generating caches anyway: The condition
effectively checked for invalid code map entries, so FEXOfflineCompiler
should reject them as bad inputs.
2026-04-28 17:31:20 +02:00
LC 821efab8aa Merge pull request #5440 from Sonicadvance1/145
OpcodeDispatcher: Ensure FIST* operations avoid TSO/GPR
2026-04-28 07:55:10 -04:00
Ryan Houdek 5295365dd0 InstcountCI: Update 2026-04-27 17:55:49 -07:00
Ryan Houdek 788959a98c OpcodeDispatcher: Ensure FIST* operations avoid TSO/GPR
Noticed this while benchmarking that the FIST* operations were
converting to a GPR, and then storing to memory using an atomic TSO
operation. This should be instead listening to the vector TSO
configuration option. This gives a 3.8x - 6.05x improvement in my
microbench.

Additionally when possible, make sure to use vector conversion
instructions when possible. It's lower cost to avoid the FPR->GPR
transfer, but we can only use it for 64-bit FIST operations. Microbench
couldn't show a difference for that on my platform, but that's because
it's float pipeline bounded regardless. Should help X-class Cortex and
newer Cortex-A.
2026-04-27 17:48:45 -07:00
LC dd145aaa88 Merge pull request #5439 from Sonicadvance1/144
Fix push/pop fs/gs segments and unittests
2026-04-27 19:43:44 -04:00
Ryan Houdek 34b3adc23d unittests/ASM: Adds unit test to ensure push/pop segment of o16 works
Only ensures we are pushing and popping the correct size, not any of the
selector data within it, as 64-bit systems with the FSGSBase extension
don't use them selectors anyway.

Can't test the 32-bit side currently because we would corrupt FS/GS in
CI and the host testharnessrunner can't fix that right now.
2026-04-27 15:21:47 -07:00
Simon Scherer ab14882761 FEXCore: Fix 2byte stack access for 0x66 PUSH/POP FS/GS 2026-04-27 15:21:11 -07:00
LC 7dc1f54fb6 Merge pull request #5435 from Sonicadvance1/143
unittests/ASM: Add unittests to ensure correct cmpxchg8b/16b flag setting
2026-04-25 10:34:07 -04:00
Ryan Houdek c09fb03eda unittests/ASM: Add unittests to ensure correct cmpxchg8b/16b flag setting 2026-04-25 00:46:40 -07:00
Ryan Houdek dbf2761fb7 InstcountCI: Update 2026-04-25 00:45:05 -07:00
Ryan Houdek fd1378f778 InstcountCI: Fix incorrect instruction 2026-04-25 00:43:50 -07:00
Simon Scherer 819dcee3ad FEXCore: Fix wrong shift value to extract NZCV in CmpPairZ 2026-04-25 00:41:24 -07:00
LC 4b02c04afc Merge pull request #5429 from Sonicadvance1/142
Steam/CompatTool: Fixes Graphics Provider path handling
2026-04-23 20:13:13 -04:00
Ryan Houdek deed99e7a3 Merge pull request #5425 from pmatos/f64-atan-fyl2x
JIT-inline F64ATAN and F64FYL2X for reduced precision x87 path
2026-04-23 15:16:04 -07:00
Ryan Houdek 7bffc4a177 Steam/CompatTool: Fixes Graphics Provider path handling
Graphics provider needs to be a path to a json file in the root of the
rootfs. Make sure to strip the filepath off to get the directory.

Misunderstood the assignment before.
2026-04-23 15:13:42 -07:00
Ryan Houdek 701555e400 Merge pull request #5428 from Sonicadvance1/141
Steam/CompatTool: Support `STEAM_COMPAT_GRAPHICS_PROVIDER` for rootfs path
2026-04-21 12:50:22 -07:00
Ryan Houdek 49fa86d0b5 Steam/CompatTool: Support STEAM_COMPAT_GRAPHICS_PROVIDER for rootfs path
If we have been provided a graphics provider path through an environment
variable, then use that path directly rather than searching.
2026-04-21 12:37:40 -07:00
Paulo Matos adbace8810 instcountci: JIT-inline F64ATAN and F64FYL2X for reduced precision x87 path 2026-04-21 20:23:45 +02:00
Paulo Matos 050138bcea asm_tests: JIT-inline F64ATAN and F64FYL2X for reduced precision x87 path 2026-04-21 20:23:45 +02:00
Paulo Matos db50b04663 JIT-inline F64ATAN and F64FYL2X for reduced precision x87 path 2026-04-21 12:09:55 +02:00
LC 59755ec115 Merge pull request #5426 from Sonicadvance1/139
Snapdragon X2 Elite fixes
2026-04-20 11:53:51 -04:00
Ryan Houdek 856fb1e636 Snapdragon X2 Elite fixes
- RNDRRS is still broken on this CPU
- Fault granularity checking needs to use loads
  - stlxp does monitor check before alignment check, use loads for all
    for consistency
  - Fault granularity is still 16B which matches LSE2 requirements.
- X2E still only ships a 19.2Mhz cycle counter, so still only
  ARMv9.0-a hardware
- Adds product name to CPUID
  - Oryon-3 being CPU PartID 2 isn't a mistake.
  - No distinction between Oryon-1 and Oryon-2, both are partid 1.

cpuinfo:
```
processor   : 0
BogoMIPS    : 38.40
Features    : fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm uscat ilrcpc flagm ssbs sb paca pacg dcpodp sve2 sveaes svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 rng ecv afp rpres
CPU implementer   : 0x51
CPU architecture: 8
CPU variant : 0x1
CPU part    : 0x002
CPU revision      : 1
```
2026-04-19 14:34:44 -07:00
Ryan Houdek d41d52b889 Merge pull request #5419 from pmatos/f64-scale-f2xm1
JIT-inline F64SCALE and F64F2XM1 for reduced precision x87 path
2026-04-17 14:38:22 -07:00
Paulo Matos 18f69fb16d instcountci: JIT-inline F64SCALE and F64F2XM1 for reduced precision x87 path 2026-04-17 17:28:13 +02:00
Paulo Matos d165711f2e asm_tests: JIT-inline F64SCALE and F64F2XM1 for reduced precision x87 path 2026-04-16 11:24:24 +02:00
Paulo Matos 6bf19ffe37 JIT-inline F64SCALE and F64F2XM1 for reduced precision x87 path 2026-04-16 11:24:19 +02:00
LC 441116e1e6 Merge pull request #5421 from Sonicadvance1/136
Scripts: Move arch check first in InstallFEX
2026-04-15 16:24:04 -04:00
Ryan Houdek ce97ef0ab1 Scripts: Move arch check first in InstallFEX
Don't give people false hope that the script might work on distros that
aren't Ubuntu.

Fixes #5420
2026-04-15 13:01:37 -07:00
LC 2ea0de92f4 Merge pull request #5418 from Sonicadvance1/135
ArchHelpers: Allow atomic memory operations in non-JIT handler
2026-04-15 07:22:25 -04:00
Ryan Houdek 14580c4675 ArchHelpers: Allow atomic memory operations in non-JIT handler
`Detroit: Become Human` decided to use unaligned CriticalSections. So
this workarounds that.
2026-04-14 13:38:41 -07:00
Ryan Houdek 9681559d56 Docs: Update for release FEX-2604 2026-04-09 13:45:35 -07:00
Ryan Houdek b478e4845f Merge pull request #5417 from tiopex/main
FEXRootFSFetcher: clear Unknown when distro is set on the CLI
2026-04-09 13:42:55 -07:00
tpietrus 1fa5104076 FEXRootFSFetcher: clear Unknown when distro is set on the CLI 2026-04-09 07:53:53 +02:00
LC ce65f5376f Merge pull request #5415 from Sonicadvance1/133
FEX: Workaround Docker seccomp bug
2026-04-08 11:39:56 -04:00
LC 0695249fc8 Merge pull request #5413 from Sonicadvance1/132
Arm64EC: Invert suspend doorbell and move out of hot path
2026-04-06 19:49:49 -04:00
LC 51144c99a7 Merge pull request #5408 from Sonicadvance1/129
FEXCore: Use WritePriorityMutex for Code Invalidation Mutex
2026-04-06 19:49:07 -04:00
LC 251398a7cb Merge pull request #5416 from Sonicadvance1/134
OpcodeDispatcher: Fixes nop encoded prefetch instruction
2026-04-06 19:45:52 -04:00
Ryan Houdek 2e6a7f869c OpcodeDispatcher: Fixes nop encoded prefetch instruction
We had a bug where nop encoded prefetch instructions were getting
flagged as illegal instructions erroneously. Fix that and add a unittest
for ensuring execution.

Fixes `Devil May Cry 4`
2026-04-06 11:04:46 -07:00
Ryan Houdek f308162334 FEX: Workaround Docker seccomp bug
Docker's seccomp filter fails to follow AAPCS64 and SysV zero-extension
rules.  For values smaller than 64-bit they were required in their
seccomp filters to truncate the value to the specific size but do not.
Instead they do a 64-bit comparison operation against smaller arguments
(in this case 32-bit). This means 64-bit -1 and 32-bit -1 passed through
have different values for this `personality` syscall.

The real fix would be for Docker to audit their seccomp filter rules and
ensure they zero-extend every argument that is smaller than 64-bit, but
we don't control that. So there is likely to be more bugs in their
filter that we encounter, this is just an easy one to resolve.
2026-04-06 09:51:15 -07:00
Ryan Houdek db4867839c Arm64EC: Invert suspend doorbell and move out of hot path
This was causing a surprisingly high amount of branch mispredicts in
Death Stranding. Suspend doorbell is fairly rare so just invert the
check and move the target down out of the hot path. Then the doorbell
handling code will trampoline to the correct location still.
2026-04-03 20:07:37 -07:00
LC 73ffff7d22 Merge pull request #5409 from Sonicadvance1/130
OpcodeDispatcher: Special case optimize a broadcast
2026-04-03 20:23:39 -04:00
LC dc48a4f73c Merge pull request #5406 from Sonicadvance1/128
FEXRootFSFetcher: Improve hashing performance
2026-04-03 00:26:44 -04:00
Ryan Houdek efbccccdc0 InstcountCI: Update 2026-04-02 18:45:22 -07:00
Ryan Houdek 3e7cd88dcc OpcodeDispatcher: Special case optimize a broadcast
Death Stranding 2 is using this instruction instead of `vbroadcastss`
for some reason. Optimize its specific case and add a note that when we
know sources match that we can optimize more patterns easily.
2026-04-02 18:45:22 -07:00
Ryan Houdek 12cfe8fc37 InstcountCI: Add instruction found in Death Stranding 2 2026-04-02 18:25:19 -07:00
Tony Wasserka c6d2ce043f Merge pull request #5388 from Sonicadvance1/123
Config: Finish wiring up Regex app overrides
2026-04-02 11:00:37 +02:00
Tony Wasserka 5c34c574c8 Merge pull request #5364 from Sonicadvance1/110
Win32: Enable support for virtual naming and THP control
2026-04-02 10:58:30 +02:00
Ryan Houdek d3cfdcb431 Win32: Enable support for virtual naming and THP control
Allows WTF to work (mostly) with Wine by letting us VirtualName things,
and also allows madvise control of THP, which significantly cuts back
memory usage.

This works around the problem of Wine not giving us control of this by
using raw syscalls when wine is detected.

Based on top of #5362 so the THP disable controls are in.
2026-04-01 10:50:40 -07:00
Ryan Houdek 4018d23c39 Config: Finish wiring up Regex app overrides
This wasn't quite wired up exactly how we wanted it. It was previously
matching against the opaque file config handle, which can be anything.

Instead compare it to the appname that now gets passed over to it for
matching.

This allows us to do the following:
```
{
    "Config": {
        "ProfileStats": "1",
        "X87ReducedPrecision": "1",
        "TSOEnabled": "1",
        "VectorTSOEnabled": "0",
        "MemcpySetTSOEnabled": "0",
        "HalfBarrierTSOEnabled":"1",
        "MaxInst": "500",
        "Multiblock": "1"
    },
    "AppOverrides" : {
        "setup*" : {
            "Comment": [
                "292030 - The Witcher 3: Wild Hunt"
            ],
            "X87ReducedPrecision": "0"
        }
     }
}
```

Based on #121 which needs to get merged first.

Code Review

Code Review: Class deletion
2026-04-01 10:46:02 -07:00
badumbatish 81d4e8fe9d Initial implementation for regex engine
Add support for question mark and plus mark in regex, supply testing for star

Added more characters to the regex alphabets, add more test case

Added support for regex matching of configs, awaiting reviews

Rename variable to CamelCase

Addresses PR reviews

Remove unnecessary features and test cases

Rewrite to naive regex with dp

Addresses PR reviews

Build fixes

Code Review
2026-04-01 10:45:59 -07:00
Tony Wasserka 34b48c4069 Merge pull request #5383 from Sonicadvance1/119
FEXGetConfig: Test for showing fault granularity
2026-04-01 10:48:18 +02:00
Ryan Houdek 01a3ab6ca7 FEXCore: Use WritePriorityMutex for Code Invalidation Mutex
The previous `ForkableSharedMutex` using `std::shared_mutex` was showing
up as significant CPU time on arm64ec. In particular it was showing up
upwards of 700ms/S of CPU time for read-contented workloads at only
~2700 locks per second. The libc++ implementation for arm64ec must be
particularly gnarly for this to be so slow.

With this swapped over, it's now only spending around 56ms/S on the
contended shared lock, but at ~8000 locks per second. So a significant
uplift.
2026-03-31 19:02:51 -07:00
Ryan Houdek 476c242d7f FEXCore: Moves SpinWaitLock and WritePriorityMutex to frontend visible includes
This will be used in a moment.
2026-03-31 19:02:51 -07:00
LC ae3fa6a836 Merge pull request #5403 from Sonicadvance1/127
IR: Adds support for printing strings
2026-03-31 16:23:42 -04:00
Ryan Houdek ba93bdd66d FEXRootFSFetcher: Improve hashing performance
Don't use pread, instead map the file and madvise larger blocks. This
removes copying overhead as its just mapping file pages in instead.
Also splits the implementation of file reading from hashing to make
tinkering less involved, as if I want more performance out of this (say
due to live hashing) then it's easier to tinker.

Improves hashing performance from ~2.2GB/s to ~3.6GB/s on my system,
which is CPU bounded by xxhash here.
2026-03-30 18:16:46 -07:00
Ryan Houdek 1fa0b37fac FEXGetConfig: Test for showing fault granularity
Useful for seeing if behaviour has changed. Useful with the
`--tso-emulation-info` option to show hardware behaviour
2026-03-30 12:32:00 -07:00
LC 5b4a5969cc Merge pull request #5401 from Sonicadvance1/126
Linux/GuestFrames: Ensure MXCSR is saved and restored on signal
2026-03-28 22:25:59 -04:00
Ryan Houdek 941a7934ef InstcountCI: Update 2026-03-28 18:06:54 -07:00
Ryan Houdek 474439ab4b InstcountCI: Update 2026-03-28 17:56:23 -07:00
Ryan Houdek 4428aea5ee IR: Adds support for printing strings
I utilize this functionality quite heavily when debugging and I need
bread crumbs spread around. Instead of reimplementing it a dozen times,
just have it upstreamed.
2026-03-28 17:53:33 -07:00
Ryan Houdek e9a9cc5bc3 unittests/FEXLinuxTests: Adds an MXCSR signal test
Ensures that the MXCSR value stays the same with a signal inbetween that
modifies it.
2026-03-28 17:49:12 -07:00
Ryan Houdek 2291c5b230 OpcodeDispatcher/Vector: Make sure MXCSR is masked
We don't support the exception bits, make sure these are masked off so
spurious exception checks don't break.
2026-03-28 17:47:55 -07:00
Ryan Houdek f78e194cf7 Linux/GuestFrames: Ensure MXCSR is saved and restored on signal
This was causing an unfortunate set of circumstances where Dark Souls
III was modifying MXCSR and we weren't saving it, cause the value to change
from 0x9fc0 to 0.

This "enabled" float exceptions by unmasking the exception masks in
MXCSR. This in turn had Dark Souls III's `expf` function to fault out,
as it checks if the MXCSR exception masks are set or not for determining
if underflow should assert or not.

Wow64/arm64ec has a similar problem where it always sets back to default
on signal. Which means game lose DAZ, but I'm not fixing that bug right
now.

Fixes #5391
2026-03-28 17:44:45 -07:00
LC b77ddcf1a7 Merge pull request #5398 from Sonicadvance1/125
Allocators: Remove legacy NOREPLACE handling
2026-03-27 08:14:05 -04:00
Ryan Houdek 3ef677537d Merge pull request #5397 from neobrain/fix_elfreads
LinuxSyscalls: Skip reading ELF files when code caching is disabled
2026-03-26 14:33:26 -07:00
Tony Wasserka 1df1265ed1 LinuxSyscalls: Skip reading ELF files when code caching is disabled
ELF headers were read unconditionally because doing so was assumed to be cheap
(as the guest app would read them anyway shortly after). However, relocation
parsing was added since then, which has less predictable performance due to
crossing page boundaries and reading larger amounts of memory. It might be
possible to make the underlying code more efficient, but until that's done
it's better to skip this logic unless needed.

Closes #5390
2026-03-26 21:40:20 +01:00
Ryan Houdek f0854a16fe Allocators: Remove legacy NOREPLACE handling
We needed this handling on old kernels that didn't understand the
NOREPLACE flag. We no longer support kernels this old, so remove some of
this vestigial code.
2026-03-26 13:36:29 -07:00
Ryan Houdek 6bd476fb03 Merge pull request #5395 from lioncash/cpuid
CPUID: Add basic stub handling for AVX10 info
2026-03-25 17:16:24 -07:00
Lioncache b9c0af7c3b CPUID: Add basic handling for AVX10 info
Just gets the feature bit handling stuff in place for various
facilities, so it can be easily expanded in the future.
2026-03-25 19:18:05 -04:00
Tony Wasserka 5149ebc70e Merge pull request #5389 from Sonicadvance1/124
SMCTracking: Remove relocation log
2026-03-24 11:12:10 +01:00
Ryan Houdek e92a6a5803 SMCTracking: Remove relocation log
Holy jeez does this thing spam.
2026-03-23 19:55:11 -07:00
Ryan Houdek 6da963a695 Merge pull request #5394 from neobrain/fix_eager_relocation_parsing
LinuxSyscalls: Don't try to parse relocations unless ELF parsing succeeded
2026-03-23 14:01:20 -07:00
Tony Wasserka 2a74489858 LinuxSyscalls: Don't try to parse relocations unless ELF parsing succeeded 2026-03-23 17:47:47 +01:00
Ryan Houdek 65a436ca98 Merge pull request #5392 from neobrain/fix_code_cache_dupfd
LinuxSyscalls: Re-open file descriptors for parsing ELF headers
2026-03-23 09:42:45 -07:00
LC bc533c8050 Merge pull request #5393 from neobrain/fix_code_cache_glibc_assert
CodeCache: Fix glibc debug mode assertion
2026-03-22 14:07:27 -04:00
Tony Wasserka fd6cea4698 CodeCache: Fix glibc debug mode assertion
If begin == end, the first vector::erase() call would invalidate the begin
iterator.
2026-03-22 10:11:53 +01:00
Tony Wasserka a1aa1658ec LinuxSyscalls: Re-open file descriptors for parsing ELF headers
File descriptors returned by dup() share state with the original FD, so we
need to use open() to create a fully independent object.

Fixes #5379
2026-03-22 10:09:37 +01:00
LC 5c4c468d13 Merge pull request #5387 from Sonicadvance1/122
github: Stop running unittests always on build failure
2026-03-20 00:58:09 -04:00
Ryan Houdek 69ef1658cc github: Stop running unittests always on build failure
This was taking too much time.
2026-03-19 20:08:10 -07:00
Ryan Houdek 8c72aa76a0 Merge pull request #5385 from lioncash/ilog
MemoryOps: Collapse duplicate add/sub in Memset
2026-03-19 18:59:17 -07:00
Lioncache 928a932a43 MemoryOps: Collapse duplicate add/sub in Memset
We can just use ilog2 to deduplicate this a bit.
2026-03-19 21:34:03 -04:00
Ryan Houdek 83601055dc Merge pull request #5368 from neobrain/feature_cc_elf_relocations
CodeCache: Support ELF relocations
2026-03-19 18:10:09 -07:00
Ryan Houdek 68480f6e43 Merge pull request #5356 from lioncash/mops
MemoryOps: Drop MOPS handling into place for MemSet/MemCpy
2026-03-19 18:04:14 -07:00
Ryan Houdek 42291540ab Merge pull request #5343 from pmatos/f64-sin-cos-tan
JIT-inline F64SIN, F64COS, and F64TAN for reduced precision x87 path
2026-03-19 17:51:50 -07:00
LC 194eb69838 Merge pull request #5384 from Sonicadvance1/120
Cmake: Default to release builds with a message
2026-03-19 19:42:39 -04:00
Ryan Houdek c1d27fa453 Cmake: Default to release builds with a message
People keep forgetting to set this and have a worse experience.
Default to a Release build, which ensures optimizations are enabled and
assertions are disabled.
2026-03-19 16:25:40 -07:00
Paulo Matos 9d5f7caa79 instcountci: JIT-inline F64SIN, F64COS, and F64TAN for reduced precision x87 path 2026-03-19 19:18:02 +01:00
Paulo Matos a1d78dceb0 asm_tests: JIT-inline F64SIN, F64COS, and F64TAN for reduced precision x87 path 2026-03-19 19:18:02 +01:00
Lioncache 2bcf435e0a MemoryOps: Handle overlapping memcpy 2026-03-19 14:13:36 -04:00
Paulo Matos 30e853305d JIT-inline F64SIN, F64COS, and F64TAN for reduced precision x87 path 2026-03-19 18:58:46 +01:00
Lioncache 600f4bb2b6 unittests: Add overlapping tests for MemCpy 2026-03-19 11:55:02 -04:00
Lioncache 5868814c91 MemoryOps: Drop in MOPS handling for MemCpy
With the MOPS featureset dropped in, we can also accelerate memcpy paths
on hardware that supports it.
2026-03-19 11:55:02 -04:00
Lioncache e862f8f86c unittests: Add specific paths for MOPS 2026-03-19 11:55:02 -04:00
Lioncache 85c1ecd035 MemoryOps: Handle inline values in MemSet() MOPS path
Lets us handle potential inline memset values.

Also fixes up the STOS tests to actually ensure all values
in the verification step pass.
2026-03-19 11:55:02 -04:00
Lioncache 68ad448672 MemoryOps: Drop 8-bit memset support into MemSet()
Can be further expanded to handle other optimization cases, but this
kicks it off for forward direction memsets at least.
2026-03-19 11:55:02 -04:00
LC c18fb3cb78 Merge pull request #5382 from Sonicadvance1/118
FEXpidof: Fixes another missing std::filesystem throw
2026-03-18 18:07:08 -04:00
Tony Wasserka 53702f989c Merge pull request #5362 from Sonicadvance1/108
FEX: Disable THP on key allocations that consume memory
2026-03-18 21:04:15 +01:00
Ryan Houdek c547b1bec3 FEX: Disable THP on key allocations that consume memory
Disables THP on some key locations that are fairly sparse
- rpmalloc
  - This is the big one as this allocates some heavy sparse buffers.
- CallRet stacks
  - These get in the hundreds of megabytes, while not being sparse they
    trend towards only using a handful of pages and ballooning to 2MB
    per thread is quite heavy.
- Lookup cache
  - L1 specifically gets hit here which adds a decent chunk of overhead
    due to sparsity.

Win32 for all of these also aren't handled, but that will need to be a
followup.
2026-03-18 12:15:39 -07:00
Ryan Houdek 740350c8ea FEXpidof: Fixes another missing std::filesystem throw
Turns out std::filesystem::exists throws as well if there was an
underlying OS API failure.
2026-03-18 12:10:22 -07:00
Tony Wasserka 9f9b20eac0 Merge pull request #5378 from Sonicadvance1/117
SMCTracking: Move read check up for ELF parsing
2026-03-18 12:04:23 +01:00
Tony Wasserka c67ffb82a8 Core: Support reporting blocks that are uncacheable due to unhandled ELF relocations 2026-03-18 11:59:44 +01:00
Tony Wasserka 1ea24f3d6c LinuxSyscalls: Enable delayed code cache load for ELF files
Specifically this is needed if any ELF relocations cover read-only code
sections, which is indicated in the ELF headers via DT_TEXTREL/DF_TEXTREL.
2026-03-18 11:59:41 +01:00
Tony Wasserka 226bd51afe LinuxSyscalls: Implement delayed cache load for binaries that require ELF/PE relocations 2026-03-18 11:58:26 +01:00
Tony Wasserka 293568be36 FEXOfflineCompiler: Apply relocations to loaded ELF binaries 2026-03-18 11:57:38 +01:00
Tony Wasserka 152fe81d16 LinuxSyscalls: Parse and provide ELF relocation information to the JIT 2026-03-18 11:57:28 +01:00
Ryan Houdek 494dd64c50 Merge pull request #5372 from CxnYusuf/add-fisttp-tests
Tests/ASM: Add FISTTP unit tests for 16, 32, 64-bit as well as negatives
2026-03-17 16:55:56 -07:00
Ryan Houdek 56de0d1ab4 Merge pull request #5377 from Sonicadvance1/116
FEXServer: Try both fusermount and fusermount3
2026-03-17 16:55:42 -07:00
Ryan Houdek 73c1f4cc54 Merge pull request #5374 from neobrain/fix_gcc_build
Fix most GCC build issues
2026-03-17 16:55:23 -07:00
Ryan Houdek 6a6a82385e Merge pull request #5381 from neobrain/fix_jit_restarts
JIT: Reset relocations on restart
2026-03-17 13:47:42 -07:00
Tony Wasserka fc8ef0e723 JIT: Reset relocations on restart 2026-03-17 21:37:01 +01:00
Ryan Houdek de11c05d2a Merge pull request #5380 from neobrain/fix_cc_32bit_constants
Arm64Emitter: Disable 32-bit constant optimization when NOP-padding is requested
2026-03-17 13:34:20 -07:00
Tony Wasserka addbc8cad8 Arm64Emitter: Disable 32-bit constant optimization when NOP-padding is requested
Code caching requires this even for simple libraries like libdl.so (as observed
in the 32-bit build of Super Meat Boy).
2026-03-17 21:06:13 +01:00
Ryan Houdek 63e37b7cbb SMCTracking: Move read check up for ELF parsing 2026-03-17 12:43:54 -07:00
Ryan Houdek 9ee329034f FEXServer: Try both fusermount and fusermount3
Apparently some distros don't symlink these, so try both with the newer
fusermount3 going first as its the common path now.

Fixes #5375
2026-03-17 12:36:49 -07:00
Ryan Houdek f3e904207b Merge pull request #5376 from OFFTKP/flag
Add test for shifts preserving flags
2026-03-17 09:40:59 -07:00
LC e4ae6ce635 Merge pull request #5373 from Sonicadvance1/115
code-format-helper: Another dependabot upgrade
2026-03-17 10:10:46 -04:00
Paris Oplopoios f7d76255ad Add test for shift preserving flags
Signed-off-by: Paris Oplopoios <21157395+OFFTKP@users.noreply.github.com>
2026-03-17 15:50:33 +02:00
Tony Wasserka ea45f9c694 FEXCore/VectorRegType: Use vector_size on GCC
GCC does not support neon_vector_type and silently ignores that attribute,
but vector_size(16) seems to have the same effect.
2026-03-16 19:15:05 +01:00
Tony Wasserka 1d449c0f58 FEXCore/Utils: Add quotes around preprocessor errors 2026-03-16 19:15:05 +01:00
Tony Wasserka 9ecc991043 LibraryForwarding: Remove unnecessary const qualifier 2026-03-16 19:15:05 +01:00
Tony Wasserka 4ef834859c SignalDelegator: Don't use the same name for two different symbols 2026-03-16 19:15:05 +01:00
Tony Wasserka fa082bc5c4 FileManagement: Fix ambiguous name reference 2026-03-16 19:15:05 +01:00
Tony Wasserka 67caab026a OpcodeDispatcher: Fix inconsistent types in ternary conditional 2026-03-16 19:15:05 +01:00
Tony Wasserka 91c55facb2 IRDumper: Avoid passing packed member data as references
Fixes "Cannot bind packed field to reference type" build errors on GCC.
2026-03-16 19:15:05 +01:00
Tony Wasserka 321d4d84d7 LinuxSyscalls: Don't cast away qualifiers 2026-03-16 19:15:05 +01:00
Tony Wasserka 22faa58e0b X86Tables: Use explicit type for SecondInstGroupOps definition
GCC considers it a "conflicting declaration" to use auto for a variable that
was already declared before.
2026-03-16 19:15:05 +01:00
Tony Wasserka ebd559f662 Core: Fix offsetof with runtime array indexes
GCC does not support this clang-specific language extension.
2026-03-16 19:15:05 +01:00
Tony Wasserka d27c9d3f98 CodeEmitter: Fix ambigious ExtendedType declaration 2026-03-16 18:50:01 +01:00
Tony Wasserka fbef482265 CMake: Link against libatomic if compiling with GCC 2026-03-16 18:50:01 +01:00
Tony Wasserka a57926ac57 CMake: Explicitly demote -Wchanges-meaning diagnostics to warnings on GCC 2026-03-16 18:50:01 +01:00
Ryan Houdek 70a7137e62 code-format-helper: Another dependabot upgrade 2026-03-15 20:45:17 -07:00
CxnYusuf 2d3a08a362 Tests/ASM: Add FISTTP unit tests for 16, 32, 64-bit as well as negatives 2026-03-16 03:10:38 +01:00
LC cc02edb3f6 Merge pull request #5371 from Sonicadvance1/114
Syscalls: Fixes crash in ELF parsing code
2026-03-15 20:36:00 -04:00
Ryan Houdek f894cd90f3 Merge pull request #5369 from Sonicadvance1/113
FEXCore: Update CPU frequency to be 64-bit
2026-03-15 15:20:28 -07:00
Ryan Houdek 24675969cc Merge pull request #5366 from Sonicadvance1/112
JIT: Use struct for Spill/Fill default arguments
2026-03-15 15:20:13 -07:00
Ryan Houdek a0cba1194e Syscalls: Fixes crash in ELF parsing code
When an application maps a file as PROT_NONE, we can't check if it is an
ELF. Was causing a crash in `Cisco Packet Tracer`.
2026-03-15 15:16:29 -07:00
Ryan Houdek 0952fa95f4 FEXCore: Update CPU frequency to be 64-bit
This annoyed by by trying to do math on a 32-bit value and it
overflowing. Just make it 64-bit.
2026-03-13 17:04:25 -07:00
Ryan Houdek 6177ab957b Merge pull request #5363 from Sonicadvance1/109
External/code-format-helper: Update dependencies
2026-03-13 11:52:07 -07:00
Ryan Houdek 5c1300a2b9 Merge pull request #5365 from Sonicadvance1/111
Config: Fix issue with config overrides
2026-03-13 11:51:51 -07:00
Ryan Houdek af9dd0827a JIT: Use struct for Spill/Fill default arguments
Cleans up the interface and makes the arguments explicit about what
they're setting. As promised from #5317
2026-03-12 19:25:12 -07:00
Ryan Houdek dc0162122f Config: Fix issue with config overrides
Accidentally was checking for Config override in the combination of
portable config and `FEX_APP_CONFIG_LOCATION`.

Fixes an early crash in PV.
2026-03-12 17:46:05 -07:00
Ryan Houdek ae491fb15b External/code-format-helper: Update dependencies
Removes dependabot alert.
2026-03-12 15:31:02 -07:00
Ryan Houdek 957c1fc420 Merge pull request #5357 from Sonicadvance1/106
Config: Enable Dynamic L1 and Disabled L2 caches by default
2026-03-11 14:12:38 -07:00
Ryan Houdek 86acfb35aa Config: Enable Dynamic L1 and Disabled L2 caches by default
Dramatically reduces memory consumption of FEX's per-thread lookup
structures. Primarily because L2 cache entirely goes away which can end
up reaching hundreds of megabytes or over a gigabyte of memory in some
cases, but also because L1 cache dynamically scales based on load.

Useful for conserving memory on systems with less than 16GB of RAM and
are UMA, like Asahi users inside of muvm.
2026-03-11 13:42:32 -07:00
LC e27d12ee5e Merge pull request #5359 from Sonicadvance1/107
OpcodeDispatcher: Fixes memory leak in IREmitter pool allocator
2026-03-11 08:40:04 -04:00
Ryan Houdek 63d52c2a1a OpcodeDispatcher: Fixes memory leak in IREmitter pool allocator
While not a leak in the traditional sense, we were causing pool
allocations to never become free until the thread was closed.

This meant in the case of a game running with >200 threads or so, these
would add up very quickly. So some minor reworking so the IREmitter
doesn't allocate a buffer until first JIT, and making sure to actually
disown the buffer on dispatch error resolved the problems.

Fixes an edge case where Ender Lilies was consuming 409MB with THP
enabled on my desktop, and now it is something like 6MB once idling for
a bit to have the pool allocations do its magic.
2026-03-10 20:55:09 -07:00
Ryan Houdek 9d4a71b57a Merge pull request #5355 from wsxarcher/patch-1
Handle zero length in ChangeProtectionFlags
2026-03-10 08:40:16 -07:00
Marco Bartoli 3462dc3e14 Handle zero length in ChangeProtectionFlags
Add a no-op for zero length in ChangeProtectionFlags.

This fixes AMD Vivado 2025.2 which tries to mprotect with 0 as size and merge strategies fails:

```
Unexpected ChangeProtectionFlags Merge strategy! [0x400000, 0x401000) Versus [0x0, 0x0)
```
2026-03-10 11:34:52 +01:00
LC d21351e66e Merge pull request #5354 from Sonicadvance1/104
Config: Fixes
2026-03-09 21:51:05 -04:00
Ryan Houdek 3204d20335 Config: Fix priorities of config paths
Fixes b4a87d8c0b

`FEX_APP_CONFIG_LOCATION` wasn't overriding paths properly anymore once
that commit landed. Instead legacy `~/.fex-emu/` path would get returned
if it existed first.

Ensures that it returns first, before `STEAM_COMPAT_DATA_PATH` even.
2026-03-09 18:20:56 -07:00
Ryan Houdek a519489d80 CMake: Make sure not to compile Steam tools on mingw 2026-03-09 17:25:59 -07:00
Ryan Houdek 5558c3a35a Merge pull request #5353 from lioncash/group
HostFeatures: Group feature ifdefs together more
2026-03-09 15:45:55 -07:00
Ryan Houdek 58d9755314 Merge pull request #5352 from Sonicadvance1/103
gitlab-ci: Update requirements
2026-03-09 15:45:48 -07:00
Lioncache 498ba0a384 HostFeatures: Group feature ifdefs together more
Makes it a little nicer to see everything grouped together.
2026-03-09 18:27:56 -04:00
Ryan Houdek 5a5477e895 gitlab-ci: Update requirements 2026-03-09 14:00:08 -07:00
Ryan Houdek a17d7ce6ba Merge pull request #5351 from lioncash/hostmops
HostFeatures: Drop in feature testing for FEAT_MOPS
2026-03-09 12:09:31 -07:00
LC afc7248912 Merge pull request #5347 from Sonicadvance1/102
CPUBackend: Enable Transparent Huge Pages on JIT buffers
2026-03-09 14:54:08 -04:00
Lioncache 6bb578fea8 HostFeatures: Drop in feature testing for FEAT_MOPS 2026-03-09 13:03:45 -04:00
Ryan Houdek bed5f293dc CPUBackend: Enable Transparent Huge Pages on JIT buffers
If/When this works, the amount of iTLB misses drop dramatically, which
reduces L2 TLB pressure, which just improves performance for our CPU
cores that have itty-bitty L1 iTLB entry counts.

We can't use `mmap(MAP_HUGETLB)` directly for terrible reasons, so we
are required to lean on madvise instead.
2026-03-06 11:51:17 -08:00
207 changed files with 21850 additions and 15338 deletions

No files matched your search

+14 -13
View File
@@ -49,6 +49,7 @@ jobs:
run: cmake --build build --target asm_files 32bit_asm_files JemallocLibs Catch2 vixl cephes_128bit
- name: Build
id: build
run: cmake --build build
- name: Install
@@ -56,40 +57,40 @@ jobs:
# GCC tests
- name: GCC64 Target Tests
if: ${{ always() }}
if: steps.build.outcome == 'success'
uses: ./.github/workflows/test
with:
target: gcc_target_tests_64
- name: GCC32 Target Tests
if: ${{ always() }}
if: steps.build.outcome == 'success'
uses: ./.github/workflows/test
with:
target: gcc_target_tests_32
# API tests
- name: API Tests
if: ${{ always() }}
if: steps.build.outcome == 'success'
uses: ./.github/workflows/test
with:
target: api_tests
- name: FEXCore API Tests
if: ${{ always() }}
if: steps.build.outcome == 'success'
uses: ./.github/workflows/test
with:
target: fexcore_apitests
# ARM emission tests
- name: ARM Emitter Tests
if: ${{ always() }}
if: steps.build.outcome == 'success'
uses: ./.github/workflows/test
with:
target: emitter_tests
# Linux tests
- name: FEX Linux Tests
if: ${{ always() }}
if: steps.build.outcome == 'success'
uses: ./.github/workflows/test
with:
target: fex_linux_tests_all
@@ -98,13 +99,13 @@ jobs:
# Thunking
- name: Thunkgen tests
if: ${{ always() }}
if: steps.build.outcome == 'success'
uses: ./.github/workflows/test
with:
target: thunkgen_tests
- name: Test GL No-Thunks
if: ${{ always() && matrix.arch[1] == 'x64' }}
if: ${{ steps.build.outcome == 'success' && matrix.arch[1] == 'x64' }}
uses: ./.github/workflows/test
with:
target: thunk_functional_tests_nothunks
@@ -112,7 +113,7 @@ jobs:
DISPLAY: ':0'
- name: Test GL Thunks
if: ${{ always() && matrix.arch[1] == 'x64' }}
if: ${{ steps.build.outcome == 'success' && matrix.arch[1] == 'x64' }}
uses: ./.github/workflows/test
with:
target: thunk_functional_tests_thunks
@@ -121,28 +122,28 @@ jobs:
# ASM tests
- name: ASM Tests
if: ${{ always() }}
if: steps.build.outcome == 'success'
uses: ./.github/workflows/test
with:
target: asm_tests
# POSIX tests
- name: POSIX Tests
if: ${{ always() }}
if: steps.build.outcome == 'success'
uses: ./.github/workflows/test
with:
target: posix_tests
# GVisor tests
- name: GVisor Tests
if: ${{ always() }}
if: steps.build.outcome == 'success'
uses: ./.github/workflows/test
with:
target: gvisor_tests
# Struct verifier tests
- name: Struct verifier tests
if: ${{ always() }}
if: steps.build.outcome == 'success'
uses: ./.github/workflows/test
with:
target: struct_verifier
+40 -1
View File
@@ -1,3 +1,17 @@
spec:
inputs:
PROMOTE_BRANCH:
description: "Branch to promote the build to. Empty means no promotion."
default: "bleeding-edge"
---
workflow:
rules:
- when: always
variables:
PROMOTE_BRANCH: $[[ inputs.PROMOTE_BRANCH ]]
variables:
DEBIAN_FRONTEND: noninteractive
GIT_SUBMODULE_STRATEGY: recursive
@@ -5,7 +19,8 @@ variables:
CC: clang
CXX: clang++
aarch64:
build:
stage: build
image: registry.gitlab.steamos.cloud/steamrt/steamrt4/sdk/arm64:4.0.20251117.183306
tags:
- docker
@@ -30,3 +45,27 @@ aarch64:
untracked: false
paths:
- install/
promote:
stage: deploy
variables:
GIT_STRATEGY: none
image: registry.gitlab.steamos.cloud/steamrt/steamrt4/sdk/arm64:4.0.20251117.183306
tags:
- docker
- linux
- arm64
- aarch64
rules:
- if: '$PROMOTE_BRANCH'
before_script:
- apt-get -y update
- apt-get install -y tmux curl
script:
# comment out to debug: SSH in via GCP, go down the container and attach to the session (with `tmux attach -t debug`)
# - tmux new-session -d -s debug
# - while tmux has-session -t debug 2>/dev/null; do sleep 1; done
# ref controls which fex-depot code runs the pipeline, while VERSION_PARAM controls which fex branch's artifacts that pipeline downloads.
- >
curl --fail --location --request POST --form token=${FEX_DEPOT_TRIGGER_TOKEN} --form ref=master --form "variables[PROMOTE_BRANCH]=${PROMOTE_BRANCH}" --form "variables[VERSION_PARAM]=${CI_COMMIT_REF_NAME}" "${CI_API_V4_URL}/projects/fex%2Ffex-depot/trigger/pipeline"
+15 -1
View File
@@ -172,6 +172,13 @@ set(TUNE_ARCH "generic" CACHE STRING "Override the Arch the build is tuned for")
set(OVERRIDE_VERSION "detect" CACHE STRING "Override the FEX version")
set(OVERRIDE_HASH "detect" CACHE STRING "Override the FEX git hash")
get_property(IS_MULTI_CONFIG GLOBAL PROPERTY GENERATOR_IS_MULTI_CONFIG)
if (NOT IS_MULTI_CONFIG AND NOT CMAKE_BUILD_TYPE)
set(CMAKE_BUILD_TYPE Release
CACHE STRING "Choose the type of build." FORCE)
message(STATUS "No build type set, defaulting to a Release build")
endif()
string(TOUPPER "${CMAKE_BUILD_TYPE}" CMAKE_BUILD_TYPE)
if (CMAKE_BUILD_TYPE MATCHES "DEBUG")
set(ENABLE_ASSERTIONS TRUE)
@@ -187,6 +194,8 @@ if (ENABLE_GDB_SYMBOLS)
add_compile_definitions(GDB_SYMBOLS_ENABLED=1)
endif()
add_compile_definitions(_LARGEFILE64_SOURCE)
set(CMAKE_CXX_STANDARD 20)
set(CMAKE_EXPORT_COMPILE_COMMANDS ON)
set(CMAKE_RUNTIME_OUTPUT_DIRECTORY ${CMAKE_BINARY_DIR}/Bin)
@@ -446,6 +455,11 @@ if(ENUM_ENUM_WARNING)
add_compile_options(-Wno-deprecated-enum-enum-conversion)
endif()
# GCC enables -Wchanges-meaning by default and treats some cases as an error
if(CMAKE_CXX_COMPILER_ID STREQUAL "GNU")
add_compile_options(-Wno-error=changes-meaning)
endif()
if(ENABLE_WERROR OR ENABLE_STRICT_WERROR)
add_compile_options(-Werror)
if (NOT ENABLE_STRICT_WERROR)
@@ -681,6 +695,6 @@ if (BUILD_THUNKS)
add_dependencies(uninstall uninstall_guest-libs-32)
endif()
if (BUILD_STEAM_SUPPORT)
if (NOT MINGW AND BUILD_STEAM_SUPPORT)
add_subdirectory(Source/Steam/)
endif()
+2 -2
View File
@@ -311,7 +311,7 @@ class ExtendedMemOperand final {
public:
ExtendedMemOperand(XRegister rn, XRegister rm = XReg::zr, ExtendedType Option = ExtendedType::LSL_64, uint32_t Shift = 0)
: rn {rn}
, MetaType {.ExtendedType {
, MetaType {.Extended {
.Header = {.MemType = TYPE_EXTENDED},
.rm = rm,
.Option = Option,
@@ -340,7 +340,7 @@ public:
Register rm;
ExtendedType Option;
uint32_t Shift;
} ExtendedType;
} Extended;
struct {
HeaderStruct Header;
IndexType Index;
+50 -50
View File
@@ -3627,8 +3627,8 @@ public:
void strb(ARMEmitter::Register rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
strb(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
strb(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
strb(rt, MemSrc.rn);
} else {
@@ -3650,8 +3650,8 @@ public:
}
void ldrb(ARMEmitter::Register rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrb(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrb(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldrb(rt, MemSrc.rn);
} else {
@@ -3673,8 +3673,8 @@ public:
}
void ldrsb(ARMEmitter::XRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrsb(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrsb(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldrsb(rt, MemSrc.rn);
} else {
@@ -3696,8 +3696,8 @@ public:
}
void ldrsb(ARMEmitter::WRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrsb(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrsb(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldrsb(rt, MemSrc.rn);
} else {
@@ -3719,8 +3719,8 @@ public:
}
void strh(ARMEmitter::Register rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
strh(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
strh(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
strh(rt, MemSrc.rn);
} else {
@@ -3742,8 +3742,8 @@ public:
}
void ldrh(ARMEmitter::Register rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrh(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrh(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldrh(rt, MemSrc.rn);
} else {
@@ -3765,8 +3765,8 @@ public:
}
void ldrsh(ARMEmitter::XRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrsh(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrsh(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldrsh(rt, MemSrc.rn);
} else {
@@ -3788,8 +3788,8 @@ public:
}
void ldrsh(ARMEmitter::WRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrsh(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrsh(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldrsh(rt, MemSrc.rn);
} else {
@@ -3811,8 +3811,8 @@ public:
}
void str(ARMEmitter::WRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
str(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
str(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
str(rt, MemSrc.rn);
} else {
@@ -3834,8 +3834,8 @@ public:
}
void ldr(ARMEmitter::WRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldr(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldr(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldr(rt, MemSrc.rn);
} else {
@@ -3857,8 +3857,8 @@ public:
}
void ldrsw(ARMEmitter::XRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrsw(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrsw(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldrsw(rt, MemSrc.rn);
} else {
@@ -3880,8 +3880,8 @@ public:
}
void str(ARMEmitter::XRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
str(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
str(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
str(rt, MemSrc.rn);
} else {
@@ -3903,8 +3903,8 @@ public:
}
void ldr(ARMEmitter::XRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldr(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldr(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldr(rt, MemSrc.rn);
} else {
@@ -3926,8 +3926,8 @@ public:
}
void prfm(ARMEmitter::Prefetch prfop, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
prfm(prfop, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
prfm(prfop, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
prfm(prfop, MemSrc.rn);
} else {
@@ -3946,9 +3946,9 @@ public:
void strb(ARMEmitter::VRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
LOGMAN_THROW_A_FMT(MemSrc.MetaType.ExtendedType.Shift == false, "Can't shift byte");
strb(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
LOGMAN_THROW_A_FMT(MemSrc.MetaType.Extended.Shift == false, "Can't shift byte");
strb(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
strb(rt, MemSrc.rn);
} else {
@@ -3970,9 +3970,9 @@ public:
}
void ldrb(ARMEmitter::VRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
LOGMAN_THROW_A_FMT(MemSrc.MetaType.ExtendedType.Shift == false, "Can't shift byte");
ldrb(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
LOGMAN_THROW_A_FMT(MemSrc.MetaType.Extended.Shift == false, "Can't shift byte");
ldrb(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldrb(rt, MemSrc.rn);
} else {
@@ -3994,8 +3994,8 @@ public:
}
void strh(ARMEmitter::VRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
strh(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
strh(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
strh(rt, MemSrc.rn);
} else {
@@ -4017,8 +4017,8 @@ public:
}
void ldrh(ARMEmitter::VRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrh(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrh(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldrh(rt, MemSrc.rn);
} else {
@@ -4040,8 +4040,8 @@ public:
}
void str(ARMEmitter::SRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
str(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
str(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
str(rt, MemSrc.rn);
} else {
@@ -4063,8 +4063,8 @@ public:
}
void ldr(ARMEmitter::SRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldr(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldr(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldr(rt, MemSrc.rn);
} else {
@@ -4086,8 +4086,8 @@ public:
}
void str(ARMEmitter::DRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
str(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
str(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
str(rt, MemSrc.rn);
} else {
@@ -4109,8 +4109,8 @@ public:
}
void ldr(ARMEmitter::DRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldr(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldr(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldr(rt, MemSrc.rn);
} else {
@@ -4132,8 +4132,8 @@ public:
}
void str(ARMEmitter::QRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
str(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
str(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
str(rt, MemSrc.rn);
} else {
@@ -4155,8 +4155,8 @@ public:
}
void ldr(ARMEmitter::QRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldr(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldr(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldr(rt, MemSrc.rn);
} else {
+81 -30
View File
@@ -4,29 +4,34 @@
#
# pip-compile --generate-hashes --output-file=requirements_formatting.txt --strip-extras requirements_formatting.txt.in
#
black==25.1.0 \
--hash=sha256:030b9759066a4ee5e5aca28c3c77f9c64789cdd4de8ac1df642c40b708be6171 \
--hash=sha256:055e59b198df7ac0b7efca5ad7ff2516bca343276c466be72eb04a3bcc1f82d7 \
--hash=sha256:0e519ecf93120f34243e6b0054db49c00a35f84f195d5bce7e9f5cfc578fc2da \
--hash=sha256:172b1dbff09f86ce6f4eb8edf9dede08b1fce58ba194c87d7a4f1a5aa2f5b3c2 \
--hash=sha256:1e2978f6df243b155ef5fa7e558a43037c3079093ed5d10fd84c43900f2d8ecc \
--hash=sha256:33496d5cd1222ad73391352b4ae8da15253c5de89b93a80b3e2c8d9a19ec2666 \
--hash=sha256:3b48735872ec535027d979e8dcb20bf4f70b5ac75a8ea99f127c106a7d7aba9f \
--hash=sha256:4b60580e829091e6f9238c848ea6750efed72140b91b048770b64e74fe04908b \
--hash=sha256:759e7ec1e050a15f89b770cefbf91ebee8917aac5c20483bc2d80a6c3a04df32 \
--hash=sha256:8f0b18a02996a836cc9c9c78e5babec10930862827b1b724ddfe98ccf2f2fe4f \
--hash=sha256:95e8176dae143ba9097f351d174fdaf0ccd29efb414b362ae3fd72bf0f710717 \
--hash=sha256:96c1c7cd856bba8e20094e36e0f948718dc688dba4a9d78c3adde52b9e6c2299 \
--hash=sha256:a1ee0a0c330f7b5130ce0caed9936a904793576ef4d2b98c40835d6a65afa6a0 \
--hash=sha256:a22f402b410566e2d1c950708c77ebf5ebd5d0d88a6a2e87c86d9fb48afa0d18 \
--hash=sha256:a39337598244de4bae26475f77dda852ea00a93bd4c728e09eacd827ec929df0 \
--hash=sha256:afebb7098bfbc70037a053b91ae8437c3857482d3a690fefc03e9ff7aa9a5fd3 \
--hash=sha256:bacabb307dca5ebaf9c118d2d2f6903da0d62c9faa82bd21a33eecc319559355 \
--hash=sha256:bce2e264d59c91e52d8000d507eb20a9aca4a778731a08cfff7e5ac4a4bb7096 \
--hash=sha256:d9e6827d563a2c820772b32ce8a42828dc6790f095f441beef18f96aa6f8294e \
--hash=sha256:db8ea9917d6f8fc62abd90d944920d95e73c83a5ee3383493e35d271aca872e9 \
--hash=sha256:ea0213189960bda9cf99be5b8c8ce66bb054af5e9e861249cd23471bd7b0b3ba \
--hash=sha256:f3df5f1bf91d36002b0a75389ca8663510cf0531cca8aa5c1ef695b46d98655f
black==26.3.1 \
--hash=sha256:0126ae5b7c09957da2bdbd91a9ba1207453feada9e9fe51992848658c6c8e01c \
--hash=sha256:0f76ff19ec5297dd8e66eb64deda23631e642c9393ab592826fd4bdc97a4bce7 \
--hash=sha256:28ef38aee69e4b12fda8dba75e21f9b4f979b490c8ac0baa7cb505369ac9e1ff \
--hash=sha256:2bd5aa94fc267d38bb21a70d7410a89f1a1d318841855f698746f8e7f51acd1b \
--hash=sha256:2c50f5063a9641c7eed7795014ba37b0f5fa227f3d408b968936e24bc0566b07 \
--hash=sha256:2d6bfaf7fd0993b420bed691f20f9492d53ce9a2bcccea4b797d34e947318a78 \
--hash=sha256:41cd2012d35b47d589cb8a16faf8a32ef7a336f56356babd9fcf70939ad1897f \
--hash=sha256:474c27574d6d7037c1bc875a81d9be0a9a4f9ee95e62800dab3cfaadbf75acd5 \
--hash=sha256:5602bdb96d52d2d0672f24f6ffe5218795736dd34807fd0fd55ccd6bf206168b \
--hash=sha256:5e9d0d86df21f2e1677cc4bd090cd0e446278bcbbe49bf3659c308c3e402843e \
--hash=sha256:5ed0ca58586c8d9a487352a96b15272b7fa55d139fc8496b519e78023a8dab0a \
--hash=sha256:6c54a4a82e291a1fee5137371ab488866b7c86a3305af4026bdd4dc78642e1ac \
--hash=sha256:6e131579c243c98f35bce64a7e08e87fb2d610544754675d4a0e73a070a5aa3a \
--hash=sha256:855822d90f884905362f602880ed8b5df1b7e3ee7d0db2502d4388a954cc8c54 \
--hash=sha256:86a8b5035fce64f5dcd1b794cf8ec4d31fe458cf6ce3986a30deb434df82a1d2 \
--hash=sha256:8a33d657f3276328ce00e4d37fe70361e1ec7614da5d7b6e78de5426cb56332f \
--hash=sha256:92c0ec1f2cc149551a2b7b47efc32c866406b6891b0ee4625e95967c8f4acfb1 \
--hash=sha256:9a5e9f45e5d5e1c5b5c29b3bd4265dcc90e8b92cf4534520896ed77f791f4da5 \
--hash=sha256:afc622538b430aa4c8c853f7f63bc582b3b8030fd8c80b70fb5fa5b834e575c2 \
--hash=sha256:b07fc0dab849d24a80a29cfab8d8a19187d1c4685d8a5e6385a5ce323c1f015f \
--hash=sha256:b5e6f89631eb88a7302d416594a32faeee9fb8fb848290da9d0a5f2903519fc1 \
--hash=sha256:bf9bf162ed91a26f1adba8efda0b573bc6924ec1408a52cc6f82cb73ec2b142c \
--hash=sha256:c7e72339f841b5a237ff14f7d3880ddd0fc7f98a1199e8c4327f9a4f478c1839 \
--hash=sha256:ddb113db38838eb9f043623ba274cfaf7d51d5b0c22ecb30afe58b1bb8322983 \
--hash=sha256:dfdd51fc3e64ea4f35873d1b3fb25326773d55d2329ff8449139ebaad7357efb \
--hash=sha256:f1cd08e99d2f9317292a311dfe578fd2a24b15dbce97792f9c4d752275c1fa56 \
--hash=sha256:f89f2ab047c76a9c03f78d0d66ca519e389519902fa27e7a91117ef7611c0568
# via
# -r requirements_formatting.txt.in
# darker
@@ -290,9 +295,9 @@ packaging==23.1 \
--hash=sha256:994793af429502c4ea2ebf6bf664629d07c1a9fe974af92966e4b8d2df7edc61 \
--hash=sha256:a392980d2b6cffa644431898be54b0045151319d1e7ec34f0cfed48767dd334f
# via black
pathspec==0.11.2 \
--hash=sha256:1d6ed233af05e679efb96b1851550ea95bbb64b7c490b0f5aa52996c11e92a20 \
--hash=sha256:e0d8d0ac2f12da61956eb2306b69f9469b42f4deb0f3cb6ed47b9cce9996ced3
pathspec==1.0.4 \
--hash=sha256:0210e2ae8a21a9137c0d470578cb0e595af87edaa6ebf12ff176f14a02e0e645 \
--hash=sha256:fb6ae2fd4e7c921a165808a552060e722767cfa526f99ca5156ed2ce45a5c723
# via black
platformdirs==3.10.0 \
--hash=sha256:b45696dab2d7cc691a3226759c0d3b00c47c8b6e293d96f6436f733303f77f6d \
@@ -306,10 +311,12 @@ pygithub==2.6.1 \
--hash=sha256:6f2fa6d076ccae475f9fc392cc6cdbd54db985d4f69b8833a28397de75ed6ca3 \
--hash=sha256:b5c035392991cca63959e9453286b41b54d83bf2de2daa7d7ff7e4312cebf3bf
# via -r requirements_formatting.txt.in
pyjwt==2.8.0 \
--hash=sha256:57e28d156e3d5c10088e0c68abb90bfac3df82b40a71bd0daa20c65ccd5c23de \
--hash=sha256:59127c392cc44c2da5bb3192169a91f429924e17aff6534d70fdc02ab3e04320
# via pygithub
pyjwt==2.12.1 \
--hash=sha256:28ca37c070cad8ba8cd9790cd940535d40274d22f80ab87f3ac6a713e6e8454c \
--hash=sha256:c74a7a2adf861c04d002db713dd85f84beb242228e671280bf709d765b03672b
# via
# -r requirements_formatting.txt.in
# pygithub
pynacl==1.6.2 \
--hash=sha256:018494d6d696ae03c7e656e5e74cdfd8ea1326962cc401bcf018f1ed8436811c \
--hash=sha256:04316d1fc625d860b6c162fff704eb8426b1a8bcd3abacea11142cbd99a6b574 \
@@ -339,6 +346,50 @@ pynacl==1.6.2 \
# via
# -r requirements_formatting.txt.in
# pygithub
pytokens==0.4.1 \
--hash=sha256:0fc71786e629cef478cbf29d7ea1923299181d0699dbe7c3c0f4a583811d9fc1 \
--hash=sha256:11edda0942da80ff58c4408407616a310adecae1ddd22eef8c692fe266fa5009 \
--hash=sha256:140709331e846b728475786df8aeb27d24f48cbcf7bcd449f8de75cae7a45083 \
--hash=sha256:24afde1f53d95348b5a0eb19488661147285ca4dd7ed752bbc3e1c6242a304d1 \
--hash=sha256:26cef14744a8385f35d0e095dc8b3a7583f6c953c2e3d269c7f82484bf5ad2de \
--hash=sha256:27b83ad28825978742beef057bfe406ad6ed524b2d28c252c5de7b4a6dd48fa2 \
--hash=sha256:292052fe80923aae2260c073f822ceba21f3872ced9a68bb7953b348e561179a \
--hash=sha256:29d1d8fb1030af4d231789959f21821ab6325e463f0503a61d204343c9b355d1 \
--hash=sha256:2a44ed93ea23415c54f3face3b65ef2b844d96aeb3455b8a69b3df6beab6acc5 \
--hash=sha256:30f51edd9bb7f85c748979384165601d028b84f7bd13fe14d3e065304093916a \
--hash=sha256:34bcc734bd2f2d5fe3b34e7b3c0116bfb2397f2d9666139988e7a3eb5f7400e3 \
--hash=sha256:3ad72b851e781478366288743198101e5eb34a414f1d5627cdd585ca3b25f1db \
--hash=sha256:3f901fe783e06e48e8cbdc82d631fca8f118333798193e026a50ce1b3757ea68 \
--hash=sha256:42f144f3aafa5d92bad964d471a581651e28b24434d184871bd02e3a0d956037 \
--hash=sha256:4a14d5f5fc78ce85e426aa159489e2d5961acf0e47575e08f35584009178e321 \
--hash=sha256:4a58d057208cb9075c144950d789511220b07636dd2e4708d5645d24de666bdc \
--hash=sha256:4e691d7f5186bd2842c14813f79f8884bb03f5995f0575272009982c5ac6c0f7 \
--hash=sha256:5502408cab1cb18e128570f8d598981c68a50d0cbd7c61312a90507cd3a1276f \
--hash=sha256:584c80c24b078eec1e227079d56dc22ff755e0ba8654d8383b2c549107528918 \
--hash=sha256:5ad948d085ed6c16413eb5fec6b3e02fa00dc29a2534f088d3302c47eb59adf9 \
--hash=sha256:670d286910b531c7b7e3c0b453fd8156f250adb140146d234a82219459b9640c \
--hash=sha256:682fa37ff4d8e95f7df6fe6fe6a431e8ed8e788023c6bcc0f0880a12eab80ad1 \
--hash=sha256:6d6c4268598f762bc8e91f5dbf2ab2f61f7b95bdc07953b602db879b3c8c18e1 \
--hash=sha256:79fc6b8699564e1f9b521582c35435f1bd32dd06822322ec44afdeba666d8cb3 \
--hash=sha256:8bdb9d0ce90cbf99c525e75a2fa415144fd570a1ba987380190e8b786bc6ef9b \
--hash=sha256:8fcb9ba3709ff77e77f1c7022ff11d13553f3c30299a9fe246a166903e9091eb \
--hash=sha256:941d4343bf27b605e9213b26bfa1c4bf197c9c599a9627eb7305b0defcfe40c1 \
--hash=sha256:967cf6e3fd4adf7de8fc73cd3043754ae79c36475c1c11d514fc72cf5490094a \
--hash=sha256:970b08dd6b86058b6dc07efe9e98414f5102974716232d10f32ff39701e841c4 \
--hash=sha256:97f50fd18543be72da51dd505e2ed20d2228c74e0464e4262e4899797803d7fa \
--hash=sha256:9bd7d7f544d362576be74f9d5901a22f317efc20046efe2034dced238cbbfe78 \
--hash=sha256:add8bf86b71a5d9fb5b89f023a80b791e04fba57960aa790cc6125f7f1d39dfe \
--hash=sha256:b35d7e5ad269804f6697727702da3c517bb8a5228afa450ab0fa787732055fc9 \
--hash=sha256:b49750419d300e2b5a3813cf229d4e5a4c728dae470bcc89867a9ad6f25a722d \
--hash=sha256:d31b97b3de0f61571a124a00ffe9a81fb9939146c122c11060725bd5aea79975 \
--hash=sha256:d70e77c55ae8380c91c0c18dea05951482e263982911fc7410b1ffd1dadd3440 \
--hash=sha256:d9907d61f15bf7261d7e775bd5d7ee4d2930e04424bab1972591918497623a16 \
--hash=sha256:da5baeaf7116dced9c6bb76dc31ba04a2dc3695f3d9f74741d7910122b456edc \
--hash=sha256:dc74c035f9bfca0255c1af77ddd2d6ae8419012805453e4b0e7513e17904545d \
--hash=sha256:dcafc12c30dbaf1e2af0490978352e0c4041a7cde31f4f81435c2a5e8b9cabb6 \
--hash=sha256:ee44d0f85b803321710f9239f335aafe16553b39106384cef8e6de40cb4ef2f6 \
--hash=sha256:f66a6bbe741bd431f6d741e617e0f39ec7257ca1f89089593479347cc4d13324
# via black
requests==2.32.4 \
--hash=sha256:27babd3cda2a6d50b30443204ee89830707d396671944c998b5975b031ac2b2c \
--hash=sha256:27d0316682c8a29834d3264820024b62a36942083d52caf2f14c0591336d3422
+2 -1
View File
@@ -1,4 +1,4 @@
black~=25.1
black>=26.3.1
darker==2.1.1
PyGithub==2.6.1
cryptography>=46.0.5
@@ -7,3 +7,4 @@ requests>=2.32.4
idna>=3.7
certifi>=2024.7.4
PyNaCl>=1.6.2
PyJWT>=2.12.1
+7 -1
View File
@@ -6,7 +6,8 @@ set(FEXCORE_BASE_SRCS
Utils/FileLoading.cpp
Utils/ForcedAssert.cpp
Utils/LogManager.cpp
Utils/SpinWaitLock.cpp)
Utils/SpinWaitLock.cpp
Utils/WildcardMatcher.cpp)
if (NOT MINGW)
list(APPEND FEXCORE_BASE_SRCS
@@ -123,6 +124,11 @@ else()
endif()
endif()
if (CMAKE_CXX_COMPILER_ID STREQUAL "GNU")
# GCC requires libatomic to use 128-bit atomics
list(APPEND LIBS atomic)
endif()
# Generate config
configure_file(${CMAKE_CURRENT_SOURCE_DIR}/Interface/Config/Config.json.in
${CMAKE_BINARY_DIR}/generated/Config/Config.json)
+19
View File
@@ -260,6 +260,10 @@ struct FEX_PACKED X80SoftFloat {
if (lhs.Top.Exponent == 0x0 && lhs.Significand == 0x0) {
return lhs;
}
// Inf/NaN pass through unchanged in the significand slot.
if (lhs.Top.Exponent == 0x7FFF) {
return lhs;
}
X80SoftFloat Tmp = lhs;
Tmp.Top.Exponent = 0x3FFF;
Tmp.Top.Sign = lhs.Top.Sign;
@@ -288,6 +292,14 @@ struct FEX_PACKED X80SoftFloat {
X80SoftFloat Result(1, 0x7FFFUL, 0x8000'0000'0000'0000UL);
return Result;
}
// +/-Inf returns +Inf in the exponent slot; NaN propagates.
if (lhs.Top.Exponent == 0x7FFF) {
if ((lhs.Significand & 0x7FFFFFFFFFFFFFFFULL) == 0) {
X80SoftFloat Result(0, 0x7FFFUL, 0x8000'0000'0000'0000UL);
return Result;
}
return lhs;
}
int32_t TrueExp = lhs.Top.Exponent - ExponentBias;
return i32_to_extF80(TrueExp);
@@ -324,6 +336,13 @@ struct FEX_PACKED X80SoftFloat {
#else
extFloat80_t Zero {0, 0};
if (extF80_eq(state, lhs, Zero)) {
// FSCALE(0, +Inf) is 0 * Inf, which is invalid. FSCALE(0, anything
// else) is still 0.
if (rhs.Top.Exponent == 0x7FFF && rhs.Top.Sign == 0 && (rhs.Significand & 0x7FFFFFFFFFFFFFFFULL) == 0) {
state->exceptionFlags |= softfloat_flag_invalid;
X80SoftFloat QNaN(0, 0x7FFFUL, 0xC000000000000000ULL);
return QNaN;
}
return lhs;
}
X80SoftFloat Int = FRNDINT(state, rhs, softfloat_round_minMag);
+4
View File
@@ -16,7 +16,11 @@ struct VectorScalarF64Pair {
#ifdef ARCHITECTURE_arm64
// Can't use uint8x16_t directly from arm_neon.h here.
// Overrides softfloat-3e's defines which causes problems.
#ifdef __clang__
using VectorRegType = __attribute__((neon_vector_type(16))) uint8_t;
#else
using VectorRegType = __attribute__((vector_size(16))) uint8_t;
#endif
struct VectorRegPairType {
VectorRegType val[2];
};
@@ -75,7 +75,9 @@
"ENABLE3DNOW": "enable3dnow",
"DISABLE3DNOW": "disable3dnow",
"ENABLESSE4A": "enablesse4a",
"DISABLESSE4A": "disablesse4a"
"DISABLESSE4A": "disablesse4a",
"ENABLEMOPS": "enablemops",
"DISABLEMOPS": "disablemops"
},
"Desc": [
"Allows controlling of the CPU features in the JIT.",
@@ -99,7 +101,8 @@
"\t{enable,disable}preserveallabi: Will force enable or disable preserve_all abi even if the host doesn't support it",
"\t{enable,disable}wfxt: Will force enable or disable wfxt even if the host doesn't support it",
"\t{enable,disable}3dnow: Will force enable or disable 3DNow! even if the host doesn't support it",
"\t{enable,disable}sse4a: Will force enable or disable SSE4a even if the host doesn't support it"
"\t{enable,disable}sse4a: Will force enable or disable SSE4a even if the host doesn't support it",
"\t{enable,disable}mops: Will force enable or disable FEAT_MOPS even if the host doesn't support it"
]
},
"SmallTSCScale": {
@@ -192,7 +195,7 @@
},
"DisableL2Cache": {
"Type": "bool",
"Default": "false",
"Default": "true",
"Desc": [
"Disables FEXCore's JIT L2 cache lookup. Saving memory.",
"Can potentially introduce more stutters."
@@ -200,7 +203,7 @@
},
"DynamicL1Cache": {
"Type": "bool",
"Default": "false",
"Default": "true",
"Desc": [
"Switches FEXCore's JIT L1 cache to be dynamically sized. Saving memory.",
"Can potentially introduce more stutters."
+3 -2
View File
@@ -127,6 +127,7 @@ public:
void ExecuteThread(FEXCore::Core::InternalThreadState* Thread) override;
bool CheckIfBlockIsCacheable(FEXCore::Core::InternalThreadState&, uint64_t GuestRIP, uint64_t MaxInst) override;
void CompileRIP(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP) override;
void CompileRIPCount(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP, uint64_t MaxInst) override;
@@ -208,7 +209,7 @@ public:
void ClearCodeCache(FEXCore::Core::InternalThreadState* Thread, bool NewCodeBuffer = true) override;
void InvalidateCodeBuffersCodeRange(uint64_t Start, uint64_t Length) override;
void InvalidateThreadCachedCodeRange(FEXCore::Core::InternalThreadState* Thread, uint64_t Start, uint64_t Length) override;
FEXCore::ForkableSharedMutex& GetCodeInvalidationMutex() override {
FEXCore::Utils::WritePriorityMutex::Mutex& GetCodeInvalidationMutex() override {
return CodeInvalidationMutex;
}
@@ -262,7 +263,7 @@ public:
FEX_CONFIG_OPT(MonoHacks, MONOHACKS);
} Config;
FEXCore::ForkableSharedMutex CodeInvalidationMutex;
FEXCore::Utils::WritePriorityMutex::Mutex CodeInvalidationMutex {};
uint32_t StrictSplitLockMutex {};
@@ -448,8 +448,10 @@ void Arm64Emitter::LoadConstant(ARMEmitter::Size s, ARMEmitter::Register Reg, ui
return;
}
if ((Constant >> 32) == 0) {
if ((Constant >> 32) == 0 && !NOPPad) {
// If the upper 32-bits is all zero, we can now switch to a 32-bit move.
// NOTE: The NOP padding code does not appropriately adjust to this yet,
// so we skip this optimization in that case
s = ARMEmitter::Size::i32Bit;
Is64Bit = false;
Segments = std::min(Segments, 2);
@@ -677,7 +679,7 @@ void Arm64Emitter::FillSpecialRegs(ARMEmitter::Register TmpReg, ARMEmitter::Regi
}
}
void Arm64Emitter::SpillStaticRegs(ARMEmitter::Register TmpReg, bool FPRs, uint32_t GPRSpillMask, uint32_t FPRSpillMask, bool NZCV) {
void Arm64Emitter::SpillStaticRegs(ARMEmitter::Register TmpReg, SpillStaticRegOptions Options) {
#ifndef VIXL_SIMULATOR
if (EmitterCTX->HostFeatures.SupportsAFP) {
// Disable AFP features when spilling registers.
@@ -698,7 +700,7 @@ void Arm64Emitter::SpillStaticRegs(ARMEmitter::Register TmpReg, bool FPRs, uint3
}
#endif
if (NZCV) {
if (Options.NZCV) {
// Regardless of what GPRs/FPRs we're spilling, we need to spill NZCV since it
// is always static and almost certainly clobbered by the subsequent code.
//
@@ -710,25 +712,25 @@ void Arm64Emitter::SpillStaticRegs(ARMEmitter::Register TmpReg, bool FPRs, uint3
// PF/AF are special, remove them from the mask
uint32_t PFAFMask = ((1u << REG_PF.Idx()) | ((1u << REG_AF.Idx())));
unsigned PFAFSpillMask = GPRSpillMask & PFAFMask;
GPRSpillMask &= ~PFAFSpillMask;
unsigned PFAFSpillMask = Options.GPRSpillMask & PFAFMask;
Options.GPRSpillMask &= ~PFAFSpillMask;
str(REG_CALLRET_SP, STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.callret_sp));
for (size_t i = 0; i < StaticRegisters.size(); i += 2) {
auto Reg1 = StaticRegisters[i];
auto Reg2 = StaticRegisters[i + 1];
if (((1U << Reg1.Idx()) & GPRSpillMask) && ((1U << Reg2.Idx()) & GPRSpillMask)) {
stp<ARMEmitter::IndexType::OFFSET>(Reg1.X(), Reg2.X(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i]));
} else if (((1U << Reg1.Idx()) & GPRSpillMask)) {
str(Reg1.X(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i]));
} else if (((1U << Reg2.Idx()) & GPRSpillMask)) {
str(Reg2.X(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i + 1]));
if (((1U << Reg1.Idx()) & Options.GPRSpillMask) && ((1U << Reg2.Idx()) & Options.GPRSpillMask)) {
stp<ARMEmitter::IndexType::OFFSET>(Reg1.X(), Reg2.X(), STATE.R(), ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, State.gregs, i));
} else if (((1U << Reg1.Idx()) & Options.GPRSpillMask)) {
str(Reg1.X(), STATE.R(), ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, State.gregs, i));
} else if (((1U << Reg2.Idx()) & Options.GPRSpillMask)) {
str(Reg2.X(), STATE.R(), ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, State.gregs, i + 1));
}
}
// Now handle PF/AF
if (NZCV && PFAFSpillMask) {
if (Options.NZCV && PFAFSpillMask) {
auto PFOffset = offsetof(FEXCore::Core::CpuStateFrame, State.pf_raw);
auto AFOffset = offsetof(FEXCore::Core::CpuStateFrame, State.af_raw);
LOGMAN_THROW_A_FMT(PFAFSpillMask == PFAFMask, "PF/AF not spilled together");
@@ -737,21 +739,21 @@ void Arm64Emitter::SpillStaticRegs(ARMEmitter::Register TmpReg, bool FPRs, uint3
stp<ARMEmitter::IndexType::OFFSET>(REG_PF.W(), REG_AF.W(), STATE.R(), PFOffset);
}
if (FPRs) {
if (Options.FPRs) {
if (EmitterCTX->HostFeatures.SupportsAVX && EmitterCTX->HostFeatures.SupportsSVE256) {
for (size_t i = 0; i < StaticFPRegisters.size(); i++) {
const auto Reg = StaticFPRegisters[i];
if (((1U << Reg.Idx()) & FPRSpillMask) != 0) {
mov(ARMEmitter::Size::i64Bit, TmpReg, offsetof(Core::CpuStateFrame, State.xmm.avx.data[i][0]));
if (((1U << Reg.Idx()) & Options.FPRSpillMask) != 0) {
mov(ARMEmitter::Size::i64Bit, TmpReg, ARRAY_OFFSETOF(Core::CpuStateFrame, State.xmm.avx.data, i));
st1b<ARMEmitter::SubRegSize::i8Bit>(Reg.Z(), PRED_TMP_32B, STATE.R(), TmpReg);
}
}
} else {
if (GPRSpillMask && FPRSpillMask == ~0U) {
if (Options.GPRSpillMask && Options.FPRSpillMask == ~0U) {
// Optimize the common case where we can spill four registers per instruction
// Load the sse offset in to the temporary register
add(ARMEmitter::Size::i64Bit, TmpReg, STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[0][0]));
add(ARMEmitter::Size::i64Bit, TmpReg, STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data));
for (size_t i = 0; i < StaticFPRegisters.size(); i += 4) {
const auto Reg1 = StaticFPRegisters[i];
const auto Reg2 = StaticFPRegisters[i + 1];
@@ -764,12 +766,12 @@ void Arm64Emitter::SpillStaticRegs(ARMEmitter::Register TmpReg, bool FPRs, uint3
const auto Reg1 = StaticFPRegisters[i];
const auto Reg2 = StaticFPRegisters[i + 1];
if (((1U << Reg1.Idx()) & FPRSpillMask) && ((1U << Reg2.Idx()) & FPRSpillMask)) {
stp<ARMEmitter::IndexType::OFFSET>(Reg1.Q(), Reg2.Q(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[i][0]));
} else if (((1U << Reg1.Idx()) & FPRSpillMask)) {
str(Reg1.Q(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[i][0]));
} else if (((1U << Reg2.Idx()) & FPRSpillMask)) {
str(Reg2.Q(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[i + 1][0]));
if (((1U << Reg1.Idx()) & Options.FPRSpillMask) && ((1U << Reg2.Idx()) & Options.FPRSpillMask)) {
stp<ARMEmitter::IndexType::OFFSET>(Reg1.Q(), Reg2.Q(), STATE.R(), ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, State.xmm.sse.data, i));
} else if (((1U << Reg1.Idx()) & Options.FPRSpillMask)) {
str(Reg1.Q(), STATE.R(), ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, State.xmm.sse.data, i));
} else if (((1U << Reg2.Idx()) & Options.FPRSpillMask)) {
str(Reg2.Q(), STATE.R(), ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, State.xmm.sse.data, i + 1));
}
}
}
@@ -777,8 +779,7 @@ void Arm64Emitter::SpillStaticRegs(ARMEmitter::Register TmpReg, bool FPRs, uint3
}
}
void Arm64Emitter::FillStaticRegs(bool FPRs, uint32_t GPRFillMask, uint32_t FPRFillMask, std::optional<ARMEmitter::Register> OptionalReg,
std::optional<ARMEmitter::Register> OptionalReg2, bool NZCV) {
void Arm64Emitter::FillStaticRegs(FillStaticRegOptions Options) {
auto FindTempReg = [this](uint32_t* GPRFillMask) -> std::optional<ARMEmitter::Register> {
for (auto Reg : StaticRegisters) {
if (((1U << Reg.Idx()) & *GPRFillMask)) {
@@ -789,20 +790,21 @@ void Arm64Emitter::FillStaticRegs(bool FPRs, uint32_t GPRFillMask, uint32_t FPRF
return std::nullopt;
};
LOGMAN_THROW_A_FMT(GPRFillMask != 0, "Must fill at least 2 GPRs for a temp");
uint32_t TempGPRFillMask = GPRFillMask;
if (!OptionalReg.has_value()) {
OptionalReg = FindTempReg(&TempGPRFillMask);
LOGMAN_THROW_A_FMT(Options.GPRFillMask != 0, "Must fill at least 2 GPRs for a temp");
uint32_t TempGPRFillMask = Options.GPRFillMask;
if (!Options.OptionalReg.has_value()) {
Options.OptionalReg = FindTempReg(&TempGPRFillMask);
}
if (!OptionalReg2.has_value()) {
OptionalReg2 = FindTempReg(&TempGPRFillMask);
if (!Options.OptionalReg2.has_value()) {
Options.OptionalReg2 = FindTempReg(&TempGPRFillMask);
}
LOGMAN_THROW_A_FMT(OptionalReg.has_value() && OptionalReg2.has_value(), "Didn't have an SRA register to use as a temporary while "
"spilling!");
LOGMAN_THROW_A_FMT(Options.OptionalReg.has_value() && Options.OptionalReg2.has_value(), "Didn't have an SRA register to use as a "
"temporary while "
"spilling!");
auto TmpReg = *OptionalReg;
auto TmpReg2 = *OptionalReg2;
auto TmpReg = *Options.OptionalReg;
auto TmpReg2 = *Options.OptionalReg2;
#ifdef ARCHITECTURE_arm64ec
// Load STATE in from the CPU area as x28 is not callee saved in the ARM64EC ABI.
@@ -812,7 +814,7 @@ void Arm64Emitter::FillStaticRegs(bool FPRs, uint32_t GPRFillMask, uint32_t FPRF
ldr(REG_CALLRET_SP, STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.callret_sp));
if (NZCV) {
if (Options.NZCV) {
// Regardless of what GPRs/FPRs we're filling, we need to fill NZCV since it
// is always static and was almost certainly clobbered.
//
@@ -822,23 +824,23 @@ void Arm64Emitter::FillStaticRegs(bool FPRs, uint32_t GPRFillMask, uint32_t FPRF
msr(ARMEmitter::SystemRegister::NZCV, TmpReg);
}
FillSpecialRegs(TmpReg, TmpReg2, true, FPRs);
FillSpecialRegs(TmpReg, TmpReg2, true, Options.FPRs);
if (FPRs) {
if (Options.FPRs) {
if (EmitterCTX->HostFeatures.SupportsAVX && EmitterCTX->HostFeatures.SupportsSVE256) {
for (size_t i = 0; i < StaticFPRegisters.size(); i++) {
const auto Reg = StaticFPRegisters[i];
if (((1U << Reg.Idx()) & FPRFillMask) != 0) {
mov(ARMEmitter::Size::i64Bit, TmpReg, offsetof(Core::CpuStateFrame, State.xmm.avx.data[i][0]));
if (((1U << Reg.Idx()) & Options.FPRFillMask) != 0) {
mov(ARMEmitter::Size::i64Bit, TmpReg, ARRAY_OFFSETOF(Core::CpuStateFrame, State.xmm.avx.data, i));
ld1b<ARMEmitter::SubRegSize::i8Bit>(Reg.Z(), PRED_TMP_32B.Zeroing(), STATE.R(), TmpReg);
}
}
} else {
if (GPRFillMask && FPRFillMask == ~0U) {
if (Options.GPRFillMask && Options.FPRFillMask == ~0U) {
// Optimize the common case where we can fill four registers per instruction.
// Use one of the filling static registers before we fill it.
// Load the sse offset in to the temporary register
add(ARMEmitter::Size::i64Bit, TmpReg, STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[0][0]));
add(ARMEmitter::Size::i64Bit, TmpReg, STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data));
for (size_t i = 0; i < StaticFPRegisters.size(); i += 4) {
const auto Reg1 = StaticFPRegisters[i];
const auto Reg2 = StaticFPRegisters[i + 1];
@@ -851,12 +853,12 @@ void Arm64Emitter::FillStaticRegs(bool FPRs, uint32_t GPRFillMask, uint32_t FPRF
const auto Reg1 = StaticFPRegisters[i];
const auto Reg2 = StaticFPRegisters[i + 1];
if (((1U << Reg1.Idx()) & FPRFillMask) && ((1U << Reg2.Idx()) & FPRFillMask)) {
ldp<ARMEmitter::IndexType::OFFSET>(Reg1.Q(), Reg2.Q(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[i][0]));
} else if (((1U << Reg1.Idx()) & FPRFillMask)) {
ldr(Reg1.Q(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[i][0]));
} else if (((1U << Reg2.Idx()) & FPRFillMask)) {
ldr(Reg2.Q(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[i + 1][0]));
if (((1U << Reg1.Idx()) & Options.FPRFillMask) && ((1U << Reg2.Idx()) & Options.FPRFillMask)) {
ldp<ARMEmitter::IndexType::OFFSET>(Reg1.Q(), Reg2.Q(), STATE.R(), ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, State.xmm.sse.data, i));
} else if (((1U << Reg1.Idx()) & Options.FPRFillMask)) {
ldr(Reg1.Q(), STATE.R(), ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, State.xmm.sse.data, i));
} else if (((1U << Reg2.Idx()) & Options.FPRFillMask)) {
ldr(Reg2.Q(), STATE.R(), ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, State.xmm.sse.data, i + 1));
}
}
}
@@ -865,23 +867,23 @@ void Arm64Emitter::FillStaticRegs(bool FPRs, uint32_t GPRFillMask, uint32_t FPRF
// PF/AF are special, remove them from the mask
uint32_t PFAFMask = ((1u << REG_PF.Idx()) | ((1u << REG_AF.Idx())));
uint32_t PFAFFillMask = GPRFillMask & PFAFMask;
GPRFillMask &= ~PFAFMask;
uint32_t PFAFFillMask = Options.GPRFillMask & PFAFMask;
Options.GPRFillMask &= ~PFAFMask;
for (size_t i = 0; i < StaticRegisters.size(); i += 2) {
auto Reg1 = StaticRegisters[i];
auto Reg2 = StaticRegisters[i + 1];
if (((1U << Reg1.Idx()) & GPRFillMask) && ((1U << Reg2.Idx()) & GPRFillMask)) {
ldp<ARMEmitter::IndexType::OFFSET>(Reg1.X(), Reg2.X(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i]));
} else if ((1U << Reg1.Idx()) & GPRFillMask) {
ldr(Reg1.X(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i]));
} else if ((1U << Reg2.Idx()) & GPRFillMask) {
ldr(Reg2.X(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i + 1]));
if (((1U << Reg1.Idx()) & Options.GPRFillMask) && ((1U << Reg2.Idx()) & Options.GPRFillMask)) {
ldp<ARMEmitter::IndexType::OFFSET>(Reg1.X(), Reg2.X(), STATE.R(), ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, State.gregs, i));
} else if ((1U << Reg1.Idx()) & Options.GPRFillMask) {
ldr(Reg1.X(), STATE.R(), ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, State.gregs, i));
} else if ((1U << Reg2.Idx()) & Options.GPRFillMask) {
ldr(Reg2.X(), STATE.R(), ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, State.gregs, i + 1));
}
}
// Now handle PF/AF
if (NZCV && PFAFFillMask) {
if (Options.NZCV && PFAFFillMask) {
LOGMAN_THROW_A_FMT(PFAFFillMask == PFAFMask, "PF/AF not filled together");
ldp<ARMEmitter::IndexType::OFFSET>(REG_PF.W(), REG_AF.W(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.pf_raw));
@@ -1056,7 +1058,10 @@ size_t Arm64Emitter::SpillForPreserveAllABICall(ARMEmitter::Register TmpReg, boo
const uint64_t SPOffset = AlignUp(GPRSize + FPRSize, 16);
// Spill the static registers.
SpillStaticRegs(TmpReg, true, PreserveSRAMask, PreserveSRAFPRMask);
SpillStaticRegs(TmpReg, {
.GPRSpillMask = PreserveSRAMask,
.FPRSpillMask = PreserveSRAFPRMask,
});
sub(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, ARMEmitter::Reg::rsp, SPOffset);
@@ -1103,7 +1108,11 @@ void Arm64Emitter::FillForPreserveAllABICall(bool FPRs) {
}
// Fill the static registers.
FillStaticRegs(FPRs, PreserveSRAMask, PreserveSRAFPRMask);
FillStaticRegs({
.GPRFillMask = PreserveSRAMask,
.FPRFillMask = PreserveSRAFPRMask,
.FPRs = FPRs,
});
// Pop the vector registers.
PopVectorRegisters(CanUseSVE256, DynamicFPRs);
@@ -135,10 +135,35 @@ protected:
// Returning REG_INVALID if there was no mapping.
FEXCore::X86State::X86Reg GetX86RegRelationToARMReg(ARMEmitter::Register Reg);
void SpillStaticRegs(ARMEmitter::Register TmpReg, bool FPRs = true, uint32_t GPRSpillMask = ~0U, uint32_t FPRSpillMask = ~0U, bool NZCV = true);
void FillStaticRegs(bool FPRs = true, uint32_t GPRFillMask = ~0U, uint32_t FPRFillMask = ~0U,
std::optional<ARMEmitter::Register> OptionalReg = std::nullopt,
std::optional<ARMEmitter::Register> OptionalReg2 = std::nullopt, bool NZCV = true);
struct SpillStaticRegOptions final {
uint32_t GPRSpillMask {~0U};
uint32_t FPRSpillMask {~0U};
bool FPRs {true};
bool NZCV {true};
};
struct FillStaticRegOptions final {
std::optional<ARMEmitter::Register> OptionalReg {std::nullopt};
std::optional<ARMEmitter::Register> OptionalReg2 {std::nullopt};
uint32_t GPRFillMask {~0U};
uint32_t FPRFillMask {~0U};
bool FPRs {true};
bool NZCV {true};
};
void SpillStaticRegs(ARMEmitter::Register TmpReg, SpillStaticRegOptions Options);
void FillStaticRegs(FillStaticRegOptions Options);
void SpillStaticRegs(ARMEmitter::Register TmpReg) {
// Work around a clang bug: https://bugs.llvm.org/show_bug.cgi?id=36684
SpillStaticRegs(TmpReg, {});
}
void FillStaticRegs() {
// Work around a clang bug: https://bugs.llvm.org/show_bug.cgi?id=36684
FillStaticRegs({});
}
// Register 0-18 + 29 + 30 are caller saved
static constexpr uint32_t CALLER_GPR_MASK = 0b0110'0000'0000'0111'1111'1111'1111'1111U;
@@ -178,7 +203,9 @@ protected:
if (SupportsPreserveAllABI) {
return SpillForPreserveAllABICall(TmpReg, FPRs);
} else {
SpillStaticRegs(TmpReg, FPRs);
SpillStaticRegs(TmpReg, {
.FPRs = FPRs,
});
return PushDynamicRegs(TmpReg);
}
}
@@ -188,7 +215,7 @@ protected:
FillForPreserveAllABICall(FPRs);
} else {
PopDynamicRegs();
FillStaticRegs(FPRs);
FillStaticRegs({.FPRs = FPRs});
}
}
+3 -1
View File
@@ -12,7 +12,6 @@
#include <cstdint>
#ifndef _WIN32
#include <linux/prctl.h>
#include <sys/prctl.h>
#endif
@@ -363,6 +362,9 @@ namespace CPU {
FEXCore::Allocator::VirtualName("FEXMemJIT", reinterpret_cast<void*>(Ptr), Size);
// Huge-pages reduce the amount of iTLB misses dramatically when it works.
FEXCore::Allocator::VirtualTHPControl(reinterpret_cast<void*>(Ptr), Size, FEXCore::Allocator::THPControl::Enable);
LookupCache = fextl::make_unique<GuestToHostMap>();
}
+116 -4
View File
@@ -92,6 +92,7 @@ namespace ProductNames {
static const char ARM_AppleSilicon[] = "Apple Silicon";
static const char ARM_ORYON_1[] = "Oryon-1";
static const char ARM_ORYON_3[] = "Oryon-3";
static const char ARM_Ampere_1[] = "AmpereOne";
static const char ARM_Ampere_1A[] = "AmpereOneA";
static const char ARM_Ampere_1B[] = "AmpereOneB";
@@ -141,7 +142,7 @@ constexpr uint32_t FAMILY_IDENTIFIER = GenerateFamily(CPUFamily {
#endif
#ifdef ARCHITECTURE_arm64
uint32_t GetCycleCounterFrequency() {
uint64_t GetCycleCounterFrequency() {
uint64_t Result {};
__asm("mrs %[Res], CNTFRQ_EL0" : [Res] "=r"(Result));
return Result;
@@ -186,8 +187,9 @@ void CPUIDEmu::SetupHostHybridFlag() {
// CPU priority order
// This is mostly arbitrary but will sort by some sort of CPU priority by performance
// Relative list so things they will commonly end up in big.little configurations sort of relate
static constexpr std::array<CPUMIDR, 67> CPUMIDRs = {{
static constexpr std::array<CPUMIDR, 68> CPUMIDRs = {{
// Typically big CPU cores
{0x51, 0x002, 1, ProductNames::ARM_ORYON_3}, // Qualcomm Oryon-3
{0x51, 0x001, 1, ProductNames::ARM_ORYON_1}, // Qualcomm Oryon-1
{0x61, 0x039, 1, ProductNames::ARM_Avalanche_M2Max}, // Apple Avalanche (M2 Max)
@@ -408,7 +410,7 @@ void CPUIDEmu::SetupHostHybridFlag() {
}
#else
uint32_t GetCycleCounterFrequency() {
uint64_t GetCycleCounterFrequency() {
return 0;
}
@@ -756,6 +758,95 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_07h(uint32_t Leaf) const {
(0 << 29) | // Arch capabilities - Speculative side channel mitigations
(0 << 30) | // Arch capabilities - MSR module specific
(0 << 31); // SSBD - Speculative Store Bypass Disable
} else if (Leaf == 1) {
Res.eax = (0U << 0) | // SHA512
(0U << 1) | // SM3
(0U << 2) | // SM4
(0U << 3) | // RAO_INT
(0U << 4) | // AVX_VNNI
(0U << 5) | // AVX512_BF16
(0U << 6) | // LASS (Linear Address Space Separation)
(0U << 7) | // CMPCCXADD
(0U << 8) | // ARCH_PERFMON_EXT
(0U << 9) | // Reserved
(0U << 10) | // FAST_REP_MOVSB
(0U << 11) | // FAST_REP_STOSB
(0U << 12) | // FAST_REP_CMPSB_SCASB
(0U << 13) | // Reserved
(0U << 14) | // Reserved
(0U << 15) | // Reserved
(0U << 16) | // Reserved
(0U << 17) | // FRED (Flexible Return and Event Delivery)
(0U << 18) | // LKGS (Load into Kernel GS Base)
(0U << 19) | // WRMSRNS
(0U << 20) | // NMI_SRC
(0U << 21) | // AMX_FP16
(0U << 22) | // HRESET
(0U << 23) | // AVX_IFMA
(0U << 24) | // Reserved
(0U << 25) | // Reserved
(0U << 26) | // LAM (Linear Address Masking)
(0U << 27) | // MSRLIST
(0U << 28) | // Reserved
(0U << 29) | // Reserved
(0U << 30) | // INVD_DISABLE_POST_BIOS_DONE
(0U << 31); // MOVRS
// Bits 4-31 currently reserved.
Res.ebx = (0U << 0) | // PPIN
(0U << 1) | // PBNDKB
(0U << 2) | // Reserved
(0U << 3); // CPUIDMAXVAL_LIM_RMV
// Bits 6-31 also reserved.
Res.ecx = (0U << 0) | // RDT_M_ASYM
(0U << 1) | // RDT_A_ASYM
(0U << 2) | // Reserved
(0U << 3) | // Reserved
(0U << 4) | // Reserved
(0U << 5); // MSR_IMM
// Bits 25-31 also reserved.
Res.edx = (0U << 0) | // Reserved
(0U << 1) | // Reserved
(0U << 2) | // Reserved
(0U << 3) | // Reserved
(0U << 4) | // AVX_VNNI_INT8
(0U << 5) | // AVX_NE_CONVERT
(0U << 6) | // Reserved
(0U << 7) | // Reserved
(0U << 8) | // AMX_COMPLEX
(0U << 9) | // Reserved
(0U << 10) | // AVX_VNNI_INT16
(0U << 11) | // Reserved
(0U << 12) | // Reserved
(0U << 13) | // UTMR (User-timer events)
(0U << 14) | // PREFETCHI
(0U << 15) | // USER_MSR
(0U << 16) | // Reserved
(0U << 17) | // UIRET_UIF
(0U << 18) | // CET_SSS
(0U << 19) | // AVX10
(0U << 20) | // Reserved
(0U << 21) | // APX_F
(0U << 22) | // SEC-TEE_ATTESTATION
(0U << 23) | // MWAIT
(0U << 24); // SLSM (Static LSM)
} else if (Leaf == 2) {
// All bits are reserved except for EDX
Res.eax = 0;
Res.ebx = 0;
Res.ecx = 0;
// Bits 8-31 are reserved.
Res.edx = (0U << 0) | // PSFD
(0U << 1) | // IPRED_CTRL
(0U << 2) | // RRSBA_CTRL
(0U << 3) | // DDPD_U
(0U << 4) | // BHI_CTRL
(0U << 5) | // MCDT_NO
(0U << 6) | // UC_LOCK_DISABLE
(0U << 7); // MONITOR_MITG_NO
}
return Res;
@@ -813,7 +904,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_0Dh(uint32_t Leaf) const {
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_15h(uint32_t Leaf) const {
FEXCore::CPUID::FunctionResults Res {};
// TSC frequency = ECX * EBX / EAX
uint32_t FrequencyHz = GetCycleCounterFrequency();
uint64_t FrequencyHz = GetCycleCounterFrequency();
if (FrequencyHz) {
Res.eax = 1;
Res.ebx = 1U << CTX->Config.TSCScale;
@@ -834,6 +925,27 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_1Ah(uint32_t Leaf) const {
return Res;
}
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_24h(uint32_t Leaf) const {
FEXCore::CPUID::FunctionResults Res {};
if (Leaf == 0) {
// EAX indicates the maximum number of subleaves.
Res.eax = 0;
// Bits 19-31 reserved
// NOTE: We return all zero here until we have a CPU with AVX10
// even if some of the fields otherwise have fixed values.
Res.ebx = (0U << 0) | // (bits 0-7 specify the vector ISA version)
(0U << 16); // Defined as always 0b111
// All bits reserved
Res.ecx = 0;
Res.edx = 0;
}
return Res;
}
// Hypervisor CPUID information leaf
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_4000_0000h(uint32_t Leaf) const {
FEXCore::CPUID::FunctionResults Res {};
+84 -2
View File
@@ -14,7 +14,7 @@ namespace Context {
class ContextImpl;
}
uint32_t GetCycleCounterFrequency();
uint64_t GetCycleCounterFrequency();
// Debugging define to switch what family of CPU we execute as.
// Might be useful if an application makes an assumption about a CPU.
@@ -176,6 +176,7 @@ private:
FEXCore::CPUID::FunctionResults Function_0Dh(uint32_t Leaf) const;
FEXCore::CPUID::FunctionResults Function_15h(uint32_t Leaf) const;
FEXCore::CPUID::FunctionResults Function_1Ah(uint32_t Leaf) const;
FEXCore::CPUID::FunctionResults Function_24h(uint32_t Leaf) const;
FEXCore::CPUID::FunctionResults Function_4000_0000h(uint32_t Leaf) const;
FEXCore::CPUID::FunctionResults Function_4000_0001h(uint32_t Leaf) const;
FEXCore::CPUID::FunctionResults Function_8000_0000h(uint32_t Leaf) const;
@@ -200,7 +201,7 @@ private:
void SetupHostHybridFlag();
void SetupFeatures();
static constexpr size_t PRIMARY_FUNCTION_COUNT = 27;
static constexpr size_t PRIMARY_FUNCTION_COUNT = 37;
static constexpr size_t HYPERVISOR_FUNCTION_COUNT = 2;
static constexpr size_t EXTENDED_FUNCTION_COUNT = 32;
static constexpr std::array<FunctionHandler, PRIMARY_FUNCTION_COUNT> Primary = {
@@ -268,7 +269,48 @@ private:
#ifndef CPUID_AMD
// 0x1A: Hybrid Information Sub-leaf
&CPUIDEmu::Function_1Ah,
// 0x1B: PCONFIG info
&CPUIDEmu::Function_Reserved,
// 0x1C: Last Branch Records (LBR) info
&CPUIDEmu::Function_Reserved,
// 0x1D: Tile info
&CPUIDEmu::Function_Reserved,
// 0x1E: TMUL info
&CPUIDEmu::Function_Reserved,
// 0x1F: V2 Extended topology
&CPUIDEmu::Function_Reserved,
// 0x20: Processor History Reset info
&CPUIDEmu::Function_Reserved,
// 0x21: Unimplemented
&CPUIDEmu::Function_Reserved,
// 0x22: Reserved
&CPUIDEmu::Function_Reserved,
// 0x23: Architectural Performance Monitoring Extended
&CPUIDEmu::Function_Reserved,
// 0x24: Converged Vector ISA
&CPUIDEmu::Function_24h,
#else
// 0x1A: Reserved
&CPUIDEmu::Function_Reserved,
// 0x1B: Reserved
&CPUIDEmu::Function_Reserved,
// 0x1C: Reserved
&CPUIDEmu::Function_Reserved,
// 0x1D: Reserved
&CPUIDEmu::Function_Reserved,
// 0x1E: Reserved
&CPUIDEmu::Function_Reserved,
// 0x1F: Reserved
&CPUIDEmu::Function_Reserved,
// 0x20: Reserved
&CPUIDEmu::Function_Reserved,
// 0x21: Reserved
&CPUIDEmu::Function_Reserved,
// 0x22: Reserved
&CPUIDEmu::Function_Reserved,
// 0x23: Reserved
&CPUIDEmu::Function_Reserved,
// 0x24: Reserved
&CPUIDEmu::Function_Reserved,
#endif
};
@@ -340,9 +382,49 @@ private:
#ifndef CPUID_AMD
// 0x1A: Hybrid Information Sub-leaf
{SupportsConstant::NONCONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x1B: PCONFIG info
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x1C: Last Branch Records (LBR) info
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x1D: Tile info
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x1E: TMUL info
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x1F: V2 Extended topology
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x20: Processor History Reset info
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x21: Unimplemented/Reserved
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x22: Reserved
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x23: Architectural Performance Monitoring Extended
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x24: Converged Vector ISA
{SupportsConstant::CONSTANT, NeedsLeafConstant::NEEDSLEAFCONSTANT},
#else
// 0x1A: Reserved
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x1B: Reserved
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x1C: Reserved
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x1D: Reserved
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x1E: Reserved
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x1F: Reserved
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x20: Reserved
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x21: Reserved
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x22: Reserved
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x23: Reserved
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x24: Reserved
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
#endif
}};
+8 -21
View File
@@ -1,5 +1,5 @@
// SPDX-License-Identifier: MIT
#include "Utils/SpinWaitLock.h"
#include <FEXCore/Utils/SpinWaitLock.h>
#include <Interface/Context/Context.h>
#include <Interface/Core/ArchHelpers/Arm64Emitter.h>
@@ -300,12 +300,10 @@ bool CodeCache::SaveData(Core::InternalThreadState& Thread, int fd, const Execut
::write(fd, Relocations.data(), Relocations.size() * sizeof(Relocations[0]));
// Pad to next page in file so that the CodeBuffer can be mmap'ed into process on load
char Zero[64] {};
auto Off = lseek(fd, 0, SEEK_CUR);
while (Off != AlignUp(Off, Utils::FEX_PAGE_SIZE)) {
auto BytesToWrite = std::min(AlignUp(Off, Utils::FEX_PAGE_SIZE) - Off, sizeof(Zero));
::write(fd, Zero, BytesToWrite);
Off += BytesToWrite;
{
auto AlignedSize = AlignUp(lseek(fd, 0, SEEK_CUR), Utils::FEX_PAGE_SIZE);
::ftruncate(fd, AlignedSize);
lseek(fd, AlignedSize, SEEK_SET);
}
// Dump the host code (relocated for position-independent serialization)
@@ -383,29 +381,18 @@ bool CodeCache::LoadData(Core::InternalThreadState* Thread, std::byte* MappedCac
MappedCacheFile += std::span {BlockPtr.second.CodePages}.size_bytes();
}
// Consistency check: VMA regions at the top and end should belong to the same file
auto [min_val, max_val] = ranges::minmax_element(BlockList, std::less {}, &decltype(BlockList)::value_type::first);
auto MinBound = CTX.SyscallHandler->LookupExecutableFileSection(Thread, min_val->first + BinarySection.FileStartVA);
auto MaxBound = CTX.SyscallHandler->LookupExecutableFileSection(Thread, max_val->first + BinarySection.FileStartVA);
if (&MinBound->FileInfo != &BinarySection.FileInfo || &MaxBound->FileInfo != &BinarySection.FileInfo) {
ERROR_AND_DIE_FMT("Cached blocks offsets {:#x}-{:#x} out of bounds for guest library {} ({:016x} @ {:#x}) while trying to load "
"section {:#x}-{:#x}!",
min_val->first, max_val->first, BinarySection.FileInfo.Filename, BinarySection.FileInfo.FileId,
BinarySection.FileStartVA, BinarySection.BeginVA, BinarySection.EndVA);
}
// Constrain BlockList to the given ExecutableFileSectionInfo
LOGMAN_THROW_A_FMT(ranges::is_sorted(BlockList, [](auto& a, auto& b) { return a.first < b.first; }), "Expected sorted block list");
auto begin = ranges::lower_bound(BlockList, BinarySection.BeginVA - BinarySection.FileStartVA, std::less {}, &BlockListEntry::first);
auto end =
ranges::upper_bound(begin, BlockList.end(), BinarySection.EndVA - BinarySection.FileStartVA - 1, std::less {}, &BlockListEntry::first);
BlockList.erase(end, BlockList.end());
BlockList.erase(BlockList.begin(), begin);
if (BlockList.empty()) {
if (begin == end) {
// Not an error since there is just no data to load
LogMan::Msg::IFmt("No blocks cached in this range, aborting");
return true;
}
BlockList.erase(end, BlockList.end());
BlockList.erase(BlockList.begin(), begin);
}
// Read relocations
+9 -4
View File
@@ -30,7 +30,7 @@ $end_info$
#include "Interface/IR/RegisterAllocationData.h"
#include "Utils/Allocator.h"
#include "Utils/Allocator/HostAllocator.h"
#include "Utils/SpinWaitLock.h"
#include <FEXCore/Utils/SpinWaitLock.h>
#include "Utils/variable_length_integer.h"
#include <FEXCore/Config/Config.h>
@@ -502,11 +502,14 @@ static void IRDumper(FEXCore::Core::InternalThreadState* Thread, IR::IREmitter*
fextl::fmt::print(FD, "IR-ShouldDump-{} 0x{:x}:\n{}\n@@@@@\n", NewIR.PostRA() ? "post" : "pre", GuestRIP, out.str());
};
bool ContextImpl::CheckIfBlockIsCacheable(FEXCore::Core::InternalThreadState& Thread, uint64_t GuestRIP, uint64_t MaxInst) {
return Thread.FrontendDecoder->CheckIfCacheable(Thread, reinterpret_cast<const uint8_t*>(GuestRIP), GuestRIP, MaxInst);
}
ContextImpl::GenerateIRResult
ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP, bool ExtendedDebugInfo, uint64_t MaxInst) {
FEXCORE_PROFILE_SCOPED("GenerateIR");
Thread->OpDispatcher->ReownOrClaimBuffer();
Thread->OpDispatcher->ResetWorkingList();
uint64_t TotalInstructions {0};
@@ -706,8 +709,8 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
// If we had a dispatch error then leave early
if (HadDispatchError && TotalInstructions == 0) {
// Couldn't handle any instruction in op dispatcher
Thread->OpDispatcher->ResetWorkingList();
return {{}, 0, 0, 0, 0};
Thread->OpDispatcher->DelayedDisownBuffer();
return {std::nullopt, 0, 0, 0, 0};
}
if (NeedsBlockEnd) {
@@ -774,6 +777,7 @@ ContextImpl::CompileCodeResult ContextImpl::CompileCode(FEXCore::Core::InternalT
auto [IRView, TotalInstructions, TotalInstructionsLength, StartAddr, Length, NeedsAddGuestCodeRanges] =
GenerateIR(Thread, GuestRIP, Config.GDBSymbols(), MaxInst);
if (!IRView) {
// OpDispatcher IR already released in this case.
return {{}, nullptr, 0, 0, false};
}
@@ -784,6 +788,7 @@ ContextImpl::CompileCodeResult ContextImpl::CompileCode(FEXCore::Core::InternalT
// as expensive and are easily reverted.
if (MaxInst != 1) {
if (auto Block = Thread->LookupCache->FindBlock(Thread, GuestRIP)) {
// Raced to compile, release the OpDispatcher IR.
Thread->OpDispatcher->DelayedDisownBuffer();
return {.CompiledCode = {.BlockBegin = reinterpret_cast<uint8_t*>(Block), .EntryPoints = {{GuestRIP, reinterpret_cast<uint8_t*>(Block)}}},
.DebugData = nullptr,
@@ -546,6 +546,14 @@ void Dispatcher::EmitDispatcher() {
LUDIVHandlerAddress = EmitLongALUOpHandler(STATE_PTR(CpuStateFrame, Pointers.LUDIV));
LDIVHandlerAddress = EmitLongALUOpHandler(STATE_PTR(CpuStateFrame, Pointers.LDIV));
EmitF64Sin();
EmitF64Cos();
EmitF64Tan();
EmitF64F2XM1();
EmitF64Scale();
EmitF64Atan();
EmitF64FYL2X();
// Interpreter fallbacks
{
constexpr static std::array<FallbackABI, FABI_UNKNOWN> ABIS {{
@@ -845,6 +853,892 @@ void Dispatcher::EmitF64ToExtF80() {
(void)Bind(&Done);
}
void Dispatcher::EmitF64Sin() {
F64SinHandlerAddress = GetCursorAddress<uint64_t>();
constexpr auto V2 = ARMEmitter::VReg::v2;
constexpr auto V3 = ARMEmitter::VReg::v3;
constexpr auto V4 = ARMEmitter::VReg::v4;
constexpr auto V5 = ARMEmitter::VReg::v5;
ARMEmitter::ForwardLabel Fallback, NonZero;
ARMEmitter::ForwardLabel InvPiPi1Label, Pi23Label;
ARMEmitter::ForwardLabel C0Label, C1Label, C2Label, C3Label, C4Label, C5Label, C6Label;
ARMEmitter::ForwardLabel RangeLabel;
// sin(+/-0) = +/-0
fmov(ARMEmitter::Size::i64Bit, TMP1, VTMP1.D());
lsl(ARMEmitter::Size::i64Bit, TMP1, TMP1, 1);
(void)cbnz(ARMEmitter::Size::i64Bit, TMP1, &NonZero);
ret();
(void)Bind(&NonZero);
// Save q2-q5.
stp<ARMEmitter::IndexType::PRE>(ARMEmitter::QReg::q2, ARMEmitter::QReg::q3, ARMEmitter::Reg::rsp, -64);
stp<ARMEmitter::IndexType::OFFSET>(ARMEmitter::QReg::q4, ARMEmitter::QReg::q5, ARMEmitter::Reg::rsp, 32);
// save nzcv
mrs(TMP1, ARMEmitter::SystemRegister::NZCV);
str(TMP1.W(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.flags[24]));
// Range check: fall back for |x| >= 2^23, NaN, and inf.
fabs(VTMP2.D(), VTMP1.D());
ldr(V2.D(), &RangeLabel);
fcmp(VTMP2.D(), V2.D());
(void)b(ARMEmitter::Condition::CC_HS, &Fallback);
// n = rint(x/pi).
ldr(V2.Q(), &InvPiPi1Label); // q2 = {inv_pi, pi_1}
fmul(VTMP2.D(), VTMP1.D(), V2.D());
frinta(VTMP2.D(), VTMP2.D());
// odd = (int(n) & 1) << 63.
fcvtzs(ARMEmitter::Size::i64Bit, TMP1, VTMP2.D());
lsl(ARMEmitter::Size::i64Bit, TMP1, TMP1, 63);
// r = x - n*pi (range reduction) via .2D lane-indexed FMLS.
ldr(V3.Q(), &Pi23Label); // q3 = {pi_2, pi_3}
fmov(V4.D(), VTMP1.D()); // r = x
fmls(ARMEmitter::SubRegSize::i64Bit, V4.Q(), VTMP2.Q(), V2.Q(), 1); // r -= n * pi_1
fmls(ARMEmitter::SubRegSize::i64Bit, V4.Q(), VTMP2.Q(), V3.Q(), 0); // r -= n * pi_2
fmls(ARMEmitter::SubRegSize::i64Bit, V4.Q(), VTMP2.Q(), V3.Q(), 1); // r -= n * pi_3
// r^2, r^4.
fmul(V5.D(), V4.D(), V4.D());
fmov(ARMEmitter::Size::i64Bit, TMP2, V4.D());
fmul(V3.D(), V5.D(), V5.D());
// Estrin polynomial: p = c0 + r2*c1 + r4*(c2 + r2*c3) + r8*(c4 + r2*c5 + r4*c6).
// Level 1 (independent FMAs).
ldr(VTMP1.D(), &C0Label);
ldr(VTMP2.D(), &C1Label);
fmadd(VTMP1.D(), V5.D(), VTMP2.D(), VTMP1.D()); // p01 = c0 + r2*c1
ldr(VTMP2.D(), &C2Label);
ldr(V2.D(), &C3Label);
fmadd(VTMP2.D(), V5.D(), V2.D(), VTMP2.D()); // p23 = c2 + r2*c3
ldr(V2.D(), &C4Label);
ldr(V4.D(), &C5Label);
fmadd(V2.D(), V5.D(), V4.D(), V2.D()); // p45 = c4 + r2*c5
// Level 2 (serial).
ldr(V4.D(), &C6Label);
fmadd(V2.D(), V3.D(), V4.D(), V2.D()); // p46 = p45 + r4*c6
fmadd(VTMP2.D(), V3.D(), V2.D(), VTMP2.D()); // p26 = p23 + r4*p46
fmadd(VTMP1.D(), V3.D(), VTMP2.D(), VTMP1.D()); // p06 = p01 + r4*p26
// y = r + r^3 * p06.
fmov(ARMEmitter::Size::i64Bit, V4.D(), TMP2);
fmul(V5.D(), V5.D(), V4.D());
fmadd(VTMP1.D(), V5.D(), VTMP1.D(), V4.D());
// result = y XOR odd.
fmov(ARMEmitter::Size::i64Bit, TMP2, VTMP1.D());
eor(ARMEmitter::Size::i64Bit, TMP2, TMP2, TMP1);
fmov(ARMEmitter::Size::i64Bit, VTMP1.D(), TMP2);
// restore nzcv
ldr(TMP1.W(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.flags[24]));
msr(ARMEmitter::SystemRegister::NZCV, TMP1);
// Restore q2-q5 and return.
ldp<ARMEmitter::IndexType::OFFSET>(ARMEmitter::QReg::q4, ARMEmitter::QReg::q5, ARMEmitter::Reg::rsp, 32);
ldp<ARMEmitter::IndexType::POST>(ARMEmitter::QReg::q2, ARMEmitter::QReg::q3, ARMEmitter::Reg::rsp, 64);
ret();
// Fallback path.
(void)Bind(&Fallback);
ldp<ARMEmitter::IndexType::OFFSET>(ARMEmitter::QReg::q4, ARMEmitter::QReg::q5, ARMEmitter::Reg::rsp, 32);
ldp<ARMEmitter::IndexType::POST>(ARMEmitter::QReg::q2, ARMEmitter::QReg::q3, ARMEmitter::Reg::rsp, 64);
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[FEXCore::Core::OPINDEX_F64SIN].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[FEXCore::Core::OPINDEX_F64SIN].Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
ret();
// Constant pool.
Align(16);
(void)Bind(&InvPiPi1Label);
dc64(0x3FD4'5F30'6DC9'C883ULL); // inv_pi
dc64(0x4009'21FB'5444'2D18ULL); // pi_1
(void)Bind(&Pi23Label);
dc64(0x3CA1'A626'3314'5C06ULL); // pi_2
dc64(0x395C'1CD1'2902'4E09ULL); // pi_3
(void)Bind(&C0Label);
dc64(0xBFC5'5555'5555'547BULL); // c0
(void)Bind(&C1Label);
dc64(0x3F81'1111'1110'8A4DULL); // c1
(void)Bind(&C2Label);
dc64(0xBF2A'01A0'1993'6F27ULL); // c2
(void)Bind(&C3Label);
dc64(0x3EC7'1DE3'7A97'D93EULL); // c3
(void)Bind(&C4Label);
dc64(0xBE5A'E633'9199'87C6ULL); // c4
(void)Bind(&C5Label);
dc64(0x3DE6'0E27'7AE0'7CECULL); // c5
(void)Bind(&C6Label);
dc64(0xBD69'E954'0300'A100ULL); // c6
(void)Bind(&RangeLabel);
dc64(0x4160'0000'0000'0000ULL); // 2^23
}
void Dispatcher::EmitF64Cos() {
F64CosHandlerAddress = GetCursorAddress<uint64_t>();
constexpr auto Accum = ARMEmitter::VReg::v2;
ARMEmitter::ForwardLabel Fallback;
ARMEmitter::ForwardLabel RangeLabel, InvPiLabel;
ARMEmitter::ForwardLabel Pi1Label, Pi2Label, Pi3Label;
ARMEmitter::ForwardLabel C0Label, C1Label, C2Label, C3Label, C4Label, C5Label, C6Label;
// Save q2 for use as accumulator
str<ARMEmitter::IndexType::PRE>(ARMEmitter::QReg::q2, ARMEmitter::Reg::rsp, -16);
// save nzcv
mrs(TMP1, ARMEmitter::SystemRegister::NZCV);
str(TMP1.W(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.flags[24]));
// Range check: fall back for |x| >= 2^23, NaN, and inf.
fabs(VTMP2.D(), VTMP1.D());
ldr(Accum.D(), &RangeLabel);
fcmp(VTMP2.D(), Accum.D());
(void)b(ARMEmitter::Condition::CC_HS, &Fallback);
// n = rint(x * (1/pi) + 0.5).
ldr(Accum.D(), &InvPiLabel);
fmov(ARMEmitter::ScalarRegSize::i64Bit, VTMP2, 0.5f);
fmadd(VTMP2.D(), VTMP1.D(), Accum.D(), VTMP2.D());
frinta(VTMP2.D(), VTMP2.D());
// odd = (int(n) & 1) << 63.
fcvtzs(ARMEmitter::Size::i64Bit, TMP1, VTMP2.D());
lsl(ARMEmitter::Size::i64Bit, TMP1, TMP1, 63);
// Save input to Accum before overwriting VTMP1.
fmov(Accum.D(), VTMP1.D());
// n = n - 0.5.
fmov(ARMEmitter::ScalarRegSize::i64Bit, VTMP1, 0.5f);
fsub(VTMP2.D(), VTMP2.D(), VTMP1.D());
// r = x - n*pi (range reduction), in extended precision.
ldr(VTMP1.D(), &Pi1Label);
fmsub(Accum.D(), VTMP2.D(), VTMP1.D(), Accum.D());
ldr(VTMP1.D(), &Pi2Label);
fmsub(Accum.D(), VTMP2.D(), VTMP1.D(), Accum.D());
ldr(VTMP1.D(), &Pi3Label);
fmsub(Accum.D(), VTMP2.D(), VTMP1.D(), Accum.D());
// sin(r) poly approx.
fmul(VTMP1.D(), Accum.D(), Accum.D());
fmov(ARMEmitter::Size::i64Bit, TMP2, Accum.D());
// Horner: p = c6 + r2*(c5 + r2*(... + r2*c0)).
ldr(VTMP2.D(), &C6Label);
ldr(Accum.D(), &C5Label);
fmadd(VTMP2.D(), VTMP1.D(), VTMP2.D(), Accum.D());
ldr(Accum.D(), &C4Label);
fmadd(VTMP2.D(), VTMP1.D(), VTMP2.D(), Accum.D());
ldr(Accum.D(), &C3Label);
fmadd(VTMP2.D(), VTMP1.D(), VTMP2.D(), Accum.D());
ldr(Accum.D(), &C2Label);
fmadd(VTMP2.D(), VTMP1.D(), VTMP2.D(), Accum.D());
ldr(Accum.D(), &C1Label);
fmadd(VTMP2.D(), VTMP1.D(), VTMP2.D(), Accum.D());
ldr(Accum.D(), &C0Label);
fmadd(VTMP2.D(), VTMP1.D(), VTMP2.D(), Accum.D());
// y = r + r^3 * p.
fmov(ARMEmitter::Size::i64Bit, Accum.D(), TMP2);
fmul(VTMP1.D(), VTMP1.D(), Accum.D());
fmadd(Accum.D(), VTMP1.D(), VTMP2.D(), Accum.D());
// result = y XOR odd.
fmov(ARMEmitter::Size::i64Bit, TMP2, Accum.D());
eor(ARMEmitter::Size::i64Bit, TMP2, TMP2, TMP1);
fmov(ARMEmitter::Size::i64Bit, VTMP1.D(), TMP2);
// restore nzcv
ldr(TMP1.W(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.flags[24]));
msr(ARMEmitter::SystemRegister::NZCV, TMP1);
// Restore q2 and return.
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::QReg::q2, ARMEmitter::Reg::rsp, 16);
ret();
// Fallback path.
(void)Bind(&Fallback);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::QReg::q2, ARMEmitter::Reg::rsp, 16);
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[FEXCore::Core::OPINDEX_F64COS].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[FEXCore::Core::OPINDEX_F64COS].Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
ret();
// Constant pool.
Align(16);
(void)Bind(&InvPiLabel);
dc64(0x3FD4'5F30'6DC9'C883ULL); // inv_pi
(void)Bind(&Pi1Label);
dc64(0x4009'21FB'5444'2D18ULL); // pi_1
(void)Bind(&Pi2Label);
dc64(0x3CA1'A626'3314'5C06ULL); // pi_2
(void)Bind(&Pi3Label);
dc64(0x395C'1CD1'2902'4E09ULL); // pi_3
(void)Bind(&C0Label);
dc64(0xBFC5'5555'5555'547BULL); // c0
(void)Bind(&C1Label);
dc64(0x3F81'1111'1110'8A4DULL); // c1
(void)Bind(&C2Label);
dc64(0xBF2A'01A0'1993'6F27ULL); // c2
(void)Bind(&C3Label);
dc64(0x3EC7'1DE3'7A97'D93EULL); // c3
(void)Bind(&C4Label);
dc64(0xBE5A'E633'9199'87C6ULL); // c4
(void)Bind(&C5Label);
dc64(0x3DE6'0E27'7AE0'7CECULL); // c5
(void)Bind(&C6Label);
dc64(0xBD69'E954'0300'A100ULL); // c6
(void)Bind(&RangeLabel);
dc64(0x4160'0000'0000'0000ULL); // 2^23
}
void Dispatcher::EmitF64Tan() {
F64TanHandlerAddress = GetCursorAddress<uint64_t>();
constexpr auto Accum = ARMEmitter::VReg::v2;
ARMEmitter::ForwardLabel Fallback, NonZero;
ARMEmitter::ForwardLabel RangeLabel, TwoOverPiLabel;
ARMEmitter::ForwardLabel HalfPi0Label, HalfPi1Label;
ARMEmitter::ForwardLabel C0Label, C1Label, C2Label, C3Label, C4Label, C5Label, C6Label, C7Label, C8Label;
// tan(+/-0) = +/-0
fmov(ARMEmitter::Size::i64Bit, TMP1, VTMP1.D());
lsl(ARMEmitter::Size::i64Bit, TMP1, TMP1, 1);
(void)cbnz(ARMEmitter::Size::i64Bit, TMP1, &NonZero);
ret();
(void)Bind(&NonZero);
// Save q2 for use as accumulator
str<ARMEmitter::IndexType::PRE>(ARMEmitter::QReg::q2, ARMEmitter::Reg::rsp, -16);
// save nzcv
mrs(TMP1, ARMEmitter::SystemRegister::NZCV);
str(TMP1.W(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.flags[24]));
// Range check: fall back for |x| >= 2^23, NaN, and inf.
fabs(VTMP2.D(), VTMP1.D());
ldr(Accum.D(), &RangeLabel);
fcmp(VTMP2.D(), Accum.D());
(void)b(ARMEmitter::Condition::CC_HS, &Fallback);
// q = nearest integer to 2 * x / pi.
ldr(VTMP2.D(), &TwoOverPiLabel);
fmul(VTMP2.D(), VTMP1.D(), VTMP2.D());
frinta(VTMP2.D(), VTMP2.D());
// qi = int(q).
fcvtzs(ARMEmitter::Size::i64Bit, TMP1, VTMP2.D());
// r = x - q * pi/2 (range reduction), in extended precision.
fmov(Accum.D(), VTMP1.D());
ldr(VTMP1.D(), &HalfPi0Label);
fmsub(Accum.D(), VTMP2.D(), VTMP1.D(), Accum.D());
ldr(VTMP1.D(), &HalfPi1Label);
fmsub(Accum.D(), VTMP2.D(), VTMP1.D(), Accum.D());
// Further reduce r to [-pi/8, pi/8].
fmov(ARMEmitter::ScalarRegSize::i64Bit, VTMP1, 0.5f);
fmul(Accum.D(), Accum.D(), VTMP1.D());
// Approximate tan(r) using order 8 polynomial.
fmul(VTMP1.D(), Accum.D(), Accum.D());
fmov(ARMEmitter::Size::i64Bit, TMP2, Accum.D());
// Horner: p = C8 + r2*(C7 + r2*(... + r2*C0)).
ldr(VTMP2.D(), &C8Label);
ldr(Accum.D(), &C7Label);
fmadd(VTMP2.D(), VTMP1.D(), VTMP2.D(), Accum.D());
ldr(Accum.D(), &C6Label);
fmadd(VTMP2.D(), VTMP1.D(), VTMP2.D(), Accum.D());
ldr(Accum.D(), &C5Label);
fmadd(VTMP2.D(), VTMP1.D(), VTMP2.D(), Accum.D());
ldr(Accum.D(), &C4Label);
fmadd(VTMP2.D(), VTMP1.D(), VTMP2.D(), Accum.D());
ldr(Accum.D(), &C3Label);
fmadd(VTMP2.D(), VTMP1.D(), VTMP2.D(), Accum.D());
ldr(Accum.D(), &C2Label);
fmadd(VTMP2.D(), VTMP1.D(), VTMP2.D(), Accum.D());
ldr(Accum.D(), &C1Label);
fmadd(VTMP2.D(), VTMP1.D(), VTMP2.D(), Accum.D());
ldr(Accum.D(), &C0Label);
fmadd(VTMP2.D(), VTMP1.D(), VTMP2.D(), Accum.D());
// p = r + r^3 * p.
fmov(ARMEmitter::Size::i64Bit, Accum.D(), TMP2);
fmul(VTMP1.D(), VTMP1.D(), Accum.D());
fmadd(Accum.D(), VTMP1.D(), VTMP2.D(), Accum.D());
// Double-angle reconstruction: tan(2x) = 2*tan(x) / (1 - tan^2(x)).
fadd(VTMP1.D(), Accum.D(), Accum.D());
fmul(VTMP2.D(), Accum.D(), Accum.D());
fmov(ARMEmitter::ScalarRegSize::i64Bit, Accum, 1.0f);
fsub(VTMP2.D(), VTMP2.D(), Accum.D());
ARMEmitter::ForwardLabel SkipSwap;
(void)tbnz(TMP1, 0, &SkipSwap);
fneg(Accum.D(), VTMP1.D());
fmov(VTMP1.D(), VTMP2.D());
fmov(VTMP2.D(), Accum.D());
(void)Bind(&SkipSwap);
// result = numerator / denominator -> VTMP1.
fdiv(VTMP1.D(), VTMP2.D(), VTMP1.D());
// restore nzcv
ldr(TMP1.W(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.flags[24]));
msr(ARMEmitter::SystemRegister::NZCV, TMP1);
// Restore q2 and return.
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::QReg::q2, ARMEmitter::Reg::rsp, 16);
ret();
// Fallback path.
(void)Bind(&Fallback);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::QReg::q2, ARMEmitter::Reg::rsp, 16);
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[FEXCore::Core::OPINDEX_F64TAN].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[FEXCore::Core::OPINDEX_F64TAN].Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
ret();
// Constant pool.
Align(16);
(void)Bind(&TwoOverPiLabel);
dc64(0x3FE4'5F30'6DC9'C883ULL); // two_over_pi
(void)Bind(&HalfPi0Label);
dc64(0x3FF9'21FB'5444'2D18ULL); // half_pi[0]
(void)Bind(&HalfPi1Label);
dc64(0x3C91'A626'3314'5C07ULL); // half_pi[1]
(void)Bind(&C0Label);
dc64(0x3FD5'5555'5555'5556ULL); // C0
(void)Bind(&C1Label);
dc64(0x3FC1'1111'1111'0A63ULL); // C1
(void)Bind(&C2Label);
dc64(0x3FAB'A1BA'1BB4'6414ULL); // C2
(void)Bind(&C3Label);
dc64(0x3F96'64F4'7E5B'5445ULL); // C3
(void)Bind(&C4Label);
dc64(0x3F82'26E5'E5EC'DFA3ULL); // C4
(void)Bind(&C5Label);
dc64(0x3F6D'6C7D'DBF8'7047ULL); // C5
(void)Bind(&C6Label);
dc64(0x3F57'EA75'D05B'583EULL); // C6
(void)Bind(&C7Label);
dc64(0x3F42'89F2'2964'A03CULL); // C7
(void)Bind(&C8Label);
dc64(0x3F34'E4FD'1414'7622ULL); // C8
(void)Bind(&RangeLabel);
dc64(0x4160'0000'0000'0000ULL); // 2^23
}
void Dispatcher::EmitF64Scale() {
// Computes result = src1 * 2^trunc(src2).
// Input: VTMP1 = base (src1), VTMP2 = exponent (src2). Output: VTMP1.
F64ScaleHandlerAddress = GetCursorAddress<uint64_t>();
ARMEmitter::ForwardLabel Fallback;
// save nzcv
mrs(TMP1, ARMEmitter::SystemRegister::NZCV);
str(TMP1.W(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.flags[24]));
// n = trunc(src2).
frintz(VTMP2.D(), VTMP2.D());
// NaN check: NaN != NaN sets V flag.
fcmp(VTMP2.D(), VTMP2.D());
(void)b(ARMEmitter::Condition::CC_VS, &Fallback);
fcvtzs(ARMEmitter::Size::i64Bit, TMP1, VTMP2.D());
// Range check: int_n in [-1022, 1023].
cmn(ARMEmitter::Size::i64Bit, TMP1, 1022);
(void)b(ARMEmitter::Condition::CC_LT, &Fallback);
cmp(ARMEmitter::Size::i64Bit, TMP1, 1023);
(void)b(ARMEmitter::Condition::CC_GT, &Fallback);
// 2^n, then result = src1 * 2^n.
add(ARMEmitter::Size::i64Bit, TMP1, TMP1, 1023);
lsl(ARMEmitter::Size::i64Bit, TMP1, TMP1, 52);
fmov(ARMEmitter::Size::i64Bit, VTMP2.D(), TMP1);
fmul(VTMP1.D(), VTMP1.D(), VTMP2.D());
// restore nzcv
ldr(TMP1.W(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.flags[24]));
msr(ARMEmitter::SystemRegister::NZCV, TMP1);
ret();
// Fallback path.
(void)Bind(&Fallback);
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[FEXCore::Core::OPINDEX_F64SCALE].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[FEXCore::Core::OPINDEX_F64SCALE].Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
ret();
}
void Dispatcher::EmitF64F2XM1() {
// JIT-inlined double-precision 2^x - 1 for x in [-1, 1].
// Uses argument reduction: split x into n = round(x) and r = x - n,
// then compute 2^x - 1 = 2^n * (2^r - 1) + (2^n - 1).
// 2^r - 1 is approximated via a 13-term Horner polynomial in r.
// Input in VTMP1.D(), output in VTMP1.D().
F64F2XM1HandlerAddress = GetCursorAddress<uint64_t>();
constexpr auto Accum = ARMEmitter::VReg::v2;
ARMEmitter::ForwardLabel Fallback;
ARMEmitter::ForwardLabel OneLabel;
ARMEmitter::ForwardLabel C1Label, C2Label, C3Label, C4Label, C5Label, C6Label;
ARMEmitter::ForwardLabel C7Label, C8Label, C9Label, C10Label, C11Label, C12Label, C13Label;
// Save q2.
str<ARMEmitter::IndexType::PRE>(ARMEmitter::QReg::q2, ARMEmitter::Reg::rsp, -16);
// save nzcv
mrs(TMP1, ARMEmitter::SystemRegister::NZCV);
str(TMP1.W(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.flags[24]));
// Range check: |x| > 1.0 -> fallback.
fabs(VTMP2.D(), VTMP1.D());
ldr(Accum.D(), &OneLabel);
fcmp(VTMP2.D(), Accum.D());
(void)b(ARMEmitter::Condition::CC_HI, &Fallback);
// Argument reduction: n = round(x), r = x - n.
frinta(VTMP2.D(), VTMP1.D());
fcvtzs(ARMEmitter::Size::i64Bit, TMP1, VTMP2.D());
fsub(Accum.D(), VTMP1.D(), VTMP2.D());
// scale = 2^n, scale_m1 = 2^n - 1.
// TMP1 = scale bits, TMP3 = scale_m1 bits, Accum = r.
add(ARMEmitter::Size::i64Bit, TMP1, TMP1, 1023);
lsl(ARMEmitter::Size::i64Bit, TMP1, TMP1, 52);
fmov(ARMEmitter::Size::i64Bit, VTMP2.D(), TMP1);
ldr(VTMP1.D(), &OneLabel);
fsub(VTMP1.D(), VTMP2.D(), VTMP1.D());
fmov(ARMEmitter::Size::i64Bit, TMP3, VTMP1.D());
// Horner polynomial: p = c1 + r * (c2 + r * (... + r * c13)).
ldr(VTMP1.D(), &C13Label);
ldr(VTMP2.D(), &C12Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C11Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C10Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C9Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C8Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C7Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C6Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C5Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C4Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C3Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C2Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C1Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
// q = r * p, then result = scale * q + scale_m1.
fmul(VTMP1.D(), Accum.D(), VTMP1.D());
fmov(ARMEmitter::Size::i64Bit, Accum.D(), TMP1);
fmov(ARMEmitter::Size::i64Bit, VTMP2.D(), TMP3);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
// restore nzcv
ldr(TMP1.W(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.flags[24]));
msr(ARMEmitter::SystemRegister::NZCV, TMP1);
// Restore q2 and return.
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::QReg::q2, ARMEmitter::Reg::rsp, 16);
ret();
// Fallback path.
(void)Bind(&Fallback);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::QReg::q2, ARMEmitter::Reg::rsp, 16);
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[FEXCore::Core::OPINDEX_F64F2XM1].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[FEXCore::Core::OPINDEX_F64F2XM1].Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
ret();
// Constant pool
Align(16);
(void)Bind(&OneLabel);
dc64(0x3FF0'0000'0000'0000ULL); // 1.0
(void)Bind(&C1Label);
dc64(0x3FE6'2E42'FEFA'39EFULL); // ln(2)
(void)Bind(&C2Label);
dc64(0x3FCE'BFBD'FF82'C58EULL);
(void)Bind(&C3Label);
dc64(0x3FAC'6B08'D704'A0BEULL);
(void)Bind(&C4Label);
dc64(0x3F83'B2AB'6FBA'4E76ULL);
(void)Bind(&C5Label);
dc64(0x3F55'D87F'E78A'672FULL);
(void)Bind(&C6Label);
dc64(0x3F24'3091'2F86'C785ULL);
(void)Bind(&C7Label);
dc64(0x3EEF'FCBF'C588'B0C2ULL);
(void)Bind(&C8Label);
dc64(0x3EB6'2C02'23A5'C821ULL);
(void)Bind(&C9Label);
dc64(0x3E7B'5253'D395'E7C0ULL);
(void)Bind(&C10Label);
dc64(0x3E3E'4CF5'158B'8EC5ULL);
(void)Bind(&C11Label);
dc64(0x3DFE'8CAC'7351'BB20ULL);
(void)Bind(&C12Label);
dc64(0x3DBC'3BD6'50FC'2981ULL);
(void)Bind(&C13Label);
dc64(0x3D78'1619'3166'D0F5ULL);
}
// JIT-inlined double-precision atan2 for the F64 reduced precision x87 path.
// Input: VTMP1 = y, VTMP2 = x. Output: VTMP1 = atan2(y, x).
// Algorithm: 20-term Horner polynomial from ARM optimized-routines atan_data.c.
// atan(z) = z + z^3 * P(z^2), with range reduction to [0,1] and quadrant adjustment.
void Dispatcher::EmitF64Atan() {
F64AtanHandlerAddress = GetCursorAddress<uint64_t>();
constexpr auto Accum = ARMEmitter::VReg::v2;
ARMEmitter::ForwardLabel Fallback;
ARMEmitter::ForwardLabel NoSwap, PosX, NegY;
ARMEmitter::ForwardLabel PiOver2Label, PiLabel;
ARMEmitter::ForwardLabel C0Label, C1Label, C2Label, C3Label, C4Label, C5Label, C6Label;
ARMEmitter::ForwardLabel C7Label, C8Label, C9Label, C10Label, C11Label, C12Label, C13Label;
ARMEmitter::ForwardLabel C14Label, C15Label, C16Label, C17Label, C18Label, C19Label;
// Stack layout: [sp] = q2 (16B), [sp+16] = y bits (8B), [sp+24] = x bits (8B).
str<ARMEmitter::IndexType::PRE>(ARMEmitter::QReg::q2, ARMEmitter::Reg::rsp, -32);
mrs(TMP1, ARMEmitter::SystemRegister::NZCV);
str(TMP1.W(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.flags[24]));
fmov(ARMEmitter::Size::i64Bit, TMP1, VTMP2.D());
fmov(ARMEmitter::Size::i64Bit, TMP2, VTMP1.D());
str(TMP2, ARMEmitter::Reg::rsp, 16);
str(TMP1, ARMEmitter::Reg::rsp, 24);
// Compute |x|, |y|, pack sign/swap flags into TMP1 = (sign_x<<2)|(sign_y<<1)|swap.
fabs(Accum.D(), VTMP2.D());
fabs(VTMP2.D(), VTMP1.D());
fcmp(VTMP2.D(), Accum.D());
lsr(ARMEmitter::Size::i64Bit, TMP1, TMP1, 63);
lsr(ARMEmitter::Size::i64Bit, TMP2, TMP2, 63);
cset(ARMEmitter::Size::i64Bit, TMP3, ARMEmitter::Condition::CC_HI);
lsl(ARMEmitter::Size::i64Bit, TMP1, TMP1, 2);
orr(ARMEmitter::Size::i64Bit, TMP1, TMP1, TMP2, ARMEmitter::ShiftType::LSL, 1);
orr(ARMEmitter::Size::i64Bit, TMP1, TMP1, TMP3);
// z = min(|y|,|x|) / max(|y|,|x|); NaN guard catches 0/0, inf/inf, NaN inputs.
fcsel(ARMEmitter::ScalarRegSize::i64Bit, VTMP1, Accum, VTMP2, ARMEmitter::Condition::CC_HI);
fcsel(ARMEmitter::ScalarRegSize::i64Bit, Accum, VTMP2, Accum, ARMEmitter::Condition::CC_HI);
fdiv(VTMP2.D(), VTMP1.D(), Accum.D());
fcmp(VTMP2.D(), VTMP2.D());
(void)b(ARMEmitter::Condition::CC_VS, &Fallback);
// P(z^2) via 20-term Horner.
fmul(Accum.D(), VTMP2.D(), VTMP2.D());
fmov(ARMEmitter::Size::i64Bit, TMP3, VTMP2.D());
ldr(VTMP1.D(), &C19Label);
ldr(VTMP2.D(), &C18Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C17Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C16Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C15Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C14Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C13Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C12Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C11Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C10Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C9Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C8Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C7Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C6Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C5Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C4Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C3Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C2Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C1Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
ldr(VTMP2.D(), &C0Label);
fmadd(VTMP1.D(), Accum.D(), VTMP1.D(), VTMP2.D());
// atan_abs = z + z^3 * P(z^2).
fmov(ARMEmitter::Size::i64Bit, VTMP2.D(), TMP3);
fmul(VTMP1.D(), Accum.D(), VTMP1.D());
fmadd(VTMP1.D(), VTMP2.D(), VTMP1.D(), VTMP2.D());
// Quadrant adjustment driven by packed flags in TMP1.
(void)tbz(TMP1, 0, &NoSwap);
ldr(VTMP2.D(), &PiOver2Label);
fsub(VTMP1.D(), VTMP2.D(), VTMP1.D());
(void)Bind(&NoSwap);
(void)tbz(TMP1, 2, &PosX);
ldr(VTMP2.D(), &PiLabel);
fsub(VTMP1.D(), VTMP2.D(), VTMP1.D());
(void)Bind(&PosX);
(void)tbz(TMP1, 1, &NegY);
fneg(VTMP1.D(), VTMP1.D());
(void)Bind(&NegY);
ldr(TMP1.W(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.flags[24]));
msr(ARMEmitter::SystemRegister::NZCV, TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::QReg::q2, ARMEmitter::Reg::rsp, 32);
ret();
// Fallback path: restore original inputs from stack stash and dispatch the ABI handler.
(void)Bind(&Fallback);
ldr(TMP1.W(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.flags[24]));
msr(ARMEmitter::SystemRegister::NZCV, TMP1);
ldr(TMP2, ARMEmitter::Reg::rsp, 16);
ldr(TMP1, ARMEmitter::Reg::rsp, 24);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::QReg::q2, ARMEmitter::Reg::rsp, 32);
fmov(ARMEmitter::Size::i64Bit, VTMP1.D(), TMP2);
fmov(ARMEmitter::Size::i64Bit, VTMP2.D(), TMP1);
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[FEXCore::Core::OPINDEX_F64ATAN].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[FEXCore::Core::OPINDEX_F64ATAN].Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
ret();
Align(16);
(void)Bind(&C19Label);
dc64(0x3EF3'5885'1160'A528ULL);
(void)Bind(&C18Label);
dc64(0xBF2A'B24D'A7BE'7402ULL);
(void)Bind(&C17Label);
dc64(0x3F51'7739'E210'171AULL);
(void)Bind(&C16Label);
dc64(0xBF6D'0062'B42F'E3BFULL);
(void)Bind(&C15Label);
dc64(0x3F81'4E9D'C19A'4A4EULL);
(void)Bind(&C14Label);
dc64(0xBF90'0513'8172'2A59ULL);
(void)Bind(&C13Label);
dc64(0x3F98'6089'7B29'E5EFULL);
(void)Bind(&C12Label);
dc64(0xBFA0'0E6E'ECE7'DE80ULL);
(void)Bind(&C11Label);
dc64(0x3FA3'38E3'1EB2'FBBCULL);
(void)Bind(&C10Label);
dc64(0xBFA5'D301'40AE'5E99ULL);
(void)Bind(&C9Label);
dc64(0x3FA8'42DB'E9B0'D916ULL);
(void)Bind(&C8Label);
dc64(0xBFAA'EBFE'7B41'8581ULL);
(void)Bind(&C7Label);
dc64(0x3FAE'1D0F'9696'F63BULL);
(void)Bind(&C6Label);
dc64(0xBFB1'1100'EE08'4227ULL);
(void)Bind(&C5Label);
dc64(0x3FB3'B139'B6A8'8BA1ULL);
(void)Bind(&C4Label);
dc64(0xBFB7'45D1'60A7'E368ULL);
(void)Bind(&C3Label);
dc64(0x3FBC'71C7'1BC3'951CULL);
(void)Bind(&C2Label);
dc64(0xBFC2'4924'9247'8F88ULL);
(void)Bind(&C1Label);
dc64(0x3FC9'9999'9999'96C1ULL);
(void)Bind(&C0Label);
dc64(0xBFD5'5555'5555'5555ULL);
(void)Bind(&PiOver2Label);
dc64(0x3FF9'21FB'5444'2D18ULL);
(void)Bind(&PiLabel);
dc64(0x4009'21FB'5444'2D18ULL);
}
// JIT-inlined double-precision y * log2(x) for the F64 reduced precision x87 path.
// Input: VTMP1 = x, VTMP2 = y. Output: VTMP1 = y * log2(x).
// Algorithm: atanh-based log via s = f/(2+f) with 9-term polynomial, scaled by 1/ln(2),
// then multiplied by y.
void Dispatcher::EmitF64FYL2X() {
F64FYL2XHandlerAddress = GetCursorAddress<uint64_t>();
constexpr auto Accum = ARMEmitter::VReg::v2;
ARMEmitter::ForwardLabel Fallback;
ARMEmitter::ForwardLabel NoNorm;
ARMEmitter::ForwardLabel Sqrt2Label, Log2eLabel;
ARMEmitter::ForwardLabel BiasLabel;
ARMEmitter::ForwardLabel P0Label, P1Label, P2Label, P3Label, P4Label, P5Label, P6Label, P7Label, P8Label;
str<ARMEmitter::IndexType::PRE>(ARMEmitter::QReg::q2, ARMEmitter::Reg::rsp, -16);
mrs(TMP1, ARMEmitter::SystemRegister::NZCV);
str(TMP1.W(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.flags[24]));
// Reject x <= 0, subnormal, inf and NaN before any FPR is clobbered,
// so VTMP1/VTMP2 still hold the original inputs at the fallback.
fmov(ARMEmitter::Size::i64Bit, TMP1, VTMP1.D());
(void)tbnz(TMP1, 63, &Fallback);
lsr(ARMEmitter::Size::i64Bit, TMP2, TMP1, 52);
(void)cbz(ARMEmitter::Size::i64Bit, TMP2, &Fallback);
cmp(ARMEmitter::Size::i64Bit, TMP2, 0x7FF);
(void)b(ARMEmitter::Condition::CC_EQ, &Fallback);
fmov(ARMEmitter::Size::i64Bit, TMP3, VTMP2.D());
// Extract k and normalize mantissa m into [1.0, 2.0).
sub(ARMEmitter::Size::i64Bit, TMP2, TMP2, 1023);
ubfx(ARMEmitter::Size::i64Bit, TMP1, TMP1, 0, 52);
ldr(TMP4, &BiasLabel);
orr(ARMEmitter::Size::i64Bit, TMP1, TMP1, TMP4);
fmov(ARMEmitter::Size::i64Bit, VTMP1.D(), TMP1);
// If m > sqrt(2), halve m and increment k.
ldr(VTMP2.D(), &Sqrt2Label);
fcmp(VTMP1.D(), VTMP2.D());
(void)b(ARMEmitter::Condition::CC_LE, &NoNorm);
fmov(ARMEmitter::ScalarRegSize::i64Bit, VTMP2, 0.5f);
fmul(VTMP1.D(), VTMP1.D(), VTMP2.D());
add(ARMEmitter::Size::i64Bit, TMP2, TMP2, 1);
(void)Bind(&NoNorm);
// f = m - 1; s = f / (2 + f); TMP1 stashes s, VTMP1 holds s^2.
fmov(ARMEmitter::ScalarRegSize::i64Bit, VTMP2, 1.0f);
fsub(VTMP1.D(), VTMP1.D(), VTMP2.D());
fmov(ARMEmitter::ScalarRegSize::i64Bit, VTMP2, 2.0f);
fadd(VTMP2.D(), VTMP1.D(), VTMP2.D());
fdiv(Accum.D(), VTMP1.D(), VTMP2.D());
fmul(VTMP1.D(), Accum.D(), Accum.D());
fmov(ARMEmitter::Size::i64Bit, TMP1, Accum.D());
// 9-term Horner from 1/19 down to 1/3.
ldr(Accum.D(), &P8Label);
ldr(VTMP2.D(), &P7Label);
fmadd(Accum.D(), VTMP1.D(), Accum.D(), VTMP2.D());
ldr(VTMP2.D(), &P6Label);
fmadd(Accum.D(), VTMP1.D(), Accum.D(), VTMP2.D());
ldr(VTMP2.D(), &P5Label);
fmadd(Accum.D(), VTMP1.D(), Accum.D(), VTMP2.D());
ldr(VTMP2.D(), &P4Label);
fmadd(Accum.D(), VTMP1.D(), Accum.D(), VTMP2.D());
ldr(VTMP2.D(), &P3Label);
fmadd(Accum.D(), VTMP1.D(), Accum.D(), VTMP2.D());
ldr(VTMP2.D(), &P2Label);
fmadd(Accum.D(), VTMP1.D(), Accum.D(), VTMP2.D());
ldr(VTMP2.D(), &P1Label);
fmadd(Accum.D(), VTMP1.D(), Accum.D(), VTMP2.D());
ldr(VTMP2.D(), &P0Label);
fmadd(Accum.D(), VTMP1.D(), Accum.D(), VTMP2.D());
// ln(1+f) = 2 * s * (1 + s^2 * P(s^2)).
fmul(Accum.D(), VTMP1.D(), Accum.D());
fmov(ARMEmitter::ScalarRegSize::i64Bit, VTMP2, 1.0f);
fadd(Accum.D(), Accum.D(), VTMP2.D());
fmov(ARMEmitter::Size::i64Bit, VTMP2.D(), TMP1);
fmul(Accum.D(), VTMP2.D(), Accum.D());
fadd(Accum.D(), Accum.D(), Accum.D());
// log2(x) = k + ln(1+f)/ln(2); multiply by y.
ldr(VTMP1.D(), &Log2eLabel);
fmul(VTMP1.D(), Accum.D(), VTMP1.D());
scvtf(ARMEmitter::Size::i64Bit, VTMP2.D(), TMP2);
fadd(VTMP1.D(), VTMP1.D(), VTMP2.D());
fmov(ARMEmitter::Size::i64Bit, VTMP2.D(), TMP3);
fmul(VTMP1.D(), VTMP1.D(), VTMP2.D());
ldr(TMP1.W(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.flags[24]));
msr(ARMEmitter::SystemRegister::NZCV, TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::QReg::q2, ARMEmitter::Reg::rsp, 16);
ret();
// Fallback path: VTMP1/VTMP2 still hold the original x/y.
(void)Bind(&Fallback);
ldr(TMP1.W(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.flags[24]));
msr(ARMEmitter::SystemRegister::NZCV, TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::QReg::q2, ARMEmitter::Reg::rsp, 16);
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[FEXCore::Core::OPINDEX_F64FYL2X].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[FEXCore::Core::OPINDEX_F64FYL2X].Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
ret();
// Constant pool: bias=1.0, sqrt(2), log2(e)=1/ln(2), P8..P0 = 1/19, 1/17, 1/15, ..., 1/3.
Align(16);
(void)Bind(&BiasLabel);
dc64(0x3FF0'0000'0000'0000ULL);
(void)Bind(&Sqrt2Label);
dc64(0x3FF6'A09E'667F'3BCDULL);
(void)Bind(&Log2eLabel);
dc64(0x3FF7'1547'652B'82FEULL);
(void)Bind(&P8Label);
dc64(0x3FAA'F286'BCA1'AF28ULL);
(void)Bind(&P7Label);
dc64(0x3FAE'1E1E'1E1E'1E1EULL);
(void)Bind(&P6Label);
dc64(0x3FB1'1111'1111'1111ULL);
(void)Bind(&P5Label);
dc64(0x3FB3'B13B'13B1'3B14ULL);
(void)Bind(&P4Label);
dc64(0x3FB7'45D1'745D'1746ULL);
(void)Bind(&P3Label);
dc64(0x3FBC'71C7'1C71'C71CULL);
(void)Bind(&P2Label);
dc64(0x3FC2'4924'9249'2492ULL);
(void)Bind(&P1Label);
dc64(0x3FC9'9999'9999'999AULL);
(void)Bind(&P0Label);
dc64(0x3FD5'5555'5555'5555ULL);
}
uint64_t Dispatcher::GenerateABICall(FallbackABI ABI) {
auto Address = GetCursorAddress<uint64_t>();
constexpr static auto FallbackPointerReg = TMP4;
@@ -1289,6 +2183,13 @@ void Dispatcher::InitThreadPointers(FEXCore::Core::InternalThreadState* Thread)
Ptrs.SignalReturnHandlerRT = SignalHandlerReturnAddressRT;
Ptrs.LUDIVHandler = LUDIVHandlerAddress;
Ptrs.LDIVHandler = LDIVHandlerAddress;
Ptrs.F64SinHandler = F64SinHandlerAddress;
Ptrs.F64CosHandler = F64CosHandlerAddress;
Ptrs.F64TanHandler = F64TanHandlerAddress;
Ptrs.F64F2XM1Handler = F64F2XM1HandlerAddress;
Ptrs.F64ScaleHandler = F64ScaleHandlerAddress;
Ptrs.F64AtanHandler = F64AtanHandlerAddress;
Ptrs.F64FYL2XHandler = F64FYL2XHandlerAddress;
// Fill in the fallback handlers
InterpreterOps::FillFallbackIndexPointers(Ptrs.FallbackHandlerPointers, &ABIPointers[0]);
@@ -28,6 +28,10 @@ class ContextImpl;
namespace FEXCore::CPU {
#define STATE_PTR(STATE_TYPE, FIELD) STATE.R(), offsetof(FEXCore::Core::STATE_TYPE, FIELD)
#define STATE_PTR_IDX(STATE_TYPE, FIELD, INDEX) STATE.R(), ARRAY_OFFSETOF(FEXCore::Core::STATE_TYPE, FIELD, INDEX)
#define FALLBACK_HANDLER_OFFSET(INDEX, FIELD) \
STATE.R(), \
(ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, Pointers.FallbackHandlerPointers, INDEX) + offsetof(FEXCore::Core::FallbackABIInfo, FIELD))
class Dispatcher final : public Arm64Emitter {
public:
@@ -95,6 +99,15 @@ private:
uint64_t LUDIVHandlerAddress {};
uint64_t LDIVHandlerAddress {};
// F64 reduced-precision shared handlers
uint64_t F64SinHandlerAddress {};
uint64_t F64CosHandlerAddress {};
uint64_t F64TanHandlerAddress {};
uint64_t F64F2XM1HandlerAddress {};
uint64_t F64ScaleHandlerAddress {};
uint64_t F64AtanHandlerAddress {};
uint64_t F64FYL2XHandlerAddress {};
void EmitDispatcher();
uint64_t GenerateABICall(FallbackABI ABI);
@@ -105,6 +118,14 @@ private:
void EmitF32ToExtF80();
void EmitF64ToExtF80();
void EmitF64Sin();
void EmitF64Cos();
void EmitF64Tan();
void EmitF64F2XM1();
void EmitF64Scale();
void EmitF64Atan();
void EmitF64FYL2X();
FEX_CONFIG_OPT(DisableL2Cache, DISABLEL2CACHE);
};
+15 -1
View File
@@ -1348,6 +1348,13 @@ const uint8_t* Decoder::AdjustAddrForSpecialRegion(const uint8_t* _InstStream, u
return _InstStream - EntryPoint + RIP;
}
bool Decoder::CheckIfCacheable(FEXCore::Core::InternalThreadState& Thread, const uint8_t* InstStream, uint64_t PC, uint64_t MaxInst) {
DecodeInstructionsAtEntry(&Thread, InstStream, PC, MaxInst);
bool Uncacheable = HitBadRelocation;
DelayedDisownBuffer();
return !Uncacheable;
}
void Decoder::DecodeInstructionsAtEntry(FEXCore::Core::InternalThreadState* Thread, const uint8_t* _InstStream, uint64_t PC, uint64_t MaxInst) {
FEXCORE_PROFILE_SCOPED("DecodeInstructions");
BlockInfo.TotalInstructionCount = 0;
@@ -1465,6 +1472,13 @@ void Decoder::DecodeInstructionsAtEntry(FEXCore::Core::InternalThreadState* Thre
}
BlockIt->BlockStatus = DecodeInstruction(OpAddress);
if (HitBadRelocation) {
BlockInfo.TotalInstructionCount = 0;
BlockInfo.Blocks = {*BlockIt};
BlockInfo.EntryPoints.clear();
BlockInfo.CodePages.clear();
return;
}
uint64_t OpEndAddress = OpAddress + DecodeInst->InstSize;
DecodedMinAddress = std::min(DecodedMinAddress, OpAddress);
@@ -1483,7 +1497,7 @@ void Decoder::DecodeInstructionsAtEntry(FEXCore::Core::InternalThreadState* Thre
// Can not continue this block at all on invalid instruction
if (BlockIt->BlockStatus != DecodedBlockStatus::SUCCESS) [[unlikely]] {
if (!EntryBlock) {
if (!EntryBlock && BlockIt->BlockStatus != DecodedBlockStatus::BAD_RELOCATION) {
// In multiblock configurations, we can early terminate any non-entrypoint blocks with the expectation that this won't get hit.
// Improves compile-times.
// Just need to undo additions that this block decoding has caused.
+1
View File
@@ -54,6 +54,7 @@ public:
};
Decoder(FEXCore::Core::InternalThreadState* Thread);
bool CheckIfCacheable(FEXCore::Core::InternalThreadState&, const uint8_t* InstStream, uint64_t PC, uint64_t MaxInst);
void DecodeInstructionsAtEntry(FEXCore::Core::InternalThreadState* Thread, const uint8_t* InstStream, uint64_t PC, uint64_t MaxInst);
const DecodedBlockInformation* GetDecodedBlockInfo() const {
+1 -1
View File
@@ -274,7 +274,7 @@ DEF_OP(CmpPairZ) {
// Restore NzCV
if (CTX->HostFeatures.SupportsFlagM) {
rmif(TMP1, 0, 0xb /* NzCV */);
rmif(TMP1, 28, 0xb /* NzCV */);
} else {
cset(ARMEmitter::Size::i32Bit, TMP2, ARMEmitter::Condition::CC_EQ);
bfi(ARMEmitter::Size::i32Bit, TMP1, TMP2, 30 /* lsb: Z */, 1);
@@ -329,7 +329,7 @@ DEF_OP(TelemetrySetValue) {
auto Op = IROp->C<IR::IROp_TelemetrySetValue>();
auto Src = GetReg(Op->Value);
ldr(TMP2, STATE_PTR(CpuStateFrame, Pointers.TelemetryValueAddresses[Op->TelemetryValueIndex]));
ldr(TMP2, STATE_PTR_IDX(CpuStateFrame, Pointers.TelemetryValueAddresses, Op->TelemetryValueIndex));
// Cortex fuses cmp+cset.
cmp(ARMEmitter::Size::i32Bit, Src, 0);
@@ -265,7 +265,10 @@ DEF_OP(Syscall) {
uint32_t GPRSpillMask = ~0U;
uint32_t FPRSpillMask = ~0U;
SpillStaticRegs(TMP1, true, GPRSpillMask, FPRSpillMask);
SpillStaticRegs(TMP1, {
.GPRSpillMask = GPRSpillMask,
.FPRSpillMask = FPRSpillMask,
});
// Now that we are spilled, store in the state that we are in a syscall
// Still without overwriting registers that matter
@@ -299,7 +302,12 @@ DEF_OP(Syscall) {
// Result is now in x0
// Fix the stack and any values that were stepped on
FillStaticRegs(true, GPRSpillMask, FPRSpillMask, ARMEmitter::Reg::r1, ARMEmitter::Reg::r2);
FillStaticRegs({
.OptionalReg = ARMEmitter::Reg::r1,
.OptionalReg2 = ARMEmitter::Reg::r2,
.GPRFillMask = GPRSpillMask,
.FPRFillMask = FPRSpillMask,
});
// Now the registers we've spilled are back in their original host registers
// We can safely claim we are no longer in a syscall
@@ -322,7 +330,10 @@ DEF_OP(Thunk) {
// X0: CTX
// X1: Args (from guest stack)
SpillStaticRegs(TMP1, true, ~0U, ~0U, false); // spill to ctx before ra64 spill
// spill to ctx before ra64 spill
SpillStaticRegs(TMP1, {
.NZCV = false,
});
PushDynamicRegs(TMP1);
@@ -337,7 +348,10 @@ DEF_OP(Thunk) {
PopDynamicRegs();
FillStaticRegs(true, ~0U, ~0U, std::nullopt, std::nullopt, false); // load from ctx after ra64 refill
// load from ctx after ra64 refill
FillStaticRegs({
.NZCV = false,
});
}
DEF_OP(ValidateCode) {
+41 -34
View File
@@ -68,6 +68,10 @@ PrintValue(uint64_t Value) {
LogMan::Msg::DFmt("Value: 0x{:x}", Value);
}
static void PrintMsg(const char* Value) {
LogMan::Msg::DFmt("{}", Value);
}
static void PrintVectorValue(uint64_t Value, uint64_t ValueUpper) {
LogMan::Msg::DFmt("Value: 0x{:016x}'{:016x}", ValueUpper, Value);
}
@@ -133,8 +137,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
const auto Src1 = GetVReg(IROp->Args[0]);
fmov(VTMP1.S(), Src1.S());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -151,8 +155,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
const auto Src1 = GetVReg(IROp->Args[0]);
fmov(VTMP1.D(), Src1.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -176,8 +180,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
mov(ARMEmitter::Size::i32Bit, TMP2, Src1);
}
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -194,8 +198,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
const auto Src1 = GetVReg(IROp->Args[0]);
mov(VTMP1.Q(), Src1.Q());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -212,8 +216,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
const auto Src1 = GetVReg(IROp->Args[0]);
mov(VTMP1.Q(), Src1.Q());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -230,8 +234,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
const auto Src1 = GetVReg(IROp->Args[0]);
fmov(VTMP1.D(), Src1.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -254,8 +258,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
fmov(VTMP1.D(), Src1.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -276,8 +280,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
fmov(VTMP1.D(), Src1.D());
fmov(VTMP2.D(), Src2.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -294,8 +298,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
const auto Src1 = GetVReg(IROp->Args[0]);
mov(VTMP1.Q(), Src1.Q());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -312,8 +316,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
const auto Src1 = GetVReg(IROp->Args[0]);
mov(VTMP1.Q(), Src1.Q());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -330,8 +334,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
const auto Src1 = GetVReg(IROp->Args[0]);
mov(VTMP1.Q(), Src1.Q());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -351,8 +355,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
mov(VTMP1.Q(), Src1.Q());
mov(VTMP2.Q(), Src2.Q());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -369,8 +373,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
const auto Src1 = GetVReg(IROp->Args[0]);
mov(VTMP1.Q(), Src1.Q());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -394,8 +398,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
mov(VTMP1.Q(), Src1.Q());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -416,8 +420,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
mov(VTMP1.Q(), Src1.Q());
mov(VTMP2.Q(), Src2.Q());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -434,8 +438,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
// tmp2 (x1/x11): source 2
// tmp3 (x2/x12): source 3
const auto Op = IROp->C<IR::IROp_VPCMPESTRX>();
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
stp<ARMEmitter::IndexType::PRE>(TMP1, ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
@@ -476,8 +480,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
mov(VTMP2.Q(), Src2.Q());
movz(ARMEmitter::Size::i32Bit, TMP1, Control);
ldr(TMP2, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP2, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP2);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -636,6 +640,8 @@ Arm64JITCore::Arm64JITCore(FEXCore::Context::ContextImpl* ctx, FEXCore::Core::In
Ptrs.PrintValue = reinterpret_cast<uint64_t>(PrintValue);
Ptrs.PrintVectorValue = reinterpret_cast<uint64_t>(PrintVectorValue);
Ptrs.PrintMsgValue = reinterpret_cast<uint64_t>(PrintMsg);
Ptrs.ThreadRemoveCodeEntryFromJIT = reinterpret_cast<uintptr_t>(&Context::ContextImpl::ThreadRemoveCodeEntryFromJit);
Ptrs.MonoBackpatcherWrite = reinterpret_cast<uint64_t>(&Context::ContextImpl::MonoBackpatcherWrite);
Ptrs.CPUIDObj = reinterpret_cast<uint64_t>(&CTX->CPUID);
@@ -841,6 +847,7 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
CallReturnTargets.clear();
PendingJumpThunks.clear();
JumpTargets.resize(IR->GetHeader()->BlockCount, {});
Relocations.resize(PrevNumAllocations, FEXCore::CPU::Relocation::Default()); // Discard any relocations generated from a previous attempt
CodeData.EntryPoints.clear();
+154 -81
View File
@@ -563,7 +563,7 @@ DEF_OP(LoadDF) {
auto Flag = X86State::RFLAG_DF_RAW_LOC;
// DF needs sign extension to turn 0x1/0xFF into 1/-1
ldrsb(Dst.X(), STATE, offsetof(FEXCore::Core::CPUState, flags[Flag]));
ldrsb(Dst.X(), STATE, ARRAY_OFFSETOF(FEXCore::Core::CPUState, flags, Flag));
}
DEF_OP(ContextClear) {
@@ -1849,13 +1849,6 @@ DEF_OP(StoreMemTSO) {
}
DEF_OP(MemSet) {
// TODO: A future looking task would be to support this with ARM's MOPS instructions.
// The 8-bit non-atomic forward path directly matches ARM's SETP/SETM/SETE instruction,
// while the backward version needs some fixup to convert it to a forward direction.
//
// Assuming non-atomicity and non-faulting behaviour, this can accelerate this implementation.
// Additionally: This is commonly used as a memset to zero. If we know up-front with an inline constant
// that the value is zero, we can optimize any operation larger than 8-bit down to 8-bit to use the MOPS implementation.
const auto Op = IROp->C<IR::IROp_MemSet>();
const bool IsAtomic = CTX->IsMemcpyAtomicTSOEnabled();
@@ -1933,8 +1926,30 @@ DEF_OP(MemSet) {
ARMEmitter::SubRegSize::i8Bit;
auto EmitMemset = [&](int32_t Direction) {
const int32_t OpSize = Size;
const int32_t SizeDirection = Size * Direction;
const bool IsBackwards = Direction == -1;
// Sets the result to the final address written depending on
// whether or not the memset is forwards or backwards.
const auto MakeFinalAddress = [&] {
if (IsBackwards) {
switch (Size) {
case 1: sub(Dst.X(), MemReg.X(), Length.X()); break;
case 2:
case 4:
case 8: sub(Dst.X(), MemReg.X(), Length.X(), ARMEmitter::ShiftType::LSL, FEXCore::ilog2(Size)); break;
default: LOGMAN_MSG_A_FMT("Unhandled MemSet size: {}", Size); break;
}
} else {
switch (Size) {
case 1: add(Dst.X(), MemReg.X(), Length.X()); break;
case 2:
case 4:
case 8: add(Dst.X(), MemReg.X(), Length.X(), ARMEmitter::ShiftType::LSL, FEXCore::ilog2(Size)); break;
default: LOGMAN_MSG_A_FMT("Unhandled MemSet size: {}", Size); break;
}
}
};
ARMEmitter::BiDirectionalLabel AgainInternal {};
ARMEmitter::ForwardLabel DoneInternal {};
@@ -1943,12 +1958,56 @@ DEF_OP(MemSet) {
(void)cbz(ARMEmitter::Size::i64Bit, TMP1, &DoneInternal);
if (!IsAtomic) {
if (CTX->HostFeatures.SupportsMOPS) {
const bool Is8Bit = SubRegSize == ARMEmitter::SubRegSize::i8Bit;
// We can handle 8-bit memsets and any other size that happens
// to be using an inlined zero value (resulting in the use of ZR).
//
// NOTE:
// Strictly speaking, this can also be trivially expanded to handle other sizes
// that happen to use any value that could fit inside a byte if the need
// arises. This does increase branching and code generation, however, since
// we'd still need to emit the fallback in the event a value for a larger size
// falls outside the range of a byte instead of only generating the MOPS code.
if (Is8Bit || Value == ARMEmitter::Reg::zr) {
// If we're performing a non-byte-sized zeroing operation then we need to
// scale the counter accordingly. (e.g. a 64-bit memset of size 2 needs to
// be turned into an 8-bit memset of size 16)
if (!Is8Bit) {
lsl(ARMEmitter::Size::i64Bit, TMP1, TMP1, FEXCore::ToUnderlying(SubRegSize));
}
// If backwards, then we need to adjust the starting address because
// set{p, m, e} memset forwards, so we need to slide this bad boy
// back like: (address - count) + 1.
//
// This lets us offset the address such that we can treat a backwards
// memset as if it were a forwards one.
if (IsBackwards) {
sub(TMP2, TMP2, TMP1);
add(ARMEmitter::Size::i64Bit, TMP2, TMP2, 1);
}
// Unfortunately set operations fiddle with NZCV, so we need to preserve it.
mrs(TMP3, ARMEmitter::SystemRegister::NZCV);
setp(TMP2, TMP1, Value.X());
setm(TMP2, TMP1, Value.X());
sete(TMP2, TMP1, Value.X());
msr(ARMEmitter::SystemRegister::NZCV, TMP3);
MakeFinalAddress();
(void)Bind(&DoneInternal);
return;
}
}
ARMEmitter::ForwardLabel AgainInternal256Exit {};
ARMEmitter::BackwardLabel AgainInternal256 {};
ARMEmitter::ForwardLabel AgainInternal128Exit {};
ARMEmitter::BackwardLabel AgainInternal128 {};
if (Direction == -1) {
if (IsBackwards) {
sub(ARMEmitter::Size::i64Bit, TMP2, TMP2, 32 - Size);
}
@@ -1986,39 +2045,23 @@ DEF_OP(MemSet) {
add(ARMEmitter::Size::i64Bit, TMP1, TMP1, 32 / Size);
(void)cbz(ARMEmitter::Size::i64Bit, TMP1, &DoneInternal);
if (Direction == -1) {
if (IsBackwards) {
add(ARMEmitter::Size::i64Bit, TMP2, TMP2, 32 - Size);
}
}
(void)Bind(&AgainInternal);
if (IsAtomic) {
MemStoreTSO(Value, OpSize, SizeDirection);
MemStoreTSO(Value, Size, SizeDirection);
} else {
MemStore(Value, OpSize, SizeDirection);
MemStore(Value, Size, SizeDirection);
}
sub(ARMEmitter::Size::i64Bit, TMP1, TMP1, 1);
(void)cbnz(ARMEmitter::Size::i64Bit, TMP1, &AgainInternal);
(void)Bind(&DoneInternal);
if (SizeDirection >= 0) {
switch (OpSize) {
case 1: add(Dst.X(), MemReg.X(), Length.X()); break;
case 2: add(Dst.X(), MemReg.X(), Length.X(), ARMEmitter::ShiftType::LSL, 1); break;
case 4: add(Dst.X(), MemReg.X(), Length.X(), ARMEmitter::ShiftType::LSL, 2); break;
case 8: add(Dst.X(), MemReg.X(), Length.X(), ARMEmitter::ShiftType::LSL, 3); break;
default: LOGMAN_MSG_A_FMT("Unhandled {} size: {}", __func__, OpSize); break;
}
} else {
switch (OpSize) {
case 1: sub(Dst.X(), MemReg.X(), Length.X()); break;
case 2: sub(Dst.X(), MemReg.X(), Length.X(), ARMEmitter::ShiftType::LSL, 1); break;
case 4: sub(Dst.X(), MemReg.X(), Length.X(), ARMEmitter::ShiftType::LSL, 2); break;
case 8: sub(Dst.X(), MemReg.X(), Length.X(), ARMEmitter::ShiftType::LSL, 3); break;
default: LOGMAN_MSG_A_FMT("Unhandled {} size: {}", __func__, OpSize); break;
}
}
MakeFinalAddress();
};
if (DirectionIsInline) {
@@ -2041,10 +2084,6 @@ DEF_OP(MemSet) {
}
DEF_OP(MemCpy) {
// TODO: A future looking task would be to support this with ARM's MOPS instructions.
// The 8-bit non-atomic path directly matches ARM's CPYP/CPYM/CPYE instruction,
//
// Assuming non-atomicity and non-faulting behaviour, this can accelerate this implementation.
const auto Op = IROp->C<IR::IROp_MemCpy>();
const bool IsAtomic = CTX->IsMemcpyAtomicTSOEnabled();
@@ -2175,8 +2214,40 @@ DEF_OP(MemCpy) {
};
auto EmitMemcpy = [&](int32_t Direction) {
const int32_t OpSize = Size;
const int32_t SizeDirection = Size * Direction;
const bool IsBackwards = Direction == -1;
const auto FinalizeAddresses = [&] {
if (IsBackwards) {
switch (Size) {
case 1:
sub(Dst0.X(), TMP1, TMP3);
sub(Dst1.X(), TMP2, TMP3);
break;
case 2:
case 4:
case 8:
sub(Dst0.X(), TMP1, TMP3, ARMEmitter::ShiftType::LSL, FEXCore::ilog2(Size));
sub(Dst1.X(), TMP2, TMP3, ARMEmitter::ShiftType::LSL, FEXCore::ilog2(Size));
break;
default: LOGMAN_MSG_A_FMT("Unhandled MemCpy size: {}", Size); break;
}
} else {
switch (Size) {
case 1:
add(Dst0.X(), TMP1, TMP3);
add(Dst1.X(), TMP2, TMP3);
break;
case 2:
case 4:
case 8:
add(Dst0.X(), TMP1, TMP3, ARMEmitter::ShiftType::LSL, FEXCore::ilog2(Size));
add(Dst1.X(), TMP2, TMP3, ARMEmitter::ShiftType::LSL, FEXCore::ilog2(Size));
break;
default: LOGMAN_MSG_A_FMT("Unhandled MemCpy size: {}", Size); break;
}
}
};
ARMEmitter::BiDirectionalLabel AgainInternal {};
ARMEmitter::ForwardLabel DoneInternal {};
@@ -2185,6 +2256,48 @@ DEF_OP(MemCpy) {
(void)cbz(ARMEmitter::Size::i64Bit, TMP1, &DoneInternal);
if (!IsAtomic) {
if (CTX->HostFeatures.SupportsMOPS) {
// In the event we have an overlap (gross), we need to fall back
// to the non-mops copy handler. Since the overlap check needs to
// make use of NZCV, we need to save it. This can be avoided with
// ARMv9.6+'s FEAT_CMPBR, but alas, we don't have access to that right now.
//
// NOTE: That we need to temporarily trash TMP1 and restore it after the
// comparison.
ARMEmitter::ForwardLabel OverlapCase;
mrs(TMP4, ARMEmitter::SystemRegister::NZCV);
sub(ARMEmitter::Size::i64Bit, TMP1, TMP2, TMP3);
cmp(ARMEmitter::Size::i64Bit, TMP1, Length.X());
mov(TMP1, Length.X());
(void)bc(ARMEmitter::Condition::CC_LT, &OverlapCase);
// If doing something larger than a byte copy, then we need to scale
// the counter value accordingly to convert it to bytes.
if (Size > 1) {
lsl(ARMEmitter::Size::i64Bit, TMP1, TMP1, FEXCore::ilog2(Size));
}
// Adjust addresses so that we treat the backward copy as a forward copy
if (IsBackwards) {
sub(ARMEmitter::Size::i64Bit, TMP2, TMP2, TMP1);
sub(ARMEmitter::Size::i64Bit, TMP3, TMP3, TMP1);
add(ARMEmitter::Size::i64Bit, TMP2, TMP2, Size);
add(ARMEmitter::Size::i64Bit, TMP3, TMP3, Size);
}
// Unfortunately copy operations fiddle with NZCV, so we need to preserve it.
cpyfp(TMP2, TMP3, TMP1);
cpyfm(TMP2, TMP3, TMP1);
cpyfe(TMP2, TMP3, TMP1);
msr(ARMEmitter::SystemRegister::NZCV, TMP4);
(void)b(&DoneInternal);
// Turns out we overlap and need to fall back. Make sure to restore NZCV.
(void)Bind(&OverlapCase);
msr(ARMEmitter::SystemRegister::NZCV, TMP4);
}
ARMEmitter::ForwardLabel AbsPos {};
ARMEmitter::ForwardLabel AgainInternal256Exit {};
ARMEmitter::ForwardLabel AgainInternal128Exit {};
@@ -2198,7 +2311,7 @@ DEF_OP(MemCpy) {
sub(ARMEmitter::Size::i64Bit, TMP4, TMP4, 32);
(void)tbnz(TMP4, 63, &AgainInternal);
if (Direction == -1) {
if (IsBackwards) {
sub(ARMEmitter::Size::i64Bit, TMP2, TMP2, 32 - Size);
sub(ARMEmitter::Size::i64Bit, TMP3, TMP3, 32 - Size);
}
@@ -2233,7 +2346,7 @@ DEF_OP(MemCpy) {
add(ARMEmitter::Size::i64Bit, TMP1, TMP1, 32 / Size);
(void)cbz(ARMEmitter::Size::i64Bit, TMP1, &DoneInternal);
if (Direction == -1) {
if (IsBackwards) {
add(ARMEmitter::Size::i64Bit, TMP2, TMP2, 32 - Size);
add(ARMEmitter::Size::i64Bit, TMP3, TMP3, 32 - Size);
}
@@ -2241,9 +2354,9 @@ DEF_OP(MemCpy) {
(void)Bind(&AgainInternal);
if (IsAtomic) {
MemCpyTSO(OpSize, SizeDirection);
MemCpyTSO(Size, SizeDirection);
} else {
MemCpy(OpSize, SizeDirection);
MemCpy(Size, SizeDirection);
}
sub(ARMEmitter::Size::i64Bit, TMP1, TMP1, 1);
(void)cbnz(ARMEmitter::Size::i64Bit, TMP1, &AgainInternal);
@@ -2255,54 +2368,14 @@ DEF_OP(MemCpy) {
mov(TMP2, MemRegSrc.X());
mov(TMP3, Length.X());
if (SizeDirection >= 0) {
switch (OpSize) {
case 1:
add(Dst0.X(), TMP1, TMP3);
add(Dst1.X(), TMP2, TMP3);
break;
case 2:
add(Dst0.X(), TMP1, TMP3, ARMEmitter::ShiftType::LSL, 1);
add(Dst1.X(), TMP2, TMP3, ARMEmitter::ShiftType::LSL, 1);
break;
case 4:
add(Dst0.X(), TMP1, TMP3, ARMEmitter::ShiftType::LSL, 2);
add(Dst1.X(), TMP2, TMP3, ARMEmitter::ShiftType::LSL, 2);
break;
case 8:
add(Dst0.X(), TMP1, TMP3, ARMEmitter::ShiftType::LSL, 3);
add(Dst1.X(), TMP2, TMP3, ARMEmitter::ShiftType::LSL, 3);
break;
default: LOGMAN_MSG_A_FMT("Unhandled {} size: {}", __func__, OpSize); break;
}
} else {
switch (OpSize) {
case 1:
sub(Dst0.X(), TMP1, TMP3);
sub(Dst1.X(), TMP2, TMP3);
break;
case 2:
sub(Dst0.X(), TMP1, TMP3, ARMEmitter::ShiftType::LSL, 1);
sub(Dst1.X(), TMP2, TMP3, ARMEmitter::ShiftType::LSL, 1);
break;
case 4:
sub(Dst0.X(), TMP1, TMP3, ARMEmitter::ShiftType::LSL, 2);
sub(Dst1.X(), TMP2, TMP3, ARMEmitter::ShiftType::LSL, 2);
break;
case 8:
sub(Dst0.X(), TMP1, TMP3, ARMEmitter::ShiftType::LSL, 3);
sub(Dst1.X(), TMP2, TMP3, ARMEmitter::ShiftType::LSL, 3);
break;
default: LOGMAN_MSG_A_FMT("Unhandled {} size: {}", __func__, OpSize); break;
}
}
FinalizeAddresses();
};
if (DirectionIsInline) {
LOGMAN_THROW_A_FMT(DirectionConstant == 1 || DirectionConstant == -1, "unexpected direction");
EmitMemcpy(DirectionConstant);
} else {
// Emit forward direction memset then backward direction memset.
// Emit forward direction memcpy then backward direction memcpy.
for (int32_t Direction : {1, -1}) {
EmitMemcpy(Direction);
if (Direction == 1) {
+30 -2
View File
@@ -210,6 +210,25 @@ DEF_OP(Print) {
PopDynamicRegs();
}
DEF_OP(PrintMsg) {
auto Op = IROp->C<IR::IROp_PrintMsg>();
PushDynamicRegs(TMP1);
SpillStaticRegs(TMP1);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, reinterpret_cast<uintptr_t>(Op->Value));
ldr(ARMEmitter::XReg::x1, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.PrintMsgValue));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<void, uint64_t>(ARMEmitter::Reg::r1);
} else {
blr(ARMEmitter::Reg::r1);
}
FillStaticRegs();
PopDynamicRegs();
}
DEF_OP(ProcessorID) {
if (CTX->HostFeatures.SupportsCPUIndexInTPIDRRO) {
mrs(GetReg(Node), ARMEmitter::SystemRegister::TPIDRRO_EL0);
@@ -227,7 +246,10 @@ DEF_OP(ProcessorID) {
// Ordering is incredibly important here
// We must spill any overlapping registers first THEN claim we are in a syscall without invalidating state at all
// Only spill the registers that intersect with our usage
SpillStaticRegs(TMP1, false, SpillMask);
SpillStaticRegs(TMP1, {
.GPRSpillMask = SpillMask,
.FPRs = false,
});
// Now that we are spilled, store in the state that we are in a syscall
// Still without overwriting registers that matter
@@ -264,7 +286,13 @@ DEF_OP(ProcessorID) {
// Now that we are done in the syscall we need to carefully peel back the state
// First unspill the registers from before
FillStaticRegs(false, SpillMask, ~0U, ARMEmitter::Reg::r8, ARMEmitter::Reg::r2);
FillStaticRegs({
.OptionalReg = ARMEmitter::Reg::r8,
.OptionalReg2 = ARMEmitter::Reg::r2,
.GPRFillMask = SpillMask,
.FPRs = false,
});
// Now the registers we've spilled are back in their original host registers
// We can safely claim we are no longer in a syscall
+112 -11
View File
@@ -977,7 +977,7 @@ DEF_OP(LoadNamedVectorConstant) {
}
// Load the pointer.
auto GenerateMemOperand = [this](IR::OpSize OpSize, uint32_t NamedConstant, ARMEmitter::Register Base) {
const auto ConstantOffset = offsetof(FEXCore::Core::CpuStateFrame, Pointers.NamedVectorConstants[NamedConstant]);
const auto ConstantOffset = ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, Pointers.NamedVectorConstants, NamedConstant);
if (ConstantOffset <= 255 || // Unscaled 9-bit signed
((ConstantOffset & (IR::OpSizeToSize(OpSize) - 1)) == 0 &&
@@ -985,13 +985,13 @@ DEF_OP(LoadNamedVectorConstant) {
return ARMEmitter::ExtendedMemOperand(Base.X(), ARMEmitter::IndexType::OFFSET, ConstantOffset);
}
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.NamedVectorConstantPointers[NamedConstant]));
ldr(TMP1, STATE_PTR_IDX(CpuStateFrame, Pointers.NamedVectorConstantPointers, NamedConstant));
return ARMEmitter::ExtendedMemOperand(TMP1, ARMEmitter::IndexType::OFFSET, 0);
};
if (OpSize == IR::OpSize::i256Bit) {
// Handle SVE 32-byte variant upfront.
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.NamedVectorConstantPointers[Op->Constant]));
ldr(TMP1, STATE_PTR_IDX(CpuStateFrame, Pointers.NamedVectorConstantPointers, Op->Constant));
ld1b<ARMEmitter::SubRegSize::i8Bit>(Dst.Z(), PRED_TMP_32B.Zeroing(), TMP1, 0);
return;
}
@@ -1013,7 +1013,7 @@ DEF_OP(LoadNamedVectorIndexedConstant) {
const auto Dst = GetVReg(Node);
// Load the pointer.
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.IndexedNamedVectorConstantPointers[Op->Constant]));
ldr(TMP1, STATE_PTR_IDX(CpuStateFrame, Pointers.IndexedNamedVectorConstantPointers, Op->Constant));
switch (OpSize) {
case IR::OpSize::i8Bit: ldrb(Dst, TMP1, Op->Index); break;
@@ -1434,8 +1434,8 @@ DEF_OP(VFMin) {
bif(Dst.Q(), Vector2.Q(), VTMP1.Q());
} else if (Dst == Vector2) {
// Destination is already Vector2, Invert arguments and insert Vector1 on false.
fcmgt(SubRegSize, VTMP1.Q(), Vector1.Q(), Vector2.Q());
bif(Dst.Q(), Vector1.Q(), VTMP1.Q());
fcmgt(SubRegSize, VTMP1.Q(), Vector2.Q(), Vector1.Q());
bit(Dst.Q(), Vector1.Q(), VTMP1.Q());
} else {
// Dst is not either source, need a move.
fcmgt(SubRegSize, VTMP1.Q(), Vector2.Q(), Vector1.Q());
@@ -1466,7 +1466,8 @@ DEF_OP(VFMax) {
const auto Mask = PRED_TMP_32B;
const auto ComparePred = ARMEmitter::PReg::p0;
fcmgt(SubRegSize, ComparePred, Mask.Zeroing(), Vector2.Z(), Vector1.Z());
fcmgt(SubRegSize, ComparePred, Mask.Zeroing(), Vector1.Z(), Vector2.Z());
not_(ComparePred, Mask.Zeroing(), ComparePred);
if (Dst == Vector1) {
// Trivial case where Vector1 is also the destination.
@@ -1488,17 +1489,17 @@ DEF_OP(VFMax) {
if (Dst == Vector1) {
// Destination is already Vector1, need to insert Vector2 on true.
fcmgt(SubRegSize, VTMP1.Q(), Vector2.Q(), Vector1.Q());
bit(Dst.Q(), Vector2.Q(), VTMP1.Q());
fcmgt(SubRegSize, VTMP1.Q(), Vector1.Q(), Vector2.Q());
bif(Dst.Q(), Vector2.Q(), VTMP1.Q());
} else if (Dst == Vector2) {
// Destination is already Vector2, Invert arguments and insert Vector1 on true.
fcmgt(SubRegSize, VTMP1.Q(), Vector1.Q(), Vector2.Q());
bit(Dst.Q(), Vector1.Q(), VTMP1.Q());
} else {
// Dst is not either source, need a move.
fcmgt(SubRegSize, VTMP1.Q(), Vector2.Q(), Vector1.Q());
fcmgt(SubRegSize, VTMP1.Q(), Vector1.Q(), Vector2.Q());
mov(Dst.Q(), Vector1.Q());
bit(Dst.Q(), Vector2.Q(), VTMP1.Q());
bif(Dst.Q(), Vector2.Q(), VTMP1.Q());
}
}
}
@@ -4611,4 +4612,104 @@ DEF_OP(VFCopySign) {
}
}
DEF_OP(F64SIN) {
const auto Op = IROp->C<IR::IROp_F64SIN>();
const auto Src = GetVReg(Op->Src);
const auto Dst = GetVReg(Node);
fmov(VTMP1.D(), Src.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.F64SinHandler));
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
fmov(Dst.D(), VTMP1.D());
}
DEF_OP(F64COS) {
const auto Op = IROp->C<IR::IROp_F64COS>();
const auto Src = GetVReg(Op->Src);
const auto Dst = GetVReg(Node);
fmov(VTMP1.D(), Src.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.F64CosHandler));
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
fmov(Dst.D(), VTMP1.D());
}
DEF_OP(F64TAN) {
const auto Op = IROp->C<IR::IROp_F64TAN>();
const auto Src = GetVReg(Op->Src);
const auto Dst = GetVReg(Node);
fmov(VTMP1.D(), Src.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.F64TanHandler));
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
fmov(Dst.D(), VTMP1.D());
}
// Src1=y(ST1), Src2=x(ST0). Marshal into VTMP1/VTMP2 and dispatch the shared handler.
DEF_OP(F64ATAN) {
const auto Op = IROp->C<IR::IROp_F64ATAN>();
const auto Src1 = GetVReg(Op->Src1);
const auto Src2 = GetVReg(Op->Src2);
const auto Dst = GetVReg(Node);
fmov(VTMP1.D(), Src1.D());
fmov(VTMP2.D(), Src2.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.F64AtanHandler));
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
fmov(Dst.D(), VTMP1.D());
}
// Src=x(ST0), Src2=y(ST1). Marshal into VTMP1/VTMP2 and dispatch the shared handler.
DEF_OP(F64FYL2X) {
const auto Op = IROp->C<IR::IROp_F64FYL2X>();
const auto Src = GetVReg(Op->Src);
const auto Src2 = GetVReg(Op->Src2);
const auto Dst = GetVReg(Node);
fmov(VTMP1.D(), Src.D());
fmov(VTMP2.D(), Src2.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.F64FYL2XHandler));
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
fmov(Dst.D(), VTMP1.D());
}
DEF_OP(F64SCALE) {
const auto Op = IROp->C<IR::IROp_F64SCALE>();
const auto Src1 = GetVReg(Op->Src1);
const auto Src2 = GetVReg(Op->Src2);
const auto Dst = GetVReg(Node);
fmov(VTMP1.D(), Src1.D());
fmov(VTMP2.D(), Src2.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.F64ScaleHandler));
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
fmov(Dst.D(), VTMP1.D());
}
DEF_OP(F64F2XM1) {
const auto Op = IROp->C<IR::IROp_F64F2XM1>();
const auto Src = GetVReg(Op->Src);
const auto Dst = GetVReg(Node);
fmov(VTMP1.D(), Src.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.F64F2XM1Handler));
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
fmov(Dst.D(), VTMP1.D());
}
} // namespace FEXCore::CPU
@@ -41,6 +41,9 @@ LookupCache::LookupCache(FEXCore::Context::ContextImpl* CTX)
PagePointer = reinterpret_cast<uintptr_t>(FEXCore::Allocator::VirtualAlloc(TotalCacheSize, false, false));
LOGMAN_THROW_A_FMT(PagePointer != -1ULL, "Failed to allocate PagePointer");
// Disable THP on the Lookup cache.
FEXCore::Allocator::VirtualTHPControl(reinterpret_cast<const void*>(PagePointer), TotalCacheSize, FEXCore::Allocator::THPControl::Disable);
FEXCore::Allocator::VirtualName("FEXMem_Lookup", reinterpret_cast<void*>(PagePointer),
ctx->Config.VirtualMemSize / FEXCore::Utils::FEX_PAGE_SIZE * 8 + CODE_SIZE);
CTX->SyscallHandler->MarkOvercommitRange(PagePointer, TotalCacheSize);
+1 -1
View File
@@ -3,7 +3,7 @@
#include "Interface/Context/Context.h"
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/SHMStats.h>
#include "Utils/WritePriorityMutex.h"
#include <FEXCore/Utils/WritePriorityMutex.h>
#include <FEXCore/fextl/map.h>
#include <FEXCore/fextl/memory_resource.h>
@@ -2643,7 +2643,10 @@ void OpDispatchBuilder::IMULOp(OpcodeArgs) {
}
// 64-bit special cased to save a move
Ref Result = Size < OpSize::i64Bit ? _Mul(OpSize::i64Bit, Src1, Src2) : nullptr;
Ref Result {};
if (Size < OpSize::i64Bit) {
Result = _Mul(OpSize::i64Bit, Src1, Src2);
}
Ref ResultHigh {};
if (Size == OpSize::i8Bit) {
// Result is stored in AX
@@ -4487,8 +4490,6 @@ void OpDispatchBuilder::StoreResult(RegClass Class, X86Tables::DecodedOp Op, Ref
OpDispatchBuilder::OpDispatchBuilder(FEXCore::Context::ContextImpl* ctx)
: IREmitter {ctx->OpDispatcherAllocator, ctx->HostFeatures.SupportsTSOImm9}
, CTX {ctx} {
ResetWorkingList();
if (CTX->HostFeatures.SupportsAVX && CTX->HostFeatures.SupportsSVE256) {
SaveAVXStateFunc = &OpDispatchBuilder::SaveAVXState;
RestoreAVXStateFunc = &OpDispatchBuilder::RestoreAVXState;
@@ -4501,7 +4502,8 @@ OpDispatchBuilder::OpDispatchBuilder(FEXCore::Context::ContextImpl* ctx)
}
void OpDispatchBuilder::ResetWorkingList() {
IREmitter::ResetWorkingList();
IREmitter::ReownOrClaimBuffer();
JumpTargets.clear();
BlockSetRIP = false;
DecodeFailure = false;
@@ -4887,6 +4889,11 @@ void OpDispatchBuilder::CLZeroOp(OpcodeArgs) {
}
void OpDispatchBuilder::Prefetch(OpcodeArgs, bool ForStore, bool Stream, uint8_t Level) {
if (Op->Src[0].IsGPR()) {
// NOP instance.
return;
}
Ref DestMem = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.LoadData = false});
_Prefetch(ForStore, Stream, Level, DestMem, Invalid(), MemOffsetType::SXTX, 1);
}
@@ -305,7 +305,9 @@ public:
OpDispatchBuilder(FEXCore::Context::ContextImpl* ctx);
// Should only be called at the start of IR Emission.
void ResetWorkingList();
void ResetDecodeFailure() {
NeedsBlockEnd = DecodeFailure = false;
}
@@ -1666,7 +1668,7 @@ private:
[[nodiscard]]
static uint32_t GPROffset(X86State::X86Reg reg) {
LOGMAN_THROW_A_FMT(reg <= X86State::X86Reg::REG_R15, "Invalid reg used");
return static_cast<uint32_t>(offsetof(Core::CPUState, gregs[static_cast<size_t>(reg)]));
return static_cast<uint32_t>(ARRAY_OFFSETOF(Core::CPUState, gregs, reg));
}
[[nodiscard]]
@@ -1885,15 +1887,15 @@ private:
// For DF, we need to transform 0/1 into 1/-1
StoreDF(_SubShift(OpSize::i64Bit, Constant(1), Value, ShiftType::LSL, 1));
} else if (BitOffset == FEXCore::X86State::RFLAG_TF_RAW_LOC) {
auto PackedTF = _LoadContextGPR(OpSize::i8Bit, offsetof(FEXCore::Core::CPUState, flags[BitOffset]));
auto PackedTF = _LoadContextGPR(OpSize::i8Bit, ARRAY_OFFSETOF(FEXCore::Core::CPUState, flags, BitOffset));
// An exception should still be raised after an instruction that unsets TF, leave the unblocked bit set but unset
// the TF bit to cause such behaviour. The handling code at the start of the next block will then unset the
// unblocked bit before raising the exception.
auto NewPackedTF =
_Select(OpSize::i64Bit, OpSize::i64Bit, CondClass::EQ, Value, Constant(0), _And(OpSize::i32Bit, PackedTF, Constant(~1)), Constant(1));
_StoreContextGPR(OpSize::i8Bit, NewPackedTF, offsetof(FEXCore::Core::CPUState, flags[BitOffset]));
_StoreContextGPR(OpSize::i8Bit, NewPackedTF, ARRAY_OFFSETOF(FEXCore::Core::CPUState, flags, BitOffset));
} else {
_StoreContextGPR(OpSize::i8Bit, Value, offsetof(FEXCore::Core::CPUState, flags[BitOffset]));
_StoreContextGPR(OpSize::i8Bit, Value, ARRAY_OFFSETOF(FEXCore::Core::CPUState, flags, BitOffset));
}
}
@@ -1948,8 +1950,8 @@ private:
[[nodiscard]]
static uint32_t CacheIndexToContextOffset(int Index) {
switch (Index) {
case MM0Index ... MM7Index: return offsetof(FEXCore::Core::CPUState, mm[Index - MM0Index]);
case AVXHigh0Index ... AVXHigh15Index: return offsetof(FEXCore::Core::CPUState, avx_high[Index - AVXHigh0Index][0]);
case MM0Index ... MM7Index: return ARRAY_OFFSETOF(FEXCore::Core::CPUState, mm, Index - MM0Index);
case AVXHigh0Index ... AVXHigh15Index: return ARRAY_OFFSETOF(FEXCore::Core::CPUState, avx_high, Index - AVXHigh0Index);
default: return ~0U;
}
}
@@ -2149,7 +2151,7 @@ private:
// Recover the sign bit, it is the logical DF value
return _Lshr(OpSize::i64Bit, LoadDF(), Constant(63));
} else {
return _LoadContextGPR(OpSize::i8Bit, offsetof(Core::CPUState, flags[BitOffset]));
return _LoadContextGPR(OpSize::i8Bit, ARRAY_OFFSETOF(Core::CPUState, flags, BitOffset));
}
}
@@ -1011,10 +1011,522 @@ Ref OpDispatchBuilder::PShufWLane(IR::OpSize Size, FEXCore::IR::IndexNamedVector
}
void OpDispatchBuilder::PSHUFW8ByteOp(OpcodeArgs) {
uint16_t Shuffle = Op->Src[1].Data.Literal.Value;
uint8_t Shuffle = Op->Src[1].Data.Literal.Value;
const auto Size = OpSizeFromSrc(Op);
const auto TBLIndex = FEXCore::IR::INDEXED_NAMED_VECTOR_PSHUFLW;
Ref Src = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
Ref Dest = PShufWLane(Size, FEXCore::IR::INDEXED_NAMED_VECTOR_PSHUFLW, true, Src, Shuffle);
// Single MMX 64-bit shuffle. Shuffle selector can fit full selection.
Ref Dest {};
switch (Shuffle) {
// Single-instruction shuffle operations.
case 0b00'00'00'00: Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 0); break;
case 0b00'10'01'00: Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 0, Src, Src); break;
case 0b01'00'01'00: Dest = _VZip(OpSize::i64Bit, OpSize::i32Bit, Src, Src); break;
case 0b00'11'10'01: Dest = _VExtr(OpSize::i64Bit, OpSize::i16Bit, Src, Src, 1); break;
case 0b01'00'11'10: Dest = _VRev64(OpSize::i64Bit, OpSize::i32Bit, Src); break;
case 0b01'01'01'01: Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 1); break;
case 0b01'10'01'00: Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 1, Src, Src); break;
case 0b10'10'01'00: Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 2, Src, Src); break;
case 0b10'10'10'10: Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 2); break;
case 0b11'00'01'00: Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 0, Src, Src); break;
case 0b11'01'01'00: Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 1, Src, Src); break;
case 0b11'10'00'00: Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 0, Src, Src); break;
case 0b11'10'01'00: Dest = Src; break;
case 0b11'10'01'01: Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 1, Src, Src); break;
case 0b11'10'01'10: Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 2, Src, Src); break;
case 0b11'10'01'11: Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 3, Src, Src); break;
case 0b11'10'10'00: Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 2, Src, Src); break;
case 0b11'10'11'00: Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 3, Src, Src); break;
case 0b11'10'11'10: Dest = _VDupElement(OpSize::i64Bit, OpSize::i32Bit, Src, 1); break;
case 0b11'11'01'00: Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 3, Src, Src); break;
case 0b11'11'11'11: Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 3); break;
// Two instruction shuffle operations.
case 0b00'00'00'01:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 0);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 1, Dest, Src);
break;
case 0b00'00'00'10:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 0);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 2, Dest, Src);
break;
case 0b00'00'00'11:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 0);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 3, Dest, Src);
break;
case 0b00'00'01'00:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 0);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 1, Dest, Src);
break;
case 0b00'00'10'00:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 0);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 2, Dest, Src);
break;
case 0b00'00'11'00:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 0);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 3, Dest, Src);
break;
case 0b00'00'11'10:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 0);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i32Bit, 0, 1, Dest, Src);
break;
case 0b00'01'00'00:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 0);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 1, Dest, Src);
break;
case 0b00'01'00'01:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i32Bit, Src, 0);
Dest = _VExtr(OpSize::i64Bit, OpSize::i16Bit, Dest, Dest, 1);
break;
case 0b00'01'00'11:
Dest = _VZip(OpSize::i64Bit, OpSize::i32Bit, Src, Src);
Dest = _VExtr(OpSize::i64Bit, OpSize::i8Bit, Dest, Src, 6);
break;
case 0b00'01'01'00:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 1, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 0, Dest, Src);
break;
case 0b00'01'01'01:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 1);
Dest = _VExtr(OpSize::i64Bit, OpSize::i16Bit, Src, Dest, 1);
break;
case 0b00'10'00'00:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 0);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 2, Dest, Src);
break;
case 0b00'10'00'10:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 0, Src, Src);
Dest = _VDupElement(OpSize::i64Bit, OpSize::i32Bit, Dest, 1);
break;
case 0b00'10'01'01:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 1, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 0, Dest, Src);
break;
case 0b00'10'01'10:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 2, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 0, Dest, Src);
break;
case 0b00'10'01'11:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 3, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 0, Dest, Src);
break;
case 0b00'10'10'00:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 2, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 0, Dest, Src);
break;
case 0b00'10'10'10:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 2);
Dest = _VExtr(OpSize::i64Bit, OpSize::i16Bit, Src, Dest, 1);
break;
case 0b00'10'11'00:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 3, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 0, Dest, Src);
break;
case 0b00'10'11'10:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i32Bit, Src, 1);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 0, Dest, Src);
break;
case 0b00'11'00'00:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 0);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 3, Dest, Src);
break;
case 0b00'11'01'00:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 3, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 0, Dest, Src);
break;
case 0b00'11'10'11:
Dest = _VZip2(OpSize::i64Bit, OpSize::i32Bit, Src, Src);
Dest = _VExtr(OpSize::i64Bit, OpSize::i16Bit, Src, Dest, 1);
break;
case 0b00'11'11'11:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 3);
Dest = _VExtr(OpSize::i64Bit, OpSize::i16Bit, Src, Dest, 1);
break;
case 0b01'00'00'00:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 0);
Dest = _VZip(OpSize::i64Bit, OpSize::i32Bit, Dest, Src);
break;
case 0b01'00'00'10:
Dest = _VRev64(OpSize::i64Bit, OpSize::i32Bit, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 0, Dest, Src);
break;
case 0b01'00'01'01:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 1);
Dest = _VZip(OpSize::i64Bit, OpSize::i32Bit, Dest, Src);
break;
case 0b01'00'01'10:
Dest = _VZip(OpSize::i64Bit, OpSize::i32Bit, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 2, Dest, Src);
break;
case 0b01'00'01'11:
Dest = _VZip(OpSize::i64Bit, OpSize::i32Bit, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 3, Dest, Src);
break;
case 0b01'00'10'00:
Dest = _VZip(OpSize::i64Bit, OpSize::i32Bit, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 2, Dest, Src);
break;
case 0b01'00'10'01:
Dest = _VExtr(OpSize::i64Bit, OpSize::i8Bit, Src, Src, 2);
Dest = _VZip(OpSize::i64Bit, OpSize::i32Bit, Dest, Src);
break;
case 0b01'00'10'10:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 2);
Dest = _VZip(OpSize::i64Bit, OpSize::i32Bit, Dest, Src);
break;
case 0b01'00'11'00:
Dest = _VZip(OpSize::i64Bit, OpSize::i32Bit, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 3, Dest, Src);
break;
case 0b01'00'11'01:
Dest = _VRev64(OpSize::i64Bit, OpSize::i32Bit, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 1, Dest, Src);
break;
case 0b01'00'11'11:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 3);
Dest = _VZip(OpSize::i64Bit, OpSize::i32Bit, Dest, Src);
break;
case 0b01'01'00'01:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 1);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 0, Dest, Src);
break;
case 0b01'01'01'00:
Dest = _VZip(OpSize::i64Bit, OpSize::i32Bit, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 1, Dest, Src);
break;
case 0b01'01'01'10:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 1);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 2, Dest, Src);
break;
case 0b01'01'01'11:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 1);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 3, Dest, Src);
break;
case 0b01'01'10'01:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 1);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 2, Dest, Src);
break;
case 0b01'01'11'01:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 1);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 3, Dest, Src);
break;
case 0b01'01'11'10:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 1);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i32Bit, 0, 1, Dest, Src);
break;
case 0b01'10'00'00:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 0, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 1, Dest, Src);
break;
case 0b01'10'01'01:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 1);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 2, Dest, Src);
break;
case 0b01'10'01'10:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 2, Src, Src);
Dest = _VZip(OpSize::i64Bit, OpSize::i32Bit, Dest, Dest);
break;
case 0b01'10'01'11:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 3, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 1, Dest, Src);
break;
case 0b01'10'10'00:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 2, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 1, Dest, Src);
break;
case 0b01'10'10'10:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 2);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 1, Dest, Src);
break;
case 0b01'10'11'00:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 3, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 1, Dest, Src);
break;
case 0b01'10'11'10:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i32Bit, Src, 1);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 1, Dest, Src);
break;
case 0b01'11'01'00:
Dest = _VZip(OpSize::i64Bit, OpSize::i32Bit, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 3, Dest, Src);
break;
case 0b01'11'01'01:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 1);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 3, Dest, Src);
break;
case 0b01'11'01'11:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 3, Src, Src);
Dest = _VZip(OpSize::i64Bit, OpSize::i32Bit, Dest, Dest);
break;
case 0b01'11'10'01:
Dest = _VExtr(OpSize::i64Bit, OpSize::i8Bit, Src, Src, 2);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 0, Dest, Dest);
break;
case 0b01'11'11'10:
Dest = _VRev64(OpSize::i64Bit, OpSize::i32Bit, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 1, Dest, Dest);
break;
case 0b01'11'11'11:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 3);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 1, Dest, Src);
break;
case 0b10'00'00'00:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 0);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 2, Dest, Src);
break;
case 0b10'00'01'00:
Dest = _VZip(OpSize::i64Bit, OpSize::i32Bit, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 2, Dest, Src);
break;
case 0b10'00'10'00:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 2, Src, Src);
Dest = _VZip(OpSize::i64Bit, OpSize::i32Bit, Dest, Dest);
break;
case 0b10'00'10'10:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 2);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 0, Dest, Src);
break;
case 0b10'00'11'10:
Dest = _VRev64(OpSize::i64Bit, OpSize::i32Bit, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 0, Dest, Dest);
break;
case 0b10'01'00'00:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 0);
Dest = _VExtr(OpSize::i64Bit, OpSize::i8Bit, Src, Dest, 6);
break;
case 0b10'01'00'01:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 1);
Dest = _VExtr(OpSize::i64Bit, OpSize::i8Bit, Src, Dest, 6);
break;
case 0b10'01'00'10:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 2);
Dest = _VExtr(OpSize::i64Bit, OpSize::i8Bit, Src, Dest, 6);
break;
case 0b10'01'00'11:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 3);
Dest = _VExtr(OpSize::i64Bit, OpSize::i8Bit, Src, Dest, 6);
break;
case 0b10'01'01'00:
Dest = _VExtr(OpSize::i64Bit, OpSize::i8Bit, Src, Src, 6);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i32Bit, 0, 0, Dest, Src);
break;
case 0b10'01'01'01:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 1);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 2, Dest, Src);
break;
case 0b10'01'10'01:
Dest = _VExtr(OpSize::i64Bit, OpSize::i8Bit, Src, Src, 2);
Dest = _VZip(OpSize::i64Bit, OpSize::i32Bit, Dest, Dest);
break;
case 0b10'01'10'10:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 2);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 1, Dest, Src);
break;
case 0b10'01'11'10:
Dest = _VExtr(OpSize::i64Bit, OpSize::i8Bit, Src, Src, 6);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i32Bit, 0, 1, Dest, Src);
break;
case 0b10'10'00'00:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 0, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 2, Dest, Src);
break;
case 0b10'10'00'10:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 2);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 0, Dest, Src);
break;
case 0b10'10'01'01:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 1, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 2, Dest, Src);
break;
case 0b10'10'01'10:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 2);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 1, Dest, Src);
break;
case 0b10'10'01'11:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 3, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 2, Dest, Src);
break;
case 0b10'10'10'00:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 2);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 0, Dest, Src);
break;
case 0b10'10'10'01:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 2);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 1, Dest, Src);
break;
case 0b10'10'10'11:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 2);
Dest = _VExtr(OpSize::i64Bit, OpSize::i8Bit, Dest, Src, 6);
break;
case 0b10'10'11'00:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 3, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 2, Dest, Src);
break;
case 0b10'10'11'10:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 2);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 3, Dest, Src);
break;
case 0b10'11'01'00:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 3, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 2, Dest, Src);
break;
case 0b10'11'10'10:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 2);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 3, Dest, Src);
break;
case 0b10'11'10'11:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i32Bit, Src, 1);
Dest = _VExtr(OpSize::i64Bit, OpSize::i16Bit, Dest, Dest, 1);
break;
case 0b10'11'11'11:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 3);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 2, Dest, Src);
break;
case 0b11'00'00'00:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 0);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 3, Dest, Src);
break;
case 0b11'00'01'01:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 1, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 0, Dest, Src);
break;
case 0b11'00'01'10:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 2, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 0, Dest, Src);
break;
case 0b11'00'01'11:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 3, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 0, Dest, Src);
break;
case 0b11'00'10'00:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 2, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 0, Dest, Src);
break;
case 0b11'00'11'00:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 3, Src, Src);
Dest = _VZip(OpSize::i64Bit, OpSize::i32Bit, Dest, Dest);
break;
case 0b11'00'11'10:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i32Bit, Src, 1);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 0, Dest, Src);
break;
case 0b11'00'11'11:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 3);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 0, Dest, Src);
break;
case 0b11'01'00'00:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 0, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 1, Dest, Src);
break;
case 0b11'01'01'01:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 1);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 3, Dest, Src);
break;
case 0b11'01'01'10:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 2, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 1, Dest, Src);
break;
case 0b11'01'01'11:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 3, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 1, Dest, Src);
break;
case 0b11'01'10'00:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 2, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 1, Dest, Src);
break;
case 0b11'01'11'00:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 3, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 1, Dest, Src);
break;
case 0b11'01'11'01:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 1, Src, Src);
Dest = _VDupElement(OpSize::i64Bit, OpSize::i32Bit, Dest, 1);
break;
case 0b11'01'11'10:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i32Bit, Src, 1);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 1, Dest, Src);
break;
case 0b11'01'11'11:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 3);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 1, Dest, Src);
break;
case 0b11'10'00'01:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 1, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 0, Dest, Src);
break;
case 0b11'10'00'10:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i32Bit, Src, 1);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 0, Dest, Src);
break;
case 0b11'10'00'11:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 3, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 0, Dest, Src);
break;
case 0b11'10'10'01:
Dest = _VExtr(OpSize::i64Bit, OpSize::i8Bit, Src, Src, 2);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i32Bit, 1, 1, Dest, Src);
break;
case 0b11'10'10'10:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 2);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 3, Dest, Src);
break;
case 0b11'10'10'11:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 3, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 2, Dest, Src);
break;
case 0b11'10'11'01:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i32Bit, Src, 1);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 1, Dest, Src);
break;
case 0b11'10'11'11:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 3);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 2, Dest, Src);
break;
case 0b11'11'00'00:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 0, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 3, Dest, Src);
break;
case 0b11'11'00'11:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 3);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 0, Dest, Src);
break;
case 0b11'11'01'01:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 1, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 3, Dest, Src);
break;
case 0b11'11'01'10:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 2, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 3, Dest, Src);
break;
case 0b11'11'01'11:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 3);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 1, Dest, Src);
break;
case 0b11'11'10'00:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 2, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 3, Dest, Src);
break;
case 0b11'11'10'11:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 3);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 2, Dest, Src);
break;
case 0b11'11'11'00:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 3);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 0, Dest, Src);
break;
case 0b11'11'11'01:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 3);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 1, Dest, Src);
break;
case 0b11'11'11'10:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 3);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 2, Dest, Src);
break;
default:
auto LookupIndexes = LoadAndCacheIndexedNamedVectorConstant(Size, TBLIndex, Shuffle * 16);
Dest = _VTBL1(Size, Src, LookupIndexes);
break;
}
StoreResultFPR(Op, Dest);
}
@@ -1347,6 +1859,11 @@ Ref OpDispatchBuilder::SHUFOpImpl(OpcodeArgs, IR::OpSize DstSize, IR::OpSize Ele
Shuffle >>= ShiftAmount;
}
} else {
if (Src1 == Src2 && Shuffle == 0) {
// TODO: We can optimize significantly more shuffles when we know the sources match.
// Special case broadcast element 0.
return _VDupElement(DstSize, ElementSize, Src1, Shuffle & SelectionMask);
}
if (ElementSize == OpSize::i32Bit) {
// We can shuffle optimally in a lot of cases.
// TODO: We can optimize more of these cases.
@@ -2756,7 +3273,7 @@ void OpDispatchBuilder::SaveSSEState(Ref MemBase) {
void OpDispatchBuilder::SaveMXCSRState(Ref MemBase) {
// Store MXCSR and the mask for all bits.
_StoreMemPairGPR(OpSize::i32Bit, GetMXCSR(), Constant(0xFFFF), MemBase, 24);
_StoreMemPairGPR(OpSize::i32Bit, GetMXCSR(), Constant(0xFFC0), MemBase, 24);
}
void OpDispatchBuilder::SaveAVXState(Ref MemBase) {
@@ -163,7 +163,7 @@ void OpDispatchBuilder::FIST(OpcodeArgs, bool Truncate) {
// Check for NaN/Infinity: exponent = 0x7fff
SaveNZCV();
_TestNZ(OpSize::i64Bit, Exponent, Constant(0x7fff));
SubWithFlags(OpSize::i64Bit, Exponent, 0x7fff);
Ref IsSpecial = _NZCVSelect01(CondClass::EQ);
// For overflow detection, check if exponent indicates a value >= 2^15
@@ -178,7 +178,8 @@ void OpDispatchBuilder::FIST(OpcodeArgs, bool Truncate) {
Data = _F80CVTInt(Size, Data, Truncate);
StoreResultGPR_WithOpSize(Op, Op->Dest, Data, Size, OpSize::i8Bit);
StoreResultGPR_WithOpSize(Op, Op->Dest, Data, Size, OpSize::i8Bit,
CTX->IsVectorAtomicTSOEnabled() ? MemoryAccessType::DEFAULT : MemoryAccessType::NONTSO);
if ((Op->TableInfo->Flags & X86Tables::InstFlags::FLAGS_POP) != 0) {
_PopStackDestroy();
@@ -106,12 +106,24 @@ void OpDispatchBuilder::FISTF64(OpcodeArgs, bool Truncate) {
const auto Size = OpSizeFromSrc(Op);
Ref data = _ReadStackValue(0);
if (Truncate) {
data = _Float_ToGPR_ZS(Size == OpSize::i32Bit ? OpSize::i32Bit : OpSize::i64Bit, OpSize::i64Bit, data);
bool CanUseFloatReg = Size == OpSize::i64Bit;
if (CanUseFloatReg) {
// If possible, it's faster to keep the data in an FPR than doing a GPR transfer.
if (Truncate) {
data = _Vector_FToZS(OpSize::i128Bit, OpSize::i64Bit, data);
} else {
data = _Vector_FToS(OpSize::i128Bit, OpSize::i64Bit, data);
}
StoreResultFPR_WithOpSize(Op, Op->Dest, data, OpSize::i64Bit, OpSize::i8Bit);
} else {
data = _Float_ToGPR_S(Size == OpSize::i32Bit ? OpSize::i32Bit : OpSize::i64Bit, OpSize::i64Bit, data);
if (Truncate) {
data = _Float_ToGPR_ZS(Size == OpSize::i32Bit ? OpSize::i32Bit : OpSize::i64Bit, OpSize::i64Bit, data);
} else {
data = _Float_ToGPR_S(Size == OpSize::i32Bit ? OpSize::i32Bit : OpSize::i64Bit, OpSize::i64Bit, data);
}
StoreResultGPR_WithOpSize(Op, Op->Dest, data, Size, OpSize::i8Bit,
CTX->IsVectorAtomicTSOEnabled() ? MemoryAccessType::DEFAULT : MemoryAccessType::NONTSO);
}
StoreResultGPR_WithOpSize(Op, Op->Dest, data, Size, OpSize::i8Bit);
if ((Op->TableInfo->Flags & X86Tables::InstFlags::FLAGS_POP) != 0) {
_PopStackDestroy();
@@ -370,6 +382,8 @@ void OpDispatchBuilder::X87FXTRACTF64(OpcodeArgs) {
// Split node into SIG and EXP while handling the special zero case.
// i.e. if val == 0.0, then sig = 0.0, exp = -inf
// if val == -0.0, then sig = -0.0, exp = -inf
// if val is +/-Inf, then sig = val, exp = +inf
// if val is NaN, then sig = val, exp = val
// otherwise we just extract the 64-bit sig and exp as normal.
Ref Node = _ReadStackValue(0);
@@ -379,6 +393,11 @@ void OpDispatchBuilder::X87FXTRACTF64(OpcodeArgs) {
Ref ExpZV = _VCastFromGPR(OpSize::i64Bit, OpSize::i64Bit, Constant(0xfff0'0000'0000'0000UL));
Ref SigZV = Node;
// Inf/NaN case
Ref ExpInfOnlyV = _VCastFromGPR(OpSize::i64Bit, OpSize::i64Bit, Constant(0x7ff0'0000'0000'0000UL));
Ref ExpNanV = Node;
Ref SigInfV = Node;
// non zero case
Ref ExpNZ = _Bfe(OpSize::i64Bit, 11, 52, Gpr);
ExpNZ = Sub(OpSize::i64Bit, ExpNZ, Constant(1023));
@@ -388,12 +407,24 @@ void OpDispatchBuilder::X87FXTRACTF64(OpcodeArgs) {
SigNZ = _Or(OpSize::i64Bit, SigNZ, Constant(0x3ff0'0000'0000'0000LL));
Ref SigNZV = _VCastFromGPR(OpSize::i64Bit, OpSize::i64Bit, SigNZ);
// Comparison and select to push onto stack
SaveNZCV();
// Mantissa non-zero => NaN (exp result = input); else Inf (exp result = +Inf)
Ref Mantissa = _And(OpSize::i64Bit, Gpr, Constant(0x000f'ffff'ffff'ffffULL));
_TestNZ(OpSize::i64Bit, Mantissa, Constant(~0ULL));
Ref ExpInfV = _NZCVSelectV(OpSize::i64Bit, CondClass::EQ, ExpInfOnlyV, ExpNanV);
// Biased exponent == 0x7ff => Inf/NaN path, else non-zero-case.
Ref BiasedExp = _Bfe(OpSize::i64Bit, 11, 52, Gpr);
SubWithFlags(OpSize::i64Bit, BiasedExp, 0x7ff);
Ref ExpNZOrInf = _NZCVSelectV(OpSize::i64Bit, CondClass::EQ, ExpInfV, ExpNZV);
Ref SigNZOrInf = _NZCVSelectV(OpSize::i64Bit, CondClass::EQ, SigInfV, SigNZV);
// Zero folds on top.
_TestNZ(OpSize::i64Bit, Gpr, Constant(0x7fff'ffff'ffff'ffffUL));
Ref Sig = _NZCVSelectV(OpSize::i64Bit, CondClass::EQ, SigZV, SigNZV);
Ref Exp = _NZCVSelectV(OpSize::i64Bit, CondClass::EQ, ExpZV, ExpNZV);
Ref Sig = _NZCVSelectV(OpSize::i64Bit, CondClass::EQ, SigZV, SigNZOrInf);
Ref Exp = _NZCVSelectV(OpSize::i64Bit, CondClass::EQ, ExpZV, ExpNZOrInf);
_PopStackDestroy();
_PushStack(Exp, Invalid(), OpSize::iInvalid);
@@ -50,7 +50,7 @@ constexpr std::array<X86InstInfo[2], ENTRY_MAX> SecondGroup_ArchSelect_LUT = {{
},
}};
constexpr auto SecondInstGroupOps = []() consteval {
constexpr std::array<X86InstInfo, MAX_INST_SECOND_GROUP_TABLE_SIZE> SecondInstGroupOps = []() consteval {
std::array<X86InstInfo, MAX_INST_SECOND_GROUP_TABLE_SIZE> Table{};
constexpr U16U8InfoStruct SecondaryExtensionOpTable[] = {
// GROUP 1
@@ -402,37 +402,37 @@ constexpr auto SecondInstGroupOps = []() consteval {
// GROUP 16
// AMD documentation claims again that this entire group is n/a to prefix
// Tooling once again fails to disassemble oens with the prefix. Disable until proven otherwise
{OPD(TYPE_GROUP_16, PF_NONE, 0), 1, X86InstInfo{"PREFETCH NTA", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_NONE, 1), 1, X86InstInfo{"PREFETCH T0", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_NONE, 2), 1, X86InstInfo{"PREFETCH T1", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_NONE, 3), 1, X86InstInfo{"PREFETCH T2", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_NONE, 0), 1, X86InstInfo{"PREFETCH NTA", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_NONE, 1), 1, X86InstInfo{"PREFETCH T0", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_NONE, 2), 1, X86InstInfo{"PREFETCH T1", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_NONE, 3), 1, X86InstInfo{"PREFETCH T2", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_NONE, 4), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_NONE, 5), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_NONE, 6), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_NONE, 7), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 0), 1, X86InstInfo{"PREFETCH NTA", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 1), 1, X86InstInfo{"PREFETCH T0", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 2), 1, X86InstInfo{"PREFETCH T1", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 3), 1, X86InstInfo{"PREFETCH T2", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 0), 1, X86InstInfo{"PREFETCH NTA", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 1), 1, X86InstInfo{"PREFETCH T0", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 2), 1, X86InstInfo{"PREFETCH T1", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 3), 1, X86InstInfo{"PREFETCH T2", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 4), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 5), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 6), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 7), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_66, 0), 1, X86InstInfo{"PREFETCH NTA", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_66, 1), 1, X86InstInfo{"PREFETCH T0", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_66, 2), 1, X86InstInfo{"PREFETCH T1", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_66, 3), 1, X86InstInfo{"PREFETCH T2", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_66, 0), 1, X86InstInfo{"PREFETCH NTA", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_66, 1), 1, X86InstInfo{"PREFETCH T0", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_66, 2), 1, X86InstInfo{"PREFETCH T1", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_66, 3), 1, X86InstInfo{"PREFETCH T2", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_66, 4), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_66, 5), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_66, 6), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_66, 7), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F2, 0), 1, X86InstInfo{"PREFETCH NTA", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_F2, 1), 1, X86InstInfo{"PREFETCH T0", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_F2, 2), 1, X86InstInfo{"PREFETCH T1", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_F2, 3), 1, X86InstInfo{"PREFETCH T2", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_F2, 0), 1, X86InstInfo{"PREFETCH NTA", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F2, 1), 1, X86InstInfo{"PREFETCH T0", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F2, 2), 1, X86InstInfo{"PREFETCH T1", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F2, 3), 1, X86InstInfo{"PREFETCH T2", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F2, 4), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F2, 5), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F2, 6), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
@@ -31,19 +31,19 @@ constexpr std::array<X86InstInfo[2], ENTRY_MAX> Secondary_ArchSelect_LUT = {{
},
{
{"PUSH FS", TYPE_INST, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_FS_PREFIX> } },
{"PUSH FS", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_FS_PREFIX> } },
{"PUSH FS", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_FS_PREFIX> } },
},
{
{"POP FS", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_DEF) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::POPSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_FS_PREFIX> } },
{"POP FS", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_64BIT) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::POPSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_FS_PREFIX> } },
{"POP FS", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_64BITDEF) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::POPSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_FS_PREFIX> } },
},
{
{"PUSH GS", TYPE_INST, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_GS_PREFIX> } },
{"PUSH GS", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_GS_PREFIX> } },
{"PUSH GS", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_GS_PREFIX> } },
},
{
{"POP GS", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_DEF) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::POPSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_GS_PREFIX> } },
{"POP GS", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_64BIT) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::POPSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_GS_PREFIX> } },
{"POP GS", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_64BITDEF) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::POPSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_GS_PREFIX> } },
},
}};
+14 -7
View File
@@ -136,6 +136,7 @@
"u16": "uint16_t",
"u32": "uint32_t",
"u64": "uint64_t",
"c_str": "const char*",
"OpSize": "FEXCore::IR::OpSize",
"SSA": "OrderedNode*",
"GPR": "OrderedNode*",
@@ -240,6 +241,12 @@
"Desc": ["Debug operation that prints an SSA value to the console",
"May only print 64bits of the value"]
},
"PrintMsg c_str:$Value": {
"HasSideEffects": true,
"Desc": ["Debug operation that prints an string to the console.",
"This is for debug only! Will break code caching!"
]
},
"GPR = AllocateGPR i1:$ForPair": {
"Desc": ["Silly pseudo-instruction to allocate a register for a future destination",
"Note: if an instruction uses allocated destinations-as-sources,",
@@ -2751,7 +2758,7 @@
"F64": {
"FPR = F64ATAN FPR:$Src1, FPR:$Src2": {
"DestSize": "OpSize::i64Bit",
"JITDispatch": false
"JITDispatch": true
},
"FPR = F64FPREM FPR:$Src1, FPR:$Src2": {
"DestSize": "OpSize::i64Bit",
@@ -2763,27 +2770,27 @@
},
"FPR = F64SCALE FPR:$Src1, FPR:$Src2": {
"DestSize": "OpSize::i64Bit",
"JITDispatch": false
"JITDispatch": true
},
"FPR = F64F2XM1 FPR:$Src": {
"DestSize": "OpSize::i64Bit",
"JITDispatch": false
"JITDispatch": true
},
"FPR = F64FYL2X FPR:$Src, FPR:$Src2": {
"DestSize": "OpSize::i64Bit",
"JITDispatch": false
"JITDispatch": true
},
"FPR = F64TAN FPR:$Src": {
"DestSize": "OpSize::i64Bit",
"JITDispatch": false
"JITDispatch": true
},
"FPR = F64SIN FPR:$Src": {
"DestSize": "OpSize::i64Bit",
"JITDispatch": false
"JITDispatch": true
},
"FPR = F64COS FPR:$Src": {
"DestSize": "OpSize::i64Bit",
"JITDispatch": false
"JITDispatch": true
},
"FPR:$Sin, FPR:$Cos = F64SINCOS FPR:$Src": {
"DestSize": "OpSize::i64Bit",
+6 -2
View File
@@ -38,6 +38,10 @@ static void PrintArg(fextl::stringstream* out, const IRListView*, uint64_t Arg)
*out << fextl::fmt::format("#{:#x}", Arg);
}
static void PrintArg(fextl::stringstream* out, const IRListView*, const char* const Arg) {
*out << fextl::fmt::format("'{}'", Arg);
}
static void PrintArg(fextl::stringstream* out, const IRListView*, CondClass Arg) {
if (Arg == CondClass::AL) {
*out << "ALWAYS";
@@ -356,8 +360,8 @@ void Dump(fextl::stringstream* out, const IRListView* IR) {
++CurrentIndent;
AddIndent();
*out << fextl::fmt::format("(%0) IRHeader %{}, #{:#x}, #{}, #{}\n", HeaderOp->Blocks.ID(), HeaderOp->OriginalRIP, HeaderOp->BlockCount,
HeaderOp->NumHostInstructions);
*out << fextl::fmt::format("(%0) IRHeader %{}, #{:#x}, #{}, #{}\n", HeaderOp->Blocks.ID(), +HeaderOp->OriginalRIP, +HeaderOp->BlockCount,
+HeaderOp->NumHostInstructions);
for (auto [BlockNode, BlockHeader] : IR->GetBlocks()) {
{
+7 -5
View File
@@ -21,15 +21,15 @@ class IREmitter {
public:
IREmitter(FEXCore::Utils::IntrusivePooledAllocator& ThreadAllocator, bool SupportsTSOImm9)
: DualListData {ThreadAllocator, 8 * 1024 * 1024}
, SupportsTSOImm9(SupportsTSOImm9) {
ReownOrClaimBuffer();
ResetWorkingList();
}
, SupportsTSOImm9(SupportsTSOImm9) {}
virtual ~IREmitter() = default;
void ReownOrClaimBuffer() {
DualListData.ReownOrClaimBuffer();
// Reset the working list on new buffer.
ResetWorkingList();
}
void DelayedDisownBuffer() {
@@ -39,7 +39,6 @@ public:
IRListView ViewIR() {
return IRListView(&DualListData);
}
void ResetWorkingList();
/**
* @name IR allocation routines
@@ -512,6 +511,9 @@ protected:
fextl::vector<Ref> CodeBlocks;
uint64_t Entry {};
bool SupportsTSOImm9 {};
private:
void ResetWorkingList();
};
} // namespace FEXCore::IR
@@ -119,11 +119,7 @@ class DualIntrusiveAllocatorThreadPool final : public DualIntrusiveAllocator {
public:
DualIntrusiveAllocatorThreadPool(FEXCore::Utils::IntrusivePooledAllocator& ThreadAllocator, size_t Size)
: DualIntrusiveAllocator {Size}
, PoolObject {ThreadAllocator, Size * 2} {
// Claim a buffer on allocation
PoolObject.ReownOrClaimBuffer();
}
, PoolObject {ThreadAllocator, Size * 2} {}
void ReownOrClaimBuffer() {
Data = PoolObject.ReownOrClaimBuffer();
List = Data + MemorySize;
@@ -51,7 +51,7 @@ void IRDumper::Run(IREmitter* IREmit) {
// DumpIRStr might be no if not dumping but ShouldDump is set in OpDisp
if (DumpToFile) {
const auto fileName = fextl::fmt::format("{}/{:x}{}", DumpIR(), HeaderOp->OriginalRIP, IR.PostRA() ? "-post.ir" : "-pre.ir");
const auto fileName = fextl::fmt::format("{}/{:x}{}", DumpIR(), +HeaderOp->OriginalRIP, IR.PostRA() ? "-post.ir" : "-pre.ir");
FD = FEXCore::File::File(fileName.c_str(),
FEXCore::File::FileModes::WRITE | FEXCore::File::FileModes::CREATE | FEXCore::File::FileModes::TRUNCATE);
}
@@ -60,9 +60,9 @@ void IRDumper::Run(IREmitter* IREmit) {
fextl::stringstream out;
FEXCore::IR::Dump(&out, &IR);
if (FD.IsValid()) {
fextl::fmt::print(FD, "IR-{} 0x{:x}:\n{}\n@@@@@\n", IR.PostRA() ? "post" : "pre", HeaderOp->OriginalRIP, out.str());
fextl::fmt::print(FD, "IR-{} 0x{:x}:\n{}\n@@@@@\n", IR.PostRA() ? "post" : "pre", +HeaderOp->OriginalRIP, out.str());
} else {
LogMan::Msg::IFmt("IR-{} 0x{:x}:\n{}\n@@@@@\n", IR.PostRA() ? "post" : "pre", HeaderOp->OriginalRIP, out.str());
LogMan::Msg::IFmt("IR-{} 0x{:x}:\n{}\n@@@@@\n", IR.PostRA() ? "post" : "pre", +HeaderOp->OriginalRIP, out.str());
}
}
}
@@ -862,25 +862,20 @@ void X87StackOptimization::Run(IREmitter* Emit) {
Ref SinValue {};
Ref CosValue {};
if (ReducedPrecisionMode) {
SinValue = IREmit->_F64SIN(St0);
CosValue = IREmit->_F64COS(St0);
}
#ifdef VIXL_SIMULATOR
if (DisableVixlIndirectCalls() == 0) {
if (ReducedPrecisionMode) {
SinValue = IREmit->_F64SIN(St0);
CosValue = IREmit->_F64COS(St0);
} else {
SinValue = IREmit->_F80SIN(St0);
CosValue = IREmit->_F80COS(St0);
}
} else
else if (DisableVixlIndirectCalls() == 0) {
SinValue = IREmit->_F80SIN(St0);
CosValue = IREmit->_F80COS(St0);
}
#endif
{
else {
SinValue = IREmit->_AllocateFPR(OpSize::i128Bit, OpSize::i128Bit);
CosValue = IREmit->_AllocateFPR(OpSize::i128Bit, OpSize::i128Bit);
if (ReducedPrecisionMode) {
IREmit->_F64SINCOS(St0, SinValue, CosValue);
} else {
IREmit->_F80SINCOS(St0, SinValue, CosValue);
}
IREmit->_F80SINCOS(St0, SinValue, CosValue);
}
// Push values
+26 -2
View File
@@ -1,5 +1,6 @@
// SPDX-License-Identifier: MIT
#include "Utils/Allocator/HostAllocator.h"
#include "Utils/Allocator.h"
#include <FEXCore/Utils/Allocator.h>
#include <FEXCore/Utils/CompilerDefs.h>
#include <FEXCore/Utils/LogManager.h>
@@ -32,8 +33,8 @@ std::pmr::memory_resource* get_default_resource() {
}
} // namespace fextl::pmr
#ifndef _WIN32
namespace FEXCore::Allocator {
#ifndef _WIN32
MMAP_Hook mmap {::mmap};
MUNMAP_Hook munmap {::munmap};
@@ -261,9 +262,19 @@ fextl::vector<MemoryRegion> StealMemoryRegion(uintptr_t Begin, uintptr_t End) {
}
// Block remaining memory gaps
bool SupportsDontDump = true;
for (auto RegionIt = Regions.begin(); RegionIt != Regions.end(); ++RegionIt) {
auto Alloc = ::mmap(RegionIt->Ptr, RegionIt->Size, PROT_NONE, MAP_ANONYMOUS | MAP_NORESERVE | MAP_PRIVATE | MAP_FIXED_NOREPLACE, -1, 0);
if (SupportsDontDump) {
// Mark these regions as don't dump so that coredump doesn't try dumping large unmapped regions.
// Ideally coredump would be smart enough to only dump resident pages, but here we are.
auto Result = madvise(RegionIt->Ptr, RegionIt->Size, MADV_DONTDUMP);
if (Result == -1) {
SupportsDontDump = false;
}
}
LogMan::Throw::AFmt(Alloc != MAP_FAILED, "StealMemoryRegion: mmap({}, {:x}) failed: {}", fmt::ptr(RegionIt->Ptr), RegionIt->Size, errno);
LogMan::Throw::AFmt(Alloc == RegionIt->Ptr, "mmap returned {} instead of {}", Alloc, fmt::ptr(RegionIt->Ptr));
}
@@ -304,5 +315,18 @@ void UnlockAfterFork(FEXCore::Core::InternalThreadState* Thread, bool Child) {
Alloc64->UnlockAfterFork(Thread, Child);
}
}
} // namespace FEXCore::Allocator
#else
void VirtualNameNOP(const char*, const void*, size_t) {}
void VirtualTHPNOP(const void* Ptr, size_t Size, THPControl Control) {}
VirtualNamePtr VirtualName {VirtualNameNOP};
VirtualTHPPtr VirtualTHPControl {VirtualTHPNOP};
void SetupHooks(size_t PageSize, HookPtrs Ptrs) {
VirtualName = Ptrs.VirtualName;
VirtualTHPControl = Ptrs.VirtualTHPControl;
}
#endif
} // namespace FEXCore::Allocator
+2
View File
@@ -1,11 +1,13 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <cstddef>
namespace FEXCore::Core {
struct InternalThreadState;
}
namespace FEXCore::Allocator {
void InitializeAllocator(size_t PageSize);
void LockBeforeFork(FEXCore::Core::InternalThreadState* Thread);
void UnlockAfterFork(FEXCore::Core::InternalThreadState* Thread, bool Child);
} // namespace FEXCore::Allocator
+3 -1
View File
@@ -2,7 +2,6 @@
#ifdef ENABLE_FEX_ALLOCATOR
#include <rpmalloc/rpmalloc.h>
#ifndef _WIN32
#include <linux/prctl.h>
#include <sys/prctl.h>
#include <sys/mman.h>
#else
@@ -109,6 +108,9 @@ static void* FEX_rp_mmap(size_t size, size_t alignment, size_t* offset, size_t*
#define PR_SET_VMA_ANON_NAME 0
#endif
prctl(PR_SET_VMA, PR_SET_VMA_ANON_NAME, ptr, map_size, global_config.page_name);
// Disable HUGEPAGE on allocation from rpmalloc.
madvise(ptr, map_size, MADV_NOHUGEPAGE);
}
if (ptr == nullptr) {
+13 -2
View File
@@ -2,7 +2,7 @@
#include "Interface/Core/CPUBackend.h"
#include "Interface/Context/Context.h"
#include "Utils/SpinWaitLock.h"
#include <FEXCore/Utils/SpinWaitLock.h>
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXCore/Utils/EnumUtils.h>
@@ -2005,8 +2005,19 @@ std::optional<int32_t> HandleUnalignedAccess(FEXCore::Core::InternalThreadState*
Thread->ExclusiveStore.Size = 0;
return 4;
}
} else if ((Instr & ArchHelpers::Arm64::ATOMIC_MEM_MASK) == ArchHelpers::Arm64::ATOMIC_MEM_INST) { // Atomic memory op
if (ArchHelpers::Arm64::HandleAtomicMemOp(Instr, GPRs, StrictSplitLockMutex)) {
// Skip this instruction now
return 4;
} else {
uint8_t Op = (PC[0] >> 12) & 0xF;
LogMan::Msg::EFmt("Unhandled JIT SIGBUS Atomic mem op 0x{:02x}: PC: 0x{:x} Instruction: 0x{:08x}\n", Op, ProgramCounter, PC[0]);
return std::nullopt;
}
}
return 0;
LogMan::Msg::EFmt("Unhandled non-JIT atomic");
return std::nullopt;
}
const auto Frame = Thread->CurrentFrame;
@@ -32,7 +32,7 @@ public:
// Differs from Itanium specification
LOGMAN_THROW_A_FMT(PMF.adj == 0, "C++ Pointer-To-Member representation didn't have adj == 0. Are you trying to cast a virtual member?");
#else
#error Don't know how to cast Member to function here. Likely just Itanium
#error "Don't know how to cast Member to function here. Likely just Itanium"
#endif
return PMF.ptr;
}
@@ -54,7 +54,7 @@ public:
"members.");
return PMF.ptr;
#else
#error Don't know how to cast Member to function here. Likely just Itanium
#error "Don't know how to cast Member to function here. Likely just Itanium"
#endif
}
+3 -3
View File
@@ -1,11 +1,11 @@
// SPDX-License-Identifier: MIT
#include "Utils/SpinWaitLock.h"
#include <FEXCore/Utils/SpinWaitLock.h>
namespace FEXCore::Utils::SpinWaitLock {
#ifdef ARCHITECTURE_arm64
constexpr uint64_t NanosecondsInSecond = 1'000'000'000ULL;
static uint32_t GetCycleCounterFrequency() {
static uint64_t GetCycleCounterFrequency() {
uint64_t Result {};
__asm("mrs %[Res], CNTFRQ_EL0" : [Res] "=r"(Result));
return Result;
@@ -21,7 +21,7 @@ static uint64_t CalculateCyclesPerNanosecond() {
return NanosecondsInSecond / CounterFrequency;
}
uint32_t CycleCounterFrequency = GetCycleCounterFrequency();
uint64_t CycleCounterFrequency = GetCycleCounterFrequency();
uint64_t CyclesPerNanosecond = CalculateCyclesPerNanosecond();
#endif
} // namespace FEXCore::Utils::SpinWaitLock
+23
View File
@@ -0,0 +1,23 @@
// SPDX-License-Identifier: MIT
#include <FEXCore/Utils/WildcardMatcher.h>
namespace FEXCore::Utils::Wildcard {
static bool matchHelper(std::string_view pattern, std::string_view text, size_t p_idx, size_t t_idx) {
if (p_idx == pattern.size()) {
// Pattern exhausted
return (t_idx == text.size());
} else if (pattern[p_idx] == '*') {
// Wildcard: Try matching zero characters, or one or more characters
return matchHelper(pattern, text, p_idx + 1, t_idx) || (t_idx < text.size() && matchHelper(pattern, text, p_idx, t_idx + 1));
} else {
// Match normally
return (t_idx < text.size() && pattern[p_idx] == text[t_idx] && matchHelper(pattern, text, p_idx + 1, t_idx + 1));
}
}
bool Matches(std::string_view pattern, std::string_view text) {
return matchHelper(pattern, text, 0, 0);
}
} // namespace FEXCore::Utils::Wildcard
+6 -1
View File
@@ -27,7 +27,12 @@ namespace HLE {
struct SourcecodeMap;
} // namespace HLE
enum class GuestRelocationType : uint32_t { Rel32, Rel64 };
enum class GuestRelocationType : uint32_t {
Rel32,
Rel64,
// Skip blocks containing this relocation
Skip,
};
// Generic information associated with an executable file.
struct ExecutableFileInfo {
+3 -2
View File
@@ -13,10 +13,10 @@
#include <FEXCore/fextl/set.h>
#include <FEXCore/fextl/string.h>
#include <FEXCore/fextl/vector.h>
#include <FEXCore/Utils/WritePriorityMutex.h>
namespace FEXCore {
struct HostFeatures;
class ForkableSharedMutex;
class ThunkHandler;
} // namespace FEXCore
@@ -73,6 +73,7 @@ public:
*/
FEX_DEFAULT_VISIBILITY virtual void ExecuteThread(FEXCore::Core::InternalThreadState* Thread) = 0;
FEX_DEFAULT_VISIBILITY virtual bool CheckIfBlockIsCacheable(FEXCore::Core::InternalThreadState& Thread, uint64_t GuestRIP, uint64_t MaxInst) = 0;
FEX_DEFAULT_VISIBILITY virtual void CompileRIP(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP) = 0;
FEX_DEFAULT_VISIBILITY virtual void CompileRIPCount(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP, uint64_t MaxInst) = 0;
@@ -143,7 +144,7 @@ public:
FEX_DEFAULT_VISIBILITY virtual void InvalidateCodeBuffersCodeRange(uint64_t Start, uint64_t Length) = 0;
FEX_DEFAULT_VISIBILITY virtual void
InvalidateThreadCachedCodeRange(FEXCore::Core::InternalThreadState* Thread, uint64_t Start, uint64_t Length) = 0;
FEX_DEFAULT_VISIBILITY virtual FEXCore::ForkableSharedMutex& GetCodeInvalidationMutex() = 0;
FEX_DEFAULT_VISIBILITY virtual FEXCore::Utils::WritePriorityMutex::Mutex& GetCodeInvalidationMutex() = 0;
FEX_DEFAULT_VISIBILITY virtual void
ConfigureAOTGen(FEXCore::Core::InternalThreadState* Thread, fextl::set<uint64_t>* ExternalBranches, uint64_t SectionMaxAddress) = 0;
+8
View File
@@ -337,6 +337,7 @@ struct JITPointers {
// Process specific
uint64_t PrintValue {};
uint64_t PrintVectorValue {};
uint64_t PrintMsgValue {};
uint64_t ThreadRemoveCodeEntryFromJIT {};
uint64_t CPUIDObj {};
uint64_t CPUIDFunction {};
@@ -375,6 +376,13 @@ struct JITPointers {
uint64_t L2Pointer {};
uint64_t LUDIVHandler {};
uint64_t LDIVHandler {};
uint64_t F64SinHandler {};
uint64_t F64CosHandler {};
uint64_t F64TanHandler {};
uint64_t F64F2XM1Handler {};
uint64_t F64ScaleHandler {};
uint64_t F64AtanHandler {};
uint64_t F64FYL2XHandler {};
/** @} */
// Copy of process-wide named vector constants data.
@@ -41,6 +41,7 @@ struct HostFeatures {
bool SupportsWFXT {};
bool Supports3DNow {};
bool SupportsSSE4a {};
bool SupportsMOPS {};
// Float exception behaviour
bool SupportsAFP {};
@@ -15,7 +15,6 @@
namespace FEXCore {
class LookupCache;
class CompileService;
struct JITSymbolBuffer;
} // namespace FEXCore
@@ -102,10 +101,6 @@ struct alignas(FEXCore::Utils::FEX_PAGE_SIZE) InternalThreadState : public FEXCo
NonMovableUniquePtr<FEXCore::IR::PassManager> PassManager;
NonMovableUniquePtr<JITSymbolBuffer> SymbolBuffer;
std::shared_ptr<FEXCore::CompileService> CompileService;
std::shared_mutex ObjectCacheRefCounter {};
// This pointer is owned by the frontend.
FEXCore::SHMStats::ThreadStats* ThreadStats {};
@@ -12,9 +12,6 @@ struct InternalThreadState;
}
namespace FEXCore::Allocator {
FEX_DEFAULT_VISIBILITY void SetupHooks(size_t PageSize);
FEX_DEFAULT_VISIBILITY void ClearHooks();
FEX_DEFAULT_VISIBILITY size_t DetermineVASize();
#ifdef GLIBC_ALLOCATOR_FAULT
+24 -3
View File
@@ -27,6 +27,24 @@ enum class ProtectOptions : uint32_t {
};
FEX_DEF_NUM_OPS(ProtectOptions)
enum class THPControl {
Enable,
Disable,
};
#ifndef _WIN32
FEX_DEFAULT_VISIBILITY void SetupHooks(size_t PageSize);
#else
using VirtualNamePtr = void (*)(const char*, const void*, size_t);
using VirtualTHPPtr = void (*)(const void*, size_t, THPControl);
struct HookPtrs {
VirtualNamePtr VirtualName;
VirtualTHPPtr VirtualTHPControl;
};
FEX_DEFAULT_VISIBILITY void SetupHooks(size_t PageSize, HookPtrs Ptrs);
#endif
FEX_DEFAULT_VISIBILITY void ClearHooks();
#ifdef _WIN32
inline void* VirtualAlloc(void* Base, size_t Size, bool Execute = false, bool Commit = true) {
// Allocate top-down to avoid polluting the lower VA space, as even on 64-bit some programs (i.e. LuaJIT) require allocations below 4GB.
@@ -82,8 +100,8 @@ inline bool VirtualProtect(void* Ptr, size_t Size, ProtectOptions options) {
return ::VirtualProtect(Ptr, Size, prot, nullptr) == 0;
}
inline void VirtualName(const char*, void*, size_t) {}
FEX_DEFAULT_VISIBILITY extern VirtualNamePtr VirtualName;
FEX_DEFAULT_VISIBILITY extern VirtualTHPPtr VirtualTHPControl;
#else
using MMAP_Hook = void* (*)(void*, size_t, int, int, int, off_t);
using MUNMAP_Hook = int (*)(void*, size_t);
@@ -123,6 +141,10 @@ inline bool VirtualProtect(void* Ptr, size_t Size, ProtectOptions options) {
return ::mprotect(Ptr, Size, prot) == 0;
}
inline void VirtualTHPControl(const void* Ptr, size_t Size, THPControl Control) {
::madvise(const_cast<void*>(Ptr), Size, Control == THPControl::Enable ? MADV_HUGEPAGE : MADV_NOHUGEPAGE);
}
#endif
// Memory allocation routines to be defined externally.
@@ -142,7 +164,6 @@ void aligned_free(void* ptr);
FEX_DEFAULT_VISIBILITY extern void InitializeThread();
#ifndef _WIN32
void InitializeAllocator(size_t PageSize);
void SetupAllocatorHooks(void* (*)(void* addr, size_t length, int prot, int flags, int fd, off_t offset), int (*)(void* addr, size_t length));
#endif
@@ -28,6 +28,9 @@
// then program behavior is undefined.
#define FEX_UNREACHABLE __builtin_unreachable()
// Like offsetof but for array members with a dynamic element index
#define ARRAY_OFFSETOF(Type, ArrayMember, Index) (offsetof(Type, ArrayMember) + sizeof(Type::ArrayMember[0]) * (Index))
namespace FEXCore::Assert {
// This function can not be inlined
[[noreturn]]
@@ -2,7 +2,6 @@
#pragma once
#ifndef _WIN32
#include <linux/prctl.h>
#include <sys/mman.h>
#include <sys/user.h>
#include <sys/prctl.h>
@@ -50,8 +50,9 @@ namespace FEXCore::Utils::SpinWaitLock {
#define SPINLOOP_32BIT SPINLOOP_BODY(ldar, w)
#define SPINLOOP_64BIT SPINLOOP_BODY(ldar, x)
extern uint32_t CycleCounterFrequency;
extern uint64_t CyclesPerNanosecond;
FEX_DEFAULT_VISIBILITY extern uint64_t CycleCounterFrequency;
FEX_DEFAULT_VISIBILITY extern uint64_t CyclesPerNanosecond;
///< Get the raw cycle counter which is synchronizing.
/// `CNTVCTSS_EL0` also does the same thing, but requires the FEAT_ECV feature.
@@ -0,0 +1,8 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <string_view>
namespace FEXCore::Utils::Wildcard {
bool Matches(std::string_view pattern, std::string_view text);
} // namespace FEXCore::Utils::Wildcard
@@ -9,11 +9,15 @@
#include <unistd.h>
#else
#include <synchapi.h>
// Don't pull in all WIN32 headers for INFINITE. Causes too many problems.
#ifndef INFINITE
#define INFINITE 0xffffffff
#endif
#endif
#include <FEXCore/Utils/LogManager.h>
#include "Utils/SpinWaitLock.h"
#include <FEXCore/Utils/SpinWaitLock.h>
namespace FEXCore::Utils::WritePriorityMutex {
+1 -1
View File
@@ -1,5 +1,5 @@
// SPDX-License-Identifier: MIT
#include "Utils/SpinWaitLock.h"
#include <FEXCore/Utils/SpinWaitLock.h>
#include <catch2/catch_test_macros.hpp>
#include <chrono>
#include <thread>
+5 -5
View File
@@ -361,6 +361,11 @@ def IsSupportedKernel():
return version_check(GetKernelVersion()) >= version_check("5.15")
def main():
if not IsSupportedDistro():
Distro = GetDistro()
print ( "'{} {}' is not a supported distro".format(Distro[0], Distro[1]))
ExitWithStatus(-1)
# Only run on supported arch
if not IsSupportedArch():
print ( "{} is not a supported architecture".format(GetArch()))
@@ -371,11 +376,6 @@ def main():
print ( "Kernel {} is too old. FEX needs 5.15 minimum".format(GetKernelVersion()))
ExitWithStatus(-1)
if not IsSupportedDistro():
Distro = GetDistro()
print ( "'{} {}' is not a supported distro".format(Distro[0], Distro[1]))
ExitWithStatus(-1)
if GetDistro()[0] == "ubuntu":
print ("Getting PPA status: {}".format(("NotInstalled", "Installed")[GetPPAStatus()]))
+2
View File
@@ -59,6 +59,7 @@ class HostFeatures(Flag) :
FEATURE_LRCPC = (1 << 14)
FEATURE_LRCPC2 = (1 << 15)
FEATURE_FRINTTS = (1 << 16)
FEATURE_MOPS = (1 << 17)
HostFeaturesLookup = {
"SVE128" : HostFeatures.FEATURE_SVE128,
@@ -78,6 +79,7 @@ HostFeaturesLookup = {
"LRCPC" : HostFeatures.FEATURE_LRCPC,
"LRCPC2" : HostFeatures.FEATURE_LRCPC2,
"FRINTTS" : HostFeatures.FEATURE_FRINTTS,
"MOPS" : HostFeatures.FEATURE_MOPS,
}
def GetHostFeatures(data):
+57 -36
View File
@@ -7,11 +7,12 @@
#include <FEXCore/fextl/fmt.h>
#include <FEXCore/fextl/map.h>
#include <FEXCore/fextl/string.h>
#include <FEXCore/fextl/vector.h>
#include <FEXCore/Utils/Allocator.h>
#include <FEXCore/Utils/FileLoading.h>
#include <FEXCore/Utils/WildcardMatcher.h>
#include <FEXHeaderUtils/Filesystem.h>
#include <FEXHeaderUtils/SymlinkChecks.h>
#include <cstring>
#include <fmt/format.h>
#include <functional>
@@ -28,7 +29,8 @@
namespace FEX::Config {
namespace JSON {
static void LoadJSonConfig(const fextl::string& Config, std::function<void(const char* Name, const char* ConfigSring)> Func) {
static void LoadJSonConfig(const fextl::string& Config, std::optional<fextl::string> AppName,
std::function<void(const char* Name, const char* ConfigString)> Func) {
fextl::vector<char> Data;
if (!FEXCore::FileLoading::LoadFile(Data, Config)) {
return;
@@ -48,21 +50,35 @@ namespace JSON {
return;
}
for (const json_t* ConfigItem = json_getChild(ConfigList); ConfigItem != nullptr; ConfigItem = json_getSibling(ConfigItem)) {
const char* ConfigName = json_getName(ConfigItem);
const char* ConfigString = json_getValue(ConfigItem);
fextl::vector<const json_t*> ConfigBlocks;
ConfigBlocks.push_back(ConfigList);
if (!ConfigName) {
LogMan::Msg::EFmt("JSON file '{}': Couldn't get config name for an item", Config);
return;
if (AppName) {
const json_t* OverrideList = json_getProperty(json, "AppOverrides");
if (OverrideList) {
for (const json_t* Item = json_getChild(OverrideList); Item != nullptr; Item = json_getSibling(Item)) {
const char* AppPattern = json_getName(Item);
// Find the first match, then break
if (FEXCore::Utils::Wildcard::Matches(AppPattern, *AppName)) {
ConfigBlocks.push_back(Item);
break;
}
}
}
}
if (!ConfigString) {
LogMan::Msg::EFmt("JSON file '{}': Couldn't get value for config item '{}'", Config, ConfigName);
return;
for (auto ConfigBlock : ConfigBlocks) {
for (const json_t* ConfigItem = json_getChild(ConfigBlock); ConfigItem != nullptr; ConfigItem = json_getSibling(ConfigItem)) {
const char* ConfigName = json_getName(ConfigItem);
const char* ConfigString = json_getValue(ConfigItem);
if (!ConfigString) {
LogMan::Msg::EFmt("JSON file '{}': Couldn't get value for config item '{}'", Config, ConfigName);
return;
}
Func(ConfigName, ConfigString);
}
Func(ConfigName, ConfigString);
}
}
} // namespace JSON
@@ -167,22 +183,24 @@ protected:
class MainLoader final : public OptionMapper {
public:
explicit MainLoader(FEXCore::Config::LayerType Type);
explicit MainLoader(fextl::string ConfigFile);
explicit MainLoader(FEXCore::Config::LayerType Type, std::optional<fextl::string> AppName = std::nullopt);
explicit MainLoader(fextl::string ConfigFile, std::optional<fextl::string> AppName = std::nullopt);
explicit MainLoader(FEXCore::Config::LayerType Type, std::string_view ConfigFile);
void Load() override;
private:
std::optional<fextl::string> AppName;
fextl::string Config;
};
class AppLoader final : public OptionMapper {
public:
explicit AppLoader(const fextl::string& Filename, FEXCore::Config::LayerType Type);
explicit AppLoader(const fextl::string& AppName, FEXCore::Config::LayerType Type);
void Load();
private:
const fextl::string AppName;
fextl::string Config;
};
@@ -221,12 +239,14 @@ void OptionMapper::MapNameToOption(const char* ConfigName, const char* ConfigStr
#include <FEXCore/Config/ConfigOptions.inl>
}
MainLoader::MainLoader(FEXCore::Config::LayerType Type)
MainLoader::MainLoader(FEXCore::Config::LayerType Type, std::optional<fextl::string> AppName)
: OptionMapper(Type)
, AppName {AppName}
, Config {FEXCore::Config::GetConfigFileLocation(Type == FEXCore::Config::LayerType::LAYER_GLOBAL_MAIN)} {}
MainLoader::MainLoader(fextl::string ConfigFile)
MainLoader::MainLoader(fextl::string ConfigFile, std::optional<fextl::string> AppName)
: OptionMapper(FEXCore::Config::LayerType::LAYER_MAIN)
, AppName {AppName}
, Config {std::move(ConfigFile)} {}
@@ -236,13 +256,14 @@ MainLoader::MainLoader(FEXCore::Config::LayerType Type, std::string_view ConfigF
void MainLoader::Load() {
SetCurrentConfigFile(Config);
JSON::LoadJSonConfig(Config, [this](const char* Name, const char* ConfigString) { MapNameToOption(Name, ConfigString); });
JSON::LoadJSonConfig(Config, AppName, [this](const char* Name, const char* ConfigString) { MapNameToOption(Name, ConfigString); });
}
AppLoader::AppLoader(const fextl::string& Filename, FEXCore::Config::LayerType Type)
: OptionMapper(Type) {
AppLoader::AppLoader(const fextl::string& AppName, FEXCore::Config::LayerType Type)
: OptionMapper(Type)
, AppName {AppName} {
const bool Global = Type == FEXCore::Config::LayerType::LAYER_GLOBAL_STEAM_APP || Type == FEXCore::Config::LayerType::LAYER_GLOBAL_APP;
Config = FEXCore::Config::GetApplicationConfig(Filename, Global);
Config = FEXCore::Config::GetApplicationConfig(AppName, Global);
// Immediately load so we can reload the meta layer
Load();
@@ -250,7 +271,7 @@ AppLoader::AppLoader(const fextl::string& Filename, FEXCore::Config::LayerType T
void AppLoader::Load() {
SetCurrentConfigFile(Config);
JSON::LoadJSonConfig(Config, [this](const char* Name, const char* ConfigString) { MapNameToOption(Name, ConfigString); });
JSON::LoadJSonConfig(Config, AppName, [this](const char* Name, const char* ConfigString) { MapNameToOption(Name, ConfigString); });
}
EnvLoader::EnvLoader(char* const _envp[])
@@ -320,11 +341,11 @@ fextl::unique_ptr<FEXCore::Config::Layer> CreateGlobalMainLayer() {
return fextl::make_unique<MainLoader>(FEXCore::Config::LayerType::LAYER_GLOBAL_MAIN);
}
fextl::unique_ptr<FEXCore::Config::Layer> CreateMainLayer(const fextl::string* File) {
fextl::unique_ptr<FEXCore::Config::Layer> CreateMainLayer(const fextl::string* File, std::optional<fextl::string> AppName) {
if (File) {
return fextl::make_unique<MainLoader>(*File);
return fextl::make_unique<MainLoader>(*File, std::move(AppName));
} else {
return fextl::make_unique<MainLoader>(FEXCore::Config::LayerType::LAYER_MAIN);
return fextl::make_unique<MainLoader>(FEXCore::Config::LayerType::LAYER_MAIN, std::move(AppName));
}
}
@@ -461,7 +482,7 @@ void LoadConfig(fextl::string ProgramName, char** const envp, const PortableInfo
if (!IsPortable) {
FEXCore::Config::AddLayer(CreateGlobalMainLayer());
}
FEXCore::Config::AddLayer(CreateMainLayer());
FEXCore::Config::AddLayer(CreateMainLayer(nullptr, ProgramName.empty() ? std::nullopt : std::optional {ProgramName}));
if (!ProgramName.empty()) {
if (!IsPortable) {
@@ -633,17 +654,10 @@ fextl::string GetDataDirectory(bool Global, const PortableInformation& PortableI
}
fextl::string GetConfigDirectory(bool Global, const PortableInformation& PortableInfo) {
#ifdef FEX_STEAM_SUPPORT
const char* SteamDataPath = getenv("STEAM_COMPAT_DATA_PATH");
if (SteamDataPath) {
return fextl::fmt::format("{}/fex-emu/", SteamDataPath);
}
#endif
const char* ConfigOverride = getenv("FEX_APP_CONFIG_LOCATION");
if (PortableInfo.IsPortable && (Global || !ConfigOverride)) {
if (PortableInfo.IsPortable && Global) {
return fextl::fmt::format("{}/fex-emu/", PortableInfo.InterpreterPath);
} else if (PortableInfo.IsPortable && ConfigOverride && !Global) {
} else if (ConfigOverride && !Global) {
fextl::string AppConfigStr = ConfigOverride;
if (FHU::Filesystem::IsRelative(AppConfigStr)) {
AppConfigStr = PortableInfo.InterpreterPath + AppConfigStr;
@@ -652,6 +666,13 @@ fextl::string GetConfigDirectory(bool Global, const PortableInformation& Portabl
return AppConfigStr;
}
#ifdef FEX_STEAM_SUPPORT
const char* SteamDataPath = getenv("STEAM_COMPAT_DATA_PATH");
if (SteamDataPath) {
return fextl::fmt::format("{}/fex-emu/", SteamDataPath);
}
#endif
fextl::string ConfigDir;
if (Global) {
return GLOBAL_DATA_DIRECTORY;
+1 -1
View File
@@ -81,7 +81,7 @@ fextl::unique_ptr<FEXCore::Config::Layer> CreateGlobalMainLayer();
*
* @return unique_ptr for that layer
*/
fextl::unique_ptr<FEXCore::Config::Layer> CreateMainLayer(const fextl::string* File = nullptr);
fextl::unique_ptr<FEXCore::Config::Layer> CreateMainLayer(const fextl::string* File = nullptr, std::optional<fextl::string> AppName = std::nullopt);
fextl::unique_ptr<FEXCore::Config::Layer> CreateUserOverrideLayer(std::string_view AppConfig);
/**
+2 -2
View File
@@ -17,9 +17,9 @@
#include <fcntl.h>
#include <linux/limits.h>
#include <unistd.h>
#include <sys/poll.h>
#include <poll.h>
#include <sys/prctl.h>
#include <sys/signal.h>
#include <signal.h>
#include <sys/socket.h>
#include <sys/stat.h>
#include <sys/types.h>
+41 -38
View File
@@ -502,6 +502,7 @@ static void OverrideFeatures(FEXCore::HostFeatures* Features, uint64_t ForceSVEW
ENABLE_DISABLE_OPTION(SupportsWFXT, WFXT, WFXT);
ENABLE_DISABLE_OPTION(Supports3DNow, 3DNOW, 3DNOW);
ENABLE_DISABLE_OPTION(SupportsSSE4a, SSE4A, SSE4A);
ENABLE_DISABLE_OPTION(SupportsMOPS, MOPS, MOPS);
GET_SINGLE_OPTION(Crypto, CRYPTO);
#undef ENABLE_DISABLE_OPTION
@@ -538,6 +539,7 @@ static void HandleErrata(FEXCore::HostFeatures* HostFeatures, uint64_t MIDR) {
constexpr uint32_t Implementer_QCOM = 0x51;
constexpr uint32_t PartNum_Oryon1 = 0x001;
constexpr uint32_t PartNum_Oryon3 = 0x002;
auto GetMIDRImplementer = [](uint32_t MIDR) -> uint32_t {
return (MIDR >> 24) & 0xFF;
@@ -551,7 +553,7 @@ static void HandleErrata(FEXCore::HostFeatures* HostFeatures, uint64_t MIDR) {
const uint32_t MIDR_PartNum = GetMIDRPartNum(MIDR);
#ifdef ARCHITECTURE_arm64
if (MIDR_Implementer == Implementer_QCOM && MIDR_PartNum == PartNum_Oryon1) {
if (MIDR_Implementer == Implementer_QCOM && (MIDR_PartNum == PartNum_Oryon1 || MIDR_PartNum == PartNum_Oryon3)) {
// Work around an errata in Qualcomm's Oryon.
// While this CPU implements the RAND extension:
// - The RNDR register works.
@@ -625,11 +627,49 @@ void FetchHostFeatures(FEX::CPUFeatures& Features, FEXCore::HostFeatures& HostFe
// Hardcode enable SVE with 256-bit wide registers.
HostFeatures.SupportsSVE128 = ForceSVEWidth() ? ForceSVEWidth() >= 128 : true;
HostFeatures.SupportsSVE256 = ForceSVEWidth() ? ForceSVEWidth() >= 256 : true;
HostFeatures.SupportsMOPS = true;
// Simulator has a hardcoded ZVA size of 64-bytes.
HostFeatures.SupportsCLZERO = true;
HostFeatures.SupportsAES = true;
HostFeatures.SupportsCRC = true;
HostFeatures.SupportsAVX = true;
HostFeatures.SupportsSHA = true;
HostFeatures.SupportsPMULL_128Bit = true;
HostFeatures.SupportsAES256 = true;
// Simulator doesn't support these
HostFeatures.SupportsRPRES = false;
HostFeatures.SupportsAFP = false;
#else
HostFeatures.SupportsSVE128 = Features.Supports(CPUFeatures::Feature::SVE2);
HostFeatures.SupportsSVE256 = Features.Supports(CPUFeatures::Feature::SVE2) && Features.GetSVEVectorLengthInBits() >= 256;
HostFeatures.SupportsMOPS = Features.Supports(CPUFeatures::Feature::MOPS);
// Check if we can support cacheline clears
if (Features.GetDCZID().SupportsDCZVA()) {
// If the DC ZVA size matches the emulated cache line size
// This means we can use the instruction
constexpr static uint64_t CACHELINE_SIZE = 64;
HostFeatures.SupportsCLZERO = Features.GetDCZID().BlockSizeInBytes() == CACHELINE_SIZE;
}
#endif
HostFeatures.SupportsAVX = true;
HostFeatures.SupportsAES256 = HostFeatures.SupportsAVX && HostFeatures.SupportsAES;
HostFeatures.SupportsPreserveAllABI = FEX_HAS_PRESERVE_ALL_ATTR;
if (CTR) {
HostFeatures.DCacheLineSize = 4 << ((CTR >> 16) & 0xF);
HostFeatures.ICacheLineSize = 4 << (CTR & 0xF);
} else {
HostFeatures.DCacheLineSize = 64;
HostFeatures.ICacheLineSize = 64;
}
if (!HostFeatures.SupportsAtomics) {
WARN_ONCE_FMT("Host CPU doesn't support atomics. Expect bad performance");
}
#ifdef _WIN32
// Disable 3DNow! by default to better match the set of extensions exposed on modern CPUs.
@@ -640,12 +680,6 @@ void FetchHostFeatures(FEX::CPUFeatures& Features, FEXCore::HostFeatures& HostFe
HostFeatures.Supports3DNow = true;
#endif
HostFeatures.SupportsAES256 = HostFeatures.SupportsAVX && HostFeatures.SupportsAES;
if (!HostFeatures.SupportsAtomics) {
WARN_ONCE_FMT("Host CPU doesn't support atomics. Expect bad performance");
}
#ifdef ARCHITECTURE_arm64
// Test if this CPU supports float exception trapping by attempting to enable
// On unsupported these bits are architecturally defined as RAZ/WI
@@ -666,36 +700,6 @@ void FetchHostFeatures(FEX::CPUFeatures& Features, FEXCore::HostFeatures& HostFe
SetFPCR(OriginalFPCR);
#endif
#ifdef VIXL_SIMULATOR
// simulator has a hardcoded ZVA size of 64-bytes.
HostFeatures.SupportsCLZERO = true;
HostFeatures.SupportsAES = true;
HostFeatures.SupportsCRC = true;
HostFeatures.SupportsAVX = true;
HostFeatures.SupportsSHA = true;
HostFeatures.SupportsPMULL_128Bit = true;
HostFeatures.SupportsAES256 = true;
// Simulator doesn't support these
HostFeatures.SupportsRPRES = false;
HostFeatures.SupportsAFP = false;
#else
// Check if we can support cacheline clears
if (Features.GetDCZID().SupportsDCZVA()) {
// If the DC ZVA size matches the emulated cache line size
// This means we can use the instruction
constexpr static uint64_t CACHELINE_SIZE = 64;
HostFeatures.SupportsCLZERO = Features.GetDCZID().BlockSizeInBytes() == CACHELINE_SIZE;
}
#endif
if (CTR) {
HostFeatures.DCacheLineSize = 4 << ((CTR >> 16) & 0xF);
HostFeatures.ICacheLineSize = 4 << (CTR & 0xF);
} else {
HostFeatures.DCacheLineSize = HostFeatures.ICacheLineSize = 64;
}
#if defined(ARCHITECTURE_x86_64) && !defined(VIXL_SIMULATOR)
FEX::X86::Features Feature {};
HostFeatures.SupportsAES = Feature.Feat_aes;
@@ -712,7 +716,6 @@ void FetchHostFeatures(FEX::CPUFeatures& Features, FEXCore::HostFeatures& HostFe
HostFeatures.SupportsAFP = true;
HostFeatures.SupportsFloatExceptions = true;
#endif
HostFeatures.SupportsPreserveAllABI = FEX_HAS_PRESERVE_ALL_ATTR;
HandleErrata(&HostFeatures, MIDR);
OverrideFeatures(&HostFeatures, ForceSVEWidth());
+3 -3
View File
@@ -30,15 +30,15 @@ public:
auto data_7 = cpuid(0x7);
Feat_fsgsbase = data_7.ebx & (1U << 0);
Feat_bmi1 = data_7.ebx & (1U << 3);
Feat_avx &= data_7.ebx & (1U << 5);
Feat_avx = Feat_avx && (data_7.ebx & (1U << 5));
Feat_bmi2 = data_7.ebx & (1U << 8);
Feat_clwb = data_7.ebx & (1U << 24);
Feat_rand &= data_7.ebx & (1U << 18);
Feat_rand = Feat_rand && (data_7.ebx & (1U << 18));
Feat_adx = data_7.ebx & (1U << 19);
Feat_clflopt = data_7.ebx & (1U << 23);
Feat_sha = data_7.ebx & (1U << 29);
Feat_vaes = data_7.ecx & (1U << 9);
Feat_pclmulqdq &= data_7.ecx & (1U << 10);
Feat_pclmulqdq = Feat_pclmulqdq && (data_7.ecx & (1U << 10));
Feat_rdpid = data_7.ecx & (1U << 22);
}
+15 -14
View File
@@ -26,24 +26,25 @@ fextl::string GenerateSteamConfigTemplate(const FEX::Config::PortableInformation
return {};
}
// Try and find a mount point.
fextl::string MountPoint {};
// If the graphics provider was provided through an environment variable, then use that.
// Otherwise, try to find a mount point.
const char* GraphicsProvider = getenv("STEAM_COMPAT_GRAPHICS_PROVIDER");
const char* RuntimeDir = getenv("XDG_RUNTIME_DIR");
if (RuntimeDir) {
const char* CacheDir = getenv("XDG_CACHE_HOME");
const auto UserDirectory = fextl::fmt::format("/run/user/{}", geteuid());
if (GraphicsProvider) {
MountPoint = FHU::Filesystem::ParentPath(GraphicsProvider);
} else if (RuntimeDir) {
MountPoint = fextl::fmt::format("{}/fexrootfs/", RuntimeDir);
} else if (FHU::Filesystem::Exists(UserDirectory)) {
MountPoint = fextl::fmt::format("{}/fexrootfs/", UserDirectory);
} else if (CacheDir) {
MountPoint = fextl::fmt::format("{}/fexrootfs/", CacheDir);
} else {
const auto UserDirectory = fextl::fmt::format("/run/user/{}", geteuid());
if (FHU::Filesystem::Exists(UserDirectory)) {
MountPoint = fextl::fmt::format("{}/fexrootfs/", UserDirectory);
} else {
const char* CacheDir = getenv("XDG_CACHE_HOME");
if (CacheDir) {
MountPoint = fextl::fmt::format("{}/fexrootfs/", CacheDir);
} else {
// We tried really hard to find a mount path.
MountPoint = "~/.cache/fexrootfs/";
}
}
// We tried really hard to find a mount path.
MountPoint = "~/.cache/fexrootfs/";
}
// Update the @FEX_COMPAT_TOOL@ config to point to the root of the depot.
+7
View File
@@ -519,6 +519,7 @@ int main(int argc, char** argv, char** const envp) {
FEATURE_LRCPC = (1U << 14),
FEATURE_LRCPC2 = (1U << 15),
FEATURE_FRINTTS = (1U << 16),
FEATURE_MOPS = (1U << 17),
};
uint64_t SVEWidth = 0;
@@ -569,6 +570,9 @@ int main(int argc, char** argv, char** const envp) {
if (TestHeaderData->EnabledHostFeatures & FEATURE_FRINTTS) {
HostFeatureControl |= static_cast<uint64_t>(FEXCore::Config::HostFeatures::ENABLEFRINTTS);
}
if (TestHeaderData->EnabledHostFeatures & FEATURE_MOPS) {
HostFeatureControl |= static_cast<uint64_t>(FEXCore::Config::HostFeatures::ENABLEMOPS);
}
if (TestHeaderData->EnabledHostFeatures & FEATURE_TSO) {
FEXCore::Config::Set(FEXCore::Config::ConfigOption::CONFIG_TSOENABLED, "1");
@@ -624,6 +628,9 @@ int main(int argc, char** argv, char** const envp) {
if (TestHeaderData->DisabledHostFeatures & FEATURE_FRINTTS) {
HostFeatureControl |= static_cast<uint64_t>(FEXCore::Config::HostFeatures::DISABLEFRINTTS);
}
if (TestHeaderData->DisabledHostFeatures & FEATURE_MOPS) {
HostFeatureControl |= static_cast<uint64_t>(FEXCore::Config::HostFeatures::DISABLEMOPS);
}
if (TestHeaderData->DisabledHostFeatures & FEATURE_TSO) {
FEXCore::Config::Set(FEXCore::Config::ConfigOption::CONFIG_TSOENABLED, "0");
@@ -1,11 +1,13 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <FEXCore/Core/CodeCache.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/fextl/string.h>
#include <FEXCore/fextl/vector.h>
#include <elf.h>
#include <fcntl.h>
#include <optional>
#include <unistd.h>
#include "Linux/Utils/ELFContainer.h"
@@ -19,6 +21,7 @@
struct ELFParser {
Elf64_Ehdr ehdr;
fextl::vector<Elf64_Phdr> phdrs;
std::optional<fextl::vector<Elf64_Shdr>> shdrs;
::ELFLoader::ELFContainer::ELFType type {::ELFLoader::ELFContainer::TYPE_NONE};
fextl::string InterpreterElf;
@@ -30,6 +33,7 @@ struct ELFParser {
fd = NewFD;
type = ::ELFLoader::ELFContainer::TYPE_NONE;
shdrs.reset();
if (fd == -1) {
// Likely just doesn't exist
@@ -236,6 +240,168 @@ struct ELFParser {
return ReadElf(NewFD);
}
/**
* Checks if DT_TEXTREL/DF_TEXTREL exist in the PT_DYNAMIC segment.
*
* These indicate that the ELF has relocations that cover to read-only code
* pages. The dynamic loader will temporarily map these pages as writeable
* to apply the relocations.
*/
bool HasCodeRelocations() const {
if (fd == -1) {
return false;
}
auto phdr_it = std::ranges::find_if(phdrs, [](auto& phdr) { return phdr.p_type == PT_DYNAMIC; });
if (phdr_it == phdrs.end()) {
return false;
}
if (type == ::ELFLoader::ELFContainer::TYPE_X86_32) {
return HasCodeRelocations<Elf32_Dyn>(*phdr_it);
} else {
return HasCodeRelocations<Elf64_Dyn>(*phdr_it);
}
}
template<typename Elf_Dyn>
bool HasCodeRelocations(const Elf64_Phdr& phdr) const {
const size_t EntryCount = phdr.p_filesz / sizeof(Elf_Dyn);
fextl::vector<Elf_Dyn> Entries(EntryCount);
if (pread(fd, Entries.data(), phdr.p_filesz, phdr.p_offset) == -1) {
return false;
}
for (auto& Entry : Entries) {
if (Entry.d_tag == DT_NULL) {
break;
}
if (Entry.d_tag == DT_TEXTREL) {
return true;
}
if (Entry.d_tag == DT_FLAGS && (Entry.d_un.d_val & DF_TEXTREL)) {
return true;
}
}
return false;
}
/**
* Parses relocation sections (SHT_REL/SHT_RELA) and returns a map of
* offsets to relocations that FEX's JIT must know about.
*/
fextl::robin_map<uint32_t, FEXCore::GuestRelocationType> PopulateRelocations() {
if (fd == -1 || !EnsureSectionHeadersLoaded()) {
return {};
}
fextl::robin_map<uint32_t, FEXCore::GuestRelocationType> Relocations;
bool Is32Bit = (type == ::ELFLoader::ELFContainer::TYPE_X86_32);
for (const auto& shdr : *shdrs) {
if (shdr.sh_entsize == 0) {
continue;
}
const size_t EntryCount = shdr.sh_size / shdr.sh_entsize;
if (!Is32Bit) {
if (shdr.sh_type == SHT_REL) {
LOGMAN_THROW_A_FMT(false, "Unexpected relocation section type");
} else if (shdr.sh_type == SHT_RELA) {
fextl::vector<Elf64_Rela> Entries(EntryCount);
if (pread(fd, Entries.data(), shdr.sh_size, shdr.sh_offset) == -1) {
LOGMAN_THROW_A_FMT(false, "Failed to read RELA section");
}
for (auto& Entry : Entries) {
auto RelocType = ClassifyRelocation64(ELF64_R_TYPE(Entry.r_info));
if (RelocType) {
Relocations.emplace(static_cast<uint32_t>(Entry.r_offset), *RelocType);
}
}
}
} else {
if (shdr.sh_type == SHT_REL) {
fextl::vector<Elf32_Rel> Entries(EntryCount);
if (pread(fd, Entries.data(), shdr.sh_size, shdr.sh_offset) == -1) {
LOGMAN_THROW_A_FMT(false, "Failed to read REL section");
}
for (auto& Entry : Entries) {
auto RelocType = ClassifyRelocation32(ELF32_R_TYPE(Entry.r_info));
if (RelocType) {
Relocations.emplace(static_cast<uint32_t>(Entry.r_offset), *RelocType);
}
}
} else if (shdr.sh_type == SHT_RELA) {
fextl::vector<Elf32_Rela> Entries(EntryCount);
if (pread(fd, Entries.data(), shdr.sh_size, shdr.sh_offset) == -1) {
LOGMAN_THROW_A_FMT(false, "Failed to read RELA section");
}
for (auto& Entry : Entries) {
auto RelocType = ClassifyRelocation32(ELF32_R_TYPE(Entry.r_info));
if (RelocType) {
Relocations.emplace(static_cast<uint32_t>(Entry.r_offset), *RelocType);
}
}
}
}
}
return Relocations;
}
/**
* Returns underlying 32-bit relocation entries.
* SHT_REL entries are implicitly converted to Elf32_Rela.
*/
fextl::vector<Elf32_Rela> ReadRawRelocations32() {
if (fd == -1 || type != ::ELFLoader::ELFContainer::TYPE_X86_32 || !EnsureSectionHeadersLoaded()) {
return {};
}
// Load dynamic symbol table (find SHT_DYNSYM section)
fextl::vector<Elf32_Sym> DynSyms;
auto DynsymHeader = std::ranges::find_if(*shdrs, [](auto& shdr) { return shdr.sh_type == SHT_DYNSYM; });
if (DynsymHeader != shdrs->end()) {
size_t SymCount = DynsymHeader->sh_size / sizeof(Elf32_Sym);
DynSyms.resize(SymCount);
if (pread(fd, DynSyms.data(), DynsymHeader->sh_size, DynsymHeader->sh_offset) == -1) {
LOGMAN_MSG_A_FMT("Could not load DYNSYM section");
}
}
fextl::vector<Elf32_Rela> Result;
for (const auto& shdr : *shdrs) {
if (shdr.sh_entsize == 0) {
continue;
}
const size_t EntryCount = shdr.sh_size / shdr.sh_entsize;
if (shdr.sh_type == SHT_REL) {
fextl::vector<Elf32_Rel> Entries(EntryCount);
if (pread(fd, Entries.data(), shdr.sh_size, shdr.sh_offset) == -1) {
LOGMAN_MSG_A_FMT("Could not load REL section");
}
for (auto& Entry : Entries) {
auto Sym = ELF32_R_SYM(Entry.r_info);
int32_t Addend = (Sym < DynSyms.size()) ? static_cast<int32_t>(DynSyms[Sym].st_value) : 0;
Result.push_back(Elf32_Rela {Entry.r_offset, Entry.r_info, Addend});
}
} else if (shdr.sh_type == SHT_RELA) {
fextl::vector<Elf32_Rela> Entries(EntryCount);
if (pread(fd, Entries.data(), shdr.sh_size, shdr.sh_offset) == -1) {
LOGMAN_MSG_A_FMT("Could not load RELA section");
}
Result.insert(Result.end(), Entries.begin(), Entries.end());
}
}
return Result;
}
void Closefd() {
if (fd != -1) {
close(fd);
@@ -246,4 +412,71 @@ struct ELFParser {
~ELFParser() {
Closefd();
}
private:
/// Returns true if loading section headers succeeded
bool EnsureSectionHeadersLoaded() {
if (shdrs.has_value()) {
return !shdrs->empty();
}
if (fd == -1 || ehdr.e_shoff == 0 || ehdr.e_shnum == 0) {
shdrs.emplace();
return false;
}
if (type == ::ELFLoader::ELFContainer::TYPE_X86_64) {
shdrs.emplace(ehdr.e_shnum);
if (pread(fd, shdrs->data(), sizeof(Elf64_Shdr) * ehdr.e_shnum, ehdr.e_shoff) == -1) {
shdrs->clear();
return false;
}
} else {
fextl::vector<Elf32_Shdr> shdrs32(ehdr.e_shnum);
if (pread(fd, shdrs32.data(), sizeof(Elf32_Shdr) * ehdr.e_shnum, ehdr.e_shoff) == -1) {
shdrs.emplace();
return false;
}
shdrs.emplace(ehdr.e_shnum);
for (int i = 0; i < ehdr.e_shnum; i++) {
#define COPY(name) (*shdrs)[i].name = shdrs32[i].name
COPY(sh_name);
COPY(sh_type);
COPY(sh_flags);
COPY(sh_addr);
COPY(sh_offset);
COPY(sh_size);
COPY(sh_link);
COPY(sh_info);
COPY(sh_addralign);
COPY(sh_entsize);
#undef COPY
}
}
return !shdrs->empty();
}
static std::optional<FEXCore::GuestRelocationType> ClassifyRelocation32(uint32_t Type) {
if (Type == R_386_RELATIVE || Type == R_386_32) {
return FEXCore::GuestRelocationType::Rel32;
} else if (Type == R_386_PC32) {
// Currently not handled
return FEXCore::GuestRelocationType::Skip;
} else if (Type == R_386_TLS_TPOFF) {
// Currently not handled
return FEXCore::GuestRelocationType::Skip;
}
return std::nullopt;
}
static std::optional<FEXCore::GuestRelocationType> ClassifyRelocation64(uint32_t Type) {
if (Type == R_X86_64_RELATIVE || Type == R_X86_64_64) {
return FEXCore::GuestRelocationType::Rel64;
} else if (Type == R_X86_64_32) {
return FEXCore::GuestRelocationType::Rel32;
}
return std::nullopt;
}
};
+1
View File
@@ -11,3 +11,4 @@ install(TARGETS FEXGetConfig RUNTIME
target_link_libraries(FEXGetConfig PRIVATE ${LIBS})
target_include_directories(FEXGetConfig PRIVATE ${CMAKE_BINARY_DIR}/generated)
target_compile_options(FEXGetConfig PRIVATE ${FEX_TUNE_COMPILE_FLAGS})
+106
View File
@@ -14,6 +14,8 @@
#include <filesystem>
#include <string>
#include <sys/prctl.h>
#include <signal.h>
#include <ucontext.h>
namespace {
struct TSOEmulationFacts {
@@ -94,6 +96,103 @@ TSOEmulationFacts GetTSOEmulationFacts() {
#endif
} // namespace
#ifdef ARCHITECTURE_arm64
namespace SIGBUSTest {
static bool* FaultArray {};
__attribute__((naked)) void atomic_load_u16(std::byte* Data) {
asm volatile(R"(
ldarh w1, [x0];
ret;
)" ::
: "x1", "memory");
}
__attribute__((naked)) void atomic_load_u32(std::byte* Data) {
asm volatile(R"(
ldar w1, [x0];
ret;
)" ::
: "x1", "memory");
}
__attribute__((naked)) void atomic_load_u64(std::byte* Data) {
asm volatile(R"(
ldar x1, [x0];
ret;
)" ::
: "x1", "memory");
}
__attribute__((naked)) void atomic_load_u128(std::byte* Data) {
asm volatile(R"(
ldaxp x1, x2, [x0];
ret;
)" ::
: "x1", "x2", "x3", "memory");
}
static void HandleSIGBUS(int, siginfo_t* info, void* context) {
FaultArray[reinterpret_cast<uintptr_t>(info->si_addr) & 63] = true;
ucontext_t* ucontext = (ucontext_t*)context;
mcontext_t* mcontext = &ucontext->uc_mcontext;
// Skip the stlr.
mcontext->pc += 4;
}
void TestSIGBUS() {
struct sigaction act {};
act.sa_sigaction = HandleSIGBUS;
act.sa_flags = SA_SIGINFO;
sigaction(SIGBUS, &act, &act);
auto ptr = reinterpret_cast<std::byte*>(mmap(nullptr, 4096, PROT_READ | PROT_WRITE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0));
auto test_fault = [](bool* FaultOffsets, auto AccessFunction, std::byte* AccessArray) {
FaultArray = FaultOffsets;
for (size_t i = 0; i < 64; ++i) {
AccessFunction(AccessArray + i);
}
};
auto print_granule = [](const char* size, bool* FaultArray) {
std::string output {};
for (size_t i = 0; i < 64; ++i) {
if (i && (i % 16 == 0)) {
output += " ";
}
if (FaultArray[i]) {
output += "\e[31m■\e[0m";
} else {
output += "\e[32m■\e[0m";
}
}
fprintf(stdout, "%s: %s\n", size, output.c_str());
};
bool FaultOffset_16bit[64] {};
bool FaultOffset_32bit[64] {};
bool FaultOffset_64bit[64] {};
bool FaultOffset_128bit[64] {};
test_fault(FaultOffset_16bit, atomic_load_u16, ptr);
test_fault(FaultOffset_32bit, atomic_load_u32, ptr);
test_fault(FaultOffset_64bit, atomic_load_u64, ptr);
test_fault(FaultOffset_128bit, atomic_load_u128, ptr);
munmap(ptr, 4096);
sigaction(SIGBUS, &act, nullptr);
fprintf(stdout, "Fault Granularity: Split every 16 bytes\n");
print_granule(" 16-bit", FaultOffset_16bit);
print_granule(" 32-bit", FaultOffset_32bit);
print_granule(" 64-bit", FaultOffset_64bit);
print_granule("128-bit", FaultOffset_128bit);
}
} // namespace SIGBUSTest
#endif
int main(int argc, char** argv, char** envp) {
FEX::Config::InitializeConfigs(FEX::Config::PortableInformation {});
FEXCore::Config::Initialize();
@@ -114,6 +213,7 @@ int main(int argc, char** argv, char** envp) {
Parser.add_option("--tso-emulation-info").action("store_true").help("Print how FEX is emulating the x86-TSO memory model.");
#ifdef ARCHITECTURE_arm64
Parser.add_option("--test-fault-granularity").action("store_true").help("Show SIGBUS fault granularity");
Parser.add_option("--identification-reg-info").action("store_true").help("Print identification registers");
#endif
@@ -146,6 +246,12 @@ int main(int argc, char** argv, char** envp) {
fprintf(stdout, GIT_DESCRIBE_STRING "\n");
}
#ifdef ARCHITECTURE_arm64
if (Options.is_set_by_user("test_fault_granularity")) {
SIGBUSTest::TestSIGBUS();
}
#endif
if (Options.is_set_by_user("install_prefix")) {
char SelfPath[PATH_MAX];
auto Result = readlink("/proc/self/exe", SelfPath, PATH_MAX);
+26 -15
View File
@@ -32,7 +32,6 @@
#include <sys/personality.h>
#include <sys/prctl.h>
#include <sys/random.h>
#include <linux/prctl.h>
#define PAGE_START(x) ((x) & ~(uintptr_t)(4095))
#define PAGE_OFFSET(x) ((x) & 4095)
@@ -411,7 +410,7 @@ public:
// Set the process personality here
// Also, what about ADDR_LIMIT_3GB & co ?
uint32_t Personality = personality(~0ULL);
uint32_t Personality = personality(~0U);
Personality |= ExecuteAll ? READ_IMPLIES_EXEC : 0;
if (-1 == personality(Personality)) {
LogMan::Msg::EFmt("Setting personality failed");
@@ -767,26 +766,38 @@ public:
// Ensure we don't read past the end into garbage data
stat_buffer[std::clamp(bytes_read, 0L, static_cast<ssize_t>(sizeof(stat_buffer)) - 1)] = '\0';
uint64_t start_code, end_code, start_stack, start_data, end_data, start_brk, arg_start, arg_end, env_start, env_end;
// See man proc_pid_stat
int items_read = sscanf(stat_buffer,
"%*d %*s %*c %*d %*d " // 1 to 5
"%*d %*d %*d %*u %*u " // 6 to 10
"%*u %*u %*u %*u %*u " // 11 to 15
"%*d %*d %*d %*d %*d " // 16 to 20
"%*d %*u %*u %*d %*u " // 21 to 25
"%llu %llu %llu %*u %*u " // 26 to 30
"%*u %*u %*u %*u %*u " // 31 to 35
"%*u %*u %*d %*d %*u " // 36 to 40
"%*u %*u %*u %*d %llu " // 40 to 45
"%llu %llu %llu %llu %llu " // 46 to 50
"%llu", // 51
&map.start_code, &map.end_code, &map.start_stack, &map.start_data, &map.end_data, &map.start_brk,
&map.arg_start, &map.arg_end, &map.env_start, &map.env_end);
"%*d %*s %*c %*d %*d " // 1 to 5
"%*d %*d %*d %*u %*u " // 6 to 10
"%*u %*u %*u %*u %*u " // 11 to 15
"%*d %*d %*d %*d %*d " // 16 to 20
"%*d %*u %*u %*d %*u " // 21 to 25
"%lu %lu %lu %*u %*u " // 26 to 30
"%*u %*u %*u %*u %*u " // 31 to 35
"%*u %*u %*d %*d %*u " // 36 to 40
"%*u %*u %*u %*d %lu " // 40 to 45
"%lu %lu %lu %lu %lu " // 46 to 50
"%lu", // 51
&start_code, &end_code, &start_stack, &start_data, &end_data, &start_brk, &arg_start, &arg_end, &env_start, &env_end);
if (items_read != 10) {
return false;
}
map.start_code = start_code;
map.end_code = end_code;
map.start_stack = start_stack;
map.start_data = start_data;
map.end_data = end_data;
map.start_brk = start_brk;
map.arg_start = arg_start;
map.arg_end = arg_end;
map.env_start = env_start;
map.env_end = env_end;
map.brk = reinterpret_cast<uint64_t>(sbrk(0));
// The kernel will leave these values unchanged, see implementation in sys.c
@@ -63,7 +63,7 @@ $end_info$
#include <utility>
#include <sys/sysinfo.h>
#include <sys/signal.h>
#include <signal.h>
namespace FEX::Logging {
static bool SilentLog {};
@@ -597,9 +597,14 @@ int main(int argc, char** argv, char** const envp) {
SyscallHandler->DefaultProgramBreak(BRKInfo.Base, BRKInfo.Size);
// Request code cache generation
if (FEXCore::Config::Get_ENABLECODECACHINGWIP()) {
// Request code cache generation
FEXServerClient::PopulateCodeCache(FEXServerClient::GetServerFD(), Loader.GetMainElfFD(), FEXCore::Config::Get_MULTIBLOCK());
if (VDSOMapping) {
// Finalize code cache for libVDSO-guest.so. This needs to be done explicitly since VDSO doesn't use LoadLib.
SyscallHandler->TriggerGuestLibWrapperCodeCacheLoad(*ParentThread->Thread, reinterpret_cast<uintptr_t>(VDSOMapping.VDSOBase));
}
}
// Pull RIP and stack pointer from loader and set the thread data to it.
+53 -8
View File
@@ -4,6 +4,7 @@
#include <PortabilityInfo.h>
#include <Thunks.h>
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/CodeCache.h>
#include <FEXCore/Core/Context.h>
#include <FEXCore/Core/HostFeatures.h>
@@ -16,6 +17,7 @@
#include <OptionParser.h>
#include <fmt/printf.h>
#include <libgen.h>
#include <fstream>
#include <optional>
@@ -50,7 +52,8 @@ public:
}
void* GuestMmap(FEXCore::Core::InternalThreadState*, void* addr, size_t Size, int prot, int Flags, int fd, off_t offset) override {
auto Ret = mmap(addr, Size, prot, Flags, fd, offset);
// Force writeable to allow applying relocations
auto Ret = mmap(addr, Size, prot | PROT_WRITE, Flags, fd, offset);
if (Ret != MAP_FAILED && VAFileStart == 0) {
VAFileStart = reinterpret_cast<uintptr_t>(Ret);
}
@@ -113,8 +116,7 @@ static FEXCore::Core::InternalThreadState* SetupCompileThread(FEXCore::Context::
}
// Returns filename of generated cache on success
static std::optional<std::string>
GenerateSingleCache(const FEXCore::ExecutableFileInfo& Binary, fextl::set<uintptr_t> BlockList, std::string_view OutDir) {
static std::optional<std::string> GenerateSingleCache(FEXCore::ExecutableFileInfo& Binary, fextl::set<uintptr_t> BlockList, std::string_view OutDir) {
uint64_t CodeCacheConfigId = 0; // TODO: Make unique to active configuration
ELFCodeLoader Loader(Binary.Filename.c_str(), -1, "", fextl::vector<fextl::string> {Binary.Filename.c_str()},
@@ -123,7 +125,19 @@ GenerateSingleCache(const FEXCore::ExecutableFileInfo& Binary, fextl::set<uintpt
fmt::print("Invalid or unsupported ELF file.\n");
return std::nullopt;
}
const bool Is64Bit = Loader.Is64BitMode();
auto SyscallOSABI = Is64Bit ? FEXCore::HLE::SyscallOSABI::OS_LINUX64 : FEXCore::HLE::SyscallOSABI::OS_LINUX32;
auto SyscallHandler = std::make_unique<AOTSyscallHandler>(SyscallOSABI);
// Populate relocations from ELF file
{
ELFParser RelocParser;
RelocParser.ReadElf(Binary.Filename);
Binary.Relocations = RelocParser.PopulateRelocations();
SyscallHandler->FileInfo.Relocations = Binary.Relocations;
}
FEXCore::Config::Set(FEXCore::Config::CONFIG_IS64BIT_MODE, Is64Bit ? "1" : "0");
// Load HostFeatures
@@ -137,13 +151,9 @@ GenerateSingleCache(const FEXCore::ExecutableFileInfo& Binary, fextl::set<uintpt
auto CTX = FEXCore::Context::Context::CreateNewContext(HostFeatures);
auto SignalDelegation = std::make_unique<FEX::DummyHandlers::DummySignalDelegator>();
auto SyscallOSABI = Is64Bit ? FEXCore::HLE::SyscallOSABI::OS_LINUX64 : FEXCore::HLE::SyscallOSABI::OS_LINUX32;
auto SyscallHandler = std::make_unique<AOTSyscallHandler>(SyscallOSABI);
Loader.CalculateHWCaps(CTX.get());
auto SignalDelegation = std::make_unique<FEX::DummyHandlers::DummySignalDelegator>();
CTX->SetSignalDelegator(SignalDelegation.get());
CTX->SetSyscallHandler(SyscallHandler.get());
auto ThunkHandler = FEX::HLE::CreateThunkHandler();
@@ -166,6 +176,25 @@ GenerateSingleCache(const FEXCore::ExecutableFileInfo& Binary, fextl::set<uintpt
if (!ElfBase.has_value()) {
ERROR_AND_DIE_FMT("Failed to load ELF file {} ({})", Binary.Filename, Binary.FileId);
}
{
ELFParser RelocParser;
RelocParser.ReadElf(Binary.Filename);
auto relocs32 = RelocParser.ReadRawRelocations32();
for (auto& reloc : relocs32) {
if (ELF32_R_TYPE(reloc.r_info) == R_386_RELATIVE) {
// The FEX-relocation is applied on top of this during cache serialization, so this must be countered
uint32_t val = *reinterpret_cast<uint32_t*>(SyscallHandler->VAFileStart + reloc.r_offset) + SyscallHandler->VAFileStart;
memcpy(reinterpret_cast<uint32_t*>(SyscallHandler->VAFileStart + reloc.r_offset), &val, sizeof(val));
} else if (ELF32_R_TYPE(reloc.r_info) == R_386_32) {
// The FEX-relocation is applied on top of this during cache serialization, so this must be countered
uint32_t* orig = reinterpret_cast<uint32_t*>(SyscallHandler->VAFileStart + reloc.r_offset);
uint32_t val = *orig + reloc.r_addend + SyscallHandler->VAFileStart;
memcpy(orig, &val, sizeof(val));
}
}
}
}
CTX->GetCodeCache().InitiateCacheGeneration();
@@ -173,8 +202,24 @@ GenerateSingleCache(const FEXCore::ExecutableFileInfo& Binary, fextl::set<uintpt
{
std::vector<std::unique_ptr<ELFCodeLoader>> LoaderMem;
// Refuse to continue if the block list contains any out-of-bounds blocks.
// This often indicates a corrupted code map.
{
auto [min_val, max_val] = std::ranges::minmax_element(BlockList, std::less {});
auto MinBound = SyscallHandler->LookupExecutableFileSection(Thread, *min_val + SyscallHandler->VAFileStart);
auto MaxBound = SyscallHandler->LookupExecutableFileSection(Thread, *max_val + SyscallHandler->VAFileStart);
LOGMAN_THROW_A_FMT(MinBound && MaxBound, "Cached blocks offsets {:#x}-{:#x} out of bounds for library {} ({:016x} @ {:#x})!",
*min_val, *max_val, Binary.Filename, Binary.FileId, SyscallHandler->VAFileStart);
}
fmt::print(stderr, "Compiling code...\n");
FEX_CONFIG_OPT(MaxInst, MAXINST);
for (auto Addr : BlockList) {
if (!CTX->CheckIfBlockIsCacheable(*Thread, Addr + SyscallHandler->VAFileStart, MaxInst)) {
continue;
}
CTX->CompileRIP(Thread, Addr + SyscallHandler->VAFileStart);
}
+4
View File
@@ -1031,6 +1031,10 @@ int32_t AskForDistroSelection(const std::span<const WebFileFetcher::FileTargets>
if (!ArgOptions::DistroVersion.empty()) {
Info.DistroVersion = ArgOptions::DistroVersion;
}
// explicit CLI selection must still run exact-match logic.
if (!ArgOptions::DistroName.empty() || !ArgOptions::DistroVersion.empty()) {
Info.Unknown = false;
}
return _AskForDistroSelection(Info, Targets);
}
+86 -18
View File
@@ -5,12 +5,74 @@
#include <fcntl.h>
#include <fmt/format.h>
#include <unistd.h>
#include <vector>
#include <xxhash.h>
#include <functional>
namespace XXFileHash {
// 32MB blocks
constexpr static size_t BLOCK_SIZE = 32 * 1024 * 1024;
class Reader {
public:
Reader(int fd, size_t Size)
: fd {fd}
, Size {Size} {}
virtual ~Reader() = default;
bool Initialized() const {
return IsInitialized;
}
using Callback = std::function<bool(const void* Data, size_t Size)>;
virtual bool Read(Callback cb) = 0;
protected:
int fd {};
size_t Size {};
bool IsInitialized {};
};
class MemoryReader final : public Reader {
public:
MemoryReader(int fd, size_t Size)
: Reader(fd, Size) {
Ptr = reinterpret_cast<std::byte*>(mmap(nullptr, Size, PROT_READ, MAP_SHARED, fd, 0));
IsInitialized = Ptr != MAP_FAILED;
}
~MemoryReader() {
munmap(reinterpret_cast<void*>(Ptr), Size);
}
bool Read(Callback cb) override {
auto ReadPtr = Ptr;
const auto ReadEndPtr = Ptr + Size;
size_t ReadSize {};
// Claim sequential access.
::madvise(reinterpret_cast<void*>(ReadPtr), Size, MADV_SEQUENTIAL);
while (ReadPtr < ReadEndPtr) {
ReadSize = std::min<size_t>(READ_BLOCK_SIZE, ReadEndPtr - ReadPtr);
if (!cb(ReadPtr, ReadSize)) {
return false;
}
// Only allow a single block read to be resident.
::madvise(reinterpret_cast<void*>(ReadPtr), ReadSize, MADV_DONTNEED);
ReadPtr += ReadSize;
}
return true;
}
private:
std::byte* Ptr {};
// Only allow 128MB in flight.
constexpr static size_t READ_BLOCK_SIZE = 128 * 1024 * 1024;
};
std::optional<uint64_t> HashFile(const fextl::string& Filepath) {
int fd = open(Filepath.c_str(), O_RDONLY);
if (fd == -1) {
@@ -48,33 +110,39 @@ std::optional<uint64_t> HashFile(const fextl::string& Filepath) {
return HadError();
}
MemoryReader Read(fd, Size);
if (!Read.Initialized()) {
return HadError();
}
const auto Start = std::chrono::high_resolution_clock::now();
auto Now = Start;
const double SizeD = Size;
std::vector<char> Data(BLOCK_SIZE);
off_t CurrentOffset = 0;
auto Now = std::chrono::high_resolution_clock::now();
size_t CurrentOffset {};
// Let the kernel know that we will be reading linearly
posix_fadvise(fd, 0, Size, POSIX_FADV_SEQUENTIAL);
while (CurrentOffset < Size) {
ssize_t Result = pread(fd, Data.data(), BLOCK_SIZE, CurrentOffset);
if (Result == -1) {
return HadError();
auto CB_XXH = [&](const void* Data, size_t BlockSize) -> bool {
if (XXH3_64bits_update(State, Data, BlockSize) == XXH_ERROR) {
return false;
}
if (XXH3_64bits_update(State, Data.data(), Result) == XXH_ERROR) {
return HadError();
}
auto Cur = std::chrono::high_resolution_clock::now();
auto Dur = Cur - Now;
if (Dur >= std::chrono::seconds(1)) {
fmt::print("{:.2}% hashed\n", (double)CurrentOffset / SizeD * 100.0);
Now = Cur;
}
CurrentOffset += Result;
CurrentOffset += BlockSize;
return true;
};
if (!Read.Read(CB_XXH)) {
return HadError();
}
const XXH64_hash_t Hash = XXH3_64bits_digest(State);
const auto Hash = XXH3_64bits_digest(State);
XXH3_freeState(State);
close(fd);
+1 -1
View File
@@ -21,7 +21,7 @@
#include <optional>
#include <poll.h>
#include <sys/prctl.h>
#include <sys/signal.h>
#include <signal.h>
#include <sys/socket.h>
#include <sys/stat.h>
#include <sys/types.h>
+10 -5
View File
@@ -12,6 +12,7 @@
#include <FEXCore/HLE/SourcecodeResolver.h>
#include <fmt/ranges.h>
#include <inttypes.h>
#include <atomic>
#include <cassert>
@@ -64,6 +65,9 @@ static std::string NewCodeMapDirectory;
// Path to directory for processed code maps (suitable for cache generation)
static std::string ReadyCodeMapDirectory;
// Path to FEXOfflineCompiler executable (inferred from FEXServer install location)
const std::string OfflineCompilerPath = (std::filesystem::read_symlink("/proc/self/exe").parent_path() / "FEXOfflineCompiler").string();
void SetWatchFD(int FD) {
WatchFD = FD;
}
@@ -94,8 +98,8 @@ void CheckRaiseFDLimit() {
if (MaxFDs.rlim_cur == MaxFDs.rlim_max) {
fprintf(stderr, "[FEXMountDaemon] Our open FD limit is already set to max and we are wanting to increase it\n");
fprintf(stderr, "[FEXMountDaemon] FEXMountDaemon will now no longer be able to track new instances of FEX\n");
fprintf(stderr, "[FEXMountDaemon] Current limit is %zd(hard %zd) FDs and we are at %zd\n", MaxFDs.rlim_cur, MaxFDs.rlim_max,
GetNumFilesOpen());
fprintf(stderr, "[FEXMountDaemon] Current limit is %" PRIuMAX "(hard %" PRIuMAX ") FDs and we are at %zu\n", (uintmax_t)MaxFDs.rlim_cur,
(uintmax_t)MaxFDs.rlim_max, GetNumFilesOpen());
fprintf(stderr, "[FEXMountDaemon] Ask your administrator to raise your kernel's hard limit on open FDs\n");
return;
}
@@ -109,7 +113,8 @@ void CheckRaiseFDLimit() {
NewLimit.rlim_cur = std::min(NewLimit.rlim_cur, NewLimit.rlim_max);
if (setrlimit(RLIMIT_NOFILE, &NewLimit) != 0) {
fprintf(stderr, "[FEXMountDaemon] Couldn't raise FD limit to %zd even though our hard limit is %zd\n", NewLimit.rlim_cur, NewLimit.rlim_max);
fprintf(stderr, "[FEXMountDaemon] Couldn't raise FD limit to %" PRIu64 " even though our hard limit is %" PRIu64 "\n",
(uintmax_t)NewLimit.rlim_cur, (uintmax_t)NewLimit.rlim_max);
} else {
// Set the new limit
MaxFDs = NewLimit;
@@ -470,8 +475,8 @@ int32_t EmbedSubprocess(const char* path, char* const* args) {
* Spawn a FEXOfflineCompiler instance to generate a code cache from the given code map
*/
static int RunOfflineCompiler(const char* CodeMap) {
const char* ExecveArgs[] = {"FEXOfflineCompiler", "generate", CodeMap, nullptr};
return EmbedSubprocess("FEXOfflineCompiler", const_cast<char* const*>(&ExecveArgs[0]));
const char* ExecveArgs[] = {OfflineCompilerPath.c_str(), "generate", CodeMap, nullptr};
return EmbedSubprocess(OfflineCompilerPath.c_str(), const_cast<char* const*>(&ExecveArgs[0]));
};
void HandleSocketData(fasio::tcp_socket& Socket) {
+9 -5
View File
@@ -9,7 +9,7 @@
#include <fcntl.h>
#include <filesystem>
#include <sys/poll.h>
#include <poll.h>
#include <sys/stat.h>
#include <sys/wait.h>
#include <thread>
@@ -202,16 +202,20 @@ void UnmountRootFS() {
if (pid == 0) {
const char* argv[5];
argv[0] = "fusermount";
argv[0] = "fusermount3";
argv[1] = "-u";
argv[2] = "-q";
argv[3] = MountFolder.c_str();
argv[4] = nullptr;
if (execvp(argv[0], (char* const*)argv) == -1) {
fprintf(stderr, "fusermount failed to execute. You may have an mount living at '%s' to clean up now\n", MountFolder.c_str());
fprintf(stderr, "Try `%s %s %s %s`\n", argv[0], argv[1], argv[2], argv[3]);
exit(1);
// Try again with `fusermount`
argv[0] = "fusermount";
if (execvp(argv[0], (char* const*)argv) == -1) {
fprintf(stderr, "fusermount{3,} failed to execute. You may have an mount living at '%s' to clean up now\n", MountFolder.c_str());
fprintf(stderr, "Try `%s %s %s %s`\n", argv[0], argv[1], argv[2], argv[3]);
exit(1);
}
}
} else {
// Wait for fusermount to leave
@@ -279,7 +279,7 @@ FileManager::FileManager(FEXCore::Context::Context* ctx)
// Using a local struct for this is slightly less ugly than using self-capturing lambdas
struct {
decltype(FileManager::ThunkOverlays)& ThunkOverlays;
decltype(ThunkDB)& ThunkDB;
decltype(ThunkDB)& DB;
const fextl::string& ThunkGuestPath;
bool Is64BitMode;
@@ -301,7 +301,7 @@ FileManager::FileManager(FEXCore::Context::Context* ctx)
void InsertDependencies(const fextl::unordered_set<fextl::string>& Depends) {
for (const auto& Depend : Depends) {
auto& DBDepend = ThunkDB.at(Depend);
auto& DBDepend = DB.at(Depend);
if (DBDepend.Enabled) {
continue;
}
@@ -35,6 +35,7 @@ $end_info$
#include <FEXCore/fextl/string.h>
#include <FEXCore/fextl/vector.h>
#include <FEXHeaderUtils/Filesystem.h>
#include <FEXHeaderUtils/Syscalls.h>
#include <atomic>
#include <cstring>
@@ -1171,7 +1172,7 @@ GdbServer::HandledPacketType GdbServer::CommandMultiLetterV(const fextl::string&
}
if (packet.starts_with("vKill")) {
tgkill(::getpid(), ::getpid(), SIGKILL);
FHU::Syscalls::tgkill(::getpid(), ::getpid(), SIGKILL);
}
// TODO: vRun
@@ -196,14 +196,7 @@ restart: {
if (MappedPtr == MAP_FAILED && errno != EEXIST) {
return reinterpret_cast<void*>(-errno);
} else if (MappedPtr == MAP_FAILED || MappedPtr >= reinterpret_cast<void*>(TOP_KEY << FEXCore::Utils::FEX_PAGE_SHIFT)) {
// Handles the case where MAP_FIXED_NOREPLACE failed with MAP_FAILED
// or if the host system's kernel isn't new enough then it returns the wrong pointer
if (MappedPtr != MAP_FAILED && MappedPtr >= reinterpret_cast<void*>(TOP_KEY << FEXCore::Utils::FEX_PAGE_SHIFT)) {
// Make sure to munmap this so we don't leak memory
::munmap(MappedPtr, length);
}
} else if (MappedPtr == MAP_FAILED) {
if (UpperPage == TOP_KEY) {
BottomPage = BASE_KEY;
Wrapped = true;
@@ -240,13 +233,7 @@ restart: {
void* MappedPtr = ::mmap(reinterpret_cast<void*>(PageAddr << FEXCore::Utils::FEX_PAGE_SHIFT),
PagesLength << FEXCore::Utils::FEX_PAGE_SHIFT, prot, flags, fd, offset);
if (MappedPtr >= reinterpret_cast<void*>(TOP_KEY << FEXCore::Utils::FEX_PAGE_SHIFT) && (flags & FEX_MAP_FIXED_NOREPLACE)) {
// Handles the case where MAP_FIXED_NOREPLACE isn't handled by the host system's
// kernel and returns the wrong pointer
// Make sure to munmap this so we don't leak memory
::munmap(MappedPtr, length);
return reinterpret_cast<void*>(-EEXIST);
} else if (MappedPtr != MAP_FAILED) {
if (MappedPtr != MAP_FAILED) {
SetUsedPages(PageAddr, PagesLength);
return MappedPtr;
} else {
@@ -90,7 +90,7 @@ uint64_t BPFEmitter::HandleLoad(uint32_t BPFIP, const sock_filter* Inst) {
// Must be smaller than scratch space size.
VALIDATE(Inst->k < 16);
EMIT_INST(ldr(DestReg, REG_SECCOMP_DATA, offsetof(WorkingBuffer, ScratchMemory[Inst->k])));
EMIT_INST(ldr(DestReg, REG_SECCOMP_DATA, ARRAY_OFFSETOF(WorkingBuffer, ScratchMemory, Inst->k)));
break;
case BPF_LEN:
// Just returns the length of seccomp_data.
@@ -114,7 +114,7 @@ uint64_t BPFEmitter::HandleStore(uint32_t BPFIP, const sock_filter* Inst) {
// Must be smaller than scratch space size.
VALIDATE(Inst->k < 16);
EMIT_INST(str(SrcReg, REG_SECCOMP_DATA, offsetof(WorkingBuffer, ScratchMemory[Inst->k])));
EMIT_INST(str(SrcReg, REG_SECCOMP_DATA, ARRAY_OFFSETOF(WorkingBuffer, ScratchMemory, Inst->k)));
RETURN_SUCCESS();
}
@@ -649,17 +649,19 @@ void SignalDelegator::HandleGuestSignal(FEX::HLE::ThreadStateObject* ThreadObjec
"capacity size. This will "
"likely crash! Asserting now!");
// Peek into sigset_t implementation details, extracting the __val member for glibc and __bits for musl
auto& [sigmask_val] = _context->uc_sigmask;
static_assert(sizeof(sigmask_val[0]) == sizeof(uint64_t), "Unknown sigset_t layout");
ThreadObject->SignalInfo.DeferredSignalFrames.emplace_back(ThreadStateObject::DeferredSignalState {
.Info = SigInfo,
.Signal = Signal,
.SigMask = _context->uc_sigmask.__val[0],
.SigMask = sigmask_val[0],
});
uint64_t NewMask = GetNewSigMask(Signal);
// Update our host signal mask so we don't hit race conditions with signals
// This allows us to maintain the expected signal mask through the guest signal handling and then all the way back again
memcpy(&_context->uc_sigmask, &NewMask, sizeof(uint64_t));
sigmask_val[0] = GetNewSigMask(Signal);
// Now update the faulting page permissions so it will fault on write.
mprotect(reinterpret_cast<void*>(&Thread->InterruptFaultPage), sizeof(Thread->InterruptFaultPage), PROT_NONE);
@@ -894,7 +896,7 @@ SignalDelegator::SignalDelegator(FEXCore::Context::Context* _CTX, const std::str
// Most signals default to termination
// These ones are slightly different
static constexpr std::array<std::pair<int, SignalDelegator::DefaultBehaviour>, 14> SignalDefaultBehaviours = {{
static constexpr std::array<std::pair<int, SignalDelegator::DefaultBehaviourType>, 14> SignalDefaultBehaviours = {{
{SIGQUIT, DEFAULT_COREDUMP},
{SIGILL, DEFAULT_COREDUMP},
{SIGTRAP, DEFAULT_COREDUMP},
@@ -174,7 +174,7 @@ private:
FEXCore::ArchHelpers::Arm64::UnalignedHandlerType UnalignedHandlerType {FEXCore::ArchHelpers::Arm64::UnalignedHandlerType::HalfBarrier};
enum DefaultBehaviour {
enum DefaultBehaviourType {
DEFAULT_TERM,
// Core dump based signals are supposed to have a coredump appear
// For FEX's behaviour we don't really care right now
@@ -201,7 +201,7 @@ private:
kernel_sigaction OldAction {};
FEX::HLE::HostSignalDelegatorFunctionForGuest GuestHandler {};
GuestSigAction GuestAction {};
DefaultBehaviour DefaultBehaviour {DEFAULT_TERM};
DefaultBehaviourType DefaultBehaviour {DEFAULT_TERM};
// Callbacks
fextl::vector<HostSignalDelegatorFunction> Handlers {};
@@ -180,6 +180,10 @@ void SignalDelegator::RestoreFrame_x64(FEXCore::Core::InternalThreadState* Threa
CTX->SetXMMRegistersFromState(Thread, fpstate->_xmm, nullptr);
}
// Technically if mxcsr contains invalid bits then rt_sigreturn should return -EINVAL.
// TODO: FEX doesn't support this today.
Frame->State.mxcsr = fpstate->mxcsr & 0xFFC0;
// FCW store default
Frame->State.FCW = fpstate->fcw;
Frame->State.AbridgedFTW = fpstate->ftw;
@@ -254,6 +258,9 @@ void SignalDelegator::RestoreFrame_ia32(FEXCore::Core::InternalThreadState* Thre
CTX->SetXMMRegistersFromState(Thread, fpstate->_xmm, nullptr);
}
// Invalid bits are silently masked off in 32-bit.
Frame->State.mxcsr = fpstate->mxcsr & 0xFFC0;
// FCW store default
Frame->State.FCW = fpstate->fcw;
Frame->State.AbridgedFTW = FEXCore::FPState::ConvertToAbridgedFTW(fpstate->ftw);
@@ -330,6 +337,9 @@ void SignalDelegator::RestoreRTFrame_ia32(FEXCore::Core::InternalThreadState* Th
CTX->SetXMMRegistersFromState(Thread, fpstate->_xmm, nullptr);
}
// Invalid bits are silently masked off in 32-bit.
Frame->State.mxcsr = fpstate->mxcsr & 0xFFC0;
// FCW store default
Frame->State.FCW = fpstate->fcw;
Frame->State.AbridgedFTW = FEXCore::FPState::ConvertToAbridgedFTW(fpstate->ftw);
@@ -471,6 +481,10 @@ uint64_t SignalDelegator::SetupFrame_x64(FEXCore::Core::InternalThreadState* Thr
CTX->ReconstructXMMRegisters(Thread, fpstate->_xmm, nullptr);
}
// Save mxcsr and the default mask.
fpstate->mxcsr = Frame->State.mxcsr;
fpstate->mxcsr_mask = 0xFFC0;
// FCW store default
fpstate->fcw = Frame->State.FCW;
fpstate->ftw = Frame->State.AbridgedFTW;
@@ -594,6 +608,8 @@ uint64_t SignalDelegator::SetupFrame_ia32(FEXCore::Core::InternalThreadState* Th
CTX->ReconstructXMMRegisters(Thread, fpstate->_xmm, nullptr);
}
fpstate->mxcsr = Frame->State.mxcsr;
// FCW store default
fpstate->fcw = Frame->State.FCW;
// Reconstruct FSW
@@ -729,6 +745,8 @@ uint64_t SignalDelegator::SetupRTFrame_ia32(FEXCore::Core::InternalThreadState*
CTX->ReconstructXMMRegisters(Thread, fpstate->_xmm, nullptr);
}
fpstate->mxcsr = Frame->State.mxcsr;
// FCW store default
fpstate->fcw = Frame->State.FCW;
// Reconstruct FSW
@@ -292,6 +292,7 @@ public:
std::optional<FEXCore::ExecutableFileSectionInfo>
LookupExecutableFileSection(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestAddr) final override;
void TriggerGuestLibWrapperCodeCacheLoad(FEXCore::Core::InternalThreadState&, uint64_t AnyAddr);
int OpenCodeMapFile() override;
FEXCore::HLE::ExecutableRangeInfo QueryGuestExecutableRange(FEXCore::Core::InternalThreadState* Thread, uint64_t Address) override;
@@ -33,7 +33,7 @@ $end_info$
#include <stdint.h>
#include <sched.h>
#include <sys/personality.h>
#include <sys/poll.h>
#include <poll.h>
#include <sys/prctl.h>
#include <sys/resource.h>
#include <sys/syscall.h>
@@ -202,7 +202,13 @@ FEXCore::HLE::ExecutableRangeInfo SyscallHandler::QueryGuestExecutableRange(FEXC
return {Entry->first, Entry->second.Length, Entry->second.Prot.Writable};
}
static fextl::vector<Elf64_Phdr> ReadELFHeaders(int FD, std::span<std::byte> HeaderData = {}) {
struct ReadELFHeadersResult {
fextl::vector<Elf64_Phdr> ProgramHeaders;
fextl::robin_map<uint32_t, FEXCore::GuestRelocationType> Relocations;
bool HasCodeRelocations;
};
static ReadELFHeadersResult ReadELFHeaders(int FD, std::span<std::byte> HeaderData = {}) {
std::string_view ELFMagic = ELFMAG;
if (HeaderData.data()) {
if (HeaderData.size_bytes() < ELFMagic.size() || std::memcmp(ELFMagic.data(), HeaderData.data(), ELFMagic.size()) != 0) {
@@ -213,9 +219,18 @@ static fextl::vector<Elf64_Phdr> ReadELFHeaders(int FD, std::span<std::byte> Hea
// Read from FD in case the caller didn't have a mapped header available
}
// Re-open the file with a fresh file descriptor (and let ELFParser close it on return).
// NOTE: FDs returned by dup() share the same cursor state, so reading from them would have observable side effects.
auto NewFD = open(fextl::fmt::format("/proc/self/fd/{}", FD).c_str(), O_RDONLY);
ELFParser Parser;
Parser.ReadElf(dup(FD));
return std::move(Parser.phdrs);
if (!Parser.ReadElf(NewFD)) {
return {};
}
auto Relocations = Parser.PopulateRelocations();
auto HasCodeRelocations = Parser.HasCodeRelocations();
return ReadELFHeadersResult {std::move(Parser.phdrs), std::move(Relocations), HasCodeRelocations};
}
static void LoadCodeCache(FEXCore::Core::InternalThreadState& Thread, FEXCore::ExecutableFileSectionInfo& Section, uint64_t CodeCacheConfigId) {
@@ -280,7 +295,7 @@ void* SyscallHandler::GuestMmap(bool Is64Bit, FEXCore::Core::InternalThreadState
InvalidateCodeRangeIfNecessary(Thread, Result, Size);
if (LateMetadata) {
auto CodeInvalidationlk = GuardSignalDeferringSectionWithFallback(CTX->GetCodeInvalidationMutex(), Thread);
auto CodeInvalidationlk = FEXCore::GuardSignalDeferringSectionWithFallback(CTX->GetCodeInvalidationMutex(), Thread);
CTX->AddForceTSOInformation(LateMetadata->VolatileValidRanges, std::move(LateMetadata->VolatileInstructions));
}
@@ -320,7 +335,7 @@ uint64_t SyscallHandler::GuestMunmap(bool Is64Bit, FEXCore::Core::InternalThread
InvalidateCodeRangeIfNecessary(Thread, reinterpret_cast<uint64_t>(addr), Size);
if (length) {
auto CodeInvalidationlk = GuardSignalDeferringSectionWithFallback(CTX->GetCodeInvalidationMutex(), Thread);
auto CodeInvalidationlk = FEXCore::GuardSignalDeferringSectionWithFallback(CTX->GetCodeInvalidationMutex(), Thread);
CTX->RemoveForceTSOInformation(reinterpret_cast<uint64_t>(addr), length);
}
@@ -351,6 +366,26 @@ uint64_t SyscallHandler::GuestMremap(bool Is64Bit, FEXCore::Core::InternalThread
return Result;
}
void SyscallHandler::TriggerGuestLibWrapperCodeCacheLoad(FEXCore::Core::InternalThreadState& Thread, uint64_t AnyAddr) {
if (!EnableCodeCaching) {
return;
}
// TODO: Instead of deferring the entire cache load, only delay applicance of sha256 relocations!
auto lk = FEXCore::GuardSignalDeferringSection<std::shared_lock>(VMATracking.Mutex, &Thread);
auto VMAEntry = VMATracking.FindVMAEntry(reinterpret_cast<uint64_t>(AnyAddr));
for (auto* VMA = VMAEntry->second.Resource->FirstVMA; VMA; VMA = VMA->ResourceNextVMA) {
if (!VMA->Prot.Executable) {
continue;
}
auto SectionInfo = BuildSectionInfo(*VMAEntry->second.Resource, VMA->Base, VMA->Length);
LoadCodeCache(Thread, SectionInfo, CodeCacheConfigId);
}
}
int SyscallHandler::OpenCodeMapFile() {
// Query from FEXServer whether this is the first instance of this executable; if it is, also enable code dumping!
FEX_CONFIG_OPT(RootFSPath, ROOTFS);
@@ -399,6 +434,34 @@ uint64_t SyscallHandler::GuestMprotect(FEXCore::Core::InternalThreadState* Threa
}
InvalidateCodeRangeIfNecessary(Thread, reinterpret_cast<uint64_t>(addr), len);
// Prepare for delayed code cache load after ld/Wine is done applying relocations.
// Hooking into mprotect is a reliable heuristic that matches behavior of ld (for ELF) and Wine (for PE).
// False-positives are avoided by setting RequiresDelayedCacheLoad in TrackMmap only for
// binaries that we know will go through this path.
fextl::vector<FEXCore::ExecutableFileSectionInfo> CachedSections;
if (EnableCodeCaching && (prot & PROT_EXEC) && (prot & PROT_WRITE) == 0) {
auto lk = FEXCore::GuardSignalDeferringSection(VMATracking.Mutex, Thread);
auto VMAEntry = VMATracking.FindVMAEntry(reinterpret_cast<uint64_t>(addr));
auto Resource = VMAEntry != VMATracking.VMAs.end() ? VMAEntry->second.Resource : nullptr;
if (Resource && Resource->MappedFile && Resource->RequiresDelayedCacheLoad) {
Resource->RequiresDelayedCacheLoad = false;
LogMan::Msg::IFmt("Triggering delayed cache load for {} after mprotect of {:#x}-{:#x}", Resource->MappedFile->Filename,
VMAEntry->first, VMAEntry->first + VMAEntry->second.Length);
for (auto VMA = Resource->FirstVMA; VMA; VMA = VMA->ResourceNextVMA) {
CachedSections.push_back(BuildSectionInfo(*Resource, VMA->Base, VMA->Length));
}
}
}
// Trigger delayed cache load. This must be done separately since
// LoadCodeCache will call interfaces that acquire the VMATracking mutex.
for (auto& CachedSection : CachedSections) {
LoadCodeCache(*Thread, CachedSection, CodeCacheConfigId);
}
return Result;
}
@@ -484,7 +547,7 @@ SyscallHandler::TrackMmap(FEXCore::Core::InternalThreadState* Thread, uint64_t a
const bool MappedELFHeaderAgain = ResourceIt != ResourceEnd && offset == 0 && !ResourceIt->second.ProgramHeaders.empty();
if (ResourceIt == ResourceEnd || MappedELFHeaderAgain) {
// Create a new MappedResource for previously unseen file and for re-mappings of an ELF header
ResourceIt = VMATracking.InsertMappedResource(mrid, {nullptr, nullptr, 0});
ResourceIt = VMATracking.InsertMappedResource(mrid, VMATracking::MappedResource {nullptr, nullptr, 0, {}, {}});
ResourceIt->second.Iterator = ResourceIt;
Inserted = true;
}
@@ -493,19 +556,34 @@ SyscallHandler::TrackMmap(FEXCore::Core::InternalThreadState* Thread, uint64_t a
// Only handle FDs that are backed by regular files that are executable
if (PathLength != -1 && S_ISREG(buf.st_mode) && (buf.st_mode & S_IXUSR)) {
// ELF files that are mapped multiple times get a separate MappedResource for each base virtual address
if (Inserted) {
if ((prot & PROT_READ) && Inserted) {
Resource->MappedFile = fextl::make_unique<FEXCore::ExecutableFileInfo>();
Resource->MappedFile->Filename = fextl::string(Tmp, PathLength);
Resource->MappedFile->FileId = CTX->GetCodeCache().ComputeCodeMapId(Resource->MappedFile->Filename, fd);
// Read ELF headers if applicable.
// Read ELF headers if applicable and needed for code caching.
// For performance, skip ELF checks if we're not mapping the file header
bool CheckForElfFile = (offset == 0);
bool CheckForElfFile = (offset == 0) && EnableCodeCaching;
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
CheckForElfFile = true;
#endif
if (CheckForElfFile) {
Resource->ProgramHeaders = ReadELFHeaders(fd, std::span {reinterpret_cast<std::byte*>(addr), length});
auto ELFResult = ReadELFHeaders(fd, std::span {reinterpret_cast<std::byte*>(addr), length});
Resource->ProgramHeaders = std::move(ELFResult.ProgramHeaders);
Resource->MappedFile->Relocations = std::move(ELFResult.Relocations);
Resource->RequiresDelayedCacheLoad = ELFResult.HasCodeRelocations;
// GuestRelocationType::Skip indicates to FEXOfflineCompiler that
// any blocks covered by the relocation may not be cached.
// At runtime, we can safely drop these relocations.
for (auto it = Resource->MappedFile->Relocations.begin(); it != Resource->MappedFile->Relocations.end();) {
if (it->second == FEXCore::GuestRelocationType::Skip) {
it = Resource->MappedFile->Relocations.erase(it);
} else {
++it;
}
}
LOGMAN_THROW_A_FMT(Resource->ProgramHeaders.empty() || offset == 0, "Expected file offset 0 for the first mapping of an ELF "
"file");
}
@@ -555,7 +633,7 @@ SyscallHandler::TrackMmap(FEXCore::Core::InternalThreadState* Thread, uint64_t a
auto [Iter, IterEnd] = VMATracking.FindResources(mrid);
LOGMAN_THROW_A_FMT(Iter == IterEnd, "VMA tracking error");
Iter = VMATracking.InsertMappedResource(mrid, {nullptr, nullptr, 0});
Iter = VMATracking.InsertMappedResource(mrid, VMATracking::MappedResource {nullptr, nullptr, 0, {}, {}});
Resource = &Iter->second;
Resource->Iterator = Iter;
}
@@ -565,8 +643,16 @@ SyscallHandler::TrackMmap(FEXCore::Core::InternalThreadState* Thread, uint64_t a
// Load code cache if present.
// FEXServer was requested to generate library caches on program launch.
if (EnableCodeCaching && Resource && Resource->MappedFile && VMATracking::VMAProt::fromProt(prot).Executable) {
if (Thread) {
CachedSection.emplace(BuildSectionInfo(*Resource, addr, Size));
if (Resource->MappedFile->Filename.ends_with("-guest.so")) {
// For guest library wrappers, cache loading must be delayed until LoadLib is called.
// Before that, we can't patch up the SHA256 function identifiers.
LogMan::Msg::IFmt("Delaying code cache load for {}", Resource->MappedFile->Filename);
} else if (Thread) {
if (!Resource->RequiresDelayedCacheLoad) {
CachedSection.emplace(BuildSectionInfo(*Resource, addr, Size));
} else {
LogMan::Msg::IFmt("Delaying code cache load for {} until mprotect {:#x}-{:#x}", Resource->MappedFile->Filename, addr, addr + Size);
}
} else {
// Cache can't be loaded with a thread; skip this for now
LogMan::Msg::DFmt("Oops, tried caching without a thread: {}", Resource->MappedFile->Filename);
@@ -627,7 +713,7 @@ void SyscallHandler::TrackShmat(FEXCore::Core::InternalThreadState* Thread, int
auto [Iter, IterEnd] = VMATracking.FindResources(mrid);
if (Iter == IterEnd) {
Iter = VMATracking.InsertMappedResource(mrid, {nullptr, nullptr, Length});
Iter = VMATracking.InsertMappedResource(mrid, VMATracking::MappedResource {nullptr, nullptr, Length, {}, {}});
Iter->second.Iterator = Iter;
}
auto Resource = &Iter->second;
@@ -281,6 +281,11 @@ void VMATracking::DeleteVMARange(FEXCore::Context::Context* CTX, uintptr_t Base,
void VMATracking::ChangeProtectionFlags(uintptr_t Base, uintptr_t Length, VMAProt NewProt) {
Mutex.check_lock_owned_by_self_as_write();
// Handle 0 size as no-op like the kernel
if (Length == 0) {
return;
}
// This needs to handle multiple split-merge strategies:
// 1) Exact overlap - No Split, no Merge. Only protection tracking changes.
// 2) Exact base overlap - Single insert, can never fail.
@@ -49,6 +49,7 @@ struct MappedResource {
uint64_t Length; // 0 if not fixed size
ContainerType::iterator Iterator;
bool RequiresDelayedCacheLoad = false;
fextl::vector<Elf64_Phdr> ProgramHeaders;
};
@@ -189,6 +189,9 @@ FEX::HLE::ThreadStateObject* ThreadManager::CreateThread(uint64_t InitialRIP, ui
FEXCore::Allocator::VirtualName("FEXMem_CallRetStacks", reinterpret_cast<void*>(AllocBase), CALLRET_STACK_ALLOC_SIZE);
// Disable HUGEPAGE on callret stacks.
FEXCore::Allocator::VirtualTHPControl(reinterpret_cast<void*>(AllocBase), CALLRET_STACK_ALLOC_SIZE, FEXCore::Allocator::THPControl::Disable);
// Set the base used for invalidation to the start past the guard pages
ThreadStateObject->Thread->CallRetStackBase = reinterpret_cast<void*>(AllocBase + FEXCore::Utils::FEX_PAGE_SIZE);
::mprotect(ThreadStateObject->Thread->CallRetStackBase, FEXCore::Core::InternalThreadState::CALLRET_STACK_SIZE, PROT_READ | PROT_WRITE);
@@ -30,7 +30,7 @@ $end_info$
#include <optional>
#include <sys/stat.h>
#include <bits/types/sigset_t.h>
#include <signal.h>
#include <linux/seccomp.h>
namespace FEX::HLE {
@@ -221,7 +221,7 @@ public:
// Potential deferred since Thread might not be valid.
// Thread object isn't valid very early in frontend's initialization.
// To be more optimal the frontend should provide this code with a valid Thread object earlier.
auto CodeInvalidationlk = GuardSignalDeferringSectionWithFallback(CTX->GetCodeInvalidationMutex(), CallingThread);
auto CodeInvalidationlk = FEXCore::GuardSignalDeferringSectionWithFallback(CTX->GetCodeInvalidationMutex(), CallingThread);
CTX->InvalidateCodeBuffersCodeRange(Start, Length);
for (auto& Thread : Threads) {
CTX->InvalidateThreadCachedCodeRange(Thread->Thread, Start, Length);
@@ -235,7 +235,7 @@ public:
// Potential deferred since Thread might not be valid.
// Thread object isn't valid very early in frontend's initialization.
// To be more optimal the frontend should provide this code with a valid Thread object earlier.
auto CodeInvalidationlk = GuardSignalDeferringSectionWithFallback(CTX->GetCodeInvalidationMutex(), CallingThread);
auto CodeInvalidationlk = FEXCore::GuardSignalDeferringSectionWithFallback(CTX->GetCodeInvalidationMutex(), CallingThread);
CTX->InvalidateCodeBuffersCodeRange(Start, Length);
for (auto& Thread : Threads) {
CTX->InvalidateThreadCachedCodeRange(Thread->Thread, Start, Length);
@@ -208,7 +208,7 @@ void RegisterThread(FEX::HLE::SyscallHandler* Handler) {
REGISTER_SYSCALL_IMPL_X32(
futex, [](FEXCore::Core::CpuStateFrame* Frame, int* uaddr, int futex_op, int val, const timespec32* timeout, int* uaddr2, uint32_t val3) -> uint64_t {
void* timeout_ptr = (void*)timeout;
const void* timeout_ptr = (const void*)timeout;
struct timespec tp64 {};
int cmd = futex_op & FUTEX_CMD_MASK;
if (timeout && (cmd == FUTEX_WAIT || cmd == FUTEX_LOCK_PI || cmd == FUTEX_WAIT_BITSET || cmd == FUTEX_WAIT_REQUEUE_PI)) {
Loaded 100 of 207 files, more files were not shown because too many files have changed in this diff. Show more