Compare commits

..
Author SHA1 Message Date
Ryan Houdek c330da4992 FEXServerClient: Workaround sun_path 108 byte limit
We really don't want to do this, but in the case that the AF_UNIX path
is longer than the 108-byte limit that sun_path provides we don't really
have a choice. The alternative choice would be to switch /entirely/ away
from AF_UNIX and instead use pipes. We need a bandage fix for now, so
throw the socket in to a temp folder if the path is too long.
2026-07-02 16:49:28 -07:00
Ryan Houdek a04b0241c2 Docs: Update for release FEX-2605 2026-05-08 19:28:30 -07:00
Ryan Houdek 670fd19d33 Merge pull request #5486 from Sonicadvance1/151
Allocator: Mark large unmapped regions as DONTDUMP
2026-05-08 19:28:13 -07:00
Ryan Houdek a66544f3f4 Allocator: Mark large unmapped regions as DONTDUMP
coredump applications aren't smart enough to only dump resident pages,
so explicitly mark our 128TB and other mapped VA ranges as DONTDUMP.

This will speed up coredumps.
2026-05-08 16:27:04 -07:00
Ryan Houdek 1bfb3aefcc Merge pull request #5485 from Sonicadvance1/150
Windows: Setup `tu_override_uncached_as_cache_coherent` inside of dlls
2026-05-08 15:58:43 -07:00
Ryan Houdek b7bfbc3fcd Windows: Setup tu_override_uncached_as_cache_coherent inside of dlls
To not have this environment variable accidently be enabled on arm64
native Wine games, we need to set it from inside of FEX.

Requires the FEX dlls to set them directly rather than launch scripts.
2026-05-08 13:39:51 -07:00
Ryan Houdek e517f3259c Merge pull request #5484 from neobrain/fix_code_cache_no_guest_wrappers
CodeCache: Fix crash when guest library wrappers aren't installed
2026-05-07 10:59:32 -07:00
Tony Wasserka 8afda92a64 CodeCache: Fix crash when guest library wrappers aren't installed 2026-05-07 18:56:35 +02:00
Ryan Houdek ed216c8d4d Merge pull request #5449 from neobrain/opt_code_cache_writing
CodeCache: Slightly optimize cache file writing
2026-05-06 18:06:51 -07:00
Ryan Houdek 7506cb4ea1 Merge pull request #5483 from neobrain/fix_guest_wrapper_code_cache
CodeCache: Delay cache loading for guest library wrappers until after LoadLib
2026-05-06 18:04:55 -07:00
Tony Wasserka 60bc5944db CodeCache: Slightly optimize cache file writing
ftruncate only requires one call (and one extra seek) instead up to 64 manual
zero writes.
2026-05-06 17:08:15 +02:00
Tony Wasserka 8f0572283a Windows/CRT: Implement ftruncate and _chsize 2026-05-06 17:08:04 +02:00
Tony Wasserka b13b46eefe CodeCache: Delay cache loading for guest library wrappers until after LoadLib
These libraries need to be initialized before relocating their caches,
since the guest function hashes won't be registered before.
2026-05-05 16:29:39 +02:00
LC 4db2a98d7f Merge pull request #5481 from Sonicadvance1/148
win32: Query DCZID_EL0 so clzero works
2026-05-05 08:42:31 -04:00
LC 05ebb07753 Merge pull request #5482 from Sonicadvance1/149
OpcodeDispatcher: Optimize MMX pshufw
2026-05-05 08:41:22 -04:00
Ryan Houdek 694e68b838 Merge pull request #5468 from peppergrayxyz/proc_self_stat
read /proc/self/stat using %lu
2026-05-04 20:21:16 -07:00
Pepper Gray 78320e1433 read /proc/self/stat using %lu
building on clang/musl causes warnings:

```
FEX/Source/Tools/FEXInterpreter/ELFCodeLoader.h:782:29: warning: format specifies type 'unsigned long long *' but the argument has type 'uint64_t *' (aka 'unsigned long *') [-Wformat]
  776 |                             "%llu %llu %llu %*u %*u "   // 26 to 30
      |                              ~~~~
      |                              %lu
  777 |                             "%*u %*u %*u %*u %*u "      // 31 to 35
  778 |                             "%*u %*u %*d %*d %*u "      // 36 to 40
  779 |                             "%*u %*u %*u %*d %llu "     // 40 to 45
  780 |                             "%llu %llu %llu %llu %llu " // 46 to 50
  781 |                             "%llu",                     // 51
  782 |                             &map.start_code, &map.end_code, &map.start_stack, &map.start_data, &map.end_data, &map.start_brk,
      |                             ^~~~~~~~~~~~~~~
```

according to the [man page](https://man7.org/linux/man-pages/man5/proc_pid_stat.5.html)
`/proc/self/stat` uses `%lu`:

read the values as unsigned long (%lu) and then write them to
prctl_mm_map (platform specific format).

fixes: #5467
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-05 04:24:31 +02:00
Ryan Houdek 082e7b2695 InstcountCI: Update 2026-05-04 17:52:54 -07:00
Ryan Houdek cdbffb80c7 unittests: Adds a full coverage pshufw test 2026-05-04 17:52:53 -07:00
Ryan Houdek cb6c8cce55 OpcodeDispatcher: Optimize MMX pshufw
Found through writing a shuffle solver rather than an LLM.

Fixes #3785
2026-05-04 17:52:53 -07:00
Ryan Houdek 47e173e549 Merge pull request #5472 from peppergrayxyz/format
use portable format specifiers
2026-05-04 14:18:28 -07:00
Ryan Houdek 162bd4be97 win32: Query DCZID_EL0 so clzero works
EL0 registers are readable without going through the registry, but we
were failing to populate this register, which was causing clzero to not
be supported.
2026-05-04 12:42:18 -07:00
Ryan Houdek f0764aeafe Merge pull request #5480 from neobrain/refactor_musl_sigmask
SignalDelegator: Simplify support for musl's sigset_t
2026-05-04 10:39:42 -07:00
Ryan Houdek 1efed71696 Merge pull request #5479 from neobrain/refactor_drop_compile_service
FEXCore: Drop unused CompileService
2026-05-04 10:36:44 -07:00
Tony Wasserka 06d77c1c19 SignalDelegator: Simplify support for musl's sigset_t 2026-05-04 16:28:47 +02:00
Tony Wasserka 942d0c631a FEXCore: Drop unused CompileService 2026-05-04 15:52:34 +02:00
Pepper Gray abae5dd93b use portable format specifiers
building with clang/musl causes these warnings:

```
FEX/Source/Tools/FEXServer/ProcessPipe.cpp:100:96: warning: format specifies type 'ssize_t' (aka 'long') but the argument has type 'rlim_t' (aka 'unsigned long long') [-Wformat]
```

- cast platform specific MaxFDs members to uintmax_t and print as PRIuMAX
- use %zu for GetNumFilesOpen (size_t)

fix: #5471
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-04 11:20:55 +02:00
Ryan Houdek 93015a0266 Merge pull request #5458 from peppergrayxyz/largefile64
make 64bit symbols visible to enhance portability (musl)
2026-05-03 23:36:59 -07:00
Ryan Houdek d238db69d3 Merge pull request #5477 from peppergrayxyz/unistd
include missing header unistd.h
2026-05-03 15:18:41 -07:00
Ryan Houdek e0ead236b6 Merge pull request #5473 from peppergrayxyz/ObjectCacheRefCounter
remove dead code (ObjectCacheRefCounter)
2026-05-03 15:16:57 -07:00
Pepper Gray c548262664 include missing header unistd.h
build on clang/musl fails with:

```
FEX/unittests/APITests/Allocator.cpp:17:5: error: use of undeclared identifier 'close'
FEX/unittests/APITests/Allocator.cpp:23:5: error: use of undeclared identifier 'lseek'; did you mean 'fseek'?
FEX/unittests/APITests/Allocator.cpp:23:11: error: cannot initialize a parameter of type 'FILE *' (aka 'struct _IO_FILE *') with an lvalue of type 'int'
FEX/unittests/APITests/Allocator.cpp:24:5: error: use of undeclared identifier 'write'; did you mean '_IO_cookie_io_functions_t::write'?
FEX/unittests/APITests/Allocator.cpp:24:5: error: invalid use of non-static data member 'write'
```

include header to provide defintions

fix: #5476
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-03 15:49:40 +02:00
Pepper Gray bfc51e577f remove dead code (ObjectCacheRefCounter)
building using libc++ failes due to shared_mutex ObjectCacheRefCounter
inflating InternalThreadState beyond FEX_PAGE_SIZE, thus triggering
the static assert:

```
FEXCore/Debug/InternalThreadState.h:133:15: error: static assertion failed
FEX/FEXCore/include/FEXCore/Debug/InternalThreadState.h:133:145: note: expression evaluates to '7680 < 4096'
FEX/FEXCore/include/FEXCore/Debug/InternalThreadState.h:136:58: note: expression evaluates to '12288 == 8192'
```

remove `ObjectCacheRefCounter` as it is not used anywhere.

fix: #5456
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-03 15:13:47 +02:00
Ryan Houdek 85773995e1 Merge pull request #5470 from peppergrayxyz/header_redirect
fix include redirect for <poll.h> and <signal.h>
2026-05-03 04:41:18 -07:00
Pepper Gray 00b4777290 fix include redirect for <poll.h> and <signal.h>
building on musl/clang causes redirecting incorrect #includes warnings:

```
warning: redirecting incorrect #include <sys/poll.h> to <poll.h> [-W#warnings]
warning: redirecting incorrect #include <sys/signal.h> to <signal.h> [-W#warnings]
```

include headers instead of sys/headers.

fix: #5469
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-03 13:22:32 +02:00
Ryan Houdek 197e6de194 Merge pull request #5462 from peppergrayxyz/uc_sigmask
determine sigset_t fieldname to enhance portability (musl)
2026-05-03 03:48:53 -07:00
Ryan Houdek 9908ea4c2f Merge pull request #5466 from peppergrayxyz/libgen_h
include <libgen.h> for basename
2026-05-03 03:46:43 -07:00
Ryan Houdek c402b15bd3 Merge pull request #5464 from peppergrayxyz/tgkill
add header and classpath for tgkill
2026-05-03 03:46:07 -07:00
Ryan Houdek a3f3118ac5 Merge pull request #5460 from peppergrayxyz/sigset_t
use <signal.h> instead of glibc header to enhance portability (musl)
2026-05-03 03:18:20 -07:00
Pepper Gray b1aab0e498 determine sigset_t fieldname to enhance portability (musl)
musl build fails due to access to internal glibc member:

```
FEX/Source/Tools/LinuxEmulation/LinuxSyscalls/SignalDelegator.cpp:655:39: error: no member named '__val' in '__sigset_t'
  655 |       .SigMask = _context->uc_sigmask.__val[0],
      |                  ~~~~~~~~~~~~~~~~~~~~ ^
1 error generated.
```

add check to determine private glibc or musl member name or throw an error.

fixes: #5461
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-03 12:17:09 +02:00
Pepper Gray 927c0ce54a include <libgen.h> for basename
building on clang/musl build fails due to missing symbol:

```
/run/host/home/pepper/dev/ghrepos/FEX/Source/Tools/FEXOfflineCompiler/Main.cpp:321:35: error: use of undeclared identifier 'basename'
  321 |   auto CommandName = std::string {basename(argv[0])} + " " + (argc > 1 ? argv[1] : "");
      |                                   ^~~~~~~~
/run/host/home/pepper/dev/ghrepos/FEX/Source/Tools/FEXOfflineCompiler/Main.cpp:327:43: error: use of undeclared identifier 'basename'
  327 |     fmt::print("Usage: {} <command>\n\n", basename(argv[0]));
      |                                           ^~~~~~~~
```

include missing header.

fixes: #5465
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-03 12:10:07 +02:00
Pepper Gray 57d9dc037d add header and classpath for tgkill
building on clang/musl fails with:

```
FEX/Source/Tools/LinuxEmulation/LinuxSyscalls/GdbServer.cpp:1174:5: error: use of undeclared identifier 'tgkill'
 1174 |     tgkill(::getpid(), ::getpid(), SIGKILL);
      |     ^~~~~~
1 error generated.
```

include and use `FHU::Syscalls::tgkill`.

fixes: #5463
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-03 11:50:23 +02:00
Pepper Gray e3cfe28848 use <signal.h> instead of glibc header to enhance portability (musl)
building using musl fails with:

```
FEX/Source/Tools/LinuxEmulation/LinuxSyscalls/ThreadManager.h:33:10: fatal error: 'bits/types/sigset_t.h' file not found
   33 | #include <bits/types/sigset_t.h>
      |          ^~~~~~~~~~~~~~~~~~~~~~~
1 error generated.
```

`<bits/types/sigset_t.h>` is a glibc internal header, use <signal.h> instead.

fixes: #5459
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-03 11:03:38 +02:00
Pepper Gray ac91f583b8 make 64bit symbols visible to enhance portability (musl)
building using clang/musl fails with:

```
FEX/Source/Tools/LinuxEmulation/LinuxSyscalls/x32/Types.h:548:5: error: member access into incomplete type 'const struct statfs64'
  548 |     COPY(f_bsize);
      |     ^
```

add `_LARGEFILE64_SOURCE` to define large-file feature macros

fix: #5457
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-03 10:52:09 +02:00
Ryan Houdek 215658bf29 Merge pull request #5455 from peppergrayxyz/sys_prctl
use <sys/prctl.h> to enhance portability (clang)
2026-05-02 15:11:33 -07:00
Pepper Gray 92b1a6ea8a use <sys/prctl.h> to enhance portability (clang)
using <sys/prctl.h> and <linux/prctl.h> simultaneously causes clang to
fail:

```
In file included from FEX/FEXCore/Source/Utils/AllocatorHooks.cpp:6:
/usr/include/sys/prctl.h:88:8: error: redefinition of 'prctl_mm_map'
   88 | struct prctl_mm_map {
      |        ^
/usr/include/linux/prctl.h:134:8: note: previous definition is here
  134 | struct prctl_mm_map {
      |        ^
1 error generated.
```

prefer <sys/prctl.h> and do not include <linux/prctl.h>

fix: #5454
Signed-off-by: Pepper Gray <hello@peppergray.xyz>
2026-05-02 15:36:40 +02:00
Ryan Houdek 8ab00758be Merge pull request #5448 from bylaws/claudefix6
X87: Fix FXTRACT with Inf and NaN inputs
2026-04-30 14:42:08 -07:00
Ryan Houdek e91bda7765 Merge pull request #5447 from bylaws/claudefix5
SoftFloat: Fix FSCALE(0, +Inf) to raise IE and return a quiet NaN
2026-04-30 14:41:24 -07:00
Ryan Houdek 9db211ac97 Merge pull request #5446 from bylaws/claudefix4
OpcodeDispatcher: Fix FIST 16-bit IE for denormal inputs
2026-04-30 14:40:41 -07:00
Ryan Houdek b1381fd3b7 Merge pull request #5444 from bylaws/claudefix2
VectorOps: Fix MAXPS/MAXPD NaN and signed-zero tie selection
2026-04-30 14:39:56 -07:00
Ryan Houdek e5f6a7d85e Merge pull request #5450 from neobrain/fix_code_cache_portable
CodeCache: Make FEXServer call FEXOfflineCompiler by absolute path
2026-04-30 14:17:16 -07:00
Ryan Houdek feae76fc4f Merge pull request #5451 from Sonicadvance1/146
arm64ec: Fixes crash in many games with SDL+Dualsense
2026-04-30 14:16:06 -07:00
Ryan Houdek 015f3cffb9 arm64ec: Fixes crash in many games with SDL+Dualsense
We were pointing to an incorrect function pointer and exploding when a
pending suspend doorbell had occured.
2026-04-30 12:59:53 -07:00
Tony Wasserka 3e5c17ae80 CodeCache: Make FEXServer call FEXOfflineCompiler by absolute path
In portable mode, FEXOfflineCompiler may not be in PATH (and if it is, it's
most likely not a compatible version). Instead, use the executable next to
the FEXServer binary.
2026-04-30 17:14:22 +02:00
Billy Laws 3e278b42f8 InstcountCI: Update 2026-04-29 02:54:34 +00:00
Billy Laws fa953445e9 InstcountCI: Update 2026-04-29 02:53:07 +00:00
Billy Laws a412b1d3b7 InstcountCI: Update 2026-04-29 02:43:19 +00:00
Billy Laws fc680388e3 X87: Fix FXTRACT with Inf and NaN inputs
Both the default F80 softfloat wrappers (FXTRACT_SIG/FXTRACT_EXP) and
the reduced-precision F64 dispatcher fell through to the generic
exponent/significand extraction for Inf and NaN inputs, producing
finite garbage. The F80 wrapper returns input unchanged in the
significand slot and +Inf (or NaN) in the exponent slot; the F64
dispatcher detects the exponent-all-ones case and selects the proper
Inf/NaN result before the existing zero-case fold.
2026-04-29 02:34:41 +00:00
Billy Laws 7c260b45e1 unittests/ASM: Adds tests for FXTRACT Inf/NaN 2026-04-29 02:29:23 +00:00
Ryan Houdek 098c4c57b4 Merge pull request #5443 from bylaws/claudefix1
X86Features: Fix AVX/RAND/PCMLULQDQ feature detection on intel
2026-04-28 19:25:15 -07:00
Billy Laws 7c826e35b4 SoftFloat: Fix FSCALE(0, +Inf) to raise IE
The lhs==0 short-circuit in X80SoftFloat::FSCALE returned lhs
unchanged without calling extF80_mul, so the 0*Inf invalid-operation
case never set softfloat_flag_invalid. Detect +Inf rhs explicitly
in the zero-lhs path and raise the flag, returning QNaN to match
hardware.
2026-04-29 02:17:03 +00:00
Billy Laws 1bd2ff3fc3 unittests/ASM: Adds test for FSCALE(0, +Inf) raising IE 2026-04-29 02:16:57 +00:00
Billy Laws e2fc91fc15 OpcodeDispatcher: Fix FIST 16-bit IE for denormal inputs
The special-value check masked the biased exponent with 0x7fff and
then used TestNZ, so it triggered on Exp==0 (denormals) instead of
Exp==0x7fff (NaN/Inf). Replace the TestNZ with SubWithFlags against
0x7fff so EQ only fires for genuine NaN/Inf.
2026-04-29 02:10:35 +00:00
Billy Laws cf20647b25 unittests/ASM: Adds test for 16-bit FIST with denormal input not setting IE 2026-04-29 02:09:54 +00:00
Billy Laws 8d7071e549 VectorOps: Fix MAXPS/MAXPD NaN and signed-zero tie selection
fcmgt returns false on NaN, so the existing polarity in the non-SVE
fcmgt+bit sequences and in the SVE predicate-merge picked the wrong
source on NaN/tie. Swap the compare operands and flip bit<->bif / add
a predicate not to match x86 second source wins behaviour.
2026-04-29 02:00:58 +00:00
Billy Laws fb006b2c6d unittests/ASM: Adds test for MAXPS/MAXPD NaN and signed-zero tie 2026-04-29 01:57:41 +00:00
Billy Laws 9039eeb3cd X86Features: Fix AVX/RAND/PCMLULQDQ feature detection on intel 2026-04-29 01:57:18 +00:00
Billy Laws dd0702d30f unittests/ASM: Mark SSE4a/CLZERO as required for tests using them 2026-04-29 01:56:35 +00:00
LC 886faf0bd4 Merge pull request #5442 from neobrain/refactor_code_cache_check
CodeCache: Move bounds check to FEXOfflineCompiler
2026-04-28 19:47:52 -04:00
Tony Wasserka 86e28c6d34 CodeCache: Move bounds check to FEXOfflineCompiler
The previous check site would easily fail when loading caches for binaries
with multiple executable sections.

It makes much more sense to refuse generating caches anyway: The condition
effectively checked for invalid code map entries, so FEXOfflineCompiler
should reject them as bad inputs.
2026-04-28 17:31:20 +02:00
LC 821efab8aa Merge pull request #5440 from Sonicadvance1/145
OpcodeDispatcher: Ensure FIST* operations avoid TSO/GPR
2026-04-28 07:55:10 -04:00
Ryan Houdek 5295365dd0 InstcountCI: Update 2026-04-27 17:55:49 -07:00
Ryan Houdek 788959a98c OpcodeDispatcher: Ensure FIST* operations avoid TSO/GPR
Noticed this while benchmarking that the FIST* operations were
converting to a GPR, and then storing to memory using an atomic TSO
operation. This should be instead listening to the vector TSO
configuration option. This gives a 3.8x - 6.05x improvement in my
microbench.

Additionally when possible, make sure to use vector conversion
instructions when possible. It's lower cost to avoid the FPR->GPR
transfer, but we can only use it for 64-bit FIST operations. Microbench
couldn't show a difference for that on my platform, but that's because
it's float pipeline bounded regardless. Should help X-class Cortex and
newer Cortex-A.
2026-04-27 17:48:45 -07:00
LC dd145aaa88 Merge pull request #5439 from Sonicadvance1/144
Fix push/pop fs/gs segments and unittests
2026-04-27 19:43:44 -04:00
Ryan Houdek 34b3adc23d unittests/ASM: Adds unit test to ensure push/pop segment of o16 works
Only ensures we are pushing and popping the correct size, not any of the
selector data within it, as 64-bit systems with the FSGSBase extension
don't use them selectors anyway.

Can't test the 32-bit side currently because we would corrupt FS/GS in
CI and the host testharnessrunner can't fix that right now.
2026-04-27 15:21:47 -07:00
Simon Scherer ab14882761 FEXCore: Fix 2byte stack access for 0x66 PUSH/POP FS/GS 2026-04-27 15:21:11 -07:00
LC 7dc1f54fb6 Merge pull request #5435 from Sonicadvance1/143
unittests/ASM: Add unittests to ensure correct cmpxchg8b/16b flag setting
2026-04-25 10:34:07 -04:00
Ryan Houdek c09fb03eda unittests/ASM: Add unittests to ensure correct cmpxchg8b/16b flag setting 2026-04-25 00:46:40 -07:00
Ryan Houdek dbf2761fb7 InstcountCI: Update 2026-04-25 00:45:05 -07:00
Ryan Houdek fd1378f778 InstcountCI: Fix incorrect instruction 2026-04-25 00:43:50 -07:00
Simon Scherer 819dcee3ad FEXCore: Fix wrong shift value to extract NZCV in CmpPairZ 2026-04-25 00:41:24 -07:00
LC 4b02c04afc Merge pull request #5429 from Sonicadvance1/142
Steam/CompatTool: Fixes Graphics Provider path handling
2026-04-23 20:13:13 -04:00
Ryan Houdek deed99e7a3 Merge pull request #5425 from pmatos/f64-atan-fyl2x
JIT-inline F64ATAN and F64FYL2X for reduced precision x87 path
2026-04-23 15:16:04 -07:00
Ryan Houdek 7bffc4a177 Steam/CompatTool: Fixes Graphics Provider path handling
Graphics provider needs to be a path to a json file in the root of the
rootfs. Make sure to strip the filepath off to get the directory.

Misunderstood the assignment before.
2026-04-23 15:13:42 -07:00
Ryan Houdek 701555e400 Merge pull request #5428 from Sonicadvance1/141
Steam/CompatTool: Support `STEAM_COMPAT_GRAPHICS_PROVIDER` for rootfs path
2026-04-21 12:50:22 -07:00
Ryan Houdek 49fa86d0b5 Steam/CompatTool: Support STEAM_COMPAT_GRAPHICS_PROVIDER for rootfs path
If we have been provided a graphics provider path through an environment
variable, then use that path directly rather than searching.
2026-04-21 12:37:40 -07:00
Paulo Matos adbace8810 instcountci: JIT-inline F64ATAN and F64FYL2X for reduced precision x87 path 2026-04-21 20:23:45 +02:00
Paulo Matos 050138bcea asm_tests: JIT-inline F64ATAN and F64FYL2X for reduced precision x87 path 2026-04-21 20:23:45 +02:00
Paulo Matos db50b04663 JIT-inline F64ATAN and F64FYL2X for reduced precision x87 path 2026-04-21 12:09:55 +02:00
LC 59755ec115 Merge pull request #5426 from Sonicadvance1/139
Snapdragon X2 Elite fixes
2026-04-20 11:53:51 -04:00
Ryan Houdek 856fb1e636 Snapdragon X2 Elite fixes
- RNDRRS is still broken on this CPU
- Fault granularity checking needs to use loads
  - stlxp does monitor check before alignment check, use loads for all
    for consistency
  - Fault granularity is still 16B which matches LSE2 requirements.
- X2E still only ships a 19.2Mhz cycle counter, so still only
  ARMv9.0-a hardware
- Adds product name to CPUID
  - Oryon-3 being CPU PartID 2 isn't a mistake.
  - No distinction between Oryon-1 and Oryon-2, both are partid 1.

cpuinfo:
```
processor   : 0
BogoMIPS    : 38.40
Features    : fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm uscat ilrcpc flagm ssbs sb paca pacg dcpodp sve2 sveaes svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 rng ecv afp rpres
CPU implementer   : 0x51
CPU architecture: 8
CPU variant : 0x1
CPU part    : 0x002
CPU revision      : 1
```
2026-04-19 14:34:44 -07:00
Ryan Houdek d41d52b889 Merge pull request #5419 from pmatos/f64-scale-f2xm1
JIT-inline F64SCALE and F64F2XM1 for reduced precision x87 path
2026-04-17 14:38:22 -07:00
Paulo Matos 18f69fb16d instcountci: JIT-inline F64SCALE and F64F2XM1 for reduced precision x87 path 2026-04-17 17:28:13 +02:00
Paulo Matos d165711f2e asm_tests: JIT-inline F64SCALE and F64F2XM1 for reduced precision x87 path 2026-04-16 11:24:24 +02:00
Paulo Matos 6bf19ffe37 JIT-inline F64SCALE and F64F2XM1 for reduced precision x87 path 2026-04-16 11:24:19 +02:00
LC 441116e1e6 Merge pull request #5421 from Sonicadvance1/136
Scripts: Move arch check first in InstallFEX
2026-04-15 16:24:04 -04:00
Ryan Houdek ce97ef0ab1 Scripts: Move arch check first in InstallFEX
Don't give people false hope that the script might work on distros that
aren't Ubuntu.

Fixes #5420
2026-04-15 13:01:37 -07:00
LC 2ea0de92f4 Merge pull request #5418 from Sonicadvance1/135
ArchHelpers: Allow atomic memory operations in non-JIT handler
2026-04-15 07:22:25 -04:00
Ryan Houdek 14580c4675 ArchHelpers: Allow atomic memory operations in non-JIT handler
`Detroit: Become Human` decided to use unaligned CriticalSections. So
this workarounds that.
2026-04-14 13:38:41 -07:00
Ryan Houdek 9681559d56 Docs: Update for release FEX-2604 2026-04-09 13:45:35 -07:00
Ryan Houdek b478e4845f Merge pull request #5417 from tiopex/main
FEXRootFSFetcher: clear Unknown when distro is set on the CLI
2026-04-09 13:42:55 -07:00
tpietrus 1fa5104076 FEXRootFSFetcher: clear Unknown when distro is set on the CLI 2026-04-09 07:53:53 +02:00
LC ce65f5376f Merge pull request #5415 from Sonicadvance1/133
FEX: Workaround Docker seccomp bug
2026-04-08 11:39:56 -04:00
LC 0695249fc8 Merge pull request #5413 from Sonicadvance1/132
Arm64EC: Invert suspend doorbell and move out of hot path
2026-04-06 19:49:49 -04:00
LC 51144c99a7 Merge pull request #5408 from Sonicadvance1/129
FEXCore: Use WritePriorityMutex for Code Invalidation Mutex
2026-04-06 19:49:07 -04:00
LC 251398a7cb Merge pull request #5416 from Sonicadvance1/134
OpcodeDispatcher: Fixes nop encoded prefetch instruction
2026-04-06 19:45:52 -04:00
Ryan Houdek 2e6a7f869c OpcodeDispatcher: Fixes nop encoded prefetch instruction
We had a bug where nop encoded prefetch instructions were getting
flagged as illegal instructions erroneously. Fix that and add a unittest
for ensuring execution.

Fixes `Devil May Cry 4`
2026-04-06 11:04:46 -07:00
Ryan Houdek f308162334 FEX: Workaround Docker seccomp bug
Docker's seccomp filter fails to follow AAPCS64 and SysV zero-extension
rules.  For values smaller than 64-bit they were required in their
seccomp filters to truncate the value to the specific size but do not.
Instead they do a 64-bit comparison operation against smaller arguments
(in this case 32-bit). This means 64-bit -1 and 32-bit -1 passed through
have different values for this `personality` syscall.

The real fix would be for Docker to audit their seccomp filter rules and
ensure they zero-extend every argument that is smaller than 64-bit, but
we don't control that. So there is likely to be more bugs in their
filter that we encounter, this is just an easy one to resolve.
2026-04-06 09:51:15 -07:00
Ryan Houdek db4867839c Arm64EC: Invert suspend doorbell and move out of hot path
This was causing a surprisingly high amount of branch mispredicts in
Death Stranding. Suspend doorbell is fairly rare so just invert the
check and move the target down out of the hot path. Then the doorbell
handling code will trampoline to the correct location still.
2026-04-03 20:07:37 -07:00
LC 73ffff7d22 Merge pull request #5409 from Sonicadvance1/130
OpcodeDispatcher: Special case optimize a broadcast
2026-04-03 20:23:39 -04:00
LC dc48a4f73c Merge pull request #5406 from Sonicadvance1/128
FEXRootFSFetcher: Improve hashing performance
2026-04-03 00:26:44 -04:00
Ryan Houdek efbccccdc0 InstcountCI: Update 2026-04-02 18:45:22 -07:00
Ryan Houdek 3e7cd88dcc OpcodeDispatcher: Special case optimize a broadcast
Death Stranding 2 is using this instruction instead of `vbroadcastss`
for some reason. Optimize its specific case and add a note that when we
know sources match that we can optimize more patterns easily.
2026-04-02 18:45:22 -07:00
Ryan Houdek 12cfe8fc37 InstcountCI: Add instruction found in Death Stranding 2 2026-04-02 18:25:19 -07:00
Tony Wasserka c6d2ce043f Merge pull request #5388 from Sonicadvance1/123
Config: Finish wiring up Regex app overrides
2026-04-02 11:00:37 +02:00
Tony Wasserka 5c34c574c8 Merge pull request #5364 from Sonicadvance1/110
Win32: Enable support for virtual naming and THP control
2026-04-02 10:58:30 +02:00
Ryan Houdek d3cfdcb431 Win32: Enable support for virtual naming and THP control
Allows WTF to work (mostly) with Wine by letting us VirtualName things,
and also allows madvise control of THP, which significantly cuts back
memory usage.

This works around the problem of Wine not giving us control of this by
using raw syscalls when wine is detected.

Based on top of #5362 so the THP disable controls are in.
2026-04-01 10:50:40 -07:00
Ryan Houdek 4018d23c39 Config: Finish wiring up Regex app overrides
This wasn't quite wired up exactly how we wanted it. It was previously
matching against the opaque file config handle, which can be anything.

Instead compare it to the appname that now gets passed over to it for
matching.

This allows us to do the following:
```
{
    "Config": {
        "ProfileStats": "1",
        "X87ReducedPrecision": "1",
        "TSOEnabled": "1",
        "VectorTSOEnabled": "0",
        "MemcpySetTSOEnabled": "0",
        "HalfBarrierTSOEnabled":"1",
        "MaxInst": "500",
        "Multiblock": "1"
    },
    "AppOverrides" : {
        "setup*" : {
            "Comment": [
                "292030 - The Witcher 3: Wild Hunt"
            ],
            "X87ReducedPrecision": "0"
        }
     }
}
```

Based on #121 which needs to get merged first.

Code Review

Code Review: Class deletion
2026-04-01 10:46:02 -07:00
badumbatish 81d4e8fe9d Initial implementation for regex engine
Add support for question mark and plus mark in regex, supply testing for star

Added more characters to the regex alphabets, add more test case

Added support for regex matching of configs, awaiting reviews

Rename variable to CamelCase

Addresses PR reviews

Remove unnecessary features and test cases

Rewrite to naive regex with dp

Addresses PR reviews

Build fixes

Code Review
2026-04-01 10:45:59 -07:00
Tony Wasserka 34b48c4069 Merge pull request #5383 from Sonicadvance1/119
FEXGetConfig: Test for showing fault granularity
2026-04-01 10:48:18 +02:00
Ryan Houdek 01a3ab6ca7 FEXCore: Use WritePriorityMutex for Code Invalidation Mutex
The previous `ForkableSharedMutex` using `std::shared_mutex` was showing
up as significant CPU time on arm64ec. In particular it was showing up
upwards of 700ms/S of CPU time for read-contented workloads at only
~2700 locks per second. The libc++ implementation for arm64ec must be
particularly gnarly for this to be so slow.

With this swapped over, it's now only spending around 56ms/S on the
contended shared lock, but at ~8000 locks per second. So a significant
uplift.
2026-03-31 19:02:51 -07:00
Ryan Houdek 476c242d7f FEXCore: Moves SpinWaitLock and WritePriorityMutex to frontend visible includes
This will be used in a moment.
2026-03-31 19:02:51 -07:00
LC ae3fa6a836 Merge pull request #5403 from Sonicadvance1/127
IR: Adds support for printing strings
2026-03-31 16:23:42 -04:00
Ryan Houdek ba93bdd66d FEXRootFSFetcher: Improve hashing performance
Don't use pread, instead map the file and madvise larger blocks. This
removes copying overhead as its just mapping file pages in instead.
Also splits the implementation of file reading from hashing to make
tinkering less involved, as if I want more performance out of this (say
due to live hashing) then it's easier to tinker.

Improves hashing performance from ~2.2GB/s to ~3.6GB/s on my system,
which is CPU bounded by xxhash here.
2026-03-30 18:16:46 -07:00
Ryan Houdek 1fa0b37fac FEXGetConfig: Test for showing fault granularity
Useful for seeing if behaviour has changed. Useful with the
`--tso-emulation-info` option to show hardware behaviour
2026-03-30 12:32:00 -07:00
LC 5b4a5969cc Merge pull request #5401 from Sonicadvance1/126
Linux/GuestFrames: Ensure MXCSR is saved and restored on signal
2026-03-28 22:25:59 -04:00
Ryan Houdek 941a7934ef InstcountCI: Update 2026-03-28 18:06:54 -07:00
Ryan Houdek 474439ab4b InstcountCI: Update 2026-03-28 17:56:23 -07:00
Ryan Houdek 4428aea5ee IR: Adds support for printing strings
I utilize this functionality quite heavily when debugging and I need
bread crumbs spread around. Instead of reimplementing it a dozen times,
just have it upstreamed.
2026-03-28 17:53:33 -07:00
Ryan Houdek e9a9cc5bc3 unittests/FEXLinuxTests: Adds an MXCSR signal test
Ensures that the MXCSR value stays the same with a signal inbetween that
modifies it.
2026-03-28 17:49:12 -07:00
Ryan Houdek 2291c5b230 OpcodeDispatcher/Vector: Make sure MXCSR is masked
We don't support the exception bits, make sure these are masked off so
spurious exception checks don't break.
2026-03-28 17:47:55 -07:00
Ryan Houdek f78e194cf7 Linux/GuestFrames: Ensure MXCSR is saved and restored on signal
This was causing an unfortunate set of circumstances where Dark Souls
III was modifying MXCSR and we weren't saving it, cause the value to change
from 0x9fc0 to 0.

This "enabled" float exceptions by unmasking the exception masks in
MXCSR. This in turn had Dark Souls III's `expf` function to fault out,
as it checks if the MXCSR exception masks are set or not for determining
if underflow should assert or not.

Wow64/arm64ec has a similar problem where it always sets back to default
on signal. Which means game lose DAZ, but I'm not fixing that bug right
now.

Fixes #5391
2026-03-28 17:44:45 -07:00
LC b77ddcf1a7 Merge pull request #5398 from Sonicadvance1/125
Allocators: Remove legacy NOREPLACE handling
2026-03-27 08:14:05 -04:00
Ryan Houdek 3ef677537d Merge pull request #5397 from neobrain/fix_elfreads
LinuxSyscalls: Skip reading ELF files when code caching is disabled
2026-03-26 14:33:26 -07:00
Tony Wasserka 1df1265ed1 LinuxSyscalls: Skip reading ELF files when code caching is disabled
ELF headers were read unconditionally because doing so was assumed to be cheap
(as the guest app would read them anyway shortly after). However, relocation
parsing was added since then, which has less predictable performance due to
crossing page boundaries and reading larger amounts of memory. It might be
possible to make the underlying code more efficient, but until that's done
it's better to skip this logic unless needed.

Closes #5390
2026-03-26 21:40:20 +01:00
Ryan Houdek f0854a16fe Allocators: Remove legacy NOREPLACE handling
We needed this handling on old kernels that didn't understand the
NOREPLACE flag. We no longer support kernels this old, so remove some of
this vestigial code.
2026-03-26 13:36:29 -07:00
Ryan Houdek 6bd476fb03 Merge pull request #5395 from lioncash/cpuid
CPUID: Add basic stub handling for AVX10 info
2026-03-25 17:16:24 -07:00
Lioncache b9c0af7c3b CPUID: Add basic handling for AVX10 info
Just gets the feature bit handling stuff in place for various
facilities, so it can be easily expanded in the future.
2026-03-25 19:18:05 -04:00
Tony Wasserka 5149ebc70e Merge pull request #5389 from Sonicadvance1/124
SMCTracking: Remove relocation log
2026-03-24 11:12:10 +01:00
Ryan Houdek e92a6a5803 SMCTracking: Remove relocation log
Holy jeez does this thing spam.
2026-03-23 19:55:11 -07:00
Ryan Houdek 6da963a695 Merge pull request #5394 from neobrain/fix_eager_relocation_parsing
LinuxSyscalls: Don't try to parse relocations unless ELF parsing succeeded
2026-03-23 14:01:20 -07:00
Tony Wasserka 2a74489858 LinuxSyscalls: Don't try to parse relocations unless ELF parsing succeeded 2026-03-23 17:47:47 +01:00
Ryan Houdek 65a436ca98 Merge pull request #5392 from neobrain/fix_code_cache_dupfd
LinuxSyscalls: Re-open file descriptors for parsing ELF headers
2026-03-23 09:42:45 -07:00
LC bc533c8050 Merge pull request #5393 from neobrain/fix_code_cache_glibc_assert
CodeCache: Fix glibc debug mode assertion
2026-03-22 14:07:27 -04:00
Tony Wasserka fd6cea4698 CodeCache: Fix glibc debug mode assertion
If begin == end, the first vector::erase() call would invalidate the begin
iterator.
2026-03-22 10:11:53 +01:00
Tony Wasserka a1aa1658ec LinuxSyscalls: Re-open file descriptors for parsing ELF headers
File descriptors returned by dup() share state with the original FD, so we
need to use open() to create a fully independent object.

Fixes #5379
2026-03-22 10:09:37 +01:00
LC 5c4c468d13 Merge pull request #5387 from Sonicadvance1/122
github: Stop running unittests always on build failure
2026-03-20 00:58:09 -04:00
Ryan Houdek 69ef1658cc github: Stop running unittests always on build failure
This was taking too much time.
2026-03-19 20:08:10 -07:00
Ryan Houdek 8c72aa76a0 Merge pull request #5385 from lioncash/ilog
MemoryOps: Collapse duplicate add/sub in Memset
2026-03-19 18:59:17 -07:00
Lioncache 928a932a43 MemoryOps: Collapse duplicate add/sub in Memset
We can just use ilog2 to deduplicate this a bit.
2026-03-19 21:34:03 -04:00
Ryan Houdek 83601055dc Merge pull request #5368 from neobrain/feature_cc_elf_relocations
CodeCache: Support ELF relocations
2026-03-19 18:10:09 -07:00
Ryan Houdek 68480f6e43 Merge pull request #5356 from lioncash/mops
MemoryOps: Drop MOPS handling into place for MemSet/MemCpy
2026-03-19 18:04:14 -07:00
Ryan Houdek 42291540ab Merge pull request #5343 from pmatos/f64-sin-cos-tan
JIT-inline F64SIN, F64COS, and F64TAN for reduced precision x87 path
2026-03-19 17:51:50 -07:00
LC 194eb69838 Merge pull request #5384 from Sonicadvance1/120
Cmake: Default to release builds with a message
2026-03-19 19:42:39 -04:00
Ryan Houdek c1d27fa453 Cmake: Default to release builds with a message
People keep forgetting to set this and have a worse experience.
Default to a Release build, which ensures optimizations are enabled and
assertions are disabled.
2026-03-19 16:25:40 -07:00
Paulo Matos 9d5f7caa79 instcountci: JIT-inline F64SIN, F64COS, and F64TAN for reduced precision x87 path 2026-03-19 19:18:02 +01:00
Paulo Matos a1d78dceb0 asm_tests: JIT-inline F64SIN, F64COS, and F64TAN for reduced precision x87 path 2026-03-19 19:18:02 +01:00
Lioncache 2bcf435e0a MemoryOps: Handle overlapping memcpy 2026-03-19 14:13:36 -04:00
Paulo Matos 30e853305d JIT-inline F64SIN, F64COS, and F64TAN for reduced precision x87 path 2026-03-19 18:58:46 +01:00
Lioncache 600f4bb2b6 unittests: Add overlapping tests for MemCpy 2026-03-19 11:55:02 -04:00
Lioncache 5868814c91 MemoryOps: Drop in MOPS handling for MemCpy
With the MOPS featureset dropped in, we can also accelerate memcpy paths
on hardware that supports it.
2026-03-19 11:55:02 -04:00
Lioncache e862f8f86c unittests: Add specific paths for MOPS 2026-03-19 11:55:02 -04:00
Lioncache 85c1ecd035 MemoryOps: Handle inline values in MemSet() MOPS path
Lets us handle potential inline memset values.

Also fixes up the STOS tests to actually ensure all values
in the verification step pass.
2026-03-19 11:55:02 -04:00
Lioncache 68ad448672 MemoryOps: Drop 8-bit memset support into MemSet()
Can be further expanded to handle other optimization cases, but this
kicks it off for forward direction memsets at least.
2026-03-19 11:55:02 -04:00
LC c18fb3cb78 Merge pull request #5382 from Sonicadvance1/118
FEXpidof: Fixes another missing std::filesystem throw
2026-03-18 18:07:08 -04:00
Tony Wasserka 53702f989c Merge pull request #5362 from Sonicadvance1/108
FEX: Disable THP on key allocations that consume memory
2026-03-18 21:04:15 +01:00
Ryan Houdek c547b1bec3 FEX: Disable THP on key allocations that consume memory
Disables THP on some key locations that are fairly sparse
- rpmalloc
  - This is the big one as this allocates some heavy sparse buffers.
- CallRet stacks
  - These get in the hundreds of megabytes, while not being sparse they
    trend towards only using a handful of pages and ballooning to 2MB
    per thread is quite heavy.
- Lookup cache
  - L1 specifically gets hit here which adds a decent chunk of overhead
    due to sparsity.

Win32 for all of these also aren't handled, but that will need to be a
followup.
2026-03-18 12:15:39 -07:00
Ryan Houdek 740350c8ea FEXpidof: Fixes another missing std::filesystem throw
Turns out std::filesystem::exists throws as well if there was an
underlying OS API failure.
2026-03-18 12:10:22 -07:00
Tony Wasserka 9f9b20eac0 Merge pull request #5378 from Sonicadvance1/117
SMCTracking: Move read check up for ELF parsing
2026-03-18 12:04:23 +01:00
Tony Wasserka c67ffb82a8 Core: Support reporting blocks that are uncacheable due to unhandled ELF relocations 2026-03-18 11:59:44 +01:00
Tony Wasserka 1ea24f3d6c LinuxSyscalls: Enable delayed code cache load for ELF files
Specifically this is needed if any ELF relocations cover read-only code
sections, which is indicated in the ELF headers via DT_TEXTREL/DF_TEXTREL.
2026-03-18 11:59:41 +01:00
Tony Wasserka 226bd51afe LinuxSyscalls: Implement delayed cache load for binaries that require ELF/PE relocations 2026-03-18 11:58:26 +01:00
Tony Wasserka 293568be36 FEXOfflineCompiler: Apply relocations to loaded ELF binaries 2026-03-18 11:57:38 +01:00
Tony Wasserka 152fe81d16 LinuxSyscalls: Parse and provide ELF relocation information to the JIT 2026-03-18 11:57:28 +01:00
Ryan Houdek 494dd64c50 Merge pull request #5372 from CxnYusuf/add-fisttp-tests
Tests/ASM: Add FISTTP unit tests for 16, 32, 64-bit as well as negatives
2026-03-17 16:55:56 -07:00
Ryan Houdek 56de0d1ab4 Merge pull request #5377 from Sonicadvance1/116
FEXServer: Try both fusermount and fusermount3
2026-03-17 16:55:42 -07:00
Ryan Houdek 73c1f4cc54 Merge pull request #5374 from neobrain/fix_gcc_build
Fix most GCC build issues
2026-03-17 16:55:23 -07:00
Ryan Houdek 6a6a82385e Merge pull request #5381 from neobrain/fix_jit_restarts
JIT: Reset relocations on restart
2026-03-17 13:47:42 -07:00
Tony Wasserka fc8ef0e723 JIT: Reset relocations on restart 2026-03-17 21:37:01 +01:00
Ryan Houdek de11c05d2a Merge pull request #5380 from neobrain/fix_cc_32bit_constants
Arm64Emitter: Disable 32-bit constant optimization when NOP-padding is requested
2026-03-17 13:34:20 -07:00
Tony Wasserka addbc8cad8 Arm64Emitter: Disable 32-bit constant optimization when NOP-padding is requested
Code caching requires this even for simple libraries like libdl.so (as observed
in the 32-bit build of Super Meat Boy).
2026-03-17 21:06:13 +01:00
Ryan Houdek 63e37b7cbb SMCTracking: Move read check up for ELF parsing 2026-03-17 12:43:54 -07:00
Ryan Houdek 9ee329034f FEXServer: Try both fusermount and fusermount3
Apparently some distros don't symlink these, so try both with the newer
fusermount3 going first as its the common path now.

Fixes #5375
2026-03-17 12:36:49 -07:00
Ryan Houdek f3e904207b Merge pull request #5376 from OFFTKP/flag
Add test for shifts preserving flags
2026-03-17 09:40:59 -07:00
LC e4ae6ce635 Merge pull request #5373 from Sonicadvance1/115
code-format-helper: Another dependabot upgrade
2026-03-17 10:10:46 -04:00
Paris Oplopoios f7d76255ad Add test for shift preserving flags
Signed-off-by: Paris Oplopoios <21157395+OFFTKP@users.noreply.github.com>
2026-03-17 15:50:33 +02:00
Tony Wasserka ea45f9c694 FEXCore/VectorRegType: Use vector_size on GCC
GCC does not support neon_vector_type and silently ignores that attribute,
but vector_size(16) seems to have the same effect.
2026-03-16 19:15:05 +01:00
Tony Wasserka 1d449c0f58 FEXCore/Utils: Add quotes around preprocessor errors 2026-03-16 19:15:05 +01:00
Tony Wasserka 9ecc991043 LibraryForwarding: Remove unnecessary const qualifier 2026-03-16 19:15:05 +01:00
Tony Wasserka 4ef834859c SignalDelegator: Don't use the same name for two different symbols 2026-03-16 19:15:05 +01:00
Tony Wasserka fa082bc5c4 FileManagement: Fix ambiguous name reference 2026-03-16 19:15:05 +01:00
Tony Wasserka 67caab026a OpcodeDispatcher: Fix inconsistent types in ternary conditional 2026-03-16 19:15:05 +01:00
Tony Wasserka 91c55facb2 IRDumper: Avoid passing packed member data as references
Fixes "Cannot bind packed field to reference type" build errors on GCC.
2026-03-16 19:15:05 +01:00
Tony Wasserka 321d4d84d7 LinuxSyscalls: Don't cast away qualifiers 2026-03-16 19:15:05 +01:00
Tony Wasserka 22faa58e0b X86Tables: Use explicit type for SecondInstGroupOps definition
GCC considers it a "conflicting declaration" to use auto for a variable that
was already declared before.
2026-03-16 19:15:05 +01:00
Tony Wasserka ebd559f662 Core: Fix offsetof with runtime array indexes
GCC does not support this clang-specific language extension.
2026-03-16 19:15:05 +01:00
Tony Wasserka d27c9d3f98 CodeEmitter: Fix ambigious ExtendedType declaration 2026-03-16 18:50:01 +01:00
Tony Wasserka fbef482265 CMake: Link against libatomic if compiling with GCC 2026-03-16 18:50:01 +01:00
Tony Wasserka a57926ac57 CMake: Explicitly demote -Wchanges-meaning diagnostics to warnings on GCC 2026-03-16 18:50:01 +01:00
Ryan Houdek 70a7137e62 code-format-helper: Another dependabot upgrade 2026-03-15 20:45:17 -07:00
CxnYusuf 2d3a08a362 Tests/ASM: Add FISTTP unit tests for 16, 32, 64-bit as well as negatives 2026-03-16 03:10:38 +01:00
LC cc02edb3f6 Merge pull request #5371 from Sonicadvance1/114
Syscalls: Fixes crash in ELF parsing code
2026-03-15 20:36:00 -04:00
Ryan Houdek f894cd90f3 Merge pull request #5369 from Sonicadvance1/113
FEXCore: Update CPU frequency to be 64-bit
2026-03-15 15:20:28 -07:00
Ryan Houdek 24675969cc Merge pull request #5366 from Sonicadvance1/112
JIT: Use struct for Spill/Fill default arguments
2026-03-15 15:20:13 -07:00
Ryan Houdek a0cba1194e Syscalls: Fixes crash in ELF parsing code
When an application maps a file as PROT_NONE, we can't check if it is an
ELF. Was causing a crash in `Cisco Packet Tracer`.
2026-03-15 15:16:29 -07:00
Ryan Houdek 0952fa95f4 FEXCore: Update CPU frequency to be 64-bit
This annoyed by by trying to do math on a 32-bit value and it
overflowing. Just make it 64-bit.
2026-03-13 17:04:25 -07:00
Ryan Houdek 6177ab957b Merge pull request #5363 from Sonicadvance1/109
External/code-format-helper: Update dependencies
2026-03-13 11:52:07 -07:00
Ryan Houdek 5c1300a2b9 Merge pull request #5365 from Sonicadvance1/111
Config: Fix issue with config overrides
2026-03-13 11:51:51 -07:00
Ryan Houdek af9dd0827a JIT: Use struct for Spill/Fill default arguments
Cleans up the interface and makes the arguments explicit about what
they're setting. As promised from #5317
2026-03-12 19:25:12 -07:00
Ryan Houdek dc0162122f Config: Fix issue with config overrides
Accidentally was checking for Config override in the combination of
portable config and `FEX_APP_CONFIG_LOCATION`.

Fixes an early crash in PV.
2026-03-12 17:46:05 -07:00
Ryan Houdek ae491fb15b External/code-format-helper: Update dependencies
Removes dependabot alert.
2026-03-12 15:31:02 -07:00
Ryan Houdek 957c1fc420 Merge pull request #5357 from Sonicadvance1/106
Config: Enable Dynamic L1 and Disabled L2 caches by default
2026-03-11 14:12:38 -07:00
Ryan Houdek 86acfb35aa Config: Enable Dynamic L1 and Disabled L2 caches by default
Dramatically reduces memory consumption of FEX's per-thread lookup
structures. Primarily because L2 cache entirely goes away which can end
up reaching hundreds of megabytes or over a gigabyte of memory in some
cases, but also because L1 cache dynamically scales based on load.

Useful for conserving memory on systems with less than 16GB of RAM and
are UMA, like Asahi users inside of muvm.
2026-03-11 13:42:32 -07:00
LC e27d12ee5e Merge pull request #5359 from Sonicadvance1/107
OpcodeDispatcher: Fixes memory leak in IREmitter pool allocator
2026-03-11 08:40:04 -04:00
Ryan Houdek 63d52c2a1a OpcodeDispatcher: Fixes memory leak in IREmitter pool allocator
While not a leak in the traditional sense, we were causing pool
allocations to never become free until the thread was closed.

This meant in the case of a game running with >200 threads or so, these
would add up very quickly. So some minor reworking so the IREmitter
doesn't allocate a buffer until first JIT, and making sure to actually
disown the buffer on dispatch error resolved the problems.

Fixes an edge case where Ender Lilies was consuming 409MB with THP
enabled on my desktop, and now it is something like 6MB once idling for
a bit to have the pool allocations do its magic.
2026-03-10 20:55:09 -07:00
Ryan Houdek 9d4a71b57a Merge pull request #5355 from wsxarcher/patch-1
Handle zero length in ChangeProtectionFlags
2026-03-10 08:40:16 -07:00
Marco Bartoli 3462dc3e14 Handle zero length in ChangeProtectionFlags
Add a no-op for zero length in ChangeProtectionFlags.

This fixes AMD Vivado 2025.2 which tries to mprotect with 0 as size and merge strategies fails:

```
Unexpected ChangeProtectionFlags Merge strategy! [0x400000, 0x401000) Versus [0x0, 0x0)
```
2026-03-10 11:34:52 +01:00
LC d21351e66e Merge pull request #5354 from Sonicadvance1/104
Config: Fixes
2026-03-09 21:51:05 -04:00
Ryan Houdek 3204d20335 Config: Fix priorities of config paths
Fixes b4a87d8c0b

`FEX_APP_CONFIG_LOCATION` wasn't overriding paths properly anymore once
that commit landed. Instead legacy `~/.fex-emu/` path would get returned
if it existed first.

Ensures that it returns first, before `STEAM_COMPAT_DATA_PATH` even.
2026-03-09 18:20:56 -07:00
Ryan Houdek a519489d80 CMake: Make sure not to compile Steam tools on mingw 2026-03-09 17:25:59 -07:00
Ryan Houdek 5558c3a35a Merge pull request #5353 from lioncash/group
HostFeatures: Group feature ifdefs together more
2026-03-09 15:45:55 -07:00
Ryan Houdek 58d9755314 Merge pull request #5352 from Sonicadvance1/103
gitlab-ci: Update requirements
2026-03-09 15:45:48 -07:00
Lioncache 498ba0a384 HostFeatures: Group feature ifdefs together more
Makes it a little nicer to see everything grouped together.
2026-03-09 18:27:56 -04:00
Ryan Houdek 5a5477e895 gitlab-ci: Update requirements 2026-03-09 14:00:08 -07:00
Ryan Houdek a17d7ce6ba Merge pull request #5351 from lioncash/hostmops
HostFeatures: Drop in feature testing for FEAT_MOPS
2026-03-09 12:09:31 -07:00
LC afc7248912 Merge pull request #5347 from Sonicadvance1/102
CPUBackend: Enable Transparent Huge Pages on JIT buffers
2026-03-09 14:54:08 -04:00
Lioncache 6bb578fea8 HostFeatures: Drop in feature testing for FEAT_MOPS 2026-03-09 13:03:45 -04:00
Ryan Houdek bed5f293dc CPUBackend: Enable Transparent Huge Pages on JIT buffers
If/When this works, the amount of iTLB misses drop dramatically, which
reduces L2 TLB pressure, which just improves performance for our CPU
cores that have itty-bitty L1 iTLB entry counts.

We can't use `mmap(MAP_HUGETLB)` directly for terrible reasons, so we
are required to lean on madvise instead.
2026-03-06 11:51:17 -08:00
Ryan Houdek 9eb639e891 Docs: Update for release FEX-2603 2026-03-04 12:44:21 -08:00
Ryan Houdek 4968dc4bab Merge pull request #5338 from cjacek/llvm22
Fix mingw builds with LLVM 22.
2026-02-27 12:05:41 -08:00
Jacek Caban e14c440795 Windows: Mark DLLEXPORT_FUNC functions as extern "C"
Fixes cases where _assert is pulled in by libc++, but is missing dllimport,
causing the linker to expect the symbol definition itself to be available.
2026-02-27 15:30:46 +01:00
Jacek Caban 4edd139967 Windows: Add GetActiveProcessorCount stub
It's required by recent libc++.
2026-02-27 15:27:36 +01:00
Ryan Houdek 366e760f7a Merge pull request #5337 from crueter/vixl-fix-commit-hash
External/vixl: Use the right commit hash
2026-02-26 18:11:20 -08:00
crueter 3b1a38621a External/vixl: Use the right commit hash
Totally forgot to update this when I opened that PR. oops!

FYI I would recommend merge-committing if I PR to other repos since it
preserves my GPG signature

Signed-off-by: crueter <crueter@eden-emu.dev>
2026-02-26 20:21:05 -05:00
Ryan Houdek 3c4de59674 Merge pull request #5334 from lioncash/fstp
X87Tables: Handle aliases for FSTP
2026-02-26 16:23:30 -08:00
LC 08a1bc8d97 Merge pull request #5335 from Sonicadvance1/98
Linux/x32: Fixes sign extension in ftruncate
2026-02-26 19:22:59 -05:00
Lioncache c8f3772753 X87Tables: Handle aliases for FSTP 2026-02-26 19:08:48 -05:00
Ryan Houdek c7b6231cd5 Linux/x32: Fixes sign extension in ftruncate
We were passing through 32-bit signed values to the 64-bit handler,
which isn't valid and we need to make sure to sign extend.

I don't think this fixes anything because negative values aren't valid,
but a negative could have been interpreted as a large positive without
this.
2026-02-26 16:04:55 -08:00
Ryan Houdek 1cec200638 Merge pull request #5333 from lioncash/xchg
X87Tables: Handle aliases for FXCH
2026-02-26 15:38:22 -08:00
Lioncache b046942a28 X87Tables: Handle aliases for FXCH 2026-02-26 17:44:12 -05:00
Ryan Houdek 4c49036b8a Merge pull request #5332 from lioncash/fcomp
X87Tables: Add handling for FCOMP DE D0 aliases
2026-02-26 13:49:31 -08:00
Lioncache 6909d49683 X87Tables: Add handling for FCOMP DE D0 aliases
Handles the single remaining alias for FCOMP.
2026-02-26 16:14:55 -05:00
Ryan Houdek 32471953ed Merge pull request #5226 from neobrain/fix_jit_consistency
JIT: Ensure code cache consistency by zero-initializing padding bytes
2026-02-26 13:13:10 -08:00
Tony Wasserka 60272479a1 Merge pull request #5307 from crueter/long-hash
Core: Use long commit hash instead of abbreviating
2026-02-26 20:02:28 +00:00
LC fb7681ddc5 Merge pull request #5331 from Sonicadvance1/97
ThreadManager: Name CallRet stacks
2026-02-26 15:01:14 -05:00
Ryan Houdek 8b581a0a0f ThreadManager: Name CallRet stacks
So WTF can see their usage. Previously didn't add this due to some weird
madvise bug that delete anon names, but that seems to be fixed in a
newer kernel.
2026-02-26 11:48:37 -08:00
crueter 33a01f2773 Core: Use long commit hash instead of abbreviating
Fixes #5230

Basically just stores it as an array of `uint8_t`, and automatically
pads it out to 24 bytes. Hopefully FEX never reaches SHA1 collisions so
this should (TM) never collide.

Note: I have no idea how fmt will handle that format string with an
array of uint8_t.

Signed-off-by: crueter <crueter@eden-emu.dev>
2026-02-26 10:08:41 -05:00
Tony Wasserka 3c517f514f Merge pull request #5064 from zeyi2/fix-clang-22-building
Fix building under clang-22
2026-02-26 15:06:05 +00:00
mtx 48680d68ed Fix building under clang-22 2026-02-26 19:55:35 +08:00
LC 871fc0a40a Merge pull request #5329 from Sonicadvance1/96
External/rpmalloc: Update for named VMA regions
2026-02-25 21:00:14 -05:00
Ryan Houdek 5a3a499f37 External/rpmalloc: Update for named VMA regions
Allows WTF to see rpmalloc tracked memory
2026-02-25 16:46:40 -08:00
LC 5ee18388c4 Merge pull request #5326 from OFFTKP/vcmpss_full
Add vcmp**_full tests
2026-02-24 17:37:46 -05:00
Paris Oplopoios c98a30f82b Fix FALSE_OS size 2026-02-24 23:51:25 +02:00
Paris Oplopoios 427206abf7 Fix TRUE_US size 2026-02-24 23:50:40 +02:00
Paris Oplopoios 2eb4267e6e Fix TRUE_US case 2026-02-24 23:36:09 +02:00
Paris Oplopoios f547b63bca Fix NGT_US vector case 2026-02-24 22:33:46 +02:00
Paris Oplopoios fd71928e3a Fix EQ_UQ cases 2026-02-24 22:15:51 +02:00
Paris Oplopoios 77c0177d23 Add vcmpsd_full, vcmpps_full and vcmppd_full tests
Signed-off-by: Paris Oplopoios <21157395+OFFTKP@users.noreply.github.com>
2026-02-24 14:49:46 +02:00
Paris Oplopoios b10bcca047 Add vcmpss_full test
Signed-off-by: Paris Oplopoios <21157395+OFFTKP@users.noreply.github.com>
2026-02-24 14:49:18 +02:00
Tony Wasserka 4618ca7b6e Merge pull request #5325 from ChanthMiao/chore/cleanup_x11_thunks
Remove unused X11 libraries
2026-02-24 10:13:16 +00:00
LC 27dea3bc57 Merge pull request #5327 from Sonicadvance1/95
FEXCore: Fixes VEX float compare operations
2026-02-23 23:53:45 -05:00
Ryan Houdek b275068569 FEXCore: Fixes VEX float compare operations
The AMD documentation about this instruction is very vague and
misleading in multiple ways. While the Intel documentation is much
cleaner and explains how we need to implement these.

8 of these "new" operations are just inverted signaling versions of the
original 8 SSE versions.
The remaining 16 new operations fill gaps in the original x86 version of
the instructions, exposing the 5 bit truth tables directly, which is why
we also have a "true" and "false" version as well.

Both scalar and vector wide.
Fixes #5326
2026-02-23 15:45:44 -08:00
Ryan Houdek 9d86c3270c IR: Implement new VOrn operation 2026-02-23 14:16:45 -08:00
Changwei Miao 7ea3d7b9b4 Docs/SourceOutline: Remove X11 thunks
Signed-off-by: Changwei Miao <chanthmiao@outlook.com>
2026-02-24 02:02:24 +08:00
Changwei Miao d81a370fca Library Forwarding: Remove unused X11 thunks
These files should be remove with libX11 in #3584.

Signed-off-by: Changwei Miao <chanthmiao@outlook.com>
2026-02-24 01:58:21 +08:00
Ryan Houdek 5849001d00 Merge pull request #5324 from neobrain/refactor_config_remove_unused
Config: Remove unused function from the old code cache skeleton implementation
2026-02-23 09:18:06 -08:00
Ryan Houdek f75ea6ea24 Merge pull request #5323 from neobrain/feature_ir_docstrings
json_ir_generator: Generate docstring for IR op emitters
2026-02-23 09:17:48 -08:00
Tony Wasserka 35e4d452c0 Config: Remove unused function from the old code cache skeleton implementation 2026-02-23 17:49:15 +01:00
Tony Wasserka 7c31ca2ac2 json_ir_generator: Generate docstring for IR op emitters 2026-02-23 17:45:05 +01:00
Tony Wasserka 35bba14b2c IR: Remove malformatted comments from json
Presumably this was done as a hack to highlight the line in red when rendering
documentation to markdown. Since it's only used for two instructions, drop this
use to ease generation of C++ docstrings.
2026-02-23 17:45:05 +01:00
Tony Wasserka 89b976121f JIT: Ensure code cache consistency by zero-initializing padding bytes 2026-02-23 16:49:02 +01:00
LC f5b08bbea6 Merge pull request #5321 from Sonicadvance1/94
AVX128: Convert vzeroupper/vzeroall zeroing to dc zva
2026-02-22 20:57:15 -05:00
LC ca2bdbca19 Merge pull request #5319 from neobrain/fix_cc_codebuffer_size
FEXCore: Default to a larger CodeBuffer size when code caching is enabled
2026-02-21 19:08:06 -05:00
LC 8dfdfdf6ef Merge pull request #5318 from neobrain/fix_cc_validation
CodeCache: Fix validation of large binaries
2026-02-21 19:07:52 -05:00
Ryan Houdek f8b8d21698 InstcountCI: Update for vzero* changes 2026-02-21 15:25:09 -08:00
Ryan Houdek d6d4f84c3b AVX128: Convert vzeroupper/vzeroall zeroing to dc zva
For the upper-half of the registers it is more efficient to zero the
context with `dc zva` on Ampere1A hardware, while Cortex implements this
as equivalent uops in their store pipeline and aren't affected one way
or the other. ARM C1-Pro and newer with FEAT_MOPS also match `dc zva`
performance with 64B/c, but theoretically slightly fewer instructions.
C1-Nano on the other hand, clearly loses to `dc zva`, where mops can
only do 16B/c, but `dc zva` does 64B/c. So we'll need to benchmark or
not if MOPS is a clear win once hardware is actually shipping.
2026-02-21 15:24:13 -08:00
Ryan Houdek 12fcf96e93 IR: Adds a ContextClear operation
This will be useful for zeroing parts of the context at CLZero alignment
and sizes.
2026-02-21 15:24:13 -08:00
Ryan Houdek 5875a121db InstcountCI: Update for alignment change
NFC but touches everything.
2026-02-21 15:24:13 -08:00
Ryan Houdek 19ff098db5 FEXCore: Ensure that CpuStateFrame is 64-byte aligned
This is our cacheline and CLZero line size. Need to ensure it is
aligned.
2026-02-21 15:11:47 -08:00
LC c0724d904f Merge pull request #5317 from Sonicadvance1/92
JIT: Support avoiding spilling CPU flags
2026-02-20 17:59:14 -05:00
LC 97bc2c5c66 Merge pull request #5320 from Sonicadvance1/93
CoreState: Add some more alignment checks
2026-02-20 17:58:55 -05:00
Ryan Houdek b130d14de4 CoreState: Add some more alignment checks
Just to ensure we don't hit performance penalties of some CPUs.
We already weren't, but make sure.
2026-02-20 12:57:21 -08:00
Tony Wasserka cec52253fb FEXCore: Default to a larger CodeBuffer size when code caching is enabled
The current CodeBuffer regrowth code discards any existing contents. This is
undesirable with code caching, since those contents can't be re-fetched from
the disk cache and instead need to be recompiled at runtime.

Additionally, FEXOfflineCompiler obviously should never discard compiled code.

Using a large enough CodeBuffer right away reduces the likelihood that FEX
runs into such scenarios.
2026-02-20 19:07:48 +01:00
Tony Wasserka 31d5805681 CodeCache: Ensure CodeBuffer is large enough for validation run
Previously, CodeBuffer regrowth would discard parts of the code compiled
during the reference run.
2026-02-20 18:49:58 +01:00
Tony Wasserka fda087e24f FEXCore: Add helper function to query available CodeBuffer size 2026-02-20 18:49:07 +01:00
Tony Wasserka 1d8935d0ab CodeCache: Drop unnecessary CodeBuffer size check
The CodeBuffer is resized as required right below.
2026-02-20 18:49:07 +01:00
Ryan Houdek b49730f255 JIT: Support avoiding spilling CPU flags
When thunks are jumping out, games are jumping /entirely/ out of their
controlled code, which means we don't need to save and restore
NZCV,PF,AF.

Some CPUs don't fully rename direct accesses to this register which adds
up during thunking. FPCR is also in the same situation where it'll force
pipeline flushes and isn't renamed away, but we can't really avoid that.

Improves performance at least in Detroit: Become human where the game
spends ~48% CPU time inside of the thunk trampoline for
`vkUpdateDescriptorSets`.

I plan on a follow-up PR where I converge all these options into a
struct argument instead, but that's a follow-up since I don't want to
burn a bunch of time right now.
2026-02-19 17:18:30 -08:00
LC 537d1599f3 Merge pull request #5316 from Sonicadvance1/91
Scripts: Add support for Ampere tuning
2026-02-19 20:00:37 -05:00
Ryan Houdek 7eb7b9f802 Scripts: Add support for Ampere tuning 2026-02-19 16:20:55 -08:00
Ryan Houdek 73822b9e54 Merge pull request #5315 from carsongoodwin32/fix_install_script
Fix InstallFEX.py not running FEXRootFSFetcher correctly
2026-02-19 16:19:21 -07:00
Carson Goodwin a71adc5950 Script: For InstallFEX.py Replace subprocess.call() with subprocess.run() and add TTY as stdin for FEXRootFSFetcher 2026-02-19 16:13:57 -06:00
Tony Wasserka 49a37c7d6f Merge pull request #5313 from neobrain/refactor_header_cleanup
Revert "CodeCache: Use defaulted dtor for ExecutableFileInfo"
2026-02-17 10:31:14 +00:00
LC df3d236408 Merge pull request #5314 from neobrain/opt_cc_generation
CodeCache: Avoid unnecessary copying of CodeBuffer contents
2026-02-16 18:48:36 -05:00
LC 8c536e4d58 Merge pull request #5311 from Sonicadvance1/89
VDSOEmulation: Adds some checks for corrupt guest VDSO ELF
2026-02-16 11:30:41 -05:00
LC 6f7f0867af Merge pull request #5312 from Sonicadvance1/90
ELFCodeLoader: Fixes accidental vector copy
2026-02-16 11:30:04 -05:00
Tony Wasserka 7bdf53291b CodeCache: Avoid unnecessary copying of CodeBuffer contents 2026-02-16 11:07:32 +01:00
Tony Wasserka 44a6bfd874 Revert "CodeCache: Use defaulted dtor for ExecutableFileInfo"
CodeCache.h has no dependency on SourcecodeResolver.h.

This reverts commit f2bbc0eccd.
2026-02-16 10:29:58 +01:00
Ryan Houdek 76551e5648 ELFCodeLoader: Fixes accidental vector copy
Fixes Coverity CID #899885
2026-02-15 15:22:58 -08:00
Ryan Houdek 64482ed07b VDSOEmulation: Adds some checks for corrupt guest VDSO ELF
Fixes Coverity CID #899901
2026-02-15 15:09:22 -08:00
LC c0ec3d92db Merge pull request #5309 from Sonicadvance1/88
Steam: Update VERSIONS.txt requirements
2026-02-14 02:36:19 -05:00
Ryan Houdek c1a2f418a2 Steam: Update VERSIONS.txt requirements
Name and object separated by a tab deliminated character.
2026-02-13 11:43:25 -08:00
Ryan Houdek 6d23625601 Merge pull request #5308 from crueter/install-all-at-once
CMake: Use a single install(FILES) command for SteamRT
2026-02-12 20:39:42 -08:00
crueter 094d429ff0 CMake: Use a single install(FILES) command for SteamRT
No need for individual calls here. While doing this I also realized
xxhash has this at one point, I guess I'll PR that there later

Signed-off-by: crueter <crueter@eden-emu.dev>
2026-02-12 22:42:20 -05:00
Ryan Houdek 2a574e8c92 Merge pull request #5306 from crueter/iwyu
CMake: Also search for `include-what-you-use` if IWYU is enabled
2026-02-12 19:03:36 -08:00
Ryan Houdek da766dd84f Merge pull request #5305 from crueter/cmake-version-style
CMake: Split version/commit identifier to two lines
2026-02-12 18:58:28 -08:00
crueter e2d11ce0fc CMake: Also search for include-what-you-use if IWYU is enabled
Gentoo installs it as this, unsure about others but should "guarantee"
compatibility here

Signed-off-by: crueter <crueter@eden-emu.dev>
2026-02-12 21:50:48 -05:00
crueter fd5c5a18a4 CMake: Split version/commit identifier to two lines
Something something style

Signed-off-by: crueter <crueter@eden-emu.dev>
2026-02-12 21:46:24 -05:00
Ryan Houdek 011fe6a8d2 Merge pull request #5304 from Sonicadvance1/87
Steam: Install a VERSIONS.txt file
2026-02-12 10:10:20 -08:00
Ryan Houdek 9275f09cf9 Merge pull request #5302 from Sonicadvance1/86
CI: Checkout with tags and non-shallow
2026-02-12 10:10:07 -08:00
Ryan Houdek 43a14c986c Merge pull request #5301 from Sonicadvance1/85
CMake: Print FEX version strings found
2026-02-12 10:10:00 -08:00
Ryan Houdek 9d5aecc51d Steam: Install a VERSIONS.txt file 2026-02-11 19:16:33 -08:00
Ryan Houdek a4498925d6 CI: Checkout with tags and non-shallow
This way `git describe` works.
2026-02-11 17:12:54 -08:00
Ryan Houdek 36ec090ee3 Merge pull request #5300 from pmatos/feat/MirrorsEdge
Inline softfloat i16/i32/f32/f64 to extF80 conversions in dispatcher
2026-02-11 16:49:29 -08:00
Ryan Houdek 079038b09e CMake: Print FEX version strings found
Useful debugging
2026-02-11 16:41:05 -08:00
Paulo Matos 11a5229f37 Inline softfloat i16/i32/f32/f64 to extF80 conversions in dispatcher
This was mainly targetted to MirrorsEdge but hopefully will help other workloads
that hit these softfloat conversions in a similar way.
2026-02-11 18:39:20 +01:00
LC 1387aecceb Merge pull request #5297 from Sonicadvance1/82
Fixes remaining quirks of the ARPL implementation
2026-02-10 22:29:25 -05:00
LC 6edde868e6 Merge pull request #5295 from Sonicadvance1/81
unittests/ASM: Adds unittests for undocumented x87 instruction encodings
2026-02-10 22:26:49 -05:00
LC 1cade38b79 Merge pull request #5298 from Sonicadvance1/83
code-format-helper: Update minimum for dependabot again
2026-02-10 22:26:14 -05:00
LC e9fd3005b9 Merge pull request #5299 from Sonicadvance1/84
External/rpmalloc: Fix rpmalloc on mingw
2026-02-10 22:25:49 -05:00
Ryan Houdek 40815c9c78 External/rpmalloc: Fix rpmalloc
Build option changed which broke rpmalloc. Fixed the cmake option.
2026-02-10 18:03:54 -08:00
Ryan Houdek 397fe29bb2 code-format-helper: Update minimum for dependabot again 2026-02-10 15:52:00 -08:00
Ryan Houdek 9e3525408f Merge pull request #5278 from Sonicadvance1/74
CI: Update containers to inherit DEBIAN_FRONTEND
2026-02-10 15:46:38 -08:00
Ryan Houdek 55a2688dc0 Merge pull request #5293 from crueter/cmake-alias-targets
CMake: Avoid global include_directories in favor of propagation
2026-02-10 14:50:09 -08:00
Ryan Houdek f4b0ecd0a0 unittests/ASM: Adds test for arpl correct set operation
arpl will only set part of the destination register if the RPL doesn't
match, it leaves the remaining bits of the register alone.
2026-02-10 14:40:21 -08:00
Ryan Houdek 6fc4ffef0a unittests/ASM: Adds test for arpl flags 2026-02-10 14:30:27 -08:00
Ryan Houdek cf3c41acc7 FEXCore: Fixup ARPL implementation
Some subtle details were wrong in the implementation, and some minor
changes to be more optimal.
2026-02-10 14:29:32 -08:00
FrontMage 9bf9c00474 x86: implement ARPL (0x63) in 32-bit mode 2026-02-10 14:12:52 -08:00
Ryan Houdek a2cdd72db0 unittests/ASM: Adds unittests for undocumented x87 instruction encodings
For more information:
https://en.wikipedia.org/wiki/List_of_x86_instructions#cite_note-x87_alias-346

For #5289
2026-02-10 13:25:53 -08:00
Ryan Houdek 8230b5faf4 Merge pull request #5289 from FrontMage/pr/x87-fcom-fcomp
x87: Implement FCOM/FCOMP st(i) decoding fixes
2026-02-10 13:25:30 -08:00
Ryan Houdek 9e934738b5 Merge pull request #5294 from Sonicadvance1/80
Steam/Toolmanifest: Update commandline tool
2026-02-10 12:25:27 -08:00
Ryan Houdek 5a3b3d797f Steam/Toolmanifest: Update commandline tool
Wrong tool was set.
2026-02-10 12:10:45 -08:00
crueter 6331635f25 CMake: Avoid global include_directories in favor of propagation
Uses ALIAS targets to propagate include directories to targets rather
than polluting the preprocessor with an extra include directory it
usually doesn't need.

Signed-off-by: crueter <crueter@eden-emu.dev>
2026-02-10 14:43:54 -05:00
Ryan Houdek 0ff9277dea Merge pull request #5292 from crueter/cmake-defs
CMake: Avoid add_definitions in favor of newer add_compile_definitions
2026-02-10 11:13:29 -08:00
Ryan Houdek 5dce162a29 Merge pull request #5291 from crueter/lognames
CI: Clean up syntax assertions and remove redundant log-name parameter
2026-02-10 11:12:13 -08:00
crueter adc45149ae CMake: Avoid add_definitions in favor of newer add_compile_definitions
CMake explicitly recommends avoiding the "old" way.

Signed-off-by: crueter <crueter@eden-emu.dev>
2026-02-10 12:47:32 -05:00
crueter 29bdbfc10f CI: Clean up syntax assertions and remove redundant log-name parameter
Now logs are uploaded directly as the name of the test  instead of a
separate name.

Signed-off-by: crueter <crueter@eden-emu.dev>
2026-02-10 12:34:35 -05:00
Ryan Houdek 57504f2359 Merge pull request #5236 from Sonicadvance1/60
Switch over to rpmalloc instead of jemalloc
2026-02-10 08:59:20 -08:00
Ryan Houdek c2a1d188c9 Move docs about allocator 2026-02-10 08:28:51 -08:00
Ryan Houdek bd8b9b3336 External: Remove jemalloc (jemalloc_glibc still exists) 2026-02-10 08:28:51 -08:00
Ryan Houdek 2bcbfe8747 Windows: rpmalloc 2026-02-10 08:28:51 -08:00
Ryan Houdek 98617a4ba5 Switch over to rpmalloc instead of jemalloc.
rpmalloc is currently very aggressively configured which causes
significant reductions in resident memory over jemalloc.

In Bayonetta's title screen it went from 963MB down to 834MB resident.
2026-02-10 08:28:51 -08:00
Ryan Houdek 0a6e13b153 External: Add rpmalloc 2026-02-10 08:28:51 -08:00
Ryan Houdek 93eee39355 Merge pull request #5288 from Sonicadvance1/79
Config: Fixes bad vector to string cast
2026-02-10 08:27:38 -08:00
Ryan Houdek f791c70643 Config: Fixes bad vector to string cast
Turns out casting a vector of data to a string is a bad idea, who knew?!
This bug was exposed by the rpmalloc PR because the
`/run/host/container-manager` file isn't null terminated. So the cast
fextl::string through std::vector::data had the /potential/ to not be
null terminated.

This was just a bug hiding in the code, convert it over to use
fextl::string directly and avoid the entire dance.

Fixes #5236
2026-02-10 08:16:11 -08:00
FrontMage 51afcc5ca8 x87: fix dispatch table sizes after adding FCOM/FCOMP 2026-02-10 18:08:41 +08:00
FrontMage 7a0ed4c462 x87: support 0xDC D0/D8 FCOM/FCOMP st(i) 2026-02-10 18:08:41 +08:00
Ryan Houdek 45376f0dab Merge pull request #5282 from Sonicadvance1/78
FEXCore/Allocator: Disable additional VMA name attempts on failure
2026-02-09 11:43:13 -08:00
Ryan Houdek 34c426a549 Merge pull request #5284 from crueter/unordered-dense-system
CMake: Allow unordered_dense to be used as a system package
2026-02-09 11:16:04 -08:00
Ryan Houdek 1ed3b22db4 Merge pull request #5286 from simon902/main
Fix Buffer Overflow in LoadFileImpl
2026-02-05 11:35:28 -08:00
Simon Scherer 085d80792f Fix Buffer Overflow in LoadFileImpl 2026-02-05 18:42:32 +01:00
LC aabddede8f Merge pull request #5279 from Sonicadvance1/75
FEXCore: Disable 48-bit VA optimization
2026-02-04 23:59:40 -05:00
LC d62038e509 Merge pull request #5281 from Sonicadvance1/77
FEXCore/Allocator: Slightly more verbose logs
2026-02-04 23:57:19 -05:00
crueter e0095d235f CMake: Allow unordered_dense to be used as a system package
Should just work (TM). Because unordered_dense is in my system include
dir I actually can't test comp with external (since it finds the right
include file anyways), so plez test on systems w/o unordered dense

Signed-off-by: crueter <crueter@eden-emu.dev>
2026-02-04 20:32:51 -05:00
Ryan Houdek 5270983dc4 FEXCore/Allocator: Disable additional VMA name attempts on failure
Reduces spam log in strace on platforms that don't support it.
2026-02-03 19:46:00 -08:00
Ryan Houdek 0cf3108af7 FEXCore/Allocator: Slightly more verbose logs
Somehow managed to hit this in a broken setup. Add some more logs.
2026-02-03 19:43:14 -08:00
Ryan Houdek 217d039bb7 Merge pull request #5277 from Sonicadvance1/73
CPUID: Add option for hiding hybrid Big.Little CPUs
2026-02-03 11:46:30 -08:00
LC fda4d386c9 Merge pull request #5280 from Sonicadvance1/76
FEXCore: Early exit for invalid VEX.vvvv on 32-bit
2026-02-02 22:50:54 -05:00
Ryan Houdek 2fc664f12a FEXCore: Early exit for invalid VEX.vvvv on 32-bit
Otherwise we hit an assert in FEXCore backend with code discovery
hitting things that look like AVX.

Fixes a crash in Uplay.

Also adds a test to just ensure that the instruction faults out and is
captured instead of crashing in FEX itself.
2026-02-02 17:44:50 -08:00
Ryan Houdek 3d989dbaff FEXCore: Disable 48-bit VA optimization
This breaks relocations currently due to not handling negatives and also
an interesting overwriting problem.

Not that big of a deal, it's only a minor optimization anyway.

Fixes #5227
2026-02-02 15:17:28 -08:00
Ryan Houdek f281b4b4cb CI: Update containers to inherit DEBIAN_FRONTEND
Just to ensure we are getting non-interactive mode inside the
containers.
2026-02-02 13:19:59 -08:00
Ryan Houdek 8bfa6b817f CPUID: Add option for hiding hybrid Big.Little CPUs
Required for Denuvo?
2026-02-02 12:59:44 -08:00
LC f295212360 Merge pull request #5275 from Sonicadvance1/72
Softfloat: Remove weird tail padding from X80SoftFloat
2026-02-01 20:14:34 -05:00
Ryan Houdek 7b2b639d1e Merge pull request #5274 from Sonicadvance1/71
FEXInterpreter: Override glibc set program invocation name
2026-01-29 12:41:00 -08:00
Ryan Houdek c6031a7806 Softfloat: Remove weird tail padding from X80SoftFloat
These are expected to match x87 registers in side. This was always a bit
weird.
2026-01-28 19:40:29 -08:00
Ryan Houdek d111518352 FEXInterpreter: Override glibc set program invocation name
Mesa uses this to determine what the executable name is in a process for
application profiles. Without thunking, the guest glibc sets this
correctly, with thunking mesa would pick up `FEX` instead of the app.

With this fixed, it fixes Dead Island rendering when thunks are enabled.
Plenty of other games in Mesa's application profiles that it would fix
as well.
2026-01-28 15:10:34 -08:00
LC a5ceb2a61d Merge pull request #5270 from Sonicadvance1/69
FEXCore: Move two functions to the frontend
2026-01-28 16:01:35 -05:00
LC bd72eb5693 Merge pull request #5266 from Sonicadvance1/68
Linux: Work around binfmt_misc bug
2026-01-28 16:00:59 -05:00
LC e3dcb4b0dc Merge pull request #5264 from Sonicadvance1/66
FEXLinuxTests: Adds execveat with mfd_cloexec test and related fixes.
2026-01-28 15:57:53 -05:00
LC b6b64c6dc5 Merge pull request #5260 from Sonicadvance1/64
FEX: Ensure VDSO and Stack are pushed as high in the VA space as possible
2026-01-28 15:55:30 -05:00
LC 030d332a79 Merge pull request #5273 from Sonicadvance1/70
unittests/FEXLinuxTests: Ensures that signals are in expected order
2026-01-28 15:53:30 -05:00
Ryan Houdek 54d3e33616 unittests/FEXLinuxTests: Ensures that signals are in expected order
This is a fairly trivial test to ensure that signals correlate to their
expected order. I keep getting spooked that signal numbers on x86 don't
always correlate to signal numbers on other architectures. So slam
through all 64 signals to ensure they are correct.

On an architecture that doesn't match signals 1:1, or an emulator that
doesn't properly remap when that occurs, then the order would be
incorrect.
2026-01-27 16:32:47 -08:00
Ryan Houdek 1c465b55e6 Merge pull request #5267 from crueter/dedup-cc-stuff
CI: Move MinGW triple set to a single step
2026-01-27 11:54:11 -08:00
Ryan Houdek 687b66da43 Merge pull request #5271 from crueter/posix
Scripts: Cleanups and POSIX compliance
2026-01-27 11:53:35 -08:00
Ryan Houdek 710ab4ae3f Merge pull request #5265 from Sonicadvance1/67
Linux: Ensure seccomp inheritence occurs unconditionally
2026-01-27 11:52:27 -08:00
Ryan Houdek 06f4ca6983 Merge pull request #5263 from crueter/ci-dedup-build-env-setup
CI: Deduplicate common build/rootFS setup steps
2026-01-27 11:44:53 -08:00
Ryan Houdek 8cedcccfb0 Merge pull request #5269 from crueter/vixl-sim-var
CI: Deduplicate VIXL_SIM_ENABLED steps
2026-01-27 11:44:34 -08:00
Ryan Houdek 70a8bea8f6 Merge pull request #5262 from Sonicadvance1/65
Linux/x32: Only copy rusage on success
2026-01-27 11:42:24 -08:00
Ryan Houdek 8512a26319 Linux: Ensure seccomp inheritence occurs unconditionally
Seccomp once enabled is a virus that passes onward to all children
processes. Ensure we enable the option if we are inheriting a filter.

For #5234
2026-01-27 11:40:26 -08:00
Ryan Houdek 778a2866e4 Merge pull request #5272 from pmatos/fix/pf-flag-initialization
FEXCore: Fix initial PF flag value
2026-01-27 10:24:05 -08:00
Tony Wasserka 08a74989b7 Merge pull request #5261 from crueter/ci-build-tests-first
CI: Move noisy build targets generation to a separate step
2026-01-27 16:55:42 +00:00
crueter 68522c82a6 CI: Move noisy build targets generation to a separate step
As requested by neobrain. Makes log parsing easier because you can see
the "noisy" and usually unchanging targets in one step
(which usually succeed, I hope), and the more error-prone "the rest of FEX"
build in separate steps. Makes parsing FEX build errors easier

Signed-off-by: crueter <crueter@eden-emu.dev>
2026-01-27 09:17:07 -05:00
Paulo Matos ec2d1547c9 asm_tests: FEXCore: Fix initial PF flag value 2026-01-27 11:25:08 +01:00
Paulo Matos eb6c050cf0 FEXCore: Fix initial PF flag value
Initialize pf_raw to 1 instead of 0 so that the reconstructed Parity
Flag matches x86 reset state (PF=0).
2026-01-27 11:25:08 +01:00
Paulo Matos 0e5ce3e07e TestHarnessRunner: Initialize flags to x86 reset state 2026-01-27 11:25:08 +01:00
crueter 14791ce7eb Scripts: Cleanups and POSIX compliance
- Makes all scripts POSIX-compliant
- General formatting and such
- shellcheck :)

Shell script is love. Shell script is life.

Signed-off-by: crueter <crueter@eden-emu.dev>
2026-01-26 23:53:33 -05:00
Ryan Houdek a8bee07d8f FEXCore: Move two functions to the frontend
Only used in the frontend and is OS specific.
2026-01-26 20:11:36 -08:00
crueter 184d3e158a CI: Deduplicate VIXL_SIM_ENABLED steps
Same deal as #5267. Is just less unique steps

Signed-off-by: crueter <crueter@eden-emu.dev>
2026-01-26 23:01:31 -05:00
crueter f8679f06aa CI: Move MinGW triple set to a single step
Long live shell script

Signed-off-by: crueter <crueter@eden-emu.dev>
2026-01-26 22:56:49 -05:00
Ryan Houdek 102e6dd2fb Linux: Work around binfmt_misc bug
When an anonymous FD is passed to execveat that has the CLOEXEC flag
set, then binfmt_misc fails with ENOENT.

Workaround this limitation by duplicating the FD, stripping its CLOEXEC
flag in the process.

For #5234
2026-01-26 19:15:47 -08:00
Ryan Houdek 7806ad8ac8 Config: Stop using getpwuid
Instead manually parse /etc/passwd in this rare case.
2026-01-26 18:23:35 -08:00
Ryan Houdek da23514424 Common: Disable glibc fault checking for getpwuid
No way to workaround this without completely reimplementing the nss
database implementation in glibc.
2026-01-26 17:37:48 -08:00
Ryan Houdek e39201147e FEXLinuxTests: Adds execveat with mfd_cloexec test
This just ensures that we have this behaviour working correctly.
binfmt_misc in the kernel has a bug that this doesn't work which is kind
of funny.
2026-01-26 17:07:57 -08:00
crueter c8d5f9edc5 CI: Deduplicate common build/rootFS setup steps
6 steps use this verbatim, and steamrt4 does as well (minus rootfs
stuff). Less LoC and whatnot, you get the drill

Signed-off-by: crueter <crueter@eden-emu.dev>
2026-01-26 19:55:41 -05:00
Ryan Houdek 9248964191 Linux/x32: Only copy rusage on success
Also removes a copy before the syscall which made no sense.
2026-01-26 16:26:37 -08:00
Ryan Houdek 5cc4d02b80 FEX: Ensure VDSO and Stack are pushed as high in the VA space as possible
On 36-bit VA systems the stack was ending up /wherever/ when it should
be at the top of the VA space (usually).

Additionally VDSO was getting mapped anywhere on 64-bit, so push that to
the top of the VA space as well.

Also removes a check for old kernels not supporting MAP_FIXED_NOREPLACE.
2026-01-26 14:33:58 -08:00
Ryan Houdek 1b18f85e42 Merge pull request #5211 from Sonicadvance1/46
Linux: Moves DRM LRU FD cache to be per-thread
2026-01-26 13:11:24 -08:00
Ryan Houdek e3cb700cf7 Review 2026-01-26 12:39:00 -08:00
Ryan Houdek 82b1763431 Linux: Moves DRM LRU FD cache to be per-thread
This was a nasty race condition where each thread could be accessing the
DRM cache at any given moment. Move it over to a per thread object that
is only allocated once it gets used.

Also removes an `atexit` registration that contributes to crashing on
exit.
2026-01-26 12:36:30 -08:00
Ryan Houdek cc54724ff1 Merge pull request #5258 from neobrain/fix_format_x86tables
Enable automatic code formatting for X86Tables.h
2026-01-23 12:08:16 -08:00
Tony Wasserka 5eebc9e093 Add previous commit to git blame ignore file 2026-01-23 15:40:53 +01:00
Tony Wasserka ba2b0ef809 Enable automatic code formatting for X86Tables.h 2026-01-23 15:39:54 +01:00
Ryan Houdek 2531d33880 Merge pull request #5255 from crueter/ci-update-deps
CI: Update actions deps
2026-01-21 11:03:40 -08:00
Ryan Houdek d66cb54a82 Merge pull request #5254 from crueter/ci-wine-dll-dedup
CI: Use a composite action for Wine DLL builds
2026-01-21 11:00:57 -08:00
crueter e28dbfb377 CI: Use a composite action for Wine DLL builds
Less duplication and less work.

Signed-off-by: crueter <crueter@eden-emu.dev>
2026-01-21 10:05:23 -05:00
crueter 83caa146d1 CI: Update actions deps
Checkout and upload-artifact. They contain general bug fixes and update
the Node runtime such that they should be generally faster. Note that
this requires Node.js 24, if that's a dealbreaker for the runners feel
free to close.

Signed-off-by: crueter <crueter@eden-emu.dev>
2026-01-21 10:03:04 -05:00
Tony Wasserka cf61d7349a Merge pull request #5233 from crueter/find-zydis
CMake: Use a Find module for Zydis/Zycore
2026-01-21 09:33:30 +00:00
Ryan Houdek 87de3ec3e3 Merge pull request #5253 from crueter/ci-tests-v2
CI: Refactor tests to use a composite action
2026-01-20 16:43:58 -08:00
crueter 311b373455 CI: Refactor tests to use a composite action
Significantly reduces the constant duplication of test run steps. Port
of #5246, but differs in that the composite action now truncates the log
file in-place.

CC: This is rebased on top of #5252 because I do NOT want to deal with
these merge conflicts again.

Signed-off-by: crueter <crueter@eden-emu.dev>
2026-01-20 19:18:25 -05:00
Ryan Houdek 7f13185f2f Merge pull request #5252 from crueter/ci-build-config
CI: Remove redundant build mkdir and `--config`
2026-01-20 16:17:48 -08:00
crueter 612138db4f CI: Remove redundant build mkdir and --config
CMake already handles the mkdir step implicitly with `-B build`. Same
with build type, where setting `CMAKE_BUILD_TYPE` will propagate it as
such to the build commands. The only exception to this is the MSBuild
generator, which we don't use here.

Signed-off-by: crueter <crueter@eden-emu.dev>
2026-01-20 18:49:10 -05:00
Ryan Houdek c39eef1d5a Merge pull request #5249 from crueter/remove-runner-workspace
CI: Remove redundant runner.workspace and working-directory directives
2026-01-20 15:06:18 -08:00
crueter 41bb09e361 CI: Remove redundant runner.workspace and working-directory directives
Signed-off-by: crueter <crueter@eden-emu.dev>
2026-01-20 17:55:05 -05:00
Ryan Houdek 6b9590a6a9 Merge pull request #5251 from crueter/ci-remove-upload-unnecessary
CI: Remove unnecessary upload steps
2026-01-20 14:12:30 -08:00
crueter 8583f38ec1 CI: Remove unnecessary upload steps
MinGW doesn't do any tests. Also forgot to remove the InstCountCI diff.

Signed-off-by: crueter <crueter@eden-emu.dev>
2026-01-20 17:00:36 -05:00
Ryan Houdek 6e11a46b2a Merge pull request #5250 from crueter/namespace-artifacts
CI: Namespace uploaded artifacts by runner label
2026-01-20 12:27:48 -08:00
crueter 801443a7d0 CI: Namespace uploaded artifacts by runner label
https://github.com/FEX-Emu/FEX/actions/runs/21184542253/job/60935403846?pr=5249

Signed-off-by: crueter <crueter@eden-emu.dev>
2026-01-20 14:51:52 -05:00
crueter 2284247b56 CMake: Use a Find module for Zydis/Zycore
Much like the prior xxhash PR. This one also has the advantage of making
it completely trivial to handle distros that only install zy{dis,core}-config.cmake
and not the PkgConfig files (Gentoo, Arch)

Working on both Gentoo and Arch (CMake), and Fedora (PkgConfig). Does
work on Ubuntu 22.04 though it gets rejected due to being too old, since
we require 4.0 or newer.

Signed-off-by: crueter <crueter@eden-emu.dev>
2026-01-20 13:56:12 -05:00
LC 33986848f3 Merge pull request #5248 from Sonicadvance1/63
unittests/ASM: Fix accidentally deleted file
2026-01-19 23:09:33 -05:00
Ryan Houdek a55928dbee unittests/ASM: Fix accidentally feleted file
This wasn't meant to be deleted.
2026-01-19 19:37:42 -08:00
Ryan Houdek 6536ccf003 Merge pull request #5239 from Sonicadvance1/61
External/vixl: Rebase on upstream
2026-01-19 13:38:18 -08:00
Ryan Houdek 548dfdc26d Merge pull request #5240 from Sonicadvance1/62
unittests/ASM: Removes superfluous warnings from nasm
2026-01-19 13:37:53 -08:00
Ryan Houdek 4428dbf885 Merge pull request #5241 from crueter/patch-1
CI: Merge InstCountCI Diff steps into one
2026-01-18 18:24:58 -08:00
crueter f11dc7e3f6 CI: Merge InstCountCI Diff steps into one
Uses `git --no-pager diff --exit-code HEAD` instead of the two-step check used before. Also shows the diff if it exists so you can see changes at a quick glance.

Signed-off-by: crueter <crueter@eden-emu.dev>
2026-01-18 21:09:04 -05:00
Ryan Houdek b91dc754b0 Merge pull request #5243 from crueter/ci-redundancy
CI: Don't specify `shell: bash` every time
2026-01-18 17:34:59 -08:00
Ryan Houdek b2c1ad7abc Merge pull request #5244 from crueter/gitignore-qtcreator
gitignore: ignore CMakeLists.txt.user
2026-01-18 16:43:30 -08:00
crueter 730f80447e CI: Don't specify shell: bash every time
Bash is already the default on Linux runners.

Just a bit cleaner and less lines. Also cleaned out some unneeded
comments from the template.

Signed-off-by: crueter <crueter@eden-emu.dev>
2026-01-18 19:36:56 -05:00
crueter 90743e68d4 gitignore: ignore CMakeLists.txt.user
I can't tell you how many times I've accidentally committed this file.
It's used by Qt Creator

Signed-off-by: crueter <crueter@eden-emu.dev>
2026-01-18 19:33:43 -05:00
Ryan Houdek 23f56e4c1b unittests/ASM: Removes more warnings 2026-01-17 13:46:13 -08:00
Ryan Houdek c16d631b58 unittests/ASM: Removes superfluous warnings from nasm 2026-01-17 13:07:19 -08:00
Ryan Houdek e000905aca unittests/Emitter: Update 2026-01-17 12:59:40 -08:00
Ryan Houdek 85685afc68 InstcountCI: Update 2026-01-17 12:51:52 -08:00
Ryan Houdek dfd81fd125 External/vixl: Rebase on upstream
Vixl has migrated its upstream repo to https://gitlab.arm.com/runtimes/vixl
without telling anyone. Rebase on their current main rather than the
dead project.

At the very least removes the obnoxious `operator""_h` warnings. Not too
much has changed since our last rebased.

Fixes #5237
2026-01-17 12:42:25 -08:00
Ryan Houdek 78fd9b4fe4 Merge pull request #5162 from crueter/unordered-dense
FEXCore: use ankerl::unordered_dense over tsl
2026-01-16 11:07:59 -08:00
Tony Wasserka 504d1e3004 Merge pull request #5232 from crueter/xdg
Config: Properly follow XDG directory specifications
2026-01-15 22:17:26 +00:00
crueter b4a87d8c0b Config: Properly follow XDG directory specifications
Previously, FEX would use `$HOME/.fex-emu` for its data and config if
`XDG_CONFIG_HOME` and/or `XDG_DATA_HOME` were unset. This doesn't follow
XDG, so instead we do a fallback to `$HOME/.config` and
`$HOME/.local/share` respectively if XDG env vars are unset. Also,
pre-emptively creates those directories since `~/.local/share` and
`~/.config` existing is technically not a guarantee if XDG dirs are unset.

Signed-off-by: crueter <crueter@eden-emu.dev>
2026-01-15 17:01:59 -05:00
crueter 884b96c300 [externals] use ankerl::unordered_dense over tsl
ankerl::unordered_dense is faster on average and has less memory usage
than tsl::robin_map. It is pretty significantly faster than std but
we'll keep that as is for now.

Obviously, this will need a lot of testing.

Signed-off-by: crueter <crueter@eden-emu.dev>
2026-01-15 14:07:08 -05:00
Ryan Houdek 14667c1756 Merge pull request #5163 from crueter/fix-pr-format
CI: fix pr-code-format
2026-01-15 11:03:20 -08:00
Tony Wasserka b324cfbfde Merge pull request #5215 from Sonicadvance1/50
FEXInterpreter: Removes three static objects
2026-01-14 20:24:09 +00:00
Ryan Houdek 7047d08ca4 FEXInterpreter: Removes three static objects
These three objects are easy enough to move out of the static namespace.
They were causing atexit registrations for deallocating memory.

Removes three `atexit` registrations that contributes to crashing on
exit.
2026-01-14 12:08:20 -08:00
Tony Wasserka 28c486dad5 Merge pull request #5128 from pmatos/feat/Zydis
Add Zydis x86/x86-64 disassembler support
2026-01-14 14:52:27 +00:00
Paulo Matos d6fd82c1ab Add Zydis x86/x86-64 disassembler support
Integrate Zydis as an optional dependency to enable x86/x86-64 guest
instruction disassembly during JIT compilation.

Build with -DENABLE_ZYDIS=TRUE.
Use FEX_X86DISASSEMBLE=1 at runtime to output guest x86 instructions
for each compiled block.
2026-01-14 15:40:28 +01:00
LC bcf53458ad Merge pull request #5229 from Sonicadvance1/57
SignalDelegator: Make sure to store XSTATE_MAGIC2
2026-01-13 19:14:00 -05:00
LC 15a6ba791e Merge pull request #5228 from Sonicadvance1/56
Linux: Updates minimum altstack requirements
2026-01-13 19:12:13 -05:00
Ryan Houdek 5bce7e0616 SignalDelegator: Make sure to store XSTATE_MAGIC2
Applications can use this to ensure they've grabbed the whole XSTATE
correctly. UML uses it to determine if the FPState was correctly saved.

Also adds a unittest to ensure the same correct behaviour.

For #5206
2026-01-13 15:33:24 -08:00
Ryan Houdek 696f69e442 Linux: Updates minimum altstack requirements
Turns out the Linux kernel's definition of `MINSIGSTKSZ` and glibc's
definition of `MINSIGSTKSZ` don't match.

The linux kernel has a massive comment about it in `arch/x86/kernel/signal.c`.

Also adds a unittest for it.

For #5206
2026-01-13 13:33:32 -08:00
crueter 96b8904d88 [ci] fix pr-code-format
Previously, changes to externals were also tracked, resulting in
absolutely hilarious abominations like https://github.com/FEX-Emu/FEX/actions/runs/20498228017/job/58901276463?pr=5162

So rather than dealing with that weird action, just fetch the main
branch and run a `diff --name-only` with the PR's merge base with main.
This is how I've done several dozen diff-based scripts (license headers,
clang-format, etc) and it works perfectly, so let's just use this.
Should run quicker too but don't put any money on it.

I also went ahead and removed a bunch of unnecessary quotations cuz
wynaut.

Signed-off-by: crueter <crueter@eden-emu.dev>
2026-01-13 15:00:02 -05:00
Tony Wasserka 6438a838ce Merge pull request #5224 from Sonicadvance1/55
CPUID: Adds AmpereOneC identifier
2026-01-13 10:31:44 +00:00
Tony Wasserka d7d870cb4c Merge pull request #5217 from Sonicadvance1/52
FEX: Move SBRK handling to the frontend
2026-01-13 09:45:40 +00:00
Ryan Houdek 6b33613bb0 CPUID: Adds AmpereOneC identifier 2026-01-12 20:04:56 -08:00
Ryan Houdek 1caa9d5294 FEX: Move SBRK handling to the frontend
This is fundamentally a frontend only problem, and also Linux only.
Moves it to the frontend where it belongs.

There's likely more things in Allocator.cpp that can be moved to the
frontend but this is the first thing.

NFC
2026-01-12 13:18:01 -08:00
Tony Wasserka 251a3babad Merge pull request #5218 from Sonicadvance1/53
Linux: Fixes personality inheritence
2026-01-12 12:24:29 +00:00
Tony Wasserka a0d3aed347 Merge pull request #5220 from Sonicadvance1/54
CI: Update tunables for artifact builds
2026-01-12 11:03:42 +00:00
Ryan Houdek 5cc074ec3b CI: Update tunables for artifact builds
These are what we want them to be.
2026-01-10 17:48:27 -08:00
Ryan Houdek 4c7b9fd6e5 Merge pull request #5219 from Sonicadvance1/22
gitlab-ci: Adds steamrt4 runner
2026-01-10 17:39:33 -08:00
Ryan Houdek f76170eab4 Linux: Fixes personality inheritence
UML disables ASLR using a personality. FEX in the ELFCodeLoader was
accidentally erasing personalities. This caused UML to do a bootloop
trying to repeat disable ASLR.

It now gets ASLR disabled, but fails shortly afterwards.

For #5206
2026-01-10 16:22:10 -08:00
Ryan Houdek a63da98032 gitlab-ci: Adds steamrt4 runner
Needs to go in.
2026-01-10 13:23:12 -08:00
Ryan Houdek 1d411df274 Merge pull request #5207 from Sonicadvance1/44
code-format-helper: Update requirements
2026-01-10 13:13:22 -08:00
Ryan Houdek 4c50a92e12 Merge pull request #5187 from svc64/accurate-stack-mapping
Accurate stack mapping, fixes #5149
2026-01-09 14:30:38 -08:00
Asaf Niv 92e71b2d08 ELFCoreLoader, unittests: don't emulate kernel < 5.8 stack mapping 2026-01-10 00:17:46 +02:00
Asaf Niv ff69959acc unittests: fix smc-exec-stack 2026-01-10 00:17:46 +02:00
Asaf Niv 10f3ce070d ELFCoreLoader: correct stack mapping behavior 2026-01-10 00:17:46 +02:00
Asaf Niv e2d3e6f02a unittests: fix formatting, change variable name 2026-01-10 00:17:46 +02:00
Asaf Niv 867b7466f4 unittests: refactor stack execution tests, move test area 2026-01-10 00:17:46 +02:00
Asaf Niv e95e465a0a unittests: add comment about using lld 2026-01-10 00:17:46 +02:00
Asaf Niv e5a8ce5ecb unittests: add tests for stack mapping 2026-01-10 00:17:46 +02:00
Asaf Niv 227256e6ba ELFCodeLoader: make stack mapping match kernel behavior
Behavior should now match the matrix in #5149
2026-01-10 00:17:46 +02:00
Asaf Niv 875956f966 gitignore: ignore .idea 2026-01-10 00:17:46 +02:00
Ryan Houdek 50ff4cc45e Merge pull request #5210 from neobrain/fix_libfwd_vulkan_update
LibraryForwarding: Redo Vulkan definitions update
2026-01-09 11:44:29 -08:00
Tony Wasserka a26aee4afe LibraryForwarding/vulkan: Update structs used by 32-bit apps 2026-01-09 12:49:51 +01:00
Tony Wasserka a76f5e4143 LibraryForwarding/vulkan: Reformat for ease of updating definitions 2026-01-09 12:35:10 +01:00
Tony Wasserka 0bff771c03 LibraryForwarding/vulkan: Redo definitions update
eb95a959bc (#5171) did not keep the interface
file in sync with the output of DefinitionExtract.py. Some definitions were
reordered in the Vulkan headers.

There also is no "conflict" about CUDA extensions; instead, the previous
update didn't reflect that the new headers fully hide beta extensions.
2026-01-09 12:35:08 +01:00
Ryan Houdek 0620f27c10 code-format-helper: Update requirements
To stop dependabot warnings.
2026-01-08 12:25:53 -08:00
Ryan Houdek 1188c90c10 Docs: Update for release FEX-2601 2026-01-07 14:20:19 -08:00
Ryan Houdek b29a78c068 Merge pull request #5205 from Sonicadvance1/43
FEXCore: Cleanup pointers structure
2026-01-07 14:17:30 -08:00
Ryan Houdek 751fd70293 InstcountCI: Update 2026-01-07 12:55:34 -08:00
Ryan Houdek d7c2b8f513 FEXCore: Cleanup pointers structure
There's no longer a distinction between AArch64 and x86 and everything
effectively falls under "Common" now. This means flattening the entire
structure just cleans it up.

NFC. (Although instcountCI will update because of a couple pointer
offsets changing)
2026-01-07 12:55:34 -08:00
LC 5627ddff8f Merge pull request #5204 from Sonicadvance1/42
FEXCore: Fixes circular dependency with thunk callback
2026-01-07 15:47:58 -05:00
Ryan Houdek 291e261a65 InstcountCI: Update 2026-01-07 12:19:22 -08:00
Ryan Houdek 817a927e31 FEXCore: Fixes circular dependency with thunk callback
Fixes crash in thunks that use callbacks, introduced in #5148.

The dispatcher would call the syscallhandler to get the VDSO thunk
callback. But due to reordering initialization, the VDSO thunk would
have not been loaded at that point. This would cause thunks that use
callbacks to crash with a nullptr exception.

Instead, defer the thunk callback pointer loading until the thread
starts executing, and load the pointer in to our thread state's pointer
struct instead.

Didn't get caught in my initial test sweep since I didn't run a Wine
game with thunks.
2026-01-07 12:03:19 -08:00
Ryan Houdek c7eb4c8447 Merge pull request #5202 from Sonicadvance1/41
Relocations: Disable 6-byte size optimization in `InsertGuestRIPMove`
2026-01-07 10:22:12 -08:00
Ryan Houdek ac3cabec07 Relocations: Disable 6-byte size optimization in InsertGuestRIPMove
This wasn't handling negatives correctly which was causing xalia.exe to
assert. Disable for now rather than further changing logic, with a TODO that it
should get fixed in the future.
2026-01-07 09:19:13 -08:00
LC 6c06f47cf5 Merge pull request #5201 from Sonicadvance1/40
LinuxSyscalls/x32: Fixes fcntl assert
2026-01-07 11:31:13 -05:00
Ryan Houdek c7df064d58 LinuxSyscalls/x32: Fixes fcntl assert
Steam started taking advantage of newer fcntl commands, in particular
`F_CREATED_QUERY`. This was causing our fcntl handler for 32-bit
processes to assert out.

Fixes the handler so it passes all other fcntl commands forward, just
like the Linux kernel does. Splits the 32-bit handler to explicitly
ignore a handful of commands just like the Linux kernel.

Fixes an assert that Steam was hitting.
2026-01-07 08:17:37 -08:00
Ryan Houdek ed1d49520f Merge pull request #5148 from Sonicadvance1/29
Interpreter: Moves around the thread and ELF initialization code
2026-01-07 08:12:09 -08:00
Ryan Houdek 651ef64617 Merge pull request #5200 from crueter/mv-cmake-modules
CMake: Move CMakeModules to Data/CMake
2026-01-05 13:32:38 -08:00
crueter f3ee822968 CMake: Move CMakeModules to Data/CMake
See
https://github.com/FEX-Emu/FEX/pull/5172#pullrequestreview-3626560058.

Signed-off-by: crueter <crueter@eden-emu.dev>
2026-01-05 16:14:46 -05:00
Ryan Houdek d592e2afb0 Merge pull request #5197 from bylaws/rwxp5
Windows: Improve handling of RWX memory
2026-01-05 13:12:01 -08:00
Ryan Houdek a25d90de8a Merge pull request #5199 from bylaws/wowsusp4
WOW64: Lock the JIT context and block suspend during context operations
2026-01-05 13:02:10 -08:00
Ryan Houdek fedebf4b66 Merge pull request #5198 from neobrain/fix_flt_gcc
unittests/FEXLinuxTests: Fix gcc build
2026-01-05 12:58:29 -08:00
Billy Laws ece38a5411 WOW64: Lock the JIT context and block suspend during context operations
Prevent suspension when flushing thread state to the WOW64 context
during BTCpuSetContext as suspension itself performs a flush. If
suspension were to occur halfway through FlushThreadStateContext then
the fact that the WOW64 context has been dirtied would be missed and the
new context would be partially overwritten with stale data from the FEX
thread state. Additionally if the context were to be dirtied by
suspension at any point during context merging in BTCpuSetContext the
input context could be lost.

Fixes occasional crashes when starting AC4.
2026-01-05 20:10:14 +00:00
Tony Wasserka 785c20c68d unittests/FEXLinuxTests: Fix gcc build 2026-01-05 19:41:43 +01:00
LC 7e4e01789d Merge pull request #5099 from Sonicadvance1/15
Tools/pidof: Fixes FEXpidof after #5097
2026-01-05 07:31:42 -05:00
Billy Laws 8e9f593c39 Windows: Handle NtReadFile calls to RX protected RWX memory
Handling this safely requires blocking all compilation for the duration
of the read, as fault-based tracking doesn't work when wine's unix side
itself is the one performing the write into the RWX region (we cannot
catch unix faults).

Fixes a startup crash in Persona 5.
2026-01-05 02:34:00 +00:00
Billy Laws ec13e5d503 Windows: Use FrontendPtr for extra TLS space on WOW64/ARM64EC
We have run out of EmulatorData slots on ARM64EC, so introduce this for
less-frequently accessed TLS data for which the additional cost doesn't
matter.
2026-01-05 02:29:34 +00:00
Ryan Houdek 721ceecd70 ELFCodeLoader: Update the thread persona if it was passed in
Execstack enables `READ_IMPLIES_EXEC` and we were accidentally
decoupling our expected persona due to the reordering of initialization.

This was causing the smc-1-dynamic test to fail because it's testing
execstack and our VMA tracking wasn't returning read-write VMAs as
executable once these decoupled.

Easy fix.
2026-01-04 17:29:11 -08:00
Ryan Houdek 3728f5f178 Interpreter: Moves around the thread and ELF initialization code
This allows the context and parent thread objects to be created earlier,
allowing the VDSO and ELFCodeLoader mapping functions to have a thread
object for tracking memory mappings through the regular guest routines.

This means that we no longer need to do any form of deferred handling
for code caching as all the state is ready early in the initialization
process.

A little bit of care needed to be taken to ensure we still close the
ELFCodeLoader's FDs later and that VDSO unmapping happens before tearing
down the parent thread, but overall this is mostly just passing the
InternalThreadState object around as normal.

I couldn't find any functional regression from this change alongside
code caching, but it would be good for @neobrain to double check this.
2026-01-04 17:29:11 -08:00
Ryan Houdek 3a3e887622 Interpreter: Moves a block of kernel testing code to its own namespace
Just keeps it out of the way during some more code changes.

NFC
2026-01-04 17:29:11 -08:00
Ryan Houdek 7eca40f093 pidof: Add wine to all found FEX instances as well 2026-01-04 17:28:59 -08:00
Ryan Houdek 3005abc3aa pidof: See through deleted links 2026-01-04 17:28:59 -08:00
Ryan Houdek 69eddaebbf Tools/pidof: Fixes FEXpidof after #5097
FEX argument will no longer exist in `/proc/<pid>/cmdline` after this
PR. Ensure this tool still works.

After #5097 is merged, one can use regular `pidof` for Linux
applications, but this is still useful as it finds Wine applications
running with FEX as well.
2026-01-04 17:28:59 -08:00
Ryan Houdek 992d4411e1 pidof: Move PID iteration to a helper 2026-01-04 17:28:59 -08:00
Billy Laws e97b18e244 InvalidationTracker: Fix potential race when reprotecting RWX ranges
Theoretically, another thread could run after the reprotection but
before the invalidation, write some code and jump to it but end up
jumping to old code as the invalidation is yet to occur. Fix this by
locking the invalidation mutex, invalidating and only then reprotecting.
2026-01-05 01:20:02 +00:00
LC e1c6a910d2 Merge pull request #5196 from Sonicadvance1/39
Config: Remove stdout from OutputLog
2026-01-04 19:33:32 -05:00
LC b40768895d Merge pull request #5195 from Sonicadvance1/38
Scripts: Have InstallFEX check kernel version
2026-01-04 04:08:14 -05:00
Ryan Houdek 96033fd225 Config: Remove stdout from OutputLog
This just causes problems with scripts. Don't allow people to output to
stdout, use stderr instead.
2026-01-03 15:15:47 -08:00
Ryan Houdek 855ef1ade9 Scripts: Have InstallFEX check kernel version
Otherwise users can get confused about when they run FEX, it just
crashes due to requiring new features early in initialization.
2026-01-03 13:47:47 -08:00
LC 62383a1c72 Merge pull request #5191 from Sonicadvance1/36
unittests/FEXLinuxTests: Force clang building for tests
2026-01-03 03:06:38 -05:00
LC eb275769a5 Merge pull request #5193 from Sonicadvance1/37
unittests/ASM: Adds test for flags clobber in `TelemetrySetValue`
2026-01-03 02:22:00 -05:00
Ryan Houdek 543a435b9f unittests/ASM: Adds test for flags clobber in TelemetrySetValue
Showcases the bug that #5192 fixed. This would have failed prior to that
PR's change.
2026-01-02 21:35:12 -08:00
Ryan Houdek 281981e619 Merge pull request #5192 from bylaws/telemfix
OpcodeDispatcher: Explicitly calculate flags after _TelemetrySetValue
2026-01-02 21:29:51 -08:00
Billy Laws 1ec8c8763e OpcodeDispatcher: Explicitly calculate flags after _TelemetrySetValue
Opcode handlers are written with the assumption that LoadSource will not
touch flags and this would be an annoying assumption to change. As this
is such an edge case anyway just don't defer flags and force a load of
the saved value before _TelemetrySetValue (which are implicitly saved
before it).

Fixes the following snippet in upc.exe:
AND        word ptr [ESP + ECX*0x1 + 0x80000000],DX
BTR        CX,DX
ADC        CX,word ptr SS:[EAX + ECX*0x1 + 0x80000000]
2026-01-03 03:52:51 +00:00
Ryan Houdek 7dba1a3552 unittests/FEXLinuxTests: Fix for building with clang 2026-01-02 18:49:21 -08:00
Ryan Houdek f4dec5d25e unittests/FEXLinuxTests: Force clang building for tests
Necessary to get #5187 passing.
2026-01-02 18:30:50 -08:00
Ryan Houdek cb7de45b48 Merge pull request #5190 from bylaws/freesys
Windows: Invalidate code in freed memory after the free syscall
2026-01-02 13:40:25 -08:00
Ryan Houdek f098b415db Merge pull request #5189 from bylaws/fixsig
Windows: Fix RtlWaitOnAddress signature
2026-01-02 13:38:24 -08:00
Billy Laws 2faf2eb5b6 Windows: Invalidate code in freed memory after the free syscall
Windows will populate Size/Address with those of the underlying free
region. This is important for cross-process free operations, where the
the 'before' callback is still called after the memory is freed so
querying the size of the section (which is not freed) fails.
2026-01-02 16:45:53 +00:00
Billy Laws a69539e583 Windows: Fix RtlWaitOnAddress signature
This led to loops of segfaults that wine handled behind the scenes.
2026-01-02 16:45:12 +00:00
Ryan Houdek a3779be9e1 Merge pull request #5185 from bylaws/aotload
ImageTracker: Load AOT images
2026-01-01 16:25:22 -08:00
Ryan Houdek 488959600e Merge pull request #5184 from bylaws/tslreloc
Relocations: Switch to robin_map to improve lookup perf
2025-12-31 09:46:29 -08:00
Ryan Houdek 9fa8148cc6 Merge pull request #5153 from Sonicadvance1/30
WritePriorityMutex: Add some more documentation
2025-12-31 09:45:54 -08:00
Billy Laws edff3dfa4a ImageTracker: Load AOT images
For now, a simple directory structure with filename based mapping is
used:
LOCALAPPDATA/cache/<main image id>/{<dll 1 image id>, <dll 2 image id>}

All cache blobs are loaded once at the start for simplicity, this will
also be useful in the future as metadata validation for settings etc can
be done just once.
2025-12-31 15:19:16 +00:00
Billy Laws 6b583ee697 Relocations: Switch to robin_map to improve lookup perf 2025-12-31 15:12:05 +00:00
LC c8d72eabe5 Merge pull request #5183 from Sonicadvance1/35
Frontend: Only decode REX if it is at the correct location
2025-12-30 01:59:51 -05:00
Ryan Houdek 063136c293 unittests/ASM: Adds REX encoding bug tests
To ensure we decode REX at the correct locations even with padding bytes
being encoded as REX.
2025-12-29 18:00:51 -08:00
Ryan Houdek 0b92d431f5 Frontend: Only decode REX if it is at the correct location
Otherwise it is a nop
2025-12-29 17:57:12 -08:00
Ryan Houdek b87bb1dec6 Merge pull request #5178 from bylaws/perelocload
ImageTracker: Load PE relocations when generating code caches
2025-12-29 13:28:44 -08:00
Ryan Houdek dbd802c85c Merge pull request #5182 from crueter/cmake-system-stuff
[cmake] explicit platform and bit-width checks
2025-12-29 13:19:20 -08:00
Ryan Houdek 1f6b3d50b6 Merge pull request #5170 from crueter/private-runtime-sameline
[cmake] more parenthesis cleanups, linker gc module, more same-line stuff
2025-12-29 13:13:50 -08:00
LC 2b4492c3f9 Merge pull request #5181 from Sonicadvance1/34
FEXCore: Switch constant emission to default to `NoPad`
2025-12-29 16:12:31 -05:00
crueter b75a2414dc [cmake] explicit platform and bit-width checks
Before compiler/architecture checks, we check:
- `sizeof(void*) == 8`: 64-bit systems have an 8-byte void pointer,
  whereas 32-bit systems (should) have a 4-byte pointer; since 32-bit
  hosts are completely unsupported might as well check for it just in
  case
- Only Windows and Linux are supported, so let's add an early check for
  that as well.

Signed-off-by: crueter <crueter@eden-emu.dev>
2025-12-29 15:53:42 -05:00
crueter 872aec20b8 [cmake] more parenthesis cleanups, linker gc module, more same-line stuff
I may have gotten carried away.

- I missed some stuff for end parenthesis because I accidentally
  searched within project files instead of the entire directory (so some
  thunk/test/windows stuff was missed), cleaned those up.
- `INTERFACE`, `PUBLIC`, `PRIVATE`, `RUNTIME`, `LIBRARY` should be on
  the same line as the target name. (I should really invest in making a
  style guide...)
- Some short statements were unnecessarily split across multiple
  lines--cleaned those up
- Made a common `LinkerGC` module that applies gc-sections etc. to a
  target in Release mode
- Usually for functions you want to have something on the first line,
  e.g. `FILES`/`DIRECTORY` for install, or the target/a positional
  argument, etc etc. Not always though, notably for some custom_command
  calls

TODO:
- What's with the `list(APPEND LIBS...)` stuff? It's used really
  inconsistently, sometimes not at all, sometimes it looks like there're
  duplicates? A more thorough cleanup is in order there.

Signed-off-by: crueter <crueter@eden-emu.dev>
2025-12-29 15:44:18 -05:00
Ryan Houdek 51f6722277 Merge pull request #5166 from crueter/cmake-arch-compiler-stuffs
[cmake] refactor: compiler and architecture handling
2025-12-29 12:41:38 -08:00
Ryan Houdek 9c0c969b48 Merge pull request #5169 from crueter/better-option-stuff-cmake
[cmake] better option descriptions + more consistent language
2025-12-29 12:14:49 -08:00
crueter 528e93c81a [cmake] handle uppercase processor names, error out if unsupported arch
Signed-off-by: crueter <crueter@eden-emu.dev>
2025-12-29 15:02:56 -05:00
Ryan Houdek 217bbf423b FEXCore: Switch constant emission to default to NoPad
Most constants don't need to be padded for relocations. So now that
these have all been audited, switch to defaulting to NoPad to reduce
verbosity.

The number of constant that need to be explicitly padded are now marked
and with all the prior changes, this allows bisecting if something has
gone wrong.
2025-12-29 11:45:51 -08:00
Ryan Houdek fd2ee4e990 Merge pull request #5146 from Sonicadvance1/27
`Constant` audit
2025-12-29 11:45:25 -08:00
Ryan Houdek 5bcdb3d478 IREmitter: Remove Pad default argument
Default argument is no longer used
2025-12-29 11:29:02 -08:00
Ryan Houdek 2f1017efed Core/Addressing: Constant audit 2025-12-29 11:29:02 -08:00
Ryan Houdek b794b9ed2c Core/Vector: Constant audit 2025-12-29 11:29:02 -08:00
Ryan Houdek 3fd86a953b Core/AVX_128: Constant audit 2025-12-29 11:29:02 -08:00
Ryan Houdek 5eb416df15 Core/OpcodeDispatcher.cpp: Constant audit 2025-12-29 11:29:02 -08:00
Ryan Houdek efa78ee0c6 Core/OpcodeDispatcher.h: Constant audit 2025-12-29 11:29:01 -08:00
Ryan Houdek 8269d04b57 IR/IREmitter: Constant audit 2025-12-29 11:29:01 -08:00
Ryan Houdek 5e782cc1c2 Core/Flags: Constant audit 2025-12-29 11:29:01 -08:00
Ryan Houdek 0ff3fb7f47 Core/X87: Constant audit 2025-12-29 11:29:01 -08:00
Ryan Houdek aa631c5585 Core/X87F64: Constant audit 2025-12-29 11:29:01 -08:00
Ryan Houdek 5c53583456 Core/Core: Constant audit 2025-12-29 11:29:01 -08:00
Ryan Houdek a1a30cd9a6 IR: Remove default argument for padding in Constant op
All use cases now pass a pad type in to this.
2025-12-29 11:29:01 -08:00
Ryan Houdek 851fbaec2d Merge pull request #5142 from Sonicadvance1/26
`_Constant` audit
2025-12-29 11:28:27 -08:00
crueter 9e8463d6d7 [cmake] refactor: compiler and architecture handling
- Do compiler/architecture checks EARLY, don't waste time doing random
  configuration stuff if the user can't even compile in the first place
- MSVC is unsupported, I assume? So add a check to disallow. There's
  literally no MSVC or MSC_VER checks anywhere, so...
- Rather than using the MSVC architecture definitions, use our own
  `ARCHITECTURE_arm64` et al. Hijacking existing "standard" definitions
  is a very bad idea. Also makes it more readable in CMake
- Change the x86 host check to `x86|amd64`. Some systems still refer to
  themselves as x86 despite being 64-bit for... reasons, and I saw one a
  very long time ago that referred to it as amd64. This should
  basically never come up, nor is it really relevant given that FEX is
  for arm64... but it kinda annoyed me so whatever.

TODOs:
- Should we check `CMAKE_SIZEOF_VOID_P (equal) 64`? I don't think anyone
  is even trying to compile this thing on armv7 or older, but might as
  well? maybe?
- What's the status of *BSD, Solaris, macOS? Technically macOS does
  support Wine, not sure about the others.

Signed-off-by: crueter <crueter@eden-emu.dev>
2025-12-29 14:05:09 -05:00
Ryan Houdek ba5fa35f09 WritePriorityMutex: Add some more documentation
Just my brain spinning as I try and determine what is causing some
hanging. Seems to be WINE specific so might not even be in FEX code.

Good to have some more documentation so when I read this again I don't
need to make some more logic deductions.
2025-12-29 11:04:50 -08:00
Ryan Houdek a480793708 Passes/RegisterAllocationPass: _Constant audit 2025-12-29 11:04:35 -08:00
Ryan Houdek 477b72ba52 OpcodeDispatcher/OpcodeDispatcher: Partial _Constant audit
`MOVGPRImmediate` is changing in another PR and we need to come back to
it.
2025-12-29 11:04:35 -08:00
Ryan Houdek f63ba7e3be OpcodeDispatcher/Vector: _Constant audit 2025-12-29 11:04:35 -08:00
Ryan Houdek 54dca47e09 OpcodeDispatcher/X87: _Constant audit 2025-12-29 11:04:35 -08:00
Ryan Houdek 53e0c8d5bf OpcodeDispatcher.h: _Constant audit 2025-12-29 11:04:35 -08:00
Ryan Houdek 8588c22170 IREmitter: _Constant audit 2025-12-29 11:04:35 -08:00
Ryan Houdek c2d5ee43c6 Passes/x87StackOptimizationPass: _Constant audit 2025-12-29 11:04:35 -08:00
Ryan Houdek c7c6855740 IREmitter: Allow the constant pool to understand padtype and bytes
Fixes an issue that a potential same constant could be padded for one
use and not padded for another.
2025-12-29 11:04:34 -08:00
Ryan Houdek 79f2832591 OpcodeDispatcher: Rename LoadConstantShift
Kind of annoying that it is overlapping with LoadConstant.

NFC
2025-12-29 11:04:34 -08:00
Ryan Houdek cb432548bf IR: Allow passing padding and MaxBytes through Constant IR 2025-12-29 11:04:34 -08:00
Ryan Houdek 900c1790d3 Merge pull request #5179 from Sonicadvance1/33
CMake: Fix mingw if host has libxxhash-dev installed
2025-12-29 11:04:02 -08:00
Ryan Houdek 1b7bce283a CMake: Fix mingw if host has libxxhash-dev installed 2025-12-29 10:47:23 -08:00
Billy Laws 702d981925 ImageTracker: Load PE relocations when generating code caches 2025-12-29 18:22:45 +00:00
Ryan Houdek 5bbbe4d2e9 Merge pull request #5140 from Sonicadvance1/25
First round of `LoadConstant` auditing
2025-12-29 10:16:23 -08:00
Ryan Houdek c54dfd9edd Merge pull request #5176 from bylaws/codmap
ImageTracker: Support codemap file generation
2025-12-29 10:14:21 -08:00
Ryan Houdek 5a47565ee1 Merge pull request #5172 from crueter/xxhash-find
[cmake] Use a Find module for xxhash
2025-12-29 10:06:00 -08:00
Billy Laws 085cf027dc ImageTracker: Do not generate codemaps when generating code caches 2025-12-29 16:25:55 +00:00
Billy Laws 299df773ce ImageTracker: Support codemap file generation 2025-12-29 16:25:54 +00:00
Ryan Houdek 9101e704ce Merge pull request #5174 from lioncash/warn
CodeCache: Fix misparenthesized expression in SaveData()
2025-12-28 12:08:32 -08:00
Ryan Houdek 212a3f45f8 Merge pull request #5164 from bylaws/gmeo
ImageTracker: Track loaded PE images for LookupExecutableFileSection
2025-12-28 12:08:07 -08:00
Lioncache 0107338020 CodeCache: Fix misparenthesized expression in SaveData()
Fixes a -Wshift-op-parentheses warning, and what seems to be a
legitimate bug, based off some quick reading.
2025-12-28 09:42:50 -05:00
LC 668e0275c3 Merge pull request #5171 from Sonicadvance1/32
Thunks/Vulkan: Update for v1.4.337
2025-12-28 07:00:23 -05:00
crueter b258546f34 [cmake] Use a Find module for xxhash
Port of #5159

Find modules are preferred for pkgconfig/otherwise non-CMake libraries.
Let's use that here for xxHash.

Signed-off-by: crueter <crueter@eden-emu.dev>
2025-12-28 01:35:46 -05:00
Ryan Houdek f24f88e46c Merge pull request #5167 from crueter/cmake-uppercase-aaaaaa
[cmake] do not use uppercase command names
2025-12-27 21:03:52 -08:00
Ryan Houdek 5747d1c5fc Merge pull request #5165 from bylaws/erebase
CodeCache: Rebase block entrypoint info
2025-12-27 21:02:14 -08:00
Ryan Houdek eb95a959bc Thunks/Vulkan: Update for v1.4.337
Fixes DXVK/vkd3d-proton on newer mesa.
2025-12-27 20:50:16 -08:00
Ryan Houdek 0edf961ce9 Merge pull request #5168 from crueter/no-tiny-vars
[cmake] reduce usage of trivial variables
2025-12-27 20:23:21 -08:00
crueter 6234297f62 [cmake] better option descriptions + more consistent language
- "Enables"/"Uses" -> Enable/Use
- Be a bit more descriptive with... the descriptions

Signed-off-by: crueter <crueter@eden-emu.dev>
2025-12-27 23:11:43 -05:00
Ryan Houdek 9c29ae486c External/Vulkan-Headers: Update to v1.4.337 2025-12-27 18:39:06 -08:00
crueter 588fec3b89 [cmake] reduce usage of trivial variables
Trivial variables like SRCS, NAME, etc. actually do more harm than good.
They *will* make your IDE mad, and are also less readable. Remember:
verbosity is not a bad thing! Usually

Also: did a few tiny cleanups that I missed from my `endpara` PR.

Signed-off-by: crueter <crueter@eden-emu.dev>
2025-12-27 21:16:51 -05:00
crueter e42dd5ddfb [cmake] do not use uppercase command names
Command names shouldn't be uppercase. A lot of LSPs scold you if you do
this, and apparently it's the official recommendation from the CMake
team to absolutely never, ever do this.

Signed-off-by: crueter <crueter@eden-emu.dev>
2025-12-27 21:09:12 -05:00
Billy Laws ecc16033be ImageTracker: Track loaded PE images for LookupExecutableFileSection 2025-12-28 00:31:50 +00:00
Billy Laws 30d0dbd2f0 Windows: Extend scoped HANDLE wrapper 2025-12-28 00:30:47 +00:00
Billy Laws f2bbc0eccd CodeCache: Use defaulted dtor for ExecutableFileInfo 2025-12-28 00:30:47 +00:00
Billy Laws 0152f3adb2 CodeCache: Rebase block entrypoint info
Entrypoint information contains guest addresses that must be rebased for
correctness.
2025-12-28 00:25:56 +00:00
Ryan Houdek ce9824a479 Merge pull request #5161 from FEX-Emu/bylaws-patch-1
WritePriorityMutex: Fix rare case of dropped read waiter wakes
2025-12-27 15:12:09 -08:00
Billy Laws 2edee2855c WritePriorityMutex: Fix rare case of dropped read waiter wakes
The Race:
1. A Reader sets `READ_WAITER_BIT` (Bit 15) and sleeps on the High 16 bits (`Futex+2`).
2. Writer A unlocks. It clears `READ_WAITER_BIT` (in Low 16 bits) and `WRITE_OWNED` (in High 16 bits).
3. Writer B immediately steals the lock. It sets `WRITE_OWNED` but preserves the now-cleared `READ_WAITER_BIT`.
4. The Reader, checking `Futex+2`, sees `WRITE_OWNED` is set. Since it cannot see that Bit 15 was unset (as it is watching High 16 bits), it assumes its wait signal is still valid and sleeps.
5. Writer B unlocks. It sees no `READ_WAITER_BIT` and wakes nobody. Deadlock.

The Fix:
Move `READ_WAITER_BIT` to Bit 30 (High 16 bits).

Now, when Writer A clears the flag, the High 16 bits change value which will prevent the wait from occurring within WaitForAddress
2025-12-27 22:52:24 +00:00
Ryan Houdek b41b967ba5 Merge pull request #5160 from crueter/cmake-endpara-sameline
[cmake] prefer end parenthesis on same line, no space after some calls
2025-12-24 16:42:50 -08:00
crueter 43173df446 [cmake] prefer end parenthesis on same line, no space after some calls
- Some CMake LSPs have aneurysms when you put the end parenthesis on a
  different line. Annoying? Yes, but this is all we can really do about
  it for now.
- `set`, `option`, and `message` should not have spaces before their
  opening parenthesis.

Signed-off-by: crueter <crueter@eden-emu.dev>
2025-12-24 19:11:42 -05:00
Ryan Houdek f153d86bce Merge pull request #5158 from crueter/obj-link-libs
[cmake] FEXCore: further reduce library redundancy
2025-12-24 16:04:40 -08:00
crueter 4ebcdf8720 [cmake] FEXCore: further reduce library redundancy
- `AddObject` doesn't need an additional Type parameter since it's
  already assumed to be an object library
- The static and shared libraries don't need any explicit compile
  options, as the object already handled this.
- They also don't need to be linked to FEXCore_Base. Object libraries
  already handle that for us since the symbols already get pulled in
  anyways.
- The object library doesn't need an output name. Only the user-facing
  libraries do

Signed-off-by: crueter <crueter@eden-emu.dev>
2025-12-24 18:45:02 -05:00
Ryan Houdek bd8f6f16aa Merge pull request #5156 from crueter/no-target-include
[cmake] propagate `-ISource` to all Tools
2025-12-24 15:35:34 -08:00
crueter ec1d9aeafa [cmake] propagate -ISource to all Tools
Rather than individually adding `${CMAKE_SOURCE_DIR}/Source` as
an include directory to each target, just use `include_directories` once
in the Tools directory and each subsequent target will have this
propagated down.

Also removed a seemingly unnecessary `-I` in LinuxEmulation--maybe
needed? But I can't test compilation right now as I don't have an ARM
development environment on hand for the next day or two.

Signed-off-by: crueter <crueter@eden-emu.dev>
2025-12-24 18:25:17 -05:00
Ryan Houdek 7cdef04fc7 Merge pull request #5157 from crueter/mingw-builtin
[cmake] use `MINGW` builtin rather than custom detection
2025-12-24 14:53:04 -08:00
crueter cbd9093e27 update jemalloc
Signed-off-by: crueter <crueter@eden-emu.dev>
2025-12-24 17:41:24 -05:00
crueter f2a1243892 [cmake] use MINGW builtin rather than custom detection
CMake has had the `MINGW` builtin to describe MinGW targets since at
least version 3.2, so it can safely be used. This variable is also set
for the MSYS2 environments, so CLANGARM64 also correctly sets `MINGW`.

Note that this depends on https://github.com/FEX-Emu/jemalloc/pull/11.

Signed-off-by: crueter <crueter@eden-emu.dev>
2025-12-24 15:13:53 -05:00
LC 144c4bf408 Merge pull request #5154 from Sonicadvance1/31
VDSO: Forgot to remove a if check
2025-12-24 09:42:07 -05:00
Ryan Houdek 499970db68 Merge pull request #5151 from bylaws/getcache
Common: Use LOCALAPPDATA for GetCacheDirectory on WOW64/ARM64EC
2025-12-23 19:02:06 -08:00
Ryan Houdek bc069f2ec8 Merge pull request #5150 from bylaws/thread_optioality
CodeCache: Make LoadData Thread argument an optional pointer
2025-12-23 19:01:57 -08:00
Ryan Houdek 440aa490ce VDSO: Forgot to remove a if check
This still finds `__fex_callback_ret` on 64-bit, saving a page of
memory.
2025-12-23 18:10:24 -08:00
Billy Laws 528afbfd3b Common: Use LOCALAPPDATA for GetCacheDirectory on WOW64/ARM64EC
Windows caches can only really work safely if they are exclusive to a
prefix due to the state for things like sharing modes that is tied to
the per-prefix wineserver and codemaps encoding prefix-local paths for
system DLLs.
2025-12-23 23:48:10 +00:00
Billy Laws 064a48e965 CodeCache: Make LoadData Thread argument an optional pointer
Windows doesn't have access to Thread when loading the main image and
ntdll.
2025-12-23 23:45:07 +00:00
Billy Laws 86211e18d7 ALookupExecutableFileSection: Take thread argument as an optional pointer 2025-12-23 23:44:58 +00:00
Ryan Houdek 974ba78a93 Merge pull request #5147 from Sonicadvance1/28
Some minor NFC
2025-12-23 11:38:13 -08:00
Ryan Houdek 6196a3a6a4 Arm64Emitter: Removes Default pad type from LoadConstant
All direct usages have been audited. Now we need to do indirect usages
through the `Constant` IR operation.
2025-12-23 11:34:34 -08:00
Ryan Houdek e7ec8e3613 JIT/BranchOps: LoadConstant audit
`ThreadRemoveCodeEntry` doesn't properly have relations wired up but it
does use the Entry. So this would be broken on code caching with
relocations.
2025-12-23 11:34:34 -08:00
Ryan Houdek a5d4ea8004 JIT/JIT: LoadConstant audit 2025-12-23 11:34:34 -08:00
Ryan Houdek 0653426793 JIT/MemoryOps: LoadConstant audit 2025-12-23 11:34:34 -08:00
Ryan Houdek ba352cebc8 JIT/MiscOps: LoadConstant audit 2025-12-23 11:34:34 -08:00
Ryan Houdek 6e712bf1b6 JIT/ALUOps: LoadConstant audit
`Constant` IR op needs the frontend to be audited and pass padding
information through.
2025-12-23 11:34:34 -08:00
Ryan Houdek b23dc6a9b3 JIT/Dispatcher: LoadConstant audit 2025-12-23 11:34:34 -08:00
Ryan Houdek c2177bff09 JIT/VectorOps: LoadConstant audit 2025-12-23 11:34:34 -08:00
Ryan Houdek 8d95172118 JIT/EncryptionOps: LoadConstant audit 2025-12-23 11:34:34 -08:00
LC d582356815 Merge pull request #5139 from Sonicadvance1/24
Arm64Emitter: Initial work for LoadConstant padding audit
2025-12-23 10:03:42 -05:00
Ryan Houdek 9d3acb362a Interpreter: Move Allocator handling to its own namespace
Gets it out of the way for some reorganizing.

NFC
2025-12-22 16:45:27 -08:00
Ryan Houdek f2d0238f84 Interpreter: Move logging to its own namespace
Gets it out of the way for some reorganizing.

NFC.
2025-12-22 16:45:27 -08:00
Ryan Houdek d6f290f6d2 Arm64Emitter: Move NOP pads to before the move instructions
Recent CPUs do nop fusion with the following instruction, this gives the
CPU the best chance to do fusion with something that actually does work.

Very trivial, doesn't do this for the more complex handling below these
as counting the number of moves before nop emitting is messy.
2025-12-22 14:14:58 -08:00
Ryan Houdek d2b9bfd6ee Arm64Emitter: Changes LoadConstant to support Pad and byte width
Instead of just a trivial pad being on or off, support a tri-state
on/off/auto where on will always pad, off will never pad, and auto will
pad only if code caching is enabled.

Further augment this by allowing a byte-width to be passed in, which can
be used with pointers to force a 48-bit VA width to only ever pad to
three instructions, reducing the common worst-case situation from 4
instructions to 3. This works because we're not going to expose a VA
width larger than 47-bit to the guest.

Fixes the handful of use-cases that explicitly chose their NOP padding,
and a bug in Arm64Relocations.cpp where it was incorrectly asking to not
receive padding even though it requires it.
2025-12-22 14:14:58 -08:00
Ryan Houdek 0a18ea8f4d Merge pull request #5138 from bylaws/eCcodecachingrelocval
Frontend: Also fetch relocations and section bounds when validating
2025-12-22 14:13:20 -08:00
Ryan Houdek f819999884 Merge pull request #5134 from bylaws/wso
Windows: Implement _[w]sopen file APIs
2025-12-22 14:11:56 -08:00
Ryan Houdek dc764db35d Merge pull request #5129 from bylaws/volttt
Windows: Introduce ImageTracker for tracking per-loaded-image data
2025-12-22 14:09:59 -08:00
Ryan Houdek 6a49b8cec8 Merge pull request #5143 from bylaws/noisb
SHMStats: Avoid ISB usage when stats are disabled
2025-12-22 13:57:15 -08:00
Tony Wasserka eb425fe640 Merge pull request #5127 from neobrain/feature_automatic_code_cache
CodeCache: Implement automatic cache generation
2025-12-22 21:55:10 +00:00
Tony Wasserka 5ca549ef5f Merge pull request #5145 from neobrain/fix_lefs
FEXOfflineCompiler: Implement SyscallHandler::LookupExecutableFileSection
2025-12-22 21:54:46 +00:00
Tony Wasserka d242ba7a52 FEXOfflineCompiler: Implement SyscallHandler::LookupExecutableFileSection
This is required for proper operation since fef1993dd7.
2025-12-22 22:37:26 +01:00
Ryan Houdek da46d51f82 Merge pull request #5137 from Sonicadvance1/23
FEXCore: Revert literal optimization from #4884
2025-12-22 13:36:55 -08:00
Billy Laws 304b0e0e8e SHMStats: Avoid ISB usage when stats are disabled
A conditional here is cheaper than unconditionally stalling the
pipeline.
2025-12-22 21:29:52 +00:00
Tony Wasserka 2573bcb90f CodeCache: Turn ExecutableFileSectionInfo::FileInfo into a const reference 2025-12-22 21:20:22 +01:00
Tony Wasserka b98b377fc0 GDBJIT: Use a const reference to read ExecutableFileInfo 2025-12-22 21:20:22 +01:00
Tony Wasserka 28029092b9 FEXServer: Centralize code map path building 2025-12-22 21:20:22 +01:00
Tony Wasserka 7eb2ce827c FEXServer: Place processed and unprocessed code maps in separate subdirectories 2025-12-22 21:20:22 +01:00
Tony Wasserka 07f9426879 FEXServer: Detect more scenarios that require cache generation 2025-12-22 21:20:22 +01:00
Tony Wasserka 7ba6cfe651 FEXServer: Implement automatic code cache generation 2025-12-22 21:20:22 +01:00
Tony Wasserka 50939df6b5 FEXServer: Implement automatic aggregation of pending code map data 2025-12-22 21:20:22 +01:00
Tony Wasserka 89e9046042 LinuxSyscalls: Use filesystem locks to protect access to code maps that are still written to 2025-12-22 21:20:22 +01:00
Billy Laws 67e3bb8596 Frontend: Also fetch relocations and section bounds when validating 2025-12-22 18:50:03 +00:00
Ryan Houdek 3025a10808 FEXCore: Revert literal optimization from #4884
This no longer does anything due to #5123
2025-12-22 10:00:55 -08:00
Billy Laws a94a9eb268 WOW64: Use ImageTracker for volatile metadata handling 2025-12-22 17:19:38 +00:00
Billy Laws 47a8fc4b49 ARM64EC: Use ImageTracker for volatile metadata handling 2025-12-22 17:19:38 +00:00
Billy Laws 64ad853a5c Windows: Introduce ImageTracker for tracking per-loaded-image data
and use it to commonise volatile metadata code (adding WOW64 image
voltmd support in the process). This will be extended to track mapped
images for code caching.
2025-12-22 17:19:38 +00:00
Billy Laws 2d6dbd5600 Windows: Add 32-bit LOAD_CONFIG structs 2025-12-22 17:19:38 +00:00
Ryan Houdek 93f6a8cb4d Merge pull request #5130 from neobrain/feature_code_cache_validation
CodeCache: Implement runtime cache validation
2025-12-22 09:17:51 -08:00
Ryan Houdek fef1993dd7 Merge pull request #5123 from bylaws/whehsajsreloc
Guest relocation support
2025-12-22 09:14:03 -08:00
Ryan Houdek 71c8436877 Merge pull request #5131 from neobrain/feature_code_cache_mainexe
CodeCache: Trigger delayed cache loading for the main executables and its interpreter
2025-12-22 09:10:10 -08:00
Ryan Houdek 9a128686a1 Merge pull request #5135 from bylaws/wfix
Dispatcher: Silence warning on ARM64EC
2025-12-22 09:07:22 -08:00
Tony Wasserka 8fcb84112b CodeCache: Implement runtime cache validation 2025-12-22 17:52:43 +01:00
Billy Laws aa73cbc8b0 Dispatcher: Silence warning on ARM64EC 2025-12-22 16:22:03 +00:00
Billy Laws 651fca36ff Windows: Implement _[w]sopen file APIs 2025-12-22 16:21:18 +00:00
Billy Laws 2f4cd33950 Frontend: Fail gracefully upon encountering unexpected relocations
As the decoder can occasionally explore non-code (e.g. after a noreturn
call) it must tolerate cases where when doing so a relocation is
encountered.
2025-12-22 16:19:10 +00:00
Billy Laws bcf48c21eb Frontend: Support relocated instruction operands when generating caches
When relocations are loaded, all immediates read from memory are checked
against the relocation map and transformed into an appropriately
sign-extended entrypoint-relative variant of the specific operands
addressing mode. As almost every case of an unhandled relocation will
lead to a later, likely harder to debug, crash at runtime just bail out
early if any such cases are encountered. Note that while this
handles/detects all cases of relocated immediates, if relocations were
applied to instructions themselves (occurs in some malware variants)
these would be missed without any errors reported.
2025-12-22 16:19:02 +00:00
Billy Laws 9571a1bc30 Frontend: Clip multiblocks to section boundaries when generating caches
When compiling code at runtime there is no harm to including jumps to
different sections within a multiblock, when enforcing as such would
introduce a lookup cost for every decode invocation (or some caching).
However when compiling offline as each cache blob is tied to a specific
library these boundaries should be enforced.
2025-12-22 16:19:02 +00:00
Billy Laws 4044b39a2f CodeCache: Track relocations in ExecutableFileInfo 2025-12-22 16:19:02 +00:00
Billy Laws 2878583627 OpcodeDispatcher: Support relocated operand type variants
In order to support code caching of 32-bit libraries, any library-base
relative relocations on the guest must be transformed into FEX
relocations so e.g. absolute jumps or loads refer to the correct
location when the library is loaded at a different base address.
2025-12-22 16:19:02 +00:00
Tony Wasserka f0c7dc48d2 CodeCache: Trigger delayed cache loading for the main executables and its interpreter
FEX isn't fully initialized by the time these are mapped into memory, so their
caches must be loaded later.
2025-12-17 15:58:03 +01:00
LC d197300be7 Merge pull request #5126 from Sonicadvance1/21
unittests/ASM: Test 32-bit displacement encoding
2025-12-16 00:42:41 -05:00
Ryan Houdek 6f846db863 unittests/ASM: Test 32-bit displacement encoding
I noticed that we weren't actually testing 32-bit displacement on
64-bit. This is an edge case on 64-bit where you can have 32-bit
displacement WITHOUT it being RIP relative. This is very uncommonly used
but is technically a possibility where 64-bit code can access an
absolute address in the lower 2GB.

With negative addresses you could technically access canonical kernel
addresses in the lower 2GB if they were mapped in userspace, but since
they aren't this is untestable (FEX also doesn't handle canonical kernel
addresses correctly anyway).

`MemoryData.asm` accidentally tested this but it didn't test both load
and store sides. So let's be a bit more stringent on this.
2025-12-15 16:21:33 -08:00
Tony Wasserka 9e8915ef9e Merge pull request #5110 from neobrain/feature_nop_padding
ARM64Emitter: Force NOP padding to be enabled
2025-12-15 19:03:44 +00:00
Ryan Houdek 805a4c1ab1 Merge pull request #5113 from neobrain/feature_request_cache_population
FEXServer: Add protocol interface to request code cache population
2025-12-15 10:39:25 -08:00
Ryan Houdek c57df7309a Merge pull request #5119 from bylaws/wheiwofjsddsjkkjdsalkj
BranchOps: Use RIP relocs for direct branch targets
2025-12-15 10:35:27 -08:00
Ryan Houdek 956f97efd4 Merge pull request #5124 from bylaws/wfnamec
Windows: Switch GetSection/ExecutableFilePath to returning full paths
2025-12-15 10:33:39 -08:00
Ryan Houdek 19d3450cb3 Merge pull request #5122 from bylaws/whehsajs
CMake: Support overriding version/hash via CMake args
2025-12-15 10:31:41 -08:00
Ryan Houdek ebdbf58474 Merge pull request #5121 from bylaws/wheiwofjsddsjkkjdsalkjkjd
Windows: Split out CRT/WinAPI reimplementation
2025-12-15 10:30:29 -08:00
Ryan Houdek 37b0e9e275 Merge pull request #5118 from bylaws/wheiwofjsddsjk
WinAPI: Implement Sleep
2025-12-15 10:29:58 -08:00
Billy Laws cd934b73ca Windows: Switch GetSection/ExecutableFilePath to returning full paths 2025-12-15 17:23:23 +00:00
Billy Laws e4816b5849 BranchOps: Use a RIP reloc for TF-checking direct branches 2025-12-15 16:09:17 +00:00
Billy Laws cf1701f1e1 BranchOps: Use a RIP reloc for EC calls exiting JIT 2025-12-15 16:08:18 +00:00
Billy Laws c578df0bf1 CMake: Support overriding version/hash via CMake args 2025-12-15 15:35:03 +00:00
Billy Laws f6dab33190 Windows: Split out CRT/WinAPI reimplementation
We can use the system CRT for the offline compiler, but still want to
pull in e.g. InvalidationTracker and other helpers.
2025-12-15 15:32:26 +00:00
Billy Laws 7c8fdc1651 WinAPI: Implement Sleep
Used by std::this_thread::sleep.
2025-12-15 15:30:00 +00:00
Tony Wasserka ec6767060a Merge pull request #5111 from neobrain/feature_code_cache_loading
CodeCache: Implement cache loading
2025-12-15 15:26:34 +00:00
Tony Wasserka 4157efaaf7 LinuxSyscalls: Move code cache loading code to a helper function
This stat-mmap-load-munmap pattern will be used a number of times in this
file.
2025-12-15 16:12:43 +01:00
Tony Wasserka 70eff81f19 CodeCache: Implement cache loading 2025-12-15 16:12:43 +01:00
Tony Wasserka 602ecece50 CodeCache: Clean up OrderedContainer checks 2025-12-15 12:21:41 +01:00
Ryan Houdek a957f1f749 Merge pull request #5114 from ChanthMiao/fix/loopupcache
LookupCache: Fix mistake in nested CacheBlockMapping call
2025-12-14 08:33:29 -08:00
Changwei Miao 98f66d5028 LookupCache: Fix mistake in nested CacheBlockMapping call
Signed-off-by: Changwei Miao <chanthmiao@outlook.com>
2025-12-14 15:43:38 +08:00
Tony Wasserka da14b12e88 FEXServer: Add protocol interface to request code cache population 2025-12-11 22:19:54 +01:00
Tony Wasserka 356f1a205b ELFCodeLoader: Add interface to get main ELF FD 2025-12-11 22:19:50 +01:00
LC bf9ab7ffbe Merge pull request #5108 from Sonicadvance1/20
github/steamrt4: Additional comments
2025-12-10 17:48:49 -05:00
Tony Wasserka e849c1b702 ARM64Emitter: Force NOP padding to be enabled
When loading code caches, these constants get patched up for the new guest
address. The new value may be larger than the original, so the padding bytes
ensure the maximum of 16 bytes of encoding space is always available.
2025-12-10 23:29:43 +01:00
Ryan Houdek a44627df29 Steamrt/ServerManager: Change loop over to a blocking read
Discard any data and wait for EOF for now. That way we know when to
exit.
2025-12-09 13:08:01 -08:00
Ryan Houdek 91fab07ef1 Steamrt: Add missing toolmanifest.vdf 2025-12-09 13:08:01 -08:00
Ryan Houdek 6f027941b4 github/steamrt4: Increase retention time
To match the wine dll artifacts
2025-12-09 10:49:55 -08:00
Ryan Houdek 53925dcc3d Merge pull request #5109 from smcv/wip/smcv/steamrt870
Steam: Don't let the FEXServer inherit FEXServerManager's original stdout
2025-12-09 10:49:48 -08:00
Simon McVittie b821022247 Steam: Don't let the FEXServer inherit FEXServerManager's original stdout
Closing the original stdout is how we signal to a larger compatibility
tool (in practice pressure-vessel) that either we are ready, or we have
crashed: when every process that held the write end open has closed it,
the compatibility tool sees EOF on the pipe. If the FEXServer inherits
stdout and holds it open, then that state will never be reached.

steamrt/tasks#870

Signed-off-by: Simon McVittie <smcv@collabora.com>
2025-12-09 18:32:33 +00:00
LC 296988be02 Merge pull request #5107 from Sonicadvance1/19
Various trivial fixes for #5106
2025-12-08 21:48:27 -05:00
Ryan Houdek 788e8a6d87 LinuxSyscalls: Don't conflict with potential define for SA_RESTORER 2025-12-07 10:08:59 -08:00
Ryan Houdek 2423849196 VDSO: Use FHU::Syscalls::getcpu 2025-12-07 10:08:59 -08:00
Ryan Houdek 609e3e2e97 SeccompEmulator: Use FHU::Syscalls::tgkill 2025-12-07 10:08:59 -08:00
Ryan Houdek 2a5c1684de AllocatorHooks: Add missing header 2025-12-07 10:08:59 -08:00
Ryan Houdek 31d89bec66 FHU: Fix copy and paste error 2025-12-07 10:08:59 -08:00
Ryan Houdek cf5435e477 AsyncNet: Use decltype to get iov count
This could differ between size_t or int on implementations.
2025-12-07 10:08:59 -08:00
Ryan Houdek e8591090f2 Merge pull request #5105 from gio3k/fix/debug-strace-fmt
Syscalls: Fix DEBUG_STRACE printing
2025-12-05 22:40:42 -08:00
Gio 486c8805ed Syscalls: fix TraceFormatString for DEBUG_STRACE
Replaces the % strings with {}
2025-12-06 14:00:44 +08:00
LC 2e2563adc0 Merge pull request #5103 from Sonicadvance1/17
code-format-helper: Update urllib3 dependency
2025-12-05 21:24:00 -05:00
LC c4258be693 Merge pull request #5104 from Sonicadvance1/18
JIT: Fixes typo
2025-12-05 21:23:29 -05:00
Ryan Houdek f7eedc1f06 JIT: Fixes typo
A previous PR fixed this in the generated SourceOutline.md file. Add the
typo fix back in the file that generates it.
2025-12-05 17:12:24 -08:00
Ryan Houdek ba1b4744c5 Docs: Update for release FEX-2512 2025-12-05 16:11:40 -08:00
Ryan Houdek bd13c02451 Merge pull request #5092 from Sonicadvance1/12
SyscallsSMCTracking: Workaround assert in ELF mapping
2025-12-05 15:44:09 -08:00
Ryan Houdek 2c63bde7b3 code-format-helper: Update urllib3 dependency
To silence dependabot warnings.
2025-12-05 15:08:54 -08:00
Ryan Houdek 5902b175f9 SyscallsSMCTracking: Workaround assert in ELF mapping
The ELF tracking thing has an expectation that only portions of ELF
files that are described in the program headers will be mapped
executable. This doesn't hold true as programs will remap random
portions of ELF files as executable. In the case that this occurs, don't
assert out and instead print a warning.

This was discovered as Node.js remaps a portion of itself executable
that isn't described as such in the program headers. I also have a local
unittest that exposes the same problem. I had discovered same problem in
some other program with #5038.

Also fixes a bug where sometimes completely anonymously mapped
executable sneak in and cause a crash, which is kind of silly.
2025-12-05 14:50:58 -08:00
Ryan Houdek d0e47f9073 Merge pull request #5096 from neobrain/feature_fexofflinecompiler
CodeCache: Implement offline compiler for cache generation
2025-12-05 14:47:58 -08:00
Ryan Houdek bd7215d36f Merge pull request #5093 from Sonicadvance1/13
SteamRT4: Adds support for building the steam depot
2025-12-05 14:39:30 -08:00
Ryan Houdek f3f134f9de FEXServer: Add support for FEX Logging control
Always enables FEXServer log thread when built for Steam so that clients
can be controlled with `STEAM_FEX_LOG=1`. FEX logs will then always go
to FEXServer and those can get directed to wherever pressure-vessel
chooses.
2025-12-04 14:06:18 -08:00
Ryan Houdek cdbf5d57bd Steam: Adds FEXServerManager
This is fairly simple. Needs to be installed alongside FEXServer, so
that it can start it in a portable config.

- Starts a FEXServer
- Tells pressure-vessel when FEXServer is ready
- Keeps FEXServer alive with the `watch_fd` as long as the process lives
- Listens for pressure-vessel to be shutting down
- Exits once pressure-vessel exits, also letting FEXServer shutdown if
  no FEX instances are alive.
2025-12-04 14:06:17 -08:00
Ryan Houdek 1ae50bd670 FEXServerClient: Split out Connect and Start
Allow an optional watch_fd to be passed to FEXServer.
2025-12-04 14:02:43 -08:00
Ryan Houdek d28c9f9843 FEXServerClient: Disallow abstract named sockets under Steam
We don't want clients connecting to random sockets.
2025-12-04 14:02:43 -08:00
Ryan Houdek fe7d52aa78 github/steamrt4: Add artifacts 2025-12-04 14:02:43 -08:00
Ryan Houdek fc0907f8c1 Steam: Add FEXCompatTool 2025-12-04 14:02:43 -08:00
Ryan Houdek e57678d7a6 Config: Move Steam configs into config system 2025-12-04 14:02:43 -08:00
Ryan Houdek 45e594e806 Utils/StringUtils: Add in-place token replace helper 2025-12-04 14:02:43 -08:00
Ryan Houdek 87e7a0effa CMake: Don't install test thunk if tests aren't enabled 2025-12-04 14:02:43 -08:00
Ryan Houdek 4fd1a35b2a CMake: Disable some installs when building for Steam 2025-12-04 14:02:42 -08:00
Ryan Houdek c460cf0678 Merge pull request #5097 from wcampbell-nv/cmdline-map
Remap /proc/pid/cmdline with PR_SET_MM_MAP
2025-12-04 14:02:19 -08:00
Tony Wasserka 983802da61 CodeCache: Add offline compiler for generating caches 2025-12-04 19:16:51 +01:00
Tony Wasserka 49273e0d59 CodeCache: Add workaround for clang-15's broken std::piecewise_construct 2025-12-04 19:16:51 +01:00
Tony Wasserka 76b8459cdc CodeCache: Zero-initialize FileId in ExecutableFileInfo 2025-12-04 19:16:51 +01:00
Tony Wasserka e269eb6f65 ELFCodeLoader: Add helper interfaces 2025-12-04 19:16:51 +01:00
Tony Wasserka 7b4774f375 ELFCodeLoader: Add option to skip interpreter loading 2025-12-04 19:16:51 +01:00
Tony Wasserka 70e9a25112 Syscalls: Move m(un)map to a dedicated interface 2025-12-04 19:16:51 +01:00
Tony Wasserka 9fb83ea56a Windows/CRT: Implement lseek 2025-12-04 19:16:51 +01:00
LC 8bb3398376 Merge pull request #5101 from Sonicadvance1/16
SVE256: Fixes AVX scalar round with insert
2025-12-04 11:24:49 -05:00
Will Campbell d42fbb3d4d Use LoadFileToBuffer + cleanup 2025-12-04 07:54:45 -08:00
Tony Wasserka 53b2245dc1 Merge pull request #5098 from Sonicadvance1/14
docs: Update ProgrammingConcerns
2025-12-04 09:52:59 +00:00
Ryan Houdek db14975828 InstcountCI: Update 2025-12-04 01:45:32 -08:00
Ryan Houdek a5139d2710 unittests/ASM: Adds test case for AVX scalar round with insert bug 2025-12-04 01:44:12 -08:00
Ryan Houdek 9f584c8014 SVE256: Fixes AVX scalar round with insert
We were using the incorrect source registers on SVE256 implementation of
these instructions.

Fixes #5100
2025-12-04 01:43:07 -08:00
Tony Wasserka 6f1b98fb52 docs: Clean up ProgrammingConcerns 2025-12-04 10:18:26 +01:00
Ryan Houdek c258a90505 docs: Update ProgrammingConcerns
Disallow all APIs that touch `FILE`, they all allocate memory that we
don't control.
2025-12-03 20:32:33 -08:00
Will Campbell eb47ef43a7 Read from a fd rather than a FILE 2025-12-03 19:59:13 -08:00
Will Campbell 1876d6b923 Remap cmdline 2025-12-03 18:15:33 -08:00
Ryan Houdek 90c8fcf393 Merge pull request #5094 from pmatos/fix/issue5084
Set current code block in x87 pass
2025-12-02 14:43:54 -08:00
Ryan Houdek 6af90575e9 Merge pull request #5095 from neobrain/feature_serialize_relocations
JIT: Add support for serializing relocations
2025-12-02 14:43:42 -08:00
Tony Wasserka 994613260c CodeCache: Support reverse application of relocations
This allows code to be serialized consistently across runs.
2025-12-02 22:27:53 +01:00
Tony Wasserka 1430fa8220 CodeCache: Move ApplyCodeRelocations 2025-12-02 18:38:59 +01:00
Tony Wasserka 33e06058c6 JIT: Move ApplyRelocations to CodeCache 2025-12-02 18:38:59 +01:00
Tony Wasserka b032d1e1f7 JIT: Make relocations relative to guest base before serialization
This ensures consistency of generated code caches across multiple runs.
2025-12-02 18:38:59 +01:00
Tony Wasserka d6b43b1fe6 Dispatcher: Add public interface to query ExitFunctionLinkerAddress 2025-12-02 17:56:56 +01:00
Paulo Matos fc771c8683 asm_tests: Set current code block in x87 pass 2025-12-02 15:33:08 +01:00
Paulo Matos f91ac09f87 Set current code block in x87 pass
This resets the constant pool in IREmit used by SelectAddressMode().

Fixes #5084.
2025-12-02 15:33:08 +01:00
Ryan Houdek e4fa399412 Merge pull request #5091 from neobrain/feature_better_relocations
JIT: Prepare FEX relocations for code caching
2025-12-01 13:21:29 -08:00
Tony Wasserka 952e949e10 JIT: Clean up block tail writing code 2025-12-01 20:13:51 +01:00
Tony Wasserka 3d093d66fb JIT: Change relocation offset base to CodeBuffer start 2025-12-01 20:13:51 +01:00
Tony Wasserka 05fe2893c7 JIT: Emit relocation from IROP_THUNK 2025-12-01 20:13:51 +01:00
Tony Wasserka 6fc17294b6 JIT: Add relocation for guest RIP stored in jump thunks 2025-12-01 20:13:51 +01:00
Tony Wasserka 6607921bee JIT: Add relocation for guest RIP stored in block tail 2025-12-01 20:13:51 +01:00
Tony Wasserka 0b0793438f JIT: Add relocation for constants relative to the guest entrypoint 2025-12-01 20:13:51 +01:00
LC 3dd591e760 Merge pull request #5088 from Sonicadvance1/11
Linux: Disable io_uring
2025-12-01 14:05:56 -05:00
LC f290d2f899 Merge pull request #5090 from neobrain/refactor_3waycomp
IR: Replace hand-written operators with three-way comparison
2025-12-01 14:05:04 -05:00
Tony Wasserka a676ad7193 JIT: Add explicit padding to relocation descriptors
This ensures zero-initialization, which is required to make code cache
generation produce consistent results.

Also consolidated header fields.
2025-12-01 19:28:39 +01:00
Tony Wasserka 096c408ef6 IR: Replace hand-written operators with three-way comparison 2025-12-01 17:10:06 +01:00
Ryan Houdek 00b65f76b6 Linux: Disable io_uring
This allows passing around `epoll_event` structs which can't be
rewritten due to queues being managed by userspace.
2025-11-30 15:53:34 -08:00
Ryan Houdek 39fb266282 Merge pull request #5087 from neobrain/refactor_simpler_calls
OpcodeDispatcher: Simplify convoluted logic for computing call offsets
2025-11-28 10:10:22 -08:00
Ryan Houdek 3969d0ac78 Merge pull request #5086 from neobrain/fix_invalid_iterators
LookupCache: Fix use of invalidated iterators
2025-11-28 09:55:20 -08:00
Tony Wasserka 439c6bb3c0 OpcodeDispatcher: Simplify convoluted logic for computing call offsets 2025-11-28 11:32:11 +01:00
Tony Wasserka 5cedbf9d34 LookupCache: Fix use of invalidated iterators 2025-11-28 11:29:46 +01:00
Tony Wasserka 427b235eb5 Merge pull request #5071 from Sonicadvance1/2
Github: Add a steamrt4 builder
2025-11-28 08:59:46 +00:00
LC 92d5ba580f Merge pull request #5085 from Sonicadvance1/10
HostFeatures: Extend LRCPC2 errata to more CPUs
2025-11-28 00:41:50 -05:00
Ryan Houdek bc2f331c8b HostFeatures: Extend LRCPC2 errata to C1 Ultra/Premium 2025-11-27 21:29:17 -08:00
Ryan Houdek ca58aef676 HostFeatures: Extend LRCPC2 errata to V3AE 2025-11-27 21:18:51 -08:00
LC 379dc405f6 Merge pull request #5080 from Sonicadvance1/9
FEXInterpreter: Fixes crash with code maps
2025-11-27 20:42:20 -05:00
LC d74b5c42da Merge pull request #5079 from Sonicadvance1/8
FEXServer: Add support for a `wait_fd`
2025-11-27 20:41:22 -05:00
LC 6004971439 Merge pull request #5077 from Sonicadvance1/6
Scripts/InstallFEX: Fixes two issues
2025-11-27 20:39:07 -05:00
LC a12b8927bc Merge pull request #5076 from Sonicadvance1/5
Utils/WritePriorityMutex: Support being forkable
2025-11-27 20:38:37 -05:00
Ryan Houdek 57e23b289a Merge pull request #5081 from esullivan-nvidia/main
HostFeatures: Disable SupportsTSOImm9 for some CPUs
2025-11-27 16:22:40 -08:00
Tony Wasserka 8214ffccf0 Merge pull request #5083 from bylaws/wheiwofjsd
JIT: Fix indirect delinker branch distance
2025-11-27 15:02:47 +00:00
Billy Laws 547135dc2d JIT: Fix indirect delinker branch distance
This is in insts not bytes.
2025-11-27 14:42:56 +00:00
Ryan Houdek 9c72113161 Merge pull request #5070 from wcampbell-nv/cmdline
Reflect application changes to argv[0] in /proc/self/cmdline
2025-11-26 19:50:22 -08:00
esullivan 8cc967fa22 HostFeatures: Disable SupportsTSOImm9 for some CPUs
This change avoids using the LDAPUR instruction with CPUs that are know to be
impacted by an ARM CPU errata that results in poor performance.
2025-11-26 21:42:29 -06:00
Ryan Houdek 4bd30bb72d HostFeatures: Fixes bug in HostFeatures where simulator doesn't support new things
We now have a machine in CI that requires this.
2025-11-26 15:43:37 -08:00
Ryan Houdek 5eeb4dabbd Github: Add a steamrt4 builder
This ensures we don't break downstream projects.
2025-11-26 14:18:08 -08:00
Tony Wasserka a27c4b3860 Merge pull request #5078 from Sonicadvance1/7
CPUID: Fixes regression from #5033
2025-11-26 15:19:51 +00:00
Ryan Houdek dfee08f74f FEXInterpreter: Fixes crash with code maps
When the realpath of a program path can't be resolved, we weren't setting
the config option. This was cascading to be a crashing in codemaps where
it was unconditionally using the optional value (with assert checks),
and causing things to crash.

Pass in the path that can't be resolved to work around a crash in PV
that can happen.
2025-11-25 16:48:47 -08:00
Will Campbell 0e2629bdd4 Address review feedback 2025-11-25 14:09:29 -08:00
Ryan Houdek 05c8630b07 FEXServer: Add support for a wait_fd
This was a requested feature. To make sure that FEXServer is running and
managed by a parent process, we need to have a way to tell FEXServer to
keep alive without any FEX clients. The best way to do this is to pass
FEXServer a Pipe (like FEX does when a client starts it), but instead of
FEXServer signaling to FEXInterpreter that it's ready. FEXServer listens
to the pipe to see if the management process is still alive.

The expectation here is that the management process passes FEXServer the
read end of a pipe, and when the management software is done (or gets
killed by the kernel!) then the write end of the pipe is closed, and
FEXServer naturally closes (As long as there's no FEX processes
remaining).
2025-11-25 12:52:16 -08:00
Ryan Houdek 1cffe618d2 CPUID: Fixes #5033
This leaf changed to being non-constant on that PR since CPUID function
1h returns APICID now.
2025-11-25 11:25:36 -08:00
Ryan Houdek 98c7bb23b5 FEX/InstallFEX: Fixes issue of installing without software-properties-common
Checks to see if the package is installed first before trying to use it.
Fixes an issue where fresh users don't have this package installed and
the script fails.
2025-11-24 14:41:48 -08:00
Ryan Houdek 922853cee1 Scripts/InstallFEX: Fixes #4972
Makes sure that stderr output doesn't cause weird interactions with
FEXRootFSFetcher.
2025-11-24 14:40:56 -08:00
Tony Wasserka a251e61859 Merge pull request #5075 from Sonicadvance1/4
Config: Document the new `FEX_APP_CACHE_LOCATION` option
2025-11-24 20:24:59 +00:00
Ryan Houdek 9c19799023 Config: Document the new FEX_APP_CACHE_LOCATION option
I forgot to document this in the man page.
2025-11-24 12:07:54 -08:00
Ryan Houdek f423b110a8 Utils/WritePriorityMutex: Support being forkable
This will be useful to fix the mutex locking mess that occurs currently
when forks occur. Instead of needing to be /very/ meticulous with many
futexes, we can instead have working threads shared_lock this one, then
when a fork occurs just only have the forker themselves unique_lock and
let the readers drain out. Since it's write-priority it'll happen quite
quickly, letting the fork get in and out relatively easily.

This is going to take some massaging to get the frontend and FEXCore to
a place that this works but we can get this simple change in early.
2025-11-24 11:50:20 -08:00
Ryan Houdek 6fd471e652 Merge pull request #5073 from discapes/unsquashfs-deco-fix
Support detecting unsquashfs>4.7.0 decompressors
2025-11-24 09:05:08 -08:00
Tony Wasserka 3d69029d33 Merge pull request #5072 from Sonicadvance1/3
Minor fixes
2025-11-24 11:24:07 +00:00
Miika Tuominen 2258f2e424 Support detecting unsquashfs>4.7.0 decompressors 2025-11-22 14:38:59 +02:00
Ryan Houdek cca5a68e20 SHMStats: Add missing header 2025-11-21 18:01:37 -08:00
Ryan Houdek da5c9bff68 Async: Add missing header. 2025-11-21 18:01:33 -08:00
Ryan Houdek 5b87f0699b pidof: Switch to using ranges 2025-11-21 18:01:28 -08:00
Will Campbell a67fe561a1 Reflect application changes to argv[0] in /proc/self/cmdline 2025-11-21 16:00:57 -08:00
Tony Wasserka e2f4065376 Merge pull request #5065 from Sonicadvance1/1
FEXCore/CodeCache: Move spin-loop over to a WFE loop
2025-11-21 14:21:45 +01:00
Ryan Houdek d3bf87f4f4 Merge pull request #4985 from Sonicadvance1/fex-atomic
Support (downstream) kernel-side unaligned atomic handling (The rebase sequel)
2025-11-20 17:09:29 -08:00
Ryan Houdek 6d351ec47f FEXCore/CodeCache: Moves spin-loop in to a WFE loop
Saves power and responds faster. Pass in the atomic to `WaitPred` with
the predicate checking if the buffer has been flushed yet. Same
behaviour as previous code but more efficient on our hardware.
2025-11-20 14:23:01 -08:00
Ryan Houdek 4dc1dd2511 FEXCore/Utils/SpinWaitLock: Adds WaitPred for waiting on a predicate
This simplifies the loop a bit and moves the non-predicated exact
matching version to use the predicated version.

We will need a predicated version for the next commit.
2025-11-20 14:23:01 -08:00
Ryan Houdek 5205ae40fa Merge pull request #5062 from pmatos/fix/address-size-handle
Implement address size modifier handling in CMPSOp and SCASOp
2025-11-20 12:19:22 -08:00
Ryan Houdek 32f1dcde7e Merge pull request #4906 from neobrain/feature_code_maps
CodeCache: Introduce code maps
2025-11-20 12:11:16 -08:00
Ryan Houdek 6772581c53 Arm64ec: Print a log when kernel unaligned atomics are used
Not having this in FEXInterpreter as it is too spammy in the general
case.
2025-11-20 12:05:55 -08:00
Ryan Houdek 387201815b Config: Add option to enable or disable kernel backpatchin on unaligned atomic
Default to enabled because this is the config we expect by default.
In the future will get some benchmarking in various games like
Assassin's Creed, and Call of Duty.
2025-11-20 12:04:59 -08:00
Billy Laws 24d61e1125 Windows: Enable downstream kernel-side unaligned atomic handling 2025-11-20 12:04:59 -08:00
Billy Laws 04259f031d FEXLoader: Enable downstream kernel-side unaligned atomic handling 2025-11-20 12:04:59 -08:00
Tony Wasserka 6403da3715 LinuxSyscalls: Shield code map FD from guest access
This prevents chromium/CEF from closing the FD.
2025-11-20 19:13:18 +01:00
Tony Wasserka 90cb76312c LinuxSyscalls: Implement code map writing for future code caching 2025-11-20 19:13:18 +01:00
Tony Wasserka b6cff01abb FEXServer: Add support for querying code maps
Managing code maps in FEXServer rather than in FEXInterpreter makes it
easier to handle multiple concurrent processes sharing code caches for
the main executable and libraries.
2025-11-20 19:13:18 +01:00
Tony Wasserka b34b711161 CodeCache: Add interfaces to describe and generate code maps
Code maps describe per-binary metadata used to generate caches. Currently,
this includes compiled block offsets and loaded shared libraries.
2025-11-20 19:13:18 +01:00
Ryan Houdek 709d767d61 FEX/Config: Allow override of cache location
This will be used.
2025-11-20 18:56:22 +01:00
Tony Wasserka e075916154 Config: Add interface to query cache directory 2025-11-20 18:56:22 +01:00
Paulo Matos de10154f29 instcountci: Implement address size modifier handling in CMPSOp and SCASOp for 64bits 2025-11-20 13:42:19 +01:00
Paulo Matos ba71e79e54 asm_tests: Implement address size modifier handling in CMPSOp and SCASOp 2025-11-20 13:42:19 +01:00
Paulo Matos 2cc70b8051 Implement address size modifier handling in CMPSOp and SCASOp for 64bits
A few games were generating "Can't handle adddress size".
I implemented 0x67 prefix handling for CMPSOp and SCASOP and improved
the error messages for the remainder. This will implement the address
modifier on 64bit systems, and keep issuing an error on 32bits.
2025-11-20 13:42:19 +01:00
Paulo Matos 9d965f94de asm_tests: Add 32-bit CMPS/SCAS tests without address size override 2025-11-20 13:42:19 +01:00
Ryan Houdek 40c2db4744 Merge pull request #5006 from bylaws/fasterrrrr
Introduce two-pass code invalidation model
2025-11-19 17:41:01 -08:00
Billy Laws 9a7285dca4 Windows: Support new two-stage invalidation model 2025-11-20 00:38:03 +00:00
Billy Laws 8c00ac78b1 Linux: Support new two-stage invalidation model 2025-11-20 00:38:03 +00:00
Billy Laws cf4478eeee LookupCache: Introduce two-pass code invalidation model
Shared code buffer support introduced the concept of having a single
GuestToHostMaps shared across many threads. In the common case all
threads will share one however if e.g. a resize recently occured and
specific thread is yet to compile any code with the new codebuffer it
will still use the old GuestToHostMap. The current invalidation
approach handles this by repeatedly calling erase for every single
thread's GuestToHostMap, even if it is repeated. An accumulator is used
to ensure when two threads share a map, the L1/L2 cache entries in the
second thread will still be invalidated even if the the iteration for
the first thread removed them from the map.

Unfortunately this is incredibly slow in cases with many threads, as
a significant number of redundant map lookups and L1/L2 cache erasures
on threads that never even observed a given block can occur. Solve this
by introducing a two-pass model:
- First, all active codebuffers (and their associated GuestToHostMaps)
  have their entries invalidated for the given range, these codebuffers
  are tracked internally within FEXCore. It is at this point that delinking
  callbacks are ran.
- Second, each thread will have its caches invalidated. But rather than
  naively invalidating the L1/L2 caches for every invalidated block for
  every thread, threads now track on their own what specific entries
  have been potentially fetched into their L1/L2 caches. This is
  aided by GuestToHostMap now tracking the pages each block touches. (an
  inverse CodePages so to speak).
2025-11-20 00:38:03 +00:00
Billy Laws 85c8e7f1bb fextl: Wrap tsl::robin_set 2025-11-20 00:38:03 +00:00
Billy Laws efd95efb40 FEXCore: Keep a list of weak refs to all allocated codebuffers
We currently rely on the frontend to keep track of threads and then
iterate over all threads to perform per-codebuffer operations. However
as codebuffers are shared between many threads (the common case is a
single code buffer across all) this ends up being inefficient. Introduce
a list of codebuffers to solve that (new codebuffers are very rare, so a
vector is plenty fine here for erasing invalid weak refs).
2025-11-20 00:38:03 +00:00
Billy Laws 99ad7ea45c LookupCache: Drop unused state frame argument for delinker cbs 2025-11-20 00:38:03 +00:00
Ryan Houdek aba0c57f73 Merge pull request #5067 from pmatos/fix/Nasm3
Fix movzx instruction syntax
2025-11-19 14:01:15 -08:00
Paulo Matos 8c4f6b648e Fix movzx instruction syntax
nasm 2.16 was happy with it but it generates a bunch of errors in nasm3.
The generated binaries remain the same.
2025-11-19 15:04:16 +01:00
Ryan Houdek 3b83bdd88d Merge pull request #5066 from neobrain/fix_async_asserts
Async: Adapt precondition checks when receiving FDs
2025-11-19 01:45:41 -08:00
Tony Wasserka 8d71e08b44 Async: Strengthen precondition check when receiving FDs
The sender might provide all requested message bytes but no FD. The receiver
interface has no simple way of indicating this scenario yet, so just assert
out for now to ensure it never happens in the first place.

If needed, this can be changed to return a new error code to indicate partial
read in the future.
2025-11-19 09:48:23 +01:00
Tony Wasserka c31063a8ef Async: Move file descriptor checks from read_some() to read()
This allows using read_some for incoming messages with an optional FD.
Doing so fits the purpose of read_some more closely, which is to read *any*
non-empty amount of data.
2025-11-19 09:35:52 +01:00
Tony Wasserka 28d101f1dd Revert "Merge pull request #5059 from Sonicadvance1/warkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwark"
This cherry-picked an unfinished patch that wasn't intended for merging.
2025-11-19 09:35:08 +01:00
Ryan Houdek aaef344ae3 Merge pull request #5039 from Sonicadvance1/warkwarkwarkwarkwarkwark
FEX: Moves FEX thunk callback function generation to the frontend
2025-11-18 14:40:53 -08:00
Ryan Houdek 42d0324304 FEX: Moves FEX thunk callback function generation to the frontend
Adds it to the VDSO handling, it's not necessarily a VDSO function but
it behaves as such as it is in every single process. This means we get
to reuse the mapped page for every process when thunks are built,
shaving a page out of 32-bit processes.

Also, fixes a bug in guest VDSO symbol loading where clang sticks all
symbols in to `.dynsym` where gcc sticks them in to `.symtab`. Search
both. This effectively meant the couple of guest VDSO symbols were
always failing to get found, causing us to allocate yet another page on
32-bit. So effectively three pages stolen.

This also means we can remove the Linux specific X86HelperGen stuff from
FEXCore, only passing a single "VDSO" function pointer to the backend
for the dispatcher. Once again moving the Linux stuff to the frontend is
good.

Fixes an assert about about untracked noexec code `NoExec
instruction in entry block: FFFFE000` whenever thunk callbacks were
used.
2025-11-18 14:15:04 -08:00
Ryan Houdek 2a0019347a Thunks: Adds FEX Thunk callback to VDSO
This isn't necessary a VDSO, but it is /always/ mapped in to every
process. Use it as such.
2025-11-18 14:08:43 -08:00
Ryan Houdek e0305ea1b9 Merge pull request #5056 from Sonicadvance1/warkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwark
FEX/VDSO: Fixes symbol lookup
2025-11-18 12:51:13 -08:00
Ryan Houdek 105ff47ae3 FEX/VDSO: Fixes symbol lookup
Fixes a bug in guest VDSO symbol loading where clang sticks all
symbols in to .dynsym where gcc sticks them in to .symtab. Search
both. This effectively meant the couple of guest VDSO symbols were
always failing to get found, causing us to allocate yet another page on
32-bit. So effectively three pages stolen.

Peeled out of #5039
2025-11-18 12:19:51 -08:00
Ryan Houdek 5ee190a41e Merge pull request #5060 from Sonicadvance1/warkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwark
FEXCore/Config: Expose GetConv members
2025-11-17 10:45:02 -08:00
Ryan Houdek cf37617c25 Merge pull request #5011 from pmatos/feat/opt-memcpyf80
Refactoring of storing code in x87 opt. stack pass
2025-11-17 10:25:39 -08:00
Tony Wasserka 2e9c8f0f51 FEXCore/Config: Expose GetConv members
From working branch commit 8244ca1666796267ce25741cdf1103eef4f7539d
`Make cache generation aware of FEX configuration`
2025-11-17 10:21:51 -08:00
Ryan Houdek 06c2319851 Merge pull request #5061 from Sonicadvance1/warkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwark
FEXCore/Common: Adds the ability to override HostFeatures registers by config.
2025-11-17 10:18:26 -08:00
Ryan Houdek da0668c7cc Merge pull request #5059 from Sonicadvance1/warkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwark
Async: Move file descriptor checks from read_some() to read()
2025-11-17 10:17:44 -08:00
Ryan Houdek b34df334cb Merge pull request #5053 from Sonicadvance1/warkwarkwarkwarkwarkwarkwarkwarkwarkwarkwark
LinuxSyscalls: Fix null pointer dereference in LookupExecutableFileSection
2025-11-17 10:16:59 -08:00
Tony Wasserka 9e9f2ccae1 Merge pull request #5050 from neobrain/refactor_new_config_getter
Config: Refactor value getter interface
2025-11-17 18:58:50 +01:00
Tony Wasserka 15b8f75730 Config: Drop unneeded namespaces from StringArrayType 2025-11-17 18:46:35 +01:00
Tony Wasserka 9d6b9aa574 Config: Allow reading config values in arbitrary C++ expressions
FEX_CONFIG_OPT can only be used as a standalone statement, which is
inconvenient for config values that are only used once. The new functions
(e.g. Get_DUMPIR()) can be used in conditions or other expressions.
2025-11-17 18:46:02 +01:00
Tony Wasserka 64724886af Config: Replace macro-based config readers with a C++ template 2025-11-17 18:46:02 +01:00
Paulo Matos c088369f4a instcountci: Refactoring of storing code in x87 opt. stack pass 2025-11-17 10:14:29 +01:00
Paulo Matos 39dbf46422 Refactoring of storing code in x87 opt. stack pass
Enables memcpy optimization of 80bit floats on reduced precision.

Also uncovered a bug where if we had done 80bit memcpy
optimization, we wouldn't have properly stored the 80bits.
This was caught by the existing tests when we enabled the optimization.
2025-11-17 10:14:29 +01:00
Paulo Matos 3c1b0bb917 Add instcountci tests for 80bit memcpy for x87 instructions 2025-11-17 10:14:29 +01:00
Ryan Houdek 2227170dbb FEXCore/Common: Adds the ability to override HostFeatures registers by config.
This is going to be necessary for the offline compiler work.
Also allow FEXGetConfig to print the same registers in the correct
format for easy fetching.
2025-11-16 16:43:01 -08:00
Ryan Houdek e08f421e1c Convert Async assert to logman assert 2025-11-16 14:56:30 -08:00
Tony Wasserka 42c58c5420 Async: Move file descriptor checks from read_some() to read()
This allows using read_some for incoming messages with an optional FD.
Doing so fits the purpose of read_some more closely, which is to read *any*
non-empty amount of data.
2025-11-16 14:54:46 -08:00
LC 4afbdd9afb Merge pull request #5055 from Sonicadvance1/warkwarkwarkwarkwarkwarkwarkwarkwarkwarkwarkwark
LinuxSyscalls: Fixes alloca use-after-free in `RecvMMsg`
2025-11-14 20:26:05 -05:00
Ryan Houdek 06b9e13904 LinuxSyscalls: Fixes alloca use-after-free in RecvMMsg
Easy enough fix, thanks to @OFFTKP for pointing this out.
2025-11-14 16:27:27 -08:00
Tony Wasserka 29473b43cb LinuxSyscalls: Fix null pointer dereference in LookupExecutableFileSection 2025-11-14 15:23:42 -08:00
Ryan Houdek 0427d48b98 Merge pull request #5033 from Sonicadvance1/warkwark
CPUID: Fixes APICID for processor count calculation.
2025-11-14 11:31:26 -08:00
Tony Wasserka b228746f1d FEXInterpreter: Fix incorrect value assignment
Previously this was overriding the cached value in the local Value object.
The global SilentLog variable never got updated, so a stale value would be
used.
2025-11-14 11:32:11 +01:00
Tony Wasserka 0be8485116 Merge pull request #5047 from antonkesy/fix_formatting
Align code with clang-format
2025-11-14 09:25:38 +01:00
Tony Wasserka 5d0279ff08 LibraryForwarding/gen: Tiny cleanup 2025-11-14 09:14:45 +01:00
Ryan Houdek 5d908d902c Merge pull request #5045 from lioncash/long
JIT: Handle long ADR/ADRP
2025-11-13 11:10:10 -08:00
Ryan Houdek 3a014f80f2 Merge pull request #5044 from antonkesy/clean_up_scripts
Scripts: Clean-up
2025-11-13 11:09:56 -08:00
Ryan Houdek fbefd7855c Merge pull request #5048 from neobrain/fix_fexconfig_string_lists
FEXConfig: Fix string list handling
2025-11-13 11:07:50 -08:00
Tony Wasserka 11f9135be6 FEXConfig: Fix string list handling 2025-11-13 17:33:21 +01:00
Lioncache 7bb0ce810e JIT: Expand LongAddressGen() to handle movz+movk sequence
This is only ever used on the path where we'd want to handle something
like this (in EmitEntryPoint()), so we can just extend the long handler
type instead of introducing a new type to handle this.
2025-11-13 11:16:34 -05:00
Lioncache 993b832771 JIT: Handle long ADR/ADRP
Wires up the long address handler into the ADR/ADRP restart
handlers.
2025-11-13 09:40:39 -05:00
Anton Kesy 013ac1e627 Align code with clang-format
Automatically done by running:
`find . \( -path './External' -prune \) -o \
  \( -iname '*.cc' -o -iname '*.cpp' -o -iname '*.hpp' -o \
     -iname '*.h' -o -iname '*.c' \) -print | \
  xargs clang-format --style=file -i`
2025-11-13 13:29:49 +01:00
Anton Kesy d7977a02fa remove semicolon 2025-11-12 21:45:09 +01:00
LC 73a32ff22c Merge pull request #5043 from antonkesy/fix_typos
Docs: fix typo
2025-11-12 15:28:24 -05:00
Anton Kesy ee4ae5390b move function comment inside function 2025-11-12 21:22:26 +01:00
Anton Kesy f71db11035 remove excess whitespaces 2025-11-12 21:22:08 +01:00
Anton Kesy 9497288b97 remove unused imports 2025-11-12 21:21:55 +01:00
Anton Kesy a00260d801 remove unused variable 2025-11-12 21:21:28 +01:00
Anton Kesy cad48e07e4 fix comment indentation 2025-11-12 21:21:05 +01:00
Anton Kesy d91e8a4278 docs: fix typo 2025-11-12 21:07:18 +01:00
LC faf74eee90 Merge pull request #5037 from Sonicadvance1/warkwarkwarkwark
FEXCore: Fixes JITGuardPage calculation in a threaded environment
2025-11-12 15:00:00 -05:00
Ryan Houdek 0b52e1cd14 Merge pull request #5042 from neobrain/fix_base_inference
LinuxSyscalls: Fix incorrectly inferred base address observed in glxtest
2025-11-12 11:18:58 -08:00
Tony Wasserka b62890f136 LinuxSyscalls: Fix incorrectly inferred base address observed in glxtest
At runtime, glxtest is mapped as follows:
0x000055fd9a030000 0x000055fd9a034000 0x4000  0x0     r--p  glxtest
0x000055fd9a034000 0x000055fd9a038000 0x4000  0x3000  r-xp  glxtest
0x000055fd9a038000 0x000055fd9a039000 0x1000  0x6000  rw-p  glxtest
0x000055fd9a039000 0x000055fd9a03a000 0x1000  0x6000  rw-p  glxtest

The problem here is that the last two sections can't be distinguished solely
by their mmap parameters. This would cause the wrong base address to be
inferred for the last mapping. To fix this, we can be more permissive by
allowing multiple candidates to be returned.

In practice, this only affects non-code sections, so it's not a big issue
either way.
2025-11-12 19:54:45 +01:00
Tony Wasserka de1d37eef8 Merge pull request #5041 from Sonicadvance1/warkwarkwarkwarkwarkwarkwarkwark
CodeEmitter: Removes a few spurious asserts
2025-11-12 14:34:16 +01:00
Ryan Houdek a57c557485 CodeEmitter: Removes a few spurious asserts
These are handled with restart.
2025-11-11 18:25:21 -08:00
LC 3c9f6c845b Merge pull request #5040 from Sonicadvance1/warkwarkwarkwarkwarkwarkwark
FEXCore: Remove usage of "remote atomic" xor
2025-11-11 21:20:18 -05:00
Ryan Houdek ff25e9a92e FEXCore: Remove usage of "remote atomic" xor
This is the only usage of LSE atomics that isn't the fetch variety.
[This article](https://www.phoronix.com/news/Linux-6.18-ARM64-Atomics-Issue)
reminded me that this was a thing and that I should double check the IR.
This was the only IR operation remaining that still didn't use the fetch
variety. Convert it over to the fetch to avoid the expectation that it
can be a "remote atomic". Change is going to fall in to noise, but might
as well as be consistent.
2025-11-11 17:16:03 -08:00
Ryan Houdek 6a60f72a9e FEXCore: Fixes JITGuardPage calculation in a threaded environment
While this worked great for the singular unit test. I remembered thatour
pool allocator returns the minimum working size asked for but will
return larger sizes if exact fitment couldn't occur.

Because we are dealing with guard pages, we need to return the full
buffer size to the "client" so they can tell the frontend where the
guard page actually lives. Otherwise the JIT will tell the frontend the
guard page is at the end of the requested size, blow past the limit,
and fault in a completely different location.

With a bit of logging I saw in a multithreaded environment that we were
basically always getting a larger requested buffer while Steam was
starting up.
2025-11-11 12:31:38 -08:00
Ryan Houdek 94b690df43 unittests/FEXLinuxTests: Adds cpu core count test to cpuid
Ensures cpuid core counts are reported correctly.
2025-11-11 11:20:26 -08:00
Ryan Houdek 94edbc3436 CPUID: Stop accidentally exposing the HTT bit
We don't support this.
2025-11-11 11:20:26 -08:00
Ryan Houdek 5eab1e559a CPUID: Fixes APICID for processor count calculation.
Primary fix here is returning the current CPU index in function 01h.
Intel Quartus uses this alongside affinity setting to check if all cores
can be used for its calculation. Since we had hardcoded apicid 0 here,
it assumed to only have one core and never generated worker threads.

Additional fix for apicid size. This is the size of the bitmask required
for apic ids, we weren't calculating this correctly at all. This mask is
a "maximum" number of APICs that the CPU reserves in power of two.
Say the core supports 256 APICs, but the processor only supports 16, or
any other combination.
2025-11-11 11:20:26 -08:00
Ryan Houdek 53db3ad6f2 Merge pull request #5036 from neobrain/fix_thunkgen_glibcxx_debug
CMake: Disable libstdc++'s debug mode when compiling thunkgen
2025-11-11 09:14:53 -08:00
Tony Wasserka 581f3263ed CMake: Disable libstdc++'s debug mode when compiling thunkgen
This allows the rest of the project to use _GLIBCXX_DEBUG.
2025-11-11 17:26:37 +01:00
LC 1e3c642be6 Merge pull request #5035 from pmatos/fix/gradual-mem-growth
Use gradual memory growth
2025-11-11 09:00:27 -05:00
LC 22c3cd553f Merge pull request #5034 from Sonicadvance1/warkwarkwark
FEXCore/Win32: Move WritePriorityMutex away from SRWLock
2025-11-11 08:59:29 -05:00
Paulo Matos e5743f8dae Use gradual memory growth
Use min instead of max, otherwise we are always using `MAX_STATS_SIZE`.
2025-11-11 09:23:40 +01:00
Ryan Houdek bddc2f227d FEXCore/Win32: Move WritePriorityMutex away from SRWLock
Turns out I was reading six year old code for Wine's implementation for
SRWLocks. It actually /doesn't/ use WAIT_BITSET in their implementation.
It's still write-priority but it's actually significantly slower than I
was expecting due to futex queue usage and some other implementation
details.

Instead of using Wine's implementation, use win32's Wait/Wake on address
functionality and reuse all our other mechanism for implementing this
futex. This grants us our regular low-overhead codepath that I tested on
Linux, while the fallback is the only "slow" path. This also allows us
to still support a pseudo `WAIT_BITSET` code-path that reduces
stampeding even on Win32. The reader side just waits on the upper-half
of the futex (the writer bits) and the `WaitOnAddress` means only the
exact match address will be woken. We also get the regular
reader<->writer hand-offs working.

While this path still uses the futex
queue, the majority of the time our mutexes get acquired in the WFE loop
already, so it's a significant win.

Dark Souls Remastered before:
```
  $RDLck Time: 4.531100 ms/second (0.04 percent)
  $WRLck Time: 2.122560 ms/second (0.02 percent)
```

after:
```
  $RDLck Time: 1.441620 ms/second (0.01 percent)
  $WRLck Time: 0.963720 ms/second (0.01 percent)
```
2025-11-10 17:43:47 -08:00
Ryan Houdek 686c04ea93 Win32: IMplement Wake/Wait by address 2025-11-10 17:43:40 -08:00
Ryan Houdek b38369199e Merge pull request #4893 from Sonicadvance1/long_long_codebuffer_pages
FEX: Implements support for JIT CodeBuffer guard page restart
2025-11-10 13:48:18 -08:00
Ryan Houdek e862c904a9 FEX: Implements support for JIT CodeBuffer guard page restart
When the JIT CodeBuffer overflows, we will now catch accesses to the
guard page and longjump while restarting the JIT with a larger buffer
request.

Fixes #4877
2025-11-10 11:55:21 -08:00
Ryan Houdek 43d9384b1c FEXCore/JIT: Add a pool allocator that understands a guard page
The size asked for has its final page guarded. It's up to the code
asking for allocations to ensure it never uses the final page if
necessary.
2025-11-10 11:54:25 -08:00
Ryan Houdek cb9af0b86a SignalDelegator: Split out SIGSEGV handler
This needs to run before the TestCodeHarness's frontend handler.
2025-11-10 11:53:40 -08:00
Ryan Houdek b8c17a843c ArchHelpers: Adds helper to get pointers to PC and FPRs 2025-11-07 15:33:47 -08:00
Ryan Houdek 7ad7f181d7 FEXCore/LongJump: Add a way to manually load from a longjump
The frontends will need this when loading a longjump buffer in to a
context.
2025-11-07 15:33:47 -08:00
Ryan Houdek eb0bf55033 FEXCore/JIT: Move the JIT long jump buffer to internalthreadstate
This will be a TLS variable that needs to be read by the frontend.
2025-11-07 15:33:47 -08:00
Ryan Houdek f4e3e4ad30 Merge pull request #5031 from lioncash/catch
Externals: Update catch2 from 3.5.3 to 3.11.0
2025-11-07 10:40:49 -08:00
Lioncache 5ae82410cc Externals: Update catch2 from 3.5.3 to 3.11.0
Updates it to the most recent release.
2025-11-07 08:31:04 -05:00
Ryan Houdek 2febb524e9 Merge pull request #5028 from lioncash/fmtup
Externals: Update fmt to 12.1.0
2025-11-06 10:04:15 -08:00
Lioncache b9e452133c Externals: Update fmt to 12.1.0
Keeps fmt updated to its latest release.
2025-11-06 10:17:22 -05:00
LC 747ea0a1f7 Merge pull request #5027 from Sbte/pr/xxhash
Update xxhash to v0.8.3
2025-11-06 07:38:56 -05:00
Sven Baars f8c52ca34a Update xxhash to v0.8.3 2025-11-06 11:53:55 +01:00
517 changed files with 39330 additions and 24063 deletions

No files matched your search

+1 -1
View File
@@ -2,7 +2,7 @@
Source/Common/cpp-optparse/*
# Files with human-indented tables for readability - don't mess with these
FEXCore/Source/Interface/Core/X86Tables/*
FEXCore/Source/Interface/Core/X86Tables/*.cpp
# Inline headers with list-like content that can't be processed individually
Source/Tools/LinuxEmulation/LinuxSyscalls/x*/SyscallsNames.inl
+3
View File
@@ -22,3 +22,6 @@
# Minor reformat with clang-format-19
9fdd96af61c969cb5732471223f00eda64b7a069
# Reformat of X86Tables.h
ba2b0ef809f66f1a6d334f000798fa2ceafab26f
+96 -195
View File
@@ -24,238 +24,139 @@ jobs:
fail-fast: false
steps:
- uses: actions/checkout@v3
- uses: actions/checkout@v6
with:
fetch-depth: '0'
fetch-tags: 'true'
- name: Set runner label
run: echo "runner_label=${{ matrix.arch[1] }}" >> $GITHUB_ENV
- name: Set rootfs paths
- name: Set runner info
run: |
echo "FEX_ROOTFS_MOUNT=/mnt/AutoNFS/rootfs/" >> $GITHUB_ENV
echo "FEX_ROOTFS_PATH=$HOME/Rootfs/" >> $GITHUB_ENV
echo "FEX_ROOTFS=$HOME/Rootfs/" >> $GITHUB_ENV
echo "runner_label=${{ matrix.arch[1] }}" >> $GITHUB_ENV
echo "runner_name=$(hostname)" >> $GITHUB_ENV
- name: Update RootFS cache
# Use a bash shell so we can use the same syntax for environment variable
# access regardless of the host operating system
shell: bash
run: $GITHUB_WORKSPACE/Scripts/CI_FetchRootFS.py
- name : submodule checkout
# Need to update submodules
run: |
git submodule sync --recursive
git submodule update --init --depth 1
- name: Clean Build Environment
run: rm -Rf ${{runner.workspace}}/build
- name: Create Build Environment
# Some projects don't allow in-source building, so create a separate build directory
# We'll use this as our working directory for all subsequent commands
run: cmake -E make_directory ${{runner.workspace}}/build
- name: Setup Build Environment
uses: ./.github/workflows/setup-env
- name: Configure CMake
# Use a bash shell so we can use the same syntax for environment variable
# access regardless of the host operating system
shell: bash
working-directory: ${{runner.workspace}}/build
# Note the current convention is to use the -S and -B options here to specify source
# and build directories, but this is only available with CMake 3.13 and higher.
# The CMake binaries on the Github Actions machines are (as of this writing) 3.12
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True -DBUILD_FEX_LINUX_TESTS=True -DBUILD_THUNKS=True -DCMAKE_INSTALL_PREFIX=${{runner.workspace}}/build/install
run: |
cmake -S . -B build -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True \
-DENABLE_X86_HOST_DEBUG=True -DBUILD_FEX_LINUX_TESTS=True -DBUILD_THUNKS=True \
-DCMAKE_INSTALL_PREFIX="$PWD"/build/install
# These steps make a lot of noise but rarely fail.
# Put them in a separate step to make normal build logs easier to parse
- name: Noisy Build Targets
run: cmake --build build --target asm_files 32bit_asm_files JemallocLibs Catch2 vixl cephes_128bit
- name: Build
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the build. You can specify a specific target with "--target <NAME>"
run: cmake --build . --config $BUILD_TYPE
id: build
run: cmake --build build
- name: Install
working-directory: ${{runner.workspace}}/build
shell: bash
run: cmake --build . --config $BUILD_TYPE --target install
run: cmake --build build --target install
- name: gcc target tests 64
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the gvisor tests
run: cmake --build . --config $BUILD_TYPE --target gcc_target_tests_64
# GCC tests
- name: GCC64 Target Tests
if: steps.build.outcome == 'success'
uses: ./.github/workflows/test
with:
target: gcc_target_tests_64
- name: GCC64 Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_GCC64.log || true
- name: GCC32 Target Tests
if: steps.build.outcome == 'success'
uses: ./.github/workflows/test
with:
target: gcc_target_tests_32
- name: gcc target tests 32
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the gvisor tests
run: cmake --build . --config $BUILD_TYPE --target gcc_target_tests_32
# API tests
- name: API Tests
if: steps.build.outcome == 'success'
uses: ./.github/workflows/test
with:
target: api_tests
- name: GCC32 Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_GCC32.log || true
- name: FEXCore API Tests
if: steps.build.outcome == 'success'
uses: ./.github/workflows/test
with:
target: fexcore_apitests
- name: APITest tests
working-directory: ${{runner.workspace}}/build
shell: bash
run: cmake --build . --config $BUILD_TYPE --target api_tests
# ARM emission tests
- name: ARM Emitter Tests
if: steps.build.outcome == 'success'
uses: ./.github/workflows/test
with:
target: emitter_tests
- name: APITest Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_APITests.log || true
- name: FEXCore APITest tests
working-directory: ${{runner.workspace}}/build
shell: bash
run: cmake --build . --config $BUILD_TYPE --target fexcore_apitests
- name: FEXCore APITest Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_FEXCoreAPITests.log || true
- name: ARMEmitter tests
working-directory: ${{runner.workspace}}/build
shell: bash
run: cmake --build . --config $BUILD_TYPE --target emitter_tests
- name: ARMEmitter Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_ARMEmitterTests.log || true
- name: FEXLinuxTests
working-directory: ${{runner.workspace}}/build
shell: bash
# Linux tests
- name: FEX Linux Tests
if: steps.build.outcome == 'success'
uses: ./.github/workflows/test
with:
target: fex_linux_tests_all
env:
# These tests require non-portable install due to thunks.
FEX_PORTABLE: 0
run: cmake --build . --config $BUILD_TYPE --target fex_linux_tests_all
- name: FEXLinuxTests Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_FEXLinuxTests.log || true
# Thunking
- name: Thunkgen tests
working-directory: ${{runner.workspace}}/build
shell: bash
run: cmake --build . --config $BUILD_TYPE --target thunkgen_tests
- name: Thunkgen Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_ThunkgenTests.log || true
if: steps.build.outcome == 'success'
uses: ./.github/workflows/test
with:
target: thunkgen_tests
- name: Test GL No-Thunks
if: matrix.arch[1] == 'x64'
working-directory: ${{runner.workspace}}/build
shell: bash
if: ${{ steps.build.outcome == 'success' && matrix.arch[1] == 'x64' }}
uses: ./.github/workflows/test
with:
target: thunk_functional_tests_nothunks
env:
DISPLAY: ":0"
run: cmake --build . --config $BUILD_TYPE --target thunk_functional_tests_nothunks
- name: No thunks Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_NoThunkResults.log || true
DISPLAY: ':0'
- name: Test GL Thunks
if: matrix.arch[1] == 'x64'
working-directory: ${{runner.workspace}}/build
shell: bash
if: ${{ steps.build.outcome == 'success' && matrix.arch[1] == 'x64' }}
uses: ./.github/workflows/test
with:
target: thunk_functional_tests_thunks
env:
DISPLAY: ":0"
run: cmake --build . --config $BUILD_TYPE --target thunk_functional_tests_thunks
- name: Thunks Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_ThunkResults.log || true
DISPLAY: ':0'
# ASM tests
- name: ASM Tests
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the unit tests
run: cmake --build . --config $BUILD_TYPE --target asm_tests
if: steps.build.outcome == 'success'
uses: ./.github/workflows/test
with:
target: asm_tests
- name: ASM Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_ASM.log || true
# POSIX tests
- name: POSIX Tests
if: steps.build.outcome == 'success'
uses: ./.github/workflows/test
with:
target: posix_tests
- name: Posix Tests
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the posixtest
run: cmake --build . --config $BUILD_TYPE --target posix_tests
- name: Posix Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_Posix.log || true
- name: gvisor tests
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the gvisor tests
run: cmake --build . --config $BUILD_TYPE --target gvisor_tests
- name: GVisor Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_GVisor.log || true
# GVisor tests
- name: GVisor Tests
if: steps.build.outcome == 'success'
uses: ./.github/workflows/test
with:
target: gvisor_tests
# Struct verifier tests
- name: Struct verifier tests
working-directory: ${{runner.workspace}}/build
shell: bash
run: cmake --build . --config $BUILD_TYPE --target struct_verifier
- name: Struct verifier Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_StructVerifier.log || true
- name: Truncate test results
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
# Cap out the log files at 20M in case something crash spins and dumps fault text
# ASM tests get quite close to 10MB
run: truncate --size="<20M" ${{runner.workspace}}/build/Testing/Temporary/LastTest_*.log || true
if: steps.build.outcome == 'success'
uses: ./.github/workflows/test
with:
target: struct_verifier
- name: Remove old SHM regions
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: cmake --build . --config $BUILD_TYPE --target remove_old_shm_regions
- name: Set runner name
if: ${{ always() }}
run: echo "runner_name=$(hostname)" >> $GITHUB_ENV
run: cmake --build build --target remove_old_shm_regions
- name: Upload results
if: ${{ always() }}
uses: 'actions/upload-artifact@v4'
uses: actions/upload-artifact@v6
timeout-minutes: 1
with:
name: Results-${{ env.runner_name }}
path: ${{runner.workspace}}/build/Testing/Temporary/LastTest_*.log
name: Results-${{ env.runner_name }}-${{ env.runner_label }}
path: results/*.log
retention-days: 3
+57 -126
View File
@@ -31,163 +31,94 @@ jobs:
fail-fast: false
steps:
- uses: actions/checkout@v3
- uses: actions/checkout@v6
with:
fetch-depth: '0'
fetch-tags: 'true'
- name: Set runner label
run: echo "runner_label=${{ matrix.arch[1] }}" >> $GITHUB_ENV
- name: Set rootfs paths
- name: Set runner info
run: |
echo "FEX_ROOTFS_MOUNT=/mnt/AutoNFS/rootfs/" >> $GITHUB_ENV
echo "FEX_ROOTFS_PATH=$HOME/Rootfs/" >> $GITHUB_ENV
echo "FEX_ROOTFS=$HOME/Rootfs/" >> $GITHUB_ENV
echo "runner_label=${{ matrix.arch[1] }}" >> $GITHUB_ENV
echo "runner_name=$(hostname)" >> $GITHUB_ENV
- name: Update RootFS cache
# Use a bash shell so we can use the same syntax for environment variable
# access regardless of the host operating system
shell: bash
run: $GITHUB_WORKSPACE/Scripts/CI_FetchRootFS.py
- name : submodule checkout
# Need to update submodules
run: |
git submodule sync --recursive
git submodule update --init --depth 1
- name: Clean Build Environment
run: rm -Rf ${{runner.workspace}}/build
- name: Create Build Environment
# Some projects don't allow in-source building, so create a separate build directory
# We'll use this as our working directory for all subsequent commands
run: cmake -E make_directory ${{runner.workspace}}/build
- name: Setup Build Environment
uses: ./.github/workflows/setup-env
- name: Configure CMake
# Use a bash shell so we can use the same syntax for environment variable
# access regardless of the host operating system
shell: bash
working-directory: ${{runner.workspace}}/build
# Note the current convention is to use the -S and -B options here to specify source
# and build directories, but this is only available with CMake 3.13 and higher.
# The CMake binaries on the Github Actions machines are (as of this writing) 3.12
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True -DBUILD_FEX_LINUX_TESTS=True -DENABLE_GLIBC_ALLOCATOR_HOOK_FAULT=True -DENABLE_JEMALLOC_GLIBC_ALLOC=False -DCMAKE_INSTALL_PREFIX=${{runner.workspace}}/build/install
run: |
cmake -S . -B build -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_LTO=False \
-DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True -DBUILD_FEX_LINUX_TESTS=True \
-DENABLE_GLIBC_ALLOCATOR_HOOK_FAULT=True -DENABLE_JEMALLOC_GLIBC_ALLOC=False \
-DCMAKE_INSTALL_PREFIX="$PWD"/build/install
# These steps make a lot of noise but rarely fail.
# Put them in a separate step to make normal build logs easier to parse
- name: Noisy Build Targets
run: cmake --build build --target asm_files 32bit_asm_files JemallocLibs Catch2 vixl cephes_128bit
- name: Build
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the build. You can specify a specific target with "--target <NAME>"
run: cmake --build . --config $BUILD_TYPE
run: cmake --build build
- name: Install
working-directory: ${{runner.workspace}}/build
shell: bash
run: cmake --build . --config $BUILD_TYPE --target install
run: cmake --build build --target install
- name: gcc target tests 64
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the gvisor tests
run: cmake --build . --config $BUILD_TYPE --target gcc_target_tests_64
- name: GCC64 Test Results move
# GCC tests
- name: GCC64 Target Tests
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_GCC64.log || true
uses: ./.github/workflows/test
with:
target: gcc_target_tests_64
- name: gcc target tests 32
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the gvisor tests
run: cmake --build . --config $BUILD_TYPE --target gcc_target_tests_32
- name: GCC32 Test Results move
- name: GCC32 Target Tests
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_GCC32.log || true
uses: ./.github/workflows/test
with:
target: gcc_target_tests_32
- name: APITest tests
working-directory: ${{runner.workspace}}/build
shell: bash
run: cmake --build . --config $BUILD_TYPE --target api_tests
- name: APITest Test Results move
# API Tests
- name: API Tests
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_APITests.log || true
uses: ./.github/workflows/test
with:
target: api_tests
- name: FEXCore APITest tests
working-directory: ${{runner.workspace}}/build
shell: bash
run: cmake --build . --config $BUILD_TYPE --target fexcore_apitests
- name: FEXCore APITest Test Results move
- name: FEXCore API Tests
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_FEXCoreAPITests.log || true
uses: ./.github/workflows/test
with:
target: fexcore_apitests
- name: FEXLinuxTests
working-directory: ${{runner.workspace}}/build
shell: bash
run: cmake --build . --config $BUILD_TYPE --target fex_linux_tests_all
- name: FEXLinuxTests Results move
# Linux tests
- name: FEX Linux Tests
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_FEXLinuxTests.log || true
uses: ./.github/workflows/test
with:
target: fex_linux_tests_all
# ASM Tests
- name: ASM Tests
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the unit tests
run: cmake --build . --config $BUILD_TYPE --target asm_tests
- name: ASM Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_ASM.log || true
uses: ./.github/workflows/test
with:
target: asm_tests
- name: Posix Tests
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the posixtest
run: cmake --build . --config $BUILD_TYPE --target posix_tests
- name: Posix Test Results move
# POSIX Tests
- name: POSIX Tests
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_Posix.log || true
- name: Truncate test results
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
# Cap out the log files at 20M in case something crash spins and dumps fault text
# ASM tests get quite close to 10MB
run: truncate --size="<20M" ${{runner.workspace}}/build/Testing/Temporary/LastTest_*.log || true
uses: ./.github/workflows/test
with:
target: posix_tests
- name: Remove old SHM regions
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: cmake --build . --config $BUILD_TYPE --target remove_old_shm_regions
- name: Set runner name
if: ${{ always() }}
run: echo "runner_name=$(hostname)" >> $GITHUB_ENV
run: cmake --build build --target remove_old_shm_regions
- name: Upload results
if: ${{ always() }}
uses: 'actions/upload-artifact@v4'
uses: actions/upload-artifact@v6
timeout-minutes: 1
with:
name: Results-${{ env.runner_name }}
path: ${{runner.workspace}}/build/Testing/Temporary/LastTest_*.log
name: Results-${{ env.runner_name }}-${{ env.runner_label }}
path: results/*.log
retention-days: 3
+25 -63
View File
@@ -24,83 +24,45 @@ jobs:
fail-fast: false
steps:
- uses: actions/checkout@v3
- uses: actions/checkout@v6
with:
fetch-depth: '0'
fetch-tags: 'true'
- name: Set runner label
run: echo "runner_label=${{ matrix.arch[1] }}" >> $GITHUB_ENV
- name: Set rootfs paths
- name: Set runner info
run: |
echo "FEX_ROOTFS_MOUNT=/mnt/AutoNFS/rootfs/" >> $GITHUB_ENV
echo "FEX_ROOTFS_PATH=$HOME/Rootfs/" >> $GITHUB_ENV
echo "FEX_ROOTFS=$HOME/Rootfs/" >> $GITHUB_ENV
echo "runner_label=${{ matrix.arch[1] }}" >> $GITHUB_ENV
echo "runner_name=$(hostname)" >> $GITHUB_ENV
- name: Update RootFS cache
# Use a bash shell so we can use the same syntax for environment variable
# access regardless of the host operating system
shell: bash
run: $GITHUB_WORKSPACE/Scripts/CI_FetchRootFS.py
- name : submodule checkout
# Need to update submodules
run: |
git submodule sync --recursive
git submodule update --init --depth 1
- name: Clean Build Environment
run: rm -Rf ${{runner.workspace}}/build
- name: Create Build Environment
# Some projects don't allow in-source building, so create a separate build directory
# We'll use this as our working directory for all subsequent commands
run: cmake -E make_directory ${{runner.workspace}}/build
- name: Setup Build Environment
uses: ./.github/workflows/setup-env
- name: Configure CMake
# Use a bash shell so we can use the same syntax for environment variable
# access regardless of the host operating system
shell: bash
working-directory: ${{runner.workspace}}/build
# Note the current convention is to use the -S and -B options here to specify source
# and build directories, but this is only available with CMake 3.13 and higher.
# The CMake binaries on the Github Actions machines are (as of this writing) 3.12
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True
run: |
cmake -S . -B build -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_LTO=False \
-DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True
# These steps make a lot of noise but rarely fail.
# Put them in a separate step to make normal build logs easier to parse
- name: Noisy Build Targets
run: cmake --build build --target asm_files 32bit_asm_files JemallocLibs Catch2 vixl cephes_128bit
- name: Build
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the build. You can specify a specific target with "--target <NAME>"
run: cmake --build . --config $BUILD_TYPE
run: cmake --build build
# ASM tests
- name: ASM Tests
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the unit tests
run: cmake --build . --config $BUILD_TYPE --target asm_tests
- name: ASM Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_ASM.log || true
- name: Truncate test results
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
# Cap out the log files at 20M in case something crash spins and dumps fault text
# ASM tests get quite close to 10MB
run: truncate --size="<20M" ${{runner.workspace}}/build/Testing/Temporary/LastTest_*.log || true
- name: Set runner name
if: ${{ always() }}
run: echo "runner_name=$(hostname)" >> $GITHUB_ENV
uses: ./.github/workflows/test
with:
target: asm_tests
- name: Upload results
if: ${{ always() }}
uses: 'actions/upload-artifact@v4'
uses: actions/upload-artifact@v6
timeout-minutes: 1
with:
name: Results-${{ env.runner_name }}
path: ${{runner.workspace}}/build/Testing/Temporary/LastTest_*.log
name: Results-${{ env.runner_name }}-${{ env.runner_label }}
path: results/*.log
retention-days: 3
+28 -95
View File
@@ -23,123 +23,56 @@ jobs:
fail-fast: false
steps:
- uses: actions/checkout@v3
- uses: actions/checkout@v6
with:
fetch-depth: '0'
fetch-tags: 'true'
- name: Set runner label
run: echo "runner_label=${{ matrix.arch[1] }}" >> $GITHUB_ENV
- name: Set rootfs paths
- name: Set runner info
run: |
echo "FEX_ROOTFS_MOUNT=/mnt/AutoNFS/rootfs/" >> $GITHUB_ENV
echo "FEX_ROOTFS_PATH=$HOME/Rootfs/" >> $GITHUB_ENV
echo "FEX_ROOTFS=$HOME/Rootfs/" >> $GITHUB_ENV
echo "runner_label=${{ matrix.arch[1] }}" >> $GITHUB_ENV
echo "runner_name=$(hostname)" >> $GITHUB_ENV
- name: Update RootFS cache
# Use a bash shell so we can use the same syntax for environment variable
# access regardless of the host operating system
shell: bash
run: $GITHUB_WORKSPACE/Scripts/CI_FetchRootFS.py
- name: Setup Build Environment
uses: ./.github/workflows/setup-env
- name : submodule checkout
# Need to update submodules
- name: Set VIXL_SIM_ENABLED
run: |
git submodule sync --recursive
git submodule update --init --depth 1
- name: Clean Build Environment
run: rm -Rf ${{runner.workspace}}/build
- name: Create Build Environment
# Some projects don't allow in-source building, so create a separate build directory
# We'll use this as our working directory for all subsequent commands
run: cmake -E make_directory ${{runner.workspace}}/build
- name: Set vixl_sim x86
if: matrix.arch[1] == 'x64'
run: |
echo "VIXL_SIM_ENABLED=True" >> $GITHUB_ENV
- name: Set vixl_sim Arm64
if: matrix.arch[1] == 'ARM64'
run: |
echo "VIXL_SIM_ENABLED=False" >> $GITHUB_ENV
case '${{ matrix.arch[1] }}' in
x64) _sim=True ;;
ARM64) _sim=False ;;
esac
echo "VIXL_SIM_ENABLED=$_sim" >> $GITHUB_ENV
- name: Configure CMake
# Use a bash shell so we can use the same syntax for environment variable
# access regardless of the host operating system
shell: bash
working-directory: ${{runner.workspace}}/build
# Note the current convention is to use the -S and -B options here to specify source
# and build directories, but this is only available with CMake 3.13 and higher.
# The CMake binaries on the Github Actions machines are (as of this writing) 3.12
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_VIXL_SIMULATOR=$VIXL_SIM_ENABLED -DENABLE_VIXL_DISASSEMBLER=True -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True
run: |
cmake -S . -B build -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_VIXL_SIMULATOR=$VIXL_SIM_ENABLED \
-DENABLE_VIXL_DISASSEMBLER=True -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True
- name: Build
working-directory: ${{runner.workspace}}/build
shell: bash
env:
FEX_DISABLETELEMETRY: 1
# Execute the build. You can specify a specific target with "--target <NAME>"
run: cmake --build . --config $BUILD_TYPE --target CodeSizeValidation instcountci_test_files
run: cmake --build build --target CodeSizeValidation instcountci_test_files
- name: Instruction Count Tests
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the unit tests
run: cmake --build . --config $BUILD_TYPE --target instcountci_tests
- name: Instruction Count Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_InstCountCI.log || true
uses: ./.github/workflows/test
with:
target: instcountci_tests
- name: Update local repo instcount
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: cmake --build . --config $BUILD_TYPE --target instcountci_update_tests
run: cmake --build build --target instcountci_update_tests
- name: Get instcountCI diff
- name: Check InstCountCI diff
if: ${{ always() }}
shell: bash
working-directory: ${{github.workspace}}/
run: git diff --output=${{runner.workspace}}/build/InstCountCI.diff
- name: Check if InstCountCI Diff exists
if: ${{ always() }}
shell: bash
working-directory: ${{github.workspace}}/
# Check if the file is empty
run: sh -c "! test -s ${{runner.workspace}}/build/InstCountCI.diff"
- name: Truncate test results
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
# Cap out the log files at 20M in case something crash spins and dumps fault text
# ASM tests get quite close to 10MB
run: truncate --size="<20M" ${{runner.workspace}}/build/Testing/Temporary/LastTest_*.log || true
- name: Set runner name
if: ${{ always() }}
run: echo "runner_name=$(hostname)" >> $GITHUB_ENV
run: git --no-pager diff --exit-code HEAD
- name: Upload results
if: ${{ always() }}
uses: 'actions/upload-artifact@v4'
uses: actions/upload-artifact@v6
timeout-minutes: 1
with:
name: Results-${{ env.runner_name }}
path: ${{runner.workspace}}/build/Testing/Temporary/LastTest_*.log
name: Results-${{ env.runner_name }}-${{ env.runner_label }}
path: results/*.log
retention-days: 3
- name: Upload results InstCountCI
if: ${{ always() }}
uses: 'actions/upload-artifact@v4'
timeout-minutes: 1
with:
name: Results-${{ env.runner_name }}-instcountci
path: ${{runner.workspace}}/build/InstCountCI.diff
retention-days: 3
+18 -65
View File
@@ -20,7 +20,10 @@ jobs:
fail-fast: false
steps:
- uses: actions/checkout@v3
- uses: actions/checkout@v6
with:
fetch-depth: '0'
fetch-tags: 'true'
- name: Set runner label
run: echo "runner_label=${{ matrix.arch[1] }}" >> $GITHUB_ENV
@@ -28,73 +31,23 @@ jobs:
- name: Add MingGW to PATH
run: echo "$HOME/llvm-mingw/build/bin/" >> $GITHUB_PATH
- name: Set CC x86
if: matrix.arch[1] == 'x64'
- name: Set CC
run: |
echo "MINGW_TRIPLE=x86_64-w64-mingw32" >> $GITHUB_ENV
case '${{ matrix.arch[1] }}' in
x64) _cpu=x86_64 ;;
ARM64) _cpu=aarch64 ;;
ARM64EC) _cpu=arm64ec ;;
esac
echo "MINGW_TRIPLE=${_cpu}-w64-mingw32" >> $GITHUB_ENV
- name: Set CC Arm64
if: matrix.arch[1] == 'ARM64'
run: |
echo "MINGW_TRIPLE=aarch64-w64-mingw32" >> $GITHUB_ENV
- name: Set CC Arm64EC
if: matrix.arch[1] == 'ARM64EC'
run: |
echo "MINGW_TRIPLE=arm64ec-w64-mingw32" >> $GITHUB_ENV
- name: Set rootfs paths
run: |
echo "FEX_ROOTFS_MOUNT=/mnt/AutoNFS/rootfs/" >> $GITHUB_ENV
echo "FEX_ROOTFS_PATH=$HOME/Rootfs/" >> $GITHUB_ENV
echo "FEX_ROOTFS=$HOME/Rootfs/" >> $GITHUB_ENV
- name: Update RootFS cache
# Use a bash shell so we can use the same syntax for environment variable
# access regardless of the host operating system
shell: bash
run: $GITHUB_WORKSPACE/Scripts/CI_FetchRootFS.py
- name : submodule checkout
# Need to update submodules
run: |
git submodule sync --recursive
git submodule update --init --depth 1
- name: Clean Build Environment
run: rm -Rf ${{runner.workspace}}/build
- name: Create Build Environment
# Some projects don't allow in-source building, so create a separate build directory
# We'll use this as our working directory for all subsequent commands
run: cmake -E make_directory ${{runner.workspace}}/build
- name: Setup Build Environment
uses: ./.github/workflows/setup-env
- name: Configure CMake
# Use a bash shell so we can use the same syntax for environment variable
# access regardless of the host operating system
shell: bash
working-directory: ${{runner.workspace}}/build
# Note the current convention is to use the -S and -B options here to specify source
# and build directories, but this is only available with CMake 3.13 and higher.
# The CMake binaries on the Github Actions machines are (as of this writing) 3.12
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -DCMAKE_TOOLCHAIN_FILE=$GITHUB_WORKSPACE/Data/CMake/toolchain_mingw.cmake -DMINGW_TRIPLE=$MINGW_TRIPLE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True -DBUILD_TESTING=False -DCMAKE_INSTALL_PREFIX=${{runner.workspace}}/build/install
run: |
cmake -S . -B build -DCMAKE_BUILD_TYPE=$BUILD_TYPE -DCMAKE_TOOLCHAIN_FILE=$GITHUB_WORKSPACE/Data/CMake/toolchain_mingw.cmake \
-DMINGW_TRIPLE=$MINGW_TRIPLE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True -DBUILD_TESTING=False \
-DCMAKE_INSTALL_PREFIX="$PWD"/build/install
- name: Build
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the build. You can specify a specific target with "--target <NAME>"
run: cmake --build . --config $BUILD_TYPE
- name: Set runner name
if: ${{ always() }}
run: echo "runner_name=$(hostname)" >> $GITHUB_ENV
- name: Upload results
if: ${{ always() }}
uses: 'actions/upload-artifact@v4'
timeout-minutes: 1
with:
name: Results-${{ env.runner_name }}
path: ${{runner.workspace}}/build/Testing/Temporary/LastTest_*.log
retention-days: 3
run: cmake --build build
+17 -23
View File
@@ -1,7 +1,7 @@
# Inspired by LLVM's pr-code-format.yml at
# Inspired by LLVM's pr-code-format.yml at
# https://github.com/llvm/llvm-project/blob/main/.github/workflows/pr-code-format.yml
name: "Check code formatting"
name: Check code formatting
on:
pull_request:
branches:
@@ -13,7 +13,7 @@ jobs:
if: github.repository == 'FEX-Emu/FEX'
steps:
- name: Fetch FEX sources
- name: Checkout
uses: actions/checkout@v4
with:
ref: ${{ github.event.pull_request.head.sha }}
@@ -27,18 +27,13 @@ jobs:
deepen_length: 500
- name: Get changed files
id: changed-files
uses: step-security/changed-files@3dbe17c78367e7d60f00d78ae6781a35be47b4a1 # v45.0.1
with:
separator: ","
skip_initial_fetch: true
- name: "Listed files"
env:
CHANGED_FILES: ${{ steps.changed-files.outputs.all_changed_files }}
run: |
echo "Formatting files:"
echo "$CHANGED_FILES"
BASE=$(git merge-base main HEAD)
FILES=$(git diff --name-only "$BASE" | tr '\n' ',' | sed 's/,$//')
echo "CHANGED_FILES=$FILES" >> $GITHUB_ENV
echo "Changed files:"
echo "$FILES"
- name: Check git-clang-format-19 exists
run: which git-clang-format-19
@@ -46,24 +41,23 @@ jobs:
- name: Setup Python env
uses: actions/setup-python@v4
with:
python-version: '3.11'
cache: 'pip'
cache-dependency-path: './External/code-format-helper/requirements_formatting.txt'
python-version: 3.11
cache: pip
cache-dependency-path: ./External/code-format-helper/requirements_formatting.txt
- name: Install python dependencies
run: pip install -r ./External/code-format-helper/requirements_formatting.txt
- name: Run code formatter
env:
CLANG_FORMAT_PATH: 'git-clang-format-19'
CLANG_FORMAT_PATH: git-clang-format-19
GITHUB_PR_NUMBER: ${{ github.event.pull_request.number }}
START_REV: ${{ github.event.pull_request.base.sha }}
END_REV: ${{ github.event.pull_request.head.sha }}
CHANGED_FILES: ${{ steps.changed-files.outputs.all_changed_files }}
run: |
python ./External/code-format-helper/code-format-helper.py \
--repo "FEX-emu/FEX" \
--issue-number $GITHUB_PR_NUMBER \
--start-rev $START_REV \
--end-rev $END_REV \
--repo "FEX-Emu/FEX" \
--issue-number "$GITHUB_PR_NUMBER" \
--start-rev "$START_REV" \
--end-rev "$END_REV" \
--changed-files "$CHANGED_FILES"
+33
View File
@@ -0,0 +1,33 @@
name: Setup Build Environment
description: Setup RootFS and build environment
inputs:
setup-rootfs:
description: 'Whether or not to set up the rootfs'
default: true
runs:
using: composite
steps:
- name: Set rootfs paths
if: ${{ inputs.setup-rootfs == 'true' }}
shell: bash
run: |
echo "FEX_ROOTFS_MOUNT=/mnt/AutoNFS/rootfs/" >> $GITHUB_ENV
echo "FEX_ROOTFS_PATH=$HOME/Rootfs/" >> $GITHUB_ENV
echo "FEX_ROOTFS=$HOME/Rootfs/" >> $GITHUB_ENV
- name: Update RootFS cache
if: ${{ inputs.setup-rootfs == 'true' }}
shell: bash
run: python3 Scripts/CI_FetchRootFS.py
- name: Checkout Submodules
shell: bash
run: |
git submodule sync --recursive
git submodule update --init --depth 1
- name: Clean Build Environment
shell: bash
run: rm -Rf build
+72
View File
@@ -0,0 +1,72 @@
name: steamrt4 build
on:
push:
branches:
- main
pull_request:
branches:
- main
env:
DEBIAN_FRONTEND: noninteractive
BUILD_TYPE: Release
CC: clang
CXX: clang++
jobs:
steamrt4_build:
runs-on: ${{ matrix.arch }}
strategy:
matrix:
arch: [[self-hosted, ARM64, distrobox]]
fail-fast: false
steps:
- uses: actions/checkout@v6
with:
fetch-depth: '0'
fetch-tags: 'true'
- name: Set runner label
run: echo "runner_label=${{ matrix.arch[1] }}" >> $GITHUB_ENV
- name: Setup Build Environment
uses: ./.github/workflows/setup-env
with:
setup-rootfs: false
# Setup everything required.
- name : distrobox setup
run: |
distrobox create -Y -i registry.gitlab.steamos.cloud/steamrt/steamrt4/sdk/arm64:4.0.20251117.183306 steamrt4 || true
distrobox upgrade steamrt4
distrobox enter --name steamrt4 -- sudo apt-get install -y \
git cmake ninja-build ccache \
lld clang \
libclang-dev llvm-dev \
libstdc++-14-dev-i386-cross libgcc-14-dev-i386-cross \
libstdc++-14-dev-amd64-cross libgcc-14-dev-amd64-cross
- name: Configure CMake
run: |
distrobox enter --name steamrt4 -- cmake -S . -B build -DCMAKE_BUILD_TYPE=$BUILD_TYPE \
-G Ninja -DBUILD_STEAM_SUPPORT=True -DENABLE_LTO=True -DENABLE_ASSERTIONS=False -DBUILD_THUNKS=True \
-DBUILD_FEXCONFIG=False -DBUILD_TESTING=False -DENABLE_CLANG_THUNKS=True -DUSE_LINKER=lld \
-DCMAKE_INSTALL_PREFIX=/usr
- name: Build
run: distrobox enter --name steamrt4 -- cmake --build build
- name: install
run: DESTDIR="$PWD"/install distrobox enter --name steamrt4 -- cmake --build build -t install
- name: Upload libraries
uses: actions/upload-artifact@v6
timeout-minutes: 1
with:
overwrite: true
name: steamrt4_steampipe_depot
path: ${{ github.workspace }}/install/*
retention-days: 60
compression-level: 9
+21
View File
@@ -0,0 +1,21 @@
name: Run Test and Store Logs
description: Run a test and store the log.
inputs:
target:
description: 'The test target to run'
required: true
runs:
using: composite
steps:
- name: Run Tests
shell: bash
run: cmake --build build --target ${{ inputs.target }}
- name: Move and Truncate Results
if: ${{ always() }}
shell: bash
run: |
mkdir -p results
mv build/Testing/Temporary/LastTest.log results/${{ inputs.target }}.log || true
truncate --size="<20M" results/${{ inputs.target }}.log || true
+32 -84
View File
@@ -25,111 +25,59 @@ jobs:
fail-fast: false
steps:
- uses: actions/checkout@v3
- uses: actions/checkout@v6
with:
fetch-depth: '0'
fetch-tags: 'true'
- name: Set runner label
run: echo "runner_label=${{ matrix.arch[1] }}" >> $GITHUB_ENV
- name: Set rootfs paths
- name: Set runner info
run: |
echo "FEX_ROOTFS_MOUNT=/mnt/AutoNFS/rootfs/" >> $GITHUB_ENV
echo "FEX_ROOTFS_PATH=$HOME/Rootfs/" >> $GITHUB_ENV
echo "FEX_ROOTFS=$HOME/Rootfs/" >> $GITHUB_ENV
echo "runner_label=${{ matrix.arch[1] }}" >> $GITHUB_ENV
echo "runner_name=$(hostname)" >> $GITHUB_ENV
- name: Update RootFS cache
# Use a bash shell so we can use the same syntax for environment variable
# access regardless of the host operating system
shell: bash
run: $GITHUB_WORKSPACE/Scripts/CI_FetchRootFS.py
- name : submodule checkout
# Need to update submodules
run: |
git submodule sync --recursive
git submodule update --init --depth 1
- name: Clean Build Environment
run: rm -Rf ${{runner.workspace}}/build
- name: Create Build Environment
# Some projects don't allow in-source building, so create a separate build directory
# We'll use this as our working directory for all subsequent commands
run: cmake -E make_directory ${{runner.workspace}}/build
- name: Setup Build Environment
uses: ./.github/workflows/setup-env
- name: Configure CMake
# Use a bash shell so we can use the same syntax for environment variable
# access regardless of the host operating system
shell: bash
working-directory: ${{runner.workspace}}/build
# Note the current convention is to use the -S and -B options here to specify source
# and build directories, but this is only available with CMake 3.13 and higher.
# The CMake binaries on the Github Actions machines are (as of this writing) 3.12
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_VIXL_SIMULATOR=True -DENABLE_VIXL_DISASSEMBLER=True -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True
run: |
cmake -S . -B build -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_VIXL_SIMULATOR=True -DENABLE_LTO=False \
-DENABLE_VIXL_DISASSEMBLER=True -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True
# These steps make a lot of noise but rarely fail.
# Put them in a separate step to make normal build logs easier to parse
- name: Noisy Build Targets
run: cmake --build build --target asm_files 32bit_asm_files JemallocLibs Catch2 vixl cephes_128bit
- name: Build
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the build. You can specify a specific target with "--target <NAME>"
run: cmake --build . --config $BUILD_TYPE
run: cmake --build build
- name: ASM Tests - SVE256
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the unit tests
run: cmake --build . --config $BUILD_TYPE --target asm_tests
- name: ASM Test SVE256 Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_ASM_SVE256Bit.log || true
uses: ./.github/workflows/test
with:
target: asm_tests
- name: ASM Tests - SVE128
working-directory: ${{runner.workspace}}/build
shell: bash
if: ${{ always() }}
uses: ./.github/workflows/test
env:
FEX_FORCESVEWIDTH: "128"
# Execute the unit tests
run: cmake --build . --config $BUILD_TYPE --target asm_tests
- name: ASM Test 128-bit Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_ASM_SVE128Bit.log || true
with:
target: asm_tests
- name: ASM Tests - ASIMD
working-directory: ${{runner.workspace}}/build
shell: bash
if: ${{ always() }}
uses: ./.github/workflows/test
env:
FEX_HOSTFEATURES: "disablesve"
# Execute the unit tests
run: cmake --build . --config $BUILD_TYPE --target asm_tests
- name: ASM Test ASIMD Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_ASM_ASIMD.log || true
- name: Truncate test results
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
# Cap out the log files at 20M in case something crash spins and dumps fault text
# ASM tests get quite close to 10MB
run: truncate --size="<20M" ${{runner.workspace}}/build/Testing/Temporary/LastTest_*.log || true
- name: Set runner name
if: ${{ always() }}
run: echo "runner_name=$(hostname)" >> $GITHUB_ENV
with:
target: asm_tests
- name: Upload results
if: ${{ always() }}
uses: 'actions/upload-artifact@v4'
uses: actions/upload-artifact@v6
timeout-minutes: 1
with:
name: Results-${{ env.runner_name }}
path: ${{runner.workspace}}/build/Testing/Temporary/LastTest_*.log
name: Results-${{ env.runner_name }}-${{ env.runner_label }}
path: results/*.log
retention-days: 3
+35
View File
@@ -0,0 +1,35 @@
name: Wine DLL Build
description: Build a wow64 or arm64ec Wine DLL
inputs:
target:
description: 'The target (arm64ec or wow64)'
required: true
runs:
using: composite
steps:
- name: Clean Build Environment
shell: bash
run: rm -Rf build_${{ inputs.target }}
- name: Configure CMake
shell: bash
run: |
case "${{ inputs.target }}" in
wow64) _cc=aarch64 ;;
arm64ec) _cc=arm64ec ;;
esac
cmake -S . -B build_${{ inputs.target }} -DCMAKE_BUILD_TYPE=$BUILD_TYPE -DCMAKE_TOOLCHAIN_FILE=Data/CMake/toolchain_mingw.cmake \
-DMINGW_TRIPLE=${_cc}-w64-mingw32 -DCMAKE_INSTALL_LIBDIR=/usr/lib/wine/aarch64-windows -G Ninja \
-DENABLE_LTO=False -DENABLE_ASSERTIONS=False -DENABLE_JEMALLOC_GLIBC_ALLOC=False \
-DBUILD_TESTING=False -DCMAKE_INSTALL_PREFIX=/usr -DTUNE_ARCH=generic -DTUNE_CPU=none
- name: Build
shell: bash
run: cmake --build build_${{ inputs.target }}
- name: Install
shell: bash
run: DESTDIR="$PWD"/install cmake --build build_${{ inputs.target }} -t install
+16 -49
View File
@@ -17,72 +17,39 @@ jobs:
fail-fast: false
steps:
- uses: actions/checkout@v3
- uses: actions/checkout@v6
with:
fetch-depth: '0'
fetch-tags: 'true'
- name: Add MingGW to PATH
run: echo "$HOME/llvm-mingw/build/bin/" >> $GITHUB_PATH
- name : submodule checkout
- name: Checkout Submodules
# Need to update submodules
run: |
git submodule sync --recursive
git submodule update --init --depth 1
- name: Clean install directory
run: |
rm -Rf ${{runner.workspace}}/build_install
mkdir ${{runner.workspace}}/build_install
run: rm -Rf install
- name: Clean Build Environment
run: |
rm -Rf ${{runner.workspace}}/build_arm64ec
rm -Rf ${{runner.workspace}}/build_wow64
- name: Build (wow64)
uses: ./.github/workflows/wine_build
with:
target: wow64
- name: Create Build Environment arm64ec
run: |
cmake -E make_directory ${{runner.workspace}}/build_arm64ec
cmake -E make_directory ${{runner.workspace}}/build_wow64
- name: Configure CMake arm64ec
shell: bash
working-directory: ${{runner.workspace}}/build_arm64ec
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -DCMAKE_TOOLCHAIN_FILE=$GITHUB_WORKSPACE/Data/CMake/toolchain_mingw.cmake -DMINGW_TRIPLE=arm64ec-w64-mingw32 -DCMAKE_INSTALL_LIBDIR=/usr/lib/wine/aarch64-windows -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=False -DENABLE_JEMALLOC_GLIBC_ALLOC=False -DCMAKE_INSTALL_PREFIX=/usr -DBUILD_TESTING=False -DCMAKE_INSTALL_PREFIX=/usr
- name: Configure CMake wow64
shell: bash
working-directory: ${{runner.workspace}}/build_wow64
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -DCMAKE_TOOLCHAIN_FILE=$GITHUB_WORKSPACE/Data/CMake/toolchain_mingw.cmake -DMINGW_TRIPLE=aarch64-w64-mingw32 -DCMAKE_INSTALL_LIBDIR=/usr/lib/wine/aarch64-windows -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=False -DENABLE_JEMALLOC_GLIBC_ALLOC=False -DCMAKE_INSTALL_PREFIX=/usr -DBUILD_TESTING=False -DCMAKE_INSTALL_PREFIX=/usr
- name: Build arm64ec
working-directory: ${{runner.workspace}}/build_arm64ec
shell: bash
run: cmake --build . --config $BUILD_TYPE
- name: install arm64ec
working-directory: ${{runner.workspace}}/build_arm64ec
shell: bash
env:
DESTDIR: ${{runner.workspace}}/build_install
run: cmake --build . --config $BUILD_TYPE -t install
- name: Build wow64
working-directory: ${{runner.workspace}}/build_wow64
shell: bash
run: cmake --build . --config $BUILD_TYPE
- name: install wow64
working-directory: ${{runner.workspace}}/build_wow64
shell: bash
env:
DESTDIR: ${{runner.workspace}}/build_install
run: cmake --build . --config $BUILD_TYPE -t install
- name: Build (arm64ec)
uses: ./.github/workflows/wine_build
with:
target: arm64ec
- name: Upload libraries
uses: 'actions/upload-artifact@v4'
uses: actions/upload-artifact@v6
timeout-minutes: 1
with:
overwrite: true
name: wine_dll_artifacts
path: ${{runner.workspace}}/build_install/usr/lib/wine/aarch64-windows/lib*.dll
path: ${{ github.workspace }}/install/usr/lib/wine/aarch64-windows/lib*.dll
retention-days: 60
compression-level: 9
+2
View File
@@ -11,3 +11,5 @@ out/
.vs/
*.pyc
.cache
.idea/
CMakeLists.txt.user
+71
View File
@@ -0,0 +1,71 @@
spec:
inputs:
PROMOTE_BRANCH:
description: "Branch to promote the build to. Empty means no promotion."
default: "bleeding-edge"
---
workflow:
rules:
- when: always
variables:
PROMOTE_BRANCH: $[[ inputs.PROMOTE_BRANCH ]]
variables:
DEBIAN_FRONTEND: noninteractive
GIT_SUBMODULE_STRATEGY: recursive
GIT_DEPTH: 0
CC: clang
CXX: clang++
build:
stage: build
image: registry.gitlab.steamos.cloud/steamrt/steamrt4/sdk/arm64:4.0.20251117.183306
tags:
- docker
- linux
- arm64
- aarch64
script:
- apt-get -y update
- apt-get install -y
git cmake ninja-build ccache
lld clang
libclang-dev llvm-dev
libstdc++-14-dev-i386-cross libgcc-14-dev-i386-cross
libstdc++-14-dev-amd64-cross libgcc-14-dev-amd64-cross
- cmake -E make_directory build/
- cmake -DCMAKE_BUILD_TYPE=Release -G Ninja -DBUILD_STEAM_SUPPORT=True -DENABLE_LTO=True -DENABLE_ASSERTIONS=False -DBUILD_THUNKS=True -DBUILD_FEXCONFIG=False -DBUILD_TESTING=False -DENABLE_CLANG_THUNKS=True -DUSE_LINKER=lld -DCMAKE_INSTALL_PREFIX=/usr -DTUNE_ARCH=armv8.2-a -DTUNE_CPU=none . -B build/
- cmake --build build/ --config Release
- DESTDIR=$(pwd)/install/ cmake --build build/ --config Release -t install
artifacts:
name: "steamrt artifacts"
untracked: false
paths:
- install/
promote:
stage: deploy
variables:
GIT_STRATEGY: none
image: registry.gitlab.steamos.cloud/steamrt/steamrt4/sdk/arm64:4.0.20251117.183306
tags:
- docker
- linux
- arm64
- aarch64
rules:
- if: '$PROMOTE_BRANCH'
before_script:
- apt-get -y update
- apt-get install -y tmux curl
script:
# comment out to debug: SSH in via GCP, go down the container and attach to the session (with `tmux attach -t debug`)
# - tmux new-session -d -s debug
# - while tmux has-session -t debug 2>/dev/null; do sleep 1; done
# ref controls which fex-depot code runs the pipeline, while VERSION_PARAM controls which fex branch's artifacts that pipeline downloads.
- >
curl --fail --location --request POST --form token=${FEX_DEPOT_TRIGGER_TOKEN} --form ref=master --form "variables[PROMOTE_BRANCH]=${PROMOTE_BRANCH}" --form "variables[VERSION_PARAM]=${CI_COMMIT_REF_NAME}" "${CI_API_V4_URL}/projects/fex%2Ffex-depot/trigger/pipeline"
+10 -7
View File
@@ -17,9 +17,6 @@
shallow = true
path = External/fex-gcc-target-tests-bins
url = https://github.com/FEX-Emu/fex-gcc-target-tests-bins.git
[submodule "External/jemalloc"]
path = External/jemalloc
url = https://github.com/FEX-Emu/jemalloc.git
[submodule "External/fmt"]
path = External/fmt
url = https://github.com/fmtlib/fmt.git
@@ -32,10 +29,6 @@
[submodule "External/Catch2"]
path = External/Catch2
url = https://github.com/catchorg/Catch2.git
[submodule "External/robin-map"]
shallow = true
path = External/robin-map
url = https://github.com/FEX-Emu/robin-map.git
[submodule "External/Vulkan-Headers"]
shallow = true
path = External/Vulkan-Headers
@@ -49,3 +42,13 @@
[submodule "External/range-v3"]
path = External/range-v3
url = https://github.com/ericniebler/range-v3.git
[submodule "External/zydis"]
shallow = true
path = External/zydis
url = https://github.com/zyantific/zydis.git
[submodule "External/unordered_dense"]
path = External/unordered_dense
url = https://github.com/martinus/unordered_dense.git
[submodule "External/rpmalloc"]
path = External/rpmalloc
url = https://github.com/FEX-Emu/rpmalloc.git
+264 -173
View File
@@ -1,46 +1,50 @@
cmake_minimum_required(VERSION 3.14)
project(FEX C CXX ASM)
INCLUDE (CheckIncludeFiles)
CHECK_INCLUDE_FILES ("gdb/jit-reader.h" HAVE_GDB_JIT_READER_H)
include(CheckIncludeFiles)
check_include_files("gdb/jit-reader.h" HAVE_GDB_JIT_READER_H)
option(BUILD_FEX_LINUX_TESTS "Build FEXLinuxTests, requires x86 compiler" FALSE)
option(BUILD_FEX_LINUX_TESTS "Build FEXLinuxTests (requires x86 compiler)" FALSE)
option(BUILD_THUNKS "Build thunks" FALSE)
option(BUILD_FEXCONFIG "Build FEXConfig" TRUE)
option(ENABLE_CLANG_THUNKS "Build thunks with clang" TRUE)
option(ENABLE_IWYU "Enables include what you use program" FALSE)
option(ENABLE_IWYU "Enable the Include What You Use sanitizer" FALSE)
option(ENABLE_LTO "Enable LTO with compilation" TRUE)
option(ENABLE_XRAY "Enable building with LLVM X-Ray" FALSE)
set(USE_LINKER "" CACHE STRING "Allow overriding the linker path directly")
option(ENABLE_UBSAN "Enables Clang UBSAN" FALSE)
option(ENABLE_ASAN "Enables Clang ASAN" FALSE)
option(ENABLE_TSAN "Enables Clang TSAN" FALSE)
option(ENABLE_COVERAGE "Enables Coverage" FALSE)
option(ENABLE_ASSERTIONS "Enables assertions in build" FALSE)
option(ENABLE_GDB_SYMBOLS "Enables GDBSymbols integration support" ${HAVE_GDB_JIT_READER_H})
option(ENABLE_STRICT_WERROR "Enables stricter -Werror for CI" FALSE)
option(ENABLE_WERROR "Enables -Werror" FALSE)
option(ENABLE_JEMALLOC "Enables jemalloc allocator" TRUE)
option(ENABLE_JEMALLOC_GLIBC_ALLOC "Enables jemalloc glibc allocator" TRUE)
option(ENABLE_OFFLINE_TELEMETRY "Enables FEX offline telemetry" TRUE)
option(ENABLE_COMPILE_TIME_TRACE "Enables time trace compile option" FALSE)
option(ENABLE_LIBCXX "Enables LLVM libc++" FALSE)
option(ENABLE_CCACHE "Enables ccache for compile caching" TRUE)
option(ENABLE_VIXL_SIMULATOR "Enable use of VIXL simulator for emulation (only useful for CI testing)" FALSE)
option(ENABLE_VIXL_DISASSEMBLER "Enables debug disassembler output with VIXL" FALSE)
option(USE_LEGACY_BINFMTMISC "Uses legacy method of setting up binfmt_misc" FALSE)
option(ENABLE_FEXCORE_PROFILER "Enables use of the FEXCore timeline profiling capabilities" FALSE)
set (FEXCORE_PROFILER_BACKEND "gpuvis" CACHE STRING "Set which backend to use for the FEXCore profiler (gpuvis, tracy)")
set(USE_LINKER "" CACHE STRING "Path to a custom linker program")
option(ENABLE_UBSAN "Enable the Clang Undefined Behavior Sanitizer" FALSE)
option(ENABLE_ASAN "Enable the Clang Address Sanitizer" FALSE)
option(ENABLE_TSAN "Enable the Clang Thread Sanitizer" FALSE)
option(ENABLE_COVERAGE "Enable Code Coverage" FALSE)
option(ENABLE_ASSERTIONS "Enable debug assertions" FALSE)
option(ENABLE_GDB_SYMBOLS "Enable GDBSymbols integration support" ${HAVE_GDB_JIT_READER_H})
option(ENABLE_STRICT_WERROR "Enable stricter -Werror" FALSE)
option(ENABLE_WERROR "Enable -Werror" FALSE)
option(ENABLE_FEX_ALLOCATOR "Enable allocator for FEX" TRUE)
option(ENABLE_JEMALLOC_GLIBC_ALLOC "Enable jemalloc glibc allocator" TRUE)
option(ENABLE_OFFLINE_TELEMETRY "Enable FEX offline telemetry" TRUE)
option(ENABLE_COMPILE_TIME_TRACE "Enable time trace compile option" FALSE)
option(ENABLE_LIBCXX "Use LLVM's libc++ instead of the GNU libstdc++" FALSE)
option(ENABLE_CCACHE "Enable ccache for build caching" TRUE)
option(ENABLE_VIXL_SIMULATOR "Use the VIXL simulator for emulation (only useful for CI testing)" FALSE)
option(ENABLE_VIXL_DISASSEMBLER "Enable debug disassembler output with VIXL" FALSE)
option(ENABLE_ZYDIS "Enable x86/x86-64 guest disassembler output with Zydis" FALSE)
option(USE_LEGACY_BINFMTMISC "Use legacy method of setting up binfmt_misc" FALSE)
option(ENABLE_FEXCORE_PROFILER "Enable FEXCore's timeline profiling capabilities" FALSE)
set(FEXCORE_PROFILER_BACKEND "gpuvis" CACHE STRING "Set which backend to use for FEXCore's profiler")
set_property(CACHE FEXCORE_PROFILER_BACKEND PROPERTY STRINGS gpuvis tracy)
option(ENABLE_GLIBC_ALLOCATOR_HOOK_FAULT "Enables glibc memory allocation hooking with fault for CI testing")
option(USE_PDB_DEBUGINFO "Builds debug info in PDB format" FALSE)
option(USE_PDB_DEBUGINFO "Build debug info in PDB format" FALSE)
option(BUILD_STEAM_SUPPORT "Enable Steam integration" FALSE)
set(X86_32_TOOLCHAIN_FILE "${CMAKE_CURRENT_SOURCE_DIR}/Data/CMake/toolchain_x86_32.cmake" CACHE FILEPATH "Toolchain file for the (cross-)compiler targeting i686")
set(X86_64_TOOLCHAIN_FILE "${CMAKE_CURRENT_SOURCE_DIR}/Data/CMake/toolchain_x86_64.cmake" CACHE FILEPATH "Toolchain file for the (cross-)compiler targeting x86_64")
set(X86_DEV_ROOTFS "/" CACHE FILEPATH "Path to the sysroot used for cross-compiling for i686 and x86_64")
set(DATA_DIRECTORY "" CACHE PATH "Global data directory (override)")
set(HOSTLIBS_DATA_DIRECTORY "" CACHE PATH "Global data directory (override)")
set (X86_32_TOOLCHAIN_FILE "${CMAKE_CURRENT_SOURCE_DIR}/Data/CMake/toolchain_x86_32.cmake" CACHE FILEPATH "Toolchain file for the (cross-)compiler targeting i686")
set (X86_64_TOOLCHAIN_FILE "${CMAKE_CURRENT_SOURCE_DIR}/Data/CMake/toolchain_x86_64.cmake" CACHE FILEPATH "Toolchain file for the (cross-)compiler targeting x86_64")
set (X86_DEV_ROOTFS "/" CACHE FILEPATH "Path to the sysroot used for cross-compiling for i686 and x86_64")
set (DATA_DIRECTORY "" CACHE PATH "Global data directory (override)")
set (HOSTLIBS_DATA_DIRECTORY "" CACHE PATH "Global data directory (override)")
if (NOT DATA_DIRECTORY)
set (DATA_DIRECTORY "${CMAKE_INSTALL_PREFIX}/share/fex-emu")
set(DATA_DIRECTORY "${CMAKE_INSTALL_PREFIX}/share/fex-emu")
endif()
include(GNUInstallDirs)
@@ -48,43 +52,95 @@ if (NOT HOSTLIBS_DATA_DIRECTORY)
set(HOSTLIBS_DATA_DIRECTORY "${CMAKE_INSTALL_FULL_LIBDIR}/fex-emu")
endif()
string(FIND ${CMAKE_BASE_NAME} mingw CONTAINS_MINGW)
if (NOT CONTAINS_MINGW EQUAL -1)
message (STATUS "Mingw build")
set (MINGW_BUILD TRUE)
set (ENABLE_JEMALLOC TRUE)
set (ENABLE_JEMALLOC_GLIBC_ALLOC FALSE)
## Platform Checks ##
# Only 64-bit Linux and Windows are supported
# NB: SIZEOF_VOID_P is in bytes, not bits
# On 32-bit systems this is set to 4
if (NOT CMAKE_SIZEOF_VOID_P EQUAL 8)
message(FATAL_ERROR "Unsupported pointer size ${CMAKE_SIZEOF_VOID_P}."
" FEX only supports 64-bit (8-byte pointer) systems."
" If you believe this is in error, file an issue.")
elseif (NOT (WIN32 OR CMAKE_SYSTEM_NAME STREQUAL "Linux"))
message(FATAL_ERROR "Unsupported system type ${CMAKE_SYSTEM_NAME}."
" FEX only supports Linux and Windows."
" If you believe this is in error, file an issue.")
endif()
if (NOT MINGW_BUILD)
message (STATUS "Clang version ${CMAKE_CXX_COMPILER_VERSION}")
set (CLANG_MINIMUM_VERSION 13.0)
## Compiler Checks ##
# GCC and MSVC are unsupported
if (CMAKE_CXX_COMPILER_ID STREQUAL "GNU")
message(FATAL_ERROR "FEX doesn't support GCC! Use Clang instead.")
elseif (MSVC)
message(FATAL_ERROR "FEX doesn't support MSVC! Use Clang on MinGW instead.")
elseif (MINGW)
message(STATUS "Building for MinGW")
set(ENABLE_FEX_ALLOCATOR TRUE)
set(ENABLE_JEMALLOC_GLIBC_ALLOC FALSE)
else ()
message(STATUS "Clang version ${CMAKE_CXX_COMPILER_VERSION}")
set(CLANG_MINIMUM_VERSION 13.0)
if (CMAKE_CXX_COMPILER_VERSION VERSION_LESS ${CLANG_MINIMUM_VERSION})
message (FATAL_ERROR "Clang version too old for FEX. Need at least ${CLANG_MINIMUM_VERSION} but has ${CMAKE_CXX_COMPILER_VERSION}")
message(FATAL_ERROR "Clang version too old for FEX. Need at least ${CLANG_MINIMUM_VERSION} but has ${CMAKE_CXX_COMPILER_VERSION}")
endif()
endif()
## Architecture Handling ##
string(TOLOWER ${CMAKE_SYSTEM_PROCESSOR} processor)
if (processor MATCHES "x86|amd64")
option(ENABLE_X86_HOST_DEBUG "Enables compiling on x86_64 host" FALSE)
if (NOT ENABLE_X86_HOST_DEBUG)
message(FATAL_ERROR
" FEX doesn't support compiling for x86-64 hosts!"
" This is /only/ a supported configuration for FEX CI and nothing else!")
else()
message(STATUS "x86_64 debug build")
endif()
set(ARCHITECTURE_x86_64 1)
add_compile_definitions(ARCHITECTURE_x86_64=1)
set(CMAKE_CXX_FLAGS "${CMAKE_CXX_FLAGS} -mcx16")
elseif (processor MATCHES "^aarch64|^arm64|^armv8\.*")
set(ARCHITECTURE_arm64 1)
add_compile_definitions(ARCHITECTURE_arm64=1)
# arm64ec needs to define both arm64 and arm64ec
if (processor MATCHES "^arm64ec")
set(ARCHITECTURE_arm64ec 1)
add_compile_definitions(ARCHITECTURE_arm64ec=1)
endif()
endif()
if (NOT (ARCHITECTURE_arm64 OR ARCHITECTURE_arm64ec OR ARCHITECTURE_x86_64))
message(FATAL_ERROR "Unsupported processor type ${processor}."
" If you believe this is in error, file an issue.")
endif()
if (BUILD_STEAM_SUPPORT)
add_compile_definitions(FEX_STEAM_SUPPORT=1)
endif()
if (ENABLE_FEXCORE_PROFILER)
add_definitions(-DENABLE_FEXCORE_PROFILER=1)
add_compile_definitions(ENABLE_FEXCORE_PROFILER=1)
string(TOUPPER "${FEXCORE_PROFILER_BACKEND}" FEXCORE_PROFILER_BACKEND)
if (FEXCORE_PROFILER_BACKEND STREQUAL "GPUVIS")
add_definitions(-DFEXCORE_PROFILER_BACKEND=1)
add_compile_definitions(FEXCORE_PROFILER_BACKEND=1)
elseif (FEXCORE_PROFILER_BACKEND STREQUAL "TRACY")
add_definitions(-DFEXCORE_PROFILER_BACKEND=2)
add_definitions(-DTRACY_ENABLE=1)
add_compile_definitions(FEXCORE_PROFILER_BACKEND=2)
add_compile_definitions(TRACY_ENABLE=1)
# Required so that Tracy will only start in the selected guest application
add_definitions(-DTRACY_MANUAL_LIFETIME=1)
add_definitions(-DTRACY_DELAYED_INIT=1)
add_compile_definitions(TRACY_MANUAL_LIFETIME=1)
add_compile_definitions(TRACY_DELAYED_INIT=1)
# This interferes with FEX's signal handling
add_definitions(-DTRACY_NO_CRASH_HANDLER=1)
add_compile_definitions(TRACY_NO_CRASH_HANDLER=1)
# Tracy can gather call stack samples in regular intervals, but this
# isn't useful for us since it would usually sample opaque JIT code
add_definitions(-DTRACY_NO_SAMPLING=1)
add_compile_definitions(TRACY_NO_SAMPLING=1)
# This pulls in libbacktrace which allocators in global constructors (before FEX can set up its allocator hooks)
add_definitions(-DTRACY_NO_CALLSTACK=1)
if (MINGW_BUILD)
message(FATAL_ERROR "Tracy profiler not supported")
add_compile_definitions(TRACY_NO_CALLSTACK=1)
if (MINGW)
message(FATAL_ERROR "Tracy profiler not supported on MinGW")
endif()
else()
message(FATAL_ERROR "Unknown FEXCore profiler backend ${FEXCORE_PROFILER_BACKEND}")
@@ -96,7 +152,7 @@ if (ENABLE_JEMALLOC_GLIBC_ALLOC AND ENABLE_GLIBC_ALLOCATOR_HOOK_FAULT)
endif()
if (ENABLE_GLIBC_ALLOCATOR_HOOK_FAULT)
add_definitions(-DGLIBC_ALLOCATOR_FAULT=1)
add_compile_definitions(GLIBC_ALLOCATOR_FAULT=1)
endif()
# uninstall target
@@ -111,9 +167,17 @@ if(NOT TARGET uninstall)
endif()
# These options are meant for package management
set (TUNE_CPU "native" CACHE STRING "Override the CPU the build is tuned for")
set (TUNE_ARCH "generic" CACHE STRING "Override the Arch the build is tuned for")
set (OVERRIDE_VERSION "detect" CACHE STRING "Override the FEX version in the format of <MMYY>{.<REV>}")
set(TUNE_CPU "native" CACHE STRING "Override the CPU the build is tuned for")
set(TUNE_ARCH "generic" CACHE STRING "Override the Arch the build is tuned for")
set(OVERRIDE_VERSION "detect" CACHE STRING "Override the FEX version")
set(OVERRIDE_HASH "detect" CACHE STRING "Override the FEX git hash")
get_property(IS_MULTI_CONFIG GLOBAL PROPERTY GENERATOR_IS_MULTI_CONFIG)
if (NOT IS_MULTI_CONFIG AND NOT CMAKE_BUILD_TYPE)
set(CMAKE_BUILD_TYPE Release
CACHE STRING "Choose the type of build." FORCE)
message(STATUS "No build type set, defaulting to a Release build")
endif()
string(TOUPPER "${CMAKE_BUILD_TYPE}" CMAKE_BUILD_TYPE)
if (CMAKE_BUILD_TYPE MATCHES "DEBUG")
@@ -122,14 +186,15 @@ endif()
if (ENABLE_ASSERTIONS)
message(STATUS "Assertions enabled")
add_definitions(-DASSERTIONS_ENABLED=1)
add_compile_definitions(ASSERTIONS_ENABLED=1)
endif()
if (ENABLE_GDB_SYMBOLS)
message(STATUS "GDBSymbols support enabled")
add_definitions(-DGDB_SYMBOLS_ENABLED=1)
add_compile_definitions(GDB_SYMBOLS_ENABLED=1)
endif()
add_compile_definitions(_LARGEFILE64_SOURCE)
set(CMAKE_CXX_STANDARD 20)
set(CMAKE_EXPORT_COMPILE_COMMANDS ON)
@@ -140,33 +205,7 @@ cmake_policy(SET CMP0083 NEW) # Follow new PIE policy
include(CheckPIESupported)
check_pie_supported()
if (ENABLE_LTO)
set(CMAKE_INTERPROCEDURAL_OPTIMIZATION TRUE)
else()
set(CMAKE_INTERPROCEDURAL_OPTIMIZATION FALSE)
endif()
if (CMAKE_SYSTEM_PROCESSOR MATCHES "x86_64")
option(ENABLE_X86_HOST_DEBUG "Enables compiling on x86_64 host" FALSE)
if (NOT ENABLE_X86_HOST_DEBUG)
message(FATAL_ERROR
" FEX-Emu doesn't support compiling for x86-64 hosts!"
" This is /only/ a supported configuration for FEX CI and nothing else!")
endif()
set(_M_X86_64 1)
add_definitions(-D_M_X86_64=1)
set (CMAKE_CXX_FLAGS "${CMAKE_CXX_FLAGS} -mcx16")
endif()
if (CMAKE_SYSTEM_PROCESSOR MATCHES "^aarch64|^arm64|^armv8\.*")
set(_M_ARM_64 1)
add_definitions(-D_M_ARM_64=1)
endif()
if (CMAKE_SYSTEM_PROCESSOR MATCHES "^arm64ec")
set(_M_ARM_64EC 1)
add_definitions(-D_M_ARM_64EC=1)
endif()
set(CMAKE_INTERPROCEDURAL_OPTIMIZATION ${ENABLE_LTO})
include(CheckCXXSourceCompiles)
set(CMAKE_REQUIRED_FLAGS "-std=c++11 -Wattributes -Werror=attributes")
@@ -182,23 +221,33 @@ check_cxx_source_compiles(
HAS_CLANG_PRESERVE_ALL)
unset(CMAKE_REQUIRED_FLAGS)
if (HAS_CLANG_PRESERVE_ALL)
if (MINGW_BUILD)
if (MINGW)
message(STATUS "Ignoring broken clang::preserve_all support")
set(HAS_CLANG_PRESERVE_ALL FALSE)
else()
message(STATUS "Has clang::preserve_all")
endif()
endif ()
if (_M_ARM_64 AND HAS_CLANG_PRESERVE_ALL)
add_definitions("-DFEX_PRESERVE_ALL_ATTR=__attribute__((preserve_all))" "-DFEX_HAS_PRESERVE_ALL_ATTR=1")
else()
add_definitions("-DFEX_PRESERVE_ALL_ATTR=" "-DFEX_HAS_PRESERVE_ALL_ATTR=0")
endif()
if (ARCHITECTURE_arm64 AND HAS_CLANG_PRESERVE_ALL)
add_compile_definitions("FEX_PRESERVE_ALL_ATTR=__attribute__((preserve_all))" "FEX_HAS_PRESERVE_ALL_ATTR=1")
else()
add_compile_definitions("FEX_PRESERVE_ALL_ATTR=" "FEX_HAS_PRESERVE_ALL_ATTR=0")
endif()
check_cxx_source_compiles(
"
#define _GNU_SOURCE
#include <errno.h>
int main() {
return program_invocation_name == nullptr;
}"
HAS_PROGRAM_INVOCATION_NAME)
add_compile_definitions("HAS_PROGRAM_INVOCATION_NAME=${HAS_PROGRAM_INVOCATION_NAME}")
if (ENABLE_VIXL_SIMULATOR)
# We can run the simulator on both x86-64 or AArch64 hosts
add_definitions(-DVIXL_SIMULATOR=1 -DVIXL_INCLUDE_SIMULATOR_AARCH64=1)
add_compile_definitions(VIXL_SIMULATOR=1 VIXL_INCLUDE_SIMULATOR_AARCH64=1)
endif()
if (ENABLE_CCACHE)
@@ -219,7 +268,7 @@ if (ENABLE_COMPILE_TIME_TRACE)
link_libraries(-ftime-trace)
endif()
set (PTHREAD_LIB pthread)
set(PTHREAD_LIB pthread)
if (USE_LINKER)
message(STATUS "Overriding linker to: ${USE_LINKER}")
@@ -234,7 +283,7 @@ endif()
if (NOT ENABLE_OFFLINE_TELEMETRY)
# Disable FEX offline telemetry entirely if asked
add_definitions(-DFEX_DISABLE_TELEMETRY=1)
add_compile_definitions(FEX_DISABLE_TELEMETRY=1)
endif()
if (ENABLE_UBSAN)
@@ -245,13 +294,13 @@ if (ENABLE_UBSAN)
# that are regularly access unaligned.
# function: syscalls cast function pointers to void (*)(unsigned long...), causing warnings
# related to this access.
add_definitions(-DENABLE_UBSAN=1)
add_compile_definitions(ENABLE_UBSAN=1)
add_compile_options(-fno-omit-frame-pointer -fsanitize=undefined -fno-sanitize=alignment -fno-sanitize=function -fno-sanitize-recover=undefined)
link_libraries(-fno-omit-frame-pointer -fsanitize=undefined -fno-sanitize=alignment -fno-sanitize=function -fno-sanitize-recover=undefined)
endif()
if (ENABLE_ASAN)
add_definitions(-DENABLE_ASAN=1)
add_compile_definitions(ENABLE_ASAN=1)
add_compile_options(-fno-omit-frame-pointer -fsanitize=address -fsanitize-address-use-after-scope)
link_libraries(-fno-omit-frame-pointer -fsanitize=address -fsanitize-address-use-after-scope)
endif()
@@ -271,20 +320,20 @@ if (ENABLE_JEMALLOC_GLIBC_ALLOC)
# Required for thunks to work.
# All host native libraries will use this allocator, while *most* other FEX internal allocations will use the other jemalloc allocator.
add_subdirectory(External/jemalloc_glibc/)
elseif (NOT MINGW_BUILD)
message (STATUS
elseif (NOT MINGW)
message(STATUS
" jemalloc glibc allocator disabled!\n"
" This is not a recommended configuration!\n"
" This will very explicitly break thunk execution!\n"
" Use at your own risk!")
endif()
if (ENABLE_JEMALLOC)
# The jemalloc subproject that all FEXCore fextl objects allocate through.
add_subdirectory(External/jemalloc/)
elseif (NOT MINGW_BUILD)
if (ENABLE_FEX_ALLOCATOR)
# The rpmalloc subproject that all FEXCore fextl objects allocate through.
add_subdirectory(External/rpmalloc/)
elseif (NOT MINGW)
message (STATUS
" jemalloc disabled!\n"
" FEX allocator is disabled!\n"
" This is not a recommended configuration!\n"
" This will very explicitly break 32-bit application execution!\n"
" Use at your own risk!")
@@ -295,45 +344,64 @@ if (USE_PDB_DEBUGINFO)
add_link_options(-g -Wl,--pdb=)
endif()
set (CMAKE_CXX_FLAGS_RELWITHDEBINFO "${CMAKE_CXX_FLAGS_RELWITHDEBINFO} -fno-omit-frame-pointer")
set (CMAKE_LINKER_FLAGS_RELWITHDEBINFO "${CMAKE_LINKER_FLAGS_RELWITHDEBINFO} -fno-omit-frame-pointer")
set(CMAKE_CXX_FLAGS_RELWITHDEBINFO "${CMAKE_CXX_FLAGS_RELWITHDEBINFO} -fno-omit-frame-pointer")
set(CMAKE_LINKER_FLAGS_RELWITHDEBINFO "${CMAKE_LINKER_FLAGS_RELWITHDEBINFO} -fno-omit-frame-pointer")
set (CMAKE_CXX_FLAGS_RELEASE "${CMAKE_CXX_FLAGS_RELEASE} -fomit-frame-pointer")
set (CMAKE_LINKER_FLAGS_RELEASE "${CMAKE_LINKER_FLAGS_RELEASE} -fomit-frame-pointer")
set(CMAKE_CXX_FLAGS_RELEASE "${CMAKE_CXX_FLAGS_RELEASE} -fomit-frame-pointer")
set(CMAKE_LINKER_FLAGS_RELEASE "${CMAKE_LINKER_FLAGS_RELEASE} -fomit-frame-pointer")
include_directories(External/robin-map/include/)
## Modules ##
list(APPEND CMAKE_MODULE_PATH ${CMAKE_SOURCE_DIR}/Data/CMake/)
include(LinkerGC)
## Externals ##
find_package(unordered_dense QUIET CONFIG)
if (NOT unordered_dense_FOUND)
add_subdirectory(External/unordered_dense)
endif()
include(CTest)
if (BUILD_TESTING OR ENABLE_VIXL_DISASSEMBLER OR ENABLE_VIXL_SIMULATOR)
add_subdirectory(External/vixl/)
include_directories(SYSTEM External/vixl/src/)
endif()
if (ENABLE_ZYDIS)
find_package(Zycore 1.5 MODULE QUIET)
find_package(Zydis 4.0 MODULE QUIET)
if (TARGET Zydis::Zydis AND TARGET Zycore::Zycore)
message(STATUS "Using system Zydis")
else()
set(ZYDIS_BUILD_TOOLS OFF CACHE BOOL "" FORCE)
set(ZYDIS_BUILD_EXAMPLES OFF CACHE BOOL "" FORCE)
message(STATUS "Using bundled Zydis")
add_subdirectory(External/zydis/)
endif()
endif()
if (ENABLE_FEXCORE_PROFILER AND FEXCORE_PROFILER_BACKEND STREQUAL "TRACY")
add_subdirectory(External/tracy)
endif()
if (CMAKE_CXX_COMPILER_ID STREQUAL "GNU")
# This means we were attempted to get compiled with GCC
message(FATAL_ERROR "FEX doesn't support getting compiled with GCC!")
endif()
find_package(PkgConfig REQUIRED)
find_package(Python 3.9 REQUIRED COMPONENTS Interpreter)
set(BUILD_SHARED_LIBS OFF)
pkg_search_module(xxhash IMPORTED_TARGET xxhash libxxhash)
if (TARGET PkgConfig::xxhash AND NOT CMAKE_CROSSCOMPILING)
add_library(xxHash::xxhash ALIAS PkgConfig::xxhash)
else()
if (NOT CMAKE_CROSSCOMPILING)
find_package(xxhash MODULE QUIET)
endif()
if (NOT TARGET xxHash::xxhash)
set(XXHASH_BUNDLED_MODE TRUE)
set(XXHASH_BUILD_XXHSUM FALSE)
add_subdirectory(External/xxhash/cmake_unofficial/)
endif()
add_definitions(-Wno-trigraphs)
add_definitions(-DGLOBAL_DATA_DIRECTORY="${DATA_DIRECTORY}/")
add_compile_options(-Wno-trigraphs)
add_compile_definitions(GLOBAL_DATA_DIRECTORY="${DATA_DIRECTORY}/")
if (BUILD_TESTING)
find_package(Catch2 3 QUIET)
@@ -364,7 +432,6 @@ if (NOT range-v3_FOUND)
endif()
add_subdirectory(External/tiny-json/)
include_directories(External/tiny-json/)
include_directories(Source/)
include_directories("${CMAKE_BINARY_DIR}/Source/")
@@ -388,6 +455,11 @@ if(ENUM_ENUM_WARNING)
add_compile_options(-Wno-deprecated-enum-enum-conversion)
endif()
# GCC enables -Wchanges-meaning by default and treats some cases as an error
if(CMAKE_CXX_COMPILER_ID STREQUAL "GNU")
add_compile_options(-Wno-error=changes-meaning)
endif()
if(ENABLE_WERROR OR ENABLE_STRICT_WERROR)
add_compile_options(-Werror)
if (NOT ENABLE_STRICT_WERROR)
@@ -407,7 +479,7 @@ if (NOT TUNE_ARCH STREQUAL "generic")
endif()
if (TUNE_CPU STREQUAL "native")
if(_M_ARM_64)
if(ARCHITECTURE_arm64)
if (CMAKE_CXX_COMPILER_VERSION VERSION_GREATER_EQUAL 999999.0)
# Clang 12.0 fixed the -mcpu=native bug with mixed big.little implementers
# Clang can not currently check for native Apple M1 type in hypervisor. Currently disabled
@@ -448,8 +520,54 @@ elseif (NOT TUNE_CPU STREQUAL "none")
endif()
endif()
set(GIT_DESCRIBE_STRING "FEX-Unknown")
if (OVERRIDE_VERSION STREQUAL "detect")
find_package(Git)
if (GIT_FOUND)
execute_process(
COMMAND ${GIT_EXECUTABLE} describe --abbrev=7
WORKING_DIRECTORY "${CMAKE_SOURCE_DIR}"
OUTPUT_VARIABLE GIT_DESCRIBE_STRING
ERROR_QUIET
OUTPUT_STRIP_TRAILING_WHITESPACE)
endif()
else()
set(GIT_DESCRIBE_STRING "${OVERRIDE_VERSION}")
endif()
set(GIT_HASH "Unknown")
if (OVERRIDE_HASH STREQUAL "detect")
find_package(Git)
if (GIT_FOUND)
execute_process(
COMMAND ${GIT_EXECUTABLE} rev-parse HEAD
WORKING_DIRECTORY "${CMAKE_SOURCE_DIR}"
OUTPUT_VARIABLE GIT_HASH
ERROR_QUIET
OUTPUT_STRIP_TRAILING_WHITESPACE)
endif()
else()
set(GIT_HASH "${OVERRIDE_HASH}")
endif()
message(STATUS "FEX version: ${GIT_DESCRIBE_STRING}")
message(STATUS "FEX commit: ${GIT_HASH}")
# Prepends 0x to every two-character sequence in the hash,
# OR the final character of the hash, to plumb it for C++ usage. e.g.:
# -DOVERRIDE_HASH=123456aa => 0x12, 0x34, 0x56, 0xaa,
# -DOVERRIDE_HASH=12345678a => 0x12, 0x34, 0x56, 0x78, 0xa,
string(REGEX
REPLACE "(..|.$)" "0x\\1, "
GIT_HASH_ARRAY "${GIT_HASH}")
if (ENABLE_IWYU)
find_program(IWYU_EXE "iwyu")
find_program(IWYU_EXE
NAMES iwyu include-what-you-use)
if (IWYU_EXE)
message(STATUS "IWYU enabled")
set(CMAKE_CXX_INCLUDE_WHAT_YOU_USE "${IWYU_EXE}")
@@ -461,7 +579,7 @@ add_compile_options(-Wall)
if (BUILD_TESTING)
message(STATUS "Unit tests are enabled")
set (TEST_JOB_COUNT "" CACHE STRING "Override number of parallel jobs to use while running tests")
set(TEST_JOB_COUNT "" CACHE STRING "Override number of parallel jobs to use while running tests")
if (TEST_JOB_COUNT)
message(STATUS "Running tests with ${TEST_JOB_COUNT} jobs")
elseif(CMAKE_VERSION VERSION_LESS "3.29")
@@ -476,13 +594,16 @@ add_subdirectory(FEXHeaderUtils/)
add_subdirectory(CodeEmitter/)
add_subdirectory(FEXCore/)
if (_M_ARM_64 AND NOT MINGW_BUILD)
if (ARCHITECTURE_arm64 AND NOT MINGW AND NOT BUILD_STEAM_SUPPORT)
# Binfmt_misc files must be installed prior to Source/ installs
add_subdirectory(Data/binfmts/)
endif()
add_subdirectory(Source/)
add_subdirectory(Data/AppConfig/)
if (NOT BUILD_STEAM_SUPPORT)
add_subdirectory(Data/AppConfig/)
endif()
# Install the ThunksDB file
file(GLOB CONFIG_SOURCES CONFIGURE_DEPENDS ${CMAKE_CURRENT_SOURCE_DIR}/Data/*.json)
@@ -499,7 +620,7 @@ if (BUILD_TESTING)
endif()
if (BUILD_THUNKS)
set (FEX_PROJECT_SOURCE_DIR ${PROJECT_SOURCE_DIR})
set(FEX_PROJECT_SOURCE_DIR ${PROJECT_SOURCE_DIR})
add_subdirectory(ThunkLibs/Generator)
# Thunk targets for both host libraries and IDE integration
@@ -526,8 +647,7 @@ if (BUILD_THUNKS)
"-DX86_DEV_ROOTFS=${X86_DEV_ROOTFS}"
INSTALL_COMMAND ""
BUILD_ALWAYS ON
DEPENDS thunkgen
)
DEPENDS thunkgen)
ExternalProject_Add(guest-libs-32
PREFIX guest-libs-32
@@ -545,65 +665,36 @@ if (BUILD_THUNKS)
"-DX86_DEV_ROOTFS=${X86_DEV_ROOTFS}"
INSTALL_COMMAND ""
BUILD_ALWAYS ON
DEPENDS thunkgen
)
DEPENDS thunkgen)
install(
CODE "MESSAGE(\"-- Installing: guest-libs\")"
CODE "message(\"-- Installing: guest-libs\")"
CODE "
EXECUTE_PROCESS(COMMAND ${CMAKE_COMMAND} --build . --target install
WORKING_DIRECTORY ${CMAKE_BINARY_DIR}/Guest
)"
execute_process(COMMAND ${CMAKE_COMMAND} --build . --target install
WORKING_DIRECTORY ${CMAKE_BINARY_DIR}/Guest)"
DEPENDS guest-libs
COMPONENT Runtime
)
COMPONENT Runtime)
install(
CODE "MESSAGE(\"-- Installing: guest-libs-32\")"
CODE "message(\"-- Installing: guest-libs-32\")"
CODE "
EXECUTE_PROCESS(COMMAND ${CMAKE_COMMAND} --build . --target install
WORKING_DIRECTORY ${CMAKE_BINARY_DIR}/Guest_32
)"
execute_process(COMMAND ${CMAKE_COMMAND} --build . --target install
WORKING_DIRECTORY ${CMAKE_BINARY_DIR}/Guest_32)"
DEPENDS guest-libs-32
COMPONENT Runtime
)
COMPONENT Runtime)
add_custom_target(uninstall_guest-libs
COMMAND ${CMAKE_COMMAND} "--build" "." "--target" "uninstall"
WORKING_DIRECTORY ${CMAKE_BINARY_DIR}/Guest
)
WORKING_DIRECTORY ${CMAKE_BINARY_DIR}/Guest)
add_custom_target(uninstall_guest-libs-32
COMMAND ${CMAKE_COMMAND} "--build" "." "--target" "uninstall"
WORKING_DIRECTORY ${CMAKE_BINARY_DIR}/Guest_32
)
WORKING_DIRECTORY ${CMAKE_BINARY_DIR}/Guest_32)
add_dependencies(uninstall uninstall_guest-libs)
add_dependencies(uninstall uninstall_guest-libs-32)
endif()
set(FEX_VERSION_MAJOR "0")
set(FEX_VERSION_MINOR "0")
set(FEX_VERSION_PATCH "0")
if (OVERRIDE_VERSION STREQUAL "detect")
find_package(Git)
if (GIT_FOUND)
execute_process(
COMMAND ${GIT_EXECUTABLE} describe --abbrev=0
WORKING_DIRECTORY "${CMAKE_SOURCE_DIR}"
OUTPUT_VARIABLE GIT_DESCRIBE_STRING
RESULT_VARIABLE GIT_ERROR
ERROR_QUIET
OUTPUT_STRIP_TRAILING_WHITESPACE
)
if (NOT ${GIT_ERROR} EQUAL 0)
# Likely built in a way that doesn't have tags
# Setup a version tag that is unknown
set(GIT_DESCRIBE_STRING "FEX-0000")
endif()
endif()
else()
set(GIT_DESCRIBE_STRING "FEX-${OVERRIDE_VERSION}")
if (NOT MINGW AND BUILD_STEAM_SUPPORT)
add_subdirectory(Source/Steam/)
endif()
+1 -1
View File
@@ -129,4 +129,4 @@
"variables": []
}
]
}
}
+19 -11
View File
@@ -38,8 +38,6 @@ public:
[[nodiscard]] BranchEncodeSucceeded adr(ARMEmitter::Register rd, const BackwardLabel* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(IsADRRange(Imm), "Unscaled offset too large");
if (IsADRRange(Imm)) {
constexpr uint32_t Op = 0b0001'0000 << 24;
DataProcessing_PCRel_Imm(Op, rd, Imm);
@@ -73,7 +71,6 @@ public:
[[nodiscard]] BranchEncodeSucceeded adrp(ARMEmitter::Register rd, const BackwardLabel* Label) {
int64_t Imm = reinterpret_cast<int64_t>(Label->Location) - (GetCursorAddress<int64_t>() & ~0xFFFLL);
LOGMAN_THROW_A_FMT(IsADRPRange(Imm) && IsADRPAligned(Imm), "Unscaled offset too large");
if (IsADRPRange(Imm) && IsADRPAligned(Imm)) {
constexpr uint32_t Op = 0b1001'0000 << 24;
@@ -103,16 +100,22 @@ public:
}
[[nodiscard]] BranchEncodeSucceeded LongAddressGen(ARMEmitter::Register rd, const BackwardLabel* Label) {
int64_t Imm = reinterpret_cast<int64_t>(Label->Location) - (GetCursorAddress<int64_t>());
const auto SLocation = reinterpret_cast<int64_t>(Label->Location);
const auto ULocation = std::bit_cast<uint64_t>(SLocation);
const int64_t Imm = SLocation - (GetCursorAddress<int64_t>());
const auto UImm = std::bit_cast<uint64_t>(Imm);
if (IsADRRange(Imm)) {
// If the range is in ADR range then we can just use ADR.
return adr(rd, Label);
} else if (IsADRPRange(Imm)) {
int64_t ADRPImm = (reinterpret_cast<int64_t>(Label->Location) & ~0xFFFLL) - (GetCursorAddress<int64_t>() & ~0xFFFLL);
}
if (IsADRPRange(Imm)) {
const int64_t ADRPImm = (SLocation & ~0xFFFLL) - (GetCursorAddress<int64_t>() & ~0xFFFLL);
// If the range is in the ADRP range then we can use ADRP.
bool NeedsOffset = !IsADRPAligned(reinterpret_cast<uint64_t>(Label->Location));
uint64_t AlignedOffset = reinterpret_cast<uint64_t>(Label->Location) & 0xFFFULL;
const bool NeedsOffset = !IsADRPAligned(ULocation);
const uint64_t AlignedOffset = ULocation & 0xFFFULL;
// First emit ADRP
adrp(rd, ADRPImm >> 12);
@@ -125,14 +128,19 @@ public:
return BranchEncodeSucceeded::Success;
}
// Can't encode.
return BranchEncodeSucceeded::Failure;
// Stinky path, we need to load the address as a sequence of movz+movk+movk
movz(ARMEmitter::Size::i64Bit, rd, (UImm >> 32) & 0xFFFF, 32);
movk(ARMEmitter::Size::i64Bit, rd, (UImm >> 16) & 0xFFFF, 16);
movk(ARMEmitter::Size::i64Bit, rd, UImm & 0xFFFF);
return BranchEncodeSucceeded::Success;
}
[[nodiscard]] BranchEncodeSucceeded LongAddressGen(ARMEmitter::Register rd, ForwardLabel* Label) {
AddLocationToLabel(Label, ForwardLabel::Reference {.Location = GetCursorAddress<uint8_t*>(), .Type = ForwardLabel::InstType::LONG_ADDRESS_GEN});
// Emit a register index and a nop. These will be backpatched.
// Emit a register index and two nops. These will be backpatched.
dc32(rd.Idx());
nop();
nop();
// Forward label doesn't know if it can encode until Bind.
return BranchEncodeSucceeded::Success;
-1
View File
@@ -301,7 +301,6 @@ public:
}
[[nodiscard]] BranchEncodeSucceeded tbnz(ARMEmitter::Register rt, uint32_t Bit, const BackwardLabel* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(Imm >= -32768 && Imm <= 32764 && ((Imm & 0b11) == 0), "Unscaled offset too large");
if (Imm >= -32768 && Imm <= 32764 && ((Imm & 0b11) == 0)) {
constexpr uint32_t Op = 0b0011'0111 << 24;
+1
View File
@@ -53,6 +53,7 @@ public:
if (!CurrentAlignment) {
return;
}
std::memset(CurrentOffset, 0, Size - CurrentAlignment);
CurrentOffset += Size - CurrentAlignment;
}
+24 -18
View File
@@ -311,7 +311,7 @@ class ExtendedMemOperand final {
public:
ExtendedMemOperand(XRegister rn, XRegister rm = XReg::zr, ExtendedType Option = ExtendedType::LSL_64, uint32_t Shift = 0)
: rn {rn}
, MetaType {.ExtendedType {
, MetaType {.Extended {
.Header = {.MemType = TYPE_EXTENDED},
.rm = rm,
.Option = Option,
@@ -340,7 +340,7 @@ public:
Register rm;
ExtendedType Option;
uint32_t Shift;
} ExtendedType;
} Extended;
struct {
HeaderStruct Header;
IndexType Index;
@@ -741,38 +741,44 @@ public:
break;
}
case ForwardLabel::InstType::LONG_ADDRESS_GEN: {
uint32_t* Instructions = reinterpret_cast<uint32_t*>(Label->Location);
int64_t ImmInstOne = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(&Instructions[0]);
int64_t ImmInstTwo = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(&Instructions[1]);
auto OriginalOffset = GetCursorOffset();
const auto* Instructions = reinterpret_cast<uint32_t*>(Label->Location);
const auto ImmInstOne = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(&Instructions[0]);
const auto ImmInstTwo = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(&Instructions[1]);
const auto ImmInstThree = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(&Instructions[2]);
const auto OriginalOffset = GetCursorOffset();
auto InstOffset = GetCursorOffsetFromAddress(Instructions);
const auto InstOffset = GetCursorOffsetFromAddress(Instructions);
SetCursorOffset(InstOffset);
// We encoded the destination register in to the first instruction space.
// Read it back.
ARMEmitter::Register DestReg(Instructions[0]);
if (IsADRRange(ImmInstTwo)) {
// If within ADR range from the second instruction, then we can emit NOP+ADR
if (IsADRRange(ImmInstThree)) {
// If within ADR range from the third instruction, then we can emit NOP+NOP+ADR
nop();
adr(DestReg, static_cast<uint32_t>(ImmInstTwo) & 0x7FFF);
} else if (IsADRPRange(ImmInstOne)) {
nop();
adr(DestReg, static_cast<uint32_t>(ImmInstThree) & 0x7FFF);
} else if (IsADRPRange(ImmInstTwo)) {
// If within ADRP range from the first instruction, then we are /definitely/ in range for the second instruction.
// First check if we are in non-offset range for second instruction.
if (IsADRPAligned(reinterpret_cast<uint64_t>(CurrentAddress))) {
// We can emit nop + adrp
// We can emit nop + nop + adrp
nop();
nop();
adrp(DestReg, static_cast<uint32_t>(ImmInstThree >> 12) & 0x7FFF);
} else {
// Not aligned, need nop + adrp + add
nop();
adrp(DestReg, static_cast<uint32_t>(ImmInstTwo >> 12) & 0x7FFF);
} else {
// Not aligned, need adrp + add
adrp(DestReg, static_cast<uint32_t>(ImmInstOne >> 12) & 0x7FFF);
add(ARMEmitter::Size::i64Bit, DestReg, DestReg, ImmInstOne & 0xFFF);
add(ARMEmitter::Size::i64Bit, DestReg, DestReg, ImmInstTwo & 0xFFF);
}
} else {
LOGMAN_MSG_A_FMT("Unscaled offset is too large");
FEX_UNREACHABLE;
// Stinky path, we need to emit a movz+movk+movk sequence.
movz(ARMEmitter::Size::i64Bit, DestReg, uint32_t(ImmInstOne >> 32) & 0x7FFF, 32);
movk(ARMEmitter::Size::i64Bit, DestReg, uint32_t(ImmInstOne >> 16) & 0xFFFF, 16);
movk(ARMEmitter::Size::i64Bit, DestReg, uint32_t(ImmInstOne) & 0xFFFF);
}
SetCursorOffset(OriginalOffset);
+50 -50
View File
@@ -3627,8 +3627,8 @@ public:
void strb(ARMEmitter::Register rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
strb(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
strb(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
strb(rt, MemSrc.rn);
} else {
@@ -3650,8 +3650,8 @@ public:
}
void ldrb(ARMEmitter::Register rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrb(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrb(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldrb(rt, MemSrc.rn);
} else {
@@ -3673,8 +3673,8 @@ public:
}
void ldrsb(ARMEmitter::XRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrsb(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrsb(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldrsb(rt, MemSrc.rn);
} else {
@@ -3696,8 +3696,8 @@ public:
}
void ldrsb(ARMEmitter::WRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrsb(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrsb(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldrsb(rt, MemSrc.rn);
} else {
@@ -3719,8 +3719,8 @@ public:
}
void strh(ARMEmitter::Register rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
strh(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
strh(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
strh(rt, MemSrc.rn);
} else {
@@ -3742,8 +3742,8 @@ public:
}
void ldrh(ARMEmitter::Register rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrh(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrh(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldrh(rt, MemSrc.rn);
} else {
@@ -3765,8 +3765,8 @@ public:
}
void ldrsh(ARMEmitter::XRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrsh(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrsh(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldrsh(rt, MemSrc.rn);
} else {
@@ -3788,8 +3788,8 @@ public:
}
void ldrsh(ARMEmitter::WRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrsh(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrsh(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldrsh(rt, MemSrc.rn);
} else {
@@ -3811,8 +3811,8 @@ public:
}
void str(ARMEmitter::WRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
str(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
str(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
str(rt, MemSrc.rn);
} else {
@@ -3834,8 +3834,8 @@ public:
}
void ldr(ARMEmitter::WRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldr(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldr(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldr(rt, MemSrc.rn);
} else {
@@ -3857,8 +3857,8 @@ public:
}
void ldrsw(ARMEmitter::XRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrsw(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrsw(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldrsw(rt, MemSrc.rn);
} else {
@@ -3880,8 +3880,8 @@ public:
}
void str(ARMEmitter::XRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
str(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
str(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
str(rt, MemSrc.rn);
} else {
@@ -3903,8 +3903,8 @@ public:
}
void ldr(ARMEmitter::XRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldr(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldr(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldr(rt, MemSrc.rn);
} else {
@@ -3926,8 +3926,8 @@ public:
}
void prfm(ARMEmitter::Prefetch prfop, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
prfm(prfop, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
prfm(prfop, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
prfm(prfop, MemSrc.rn);
} else {
@@ -3946,9 +3946,9 @@ public:
void strb(ARMEmitter::VRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
LOGMAN_THROW_A_FMT(MemSrc.MetaType.ExtendedType.Shift == false, "Can't shift byte");
strb(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
LOGMAN_THROW_A_FMT(MemSrc.MetaType.Extended.Shift == false, "Can't shift byte");
strb(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
strb(rt, MemSrc.rn);
} else {
@@ -3970,9 +3970,9 @@ public:
}
void ldrb(ARMEmitter::VRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
LOGMAN_THROW_A_FMT(MemSrc.MetaType.ExtendedType.Shift == false, "Can't shift byte");
ldrb(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
LOGMAN_THROW_A_FMT(MemSrc.MetaType.Extended.Shift == false, "Can't shift byte");
ldrb(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldrb(rt, MemSrc.rn);
} else {
@@ -3994,8 +3994,8 @@ public:
}
void strh(ARMEmitter::VRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
strh(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
strh(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
strh(rt, MemSrc.rn);
} else {
@@ -4017,8 +4017,8 @@ public:
}
void ldrh(ARMEmitter::VRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrh(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldrh(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldrh(rt, MemSrc.rn);
} else {
@@ -4040,8 +4040,8 @@ public:
}
void str(ARMEmitter::SRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
str(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
str(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
str(rt, MemSrc.rn);
} else {
@@ -4063,8 +4063,8 @@ public:
}
void ldr(ARMEmitter::SRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldr(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldr(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldr(rt, MemSrc.rn);
} else {
@@ -4086,8 +4086,8 @@ public:
}
void str(ARMEmitter::DRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
str(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
str(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
str(rt, MemSrc.rn);
} else {
@@ -4109,8 +4109,8 @@ public:
}
void ldr(ARMEmitter::DRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldr(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldr(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldr(rt, MemSrc.rn);
} else {
@@ -4132,8 +4132,8 @@ public:
}
void str(ARMEmitter::QRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
str(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
str(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
str(rt, MemSrc.rn);
} else {
@@ -4155,8 +4155,8 @@ public:
}
void ldr(ARMEmitter::QRegister rt, ARMEmitter::ExtendedMemOperand MemSrc) {
if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED &&
MemSrc.MetaType.ExtendedType.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldr(rt, MemSrc.rn, MemSrc.MetaType.ExtendedType.rm, MemSrc.MetaType.ExtendedType.Option, MemSrc.MetaType.ExtendedType.Shift);
MemSrc.MetaType.Extended.rm.Idx() != ARMEmitter::Reg::r31.Idx()) {
ldr(rt, MemSrc.rn, MemSrc.MetaType.Extended.rm, MemSrc.MetaType.Extended.Option, MemSrc.MetaType.Extended.Shift);
} else if (MemSrc.MetaType.Header.MemType == ARMEmitter::ExtendedMemOperand::Type::TYPE_EXTENDED) {
ldr(rt, MemSrc.rn);
} else {
+2 -5
View File
@@ -15,13 +15,10 @@ foreach(GEN_CONFIG_SRC ${GEN_CONFIG_SOURCES})
get_filename_component(CONFIG_NAME ${GEN_CONFIG_SRC} NAME_WLE)
# Configure it
configure_file(
${GEN_CONFIG_SRC}
${CMAKE_BINARY_DIR}/Data/AppConfig/${CONFIG_NAME})
configure_file(${GEN_CONFIG_SRC} ${CMAKE_BINARY_DIR}/Data/AppConfig/${CONFIG_NAME})
# Then install the configured json
install(
FILES ${CMAKE_BINARY_DIR}/Data/AppConfig/${CONFIG_NAME}
install(FILES ${CMAKE_BINARY_DIR}/Data/AppConfig/${CONFIG_NAME}
DESTINATION ${DATA_DIRECTORY}/AppConfig/
COMPONENT Runtime)
endforeach()
+23
View File
@@ -0,0 +1,23 @@
# SPDX-License-Identifier: MIT
if (CMAKE_CROSSCOMPILING)
return()
endif()
include(FindPackageHandleStandardArgs)
find_package(Zycore QUIET CONFIG)
if (Zycore_CONSIDERED_CONFIGS)
find_package_handle_standard_args(Zycore CONFIG_MODE)
else()
find_package(PkgConfig QUIET)
pkg_search_module(Zycore QUIET IMPORTED_TARGET zycore)
find_package_handle_standard_args(Zycore
REQUIRED_VARS zycore_LINK_LIBRARIES
VERSION_VAR zycore_VERSION)
if (TARGET PkgConfig::zycore)
add_library(Zycore::Zycore ALIAS PkgConfig::zycore)
endif()
endif()
+23
View File
@@ -0,0 +1,23 @@
# SPDX-License-Identifier: MIT
if (CMAKE_CROSSCOMPILING)
return()
endif()
include(FindPackageHandleStandardArgs)
find_package(Zydis QUIET CONFIG)
if (Zydis_CONSIDERED_CONFIGS)
find_package_handle_standard_args(Zydis CONFIG_MODE)
else()
find_package(PkgConfig QUIET)
pkg_search_module(Zydis QUIET IMPORTED_TARGET zydis)
find_package_handle_standard_args(Zydis
REQUIRED_VARS zydis_LINK_LIBRARIES
VERSION_VAR zydis_VERSION)
if (TARGET PkgConfig::zydis)
add_library(Zydis::Zydis ALIAS PkgConfig::zydis)
endif()
endif()
+18
View File
@@ -0,0 +1,18 @@
# SPDX-License-Identifier: MIT
include(FindPackageHandleStandardArgs)
find_package(PkgConfig QUIET)
pkg_search_module(xxhash QUIET IMPORTED_TARGET xxhash libxxhash)
find_package_handle_standard_args(xxhash
REQUIRED_VARS xxhash_LINK_LIBRARIES
VERSION_VAR xxhash_VERSION
)
if (xxhash_FOUND AND NOT TARGET xxHash::xxhash)
if (TARGET PkgConfig::xxhash)
add_library(xxHash::xxhash ALIAS PkgConfig::xxhash)
else()
add_library(xxHash::xxhash ALIAS xxhash)
endif()
endif()
+15
View File
@@ -0,0 +1,15 @@
# SPDX-License-Identifier: MIT
# This applies some common linker options that reduce code size and linking time in Release mode. Namely:
# --gc-sections: Linktime garbage collection, discards unused sections from the final output
# --strip-all : Similar to running `strip`, discards the symbol table from the final output
# --as-needed : Only includes libraries that are actually needed in the final output.
macro(LinkerGC target)
if (CMAKE_BUILD_TYPE MATCHES "RELEASE")
target_link_options(${target} PRIVATE
"LINKER:--gc-sections"
"LINKER:--strip-all"
"LINKER:--as-needed")
endif()
endmacro()
+3 -7
View File
@@ -3,13 +3,10 @@ function(GenBinFmt Name)
get_filename_component(FMT_NAME ${Name} NAME_WE)
# Configure it
configure_file(
${Name}
${CMAKE_BINARY_DIR}/Data/binfmts/${FMT_NAME})
configure_file(${Name} ${CMAKE_BINARY_DIR}/Data/binfmts/${FMT_NAME})
# Then install the configured binfmt
install(
FILES ${CMAKE_BINARY_DIR}/Data/binfmts/${FMT_NAME}
install(FILES ${CMAKE_BINARY_DIR}/Data/binfmts/${FMT_NAME}
DESTINATION ${CMAKE_INSTALL_PREFIX}/share/binfmts/
COMPONENT Runtime)
endfunction()
@@ -18,8 +15,7 @@ if (NOT USE_LEGACY_BINFMTMISC)
configure_file(FEX-x86.conf.in ${CMAKE_BINARY_DIR}/Data/binfmts/FEX-x86.conf)
configure_file(FEX-x86_64.conf.in ${CMAKE_BINARY_DIR}/Data/binfmts/FEX-x86_64.conf)
install(
FILES ${CMAKE_BINARY_DIR}/Data/binfmts/FEX-x86.conf ${CMAKE_BINARY_DIR}/Data/binfmts/FEX-x86_64.conf
install(FILES ${CMAKE_BINARY_DIR}/Data/binfmts/FEX-x86.conf ${CMAKE_BINARY_DIR}/Data/binfmts/FEX-x86_64.conf
DESTINATION ${CMAKE_INSTALL_PREFIX}/lib/binfmt.d/
COMPONENT Runtime)
else()
+1 -1
+2 -3
View File
@@ -1,5 +1,5 @@
set (SRCS
add_library(softfloat_3e STATIC
# F80 support
src/extF80_add.c
src/extF80_div.c
@@ -84,7 +84,7 @@ set (SRCS
src/s_normSubnormalF32Sig.c
src/s_f32UIToCommonNaN.c)
if (_M_ARM_64 AND HAS_CLANG_PRESERVE_ALL)
if (ARCHITECTURE_arm64 AND HAS_CLANG_PRESERVE_ALL)
list(APPEND DEFINES "-DFEXCORE_PRESERVE_ALL_ATTR=__attribute__((preserve_all));-DFEXCORE_HAS_PRESERVE_ALL_ATTR=1")
else()
list(APPEND DEFINES "-DFEXCORE_PRESERVE_ALL_ATTR=;-DFEXCORE_HAS_PRESERVE_ALL_ATTR=0")
@@ -92,7 +92,6 @@ endif()
list(APPEND DEFINES "-DSOFTFLOAT_BUILTIN_CLZ=1;-DINLINE=static inline;-DINLINE_LEVEL=4;-DSOFTFLOAT_FAST_INT64=1;-DSOFTFLOAT_FAST_DIV32TO16=1;-DSOFTFLOAT_FAST_DIV64TO32=1")
add_library(softfloat_3e STATIC ${SRCS})
target_include_directories(softfloat_3e PUBLIC ${CMAKE_CURRENT_SOURCE_DIR}/include/)
target_include_directories(softfloat_3e PUBLIC ${CMAKE_CURRENT_SOURCE_DIR}/include/SoftFloat-3e/)
target_compile_definitions(softfloat_3e PUBLIC ${DEFINES})
+1 -2
View File
@@ -1,4 +1,4 @@
set(SRCS_128BIT
add_library(cephes_128bit STATIC
src/128bit/Impl.cpp
src/128bit/atanll.c
src/128bit/constll.c
@@ -11,7 +11,6 @@ set(SRCS_128BIT
src/128bit/tanll.c)
# 128-bit library
add_library(cephes_128bit STATIC ${SRCS_128BIT})
target_link_libraries(cephes_128bit softfloat_3e)
target_include_directories(cephes_128bit PUBLIC ${CMAKE_CURRENT_SOURCE_DIR}/include/)
target_compile_options(cephes_128bit PRIVATE -fno-builtin)
+248 -148
View File
@@ -4,29 +4,34 @@
#
# pip-compile --generate-hashes --output-file=requirements_formatting.txt --strip-extras requirements_formatting.txt.in
#
black==25.1.0 \
--hash=sha256:030b9759066a4ee5e5aca28c3c77f9c64789cdd4de8ac1df642c40b708be6171 \
--hash=sha256:055e59b198df7ac0b7efca5ad7ff2516bca343276c466be72eb04a3bcc1f82d7 \
--hash=sha256:0e519ecf93120f34243e6b0054db49c00a35f84f195d5bce7e9f5cfc578fc2da \
--hash=sha256:172b1dbff09f86ce6f4eb8edf9dede08b1fce58ba194c87d7a4f1a5aa2f5b3c2 \
--hash=sha256:1e2978f6df243b155ef5fa7e558a43037c3079093ed5d10fd84c43900f2d8ecc \
--hash=sha256:33496d5cd1222ad73391352b4ae8da15253c5de89b93a80b3e2c8d9a19ec2666 \
--hash=sha256:3b48735872ec535027d979e8dcb20bf4f70b5ac75a8ea99f127c106a7d7aba9f \
--hash=sha256:4b60580e829091e6f9238c848ea6750efed72140b91b048770b64e74fe04908b \
--hash=sha256:759e7ec1e050a15f89b770cefbf91ebee8917aac5c20483bc2d80a6c3a04df32 \
--hash=sha256:8f0b18a02996a836cc9c9c78e5babec10930862827b1b724ddfe98ccf2f2fe4f \
--hash=sha256:95e8176dae143ba9097f351d174fdaf0ccd29efb414b362ae3fd72bf0f710717 \
--hash=sha256:96c1c7cd856bba8e20094e36e0f948718dc688dba4a9d78c3adde52b9e6c2299 \
--hash=sha256:a1ee0a0c330f7b5130ce0caed9936a904793576ef4d2b98c40835d6a65afa6a0 \
--hash=sha256:a22f402b410566e2d1c950708c77ebf5ebd5d0d88a6a2e87c86d9fb48afa0d18 \
--hash=sha256:a39337598244de4bae26475f77dda852ea00a93bd4c728e09eacd827ec929df0 \
--hash=sha256:afebb7098bfbc70037a053b91ae8437c3857482d3a690fefc03e9ff7aa9a5fd3 \
--hash=sha256:bacabb307dca5ebaf9c118d2d2f6903da0d62c9faa82bd21a33eecc319559355 \
--hash=sha256:bce2e264d59c91e52d8000d507eb20a9aca4a778731a08cfff7e5ac4a4bb7096 \
--hash=sha256:d9e6827d563a2c820772b32ce8a42828dc6790f095f441beef18f96aa6f8294e \
--hash=sha256:db8ea9917d6f8fc62abd90d944920d95e73c83a5ee3383493e35d271aca872e9 \
--hash=sha256:ea0213189960bda9cf99be5b8c8ce66bb054af5e9e861249cd23471bd7b0b3ba \
--hash=sha256:f3df5f1bf91d36002b0a75389ca8663510cf0531cca8aa5c1ef695b46d98655f
black==26.3.1 \
--hash=sha256:0126ae5b7c09957da2bdbd91a9ba1207453feada9e9fe51992848658c6c8e01c \
--hash=sha256:0f76ff19ec5297dd8e66eb64deda23631e642c9393ab592826fd4bdc97a4bce7 \
--hash=sha256:28ef38aee69e4b12fda8dba75e21f9b4f979b490c8ac0baa7cb505369ac9e1ff \
--hash=sha256:2bd5aa94fc267d38bb21a70d7410a89f1a1d318841855f698746f8e7f51acd1b \
--hash=sha256:2c50f5063a9641c7eed7795014ba37b0f5fa227f3d408b968936e24bc0566b07 \
--hash=sha256:2d6bfaf7fd0993b420bed691f20f9492d53ce9a2bcccea4b797d34e947318a78 \
--hash=sha256:41cd2012d35b47d589cb8a16faf8a32ef7a336f56356babd9fcf70939ad1897f \
--hash=sha256:474c27574d6d7037c1bc875a81d9be0a9a4f9ee95e62800dab3cfaadbf75acd5 \
--hash=sha256:5602bdb96d52d2d0672f24f6ffe5218795736dd34807fd0fd55ccd6bf206168b \
--hash=sha256:5e9d0d86df21f2e1677cc4bd090cd0e446278bcbbe49bf3659c308c3e402843e \
--hash=sha256:5ed0ca58586c8d9a487352a96b15272b7fa55d139fc8496b519e78023a8dab0a \
--hash=sha256:6c54a4a82e291a1fee5137371ab488866b7c86a3305af4026bdd4dc78642e1ac \
--hash=sha256:6e131579c243c98f35bce64a7e08e87fb2d610544754675d4a0e73a070a5aa3a \
--hash=sha256:855822d90f884905362f602880ed8b5df1b7e3ee7d0db2502d4388a954cc8c54 \
--hash=sha256:86a8b5035fce64f5dcd1b794cf8ec4d31fe458cf6ce3986a30deb434df82a1d2 \
--hash=sha256:8a33d657f3276328ce00e4d37fe70361e1ec7614da5d7b6e78de5426cb56332f \
--hash=sha256:92c0ec1f2cc149551a2b7b47efc32c866406b6891b0ee4625e95967c8f4acfb1 \
--hash=sha256:9a5e9f45e5d5e1c5b5c29b3bd4265dcc90e8b92cf4534520896ed77f791f4da5 \
--hash=sha256:afc622538b430aa4c8c853f7f63bc582b3b8030fd8c80b70fb5fa5b834e575c2 \
--hash=sha256:b07fc0dab849d24a80a29cfab8d8a19187d1c4685d8a5e6385a5ce323c1f015f \
--hash=sha256:b5e6f89631eb88a7302d416594a32faeee9fb8fb848290da9d0a5f2903519fc1 \
--hash=sha256:bf9bf162ed91a26f1adba8efda0b573bc6924ec1408a52cc6f82cb73ec2b142c \
--hash=sha256:c7e72339f841b5a237ff14f7d3880ddd0fc7f98a1199e8c4327f9a4f478c1839 \
--hash=sha256:ddb113db38838eb9f043623ba274cfaf7d51d5b0c22ecb30afe58b1bb8322983 \
--hash=sha256:dfdd51fc3e64ea4f35873d1b3fb25326773d55d2329ff8449139ebaad7357efb \
--hash=sha256:f1cd08e99d2f9317292a311dfe578fd2a24b15dbce97792f9c4d752275c1fa56 \
--hash=sha256:f89f2ab047c76a9c03f78d0d66ca519e389519902fa27e7a91117ef7611c0568
# via
# -r requirements_formatting.txt.in
# darker
@@ -36,71 +41,91 @@ certifi==2025.7.14 \
# via
# -r requirements_formatting.txt.in
# requests
cffi==1.15.1 \
--hash=sha256:00a9ed42e88df81ffae7a8ab6d9356b371399b91dbdf0c3cb1e84c03a13aceb5 \
--hash=sha256:03425bdae262c76aad70202debd780501fabeaca237cdfddc008987c0e0f59ef \
--hash=sha256:04ed324bda3cda42b9b695d51bb7d54b680b9719cfab04227cdd1e04e5de3104 \
--hash=sha256:0e2642fe3142e4cc4af0799748233ad6da94c62a8bec3a6648bf8ee68b1c7426 \
--hash=sha256:173379135477dc8cac4bc58f45db08ab45d228b3363adb7af79436135d028405 \
--hash=sha256:198caafb44239b60e252492445da556afafc7d1e3ab7a1fb3f0584ef6d742375 \
--hash=sha256:1e74c6b51a9ed6589199c787bf5f9875612ca4a8a0785fb2d4a84429badaf22a \
--hash=sha256:2012c72d854c2d03e45d06ae57f40d78e5770d252f195b93f581acf3ba44496e \
--hash=sha256:21157295583fe8943475029ed5abdcf71eb3911894724e360acff1d61c1d54bc \
--hash=sha256:2470043b93ff09bf8fb1d46d1cb756ce6132c54826661a32d4e4d132e1977adf \
--hash=sha256:285d29981935eb726a4399badae8f0ffdff4f5050eaa6d0cfc3f64b857b77185 \
--hash=sha256:30d78fbc8ebf9c92c9b7823ee18eb92f2e6ef79b45ac84db507f52fbe3ec4497 \
--hash=sha256:320dab6e7cb2eacdf0e658569d2575c4dad258c0fcc794f46215e1e39f90f2c3 \
--hash=sha256:33ab79603146aace82c2427da5ca6e58f2b3f2fb5da893ceac0c42218a40be35 \
--hash=sha256:3548db281cd7d2561c9ad9984681c95f7b0e38881201e157833a2342c30d5e8c \
--hash=sha256:3799aecf2e17cf585d977b780ce79ff0dc9b78d799fc694221ce814c2c19db83 \
--hash=sha256:39d39875251ca8f612b6f33e6b1195af86d1b3e60086068be9cc053aa4376e21 \
--hash=sha256:3b926aa83d1edb5aa5b427b4053dc420ec295a08e40911296b9eb1b6170f6cca \
--hash=sha256:3bcde07039e586f91b45c88f8583ea7cf7a0770df3a1649627bf598332cb6984 \
--hash=sha256:3d08afd128ddaa624a48cf2b859afef385b720bb4b43df214f85616922e6a5ac \
--hash=sha256:3eb6971dcff08619f8d91607cfc726518b6fa2a9eba42856be181c6d0d9515fd \
--hash=sha256:40f4774f5a9d4f5e344f31a32b5096977b5d48560c5592e2f3d2c4374bd543ee \
--hash=sha256:4289fc34b2f5316fbb762d75362931e351941fa95fa18789191b33fc4cf9504a \
--hash=sha256:470c103ae716238bbe698d67ad020e1db9d9dba34fa5a899b5e21577e6d52ed2 \
--hash=sha256:4f2c9f67e9821cad2e5f480bc8d83b8742896f1242dba247911072d4fa94c192 \
--hash=sha256:50a74364d85fd319352182ef59c5c790484a336f6db772c1a9231f1c3ed0cbd7 \
--hash=sha256:54a2db7b78338edd780e7ef7f9f6c442500fb0d41a5a4ea24fff1c929d5af585 \
--hash=sha256:5635bd9cb9731e6d4a1132a498dd34f764034a8ce60cef4f5319c0541159392f \
--hash=sha256:59c0b02d0a6c384d453fece7566d1c7e6b7bae4fc5874ef2ef46d56776d61c9e \
--hash=sha256:5d598b938678ebf3c67377cdd45e09d431369c3b1a5b331058c338e201f12b27 \
--hash=sha256:5df2768244d19ab7f60546d0c7c63ce1581f7af8b5de3eb3004b9b6fc8a9f84b \
--hash=sha256:5ef34d190326c3b1f822a5b7a45f6c4535e2f47ed06fec77d3d799c450b2651e \
--hash=sha256:6975a3fac6bc83c4a65c9f9fcab9e47019a11d3d2cf7f3c0d03431bf145a941e \
--hash=sha256:6c9a799e985904922a4d207a94eae35c78ebae90e128f0c4e521ce339396be9d \
--hash=sha256:70df4e3b545a17496c9b3f41f5115e69a4f2e77e94e1d2a8e1070bc0c38c8a3c \
--hash=sha256:7473e861101c9e72452f9bf8acb984947aa1661a7704553a9f6e4baa5ba64415 \
--hash=sha256:8102eaf27e1e448db915d08afa8b41d6c7ca7a04b7d73af6514df10a3e74bd82 \
--hash=sha256:87c450779d0914f2861b8526e035c5e6da0a3199d8f1add1a665e1cbc6fc6d02 \
--hash=sha256:8b7ee99e510d7b66cdb6c593f21c043c248537a32e0bedf02e01e9553a172314 \
--hash=sha256:91fc98adde3d7881af9b59ed0294046f3806221863722ba7d8d120c575314325 \
--hash=sha256:94411f22c3985acaec6f83c6df553f2dbe17b698cc7f8ae751ff2237d96b9e3c \
--hash=sha256:98d85c6a2bef81588d9227dde12db8a7f47f639f4a17c9ae08e773aa9c697bf3 \
--hash=sha256:9ad5db27f9cabae298d151c85cf2bad1d359a1b9c686a275df03385758e2f914 \
--hash=sha256:a0b71b1b8fbf2b96e41c4d990244165e2c9be83d54962a9a1d118fd8657d2045 \
--hash=sha256:a0f100c8912c114ff53e1202d0078b425bee3649ae34d7b070e9697f93c5d52d \
--hash=sha256:a591fe9e525846e4d154205572a029f653ada1a78b93697f3b5a8f1f2bc055b9 \
--hash=sha256:a5c84c68147988265e60416b57fc83425a78058853509c1b0629c180094904a5 \
--hash=sha256:a66d3508133af6e8548451b25058d5812812ec3798c886bf38ed24a98216fab2 \
--hash=sha256:a8c4917bd7ad33e8eb21e9a5bbba979b49d9a97acb3a803092cbc1133e20343c \
--hash=sha256:b3bbeb01c2b273cca1e1e0c5df57f12dce9a4dd331b4fa1635b8bec26350bde3 \
--hash=sha256:cba9d6b9a7d64d4bd46167096fc9d2f835e25d7e4c121fb2ddfc6528fb0413b2 \
--hash=sha256:cc4d65aeeaa04136a12677d3dd0b1c0c94dc43abac5860ab33cceb42b801c1e8 \
--hash=sha256:ce4bcc037df4fc5e3d184794f27bdaab018943698f4ca31630bc7f84a7b69c6d \
--hash=sha256:cec7d9412a9102bdc577382c3929b337320c4c4c4849f2c5cdd14d7368c5562d \
--hash=sha256:d400bfb9a37b1351253cb402671cea7e89bdecc294e8016a707f6d1d8ac934f9 \
--hash=sha256:d61f4695e6c866a23a21acab0509af1cdfd2c013cf256bbf5b6b5e2695827162 \
--hash=sha256:db0fbb9c62743ce59a9ff687eb5f4afbe77e5e8403d6697f7446e5f609976f76 \
--hash=sha256:dd86c085fae2efd48ac91dd7ccffcfc0571387fe1193d33b6394db7ef31fe2a4 \
--hash=sha256:e00b098126fd45523dd056d2efba6c5a63b71ffe9f2bbe1a4fe1716e1d0c331e \
--hash=sha256:e229a521186c75c8ad9490854fd8bbdd9a0c9aa3a524326b55be83b54d4e0ad9 \
--hash=sha256:e263d77ee3dd201c3a142934a086a4450861778baaeeb45db4591ef65550b0a6 \
--hash=sha256:ed9cb427ba5504c1dc15ede7d516b84757c3e3d7868ccc85121d9310d27eed0b \
--hash=sha256:fa6693661a4c91757f4412306191b6dc88c1703f780c8234035eac011922bc01 \
--hash=sha256:fcd131dd944808b5bdb38e6f5b53013c5aa4f334c5cad0c72742f6eba4b73db0
cffi==2.0.0 \
--hash=sha256:00bdf7acc5f795150faa6957054fbbca2439db2f775ce831222b66f192f03beb \
--hash=sha256:07b271772c100085dd28b74fa0cd81c8fb1a3ba18b21e03d7c27f3436a10606b \
--hash=sha256:087067fa8953339c723661eda6b54bc98c5625757ea62e95eb4898ad5e776e9f \
--hash=sha256:0a1527a803f0a659de1af2e1fd700213caba79377e27e4693648c2923da066f9 \
--hash=sha256:0cf2d91ecc3fcc0625c2c530fe004f82c110405f101548512cce44322fa8ac44 \
--hash=sha256:0f6084a0ea23d05d20c3edcda20c3d006f9b6f3fefeac38f59262e10cef47ee2 \
--hash=sha256:12873ca6cb9b0f0d3a0da705d6086fe911591737a59f28b7936bdfed27c0d47c \
--hash=sha256:19f705ada2530c1167abacb171925dd886168931e0a7b78f5bffcae5c6b5be75 \
--hash=sha256:1cd13c99ce269b3ed80b417dcd591415d3372bcac067009b6e0f59c7d4015e65 \
--hash=sha256:1e3a615586f05fc4065a8b22b8152f0c1b00cdbc60596d187c2a74f9e3036e4e \
--hash=sha256:1f72fb8906754ac8a2cc3f9f5aaa298070652a0ffae577e0ea9bd480dc3c931a \
--hash=sha256:1fc9ea04857caf665289b7a75923f2c6ed559b8298a1b8c49e59f7dd95c8481e \
--hash=sha256:203a48d1fb583fc7d78a4c6655692963b860a417c0528492a6bc21f1aaefab25 \
--hash=sha256:2081580ebb843f759b9f617314a24ed5738c51d2aee65d31e02f6f7a2b97707a \
--hash=sha256:21d1152871b019407d8ac3985f6775c079416c282e431a4da6afe7aefd2bccbe \
--hash=sha256:24b6f81f1983e6df8db3adc38562c83f7d4a0c36162885ec7f7b77c7dcbec97b \
--hash=sha256:256f80b80ca3853f90c21b23ee78cd008713787b1b1e93eae9f3d6a7134abd91 \
--hash=sha256:28a3a209b96630bca57cce802da70c266eb08c6e97e5afd61a75611ee6c64592 \
--hash=sha256:2c8f814d84194c9ea681642fd164267891702542f028a15fc97d4674b6206187 \
--hash=sha256:2de9a304e27f7596cd03d16f1b7c72219bd944e99cc52b84d0145aefb07cbd3c \
--hash=sha256:38100abb9d1b1435bc4cc340bb4489635dc2f0da7456590877030c9b3d40b0c1 \
--hash=sha256:3925dd22fa2b7699ed2617149842d2e6adde22b262fcbfada50e3d195e4b3a94 \
--hash=sha256:3e17ed538242334bf70832644a32a7aae3d83b57567f9fd60a26257e992b79ba \
--hash=sha256:3e837e369566884707ddaf85fc1744b47575005c0a229de3327f8f9a20f4efeb \
--hash=sha256:3f4d46d8b35698056ec29bca21546e1551a205058ae1a181d871e278b0b28165 \
--hash=sha256:44d1b5909021139fe36001ae048dbdde8214afa20200eda0f64c068cac5d5529 \
--hash=sha256:45d5e886156860dc35862657e1494b9bae8dfa63bf56796f2fb56e1679fc0bca \
--hash=sha256:4647afc2f90d1ddd33441e5b0e85b16b12ddec4fca55f0d9671fef036ecca27c \
--hash=sha256:4671d9dd5ec934cb9a73e7ee9676f9362aba54f7f34910956b84d727b0d73fb6 \
--hash=sha256:53f77cbe57044e88bbd5ed26ac1d0514d2acf0591dd6bb02a3ae37f76811b80c \
--hash=sha256:5eda85d6d1879e692d546a078b44251cdd08dd1cfb98dfb77b670c97cee49ea0 \
--hash=sha256:5fed36fccc0612a53f1d4d9a816b50a36702c28a2aa880cb8a122b3466638743 \
--hash=sha256:61d028e90346df14fedc3d1e5441df818d095f3b87d286825dfcbd6459b7ef63 \
--hash=sha256:66f011380d0e49ed280c789fbd08ff0d40968ee7b665575489afa95c98196ab5 \
--hash=sha256:6824f87845e3396029f3820c206e459ccc91760e8fa24422f8b0c3d1731cbec5 \
--hash=sha256:6c6c373cfc5c83a975506110d17457138c8c63016b563cc9ed6e056a82f13ce4 \
--hash=sha256:6d02d6655b0e54f54c4ef0b94eb6be0607b70853c45ce98bd278dc7de718be5d \
--hash=sha256:6d50360be4546678fc1b79ffe7a66265e28667840010348dd69a314145807a1b \
--hash=sha256:730cacb21e1bdff3ce90babf007d0a0917cc3e6492f336c2f0134101e0944f93 \
--hash=sha256:737fe7d37e1a1bffe70bd5754ea763a62a066dc5913ca57e957824b72a85e205 \
--hash=sha256:74a03b9698e198d47562765773b4a8309919089150a0bb17d829ad7b44b60d27 \
--hash=sha256:7553fb2090d71822f02c629afe6042c299edf91ba1bf94951165613553984512 \
--hash=sha256:7a66c7204d8869299919db4d5069a82f1561581af12b11b3c9f48c584eb8743d \
--hash=sha256:7cc09976e8b56f8cebd752f7113ad07752461f48a58cbba644139015ac24954c \
--hash=sha256:81afed14892743bbe14dacb9e36d9e0e504cd204e0b165062c488942b9718037 \
--hash=sha256:8941aaadaf67246224cee8c3803777eed332a19d909b47e29c9842ef1e79ac26 \
--hash=sha256:89472c9762729b5ae1ad974b777416bfda4ac5642423fa93bd57a09204712322 \
--hash=sha256:8ea985900c5c95ce9db1745f7933eeef5d314f0565b27625d9a10ec9881e1bfb \
--hash=sha256:8eca2a813c1cb7ad4fb74d368c2ffbbb4789d377ee5bb8df98373c2cc0dee76c \
--hash=sha256:92b68146a71df78564e4ef48af17551a5ddd142e5190cdf2c5624d0c3ff5b2e8 \
--hash=sha256:9332088d75dc3241c702d852d4671613136d90fa6881da7d770a483fd05248b4 \
--hash=sha256:94698a9c5f91f9d138526b48fe26a199609544591f859c870d477351dc7b2414 \
--hash=sha256:9a67fc9e8eb39039280526379fb3a70023d77caec1852002b4da7e8b270c4dd9 \
--hash=sha256:9de40a7b0323d889cf8d23d1ef214f565ab154443c42737dfe52ff82cf857664 \
--hash=sha256:a05d0c237b3349096d3981b727493e22147f934b20f6f125a3eba8f994bec4a9 \
--hash=sha256:afb8db5439b81cf9c9d0c80404b60c3cc9c3add93e114dcae767f1477cb53775 \
--hash=sha256:b18a3ed7d5b3bd8d9ef7a8cb226502c6bf8308df1525e1cc676c3680e7176739 \
--hash=sha256:b1e74d11748e7e98e2f426ab176d4ed720a64412b6a15054378afdb71e0f37dc \
--hash=sha256:b21e08af67b8a103c71a250401c78d5e0893beff75e28c53c98f4de42f774062 \
--hash=sha256:b4c854ef3adc177950a8dfc81a86f5115d2abd545751a304c5bcf2c2c7283cfe \
--hash=sha256:b882b3df248017dba09d6b16defe9b5c407fe32fc7c65a9c69798e6175601be9 \
--hash=sha256:baf5215e0ab74c16e2dd324e8ec067ef59e41125d3eade2b863d294fd5035c92 \
--hash=sha256:c649e3a33450ec82378822b3dad03cc228b8f5963c0c12fc3b1e0ab940f768a5 \
--hash=sha256:c654de545946e0db659b3400168c9ad31b5d29593291482c43e3564effbcee13 \
--hash=sha256:c6638687455baf640e37344fe26d37c404db8b80d037c3d29f58fe8d1c3b194d \
--hash=sha256:c8d3b5532fc71b7a77c09192b4a5a200ea992702734a2e9279a37f2478236f26 \
--hash=sha256:cb527a79772e5ef98fb1d700678fe031e353e765d1ca2d409c92263c6d43e09f \
--hash=sha256:cf364028c016c03078a23b503f02058f1814320a56ad535686f90565636a9495 \
--hash=sha256:d48a880098c96020b02d5a1f7d9251308510ce8858940e6fa99ece33f610838b \
--hash=sha256:d68b6cef7827e8641e8ef16f4494edda8b36104d79773a334beaa1e3521430f6 \
--hash=sha256:d9b29c1f0ae438d5ee9acb31cadee00a58c46cc9c0b2f9038c6b0b3470877a8c \
--hash=sha256:d9b97165e8aed9272a6bb17c01e3cc5871a594a446ebedc996e2397a1c1ea8ef \
--hash=sha256:da68248800ad6320861f129cd9c1bf96ca849a2771a59e0344e88681905916f5 \
--hash=sha256:da902562c3e9c550df360bfa53c035b2f241fed6d9aef119048073680ace4a18 \
--hash=sha256:dbd5c7a25a7cb98f5ca55d258b103a2054f859a46ae11aaf23134f9cc0d356ad \
--hash=sha256:dd4f05f54a52fb558f1ba9f528228066954fee3ebe629fc1660d874d040ae5a3 \
--hash=sha256:de8dad4425a6ca6e4e5e297b27b5c824ecc7581910bf9aee86cb6835e6812aa7 \
--hash=sha256:e11e82b744887154b182fd3e7e8512418446501191994dbf9c9fc1f32cc8efd5 \
--hash=sha256:e6e73b9e02893c764e7e8d5bb5ce277f1a009cd5243f8228f75f842bf937c534 \
--hash=sha256:f73b96c41e3b2adedc34a7356e64c8eb96e03a3782b535e043a986276ce12a49 \
--hash=sha256:f93fd8e5c8c0a4aa1f424d6173f14a892044054871c771f8566e4008eaa359d2 \
--hash=sha256:fc33c5141b55ed366cfaad382df24fe7dcbc686de5be719b207bb248e3053dc5 \
--hash=sha256:fc7de24befaeae77ba923797c7c87834c73648a05a4bde34b3b7e5588973a453 \
--hash=sha256:fe562eb1a64e67dd297ccc4f5addea2501664954f2692b69a76449ec7913ecbf
# via
# cryptography
# pynacl
@@ -185,44 +210,56 @@ click==8.1.7 \
--hash=sha256:ae74fb96c20a0277a1d615f1e4d73c8414f5a98db8b799a7931d1582f3390c28 \
--hash=sha256:ca9853ad459e787e2192211578cc907e7594e294c7ccc834310722b41b9ca6de
# via black
cryptography==45.0.5 \
--hash=sha256:0027d566d65a38497bc37e0dd7c2f8ceda73597d2ac9ba93810204f56f52ebc7 \
--hash=sha256:101ee65078f6dd3e5a028d4f19c07ffa4dd22cce6a20eaa160f8b5219911e7d8 \
--hash=sha256:12e55281d993a793b0e883066f590c1ae1e802e3acb67f8b442e721e475e6463 \
--hash=sha256:14d96584701a887763384f3c47f0ca7c1cce322aa1c31172680eb596b890ec30 \
--hash=sha256:1e1da5accc0c750056c556a93c3e9cb828970206c68867712ca5805e46dc806f \
--hash=sha256:206210d03c1193f4e1ff681d22885181d47efa1ab3018766a7b32a7b3d6e6afd \
--hash=sha256:2089cc8f70a6e454601525e5bf2779e665d7865af002a5dec8d14e561002e135 \
--hash=sha256:3a264aae5f7fbb089dbc01e0242d3b67dffe3e6292e1f5182122bdf58e65215d \
--hash=sha256:3af26738f2db354aafe492fb3869e955b12b2ef2e16908c8b9cb928128d42c57 \
--hash=sha256:3fcfbefc4a7f332dece7272a88e410f611e79458fab97b5efe14e54fe476f4fd \
--hash=sha256:460f8c39ba66af7db0545a8c6f2eabcbc5a5528fc1cf6c3fa9a1e44cec33385e \
--hash=sha256:57c816dfbd1659a367831baca4b775b2a5b43c003daf52e9d57e1d30bc2e1b0e \
--hash=sha256:5aa1e32983d4443e310f726ee4b071ab7569f58eedfdd65e9675484a4eb67bd1 \
--hash=sha256:6ff8728d8d890b3dda5765276d1bc6fb099252915a2cd3aff960c4c195745dd0 \
--hash=sha256:7259038202a47fdecee7e62e0fd0b0738b6daa335354396c6ddebdbe1206af2a \
--hash=sha256:72e76caa004ab63accdf26023fccd1d087f6d90ec6048ff33ad0445abf7f605a \
--hash=sha256:7760c1c2e1a7084153a0f68fab76e754083b126a47d0117c9ed15e69e2103492 \
--hash=sha256:8c4a6ff8a30e9e3d38ac0539e9a9e02540ab3f827a3394f8852432f6b0ea152e \
--hash=sha256:9024beb59aca9d31d36fcdc1604dd9bbeed0a55bface9f1908df19178e2f116e \
--hash=sha256:90cb0a7bb35959f37e23303b7eed0a32280510030daba3f7fdfbb65defde6a97 \
--hash=sha256:91098f02ca81579c85f66df8a588c78f331ca19089763d733e34ad359f474174 \
--hash=sha256:926c3ea71a6043921050eaa639137e13dbe7b4ab25800932a8498364fc1abec9 \
--hash=sha256:982518cd64c54fcada9d7e5cf28eabd3ee76bd03ab18e08a48cad7e8b6f31b18 \
--hash=sha256:9b4cf6318915dccfe218e69bbec417fdd7c7185aa7aab139a2c0beb7468c89f0 \
--hash=sha256:ad0caded895a00261a5b4aa9af828baede54638754b51955a0ac75576b831b27 \
--hash=sha256:b85980d1e345fe769cfc57c57db2b59cff5464ee0c045d52c0df087e926fbe63 \
--hash=sha256:b8fa8b0a35a9982a3c60ec79905ba5bb090fc0b9addcfd3dc2dd04267e45f25e \
--hash=sha256:b9e38e0a83cd51e07f5a48ff9691cae95a79bea28fe4ded168a8e5c6c77e819d \
--hash=sha256:bd4c45986472694e5121084c6ebbd112aa919a25e783b87eb95953c9573906d6 \
--hash=sha256:be97d3a19c16a9be00edf79dca949c8fa7eff621763666a145f9f9535a5d7f42 \
--hash=sha256:c648025b6840fe62e57107e0a25f604db740e728bd67da4f6f060f03017d5097 \
--hash=sha256:d05a38884db2ba215218745f0781775806bde4f32e07b135348355fe8e4991d9 \
--hash=sha256:dd420e577921c8c2d31289536c386aaa30140b473835e97f83bc71ea9d2baf2d \
--hash=sha256:e357286c1b76403dd384d938f93c46b2b058ed4dfcdce64a770f0537ed3feb6f \
--hash=sha256:e6c00130ed423201c5bc5544c23359141660b07999ad82e34e7bb8f882bb78e0 \
--hash=sha256:e74d30ec9c7cb2f404af331d5b4099a9b322a8a6b25c4632755c8757345baac5 \
--hash=sha256:f3562c2f23c612f2e4a6964a61d942f891d29ee320edb62ff48ffb99f3de9ae8
cryptography==46.0.5 \
--hash=sha256:02f547fce831f5096c9a567fd41bc12ca8f11df260959ecc7c3202555cc47a72 \
--hash=sha256:039917b0dc418bb9f6edce8a906572d69e74bd330b0b3fea4f79dab7f8ddd235 \
--hash=sha256:1abfdb89b41c3be0365328a410baa9df3ff8a9110fb75e7b52e66803ddabc9a9 \
--hash=sha256:2ae6971afd6246710480e3f15824ed3029a60fc16991db250034efd0b9fb4356 \
--hash=sha256:2b7a67c9cd56372f3249b39699f2ad479f6991e62ea15800973b956f4b73e257 \
--hash=sha256:351695ada9ea9618b3500b490ad54c739860883df6c1f555e088eaf25b1bbaad \
--hash=sha256:38946c54b16c885c72c4f59846be9743d699eee2b69b6988e0a00a01f46a61a4 \
--hash=sha256:3b4995dc971c9fb83c25aa44cf45f02ba86f71ee600d81091c2f0cbae116b06c \
--hash=sha256:3ce58ba46e1bc2aac4f7d9290223cead56743fa6ab94a5d53292ffaac6a91614 \
--hash=sha256:3ee190460e2fbe447175cda91b88b84ae8322a104fc27766ad09428754a618ed \
--hash=sha256:4108d4c09fbbf2789d0c926eb4152ae1760d5a2d97612b92d508d96c861e4d31 \
--hash=sha256:420d0e909050490d04359e7fdb5ed7e667ca5c3c402b809ae2563d7e66a92229 \
--hash=sha256:47fb8a66058b80e509c47118ef8a75d14c455e81ac369050f20ba0d23e77fee0 \
--hash=sha256:4c3341037c136030cb46e4b1e17b7418ea4cbd9dd207e4a6f3b2b24e0d4ac731 \
--hash=sha256:4d7e3d356b8cd4ea5aff04f129d5f66ebdc7b6f8eae802b93739ed520c47c79b \
--hash=sha256:4d8ae8659ab18c65ced284993c2265910f6c9e650189d4e3f68445ef82a810e4 \
--hash=sha256:4e817a8920bfbcff8940ecfd60f23d01836408242b30f1a708d93198393a80b4 \
--hash=sha256:50bfb6925eff619c9c023b967d5b77a54e04256c4281b0e21336a130cd7fc263 \
--hash=sha256:556e106ee01aa13484ce9b0239bca667be5004efb0aabbed28d353df86445595 \
--hash=sha256:582f5fcd2afa31622f317f80426a027f30dc792e9c80ffee87b993200ea115f1 \
--hash=sha256:5be7bf2fb40769e05739dd0046e7b26f9d4670badc7b032d6ce4db64dddc0678 \
--hash=sha256:60ee7e19e95104d4c03871d7d7dfb3d22ef8a9b9c6778c94e1c8fcc8365afd48 \
--hash=sha256:61aa400dce22cb001a98014f647dc21cda08f7915ceb95df0c9eaf84b4b6af76 \
--hash=sha256:68f68d13f2e1cb95163fa3b4db4bf9a159a418f5f6e7242564fc75fcae667fd0 \
--hash=sha256:7d1f30a86d2757199cb2d56e48cce14deddf1f9c95f1ef1b64ee91ea43fe2e18 \
--hash=sha256:7d731d4b107030987fd61a7f8ab512b25b53cef8f233a97379ede116f30eb67d \
--hash=sha256:803812e111e75d1aa73690d2facc295eaefd4439be1023fefc4995eaea2af90d \
--hash=sha256:80a8d7bfdf38f87ca30a5391c0c9ce4ed2926918e017c29ddf643d0ed2778ea1 \
--hash=sha256:8293f3dea7fc929ef7240796ba231413afa7b68ce38fd21da2995549f5961981 \
--hash=sha256:8456928655f856c6e1533ff59d5be76578a7157224dbd9ce6872f25055ab9ab7 \
--hash=sha256:890bcb4abd5a2d3f852196437129eb3667d62630333aacc13dfd470fad3aaa82 \
--hash=sha256:94a76daa32eb78d61339aff7952ea819b1734b46f73646a07decb40e5b3448e2 \
--hash=sha256:9f16fbdf4da055efb21c22d81b89f155f02ba420558db21288b3d0035bafd5f4 \
--hash=sha256:a3d1fae9863299076f05cb8a778c467578262fae09f9dc0ee9b12eb4268ce663 \
--hash=sha256:a3d507bb6a513ca96ba84443226af944b0f7f47dcc9a399d110cd6146481d24c \
--hash=sha256:abace499247268e3757271b2f1e244b36b06f8515cf27c4d49468fc9eb16e93d \
--hash=sha256:ba2a27ff02f48193fc4daeadf8ad2590516fa3d0adeeb34336b96f7fa64c1e3a \
--hash=sha256:bc84e875994c3b445871ea7181d424588171efec3e185dced958dad9e001950a \
--hash=sha256:bfd56bb4b37ed4f330b82402f6f435845a5f5648edf1ad497da51a8452d5d62d \
--hash=sha256:c18ff11e86df2e28854939acde2d003f7984f721eba450b56a200ad90eeb0e6b \
--hash=sha256:c3bcce8521d785d510b2aad26ae2c966092b7daa8f45dd8f44734a104dc0bc1a \
--hash=sha256:c4143987a42a2397f2fc3b4d7e3a7d313fbe684f67ff443999e803dd75a76826 \
--hash=sha256:c69fd885df7d089548a42d5ec05be26050ebcd2283d89b3d30676eb32ff87dee \
--hash=sha256:ced80795227d70549a411a4ab66e8ce307899fad2220ce5ab2f296e687eacde9 \
--hash=sha256:d66e421495fdb797610a08f43b05269e0a5ea7f5e652a89bfd5a7d3c1dee3648 \
--hash=sha256:d861ee9e76ace6cf36a6a89b959ec08e7bc2493ee39d07ffe5acb23ef46d27da \
--hash=sha256:e9251e3be159d1020c4030bd2e5f84d6a43fe54b6c19c12f51cde9542a2817b2 \
--hash=sha256:f145bba11b878005c496e93e257c1e88f154d278d2638e6450d17e0f31e558d2 \
--hash=sha256:fe346b143ff9685e40192a4960938545c699054ba11d4f9029f94751e3f71d87
# via
# -r requirements_formatting.txt.in
# pyjwt
@@ -258,9 +295,9 @@ packaging==23.1 \
--hash=sha256:994793af429502c4ea2ebf6bf664629d07c1a9fe974af92966e4b8d2df7edc61 \
--hash=sha256:a392980d2b6cffa644431898be54b0045151319d1e7ec34f0cfed48767dd334f
# via black
pathspec==0.11.2 \
--hash=sha256:1d6ed233af05e679efb96b1851550ea95bbb64b7c490b0f5aa52996c11e92a20 \
--hash=sha256:e0d8d0ac2f12da61956eb2306b69f9469b42f4deb0f3cb6ed47b9cce9996ced3
pathspec==1.0.4 \
--hash=sha256:0210e2ae8a21a9137c0d470578cb0e595af87edaa6ebf12ff176f14a02e0e645 \
--hash=sha256:fb6ae2fd4e7c921a165808a552060e722767cfa526f99ca5156ed2ce45a5c723
# via black
platformdirs==3.10.0 \
--hash=sha256:b45696dab2d7cc691a3226759c0d3b00c47c8b6e293d96f6436f733303f77f6d \
@@ -274,22 +311,85 @@ pygithub==2.6.1 \
--hash=sha256:6f2fa6d076ccae475f9fc392cc6cdbd54db985d4f69b8833a28397de75ed6ca3 \
--hash=sha256:b5c035392991cca63959e9453286b41b54d83bf2de2daa7d7ff7e4312cebf3bf
# via -r requirements_formatting.txt.in
pyjwt==2.8.0 \
--hash=sha256:57e28d156e3d5c10088e0c68abb90bfac3df82b40a71bd0daa20c65ccd5c23de \
--hash=sha256:59127c392cc44c2da5bb3192169a91f429924e17aff6534d70fdc02ab3e04320
# via pygithub
pynacl==1.5.0 \
--hash=sha256:06b8f6fa7f5de8d5d2f7573fe8c863c051225a27b61e6860fd047b1775807858 \
--hash=sha256:0c84947a22519e013607c9be43706dd42513f9e6ae5d39d3613ca1e142fba44d \
--hash=sha256:20f42270d27e1b6a29f54032090b972d97f0a1b0948cc52392041ef7831fee93 \
--hash=sha256:401002a4aaa07c9414132aaed7f6836ff98f59277a234704ff66878c2ee4a0d1 \
--hash=sha256:52cb72a79269189d4e0dc537556f4740f7f0a9ec41c1322598799b0bdad4ef92 \
--hash=sha256:61f642bf2378713e2c2e1de73444a3778e5f0a38be6fee0fe532fe30060282ff \
--hash=sha256:8ac7448f09ab85811607bdd21ec2464495ac8b7c66d146bf545b0f08fb9220ba \
--hash=sha256:a36d4a9dda1f19ce6e03c9a784a2921a4b726b02e1c736600ca9c22029474394 \
--hash=sha256:a422368fc821589c228f4c49438a368831cb5bbc0eab5ebe1d7fac9dded6567b \
--hash=sha256:e46dae94e34b085175f8abb3b0aaa7da40767865ac82c928eeb9e57e1ea8a543
# via pygithub
pyjwt==2.12.1 \
--hash=sha256:28ca37c070cad8ba8cd9790cd940535d40274d22f80ab87f3ac6a713e6e8454c \
--hash=sha256:c74a7a2adf861c04d002db713dd85f84beb242228e671280bf709d765b03672b
# via
# -r requirements_formatting.txt.in
# pygithub
pynacl==1.6.2 \
--hash=sha256:018494d6d696ae03c7e656e5e74cdfd8ea1326962cc401bcf018f1ed8436811c \
--hash=sha256:04316d1fc625d860b6c162fff704eb8426b1a8bcd3abacea11142cbd99a6b574 \
--hash=sha256:22de65bb9010a725b0dac248f353bb072969c94fa8d6b1f34b87d7953cf7bbe4 \
--hash=sha256:26bfcd00dcf2cf160f122186af731ae30ab120c18e8375684ec2670dccd28130 \
--hash=sha256:2fef529ef3ee487ad8113d287a593fa26f48ee3620d92ecc6f1d09ea38e0709b \
--hash=sha256:320ef68a41c87547c91a8b58903c9caa641ab01e8512ce291085b5fe2fcb7590 \
--hash=sha256:3bffb6d0f6becacb6526f8f42adfb5efb26337056ee0831fb9a7044d1a964444 \
--hash=sha256:44081faff368d6c5553ccf55322ef2819abb40e25afaec7e740f159f74813634 \
--hash=sha256:46065496ab748469cdd999246d17e301b2c24ae2fdf739132e580a0e94c94a87 \
--hash=sha256:5811c72b473b2f38f7e2a3dc4f8642e3a3e9b5e7317266e4ced1fba85cae41aa \
--hash=sha256:622d7b07cc5c02c666795792931b50c91f3ce3c2649762efb1ef0d5684c81594 \
--hash=sha256:62985f233210dee6548c223301b6c25440852e13d59a8b81490203c3227c5ba0 \
--hash=sha256:68be3a09455743ff9505491220b64440ced8973fe930f270c8e07ccfa25b1f9e \
--hash=sha256:834a43af110f743a754448463e8fd61259cd4ab5bbedcf70f9dabad1d28a394c \
--hash=sha256:8845c0631c0be43abdd865511c41eab235e0be69c81dc66a50911594198679b0 \
--hash=sha256:8a66d6fb6ae7661c58995f9c6435bda2b1e68b54b598a6a10247bfcdadac996c \
--hash=sha256:8b097553b380236d51ed11356c953bf8ce36a29a3e596e934ecabe76c985a577 \
--hash=sha256:a84bf1c20339d06dc0c85d9aea9637a24f718f375d861b2668b2f9f96fa51145 \
--hash=sha256:a9f9932d8d2811ce1a8ffa79dcbdf3970e7355b5c8eb0c1a881a57e7f7d96e88 \
--hash=sha256:bc4a36b28dd72fb4845e5d8f9760610588a96d5a51f01d84d8c6ff9849968c14 \
--hash=sha256:c8a231e36ec2cab018c4ad4358c386e36eede0319a0c41fed24f840b1dac59f6 \
--hash=sha256:c949ea47e4206af7c8f604b8278093b674f7c79ed0d4719cc836902bf4517465 \
--hash=sha256:d071c6a9a4c94d79eb665db4ce5cedc537faf74f2355e4d502591d850d3913c0 \
--hash=sha256:d29bfe37e20e015a7d8b23cfc8bd6aa7909c92a1b8f41ee416bbb3e79ef182b2 \
--hash=sha256:fe9847ca47d287af41e82be1dd5e23023d3c31a951da134121ab02e42ac218c9
# via
# -r requirements_formatting.txt.in
# pygithub
pytokens==0.4.1 \
--hash=sha256:0fc71786e629cef478cbf29d7ea1923299181d0699dbe7c3c0f4a583811d9fc1 \
--hash=sha256:11edda0942da80ff58c4408407616a310adecae1ddd22eef8c692fe266fa5009 \
--hash=sha256:140709331e846b728475786df8aeb27d24f48cbcf7bcd449f8de75cae7a45083 \
--hash=sha256:24afde1f53d95348b5a0eb19488661147285ca4dd7ed752bbc3e1c6242a304d1 \
--hash=sha256:26cef14744a8385f35d0e095dc8b3a7583f6c953c2e3d269c7f82484bf5ad2de \
--hash=sha256:27b83ad28825978742beef057bfe406ad6ed524b2d28c252c5de7b4a6dd48fa2 \
--hash=sha256:292052fe80923aae2260c073f822ceba21f3872ced9a68bb7953b348e561179a \
--hash=sha256:29d1d8fb1030af4d231789959f21821ab6325e463f0503a61d204343c9b355d1 \
--hash=sha256:2a44ed93ea23415c54f3face3b65ef2b844d96aeb3455b8a69b3df6beab6acc5 \
--hash=sha256:30f51edd9bb7f85c748979384165601d028b84f7bd13fe14d3e065304093916a \
--hash=sha256:34bcc734bd2f2d5fe3b34e7b3c0116bfb2397f2d9666139988e7a3eb5f7400e3 \
--hash=sha256:3ad72b851e781478366288743198101e5eb34a414f1d5627cdd585ca3b25f1db \
--hash=sha256:3f901fe783e06e48e8cbdc82d631fca8f118333798193e026a50ce1b3757ea68 \
--hash=sha256:42f144f3aafa5d92bad964d471a581651e28b24434d184871bd02e3a0d956037 \
--hash=sha256:4a14d5f5fc78ce85e426aa159489e2d5961acf0e47575e08f35584009178e321 \
--hash=sha256:4a58d057208cb9075c144950d789511220b07636dd2e4708d5645d24de666bdc \
--hash=sha256:4e691d7f5186bd2842c14813f79f8884bb03f5995f0575272009982c5ac6c0f7 \
--hash=sha256:5502408cab1cb18e128570f8d598981c68a50d0cbd7c61312a90507cd3a1276f \
--hash=sha256:584c80c24b078eec1e227079d56dc22ff755e0ba8654d8383b2c549107528918 \
--hash=sha256:5ad948d085ed6c16413eb5fec6b3e02fa00dc29a2534f088d3302c47eb59adf9 \
--hash=sha256:670d286910b531c7b7e3c0b453fd8156f250adb140146d234a82219459b9640c \
--hash=sha256:682fa37ff4d8e95f7df6fe6fe6a431e8ed8e788023c6bcc0f0880a12eab80ad1 \
--hash=sha256:6d6c4268598f762bc8e91f5dbf2ab2f61f7b95bdc07953b602db879b3c8c18e1 \
--hash=sha256:79fc6b8699564e1f9b521582c35435f1bd32dd06822322ec44afdeba666d8cb3 \
--hash=sha256:8bdb9d0ce90cbf99c525e75a2fa415144fd570a1ba987380190e8b786bc6ef9b \
--hash=sha256:8fcb9ba3709ff77e77f1c7022ff11d13553f3c30299a9fe246a166903e9091eb \
--hash=sha256:941d4343bf27b605e9213b26bfa1c4bf197c9c599a9627eb7305b0defcfe40c1 \
--hash=sha256:967cf6e3fd4adf7de8fc73cd3043754ae79c36475c1c11d514fc72cf5490094a \
--hash=sha256:970b08dd6b86058b6dc07efe9e98414f5102974716232d10f32ff39701e841c4 \
--hash=sha256:97f50fd18543be72da51dd505e2ed20d2228c74e0464e4262e4899797803d7fa \
--hash=sha256:9bd7d7f544d362576be74f9d5901a22f317efc20046efe2034dced238cbbfe78 \
--hash=sha256:add8bf86b71a5d9fb5b89f023a80b791e04fba57960aa790cc6125f7f1d39dfe \
--hash=sha256:b35d7e5ad269804f6697727702da3c517bb8a5228afa450ab0fa787732055fc9 \
--hash=sha256:b49750419d300e2b5a3813cf229d4e5a4c728dae470bcc89867a9ad6f25a722d \
--hash=sha256:d31b97b3de0f61571a124a00ffe9a81fb9939146c122c11060725bd5aea79975 \
--hash=sha256:d70e77c55ae8380c91c0c18dea05951482e263982911fc7410b1ffd1dadd3440 \
--hash=sha256:d9907d61f15bf7261d7e775bd5d7ee4d2930e04424bab1972591918497623a16 \
--hash=sha256:da5baeaf7116dced9c6bb76dc31ba04a2dc3695f3d9f74741d7910122b456edc \
--hash=sha256:dc74c035f9bfca0255c1af77ddd2d6ae8419012805453e4b0e7513e17904545d \
--hash=sha256:dcafc12c30dbaf1e2af0490978352e0c4041a7cde31f4f81435c2a5e8b9cabb6 \
--hash=sha256:ee44d0f85b803321710f9239f335aafe16553b39106384cef8e6de40cb4ef2f6 \
--hash=sha256:f66a6bbe741bd431f6d741e617e0f39ec7257ca1f89089593479347cc4d13324
# via black
requests==2.32.4 \
--hash=sha256:27babd3cda2a6d50b30443204ee89830707d396671944c998b5975b031ac2b2c \
--hash=sha256:27d0316682c8a29834d3264820024b62a36942083d52caf2f14c0591336d3422
@@ -306,9 +406,9 @@ typing-extensions==4.14.1 \
--hash=sha256:38b39f4aeeab64884ce9f74c94263ef78f3c22467c8724005483154c26648d36 \
--hash=sha256:d1e1e3b58374dc93031d6eda2420a48ea44a36c2b4766a4fdeb3710755731d76
# via pygithub
urllib3==2.5.0 \
--hash=sha256:3fc47733c7e419d4bc3f6b3dc2b4f890bb743906a30d56ba4a5bfa4bbff92760 \
--hash=sha256:e6b01673c0fa6a13e374b50871808eb3bf7046c4b125b216f6bf1cc604cff0dc
urllib3==2.6.3 \
--hash=sha256:1b62b6884944a57dbe321509ab94fd4d3b307075e0c2eae991ac71ee15ad38ed \
--hash=sha256:bf272323e553dfb2e87d9bfd225ca7b0f467b919d7bbd355436d3fd37cb0acd4
# via
# -r requirements_formatting.txt.in
# pygithub
+5 -3
View File
@@ -1,8 +1,10 @@
black~=25.1
black>=26.3.1
darker==2.1.1
PyGithub==2.6.1
cryptography>=43.0.1
urllib3>=2.5.0
cryptography>=46.0.5
urllib3>=2.6.3
requests>=2.32.4
idna>=3.7
certifi>=2024.7.4
PyNaCl>=1.6.2
PyJWT>=2.12.1
+1 -1
Submodule External/jemalloc deleted from ce24593018.
Submodule External/robin-map deleted from d5683d9f18.
Vendored Submodule
+1
Submodule External/rpmalloc added at 1f6fb494f2.
+3
View File
@@ -1,3 +1,6 @@
set(NAME tiny-json)
set(SRCS tiny-json.c)
add_library(${NAME} STATIC ${SRCS})
target_include_directories(${NAME} PUBLIC ${CMAKE_CURRENT_LIST_DIR})
add_library(${NAME}::${NAME} ALIAS ${NAME})
Vendored Submodule
+1
Submodule External/unordered_dense added at 3234af2c03.
+1 -1
+1 -1
Vendored Submodule
+1
Submodule External/zydis added at 9bfadd6a55.
+9 -42
View File
@@ -1,16 +1,16 @@
cmake_minimum_required(VERSION 3.14)
set (PROJECT_NAME FEXCore)
set(PROJECT_NAME FEXCore)
project(${PROJECT_NAME}
VERSION 0.01
LANGUAGES CXX)
if (CMAKE_SYSTEM_PROCESSOR MATCHES "x86_64")
set(_M_X86_64 1)
set(ARCHITECTURE_x86_64 1)
set(CMAKE_CXX_FLAGS "${CMAKE_CXX_FLAGS} -mcx16")
endif()
if (CMAKE_SYSTEM_PROCESSOR MATCHES "^aarch64|^arm64|^armv8\.*")
set(_M_ARM_64 1)
set(ARCHITECTURE_arm64 1)
endif()
set(CMAKE_POSITION_INDEPENDENT_CODE ON)
@@ -24,45 +24,10 @@ include(CheckCXXCompilerFlag)
include(CheckIncludeFileCXX)
include(CheckCXXSourceCompiles)
if (EXISTS ${CMAKE_CURRENT_DIR}/External/vixl/)
# Useful to have for freestanding libFEXCore
add_subdirectory(External/vixl/)
include_directories(External/vixl/src/)
endif()
set(CMAKE_CXX_STANDARD 20)
set(CMAKE_EXPORT_COMPILE_COMMANDS ON)
set(GIT_SHORT_HASH "Unknown")
set(GIT_DESCRIBE_STRING "FEX-Unknown")
if (OVERRIDE_VERSION STREQUAL "detect")
# Find our git hash
find_package(Git)
if (GIT_FOUND)
execute_process(
COMMAND ${GIT_EXECUTABLE} rev-parse --short=7 HEAD
WORKING_DIRECTORY "${CMAKE_SOURCE_DIR}"
OUTPUT_VARIABLE GIT_SHORT_HASH
ERROR_QUIET
OUTPUT_STRIP_TRAILING_WHITESPACE
)
execute_process(
COMMAND ${GIT_EXECUTABLE} describe --abbrev=7
WORKING_DIRECTORY "${CMAKE_SOURCE_DIR}"
OUTPUT_VARIABLE GIT_DESCRIBE_STRING
ERROR_QUIET
OUTPUT_STRIP_TRAILING_WHITESPACE
)
endif()
else()
set(GIT_SHORT_HASH "${OVERRIDE_VERSION}")
set(GIT_DESCRIBE_STRING "FEX-${OVERRIDE_VERSION}")
endif()
configure_file(
${CMAKE_CURRENT_SOURCE_DIR}/include/git_version.h.in
configure_file(${CMAKE_CURRENT_SOURCE_DIR}/include/git_version.h.in
${CMAKE_BINARY_DIR}/generated/git_version.h)
include_directories(${CMAKE_BINARY_DIR}/generated)
@@ -74,9 +39,11 @@ add_compile_options($<$<COMPILE_LANGUAGE:CXX>:-fno-strict-aliasing> $<$<COMPILE_
add_subdirectory(Source/)
install (DIRECTORY include/FEXCore ${CMAKE_BINARY_DIR}/include/FEXCore
DESTINATION include
COMPONENT Development)
if (NOT BUILD_STEAM_SUPPORT)
install (DIRECTORY include/FEXCore ${CMAKE_BINARY_DIR}/include/FEXCore
DESTINATION include
COMPONENT Development)
endif()
if (BUILD_TESTING)
add_subdirectory(unittests/)
+13 -4
View File
@@ -156,7 +156,7 @@ def print_man_environment_tail():
"APP_CONFIG_LOCATION",
[
"Allows the user to override where FEX looks for configuration files",
"By default FEX will look in {$HOME, $XDG_CONFIG_HOME}/.fex-emu/",
"By default FEX will look in ${XDG_CONFIG_HOME, $HOME/.config}/fex-emu/",
"This will override the full path",
"If FEX_PORTABLE is declared then relative paths are also supported",
"For FEX: Relative to the FEX binary",
@@ -168,7 +168,7 @@ def print_man_environment_tail():
"APP_CONFIG",
[
"Allows the user to override where FEX looks for only the application config file",
"By default FEX will look in {$HOME, $XDG_CONFIG_HOME}/.fex-emu/Config.json",
"By default FEX will look in ${XDG_CONFIG_HOME, $HOME/.config}/fex-emu/Config.json",
"This will override this file location",
"One must be careful with this option as it will override any applications that load with execve as well"
"If you need to support applications that execve then use FEX_APP_CONFIG_LOCATION instead"
@@ -182,7 +182,7 @@ def print_man_environment_tail():
"APP_DATA_LOCATION",
[
"Allows the user to override where FEX looks for data files",
"By default FEX will look in {$HOME, $XDG_DATA_HOME}/.fex-emu/",
"By default FEX will look in {$XDG_DATA_HOME, $HOME/.local/share}/fex-emu/",
"This will override the full path",
"This is the folder where FEX stores generated files like IR cache"
],
@@ -200,6 +200,15 @@ def print_man_environment_tail():
],
"''", True)
print_man_env_option(
"APP_CACHE_LOCATION",
[
"Allows the user to override where FEX stores and loads cache files",
"By default FEX will look in ${XDG_CACHE_HOME, $HOME/.cache}/fex-emu/",
"This will override the full path, trailing forward-slash is expected to exist",
],
"''", True)
def print_man_header():
header ='''.Dd {0}
.Dt FEX
@@ -225,7 +234,7 @@ FEX is very much work in progress, so expect things to change.
def print_man_tail():
tail ='''.Sh FILES
.Bl -tag -width "$prefix/share/fex-emu/GuestThunks" -compact
.It Pa $XDG_HOME_DIR/.fex-emu
.It Pa $XDG_CONFIG_DIR/fex-emu
Default FEX user configuration directory
.It Pa $prefix/share/fex-emu/AppConfig
System level application configuration files
+2
View File
@@ -678,6 +678,7 @@ def print_ir_allocator_helpers():
# Generate helpers with operands
for op in IROps:
if op.Name != "Last":
output_file.write("\t///\n".join(["\t/// {}\n" .format(comment) for comment in op.Desc]))
output_file.write("\tIRPair<IROp_{}> _{}(" .format(op.Name, op.Name))
# Output SSA args first
@@ -751,6 +752,7 @@ def print_ir_allocator_helpers():
# Now do the OrderedNode * version if necessary
if op.SSAArgNum:
output_file.write("\t///\n".join(["\t/// {}\n" .format(comment) for comment in op.Desc]))
output_file.write("\tIRPair<IROp_{}> _{}(" .format(op.Name, op.Name))
for i, arg in enumerate(op.Arguments):
+74 -77
View File
@@ -1,20 +1,20 @@
set (MAN_DIR share/man CACHE PATH "MAN_DIR")
set(MAN_DIR share/man CACHE PATH "MAN_DIR")
set (FEXCORE_BASE_SRCS
set(FEXCORE_BASE_SRCS
Interface/Config/Config.cpp
Utils/Allocator.cpp
Utils/FileLoading.cpp
Utils/ForcedAssert.cpp
Utils/LogManager.cpp
Utils/SpinWaitLock.cpp
)
Utils/WildcardMatcher.cpp)
if (NOT MINGW_BUILD)
if (NOT MINGW)
list(APPEND FEXCORE_BASE_SRCS
Utils/Allocator/64BitAllocator.cpp)
endif()
set (SRCS
set(SRCS
Common/JitSymbols.cpp
Interface/Context/Context.cpp
Interface/Core/LookupCache.cpp
@@ -31,7 +31,6 @@ set (SRCS
Interface/Core/OpcodeDispatcher/X87.cpp
Interface/Core/OpcodeDispatcher/X87F64.cpp
Interface/Core/OpcodeDispatcher.cpp
Interface/Core/X86HelperGen.cpp
Interface/Core/ArchHelpers/Arm64Emitter.cpp
Interface/Core/Dispatcher/Dispatcher.cpp
Interface/Core/Interpreter/Fallbacks/InterpreterFallbacks.cpp
@@ -69,10 +68,9 @@ set (SRCS
Utils/LongJump.cpp
Utils/Telemetry.cpp
Utils/Threads.cpp
Utils/Profiler.cpp
)
Utils/Profiler.cpp)
if (_M_ARM_64)
if (ARCHITECTURE_arm64)
list(APPEND SRCS Utils/ArchHelpers/Arm64.cpp)
else()
list(APPEND SRCS Utils/ArchHelpers/Arm64_stubs.cpp)
@@ -85,42 +83,54 @@ endif()
set(DEFINES -DJIT_ARM64)
if (_M_X86_64)
list(APPEND DEFINES -D_M_X86_64=1)
if (ARCHITECTURE_x86_64)
list(APPEND DEFINES -DARCHITECTURE_x86_64=1)
endif()
if (_M_ARM_64)
list(APPEND DEFINES -D_M_ARM_64=1)
if (ARCHITECTURE_arm64)
list(APPEND DEFINES -DARCHITECTURE_arm64=1)
endif()
if (ENABLE_VIXL_DISASSEMBLER)
list(APPEND DEFINES -DVIXL_DISASSEMBLER=1)
endif()
if (_M_ARM_64 AND HAS_CLANG_PRESERVE_ALL)
if (ENABLE_ZYDIS)
list(APPEND DEFINES -DZYDIS_DISASSEMBLER=1)
endif()
if (ARCHITECTURE_arm64 AND HAS_CLANG_PRESERVE_ALL)
list(APPEND DEFINES "-DFEXCORE_PRESERVE_ALL_ATTR=__attribute__((preserve_all));-DFEXCORE_HAS_PRESERVE_ALL_ATTR=1")
else()
list(APPEND DEFINES "-DFEXCORE_PRESERVE_ALL_ATTR=;-DFEXCORE_HAS_PRESERVE_ALL_ATTR=0")
endif()
set (LIBS fmt::fmt xxHash::xxhash FEXHeaderUtils CodeEmitter cephes_128bit)
set(LIBS fmt::fmt xxHash::xxhash FEXHeaderUtils CodeEmitter cephes_128bit)
if (ENABLE_VIXL_DISASSEMBLER OR ENABLE_VIXL_SIMULATOR)
list (APPEND LIBS vixl)
list(APPEND LIBS vixl::vixl)
endif()
if (NOT MINGW_BUILD)
list (APPEND LIBS dl)
if (ENABLE_ZYDIS)
list(APPEND LIBS Zydis::Zydis)
endif()
if (NOT MINGW)
list(APPEND LIBS dl)
else()
list (APPEND LIBS synchronization)
if (_M_ARM_64EC)
list (APPEND LIBS mincore)
list(APPEND LIBS synchronization)
if (ARCHITECTURE_arm64ec)
list(APPEND LIBS mincore)
endif()
endif()
if (CMAKE_CXX_COMPILER_ID STREQUAL "GNU")
# GCC requires libatomic to use 128-bit atomics
list(APPEND LIBS atomic)
endif()
# Generate config
configure_file(
${CMAKE_CURRENT_SOURCE_DIR}/Interface/Config/Config.json.in
configure_file(${CMAKE_CURRENT_SOURCE_DIR}/Interface/Config/Config.json.in
${CMAKE_BINARY_DIR}/generated/Config/Config.json)
# Generate IR include file
@@ -135,11 +145,10 @@ add_custom_command(
OUTPUT "${OUTPUT_NAME}" "${OUTPUT_DISPATCHER_NAME}"
DEPENDS "${INPUT_NAME}"
DEPENDS "${CMAKE_CURRENT_SOURCE_DIR}/../Scripts/json_ir_generator.py"
COMMAND "python3" "${CMAKE_CURRENT_SOURCE_DIR}/../Scripts/json_ir_generator.py" "${INPUT_NAME}" "${OUTPUT_NAME}" "${OUTPUT_DISPATCHER_NAME}"
)
COMMAND "python3" "${CMAKE_CURRENT_SOURCE_DIR}/../Scripts/json_ir_generator.py"
"${INPUT_NAME}" "${OUTPUT_NAME}" "${OUTPUT_DISPATCHER_NAME}")
set_source_files_properties(${OUTPUT_NAME} PROPERTIES
GENERATED TRUE)
set_source_files_properties(${OUTPUT_NAME} PROPERTIES GENERATED TRUE)
# Generate IR documentation
set(OUTPUT_IR_DOC "${CMAKE_BINARY_DIR}/IR.md")
@@ -148,11 +157,10 @@ add_custom_command(
OUTPUT "${OUTPUT_IR_DOC}"
DEPENDS "${INPUT_NAME}"
DEPENDS "${CMAKE_CURRENT_SOURCE_DIR}/../Scripts/json_ir_doc_generator.py"
COMMAND "python3" "${CMAKE_CURRENT_SOURCE_DIR}/../Scripts/json_ir_doc_generator.py" "${INPUT_NAME}" "${OUTPUT_IR_DOC}"
)
COMMAND "python3" "${CMAKE_CURRENT_SOURCE_DIR}/../Scripts/json_ir_doc_generator.py"
"${INPUT_NAME}" "${OUTPUT_IR_DOC}")
set_source_files_properties(${OUTPUT_IR_NAME} PROPERTIES
GENERATED TRUE)
set_source_files_properties(${OUTPUT_IR_NAME} PROPERTIES GENERATED TRUE)
# Create the target
add_custom_target(IR_INC
@@ -176,14 +184,12 @@ add_custom_command(
DEPENDS "${INPUT_CONFIG_NAME}"
DEPENDS "${CMAKE_CURRENT_SOURCE_DIR}/../Scripts/config_generator.py"
COMMAND "python3" "${CMAKE_CURRENT_SOURCE_DIR}/../Scripts/config_generator.py" "${INPUT_CONFIG_NAME}" "${OUTPUT_CONFIG_NAME}" "${OUTPUT_MAN_NAME}"
"${OUTPUT_CONFIG_OPTION_NAME}"
)
"${OUTPUT_CONFIG_OPTION_NAME}")
add_custom_command(
OUTPUT "${OUTPUT_MAN_NAME_COMPRESS}"
DEPENDS "${OUTPUT_MAN_NAME}"
COMMAND "gzip" "-kf9n" "${OUTPUT_MAN_NAME}"
)
COMMAND "gzip" "-kf9n" "${OUTPUT_MAN_NAME}")
set_source_files_properties(${OUTPUT_CONFIG_NAME} PROPERTIES
GENERATED TRUE)
@@ -202,8 +208,10 @@ add_custom_target(CONFIG_INC
DEPENDS "${OUTPUT_MAN_NAME}"
DEPENDS "${OUTPUT_MAN_NAME_COMPRESS}")
# Install the compressed man page
install(FILES ${OUTPUT_MAN_NAME_COMPRESS} COMPONENT Runtime DESTINATION ${MAN_DIR}/man1)
if (NOT BUILD_STEAM_SUPPORT)
# Install the compressed man page
install(FILES ${OUTPUT_MAN_NAME_COMPRESS} COMPONENT Runtime DESTINATION ${MAN_DIR}/man1)
endif()
# Add in diagnostic colours if the option is available.
# Ninja code generator will kill colours if this isn't here
@@ -225,8 +233,7 @@ function(AddDefaultOptionsToTarget Name)
target_compile_definitions(${Name} PRIVATE ${DEFINES})
add_dependencies(${Name} CONFIG_INC IR_INC)
target_compile_options(${Name}
PRIVATE
target_compile_options(${Name} PRIVATE
-Wall
-Werror=cast-qual
-Werror=ignored-qualifiers
@@ -234,82 +241,72 @@ function(AddDefaultOptionsToTarget Name)
-Wno-trigraphs
-ffunction-sections
-fwrapv
)
-fwrapv)
if (GCC_COLOR)
target_compile_options(${Name}
PRIVATE
"-fdiagnostics-color=always")
endif()
if (CLANG_COLOR)
target_compile_options(${Name}
PRIVATE
"-fcolor-diagnostics")
target_compile_options(${Name} PRIVATE "-fdiagnostics-color=always")
endif()
if (CMAKE_BUILD_TYPE MATCHES "RELEASE")
target_link_options(${Name}
PRIVATE
"LINKER:--gc-sections"
"LINKER:--strip-all"
"LINKER:--as-needed"
)
if (CLANG_COLOR)
target_compile_options(${Name} PRIVATE "-fcolor-diagnostics")
endif()
LinkerGC(${Name})
target_link_libraries(${Name} PUBLIC unordered_dense::unordered_dense)
endfunction()
# Build FEXCore_Config static library
# Build FEXCore_Base static library
add_library(FEXCore_Base STATIC ${FEXCORE_BASE_SRCS})
target_link_libraries(FEXCore_Base ${LIBS})
target_link_libraries(FEXCore_Base PUBLIC ${LIBS})
AddDefaultOptionsToTarget(FEXCore_Base)
if (ENABLE_FEXCORE_PROFILER AND FEXCORE_PROFILER_BACKEND STREQUAL "TRACY")
target_link_libraries(FEXCore_Base TracyClient)
target_link_libraries(FEXCore_Base PUBLIC TracyClient)
endif()
function(AddObject Name Type)
add_library(${Name} ${Type} ${SRCS})
function(AddObject Name)
add_library(${Name} OBJECT ${SRCS})
target_link_libraries(${Name} FEXCore_Base)
target_link_libraries(${Name} PRIVATE FEXCore_Base)
target_compile_options(${Name} PRIVATE ${FEX_TUNE_COMPILE_FLAGS})
AddDefaultOptionsToTarget(${Name})
set_target_properties(${Name} PROPERTIES OUTPUT_NAME FEXCore)
endfunction()
function(AddLibrary Name Type)
add_library(${Name} ${Type} $<TARGET_OBJECTS:${PROJECT_NAME}_object>)
target_link_libraries(${Name} FEXCore_Base)
target_compile_options(${Name} PRIVATE ${FEX_TUNE_COMPILE_FLAGS})
set_target_properties(${Name} PROPERTIES OUTPUT_NAME FEXCore)
# During generation of the import library (dll.a), MinGW needs some extra symbols from libraries
# such as fmt, which are propagated by FEXCore_Base. Wonderful.
if (MINGW)
target_link_libraries(${Name} PRIVATE FEXCore_Base)
endif()
AddDefaultOptionsToTarget(${Name})
endfunction()
AddObject(${PROJECT_NAME}_object OBJECT)
AddObject(${PROJECT_NAME}_object)
AddLibrary(${PROJECT_NAME} STATIC)
AddLibrary(${PROJECT_NAME}_shared SHARED)
if (NOT MINGW_BUILD)
install(TARGETS ${PROJECT_NAME}_shared
LIBRARY
DESTINATION ${CMAKE_INSTALL_LIBDIR}
COMPONENT Libraries)
if (NOT MINGW AND NOT BUILD_STEAM_SUPPORT)
install(TARGETS ${PROJECT_NAME}_shared LIBRARY
DESTINATION ${CMAKE_INSTALL_LIBDIR}
COMPONENT Libraries)
endif()
# Meta-library to link jemalloc libraries enabled in the build configuration.
# Only needed for targets that run emulation. For others, use JemallocDummy.
add_library(JemallocLibs STATIC Utils/AllocatorHooks.cpp)
if (ENABLE_JEMALLOC)
target_compile_definitions(JemallocLibs PRIVATE ENABLE_JEMALLOC=1 JEMALLOC_NO_RENAME=1)
target_link_libraries(JemallocLibs PUBLIC FEX_jemalloc)
if (ENABLE_FEX_ALLOCATOR)
target_compile_definitions(JemallocLibs PRIVATE ENABLE_FEX_ALLOCATOR=1)
target_link_libraries(JemallocLibs PUBLIC rpmalloc)
endif()
if (ENABLE_JEMALLOC_GLIBC_ALLOC)
set_source_files_properties(Interface/HLE/Thunks/Thunks.cpp PROPERTIES COMPILE_DEFINITIONS ENABLE_JEMALLOC_GLIBC=1)
target_link_libraries(JemallocLibs INTERFACE FEX_jemalloc_glibc)
endif()
if (NOT MINGW_BUILD)
if (NOT MINGW)
# Dummy project to use for host tools.
# This overrides use of jemalloc in FEXCore with the normal glibc allocator.
add_library(JemallocDummy STATIC Utils/AllocatorHooks.cpp)
@@ -317,4 +314,4 @@ if (NOT MINGW_BUILD)
endif()
# The shared library should always link enabled jemalloc libraries
target_link_libraries(${PROJECT_NAME}_shared JemallocLibs)
target_link_libraries(${PROJECT_NAME}_shared PRIVATE JemallocLibs)
+49 -33
View File
@@ -19,7 +19,7 @@ extern "C" {
}
struct FEX_PACKED X80SoftFloat {
#ifdef _M_X86_64
#ifdef ARCHITECTURE_x86_64
// Define this to push some operations to x87
// Only useful to see if precision loss is killing something
// #define DEBUG_X86_FLOAT
@@ -30,29 +30,33 @@ struct FEX_PACKED X80SoftFloat {
#define BIGFLOAT float128_t
#define BIGFLOATSIZE 16
#endif
#elif defined(_M_ARM_64)
#elif defined(ARCHITECTURE_arm64)
#define BIGFLOAT float128_t
#define BIGFLOATSIZE 16
#else
#error No 128bit float for this target!
#endif
uint64_t Significand : 64;
uint16_t Exponent : 15;
uint16_t Sign : 1;
uint64_t Significand;
union {
uint16_t Raw;
struct {
uint16_t Exponent : 15;
uint16_t Sign : 1;
};
} Top;
X80SoftFloat() {
memset(this, 0, sizeof(*this));
}
X80SoftFloat(uint16_t _Sign, uint16_t _Exponent, uint64_t _Significand)
: Significand {_Significand}
, Exponent {_Exponent}
, Sign {_Sign} {}
, Top {.Raw = static_cast<uint16_t>((_Exponent & 0x7FFF) | (_Sign << 15))} {}
fextl::string str() const {
fextl::ostringstream string;
string << std::hex << Sign;
string << "_" << Exponent;
string << std::hex << Top.Sign;
string << "_" << Top.Exponent;
string << "_" << (Significand >> 63);
string << "_" << (Significand & ((1ULL << 63) - 1));
return string.str();
@@ -163,18 +167,18 @@ struct FEX_PACKED X80SoftFloat {
X80SoftFloat result = 0;
if (HandleInfinityOp(state, lhs, result)) {
return result;
} else if (lhs.Exponent == 0x7FFF && (lhs.Significand & 0x7FFFFFFFFFFFFFFFULL)) { // NaN
} else if (lhs.Top.Exponent == 0x7FFF && (lhs.Significand & 0x7FFFFFFFFFFFFFFFULL)) { // NaN
// propagate NaN
state->exceptionFlags |= softfloat_flag_invalid;
return lhs;
}
// Check for zero divisor - fprem(x, 0) is invalid operation
if (rhs.Exponent == 0 && rhs.Significand == 0) {
if (rhs.Top.Exponent == 0 && rhs.Significand == 0) {
state->exceptionFlags |= softfloat_flag_invalid;
// Return QNaN
result.Sign = 0;
result.Exponent = 0x7FFF;
result.Top.Sign = 0;
result.Top.Exponent = 0x7FFF;
result.Significand = 0xC000000000000000ULL;
return result;
}
@@ -253,12 +257,16 @@ struct FEX_PACKED X80SoftFloat {
return Result;
#else
// Zero is a special case, the significand for +/- 0 is +/- zero.
if (lhs.Exponent == 0x0 && lhs.Significand == 0x0) {
if (lhs.Top.Exponent == 0x0 && lhs.Significand == 0x0) {
return lhs;
}
// Inf/NaN pass through unchanged in the significand slot.
if (lhs.Top.Exponent == 0x7FFF) {
return lhs;
}
X80SoftFloat Tmp = lhs;
Tmp.Exponent = 0x3FFF;
Tmp.Sign = lhs.Sign;
Tmp.Top.Exponent = 0x3FFF;
Tmp.Top.Sign = lhs.Top.Sign;
return Tmp;
#endif
}
@@ -280,12 +288,20 @@ struct FEX_PACKED X80SoftFloat {
return Result;
#else
// Zero is a special case, the exponent is always -inf
if (lhs.Exponent == 0x0 && lhs.Significand == 0x0) {
if (lhs.Top.Exponent == 0x0 && lhs.Significand == 0x0) {
X80SoftFloat Result(1, 0x7FFFUL, 0x8000'0000'0000'0000UL);
return Result;
}
// +/-Inf returns +Inf in the exponent slot; NaN propagates.
if (lhs.Top.Exponent == 0x7FFF) {
if ((lhs.Significand & 0x7FFFFFFFFFFFFFFFULL) == 0) {
X80SoftFloat Result(0, 0x7FFFUL, 0x8000'0000'0000'0000UL);
return Result;
}
return lhs;
}
int32_t TrueExp = lhs.Exponent - ExponentBias;
int32_t TrueExp = lhs.Top.Exponent - ExponentBias;
return i32_to_extF80(TrueExp);
#endif
}
@@ -320,6 +336,13 @@ struct FEX_PACKED X80SoftFloat {
#else
extFloat80_t Zero {0, 0};
if (extF80_eq(state, lhs, Zero)) {
// FSCALE(0, +Inf) is 0 * Inf, which is invalid. FSCALE(0, anything
// else) is still 0.
if (rhs.Top.Exponent == 0x7FFF && rhs.Top.Sign == 0 && (rhs.Significand & 0x7FFFFFFFFFFFFFFFULL) == 0) {
state->exceptionFlags |= softfloat_flag_invalid;
X80SoftFloat QNaN(0, 0x7FFFUL, 0xC000000000000000ULL);
return QNaN;
}
return lhs;
}
X80SoftFloat Int = FRNDINT(state, rhs, softfloat_round_minMag);
@@ -572,8 +595,7 @@ struct FEX_PACKED X80SoftFloat {
X80SoftFloat(extFloat80_t rhs) {
Significand = rhs.signif;
Exponent = rhs.signExp & 0x7FFF;
Sign = rhs.signExp >> 15;
Top.Raw = rhs.signExp;
}
X80SoftFloat(softfloat_state* state, const float rhs) {
@@ -606,8 +628,7 @@ struct FEX_PACKED X80SoftFloat {
void operator=(extFloat80_t rhs) {
Significand = rhs.signif;
Exponent = rhs.signExp & 0x7FFF;
Sign = rhs.signExp >> 15;
Top.Raw = rhs.signExp;
}
operator FEXCore::VectorRegType() const {
@@ -617,16 +638,16 @@ struct FEX_PACKED X80SoftFloat {
operator extFloat80_t() const {
extFloat80_t Result {};
Result.signif = Significand;
Result.signExp = Exponent | (Sign << 15);
Result.signExp = Top.Raw;
return Result;
}
static bool IsNan(const X80SoftFloat& lhs) {
return (lhs.Exponent == 0x7FFF) && (lhs.Significand & IntegerBit) && (lhs.Significand & Bottom62Significand);
return (lhs.Top.Exponent == 0x7FFF) && (lhs.Significand & IntegerBit) && (lhs.Significand & Bottom62Significand);
}
static bool SignBit(const X80SoftFloat& lhs) {
return lhs.Sign;
return lhs.Top.Sign;
}
private:
@@ -637,11 +658,11 @@ private:
// Helper function to check for infinity and set invalid operation flag.
// Returns true if infinity is dealt with, false otherwise.
FEXCORE_PRESERVE_ALL_ATTR static bool HandleInfinityOp(softfloat_state* state, const X80SoftFloat& arg, X80SoftFloat& result) {
if (arg.Exponent == 0x7FFF && arg.Significand == 0x8000000000000000ULL) {
if (arg.Top.Exponent == 0x7FFF && arg.Significand == 0x8000000000000000ULL) {
state->exceptionFlags |= softfloat_flag_invalid;
// Return QNaN.
result.Sign = 0;
result.Exponent = 0x7FFF;
result.Top.Sign = 0;
result.Top.Exponent = 0x7FFF;
result.Significand = 0xC000000000000000ULL;
return true;
}
@@ -649,9 +670,4 @@ private:
}
};
#ifndef _WIN32
static_assert(sizeof(X80SoftFloat) == 10, "tword must be 10bytes in size");
#else
// Padding on this extends to 16-bytes rather than 10-bytes on WIN32.
static_assert(sizeof(X80SoftFloat) == 16, "tword must be 16bytes in size");
#endif
+7 -3
View File
@@ -1,7 +1,7 @@
// SPDX-License-Identifier: MIT
#pragma once
#ifdef _M_X86_64
#ifdef ARCHITECTURE_x86_64
#include <xmmintrin.h>
#include <immintrin.h>
#else
@@ -13,10 +13,14 @@ struct VectorScalarF64Pair {
double val[2];
};
#ifdef _M_ARM_64
#ifdef ARCHITECTURE_arm64
// Can't use uint8x16_t directly from arm_neon.h here.
// Overrides softfloat-3e's defines which causes problems.
#ifdef __clang__
using VectorRegType = __attribute__((neon_vector_type(16))) uint8_t;
#else
using VectorRegType = __attribute__((vector_size(16))) uint8_t;
#endif
struct VectorRegPairType {
VectorRegType val[2];
};
@@ -25,7 +29,7 @@ static inline VectorRegPairType MakeVectorRegPair(VectorRegType low, VectorRegTy
return VectorRegPairType {low, high};
}
#elif defined(_M_X86_64)
#elif defined(ARCHITECTURE_x86_64)
using VectorRegType = __m128i;
using VectorRegPairType = __m256i;
+23 -22
View File
@@ -30,14 +30,14 @@ class Context;
}
namespace FEXCore::Config {
namespace DefaultValues {
namespace detail {
#define P(x) x
#define OPT_BASE(type, group, enum, json, default) const P(type) P(enum) = P(default);
#define OPT_STR(group, enum, json, default) const std::string_view P(enum) = P(default);
#define OPT_STRARRAY(group, enum, json, default) OPT_STR(group, enum, json, default)
#define OPT_STRENUM(group, enum, json, default) const uint64_t P(enum) = FEXCore::ToUnderlying(P(default));
#include <FEXCore/Config/ConfigValues.inl>
} // namespace DefaultValues
} // namespace detail
enum Paths {
PATH_DATA_DIR_LOCAL = 0,
@@ -134,7 +134,7 @@ public:
void Load();
template<typename T>
requires (!std::is_same_v<fextl::string, T> && !std::is_same_v<DefaultValues::Type::StringArrayType, T>)
requires (!std::is_same_v<fextl::string, T> && !std::is_same_v<StringArrayType, T>)
std::optional<T> GetConv(ConfigOption Option) {
const auto it = OptionMap.find(Option);
if (it == OptionMap.end()) {
@@ -142,7 +142,7 @@ public:
}
const auto& Value = it->second;
LOGMAN_THROW_A_FMT(!std::holds_alternative<DefaultValues::Type::StringArrayType>(Value), "Tried to get config of invalid type!");
LOGMAN_THROW_A_FMT(!std::holds_alternative<StringArrayType>(Value), "Tried to get config of invalid type!");
if (std::holds_alternative<T>(Value)) [[likely]] {
return std::get<T>(Value);
@@ -165,7 +165,7 @@ public:
private:
void MergeConfigMap(const LayerOptions& Options);
void MergeEnvironmentVariables(const ConfigOption& Option, const DefaultValues::Type::StringArrayType& Value);
void MergeEnvironmentVariables(const ConfigOption& Option, const StringArrayType& Value);
};
void MetaLayer::Load() {
@@ -181,7 +181,7 @@ void MetaLayer::Load() {
}
void MetaLayer::MergeEnvironmentVariables(const ConfigOption& Option, const DefaultValues::Type::StringArrayType& Value) {
void MetaLayer::MergeEnvironmentVariables(const ConfigOption& Option, const StringArrayType& Value) {
// Environment variables need a bit of additional work
// We want to merge the arrays rather than overwrite entirely
auto MetaEnvironment = OptionMap.find(Option);
@@ -193,7 +193,7 @@ void MetaLayer::MergeEnvironmentVariables(const ConfigOption& Option, const Defa
// If an environment variable exists in both current meta and in the incoming layer then the meta layer value is overwritten
fextl::unordered_map<fextl::string, fextl::string> LookupMap;
const auto AddToMap = [&LookupMap](const DefaultValues::Type::StringArrayType& Value) {
const auto AddToMap = [&LookupMap](const StringArrayType& Value) {
for (const auto& EnvVar : Value) {
const auto ItEq = EnvVar.find_first_of('=');
if (ItEq == fextl::string::npos) {
@@ -209,7 +209,7 @@ void MetaLayer::MergeEnvironmentVariables(const ConfigOption& Option, const Defa
}
};
AddToMap(std::get<DefaultValues::Type::StringArrayType>(MetaEnvironment->second));
AddToMap(std::get<StringArrayType>(MetaEnvironment->second));
AddToMap(Value);
// Now with the two layers merged in the map
@@ -225,8 +225,8 @@ void MetaLayer::MergeConfigMap(const LayerOptions& Options) {
// Insert this layer's options, overlaying previous options that exist here
for (auto& it : Options) {
if (it.first == FEXCore::Config::ConfigOption::CONFIG_ENV || it.first == FEXCore::Config::ConfigOption::CONFIG_HOSTENV) {
LOGMAN_THROW_A_FMT(std::holds_alternative<DefaultValues::Type::StringArrayType>(it.second), "Tried to get config of invalid type!");
MergeEnvironmentVariables(it.first, std::get<DefaultValues::Type::StringArrayType>(it.second));
LOGMAN_THROW_A_FMT(std::holds_alternative<StringArrayType>(it.second), "Tried to get config of invalid type!");
MergeEnvironmentVariables(it.first, std::get<StringArrayType>(it.second));
} else {
OptionMap.insert_or_assign(it.first, it.second);
}
@@ -307,12 +307,10 @@ constexpr char ContainerManager[] = "/run/host/container-manager";
fextl::string FindContainer() {
// We only support pressure-vessel at the moment
if (FHU::Filesystem::Exists(ContainerManager)) {
fextl::vector<char> Manager {};
fextl::string Manager {};
if (FEXCore::FileLoading::LoadFile(Manager, ContainerManager)) {
// Trim the whitespace, may contain a newline
fextl::string ManagerStr = Manager.data();
ManagerStr = FEXCore::StringUtils::Trim(ManagerStr);
return ManagerStr;
return FEXCore::StringUtils::Trim(Manager);
}
}
return {};
@@ -321,12 +319,10 @@ fextl::string FindContainer() {
fextl::string FindContainerPrefix() {
// We only support pressure-vessel at the moment
if (FHU::Filesystem::Exists(ContainerManager)) {
fextl::vector<char> Manager {};
fextl::string Manager {};
if (FEXCore::FileLoading::LoadFile(Manager, ContainerManager)) {
// Trim the whitespace, may contain a newline
fextl::string ManagerStr = Manager.data();
ManagerStr = FEXCore::StringUtils::Trim(ManagerStr);
if (strncmp(ManagerStr.data(), "pressure-vessel", Manager.size()) == 0) {
if (FEXCore::StringUtils::Trim(Manager) == "pressure-vessel") {
// We are running inside of pressure vessel
// Our $CMAKE_INSTALL_PREFIX paths are now inside of /run/host/$CMAKE_INSTALL_PREFIX
return "/run/host/";
@@ -423,7 +419,7 @@ bool Exists(ConfigOption Option) {
return Meta->OptionExists(Option);
}
std::optional<DefaultValues::Type::StringArrayType*> All(ConfigOption Option) {
std::optional<StringArrayType*> All(ConfigOption Option) {
return Meta->All(Option);
}
@@ -436,6 +432,12 @@ std::optional<T> GetConv(ConfigOption Option) {
return Meta->GetConv<T>(Option);
}
template std::optional<bool> GetConv(ConfigOption Option);
template std::optional<uint8_t> GetConv(ConfigOption Option);
template std::optional<int32_t> GetConv(ConfigOption Option);
template std::optional<uint32_t> GetConv(ConfigOption Option);
template std::optional<uint64_t> GetConv(ConfigOption Option);
void Set(ConfigOption Option, std::string_view Data) {
Meta->Set(Option, Data);
}
@@ -491,13 +493,12 @@ template Value<uint8_t>::Value(FEXCore::Config::ConfigOption _Option, uint8_t De
template Value<uint64_t>::Value(FEXCore::Config::ConfigOption _Option, uint64_t Default);
template<typename T>
void Value<T>::GetListIfExists(FEXCore::Config::ConfigOption Option, DefaultValues::Type::StringArrayType* List) {
void Value<T>::GetListIfExists(FEXCore::Config::ConfigOption Option, StringArrayType* List) {
auto Value = FEXCore::Config::All(Option);
List->clear();
if (Value) {
*List = **Value;
}
}
template void Value<DefaultValues::Type::StringArrayType>::GetListIfExists(FEXCore::Config::ConfigOption Option,
DefaultValues::Type::StringArrayType* List);
template void Value<StringArrayType>::GetListIfExists(FEXCore::Config::ConfigOption Option, StringArrayType* List);
} // namespace FEXCore::Config
+59 -13
View File
@@ -16,6 +16,20 @@
"Maximum number of instruction to store in a block"
]
},
"EnableCodeCachingWIP": {
"Type": "bool",
"Default": "false",
"Desc": [
"Enable the code caching subsystem"
]
},
"EnableCodeCacheValidation": {
"Type": "bool",
"Default": "false",
"Desc": [
"Enable expensive validation when loading code caches"
]
},
"HostFeatures": {
"Type": "strenum",
"Default": "FEXCore::Config::HostFeatures::OFF",
@@ -61,7 +75,9 @@
"ENABLE3DNOW": "enable3dnow",
"DISABLE3DNOW": "disable3dnow",
"ENABLESSE4A": "enablesse4a",
"DISABLESSE4A": "disablesse4a"
"DISABLESSE4A": "disablesse4a",
"ENABLEMOPS": "enablemops",
"DISABLEMOPS": "disablemops"
},
"Desc": [
"Allows controlling of the CPU features in the JIT.",
@@ -85,7 +101,8 @@
"\t{enable,disable}preserveallabi: Will force enable or disable preserve_all abi even if the host doesn't support it",
"\t{enable,disable}wfxt: Will force enable or disable wfxt even if the host doesn't support it",
"\t{enable,disable}3dnow: Will force enable or disable 3DNow! even if the host doesn't support it",
"\t{enable,disable}sse4a: Will force enable or disable SSE4a even if the host doesn't support it"
"\t{enable,disable}sse4a: Will force enable or disable SSE4a even if the host doesn't support it",
"\t{enable,disable}mops: Will force enable or disable FEAT_MOPS even if the host doesn't support it"
]
},
"SmallTSCScale": {
@@ -94,6 +111,20 @@
"Desc": [
"Scales the cycle counter on systems that have low frequencies."
]
},
"HideHybrid": {
"Type": "bool",
"Default": "true",
"Desc": [
"Hides hybrid CPU core arrangement."
]
},
"CPUFeatureRegisters": {
"Type": "str",
"Default": "",
"Desc": [
"Allows overriding cpu feature flags for manual testing"
]
}
},
"Emulation": {
@@ -106,9 +137,9 @@
"\teg: ~/RootFS/Debian_x86_64",
"Or this can be a name of a rootfs",
"If the named rootfs exists in the FEX data folder then it will use that one",
"\teg: $HOME/.fex-emu/RootFS/<RootFS name>/",
"Or if you have XDG_DATA_HOME the config will search in that directory",
"\teg: $XDG_DATA_HOME/.fex-emu/RootFS/<RootFS name>/"
"\teg: $XDG_DATA_HOME/fex-emu/RootFS/<RootFS name>/",
"If XDG_DATA_HOME is unset, ~/.local/share will be used in its place.",
"\teg: $HOME/.local/share/fex-emu/RootFS/<RootFS name>/"
]
},
"ThunkHostLibs": {
@@ -134,9 +165,9 @@
"\teg: ~/MyThunkConfig.json",
"Or this can be a named of a Thunk config file",
"If the named config file exists in the FEX data folder folder the it will use that one",
"\teg: $HOME/.fex-emu/ThunkConfigs/<ThunkConfig name>",
"Or if you have XDG_DATA_HOME the config will search in that directory",
"\teg: $XDG_DATA_HOME/.fex-emu/ThunkConfigs/<ThunkConfig name>"
"\teg: $XDG_DATA_HOME/fex-emu/ThunkConfigs/<ThunkConfig name>",
"If XDG_DATA_HOME is unset, ~/.local/share will be used in its place.",
"\teg: $HOME/.local/share/fex-emu/ThunkConfigs/<ThunkConfig name>"
]
},
"Env": {
@@ -164,7 +195,7 @@
},
"DisableL2Cache": {
"Type": "bool",
"Default": "false",
"Default": "true",
"Desc": [
"Disables FEXCore's JIT L2 cache lookup. Saving memory.",
"Can potentially introduce more stutters."
@@ -172,7 +203,7 @@
},
"DynamicL1Cache": {
"Type": "bool",
"Default": "false",
"Default": "true",
"Desc": [
"Switches FEXCore's JIT L1 cache to be dynamically sized. Saving memory.",
"Can potentially introduce more stutters."
@@ -313,13 +344,21 @@
"STATS": "stats"
},
"Desc": [
"Allows controlling of the vixl disassembler.",
"Allows controlling of the vixl disassembler for generated ARM code.",
"\toff: No disassembly will be output",
"\tdispatcher: Will enable disassembly of the JIT dispatcher loop",
"\tblocks: Will enable disassembly of the translated instruction code blocks",
"\tstats: Will print stats when disassembling the code"
]
},
"X86Disassemble": {
"Type": "bool",
"Default": "false",
"Desc": [
"Enables x86/x86-64 guest disassembly output for compiled blocks.",
"Requires FEX to be built with -DENABLE_ZYDIS=TRUE"
]
},
"ForceSVEWidth": {
"Type": "uint32",
"Default": "0",
@@ -350,7 +389,7 @@
"Default": "server",
"Desc": [
"File to write FEX output to.",
"[stdout, stderr, server, <Filename>]"
"[stderr, server, <Filename>]"
]
},
"TelemetryDirectory": {
@@ -358,7 +397,7 @@
"Default": "",
"Desc": [
"Redirects the telemetry folder that FEX usually writes to.",
"By default telemetry data is stored in {$FEX_APP_DATA_LOCATION,{$XDG_DATA_HOME,$HOME}/.fex-emu/Telemetry/}"
"By default telemetry data is stored in {$FEX_APP_DATA_LOCATION,{$XDG_DATA_HOME,$HOME}/fex-emu/Telemetry/}"
]
},
"ProfileStats": {
@@ -429,6 +468,13 @@
"This is required to ensure a split-lock doesn't tear inside the process"
]
},
"KernelUnalignedAtomicBackpatching": {
"Type": "bool",
"Default": "true",
"Desc": [
"When the kernel unaligned atomic handler is enabled, use backpatching to reduce kernel context switches."
]
},
"VolatileMetadata": {
"Type": "bool",
"Default": "true",
+63 -12
View File
@@ -4,7 +4,6 @@
#include "Common/JitSymbols.h"
#include "Interface/Core/CPUBackend.h"
#include "Interface/Core/CPUID.h"
#include "Interface/Core/X86HelperGen.h"
#include <Interface/IR/IntrusiveIRList.h>
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/Context.h>
@@ -62,7 +61,7 @@ struct CustomIRResult {
, Data(Data) {}
};
using BlockDelinkerFunc = void (*)(FEXCore::Core::CpuStateFrame* Frame, FEXCore::Context::ExitFunctionLinkData* Record);
using BlockDelinkerFunc = void (*)(FEXCore::Context::ExitFunctionLinkData* Record);
constexpr uint32_t TSC_SCALE_MAXIMUM = 1'000'000'000; ///< 1Ghz
class CodeCache : public AbstractCodeCache {
@@ -71,14 +70,54 @@ public:
~CodeCache();
ContextImpl& CTX;
fextl::unique_ptr<ContextImpl> ValidationCTX;
fextl::unique_ptr<Core::InternalThreadState> ValidationThread;
FEXCore::Core::CPUState::gdt_segment ValidationGDT[32] {};
bool IsGeneratingCache = false;
void LoadData(Core::InternalThreadState&, std::byte* MappedCacheFile, const ExecutableFileSectionInfo&) override;
FEX_CONFIG_OPT(EnableCodeCaching, ENABLECODECACHINGWIP);
FEX_CONFIG_OPT(EnableCodeCacheValidation, ENABLECODECACHEVALIDATION);
uint64_t ComputeCodeMapId(std::string_view Filename, int FD) override;
bool SaveData(Core::InternalThreadState&, int TargetFD, const ExecutableFileSectionInfo&, uint64_t SerializedBaseAddress) override;
bool LoadData(Core::InternalThreadState*, std::byte* MappedCacheFile, const ExecutableFileSectionInfo&) override;
/**
* Performs expensive extra validation on the loaded code cache data.
*
* This kicks off an in-process recompile of all cached blocks and compares
* them with the cached data. Differences will be reported as fatal errors,
* which can uncover bugs like for example:
* - mismatches of the JIT configuration used during cache generation
* - hidden position dependencies due to missing FEX relocations
* - incorrect instruction padding
*/
void Validate(const ExecutableFileSectionInfo&, fextl::set<uint64_t> GuestBlocks, const fextl::set<uint64_t>& HostBlocks,
std::span<std::byte> CachedCode);
void InitiateCacheGeneration() override {
IsGeneratingCache = true;
}
/**
* Applies a set of FEX relocations to the given code section.
*
* FEX relocations describe runtime-dependencies of FEX-generated code.
* When loading a code cache, they are used to move cached code to the
* dynamically chosen base address of the guest binary.
*
* Conversely, relocations are applied in reverse when writing code caches
* to ensure consistency across generation runs.
*
* Note that FEX relocations are unrelated to ELF/PE relocations.
*
* @param GuestDelta Guest address offset to apply to RIP-relative data
* @param ForStorage True for serializing data (producing deterministic output); false for de-serializing it (resolving dynamic symbols)
*
* @return Returns true on success
*/
[[nodiscard]]
bool ApplyCodeRelocations(uint64_t GuestDelta, std::span<std::byte> Code, std::span<const CPU::Relocation> Relocations, bool ForStorage);
};
class ContextImpl final : public FEXCore::Context::Context, public CPU::CodeBufferManager {
@@ -88,6 +127,7 @@ public:
void ExecuteThread(FEXCore::Core::InternalThreadState* Thread) override;
bool CheckIfBlockIsCacheable(FEXCore::Core::InternalThreadState&, uint64_t GuestRIP, uint64_t MaxInst) override;
void CompileRIP(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP) override;
void CompileRIPCount(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP, uint64_t MaxInst) override;
@@ -155,11 +195,21 @@ public:
return CodeCache;
}
void OnCodeBufferAllocated(CPU::CodeBuffer&) override;
void SetCodeMapWriter(fextl::unique_ptr<CodeMapWriter> Writer) override {
CodeMapWriter = std::move(Writer);
}
void FlushAndCloseCodeMap() override {
if (CodeMapWriter) {
CodeMapWriter.reset();
}
}
void OnCodeBufferAllocated(const std::shared_ptr<CPU::CodeBuffer>&) override;
void ClearCodeCache(FEXCore::Core::InternalThreadState* Thread, bool NewCodeBuffer = true) override;
void InvalidateGuestCodeRange(FEXCore::Core::InternalThreadState* Thread, InvalidatedEntryAccumulator& Accumulator, uint64_t Start,
uint64_t Length) override;
FEXCore::ForkableSharedMutex& GetCodeInvalidationMutex() override {
void InvalidateCodeBuffersCodeRange(uint64_t Start, uint64_t Length) override;
void InvalidateThreadCachedCodeRange(FEXCore::Core::InternalThreadState* Thread, uint64_t Start, uint64_t Length) override;
FEXCore::Utils::WritePriorityMutex::Mutex& GetCodeInvalidationMutex() override {
return CodeInvalidationMutex;
}
@@ -213,7 +263,7 @@ public:
FEX_CONFIG_OPT(MonoHacks, MONOHACKS);
} Config;
FEXCore::ForkableSharedMutex CodeInvalidationMutex;
FEXCore::Utils::WritePriorityMutex::Mutex CodeInvalidationMutex {};
uint32_t StrictSplitLockMutex {};
@@ -225,14 +275,12 @@ public:
FEXCore::ThunkHandler* ThunkHandler {};
fextl::unique_ptr<FEXCore::CPU::Dispatcher> Dispatcher;
CodeCache CodeCache;
fextl::unique_ptr<CodeMapWriter> CodeMapWriter;
SignalDelegator* SignalDelegation {};
X86GeneratedCode X86CodeGen;
ContextImpl(const FEXCore::HostFeatures& Features);
static bool ThreadRemoveCodeEntry(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP, const FEXCore::LookupCacheWriteLockToken& lk);
static void ThreadRemoveCodeEntryFromJit(FEXCore::Core::CpuStateFrame* Frame, uint64_t GuestRIP);
// This is used as a replacement for the SMC writes in the mono callsite backpatcher that avoids atomic operations
@@ -269,7 +317,7 @@ public:
FEXCore::Utils::PooledAllocatorVirtual OpDispatcherAllocator {"FEXMem_OpDispatcher"};
FEXCore::Utils::PooledAllocatorVirtual FrontendAllocator {"FEXMem_Frontend"};
FEXCore::Utils::PooledAllocatorVirtual CPUBackendAllocator {"FEXMem_CPUBackend"};
FEXCore::Utils::PooledAllocatorVirtualWithGuard CPUBackendAllocator {"FEXMem_CPUBackend"};
// If Atomic-based TSO emulation is enabled or not.
bool IsAtomicTSOEnabled() const {
@@ -348,5 +396,8 @@ private:
bool MonoDetected = false;
std::atomic<uint64_t> MonoBackpatcherBlock;
std::mutex CodeBufferListLock;
fextl::vector<std::weak_ptr<CPU::CodeBuffer>> CodeBufferList;
};
} // namespace FEXCore::Context
@@ -41,7 +41,7 @@ namespace FEXCore::CPU {
// r19-r29 and SP.
namespace x64 {
#ifndef _M_ARM_64EC
#ifndef ARCHITECTURE_arm64ec
// All but x19 and x29 are caller saved
// Note that rax/rdx are rearranged here so we can coalesce cmpxchg.
constexpr std::array<ARMEmitter::Register, 18> SRA = {
@@ -417,36 +417,54 @@ FEXCore::X86State::X86Reg Arm64Emitter::GetX86RegRelationToARMReg(ARMEmitter::Re
return FEXCore::X86State::X86Reg::REG_INVALID;
}
void Arm64Emitter::LoadConstant(ARMEmitter::Size s, ARMEmitter::Register Reg, uint64_t Constant, bool NOPPad) {
void Arm64Emitter::LoadConstant(ARMEmitter::Size s, ARMEmitter::Register Reg, uint64_t Constant, PadType Pad, int MaxBytes) {
bool NOPPad = false;
if (Pad == PadType::DOPAD) {
NOPPad = true;
} else if (Pad == PadType::NOPAD) {
NOPPad = false;
} else if (Pad == PadType::AUTOPAD) {
// Force NOP padding to ensure relocated constants always have enough encoding space available
NOPPad = EnableCodeCaching;
}
bool Is64Bit = s == ARMEmitter::Size::i64Bit;
int Segments = Is64Bit ? 4 : 2;
const auto UpperBound = Is64Bit ? 4 : 2;
int Segments = MaxBytes ? (MaxBytes / 2) : UpperBound;
LOGMAN_THROW_A_FMT(MaxBytes >= 0 && MaxBytes <= (UpperBound * 2) && (MaxBytes & 1) == 0,
"MaxBytes must be bounded in the range of [0, {}] and 16-bit aligned", UpperBound);
// If MaxBytes specified then make sure to sanity check incoming data.
LOGMAN_THROW_A_FMT(MaxBytes == 0 || (Constant >> (MaxBytes * 8)) == 0, "MaxBytes provided but data can't fit within provided range.");
if (Is64Bit && ((~Constant) >> 16) == 0) {
movn(s, Reg, (~Constant) & 0xFFFF);
if (NOPPad) {
nop();
nop();
nop();
}
movn(s, Reg, (~Constant) & 0xFFFF);
return;
}
if ((Constant >> 32) == 0) {
if ((Constant >> 32) == 0 && !NOPPad) {
// If the upper 32-bits is all zero, we can now switch to a 32-bit move.
// NOTE: The NOP padding code does not appropriately adjust to this yet,
// so we skip this optimization in that case
s = ARMEmitter::Size::i32Bit;
Is64Bit = false;
Segments = 2;
Segments = std::min(Segments, 2);
}
if (!Is64Bit && ((~Constant) & 0xFFFF0000) == 0) {
movn(s, Reg.W(), (~Constant) & 0xFFFF);
if (NOPPad) {
nop();
nop();
nop();
}
movn(s, Reg.W(), (~Constant) & 0xFFFF);
return;
}
@@ -467,24 +485,24 @@ void Arm64Emitter::LoadConstant(ARMEmitter::Size s, ARMEmitter::Register Reg, ui
// `movz` is better than `orr` since hardware will rename or merge if possible when `movz` is used.
const auto IsImm = ARMEmitter::Emitter::IsImmLogical(Constant, RegSizeInBits(s));
if (IsImm) {
orr(s, Reg, ARMEmitter::Reg::zr, Constant);
if (NOPPad) {
nop();
nop();
nop();
}
orr(s, Reg, ARMEmitter::Reg::zr, Constant);
return;
}
}
// If we can't handle negatives with the orr, try with movn+movk
if (Is64Bit && ((~Constant) >> 32) == 0) {
movn(s, Reg, (~Constant) & 0xFFFF);
movk(s, Reg, (Constant >> 16) & 0xFFFF, 16);
if (NOPPad) {
nop();
nop();
}
movn(s, Reg, (~Constant) & 0xFFFF);
movk(s, Reg, (Constant >> 16) & 0xFFFF, 16);
return;
}
@@ -661,7 +679,7 @@ void Arm64Emitter::FillSpecialRegs(ARMEmitter::Register TmpReg, ARMEmitter::Regi
}
}
void Arm64Emitter::SpillStaticRegs(ARMEmitter::Register TmpReg, bool FPRs, uint32_t GPRSpillMask, uint32_t FPRSpillMask) {
void Arm64Emitter::SpillStaticRegs(ARMEmitter::Register TmpReg, SpillStaticRegOptions Options) {
#ifndef VIXL_SIMULATOR
if (EmitterCTX->HostFeatures.SupportsAFP) {
// Disable AFP features when spilling registers.
@@ -682,35 +700,37 @@ void Arm64Emitter::SpillStaticRegs(ARMEmitter::Register TmpReg, bool FPRs, uint3
}
#endif
// Regardless of what GPRs/FPRs we're spilling, we need to spill NZCV since it
// is always static and almost certainly clobbered by the subsequent code.
//
// TODO: Can we prove that NZCV is not used across a call in some cases and
// omit this? Might help x87 perf? Future idea.
mrs(TmpReg, ARMEmitter::SystemRegister::NZCV);
str(TmpReg.W(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.flags[24]));
if (Options.NZCV) {
// Regardless of what GPRs/FPRs we're spilling, we need to spill NZCV since it
// is always static and almost certainly clobbered by the subsequent code.
//
// TODO: Can we prove that NZCV is not used across a call in some cases and
// omit this? Might help x87 perf? Future idea.
mrs(TmpReg, ARMEmitter::SystemRegister::NZCV);
str(TmpReg.W(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.flags[24]));
}
// PF/AF are special, remove them from the mask
uint32_t PFAFMask = ((1u << REG_PF.Idx()) | ((1u << REG_AF.Idx())));
unsigned PFAFSpillMask = GPRSpillMask & PFAFMask;
GPRSpillMask &= ~PFAFSpillMask;
unsigned PFAFSpillMask = Options.GPRSpillMask & PFAFMask;
Options.GPRSpillMask &= ~PFAFSpillMask;
str(REG_CALLRET_SP, STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.callret_sp));
for (size_t i = 0; i < StaticRegisters.size(); i += 2) {
auto Reg1 = StaticRegisters[i];
auto Reg2 = StaticRegisters[i + 1];
if (((1U << Reg1.Idx()) & GPRSpillMask) && ((1U << Reg2.Idx()) & GPRSpillMask)) {
stp<ARMEmitter::IndexType::OFFSET>(Reg1.X(), Reg2.X(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i]));
} else if (((1U << Reg1.Idx()) & GPRSpillMask)) {
str(Reg1.X(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i]));
} else if (((1U << Reg2.Idx()) & GPRSpillMask)) {
str(Reg2.X(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i + 1]));
if (((1U << Reg1.Idx()) & Options.GPRSpillMask) && ((1U << Reg2.Idx()) & Options.GPRSpillMask)) {
stp<ARMEmitter::IndexType::OFFSET>(Reg1.X(), Reg2.X(), STATE.R(), ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, State.gregs, i));
} else if (((1U << Reg1.Idx()) & Options.GPRSpillMask)) {
str(Reg1.X(), STATE.R(), ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, State.gregs, i));
} else if (((1U << Reg2.Idx()) & Options.GPRSpillMask)) {
str(Reg2.X(), STATE.R(), ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, State.gregs, i + 1));
}
}
// Now handle PF/AF
if (PFAFSpillMask) {
if (Options.NZCV && PFAFSpillMask) {
auto PFOffset = offsetof(FEXCore::Core::CpuStateFrame, State.pf_raw);
auto AFOffset = offsetof(FEXCore::Core::CpuStateFrame, State.af_raw);
LOGMAN_THROW_A_FMT(PFAFSpillMask == PFAFMask, "PF/AF not spilled together");
@@ -719,21 +739,21 @@ void Arm64Emitter::SpillStaticRegs(ARMEmitter::Register TmpReg, bool FPRs, uint3
stp<ARMEmitter::IndexType::OFFSET>(REG_PF.W(), REG_AF.W(), STATE.R(), PFOffset);
}
if (FPRs) {
if (Options.FPRs) {
if (EmitterCTX->HostFeatures.SupportsAVX && EmitterCTX->HostFeatures.SupportsSVE256) {
for (size_t i = 0; i < StaticFPRegisters.size(); i++) {
const auto Reg = StaticFPRegisters[i];
if (((1U << Reg.Idx()) & FPRSpillMask) != 0) {
mov(ARMEmitter::Size::i64Bit, TmpReg, offsetof(Core::CpuStateFrame, State.xmm.avx.data[i][0]));
if (((1U << Reg.Idx()) & Options.FPRSpillMask) != 0) {
mov(ARMEmitter::Size::i64Bit, TmpReg, ARRAY_OFFSETOF(Core::CpuStateFrame, State.xmm.avx.data, i));
st1b<ARMEmitter::SubRegSize::i8Bit>(Reg.Z(), PRED_TMP_32B, STATE.R(), TmpReg);
}
}
} else {
if (GPRSpillMask && FPRSpillMask == ~0U) {
if (Options.GPRSpillMask && Options.FPRSpillMask == ~0U) {
// Optimize the common case where we can spill four registers per instruction
// Load the sse offset in to the temporary register
add(ARMEmitter::Size::i64Bit, TmpReg, STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[0][0]));
add(ARMEmitter::Size::i64Bit, TmpReg, STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data));
for (size_t i = 0; i < StaticFPRegisters.size(); i += 4) {
const auto Reg1 = StaticFPRegisters[i];
const auto Reg2 = StaticFPRegisters[i + 1];
@@ -746,12 +766,12 @@ void Arm64Emitter::SpillStaticRegs(ARMEmitter::Register TmpReg, bool FPRs, uint3
const auto Reg1 = StaticFPRegisters[i];
const auto Reg2 = StaticFPRegisters[i + 1];
if (((1U << Reg1.Idx()) & FPRSpillMask) && ((1U << Reg2.Idx()) & FPRSpillMask)) {
stp<ARMEmitter::IndexType::OFFSET>(Reg1.Q(), Reg2.Q(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[i][0]));
} else if (((1U << Reg1.Idx()) & FPRSpillMask)) {
str(Reg1.Q(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[i][0]));
} else if (((1U << Reg2.Idx()) & FPRSpillMask)) {
str(Reg2.Q(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[i + 1][0]));
if (((1U << Reg1.Idx()) & Options.FPRSpillMask) && ((1U << Reg2.Idx()) & Options.FPRSpillMask)) {
stp<ARMEmitter::IndexType::OFFSET>(Reg1.Q(), Reg2.Q(), STATE.R(), ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, State.xmm.sse.data, i));
} else if (((1U << Reg1.Idx()) & Options.FPRSpillMask)) {
str(Reg1.Q(), STATE.R(), ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, State.xmm.sse.data, i));
} else if (((1U << Reg2.Idx()) & Options.FPRSpillMask)) {
str(Reg2.Q(), STATE.R(), ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, State.xmm.sse.data, i + 1));
}
}
}
@@ -759,8 +779,7 @@ void Arm64Emitter::SpillStaticRegs(ARMEmitter::Register TmpReg, bool FPRs, uint3
}
}
void Arm64Emitter::FillStaticRegs(bool FPRs, uint32_t GPRFillMask, uint32_t FPRFillMask, std::optional<ARMEmitter::Register> OptionalReg,
std::optional<ARMEmitter::Register> OptionalReg2) {
void Arm64Emitter::FillStaticRegs(FillStaticRegOptions Options) {
auto FindTempReg = [this](uint32_t* GPRFillMask) -> std::optional<ARMEmitter::Register> {
for (auto Reg : StaticRegisters) {
if (((1U << Reg.Idx()) & *GPRFillMask)) {
@@ -771,22 +790,23 @@ void Arm64Emitter::FillStaticRegs(bool FPRs, uint32_t GPRFillMask, uint32_t FPRF
return std::nullopt;
};
LOGMAN_THROW_A_FMT(GPRFillMask != 0, "Must fill at least 2 GPRs for a temp");
uint32_t TempGPRFillMask = GPRFillMask;
if (!OptionalReg.has_value()) {
OptionalReg = FindTempReg(&TempGPRFillMask);
LOGMAN_THROW_A_FMT(Options.GPRFillMask != 0, "Must fill at least 2 GPRs for a temp");
uint32_t TempGPRFillMask = Options.GPRFillMask;
if (!Options.OptionalReg.has_value()) {
Options.OptionalReg = FindTempReg(&TempGPRFillMask);
}
if (!OptionalReg2.has_value()) {
OptionalReg2 = FindTempReg(&TempGPRFillMask);
if (!Options.OptionalReg2.has_value()) {
Options.OptionalReg2 = FindTempReg(&TempGPRFillMask);
}
LOGMAN_THROW_A_FMT(OptionalReg.has_value() && OptionalReg2.has_value(), "Didn't have an SRA register to use as a temporary while "
"spilling!");
LOGMAN_THROW_A_FMT(Options.OptionalReg.has_value() && Options.OptionalReg2.has_value(), "Didn't have an SRA register to use as a "
"temporary while "
"spilling!");
auto TmpReg = *OptionalReg;
auto TmpReg2 = *OptionalReg2;
auto TmpReg = *Options.OptionalReg;
auto TmpReg2 = *Options.OptionalReg2;
#ifdef _M_ARM_64EC
#ifdef ARCHITECTURE_arm64ec
// Load STATE in from the CPU area as x28 is not callee saved in the ARM64EC ABI.
ldr(TmpReg.X(), ARMEmitter::Reg::r18, TEB_CPU_AREA_OFFSET);
ldr(STATE, TmpReg, CPU_AREA_EMULATOR_DATA_OFFSET);
@@ -794,31 +814,33 @@ void Arm64Emitter::FillStaticRegs(bool FPRs, uint32_t GPRFillMask, uint32_t FPRF
ldr(REG_CALLRET_SP, STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.callret_sp));
// Regardless of what GPRs/FPRs we're filling, we need to fill NZCV since it
// is always static and was almost certainly clobbered.
//
// TODO: Can we prove that NZCV is not used across a call in some cases and
// omit this? Might help x87 perf? Future idea.
ldr(TmpReg.W(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.flags[24]));
msr(ARMEmitter::SystemRegister::NZCV, TmpReg);
if (Options.NZCV) {
// Regardless of what GPRs/FPRs we're filling, we need to fill NZCV since it
// is always static and was almost certainly clobbered.
//
// TODO: Can we prove that NZCV is not used across a call in some cases and
// omit this? Might help x87 perf? Future idea.
ldr(TmpReg.W(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.flags[24]));
msr(ARMEmitter::SystemRegister::NZCV, TmpReg);
}
FillSpecialRegs(TmpReg, TmpReg2, true, FPRs);
FillSpecialRegs(TmpReg, TmpReg2, true, Options.FPRs);
if (FPRs) {
if (Options.FPRs) {
if (EmitterCTX->HostFeatures.SupportsAVX && EmitterCTX->HostFeatures.SupportsSVE256) {
for (size_t i = 0; i < StaticFPRegisters.size(); i++) {
const auto Reg = StaticFPRegisters[i];
if (((1U << Reg.Idx()) & FPRFillMask) != 0) {
mov(ARMEmitter::Size::i64Bit, TmpReg, offsetof(Core::CpuStateFrame, State.xmm.avx.data[i][0]));
if (((1U << Reg.Idx()) & Options.FPRFillMask) != 0) {
mov(ARMEmitter::Size::i64Bit, TmpReg, ARRAY_OFFSETOF(Core::CpuStateFrame, State.xmm.avx.data, i));
ld1b<ARMEmitter::SubRegSize::i8Bit>(Reg.Z(), PRED_TMP_32B.Zeroing(), STATE.R(), TmpReg);
}
}
} else {
if (GPRFillMask && FPRFillMask == ~0U) {
if (Options.GPRFillMask && Options.FPRFillMask == ~0U) {
// Optimize the common case where we can fill four registers per instruction.
// Use one of the filling static registers before we fill it.
// Load the sse offset in to the temporary register
add(ARMEmitter::Size::i64Bit, TmpReg, STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[0][0]));
add(ARMEmitter::Size::i64Bit, TmpReg, STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data));
for (size_t i = 0; i < StaticFPRegisters.size(); i += 4) {
const auto Reg1 = StaticFPRegisters[i];
const auto Reg2 = StaticFPRegisters[i + 1];
@@ -831,12 +853,12 @@ void Arm64Emitter::FillStaticRegs(bool FPRs, uint32_t GPRFillMask, uint32_t FPRF
const auto Reg1 = StaticFPRegisters[i];
const auto Reg2 = StaticFPRegisters[i + 1];
if (((1U << Reg1.Idx()) & FPRFillMask) && ((1U << Reg2.Idx()) & FPRFillMask)) {
ldp<ARMEmitter::IndexType::OFFSET>(Reg1.Q(), Reg2.Q(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[i][0]));
} else if (((1U << Reg1.Idx()) & FPRFillMask)) {
ldr(Reg1.Q(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[i][0]));
} else if (((1U << Reg2.Idx()) & FPRFillMask)) {
ldr(Reg2.Q(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[i + 1][0]));
if (((1U << Reg1.Idx()) & Options.FPRFillMask) && ((1U << Reg2.Idx()) & Options.FPRFillMask)) {
ldp<ARMEmitter::IndexType::OFFSET>(Reg1.Q(), Reg2.Q(), STATE.R(), ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, State.xmm.sse.data, i));
} else if (((1U << Reg1.Idx()) & Options.FPRFillMask)) {
ldr(Reg1.Q(), STATE.R(), ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, State.xmm.sse.data, i));
} else if (((1U << Reg2.Idx()) & Options.FPRFillMask)) {
ldr(Reg2.Q(), STATE.R(), ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, State.xmm.sse.data, i + 1));
}
}
}
@@ -845,23 +867,23 @@ void Arm64Emitter::FillStaticRegs(bool FPRs, uint32_t GPRFillMask, uint32_t FPRF
// PF/AF are special, remove them from the mask
uint32_t PFAFMask = ((1u << REG_PF.Idx()) | ((1u << REG_AF.Idx())));
uint32_t PFAFFillMask = GPRFillMask & PFAFMask;
GPRFillMask &= ~PFAFMask;
uint32_t PFAFFillMask = Options.GPRFillMask & PFAFMask;
Options.GPRFillMask &= ~PFAFMask;
for (size_t i = 0; i < StaticRegisters.size(); i += 2) {
auto Reg1 = StaticRegisters[i];
auto Reg2 = StaticRegisters[i + 1];
if (((1U << Reg1.Idx()) & GPRFillMask) && ((1U << Reg2.Idx()) & GPRFillMask)) {
ldp<ARMEmitter::IndexType::OFFSET>(Reg1.X(), Reg2.X(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i]));
} else if ((1U << Reg1.Idx()) & GPRFillMask) {
ldr(Reg1.X(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i]));
} else if ((1U << Reg2.Idx()) & GPRFillMask) {
ldr(Reg2.X(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i + 1]));
if (((1U << Reg1.Idx()) & Options.GPRFillMask) && ((1U << Reg2.Idx()) & Options.GPRFillMask)) {
ldp<ARMEmitter::IndexType::OFFSET>(Reg1.X(), Reg2.X(), STATE.R(), ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, State.gregs, i));
} else if ((1U << Reg1.Idx()) & Options.GPRFillMask) {
ldr(Reg1.X(), STATE.R(), ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, State.gregs, i));
} else if ((1U << Reg2.Idx()) & Options.GPRFillMask) {
ldr(Reg2.X(), STATE.R(), ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, State.gregs, i + 1));
}
}
// Now handle PF/AF
if (PFAFFillMask) {
if (Options.NZCV && PFAFFillMask) {
LOGMAN_THROW_A_FMT(PFAFFillMask == PFAFMask, "PF/AF not filled together");
ldp<ARMEmitter::IndexType::OFFSET>(REG_PF.W(), REG_AF.W(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.pf_raw));
@@ -1036,7 +1058,10 @@ size_t Arm64Emitter::SpillForPreserveAllABICall(ARMEmitter::Register TmpReg, boo
const uint64_t SPOffset = AlignUp(GPRSize + FPRSize, 16);
// Spill the static registers.
SpillStaticRegs(TmpReg, true, PreserveSRAMask, PreserveSRAFPRMask);
SpillStaticRegs(TmpReg, {
.GPRSpillMask = PreserveSRAMask,
.FPRSpillMask = PreserveSRAFPRMask,
});
sub(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, ARMEmitter::Reg::rsp, SPOffset);
@@ -1083,7 +1108,11 @@ void Arm64Emitter::FillForPreserveAllABICall(bool FPRs) {
}
// Fill the static registers.
FillStaticRegs(FPRs, PreserveSRAMask, PreserveSRAFPRMask);
FillStaticRegs({
.GPRFillMask = PreserveSRAMask,
.FPRFillMask = PreserveSRAFPRMask,
.FPRs = FPRs,
});
// Pop the vector registers.
PopVectorRegisters(CanUseSVE256, DynamicFPRs);
@@ -1,9 +1,10 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <FEXCore/Config/Config.h>
#ifdef VIXL_DISASSEMBLER
#include <aarch64/disasm-aarch64.h>
#include <FEXCore/Config/Config.h>
#include <FEXCore/fextl/memory.h>
#include <FEXCore/fextl/vector.h>
#endif
@@ -31,7 +32,7 @@ namespace FEXCore::CPU {
// Contains the address to the currently available CPU state
constexpr auto STATE = ARMEmitter::XReg::x28;
#ifndef _M_ARM_64EC
#ifndef ARCHITECTURE_arm64ec
// GPR temporaries. Only x3 can be used across spill boundaries
// so if these ever need to change, be very careful about that.
constexpr auto TMP1 = ARMEmitter::XReg::x0;
@@ -105,9 +106,20 @@ constexpr ARMEmitter::PRegister PRED_TMP_32B = ARMEmitter::PReg::p7;
// This class contains common emitter utility functions that can
// be used by both Arm64 JIT and ARM64 Dispatcher
class Arm64Emitter : public ARMEmitter::Emitter {
protected:
public:
Arm64Emitter(FEXCore::Context::ContextImpl* ctx, void* EmissionPtr = nullptr, size_t size = 0);
enum class PadType {
// Explicitly does not need padding, even if code-caching is enabled.
NOPAD,
// Explicitly needs padding, even if code-caching is disabled.
DOPAD,
// Choose to pad or not depending on if code-caching is enabled.
AUTOPAD,
};
void LoadConstant(ARMEmitter::Size s, ARMEmitter::Register Reg, uint64_t Constant, PadType Pad = PadType::NOPAD, int MaxBytes = 0);
protected:
FEXCore::Context::ContextImpl* EmitterCTX;
std::span<const ARMEmitter::Register> StaticRegisters {};
@@ -117,18 +129,41 @@ protected:
std::span<const ARMEmitter::VRegister> GeneralFPRegisters {};
uint32_t PairRegisters = 0;
void LoadConstant(ARMEmitter::Size s, ARMEmitter::Register Reg, uint64_t Constant, bool NOPPad = false);
void FillSpecialRegs(ARMEmitter::Register TmpReg, ARMEmitter::Register TmpReg2, bool SetFIZ, bool SetPredRegs);
// Correlate an ARM register back to an x86 register index.
// Returning REG_INVALID if there was no mapping.
FEXCore::X86State::X86Reg GetX86RegRelationToARMReg(ARMEmitter::Register Reg);
void SpillStaticRegs(ARMEmitter::Register TmpReg, bool FPRs = true, uint32_t GPRSpillMask = ~0U, uint32_t FPRSpillMask = ~0U);
void FillStaticRegs(bool FPRs = true, uint32_t GPRFillMask = ~0U, uint32_t FPRFillMask = ~0U,
std::optional<ARMEmitter::Register> OptionalReg = std::nullopt,
std::optional<ARMEmitter::Register> OptionalReg2 = std::nullopt);
struct SpillStaticRegOptions final {
uint32_t GPRSpillMask {~0U};
uint32_t FPRSpillMask {~0U};
bool FPRs {true};
bool NZCV {true};
};
struct FillStaticRegOptions final {
std::optional<ARMEmitter::Register> OptionalReg {std::nullopt};
std::optional<ARMEmitter::Register> OptionalReg2 {std::nullopt};
uint32_t GPRFillMask {~0U};
uint32_t FPRFillMask {~0U};
bool FPRs {true};
bool NZCV {true};
};
void SpillStaticRegs(ARMEmitter::Register TmpReg, SpillStaticRegOptions Options);
void FillStaticRegs(FillStaticRegOptions Options);
void SpillStaticRegs(ARMEmitter::Register TmpReg) {
// Work around a clang bug: https://bugs.llvm.org/show_bug.cgi?id=36684
SpillStaticRegs(TmpReg, {});
}
void FillStaticRegs() {
// Work around a clang bug: https://bugs.llvm.org/show_bug.cgi?id=36684
FillStaticRegs({});
}
// Register 0-18 + 29 + 30 are caller saved
static constexpr uint32_t CALLER_GPR_MASK = 0b0110'0000'0000'0111'1111'1111'1111'1111U;
@@ -168,7 +203,9 @@ protected:
if (SupportsPreserveAllABI) {
return SpillForPreserveAllABICall(TmpReg, FPRs);
} else {
SpillStaticRegs(TmpReg, FPRs);
SpillStaticRegs(TmpReg, {
.FPRs = FPRs,
});
return PushDynamicRegs(TmpReg);
}
}
@@ -178,7 +215,7 @@ protected:
FillForPreserveAllABICall(FPRs);
} else {
PopDynamicRegs();
FillStaticRegs(FPRs);
FillStaticRegs({.FPRs = FPRs});
}
}
@@ -271,6 +308,8 @@ protected:
FEX_CONFIG_OPT(Disassemble, DISASSEMBLE);
#endif
FEX_CONFIG_OPT(EnableCodeCaching, ENABLECODECACHINGWIP);
};
} // namespace FEXCore::CPU
+27 -18
View File
@@ -1,4 +1,5 @@
// SPDX-License-Identifier: MIT
#include "FEXCore/Config/Config.h"
#include "Interface/Context/Context.h"
#include "Interface/Core/CPUBackend.h"
#include "Interface/Core/LookupCache.h"
@@ -11,7 +12,6 @@
#include <cstdint>
#ifndef _WIN32
#include <linux/prctl.h>
#include <sys/prctl.h>
#endif
@@ -277,37 +277,37 @@ namespace CPU {
: ThreadState(ThreadState)
, CodeBuffers(CodeBuffers) {
auto& Common = ThreadState->CurrentFrame->Pointers.Common;
auto& Ptrs = ThreadState->CurrentFrame->Pointers;
// Initialize named vector constants.
for (size_t i = 0; i < FEXCore::IR::NamedVectorConstant::NAMED_VECTOR_CONST_POOL_MAX; ++i) {
Common.NamedVectorConstantPointers[i] = reinterpret_cast<uint64_t>(NamedVectorConstants[i]);
Ptrs.NamedVectorConstantPointers[i] = reinterpret_cast<uint64_t>(NamedVectorConstants[i]);
}
// Copy named vector constants.
memcpy(Common.NamedVectorConstants, NamedVectorConstants, sizeof(NamedVectorConstants));
memcpy(Ptrs.NamedVectorConstants, NamedVectorConstants, sizeof(NamedVectorConstants));
// Initialize Indexed named vector constants.
Common.IndexedNamedVectorConstantPointers[FEXCore::IR::IndexNamedVectorConstant::INDEXED_NAMED_VECTOR_PSHUFLW] =
Ptrs.IndexedNamedVectorConstantPointers[FEXCore::IR::IndexNamedVectorConstant::INDEXED_NAMED_VECTOR_PSHUFLW] =
reinterpret_cast<uint64_t>(PSHUFLW_LUT.data());
Common.IndexedNamedVectorConstantPointers[FEXCore::IR::IndexNamedVectorConstant::INDEXED_NAMED_VECTOR_PSHUFHW] =
Ptrs.IndexedNamedVectorConstantPointers[FEXCore::IR::IndexNamedVectorConstant::INDEXED_NAMED_VECTOR_PSHUFHW] =
reinterpret_cast<uint64_t>(PSHUFHW_LUT.data());
Common.IndexedNamedVectorConstantPointers[FEXCore::IR::IndexNamedVectorConstant::INDEXED_NAMED_VECTOR_PSHUFD] =
Ptrs.IndexedNamedVectorConstantPointers[FEXCore::IR::IndexNamedVectorConstant::INDEXED_NAMED_VECTOR_PSHUFD] =
reinterpret_cast<uint64_t>(PSHUFD_LUT.data());
Common.IndexedNamedVectorConstantPointers[FEXCore::IR::IndexNamedVectorConstant::INDEXED_NAMED_VECTOR_SHUFPS] =
Ptrs.IndexedNamedVectorConstantPointers[FEXCore::IR::IndexNamedVectorConstant::INDEXED_NAMED_VECTOR_SHUFPS] =
reinterpret_cast<uint64_t>(SHUFPS_LUT.data());
Common.IndexedNamedVectorConstantPointers[FEXCore::IR::IndexNamedVectorConstant::INDEXED_NAMED_VECTOR_DPPS_MASK] =
Ptrs.IndexedNamedVectorConstantPointers[FEXCore::IR::IndexNamedVectorConstant::INDEXED_NAMED_VECTOR_DPPS_MASK] =
reinterpret_cast<uint64_t>(DPPS_MASK.data());
Common.IndexedNamedVectorConstantPointers[FEXCore::IR::IndexNamedVectorConstant::INDEXED_NAMED_VECTOR_DPPD_MASK] =
Ptrs.IndexedNamedVectorConstantPointers[FEXCore::IR::IndexNamedVectorConstant::INDEXED_NAMED_VECTOR_DPPD_MASK] =
reinterpret_cast<uint64_t>(DPPD_MASK.data());
Common.IndexedNamedVectorConstantPointers[FEXCore::IR::IndexNamedVectorConstant::INDEXED_NAMED_VECTOR_PBLENDW] =
Ptrs.IndexedNamedVectorConstantPointers[FEXCore::IR::IndexNamedVectorConstant::INDEXED_NAMED_VECTOR_PBLENDW] =
reinterpret_cast<uint64_t>(PBLENDW_LUT.data());
#ifndef FEX_DISABLE_TELEMETRY
// Fill in telemetry values
for (size_t i = 0; i < FEXCore::Telemetry::TYPE_LAST; ++i) {
auto& Telem = FEXCore::Telemetry::GetTelemetryValue(static_cast<FEXCore::Telemetry::TelemetryType>(i));
Common.TelemetryValueAddresses[i] = reinterpret_cast<uint64_t>(&Telem);
Ptrs.TelemetryValueAddresses[i] = reinterpret_cast<uint64_t>(&Telem);
}
#endif
}
@@ -349,7 +349,7 @@ namespace CPU {
}
CodeBuffer::CodeBuffer(size_t Size)
: Size(Size) {
: AllocatedSize(Size) {
Ptr = static_cast<uint8_t*>(FEXCore::Allocator::VirtualAlloc(Size, true));
LOGMAN_THROW_A_FMT(!!Ptr, "Couldn't allocate code buffer");
@@ -362,11 +362,14 @@ namespace CPU {
FEXCore::Allocator::VirtualName("FEXMemJIT", reinterpret_cast<void*>(Ptr), Size);
// Huge-pages reduce the amount of iTLB misses dramatically when it works.
FEXCore::Allocator::VirtualTHPControl(reinterpret_cast<void*>(Ptr), Size, FEXCore::Allocator::THPControl::Enable);
LookupCache = fextl::make_unique<GuestToHostMap>();
}
CodeBuffer::~CodeBuffer() {
FEXCore::Allocator::VirtualFree(Ptr, Size);
FEXCore::Allocator::VirtualFree(Ptr, AllocatedSize);
}
auto CodeBufferManager::AllocateNew(size_t Size) -> fextl::shared_ptr<CodeBuffer> {
@@ -400,14 +403,20 @@ namespace CPU {
Latest = Buffer;
LatestOffset = 0;
OnCodeBufferAllocated(*Buffer);
OnCodeBufferAllocated(Buffer);
return Buffer;
}
fextl::shared_ptr<CodeBuffer> CodeBufferManager::GetLatest() {
if (!Latest) {
AllocateNew(INITIAL_CODE_SIZE);
if (FEXCore::Config::Get_ENABLECODECACHINGWIP()) {
// Start with a larger code buffer to avoid resizes that would discard
// code loaded from caches
AllocateNew(MAX_CODE_SIZE);
} else {
AllocateNew(INITIAL_CODE_SIZE);
}
}
return Latest;
}
@@ -418,7 +427,7 @@ namespace CPU {
return GetLatest();
}
auto NewCodeBufferSize = GetLatest()->Size;
auto NewCodeBufferSize = GetLatest()->AllocatedSize;
NewCodeBufferSize = std::min<size_t>(NewCodeBufferSize * 2, MAX_CODE_SIZE);
return AllocateNew(NewCodeBufferSize);
}
@@ -428,7 +437,7 @@ namespace CPU {
auto CheckCodeBuffer = [](CodeBuffer& Buffer, uintptr_t Address) {
// The last page of the code buffer is protected, so we need to exclude it from the valid range
// when checking if the address is in the code buffer.
uintptr_t LastPageAddr = AlignDown(reinterpret_cast<uintptr_t>(Buffer.Ptr) + Buffer.Size - 1, FEXCore::Utils::FEX_PAGE_SIZE);
uintptr_t LastPageAddr = AlignDown(reinterpret_cast<uintptr_t>(Buffer.Ptr) + Buffer.AllocatedSize - 1, FEXCore::Utils::FEX_PAGE_SIZE);
return (Address >= reinterpret_cast<uintptr_t>(Buffer.Ptr) && Address < LastPageAddr);
};
+8 -3
View File
@@ -43,7 +43,7 @@ struct GuestToHostMap;
namespace CPU {
struct CodeBuffer {
uint8_t* Ptr;
size_t Size;
size_t AllocatedSize; // including guard page; see UsableSize()
fextl::unique_ptr<GuestToHostMap> LookupCache;
@@ -54,6 +54,11 @@ namespace CPU {
CodeBuffer& operator=(CodeBuffer&&) = delete;
~CodeBuffer();
/// Returns the number of bytes available for storing code
size_t UsableSize() const {
return AllocatedSize - FEXCore::Utils::FEX_PAGE_SIZE;
}
};
/**
@@ -81,7 +86,7 @@ namespace CPU {
// Protects writes to the latest CodeBuffer and changes to LatestOffset
FEXCore::ForkableUniqueMutex CodeBufferWriteMutex;
virtual void OnCodeBufferAllocated(CodeBuffer&) {};
virtual void OnCodeBufferAllocated(const std::shared_ptr<CodeBuffer>&) {};
private:
fextl::shared_ptr<CodeBuffer> Latest;
@@ -161,7 +166,7 @@ namespace CPU {
virtual CompiledCode CompileCode(uint64_t Entry, uint64_t Size, bool SingleInst, const FEXCore::IR::IRListView* IR,
FEXCore::Core::DebugData* DebugData, bool CheckTF) = 0;
virtual fextl::vector<FEXCore::CPU::Relocation> TakeRelocations() = 0;
virtual fextl::vector<FEXCore::CPU::Relocation> TakeRelocations(uint64_t GuestBaseAddress) = 0;
virtual void ClearCache() {}
+140 -19
View File
@@ -14,6 +14,7 @@ $end_info$
#include <FEXCore/Core/CPUID.h>
#include <FEXCore/Core/HostFeatures.h>
#include <FEXCore/Utils/FileLoading.h>
#include <FEXCore/Utils/MathUtils.h>
#include <FEXCore/fextl/string.h>
#include <FEXHeaderUtils/Syscalls.h>
@@ -23,7 +24,7 @@ $end_info$
namespace FEXCore {
namespace ProductNames {
#ifdef _M_ARM_64
#ifdef ARCHITECTURE_arm64
static const char ARM_UNKNOWN[] = "Unknown ARM CPU";
static const char ARM_A57[] = "Cortex-A57";
static const char ARM_A72[] = "Cortex-A72";
@@ -91,10 +92,11 @@ namespace ProductNames {
static const char ARM_AppleSilicon[] = "Apple Silicon";
static const char ARM_ORYON_1[] = "Oryon-1";
static const char ARM_ORYON_3[] = "Oryon-3";
static const char ARM_Ampere_1[] = "AmpereOne";
static const char ARM_Ampere_1A[] = "AmpereOneA";
static const char ARM_Ampere_1B[] = "AmpereOneB";
#else
static const char ARM_Ampere_1C[] = "AmpereOneC";
#endif
} // namespace ProductNames
@@ -139,8 +141,8 @@ constexpr uint32_t FAMILY_IDENTIFIER = GenerateFamily(CPUFamily {
});
#endif
#ifdef _M_ARM_64
uint32_t GetCycleCounterFrequency() {
#ifdef ARCHITECTURE_arm64
uint64_t GetCycleCounterFrequency() {
uint64_t Result {};
__asm("mrs %[Res], CNTFRQ_EL0" : [Res] "=r"(Result));
return Result;
@@ -153,6 +155,7 @@ uint32_t GetCPUID_TPIDRRO() {
}
void CPUIDEmu::SetupHostHybridFlag() {
FEX_CONFIG_OPT(HideHybrid, HIDEHYBRID);
PerCPUData.resize(Cores);
uint64_t MIDR {};
@@ -169,6 +172,11 @@ void CPUIDEmu::SetupHostHybridFlag() {
MIDR = NewMIDR;
}
if (HideHybrid()) {
// Hide the hybrid flag.
Hybrid = false;
}
struct CPUMIDR {
uint8_t Implementer;
uint16_t Part;
@@ -179,8 +187,9 @@ void CPUIDEmu::SetupHostHybridFlag() {
// CPU priority order
// This is mostly arbitrary but will sort by some sort of CPU priority by performance
// Relative list so things they will commonly end up in big.little configurations sort of relate
static constexpr std::array<CPUMIDR, 66> CPUMIDRs = {{
static constexpr std::array<CPUMIDR, 68> CPUMIDRs = {{
// Typically big CPU cores
{0x51, 0x002, 1, ProductNames::ARM_ORYON_3}, // Qualcomm Oryon-3
{0x51, 0x001, 1, ProductNames::ARM_ORYON_1}, // Qualcomm Oryon-1
{0x61, 0x039, 1, ProductNames::ARM_Avalanche_M2Max}, // Apple Avalanche (M2 Max)
@@ -228,6 +237,7 @@ void CPUIDEmu::SetupHostHybridFlag() {
{0xc0, 0xac3, 1, ProductNames::ARM_Ampere_1}, // AmpereOne
{0xc0, 0xac4, 1, ProductNames::ARM_Ampere_1A}, // AmpereOneA
{0xc0, 0xac5, 1, ProductNames::ARM_Ampere_1B}, // AmpereOneB
{0xc0, 0xac7, 1, ProductNames::ARM_Ampere_1C}, // AmpereOneC
{0x4e, 0x010, 1, ProductNames::ARM_Olympus}, // Olympus
{0x4e, 0x004, 1, ProductNames::ARM_Carmel}, // Carmel
@@ -385,7 +395,8 @@ void CPUIDEmu::SetupHostHybridFlag() {
} else {
// If we aren't hybrid then just claim everything is big
for (size_t i = 0; i < Cores; ++i) {
uint32_t MIDR = PerCPUData[i].MIDR;
const auto MIDRIndex = HideHybrid() ? 0 : i;
uint32_t MIDR = PerCPUData[MIDRIndex].MIDR;
auto MIDROption = FindDefinedMIDR(MIDR);
PerCPUData[i].IsBig = true;
@@ -399,7 +410,7 @@ void CPUIDEmu::SetupHostHybridFlag() {
}
#else
uint32_t GetCycleCounterFrequency() {
uint64_t GetCycleCounterFrequency() {
return 0;
}
@@ -443,10 +454,10 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_01h(uint32_t Leaf) const {
Res.eax = FAMILY_IDENTIFIER;
Res.ebx = 0 | // Brand index
(8 << 8) | // Cache line size in bytes
(Cores << 16) | // Number of addressable IDs for the logical cores in the physical CPU
(0 << 24); // Local APIC ID
Res.ebx = 0 | // Brand index
(8 << 8) | // Cache line size in bytes
(Cores << 16) | // Number of addressable IDs for the logical cores in the physical CPU
(GetCPUID() << 24); // Local APIC ID
Res.ecx = (1 << 0) | // SSE3
(CTX->HostFeatures.SupportsPMULL_128Bit << 1) | // PCLMULQDQ
@@ -509,7 +520,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_01h(uint32_t Leaf) const {
(1 << 25) | // SSE
(1 << 26) | // SSE2
(0 << 27) | // Self Snoop
(1 << 28) | // Max APIC IDs reserved field is valid
(0 << 28) | // (HTT) Max APIC IDs reserved field is valid
(1 << 29) | // Thermal monitor
(0 << 30) | // Reserved
(0 << 31); // Pending break enable
@@ -747,6 +758,95 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_07h(uint32_t Leaf) const {
(0 << 29) | // Arch capabilities - Speculative side channel mitigations
(0 << 30) | // Arch capabilities - MSR module specific
(0 << 31); // SSBD - Speculative Store Bypass Disable
} else if (Leaf == 1) {
Res.eax = (0U << 0) | // SHA512
(0U << 1) | // SM3
(0U << 2) | // SM4
(0U << 3) | // RAO_INT
(0U << 4) | // AVX_VNNI
(0U << 5) | // AVX512_BF16
(0U << 6) | // LASS (Linear Address Space Separation)
(0U << 7) | // CMPCCXADD
(0U << 8) | // ARCH_PERFMON_EXT
(0U << 9) | // Reserved
(0U << 10) | // FAST_REP_MOVSB
(0U << 11) | // FAST_REP_STOSB
(0U << 12) | // FAST_REP_CMPSB_SCASB
(0U << 13) | // Reserved
(0U << 14) | // Reserved
(0U << 15) | // Reserved
(0U << 16) | // Reserved
(0U << 17) | // FRED (Flexible Return and Event Delivery)
(0U << 18) | // LKGS (Load into Kernel GS Base)
(0U << 19) | // WRMSRNS
(0U << 20) | // NMI_SRC
(0U << 21) | // AMX_FP16
(0U << 22) | // HRESET
(0U << 23) | // AVX_IFMA
(0U << 24) | // Reserved
(0U << 25) | // Reserved
(0U << 26) | // LAM (Linear Address Masking)
(0U << 27) | // MSRLIST
(0U << 28) | // Reserved
(0U << 29) | // Reserved
(0U << 30) | // INVD_DISABLE_POST_BIOS_DONE
(0U << 31); // MOVRS
// Bits 4-31 currently reserved.
Res.ebx = (0U << 0) | // PPIN
(0U << 1) | // PBNDKB
(0U << 2) | // Reserved
(0U << 3); // CPUIDMAXVAL_LIM_RMV
// Bits 6-31 also reserved.
Res.ecx = (0U << 0) | // RDT_M_ASYM
(0U << 1) | // RDT_A_ASYM
(0U << 2) | // Reserved
(0U << 3) | // Reserved
(0U << 4) | // Reserved
(0U << 5); // MSR_IMM
// Bits 25-31 also reserved.
Res.edx = (0U << 0) | // Reserved
(0U << 1) | // Reserved
(0U << 2) | // Reserved
(0U << 3) | // Reserved
(0U << 4) | // AVX_VNNI_INT8
(0U << 5) | // AVX_NE_CONVERT
(0U << 6) | // Reserved
(0U << 7) | // Reserved
(0U << 8) | // AMX_COMPLEX
(0U << 9) | // Reserved
(0U << 10) | // AVX_VNNI_INT16
(0U << 11) | // Reserved
(0U << 12) | // Reserved
(0U << 13) | // UTMR (User-timer events)
(0U << 14) | // PREFETCHI
(0U << 15) | // USER_MSR
(0U << 16) | // Reserved
(0U << 17) | // UIRET_UIF
(0U << 18) | // CET_SSS
(0U << 19) | // AVX10
(0U << 20) | // Reserved
(0U << 21) | // APX_F
(0U << 22) | // SEC-TEE_ATTESTATION
(0U << 23) | // MWAIT
(0U << 24); // SLSM (Static LSM)
} else if (Leaf == 2) {
// All bits are reserved except for EDX
Res.eax = 0;
Res.ebx = 0;
Res.ecx = 0;
// Bits 8-31 are reserved.
Res.edx = (0U << 0) | // PSFD
(0U << 1) | // IPRED_CTRL
(0U << 2) | // RRSBA_CTRL
(0U << 3) | // DDPD_U
(0U << 4) | // BHI_CTRL
(0U << 5) | // MCDT_NO
(0U << 6) | // UC_LOCK_DISABLE
(0U << 7); // MONITOR_MITG_NO
}
return Res;
@@ -804,7 +904,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_0Dh(uint32_t Leaf) const {
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_15h(uint32_t Leaf) const {
FEXCore::CPUID::FunctionResults Res {};
// TSC frequency = ECX * EBX / EAX
uint32_t FrequencyHz = GetCycleCounterFrequency();
uint64_t FrequencyHz = GetCycleCounterFrequency();
if (FrequencyHz) {
Res.eax = 1;
Res.ebx = 1U << CTX->Config.TSCScale;
@@ -825,6 +925,27 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_1Ah(uint32_t Leaf) const {
return Res;
}
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_24h(uint32_t Leaf) const {
FEXCore::CPUID::FunctionResults Res {};
if (Leaf == 0) {
// EAX indicates the maximum number of subleaves.
Res.eax = 0;
// Bits 19-31 reserved
// NOTE: We return all zero here until we have a CPU with AVX10
// even if some of the fields otherwise have fixed values.
Res.ebx = (0U << 0) | // (bits 0-7 specify the vector ISA version)
(0U << 16); // Defined as always 0b111
// All bits reserved
Res.ecx = 0;
Res.edx = 0;
}
return Res;
}
// Hypervisor CPUID information leaf
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_4000_0000h(uint32_t Leaf) const {
FEXCore::CPUID::FunctionResults Res {};
@@ -856,10 +977,10 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_4000_0001h(uint32_t Leaf) con
constexpr uint32_t MaximumSubLeafNumber = 2;
if (Leaf == 0) {
// EAX[3:0] Is the host architecture that FEX is running under
#ifdef _M_X86_64
#ifdef ARCHITECTURE_x86_64
// EAX[3:0] = 1 = x86_64 host architecture
Res.eax |= 0b0001;
#elif defined(_M_ARM_64)
#elif defined(ARCHITECTURE_arm64)
// EAX[3:0] = 2 = AArch64 host architecture
Res.eax |= 0b0010;
#else
@@ -1096,9 +1217,9 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0008h(uint32_t Leaf) con
(CTX->HostFeatures.SupportsCLZERO << 0); // CLZERO support
uint32_t CoreCount = Cores - 1;
Res.ecx = (0 << 16) | // PerfTscSize: Performance timestamp count size
((uint32_t)std::log2(CoreCount + 1) << 12) | // ApicIdSize: Number of bits in ApicID
(CoreCount << 0); // Count count subtract one
Res.ecx = (0 << 16) | // PerfTscSize: Performance timestamp count size
(std::bit_ceil(Cores) << 12) | // ApicIdSize: Number of bits in ApicID
(CoreCount << 0); // Count count subtract one
return Res;
}
@@ -1229,7 +1350,7 @@ CPUIDEmu::CPUIDEmu(const FEXCore::Context::ContextImpl* ctx)
SetupFeatures();
#ifdef _M_ARM_64
#ifdef ARCHITECTURE_arm64
if (SupportsCPUIndexInTPIDRRO) {
GetCPUID = GetCPUID_TPIDRRO;
}
+86 -4
View File
@@ -14,7 +14,7 @@ namespace Context {
class ContextImpl;
}
uint32_t GetCycleCounterFrequency();
uint64_t GetCycleCounterFrequency();
// Debugging define to switch what family of CPU we execute as.
// Might be useful if an application makes an assumption about a CPU.
@@ -159,7 +159,7 @@ private:
struct CPUData {
const char* ProductName {};
#ifdef _M_ARM_64
#ifdef ARCHITECTURE_arm64
uint32_t MIDR {};
#endif
bool IsBig {};
@@ -176,6 +176,7 @@ private:
FEXCore::CPUID::FunctionResults Function_0Dh(uint32_t Leaf) const;
FEXCore::CPUID::FunctionResults Function_15h(uint32_t Leaf) const;
FEXCore::CPUID::FunctionResults Function_1Ah(uint32_t Leaf) const;
FEXCore::CPUID::FunctionResults Function_24h(uint32_t Leaf) const;
FEXCore::CPUID::FunctionResults Function_4000_0000h(uint32_t Leaf) const;
FEXCore::CPUID::FunctionResults Function_4000_0001h(uint32_t Leaf) const;
FEXCore::CPUID::FunctionResults Function_8000_0000h(uint32_t Leaf) const;
@@ -200,7 +201,7 @@ private:
void SetupHostHybridFlag();
void SetupFeatures();
static constexpr size_t PRIMARY_FUNCTION_COUNT = 27;
static constexpr size_t PRIMARY_FUNCTION_COUNT = 37;
static constexpr size_t HYPERVISOR_FUNCTION_COUNT = 2;
static constexpr size_t EXTENDED_FUNCTION_COUNT = 32;
static constexpr std::array<FunctionHandler, PRIMARY_FUNCTION_COUNT> Primary = {
@@ -268,7 +269,48 @@ private:
#ifndef CPUID_AMD
// 0x1A: Hybrid Information Sub-leaf
&CPUIDEmu::Function_1Ah,
// 0x1B: PCONFIG info
&CPUIDEmu::Function_Reserved,
// 0x1C: Last Branch Records (LBR) info
&CPUIDEmu::Function_Reserved,
// 0x1D: Tile info
&CPUIDEmu::Function_Reserved,
// 0x1E: TMUL info
&CPUIDEmu::Function_Reserved,
// 0x1F: V2 Extended topology
&CPUIDEmu::Function_Reserved,
// 0x20: Processor History Reset info
&CPUIDEmu::Function_Reserved,
// 0x21: Unimplemented
&CPUIDEmu::Function_Reserved,
// 0x22: Reserved
&CPUIDEmu::Function_Reserved,
// 0x23: Architectural Performance Monitoring Extended
&CPUIDEmu::Function_Reserved,
// 0x24: Converged Vector ISA
&CPUIDEmu::Function_24h,
#else
// 0x1A: Reserved
&CPUIDEmu::Function_Reserved,
// 0x1B: Reserved
&CPUIDEmu::Function_Reserved,
// 0x1C: Reserved
&CPUIDEmu::Function_Reserved,
// 0x1D: Reserved
&CPUIDEmu::Function_Reserved,
// 0x1E: Reserved
&CPUIDEmu::Function_Reserved,
// 0x1F: Reserved
&CPUIDEmu::Function_Reserved,
// 0x20: Reserved
&CPUIDEmu::Function_Reserved,
// 0x21: Reserved
&CPUIDEmu::Function_Reserved,
// 0x22: Reserved
&CPUIDEmu::Function_Reserved,
// 0x23: Reserved
&CPUIDEmu::Function_Reserved,
// 0x24: Reserved
&CPUIDEmu::Function_Reserved,
#endif
};
@@ -277,7 +319,7 @@ private:
// 0: Highest function parameter and ID
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 1: Processor info
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
{SupportsConstant::NONCONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 2: Cache and TLB info
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 3: Serial Number(previously), now reserved
@@ -340,9 +382,49 @@ private:
#ifndef CPUID_AMD
// 0x1A: Hybrid Information Sub-leaf
{SupportsConstant::NONCONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x1B: PCONFIG info
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x1C: Last Branch Records (LBR) info
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x1D: Tile info
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x1E: TMUL info
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x1F: V2 Extended topology
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x20: Processor History Reset info
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x21: Unimplemented/Reserved
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x22: Reserved
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x23: Architectural Performance Monitoring Extended
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x24: Converged Vector ISA
{SupportsConstant::CONSTANT, NeedsLeafConstant::NEEDSLEAFCONSTANT},
#else
// 0x1A: Reserved
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x1B: Reserved
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x1C: Reserved
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x1D: Reserved
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x1E: Reserved
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x1F: Reserved
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x20: Reserved
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x21: Reserved
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x22: Reserved
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x23: Reserved
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x24: Reserved
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
#endif
}};
+630 -4
View File
@@ -1,12 +1,218 @@
// SPDX-License-Identifier: MIT
#include <Interface/Context/Context.h>
#include <FEXCore/Utils/SpinWaitLock.h>
#include <Interface/Context/Context.h>
#include <Interface/Core/ArchHelpers/Arm64Emitter.h>
#include <Interface/Core/Dispatcher/Dispatcher.h>
#include <Interface/Core/JIT/DebugData.h>
#include <Interface/Core/JIT/Relocations.h>
#include <Interface/Core/LookupCache.h>
#include <Interface/Core/OpcodeDispatcher.h>
#include <Interface/IR/PassManager.h>
#include <FEXCore/Core/Thunks.h>
#include <FEXCore/HLE/SourcecodeResolver.h>
#include <FEXCore/HLE/SyscallHandler.h>
#include <FEXHeaderUtils/Filesystem.h>
#include <git_version.h>
#include <xxhash.h>
#include <fstream>
namespace FEXCore {
#if __clang_major__ < 16
ExecutableFileInfo::ExecutableFileInfo(fextl::unique_ptr<HLE::SourcecodeMap> Map, uint64_t FileId, fextl::string Filename)
: SourcecodeMap(std::move(Map))
, FileId(FileId)
, Filename(Filename) {}
#endif
ExecutableFileInfo::~ExecutableFileInfo() = default;
fextl::string CodeMap::GetBaseFilename(const ExecutableFileInfo& MainExecutable, bool AddNombSuffix) {
auto FileId = MainExecutable.FileId;
std::string_view base_filename = FHU::Filesystem::GetFilename(std::string_view {MainExecutable.Filename});
if (FileId != 0xffff'ffff'ffff'ffff) {
return fextl::fmt::format("{}-{:016x}{}", base_filename, MainExecutable.FileId, AddNombSuffix ? "-nomb" : "");
}
return "";
}
fextl::map<CodeMapFileId, CodeMap::ParsedContents> CodeMap::ParseCodeMap(std::ifstream& File) {
fextl::map<CodeMapFileId, CodeMap::ParsedContents> Ret;
while (true) {
Entry Entry;
File.read(reinterpret_cast<char*>(&Entry), sizeof(Entry));
if (!File) {
break;
}
if (Entry.FileId == LoadExternalLibrary.FileId && Entry.BlockOffset == LoadExternalLibrary.BlockOffset) {
ExternalLibraryInfo Info;
File.read(reinterpret_cast<char*>(&Info), sizeof(Info));
fextl::string Filename;
std::getline(File, Filename, '\0');
// Align to 4-byte boundary
char Null[4];
File.read(Null, AlignUp(Filename.size() + 1, 4) - Filename.size() - 1);
if (!File) {
break;
}
Ret[Info.ExternalFileId].Filename = std::move(Filename);
} else if (Entry.FileId == SetExecutableFileId {}.Marker.FileId && Entry.BlockOffset == SetExecutableFileId {}.Marker.BlockOffset) {
CodeMapFileId ExecutableFileId;
File.read(reinterpret_cast<char*>(&ExecutableFileId), sizeof(ExecutableFileId));
if (!File) {
break;
}
Ret[ExecutableFileId].IsExecutable = true;
} else {
if (!Ret.contains(Entry.FileId)) {
LogMan::Msg::EFmt("Code map referenced unknown file id {:016x}", Entry.FileId);
} else {
Ret[Entry.FileId].Blocks.insert(Entry.BlockOffset);
}
}
if (!File) {
break;
}
}
return Ret;
}
CodeMapWriter::CodeMapWriter(CodeMapOpener& Opener, bool OpenEagerly)
: Buffer(4096)
, FileOpener(Opener) {
if (OpenEagerly) {
CodeMapFD = FileOpener.OpenCodeMapFile();
}
}
CodeMapWriter::~CodeMapWriter() {
if (CodeMapFD.value_or(-1) != -1) {
Flush(BufferOffset);
close(*CodeMapFD);
}
}
bool CodeMapWriter::IsWriteEnabled(const ExecutableFileSectionInfo& Section) {
if (CodeMapFD == -1) {
return false;
}
// PV libraries can't yet be read by FEXServer, so skip dumping them
if (Section.FileInfo.Filename.starts_with("/run/pressure-vessel")) {
return false;
}
if (CodeMapFD) {
return true;
}
// Acquire mutex and re-check CodeMapFD to avoid race conditions
auto lk = std::unique_lock {Mutex};
if (!CodeMapFD) {
CodeMapFD = FileOpener.OpenCodeMapFile();
}
return CodeMapFD != -1;
}
void CodeMapWriter::Flush(size_t Offset) {
// Acquire exclusive lock and flush circular buffer
std::unique_lock Lock {Mutex};
Flush(Offset, Lock);
}
void CodeMapWriter::Flush(size_t Offset, std::unique_lock<std::shared_mutex>&) {
write(*CodeMapFD, Buffer.data(), Offset);
BufferOffset = 0;
}
void CodeMapWriter::AppendBlock(const FEXCore::ExecutableFileSectionInfo& SectionInfo, uint64_t BlockEntry) {
if (!IsWriteEnabled(SectionInfo)) {
return;
}
BlockEntry -= SectionInfo.FileStartVA;
if (BlockEntry > std::numeric_limits<uint32_t>::max()) {
ERROR_AND_DIE_FMT("Cannot write code map");
}
// Register new library if not already known
bool NewLibraryLoad = false;
{
// Check prior registration with shared lock
std::shared_lock Lock {Mutex};
NewLibraryLoad = !KnownFileIds.contains(SectionInfo.FileInfo.FileId);
}
if (NewLibraryLoad) {
// Register to map with exclusive lock
std::unique_lock Lock {Mutex};
NewLibraryLoad &= KnownFileIds.insert(SectionInfo.FileInfo.FileId).second;
}
if (NewLibraryLoad) {
// Add entry to code map
AppendLibraryLoad(SectionInfo.FileInfo);
}
// Register the actual code block
CodeMap::Entry DataEntry {SectionInfo.FileInfo.FileId, static_cast<uint32_t>(BlockEntry)};
AppendData(std::as_bytes(std::span {&DataEntry, 1}));
}
void CodeMapWriter::AppendLibraryLoad(const FEXCore::ExecutableFileInfo& FileInfo) {
// See CodeMap::ExternalLibraryInfo
auto ExternalFileId = FileInfo.FileId;
auto TotalSize = AlignUp(sizeof(CodeMap::LoadExternalLibrary) + sizeof(ExternalFileId) + FileInfo.Filename.size() + 1, 4);
const auto Data = reinterpret_cast<char*>(alloca(TotalSize));
auto WritePtr = std::copy_n(reinterpret_cast<const char*>(&CodeMap::LoadExternalLibrary), sizeof(CodeMap::LoadExternalLibrary), Data);
WritePtr = std::copy_n(reinterpret_cast<const char*>(&ExternalFileId), sizeof(ExternalFileId), WritePtr);
WritePtr = std::copy(FileInfo.Filename.begin(), FileInfo.Filename.end(), WritePtr);
std::fill(WritePtr, Data + TotalSize, 0);
AppendData(std::as_bytes(std::span {Data, TotalSize}));
}
void CodeMapWriter::AppendSetMainExecutable(const FEXCore::ExecutableFileInfo& FileInfo) {
CodeMap::SetExecutableFileId Data {.ExecutableFileId = FileInfo.FileId};
AppendData(std::span {reinterpret_cast<const std::byte*>(&Data), sizeof(Data)});
}
void CodeMapWriter::AppendData(std::span<const std::byte> Data) {
std::shared_lock Lock {Mutex};
auto Offset = BufferOffset.fetch_add(Data.size_bytes());
if (Offset + Data.size_bytes() > Buffer.size()) {
// Acquire exclusive lock and flush the buffer.
// Under heavy pressure, multiple threads may observe an exhausted buffer simultaneously.
// The thread with the last in-bounds Offset is responsible for flushing the buffer.
Lock.unlock();
bool IsResponsibleForFlush = false;
{
std::unique_lock ExclusiveLock {Mutex};
IsResponsibleForFlush = (Offset <= Buffer.size());
if (IsResponsibleForFlush) {
Flush(Offset, ExclusiveLock);
}
}
if (!IsResponsibleForFlush) {
// Wait for the buffer to be flushed on the responsible thread
Utils::SpinWaitLock::WaitPred<std::less_equal<>, size_t>(reinterpret_cast<size_t*>(&BufferOffset), Buffer.size());
}
AppendData(Data);
return;
}
memcpy(&Buffer.at(Offset), Data.data(), Data.size_bytes());
}
} // namespace FEXCore
namespace FEXCore::Context {
@@ -15,12 +221,432 @@ CodeCache::CodeCache(ContextImpl& CTX_)
: CTX(CTX_) {}
CodeCache::~CodeCache() = default;
void CodeCache::LoadData(Core::InternalThreadState& Thread, std::byte* MappedCacheFile, const ExecutableFileSectionInfo& GuestRIPLookup) {
// TODO
uint64_t CodeCache::ComputeCodeMapId(std::string_view Filename, int FD) {
if (Filename.empty()) {
return 0xffff'ffff'ffff'ffff;
}
// For now, we just use the file path as an identifier.
// TODO: Ensure the hash is unique enough to distinguish executables while remaining independent of the installation location
return XXH3_64bits(Filename.data(), Filename.size());
}
struct CodeCacheHeader {
std::array<char, 4> Magic = ExpectedMagic;
uint32_t FormatVersion = 1;
uint8_t FEXVersion[20] = {};
uint32_t NumBlocks;
uint32_t NumCodePages;
uint32_t CodeBufferSize;
uint32_t NumRelocations;
uint32_t padding;
uint64_t SerializedBaseAddress;
// TODO: Consider including information from LookupCache.BlockLinks
static constexpr std::array<char, 4> ExpectedMagic = {'F', 'X', 'C', 'C'};
};
template<typename T>
concept OrderedContainer = requires { typename T::key_compare; };
bool CodeCache::SaveData(Core::InternalThreadState& Thread, int fd, const ExecutableFileSectionInfo& SourceBinary, uint64_t SerializedBaseAddress) {
// TODO
auto CodeBuffer = CTX.GetLatest();
auto& LookupCache = *Thread.LookupCache->Shared;
auto Relocations = Thread.CPUBackend->TakeRelocations(SourceBinary.FileStartVA);
// Write file header
CodeCacheHeader header {};
static_assert(GIT_HASH.size() == sizeof(header.FEXVersion));
std::ranges::copy(GIT_HASH, header.FEXVersion);
header.NumBlocks = LookupCache.BlockList.size();
header.NumCodePages = LookupCache.CodePages.size();
header.CodeBufferSize = CTX.LatestOffset;
header.NumRelocations = Relocations.size();
header.SerializedBaseAddress = SerializedBaseAddress;
::write(fd, &header, sizeof(header));
// Dump guest<->host block mappings
{
// Cache contents must be deterministic, so copy the unordered block list and then sort by key
static_assert(!OrderedContainer<decltype(LookupCache.BlockList)>, "Already deterministic; drop temporary container");
fextl::vector<std::pair<uint64_t, const GuestToHostMap::BlockEntry*>> BlockList;
BlockList.reserve(LookupCache.BlockList.size());
for (auto& [Guest, BlockEntry] : LookupCache.BlockList) {
static_assert(sizeof(Guest) == 8, "Breaking change in code cache data layout");
BlockList.emplace_back(Guest, &BlockEntry);
}
std::ranges::sort(BlockList);
for (auto [Guest, Host] : BlockList) {
static_assert(sizeof(Host->HostCode) == 8, "Breaking change in code cache data layout");
static_assert(sizeof(Host->CodePages[0]) == 8, "Breaking change in code cache data layout");
Guest -= SourceBinary.FileStartVA;
::write(fd, &Guest, sizeof(Guest));
uint64_t HostCode = Host->HostCode - reinterpret_cast<uintptr_t>(CodeBuffer->Ptr);
::write(fd, &HostCode, sizeof(HostCode));
uint64_t NumCodePages = Host->CodePages.size();
::write(fd, &NumCodePages, sizeof(NumCodePages));
LOGMAN_THROW_A_FMT(std::ranges::is_sorted(Host->CodePages), "Code pages aren't sorted");
for (auto CodePage : Host->CodePages) {
CodePage -= SourceBinary.FileStartVA;
::write(fd, &CodePage, sizeof(CodePage));
}
}
}
// Dump relocations
static_assert(sizeof(Relocations[0]) == 48, "Breaking change in code cache data layout");
::write(fd, Relocations.data(), Relocations.size() * sizeof(Relocations[0]));
// Pad to next page in file so that the CodeBuffer can be mmap'ed into process on load
{
auto AlignedSize = AlignUp(lseek(fd, 0, SEEK_CUR), Utils::FEX_PAGE_SIZE);
::ftruncate(fd, AlignedSize);
lseek(fd, AlignedSize, SEEK_SET);
}
// Dump the host code (relocated for position-independent serialization)
std::span CodeBufferData(reinterpret_cast<std::byte*>(CodeBuffer->Ptr), reinterpret_cast<std::byte*>(CodeBuffer->Ptr) + CTX.LatestOffset);
if (!ApplyCodeRelocations(SerializedBaseAddress, CodeBufferData, Relocations, true)) {
LOGMAN_THROW_A_FMT(false, "Failed to apply code relocations");
return false;
}
::write(fd, CodeBufferData.data(), CodeBufferData.size());
// Dump code pages
static_assert(OrderedContainer<decltype(LookupCache.CodePages)>, "Non-deterministic data source");
for (const auto& [PageIndex, Entrypoints] : LookupCache.CodePages) {
uint64_t PageAddr = (PageIndex << 12) - SourceBinary.FileStartVA;
::write(fd, &PageAddr, sizeof(PageAddr));
uint64_t NumEntrypoints = Entrypoints.size();
::write(fd, &NumEntrypoints, sizeof(NumEntrypoints));
for (uint64_t Entrypoint : Entrypoints) {
Entrypoint -= SourceBinary.FileStartVA;
::write(fd, &Entrypoint, sizeof(Entrypoint));
}
}
return true;
}
bool CodeCache::LoadData(Core::InternalThreadState* Thread, std::byte* MappedCacheFile, const ExecutableFileSectionInfo& BinarySection) {
if (!EnableCodeCaching) {
return true;
}
namespace ranges = std::ranges;
// Read file header
CodeCacheHeader header {};
::memcpy(&header, MappedCacheFile, sizeof(header));
MappedCacheFile += sizeof(header);
LogMan::Msg::IFmt("Cache load: {:5} blocks; base={:#14x}; off={:#9x}-{:#09x}; {:016x} {}", header.NumBlocks, BinarySection.FileStartVA,
BinarySection.BeginVA - BinarySection.FileStartVA, BinarySection.EndVA - BinarySection.FileStartVA,
BinarySection.FileInfo.FileId, BinarySection.FileInfo.Filename);
if (!ranges::equal(header.Magic, header.ExpectedMagic)) {
LogMan::Msg::EFmt("Invalid cache file header");
return false;
}
if (!ranges::equal(header.FEXVersion, GIT_HASH)) {
LogMan::Msg::IFmt("Cache generated from old FEX version {:02x}, current is {:02x}; skipping", fmt::join(header.FEXVersion, ""),
fmt::join(GIT_HASH, ""));
return false;
}
if (header.NumBlocks == 0) {
// Valid caches are never empty
LogMan::Msg::IFmt("Code cache empty, aborting");
return false;
}
// Read guest<->host block mappings
using BlockListEntry = decltype(GuestToHostMap::BlockList)::value_type;
fextl::vector<BlockListEntry> BlockList(header.NumBlocks);
{
for (auto& BlockPtr : BlockList) {
::memcpy(&BlockPtr.first, MappedCacheFile, sizeof(BlockPtr.first));
MappedCacheFile += sizeof(BlockPtr.first);
::memcpy(&BlockPtr.second.HostCode, MappedCacheFile, sizeof(BlockPtr.second.HostCode));
MappedCacheFile += sizeof(BlockPtr.second.HostCode);
uint64_t NumGuestPages;
::memcpy(&NumGuestPages, MappedCacheFile, sizeof(NumGuestPages));
MappedCacheFile += sizeof(NumGuestPages);
BlockPtr.second.CodePages.resize(NumGuestPages);
::memcpy(BlockPtr.second.CodePages.data(), MappedCacheFile, std::span {BlockPtr.second.CodePages}.size_bytes());
MappedCacheFile += std::span {BlockPtr.second.CodePages}.size_bytes();
}
// Constrain BlockList to the given ExecutableFileSectionInfo
LOGMAN_THROW_A_FMT(ranges::is_sorted(BlockList, [](auto& a, auto& b) { return a.first < b.first; }), "Expected sorted block list");
auto begin = ranges::lower_bound(BlockList, BinarySection.BeginVA - BinarySection.FileStartVA, std::less {}, &BlockListEntry::first);
auto end =
ranges::upper_bound(begin, BlockList.end(), BinarySection.EndVA - BinarySection.FileStartVA - 1, std::less {}, &BlockListEntry::first);
if (begin == end) {
// Not an error since there is just no data to load
LogMan::Msg::IFmt("No blocks cached in this range, aborting");
return true;
}
BlockList.erase(end, BlockList.end());
BlockList.erase(BlockList.begin(), begin);
}
// Read relocations
fextl::vector<FEXCore::CPU::Relocation> Relocations(header.NumRelocations, FEXCore::CPU::Relocation::Default());
::memcpy(Relocations.data(), MappedCacheFile, Relocations.size() * sizeof(Relocations[0]));
MappedCacheFile += Relocations.size() * sizeof(Relocations[0]);
// Pad to next page in file, which contains CodeBuffer data
MappedCacheFile = reinterpret_cast<std::byte*>(AlignUp(reinterpret_cast<uintptr_t>(MappedCacheFile), Utils::FEX_PAGE_SIZE));
// Prepare CodeBuffer: Page aligned and big enough to hold all cached data
auto Lock = std::unique_lock {CTX.CodeBufferWriteMutex};
if (Thread) {
if (auto Prev = Thread->CPUBackend->CheckCodeBufferUpdate()) {
Allocator::VirtualDontNeed(Thread->CallRetStackBase, FEXCore::Core::InternalThreadState::CALLRET_STACK_SIZE);
auto lk = Thread->LookupCache->AcquireWriteLock();
Thread->LookupCache->ChangeGuestToHostMapping(*Prev, *CTX.GetLatest()->LookupCache, lk);
}
}
auto CodeBuffer = CTX.GetLatest();
LOGMAN_THROW_A_FMT(reinterpret_cast<uintptr_t>(CodeBuffer->Ptr) % 0x1000 == 0, "Expected CodeBuffer base to be page-aligned");
const auto Delta = AlignUp(CTX.LatestOffset, 0x1000) - CTX.LatestOffset;
CTX.LatestOffset += Delta;
while (CTX.LatestOffset + header.CodeBufferSize > CodeBuffer->UsableSize()) {
if (Thread) {
CTX.ClearCodeCache(Thread);
CodeBuffer = CTX.GetLatest();
LogMan::Msg::IFmt("Increased code buffer size to {} MiB for cache load", CodeBuffer->AllocatedSize / 1024 / 1024);
} else {
ERROR_AND_DIE_FMT("Cannot extend codebuffer without thread!");
}
}
// Read CodeBuffer data from file. Make sure the destination is page-aligned.
// TODO: Only load the data needed for the selected section
auto CodeBufferRange =
std::as_writable_bytes(std::span {CodeBuffer->Ptr, CodeBuffer->UsableSize()}).subspan(CTX.LatestOffset, header.CodeBufferSize);
::memcpy(CodeBufferRange.data(), MappedCacheFile, header.CodeBufferSize);
MappedCacheFile += header.CodeBufferSize;
CTX.LatestOffset += header.CodeBufferSize;
// Apply FEX relocations
auto Ret = ApplyCodeRelocations(BinarySection.FileStartVA, CodeBufferRange, Relocations, false);
LOGMAN_THROW_A_FMT(Ret == true, "Failed to apply code cache relocations");
{
auto& LookupCache = *CodeBuffer->LookupCache;
auto WriteLock = LookupCache.AcquireWriteLock();
// Register blocks to LookupCache
for (auto& [Guest, Host] : BlockList) {
for (auto& CodePage : Host.CodePages) {
CodePage += BinarySection.FileStartVA;
}
auto HostCode = reinterpret_cast<void*>(Host.HostCode + reinterpret_cast<uintptr_t>(CodeBufferRange.data()));
LookupCache.AddBlockMapping(Guest + BinarySection.FileStartVA, std::move(Host.CodePages), HostCode, WriteLock);
}
// Register loaded code ranges
fextl::vector<uint64_t> Entrypoints;
for (uint32_t i = 0; i < header.NumCodePages; ++i) {
uint64_t CodePage;
memcpy(&CodePage, MappedCacheFile, sizeof(CodePage));
CodePage += BinarySection.FileStartVA;
MappedCacheFile += sizeof(CodePage);
uint64_t NumEntrypoints;
memcpy(&NumEntrypoints, MappedCacheFile, sizeof(NumEntrypoints));
MappedCacheFile += sizeof(NumEntrypoints);
Entrypoints.resize(NumEntrypoints);
memcpy(Entrypoints.data(), MappedCacheFile, NumEntrypoints * sizeof(Entrypoints[0]));
MappedCacheFile += NumEntrypoints * sizeof(Entrypoints[0]);
for (auto& Entrypoint : Entrypoints) {
Entrypoint += BinarySection.FileStartVA;
}
if (LookupCache.AddBlockExecutableRange(Entrypoints, CodePage, FEXCore::Utils::FEX_PAGE_SIZE, WriteLock)) {
CTX.SyscallHandler->MarkGuestExecutableRange(Thread, CodePage, FEXCore::Utils::FEX_PAGE_SIZE);
}
}
}
if (EnableCodeCacheValidation) {
fextl::set<uint64_t> GuestBlocks, HostBlocks;
for (auto& [Guest, Host] : BlockList) {
GuestBlocks.insert(Guest + BinarySection.FileStartVA);
HostBlocks.insert(Host.HostCode);
}
Validate(BinarySection, std::move(GuestBlocks), HostBlocks, CodeBufferRange);
}
return true;
}
void CodeCache::Validate(const ExecutableFileSectionInfo& Section, fextl::set<uint64_t> GuestBlocks, const fextl::set<uint64_t>& HostBlocks,
std::span<std::byte> CachedCode) {
LOGMAN_THROW_A_FMT(!HostBlocks.empty(), "Tried to validate without any host blocks");
// Skip any cached data before the first host block
CachedCode = CachedCode.subspan(*HostBlocks.begin() - sizeof(CPU::CPUBackend::JITCodeHeader));
if (!ValidationCTX) {
ValidationCTX.reset(static_cast<ContextImpl*>(FEXCore::Context::Context::CreateNewContext(CTX.HostFeatures).release()));
ValidationCTX->SetSignalDelegator(CTX.SignalDelegation);
ValidationCTX->SetSyscallHandler(CTX.SyscallHandler);
ValidationCTX->SetThunkHandler(CTX.ThunkHandler);
if (!ValidationCTX->InitCore()) {
ERROR_AND_DIE_FMT("Failed to create cache load validation context");
}
ValidationThread.reset(ValidationCTX->CreateThread(0, 0, nullptr));
auto Frame = ValidationThread->CurrentFrame;
Frame->State.segment_arrays[FEXCore::Core::CPUState::SEGMENT_ARRAY_INDEX_GDT] = &ValidationGDT[0];
Frame->State.segment_arrays[FEXCore::Core::CPUState::SEGMENT_ARRAY_INDEX_LDT] = &ValidationGDT[0];
Frame->State.cs_idx = 0;
Frame->State.cs_cached = 0;
if (ValidationCTX->Config.Is64BitMode()) {
ValidationGDT[0].L = 1; // L = Long Mode = 64-bit
ValidationGDT[0].D = 0; // D = Default Operand Size = Reserved
} else {
ValidationGDT[0].L = 0; // L = Long Mode = 32-bit
ValidationGDT[0].D = 1; // D = Default Operand Size = 32-bit
}
}
auto NewCodeBuffer = ValidationCTX->GetLatest();
while (CachedCode.size_bytes() > NewCodeBuffer->UsableSize()) {
ValidationCTX->ClearCodeCache(ValidationThread.get());
NewCodeBuffer = ValidationCTX->GetLatest();
LogMan::Msg::IFmt("Increased cache validation code buffer size to {} MiB", NewCodeBuffer->AllocatedSize / 1024 / 1024);
}
std::span<std::byte> CodeBufferRangeRef =
std::as_writable_bytes(std::span {NewCodeBuffer->Ptr, NewCodeBuffer->Ptr + NewCodeBuffer->UsableSize()}).subspan(0, CachedCode.size_bytes());
while (!GuestBlocks.empty()) {
auto [CompiledBlocks, _, _2, _3, _4] = ValidationCTX->CompileCode(ValidationThread.get(), *GuestBlocks.begin(), 0 /* TODO: Set MaxInst? */);
for (auto& Entry : CompiledBlocks.EntryPoints) {
GuestBlocks.erase(Entry.first);
}
}
// Patch FEX-internal function addresses with values from the main Context to ensure the code blocks are comparable
auto NewRelocations = ValidationThread->CPUBackend->TakeRelocations(Section.FileStartVA);
NewRelocations.erase(std::remove_if(NewRelocations.begin(), NewRelocations.end(), [](const CPU::Relocation& Reloc) {
return Reloc.Header.Type != CPU::RelocationTypes::RELOC_NAMED_SYMBOL_LITERAL && Reloc.Header.Type != CPU::RelocationTypes::RELOC_NAMED_THUNK_MOVE;
}));
(void)ApplyCodeRelocations(Section.FileStartVA, CodeBufferRangeRef, NewRelocations, false);
if (ValidationCTX->LatestOffset <= CodeBufferRangeRef.size()) {
// Reference compilation produced fewer bytes than our cache, so validation is going to fail.
// Make sure we don't output any garbage bytes though.
CodeBufferRangeRef = CodeBufferRangeRef.subspan(0, ValidationCTX->LatestOffset);
}
auto [Mismatch, _] = std::mismatch(CodeBufferRangeRef.begin(), CodeBufferRangeRef.end(), CachedCode.begin());
if (Mismatch != CodeBufferRangeRef.end()) {
// Align down to instruction size
auto Idx = AlignDown(std::distance(CodeBufferRangeRef.begin(), Mismatch), 4);
auto BlockIt = std::prev(HostBlocks.lower_bound(*HostBlocks.begin() + Idx + 1));
std::optional<uint64_t> GuestBlockAddr;
std::optional<uint64_t> GuestBlockAddrRef;
if (BlockIt != HostBlocks.end()) {
for (int i : {0, 1}) {
std::span Buffer = (i == 0 ? CachedCode : CodeBufferRangeRef);
// Second instruction is always a constant load for relative offset to the (multi)block start
int32_t addr = (*reinterpret_cast<uint32_t*>(&Buffer[*BlockIt - *HostBlocks.begin() + 4]) & 0x3ff'ffe0) << 11;
addr >>= 14;
auto header = reinterpret_cast<CPU::CPUBackend::JITCodeHeader*>(&Buffer[*BlockIt - *HostBlocks.begin() + 4 + addr]);
auto tail = reinterpret_cast<CPU::CPUBackend::JITCodeTail*>(reinterpret_cast<uintptr_t>(header) + header->OffsetToBlockTail);
(i == 0 ? GuestBlockAddr : GuestBlockAddrRef) = tail->RIP - Section.FileStartVA;
LogMan::Msg::EFmt("Recorded rip {}: {:#x} (offset {:#x})", i, tail->RIP, tail->RIP - Section.FileStartVA);
if (i == 1) {
if (tail->RIP >= Section.BeginVA && tail->RIP < Section.EndVA) {
auto [IRView, TotalInstructions, TotalInstructionsLength, StartAddr, Length, _] =
ValidationCTX->GenerateIR(ValidationThread.get(), tail->RIP, false, FEXCore::Config::Get_MAXINST());
fextl::stringstream ss;
FEXCore::IR::Dump(&ss, &*IRView);
LogMan::Msg::EFmt("IR:\n{}", ss.str());
} else {
LogMan::Msg::EFmt("Can't dump IR for out-of-range RIP {:#x}", tail->RIP);
}
}
}
}
fextl::string GuestBlockInfo = "UNKNOWN";
if (GuestBlockAddr) {
GuestBlockInfo = fextl::fmt::format("{:#x}", GuestBlockAddr.value());
}
if (GuestBlockAddr != GuestBlockAddrRef) {
GuestBlockInfo += " (MISMATCH)";
}
ERROR_AND_DIE_FMT("Cache validation failed at offset {:#x}: {:02x} <-> {:02x} (at {} <-> {}, guest block {})", Idx,
fmt::join(CachedCode.subspan(Idx, 4), ""), fmt::join(CodeBufferRangeRef.subspan(Idx, 4), ""),
fmt::ptr(CachedCode.data()), fmt::ptr(CodeBufferRangeRef.data()), GuestBlockInfo);
}
// Reset Context state for next validation
ValidationThread->LookupCache->ClearCache(ValidationThread->LookupCache->AcquireWriteLock());
ValidationCTX->LatestOffset = 0;
LogMan::Msg::IFmt("\tSuccessfully validated cache");
}
bool CodeCache::ApplyCodeRelocations(uint64_t GuestEntry, std::span<std::byte> Code,
std::span<const FEXCore::CPU::Relocation> EntryRelocations, bool ForStorage) {
CPU::Arm64Emitter Emitter(&CTX, Code.data(), Code.size_bytes());
for (size_t j = 0; j < EntryRelocations.size(); ++j) {
const FEXCore::CPU::Relocation& Reloc = EntryRelocations[j];
Emitter.SetCursorOffset(Reloc.Header.Offset);
switch (Reloc.Header.Type) {
case FEXCore::CPU::RelocationTypes::RELOC_NAMED_SYMBOL_LITERAL: {
// Generate a literal so we can place it
uint64_t Pointer = ForStorage ? 0 : GetNamedSymbolLiteral(CTX, Reloc.NamedSymbolLiteral.Symbol);
Emitter.dc64(Pointer);
break;
}
case FEXCore::CPU::RelocationTypes::RELOC_NAMED_THUNK_MOVE: {
uint64_t Pointer = ForStorage ? 0 : reinterpret_cast<uint64_t>(CTX.ThunkHandler->LookupThunk(Reloc.NamedThunkMove.Symbol));
if (Pointer == ~0ULL) {
return false;
}
// TODO: Pointers are required to fit within 48-bit VA space.
// But forcing 6-byte broke relocations.
Emitter.LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Register(Reloc.NamedThunkMove.RegisterIndex), Pointer,
CPU::Arm64Emitter::PadType::DOPAD);
break;
}
case FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_LITERAL: {
Emitter.dc64(GuestEntry + Reloc.GuestRIP.GuestRIP);
break;
}
case FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_MOVE: {
uint64_t Pointer = Reloc.GuestRIP.GuestRIP + GuestEntry;
// TODO: Pointers are required to fit within 48-bit VA space.
// But forcing 6-byte broke relocations.
Emitter.LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Register(Reloc.GuestRIP.RegisterIndex), Pointer, CPU::Arm64Emitter::PadType::DOPAD);
break;
}
default: ERROR_AND_DIE_FMT("Unknown relocation type {}", ToUnderlying(Reloc.Header.Type));
}
}
return true;
}
+104 -49
View File
@@ -9,6 +9,9 @@ $end_info$
*/
#include <cstdint>
#ifdef ZYDIS_DISASSEMBLER
#include <Zydis/Zydis.h>
#endif
#include "Interface/Core/ArchHelpers/Arm64Emitter.h"
#include "Interface/Core/LookupCache.h"
#include "Interface/Core/CPUBackend.h"
@@ -27,7 +30,7 @@ $end_info$
#include "Interface/IR/RegisterAllocationData.h"
#include "Utils/Allocator.h"
#include "Utils/Allocator/HostAllocator.h"
#include "Utils/SpinWaitLock.h"
#include <FEXCore/Utils/SpinWaitLock.h>
#include "Utils/variable_length_integer.h"
#include <FEXCore/Config/Config.h>
@@ -345,7 +348,7 @@ bool ContextImpl::InitCore() {
// Set up the SignalDelegator config since core is initialized.
SignalDelegation->SetConfig(Dispatcher->MakeSignalDelegatorConfig());
#if defined(_WIN32) && !defined(_M_ARM_64EC)
#if defined(_WIN32) && !defined(ARCHITECTURE_arm64ec)
// WOW64 always needs the interrupt fault check to be enabled.
Config.NeedsPendingInterruptFaultCheck = true;
#endif
@@ -363,6 +366,9 @@ void ContextImpl::HandleCallback(FEXCore::Core::InternalThreadState* Thread, uin
}
void ContextImpl::ExecuteThread(FEXCore::Core::InternalThreadState* Thread) {
// Update the thread pointer for Thunk return to the latest.
Thread->CurrentFrame->Pointers.ThunkCallbackRet = SignalDelegation->GetThunkCallbackRET();
Dispatcher->ExecuteDispatch(Thread->CurrentFrame);
// If it is the parent thread that died then just leave
@@ -379,7 +385,7 @@ void ContextImpl::InitializeCompiler(FEXCore::Core::InternalThreadState* Thread)
Thread->CurrentFrame->State.L1Pointer = Thread->LookupCache->GetL1Pointer();
Thread->CurrentFrame->State.L1Mask = Thread->LookupCache->GetScaledL1PointerMask();
Thread->CurrentFrame->Pointers.Common.L2Pointer = Thread->LookupCache->GetPagePointer();
Thread->CurrentFrame->Pointers.L2Pointer = Thread->LookupCache->GetPagePointer();
Dispatcher->InitThreadPointers(Thread);
@@ -437,6 +443,10 @@ void ContextImpl::UnlockAfterFork(FEXCore::Core::InternalThreadState* LiveThread
Profiler::PostForkAction(Child);
if (Child) {
if (CodeMapWriter) {
CodeMapWriter->ResetAfterFork();
}
CodeInvalidationMutex.StealAndDropActiveLocks();
if (Config.StrictInProcessSplitLocks) {
StrictSplitLockMutex = 0;
@@ -459,9 +469,14 @@ void ContextImpl::LockBeforeFork(FEXCore::Core::InternalThreadState* Thread) {
}
#endif
void ContextImpl::OnCodeBufferAllocated(CPU::CodeBuffer& Buffer) {
void ContextImpl::OnCodeBufferAllocated(const fextl::shared_ptr<CPU::CodeBuffer>& Buffer) {
if (Config.GlobalJITNaming()) {
Symbols.RegisterJITSpace(Buffer.Ptr, Buffer.Size);
Symbols.RegisterJITSpace(Buffer->Ptr, Buffer->AllocatedSize);
}
{
std::scoped_lock lk {CodeBufferListLock};
CodeBufferList.emplace_back(Buffer);
}
}
@@ -487,11 +502,14 @@ static void IRDumper(FEXCore::Core::InternalThreadState* Thread, IR::IREmitter*
fextl::fmt::print(FD, "IR-ShouldDump-{} 0x{:x}:\n{}\n@@@@@\n", NewIR.PostRA() ? "post" : "pre", GuestRIP, out.str());
};
bool ContextImpl::CheckIfBlockIsCacheable(FEXCore::Core::InternalThreadState& Thread, uint64_t GuestRIP, uint64_t MaxInst) {
return Thread.FrontendDecoder->CheckIfCacheable(Thread, reinterpret_cast<const uint8_t*>(GuestRIP), GuestRIP, MaxInst);
}
ContextImpl::GenerateIRResult
ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP, bool ExtendedDebugInfo, uint64_t MaxInst) {
FEXCORE_PROFILE_SCOPED("GenerateIR");
Thread->OpDispatcher->ReownOrClaimBuffer();
Thread->OpDispatcher->ResetWorkingList();
uint64_t TotalInstructions {0};
@@ -527,9 +545,24 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
const auto GPRSize = Thread->OpDispatcher->GetGPROpSize();
#ifdef ZYDIS_DISASSEMBLER
const auto ZydisMachineMode = Config.Is64BitMode ? ZYDIS_MACHINE_MODE_LONG_64 : ZYDIS_MACHINE_MODE_LEGACY_32;
if (FEXCore::Config::Get_X86DISASSEMBLE()) {
const uint64_t DecodedMin = Thread->FrontendDecoder->DecodedMinAddress;
const uint64_t DecodedMax = Thread->FrontendDecoder->DecodedMaxAddress;
LogMan::Msg::IFmt("Guest x86 Begin (RIP={:#x}, {:#x}-{:#x})", GuestRIP, DecodedMin, DecodedMax);
}
#endif
for (size_t j = 0; j < CodeBlocks->size(); ++j) {
const FEXCore::Frontend::Decoder::DecodedBlocks& Block = CodeBlocks->at(j);
#ifdef ZYDIS_DISASSEMBLER
if (FEXCore::Config::Get_X86DISASSEMBLE() && CodeBlocks->size() > 1) {
LogMan::Msg::IFmt(" Block {} Entry={:#x} NumInsts={}", j, Block.Entry, Block.NumInstructions);
}
#endif
bool BlockInForceTSOValidRange = false;
auto InstForceTSOIt = ForceTSOInstructions.end();
if (ForceTSOValidRanges.Contains({Block.Entry, Block.Entry + Block.Size})) {
@@ -561,6 +594,19 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
TableInfo = Block.DecodedInstructions[i].TableInfo;
DecodedInfo = &Block.DecodedInstructions[i];
#ifdef ZYDIS_DISASSEMBLER
if (FEXCore::Config::Get_X86DISASSEMBLE()) {
const uint8_t* InstBytes = reinterpret_cast<const uint8_t*>(InstAddress);
ZydisDisassembledInstruction ZydisInst;
if (ZYAN_SUCCESS(ZydisDisassembleIntel(ZydisMachineMode, InstAddress, InstBytes, DecodedInfo->InstSize, &ZydisInst))) {
LogMan::Msg::IFmt(" {:#x}: {}", InstAddress, ZydisInst.text);
} else {
LogMan::Msg::IFmt(" {:#x}: (decode failed, {} bytes)", InstAddress, DecodedInfo->InstSize);
}
}
#endif
bool IsLocked = DecodedInfo->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_LOCK;
// Do a partial register cache flush before every instruction. This
@@ -646,7 +692,8 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
LogMan::Msg::EFmt("Invalid or Unknown instruction: {} 0x{:x}", TableInfo->Name ?: "UND", Block.Entry - GuestRIP);
}
if (Block.BlockStatus == Frontend::Decoder::DecodedBlockStatus::INVALID_INST) {
if (Block.BlockStatus == Frontend::Decoder::DecodedBlockStatus::INVALID_INST ||
Block.BlockStatus == Frontend::Decoder::DecodedBlockStatus::BAD_RELOCATION) {
Thread->OpDispatcher->InvalidOp(DecodedInfo);
} else {
Thread->OpDispatcher->NoExecOp(DecodedInfo);
@@ -662,8 +709,8 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
// If we had a dispatch error then leave early
if (HadDispatchError && TotalInstructions == 0) {
// Couldn't handle any instruction in op dispatcher
Thread->OpDispatcher->ResetWorkingList();
return {{}, 0, 0, 0, 0};
Thread->OpDispatcher->DelayedDisownBuffer();
return {std::nullopt, 0, 0, 0, 0};
}
if (NeedsBlockEnd) {
@@ -680,6 +727,12 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
}
}
#ifdef ZYDIS_DISASSEMBLER
if (FEXCore::Config::Get_X86DISASSEMBLE()) {
LogMan::Msg::IFmt("Guest x86 End");
}
#endif
Thread->OpDispatcher->Finalize();
Thread->FrontendDecoder->DelayedDisownBuffer();
@@ -713,9 +766,10 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
ContextImpl::CompileCodeResult ContextImpl::CompileCode(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP, uint64_t MaxInst) {
if (SourcecodeResolver && Config.GDBSymbols()) {
auto MappedSection = SyscallHandler->LookupExecutableFileSection(*Thread, GuestRIP);
auto MappedSection = SyscallHandler->LookupExecutableFileSection(Thread, GuestRIP);
if (MappedSection) {
MappedSection->FileInfo.SourcecodeMap = SourcecodeResolver->GenerateMap(MappedSection->FileInfo.Filename, MappedSection->FileInfo.FileId);
MappedSection->FileInfo.SourcecodeMap =
SourcecodeResolver->GenerateMap(MappedSection->FileInfo.Filename, CodeMap::GetBaseFilename(MappedSection->FileInfo, false));
}
}
@@ -723,6 +777,7 @@ ContextImpl::CompileCodeResult ContextImpl::CompileCode(FEXCore::Core::InternalT
auto [IRView, TotalInstructions, TotalInstructionsLength, StartAddr, Length, NeedsAddGuestCodeRanges] =
GenerateIR(Thread, GuestRIP, Config.GDBSymbols(), MaxInst);
if (!IRView) {
// OpDispatcher IR already released in this case.
return {{}, nullptr, 0, 0, false};
}
@@ -733,6 +788,7 @@ ContextImpl::CompileCodeResult ContextImpl::CompileCode(FEXCore::Core::InternalT
// as expensive and are easily reverted.
if (MaxInst != 1) {
if (auto Block = Thread->LookupCache->FindBlock(Thread, GuestRIP)) {
// Raced to compile, release the OpDispatcher IR.
Thread->OpDispatcher->DelayedDisownBuffer();
return {.CompiledCode = {.BlockBegin = reinterpret_cast<uint8_t*>(Block), .EntryPoints = {{GuestRIP, reinterpret_cast<uint8_t*>(Block)}}},
.DebugData = nullptr,
@@ -793,7 +849,7 @@ uintptr_t ContextImpl::CompileBlock(FEXCore::Core::CpuStateFrame* Frame, uint64_
if (Config.BlockJITNaming()) {
auto FragmentBasePtr = CompiledCode.BlockBegin;
auto GuestRIPLookup = SyscallHandler->LookupExecutableFileSection(*Thread, GuestRIP);
auto GuestRIPLookup = SyscallHandler->LookupExecutableFileSection(Thread, GuestRIP);
if (DebugData->Subblocks.size()) {
for (auto& Subblock : DebugData->Subblocks) {
@@ -816,7 +872,7 @@ uintptr_t ContextImpl::CompileBlock(FEXCore::Core::CpuStateFrame* Frame, uint64_
}
if (Config.LibraryJITNaming() || Config.GDBSymbols()) {
auto MappedSection = SyscallHandler->LookupExecutableFileSection(*Thread, GuestRIP);
auto MappedSection = SyscallHandler->LookupExecutableFileSection(Thread, GuestRIP);
if (MappedSection) {
if (Config.LibraryJITNaming()) {
Symbols.RegisterNamedRegion(Thread->SymbolBuffer.get(), CodePtr, DebugData->HostCodeSize, MappedSection->FileInfo.Filename);
@@ -833,10 +889,14 @@ uintptr_t ContextImpl::CompileBlock(FEXCore::Core::CpuStateFrame* Frame, uint64_
Thread->CPUBackend->ClearRelocations();
}
fextl::vector<uint64_t> CodePages;
if (NeedsAddGuestCodeRanges) {
// Track in the guest to host map all entrypoints for all pages the compiled block touches, if any page didn't previously
// contain code, inform the frontend so it can setup SMC detection.
auto BlockInfo = Thread->FrontendDecoder->GetDecodedBlockInfo();
CodePages.reserve(BlockInfo->CodePages.size());
CodePages.insert(CodePages.end(), BlockInfo->CodePages.begin(), BlockInfo->CodePages.end());
for (auto CodePage : BlockInfo->CodePages) {
if (Thread->LookupCache->AddBlockExecutableRange(Thread, BlockInfo->EntryPoints, CodePage, FEXCore::Utils::FEX_PAGE_SIZE)) {
SyscallHandler->MarkGuestExecutableRange(Thread, CodePage, FEXCore::Utils::FEX_PAGE_SIZE);
@@ -845,8 +905,16 @@ uintptr_t ContextImpl::CompileBlock(FEXCore::Core::CpuStateFrame* Frame, uint64_
}
// Insert to lookup cache
for (auto [GuestAddr, HostAddr] : CompiledCode.EntryPoints) {
Thread->LookupCache->AddBlockMapping(Thread, GuestAddr, HostAddr);
Thread->LookupCache->AddBlockMapping(Thread, GuestAddr, CodePages, HostAddr);
}
if (CodeMapWriter) {
auto Region = SyscallHandler->LookupExecutableFileSection(Thread, GuestRIP);
if (Region && Region->FileStartVA != 0) {
CodeMapWriter->AppendBlock(*Region, GuestRIP);
}
}
return (uintptr_t)CodePtr;
@@ -873,50 +941,37 @@ uintptr_t ContextImpl::CompileSingleStep(FEXCore::Core::CpuStateFrame* Frame, ui
return (uintptr_t)CodePtr;
}
static void InvalidateGuestThreadCodeRange(FEXCore::Core::InternalThreadState* Thread, InvalidatedEntryAccumulator& Accumulator,
uint64_t Start, uint64_t Length) {
void ContextImpl::InvalidateCodeBuffersCodeRange(uint64_t Start, uint64_t Length) {
FEXCORE_PROFILE_SCOPED("InvalidateCodeBuffersCodeRange");
LOGMAN_THROW_A_FMT(CodeInvalidationMutex.try_lock() == false, "CodeInvalidationMutex needs to be unique_locked here");
std::scoped_lock lk {CodeBufferListLock};
auto it = CodeBufferList.begin();
while (it != CodeBufferList.end()) {
if (auto Strong = it->lock()) {
Strong->LookupCache->InvalidateRange(Start, Length);
it++;
} else {
it = CodeBufferList.erase(it);
}
}
}
void ContextImpl::InvalidateThreadCachedCodeRange(FEXCore::Core::InternalThreadState* Thread, uint64_t Start, uint64_t Length) {
LOGMAN_THROW_A_FMT(CodeInvalidationMutex.try_lock() == false, "CodeInvalidationMutex needs to be unique_locked here");
// Ensures now-modified mappings aren't cached as being in their previous non-executable state.
// Accessing FrontendDecoder is safe as the thread's code invalidation mutex must be locked here.
Thread->FrontendDecoder->ResetExecutableRangeCache();
auto lk = Thread->LookupCache->AcquireWriteLock();
auto& CodePages = Thread->LookupCache->Shared->CodePages;
if (Thread->LookupCache->InvalidateCacheRange(Start, Length)) {
FEXCORE_PROFILE_SCOPED("InvalidateCallRet");
auto lower = CodePages.lower_bound(Start >> 12);
auto upper = CodePages.upper_bound((Start + Length - 1) >> 12);
for (auto it = lower; it != upper; it++) {
Accumulator.emplace_back(std::move(it->second));
}
bool InvalidatedAnyEntries = false;
for (const auto& PageEntries : Accumulator) {
for (const auto& Entry : PageEntries) {
if (ContextImpl::ThreadRemoveCodeEntry(Thread, Entry, lk)) {
InvalidatedAnyEntries = true;
}
}
}
if (InvalidatedAnyEntries) {
// This may cause access violations in the thread on Windows as zeroing is not atomic, this is handled by the frontend
Allocator::VirtualDontNeed(Thread->CallRetStackBase, FEXCore::Core::InternalThreadState::CALLRET_STACK_SIZE);
}
}
void ContextImpl::InvalidateGuestCodeRange(FEXCore::Core::InternalThreadState* Thread, InvalidatedEntryAccumulator& Accumulator,
uint64_t Start, uint64_t Length) {
InvalidateGuestThreadCodeRange(Thread, Accumulator, Start, Length);
}
bool ContextImpl::ThreadRemoveCodeEntry(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP,
const FEXCore::LookupCacheWriteLockToken& lk) {
LogMan::Throw::AFmt(static_cast<ContextImpl*>(Thread->CTX)->CodeInvalidationMutex.try_lock() == false, "CodeInvalidationMutex needs to "
"be unique_locked here");
return Thread->LookupCache->Erase(Thread->CurrentFrame, GuestRIP, lk);
}
void ContextImpl::ThreadRemoveCodeEntryFromJit(FEXCore::Core::CpuStateFrame* Frame, uint64_t GuestRIP) {
static_cast<ContextImpl*>(Frame->Thread->CTX)->SyscallHandler->InvalidateGuestCodeRange(Frame->Thread, GuestRIP, 1);
}
@@ -958,6 +1013,7 @@ void ContextImpl::AddThunkTrampolineIRHandler(uintptr_t Entrypoint, uintptr_t Gu
const auto GPRSize = this->Config.Is64BitMode ? IR::OpSize::i64Bit : IR::OpSize::i32Bit;
// Thunk entry-points don't get cached, don't need to be padded.
if (GPRSize == IR::OpSize::i64Bit) {
IR::Ref R = emit->_StoreRegister(emit->Constant(Entrypoint), GPRSize);
R->Reg = IR::PhysicalRegister(IR::RegClass::GPRFixed, X86State::REG_R11).Raw;
@@ -1028,6 +1084,5 @@ void ContextImpl::MonoBackpatcherWrite(FEXCore::Core::CpuStateFrame* Frame, uint
void ContextImpl::ConfigureAOTGen(FEXCore::Core::InternalThreadState* Thread, fextl::set<uint64_t>* ExternalBranches, uint64_t SectionMaxAddress) {
Thread->FrontendDecoder->SetExternalBranches(ExternalBranches);
Thread->FrontendDecoder->SetSectionMaxAddress(SectionMaxAddress);
}
} // namespace FEXCore::Context
File diff suppressed because it is too large. Load diff
@@ -28,6 +28,10 @@ class ContextImpl;
namespace FEXCore::CPU {
#define STATE_PTR(STATE_TYPE, FIELD) STATE.R(), offsetof(FEXCore::Core::STATE_TYPE, FIELD)
#define STATE_PTR_IDX(STATE_TYPE, FIELD, INDEX) STATE.R(), ARRAY_OFFSETOF(FEXCore::Core::STATE_TYPE, FIELD, INDEX)
#define FALLBACK_HANDLER_OFFSET(INDEX, FIELD) \
STATE.R(), \
(ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, Pointers.FallbackHandlerPointers, INDEX) + offsetof(FEXCore::Core::FallbackABIInfo, FIELD))
class Dispatcher final : public Arm64Emitter {
public:
@@ -51,6 +55,10 @@ public:
}
#endif
uint64_t GetExitFunctionLinkerAddress() const {
return ExitFunctionLinkerAddress;
}
SignalDelegatorConfig MakeSignalDelegatorConfig() const;
protected:
@@ -91,9 +99,33 @@ private:
uint64_t LUDIVHandlerAddress {};
uint64_t LDIVHandlerAddress {};
// F64 reduced-precision shared handlers
uint64_t F64SinHandlerAddress {};
uint64_t F64CosHandlerAddress {};
uint64_t F64TanHandlerAddress {};
uint64_t F64F2XM1HandlerAddress {};
uint64_t F64ScaleHandlerAddress {};
uint64_t F64AtanHandlerAddress {};
uint64_t F64FYL2XHandlerAddress {};
void EmitDispatcher();
uint64_t GenerateABICall(FallbackABI ABI);
// Inline softfloat conversion emitters - avoid FPCR save/restore overhead
// These emit ARM64 code that performs the conversion using only integer ops
void EmitI16ToExtF80();
void EmitI32ToExtF80();
void EmitF32ToExtF80();
void EmitF64ToExtF80();
void EmitF64Sin();
void EmitF64Cos();
void EmitF64Tan();
void EmitF64F2XM1();
void EmitF64Scale();
void EmitF64Atan();
void EmitF64FYL2X();
FEX_CONFIG_OPT(DisableL2Cache, DISABLEL2CACHE);
};
+133 -78
View File
@@ -9,7 +9,6 @@ $end_info$
#include "Interface/Context/Context.h"
#include "Interface/Core/Frontend.h"
#include "Interface/Core/X86Tables/X86Tables.h"
#include "Interface/Core/X86HelperGen.h"
#include "Interface/Core/LookupCache.h"
#include <array>
@@ -90,11 +89,6 @@ Decoder::Decoder(FEXCore::Core::InternalThreadState* Thread)
}
bool Decoder::CheckRangeExecutable(uint64_t Address, uint64_t Size) {
// Treat FEX-internal X86 callbacks as always executable
if (EntryPoint == CTX->X86CodeGen.CallbackReturn) {
return true;
}
while (Address < ExecutableRangeBase || Address + Size > ExecutableRangeEnd) {
auto RangeInfo = CTX->SyscallHandler->QueryGuestExecutableRange(Thread, Address);
ExecutableRangeBase = RangeInfo.Base;
@@ -138,7 +132,7 @@ std::optional<uint8_t> Decoder::PeekByte(uint8_t Offset) {
}
}
uint64_t Decoder::ReadData(uint8_t Size) {
std::pair<uint64_t, bool> Decoder::ReadData(uint8_t Size) {
LOGMAN_THROW_A_FMT(Size != 0 && Size <= sizeof(uint64_t), "Unknown data size to read");
uint64_t Res = 0;
@@ -160,7 +154,21 @@ uint64_t Decoder::ReadData(uint8_t Size) {
SkipBytes(Size);
#endif
return Res;
if (Relocations) {
uint32_t SectionOffset = static_cast<uint32_t>(Address - SectionMinAddress);
if (auto It = Relocations->find(SectionOffset); It != Relocations->end()) {
if (It->second == GuestRelocationType::Rel32 && Size == 4) {
return {static_cast<int64_t>(static_cast<int32_t>(Res) - static_cast<int32_t>(EntryPoint)), true};
} else if (It->second == GuestRelocationType::Rel64 && Size == 8) {
return {static_cast<int64_t>(Res) - static_cast<int64_t>(EntryPoint), true};
} else {
HitBadRelocation = true;
Res = 0;
}
}
}
return {Res, false};
}
void Decoder::DecodeModRM_16(X86Tables::DecodedOperand* Operand, X86Tables::ModRMDecoded ModRM) {
@@ -192,7 +200,9 @@ void Decoder::DecodeModRM_16(X86Tables::DecodedOperand* Operand, X86Tables::ModR
DisplacementSize = 1;
}
if (DisplacementSize) {
Literal = ReadData(DisplacementSize);
bool IsRelocation = false;
std::tie(Literal, IsRelocation) = ReadData(DisplacementSize);
LOGMAN_THROW_A_FMT(!IsRelocation, "1/2 byte relocations unsupported");
if (DisplacementSize == 1) {
Literal = static_cast<int8_t>(Literal);
}
@@ -298,7 +308,10 @@ void Decoder::DecodeModRM_64(X86Tables::DecodedOperand* Operand, X86Tables::ModR
LOGMAN_THROW_A_FMT(Displacement <= 4, "Number of bytes should be <= 4 for literal src");
if (Displacement) {
uint64_t Literal = ReadData(Displacement);
auto [Literal, IsRelocation] = ReadData(Displacement);
if (IsRelocation) {
Operand->Type = DecodedOperand::OpType::SIBRelocation;
}
if (Displacement == 1) {
Literal = static_cast<int8_t>(Literal);
}
@@ -308,10 +321,9 @@ void Decoder::DecodeModRM_64(X86Tables::DecodedOperand* Operand, X86Tables::ModR
// Explained in Table 1-14. "Operand Addressing Using ModRM and SIB Bytes"
if (ModRM.rm == 0b101) {
// 32bit Displacement
const uint32_t Literal = ReadData(4);
Operand->Type = DecodedOperand::OpType::RIPRelative;
Operand->Data.RIPLiteral.Value.u = Literal;
auto [Literal, IsRelocation] = ReadData(4);
Operand->Type = IsRelocation ? DecodedOperand::OpType::RIPRelativeRelocation : DecodedOperand::OpType::RIPRelative;
Operand->Data.RIPLiteral.Value = Literal;
} else {
// Register-direct addressing
Operand->Type = DecodedOperand::OpType::GPRDirect;
@@ -319,12 +331,12 @@ void Decoder::DecodeModRM_64(X86Tables::DecodedOperand* Operand, X86Tables::ModR
}
} else {
uint8_t DisplacementSize = ModRM.mod == 1 ? 1 : 4;
uint32_t Literal = ReadData(DisplacementSize);
auto [Literal, IsRelocation] = ReadData(DisplacementSize);
if (DisplacementSize == 1) {
Literal = static_cast<int8_t>(Literal);
}
Operand->Type = DecodedOperand::OpType::GPRIndirect;
Operand->Type = IsRelocation ? DecodedOperand::OpType::GPRIndirectRelocation : DecodedOperand::OpType::GPRIndirect;
Operand->Data.GPRIndirect.GPR = MapModRMToReg(DecodeInst->Flags & DecodeFlags::FLAG_REX_XGPR_B ? 1 : 0, ModRM.rm, false, false, false, false);
Operand->Data.GPRIndirect.Displacement = Literal;
}
@@ -620,31 +632,29 @@ bool Decoder::NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op,
if (Bytes != 0) {
LOGMAN_THROW_A_FMT(Bytes <= 8, "Number of bytes should be <= 8 for literal src");
DecodeInst->Src[CurrentSrc].Data.Literal.Size = Bytes;
uint64_t Literal = ReadData(Bytes);
auto [Literal, IsRelocation] = ReadData(Bytes);
if (IsRelocation) {
DecodeInst->Src[CurrentSrc].Type = DecodedOperand::OpType::LiteralRelocation;
DecodeInst->Src[CurrentSrc].Data.LiteralRelocation.EntrypointOffset = Literal;
} else {
DecodeInst->Src[CurrentSrc].Data.Literal.Size = Bytes;
if ((Info->Flags & FEXCore::X86Tables::InstFlags::FLAGS_SRC_SEXT) || (DecodeFlags::GetSizeDstFlags(DecodeInst->Flags) == DecodeFlags::SIZE_64BIT &&
Info->Flags & FEXCore::X86Tables::InstFlags::FLAGS_SRC_SEXT64BIT)) {
if (Bytes == 1) {
Literal = static_cast<int8_t>(Literal);
} else if (Bytes == 2) {
Literal = static_cast<int16_t>(Literal);
} else {
Literal = static_cast<int32_t>(Literal);
if ((Info->Flags & FEXCore::X86Tables::InstFlags::FLAGS_SRC_SEXT) ||
(DecodeFlags::GetSizeDstFlags(DecodeInst->Flags) == DecodeFlags::SIZE_64BIT &&
Info->Flags & FEXCore::X86Tables::InstFlags::FLAGS_SRC_SEXT64BIT)) {
if (Bytes == 1) {
Literal = static_cast<int8_t>(Literal);
} else if (Bytes == 2) {
Literal = static_cast<int16_t>(Literal);
} else {
Literal = static_cast<int32_t>(Literal);
}
DecodeInst->Src[CurrentSrc].Data.Literal.Size = DestSize;
}
DecodeInst->Src[CurrentSrc].Data.Literal.Size = DestSize;
DecodeInst->Src[CurrentSrc].Data.Literal.SignExtend = true;
}
DecodeInst->Src[CurrentSrc].Type = DecodedOperand::OpType::Literal;
DecodeInst->Src[CurrentSrc].Data.Literal.Value = Literal;
++CurrentSrc;
if (Bytes == 8) [[unlikely]] {
DecodeInst->Src[CurrentSrc].Data.Literal.Size = 4;
DecodeInst->Src[CurrentSrc].Type = DecodedOperand::OpType::Literal;
DecodeInst->Src[CurrentSrc].Data.Literal.Value = Literal >> 32;
DecodeInst->Src[CurrentSrc].Data.Literal.Value = Literal;
}
Bytes = 0;
@@ -759,13 +769,23 @@ bool Decoder::NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16
if (Op == 0xC5) { // Two byte VEX
pp = Byte1 & 0b11;
options.vvvv = 15 - ((Byte1 & 0b01111000) >> 3);
const uint8_t vvvv = ((Byte1 & 0b01111000) >> 3);
if (!BlockInfo.Is64BitMode && vvvv <= 0b0111) {
// Invalid on 32-bit, can't use the high registers.
return false;
}
options.vvvv = 15 - vvvv;
options.L = (Byte1 & 0b100) != 0;
} else { // 0xC4 = Three byte VEX
const uint8_t Byte2 = ReadByte();
pp = Byte2 & 0b11;
map_select = Byte1 & 0b11111;
options.vvvv = 15 - ((Byte2 & 0b01111000) >> 3);
const uint8_t vvvv = ((Byte2 & 0b01111000) >> 3);
if (!BlockInfo.Is64BitMode && vvvv <= 0b0111) {
// Invalid on 32-bit, can't use the high registers.
return false;
}
options.vvvv = 15 - vvvv;
options.w = (Byte2 & 0b10000000) != 0;
options.L = (Byte2 & 0b100) != 0;
if ((Byte1 & 0b01000000) == 0) {
@@ -838,6 +858,7 @@ bool Decoder::DecodeInstructionImpl(uint64_t PC) {
switch (EscapeOp) {
case 0x0F:
[[unlikely]] { // 3DNow!
DecodeREXIfValid(-2);
// 3DNow! Instruction Encoding: 0F 0F [ModRM] [SIB] [Displacement] [Opcode]
// Decode ModRM
uint8_t ModRMByte = ReadByte();
@@ -862,6 +883,7 @@ bool Decoder::DecodeInstructionImpl(uint64_t PC) {
break;
}
case 0x38: { // F38 Table!
DecodeREXIfValid(-2);
constexpr uint16_t PF_38_NONE = 0;
constexpr uint16_t PF_38_66 = (1U << 0);
constexpr uint16_t PF_38_F2 = (1U << 1);
@@ -887,11 +909,11 @@ bool Decoder::DecodeInstructionImpl(uint64_t PC) {
DecodeInst->Flags &= ~DecodeFlags::FLAG_OPERAND_SIZE;
DecodeFlags::PopOpAddrIf(&DecodeInst->Flags, DecodeFlags::FLAG_OPERAND_SIZE_LAST);
}
return NormalOpHeader(&FEXCore::X86Tables::H0F38TableOps[LocalOp], LocalOp);
break;
}
case 0x3A: { // F3A Table!
DecodeREXIfValid(-2);
constexpr uint16_t PF_3A_NONE = 0;
constexpr uint16_t PF_3A_66 = (1 << 0);
constexpr uint16_t PF_3A_REX = (1 << 1);
@@ -921,6 +943,7 @@ bool Decoder::DecodeInstructionImpl(uint64_t PC) {
bool NoOverlay = (FEXCore::X86Tables::SecondBaseOps[EscapeOp].Flags & InstFlags::FLAGS_NO_OVERLAY) != 0;
bool NoOverlay66 = (FEXCore::X86Tables::SecondBaseOps[EscapeOp].Flags & InstFlags::FLAGS_NO_OVERLAY66) != 0;
DecodeREXIfValid(-2);
if (NoOverlay) { // This section of the table ignores prefix extention
return NormalOpHeader(&FEXCore::X86Tables::SecondBaseOps[EscapeOp], EscapeOp);
} else if (LastEscapePrefix == 0xF3) { // REP
@@ -1000,29 +1023,9 @@ bool Decoder::DecodeInstructionImpl(uint64_t PC) {
}
if (Info->Type == FEXCore::X86Tables::TYPE_REX_PREFIX) {
DecodeInst->Flags |= DecodeFlags::FLAG_REX_PREFIX;
// Widening displacement
if (Op & 0b1000) {
DecodeInst->Flags |= DecodeFlags::FLAG_REX_WIDENING;
DecodeFlags::PushOpAddr(&DecodeInst->Flags, DecodeFlags::FLAG_WIDENING_SIZE_LAST);
}
// XGPR_B bit set
if (Op & 0b0001) {
DecodeInst->Flags |= DecodeFlags::FLAG_REX_XGPR_B;
}
// XGPR_X bit set
if (Op & 0b0010) {
DecodeInst->Flags |= DecodeFlags::FLAG_REX_XGPR_X;
}
// XGPR_R bit set
if (Op & 0b0100) {
DecodeInst->Flags |= DecodeFlags::FLAG_REX_XGPR_R;
}
DecodeInst->REXIndex = InstructionSize;
} else {
DecodeREXIfValid();
return NormalOpHeader(Info, Op);
}
@@ -1038,18 +1041,51 @@ bool Decoder::DecodeInstructionImpl(uint64_t PC) {
return true;
}
void Decoder::DecodeREXIfValid(int8_t ExpectedOffset) {
LOGMAN_THROW_A_FMT(ExpectedOffset < 0, "Expecting an negative offset for the REX offset!");
const int8_t REXIndex = InstructionSize + ExpectedOffset;
if (DecodeInst->REXIndex != 0 && DecodeInst->REXIndex == REXIndex) {
const uint8_t Op = Instruction[REXIndex - 1];
DecodeInst->Flags |= DecodeFlags::FLAG_REX_PREFIX;
// Widening displacement
if (Op & 0b1000) {
DecodeInst->Flags |= DecodeFlags::FLAG_REX_WIDENING;
DecodeFlags::PushOpAddr(&DecodeInst->Flags, DecodeFlags::FLAG_WIDENING_SIZE_LAST);
}
// XGPR_B bit set
if (Op & 0b0001) {
DecodeInst->Flags |= DecodeFlags::FLAG_REX_XGPR_B;
}
// XGPR_X bit set
if (Op & 0b0010) {
DecodeInst->Flags |= DecodeFlags::FLAG_REX_XGPR_X;
}
// XGPR_R bit set
if (Op & 0b0100) {
DecodeInst->Flags |= DecodeFlags::FLAG_REX_XGPR_R;
}
}
}
Decoder::DecodedBlockStatus Decoder::DecodeInstruction(uint64_t PC) {
// Will be set if DecodeInstructionImpl tries to read non-executable memory
HitNonExecutableRange = false;
HitBadRelocation = false;
bool ErrorDuringDecoding = !DecodeInstructionImpl(PC);
if (ErrorDuringDecoding || HitNonExecutableRange) [[unlikely]] {
if (ErrorDuringDecoding || HitNonExecutableRange || HitBadRelocation) [[unlikely]] {
// Put an invalid instruction in the stream so the core can raise SIGILL if hit
// Error while decoding instruction. We don't know the table or instruction size
DecodeInst->TableInfo = nullptr;
auto Result = ErrorDuringDecoding ? DecodedBlockStatus::INVALID_INST :
DecodeInst->InstSize ? DecodedBlockStatus::PARTIAL_DECODE_INST :
DecodedBlockStatus::NOEXEC_INST;
auto Result = ErrorDuringDecoding ? DecodedBlockStatus::INVALID_INST :
DecodeInst->InstSize ? DecodedBlockStatus::PARTIAL_DECODE_INST :
HitNonExecutableRange ? DecodedBlockStatus::NOEXEC_INST :
DecodedBlockStatus::BAD_RELOCATION;
DecodeInst->InstSize = 0;
return Result;
} else if (!DecodeInst->TableInfo || (DecodeInst->TableInfo->Type == TYPE_INST && !DecodeInst->TableInfo->OpcodeDispatcher.OpDispatch)) {
@@ -1135,9 +1171,9 @@ void Decoder::BranchTargetInMultiblockRange() {
// Forbid distant branches to have the cost code better match the guest code layout, avoiding massive (range-wise) code
// blocks in highly fragmented guest code. Such branches are often not-taken branches to garbage in obfuscated code.
constexpr uint64_t MAX_FORWARD_BRANCH_DIST = FEXCore::Utils::FEX_PAGE_SIZE * 4;
bool ValidMultiblockMember = TargetRIP >= SymbolMinAddress && TargetRIP < std::min(InstEnd + MAX_FORWARD_BRANCH_DIST, SymbolMaxAddress);
bool ValidMultiblockMember = TargetRIP >= EntryPoint && TargetRIP < std::min(InstEnd + MAX_FORWARD_BRANCH_DIST, SectionMaxAddress);
#ifdef _M_ARM_64EC
#ifdef ARCHITECTURE_arm64ec
ValidMultiblockMember = ValidMultiblockMember && !RtlIsEcCode(TargetRIP);
#endif
@@ -1312,6 +1348,13 @@ const uint8_t* Decoder::AdjustAddrForSpecialRegion(const uint8_t* _InstStream, u
return _InstStream - EntryPoint + RIP;
}
bool Decoder::CheckIfCacheable(FEXCore::Core::InternalThreadState& Thread, const uint8_t* InstStream, uint64_t PC, uint64_t MaxInst) {
DecodeInstructionsAtEntry(&Thread, InstStream, PC, MaxInst);
bool Uncacheable = HitBadRelocation;
DelayedDisownBuffer();
return !Uncacheable;
}
void Decoder::DecodeInstructionsAtEntry(FEXCore::Core::InternalThreadState* Thread, const uint8_t* _InstStream, uint64_t PC, uint64_t MaxInst) {
FEXCORE_PROFILE_SCOPED("DecodeInstructions");
BlockInfo.TotalInstructionCount = 0;
@@ -1328,19 +1371,23 @@ void Decoder::DecodeInstructionsAtEntry(FEXCore::Core::InternalThreadState* Thre
BlockInfo.Is64BitMode = CSSegment->L == 1;
LOGMAN_THROW_A_FMT(BlockInfo.Is64BitMode == CTX->Config.Is64BitMode, "Expected operating mode to not change at runtime!");
// XXX: Load symbol data
SymbolAvailable = false;
EntryPoint = PC;
BlockInfo.EntryPoints = {PC};
InstStream = _InstStream;
uint64_t TotalInstructions {};
// If we don't have symbols available then we become a bit optimistic about multiblock ranges
if (!SymbolAvailable) {
// If we don't have a symbol available then assume all branches are valid for multiblock
SymbolMaxAddress = SectionMaxAddress;
SymbolMinAddress = EntryPoint;
SectionMinAddress = 0;
SectionMaxAddress = ~0ULL;
Relocations = nullptr;
if (CTX->GetCodeCache().IsGeneratingCache || EnableCodeCacheValidation) {
// If generating cache, attempt to load section bounds and relocations
if (auto SectionInfo = CTX->SyscallHandler->LookupExecutableFileSection(Thread, EntryPoint)) {
SectionMinAddress = SectionInfo->FileStartVA;
SectionMaxAddress = SectionInfo->EndVA;
Relocations = &SectionInfo->FileInfo.Relocations;
}
}
DecodedMinAddress = EntryPoint;
@@ -1425,6 +1472,13 @@ void Decoder::DecodeInstructionsAtEntry(FEXCore::Core::InternalThreadState* Thre
}
BlockIt->BlockStatus = DecodeInstruction(OpAddress);
if (HitBadRelocation) {
BlockInfo.TotalInstructionCount = 0;
BlockInfo.Blocks = {*BlockIt};
BlockInfo.EntryPoints.clear();
BlockInfo.CodePages.clear();
return;
}
uint64_t OpEndAddress = OpAddress + DecodeInst->InstSize;
DecodedMinAddress = std::min(DecodedMinAddress, OpAddress);
@@ -1443,7 +1497,7 @@ void Decoder::DecodeInstructionsAtEntry(FEXCore::Core::InternalThreadState* Thre
// Can not continue this block at all on invalid instruction
if (BlockIt->BlockStatus != DecodedBlockStatus::SUCCESS) [[unlikely]] {
if (!EntryBlock) {
if (!EntryBlock && BlockIt->BlockStatus != DecodedBlockStatus::BAD_RELOCATION) {
// In multiblock configurations, we can early terminate any non-entrypoint blocks with the expectation that this won't get hit.
// Improves compile-times.
// Just need to undo additions that this block decoding has caused.
@@ -1453,9 +1507,10 @@ void Decoder::DecodeInstructionsAtEntry(FEXCore::Core::InternalThreadState* Thre
EraseBlock = true;
} else {
LogMan::Msg::EFmt("{} instruction in entry block: {:X}",
BlockIt->BlockStatus == DecodedBlockStatus::INVALID_INST ? "Invalid" :
BlockIt->BlockStatus == DecodedBlockStatus::NOEXEC_INST ? "NoExec" :
"PartialDecode",
BlockIt->BlockStatus == DecodedBlockStatus::INVALID_INST ? "Invalid" :
BlockIt->BlockStatus == DecodedBlockStatus::NOEXEC_INST ? "NoExec" :
BlockIt->BlockStatus == DecodedBlockStatus::BAD_RELOCATION ? "BadRelocation" :
"PartialDecode",
OpAddress);
}
break;
+15 -7
View File
@@ -4,9 +4,12 @@
#include "Interface/Core/X86Tables/X86Tables.h"
#include "Interface/IR/IR.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/CodeCache.h>
#include <FEXCore/Utils/ThreadPoolAllocator.h>
#include <FEXCore/fextl/set.h>
#include <FEXCore/fextl/vector.h>
#include <FEXCore/fextl/robin_map.h>
#include <array>
#include <cstddef>
@@ -28,6 +31,7 @@ public:
INVALID_INST,
NOEXEC_INST,
PARTIAL_DECODE_INST,
BAD_RELOCATION,
};
// New Frontend decoding
@@ -50,6 +54,7 @@ public:
};
Decoder(FEXCore::Core::InternalThreadState* Thread);
bool CheckIfCacheable(FEXCore::Core::InternalThreadState&, const uint8_t* InstStream, uint64_t PC, uint64_t MaxInst);
void DecodeInstructionsAtEntry(FEXCore::Core::InternalThreadState* Thread, const uint8_t* InstStream, uint64_t PC, uint64_t MaxInst);
const DecodedBlockInformation* GetDecodedBlockInfo() const {
@@ -59,9 +64,6 @@ public:
uint64_t DecodedMinAddress {};
uint64_t DecodedMaxAddress {~0ULL};
void SetSectionMaxAddress(uint64_t v) {
SectionMaxAddress = v;
}
void SetExternalBranches(fextl::set<uint64_t>* v) {
ExternalBranches = v;
}
@@ -87,6 +89,8 @@ private:
FEXCore::Context::ContextImpl* CTX;
const FEXCore::HLE::SyscallOSABI OSABI {};
FEX_CONFIG_OPT(EnableCodeCacheValidation, ENABLECODECACHEVALIDATION);
bool DecodeInstructionImpl(uint64_t PC);
DecodedBlockStatus DecodeInstruction(uint64_t PC);
@@ -100,7 +104,8 @@ private:
uint8_t ReadByte();
std::optional<uint8_t> PeekByte(uint8_t Offset);
uint64_t ReadData(uint8_t Size);
std::pair<uint64_t, bool> ReadData(uint8_t Size);
void SkipBytes(uint8_t Size) {
InstructionSize += Size;
}
@@ -108,6 +113,8 @@ private:
bool NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op, DecodedHeader Options = {});
bool NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op);
void DecodeREXIfValid(int8_t ExpectedOffset = -1);
static constexpr size_t DefaultDecodedBufferSize = 0x10000;
FEXCore::X86Tables::DecodedInst* DecodedBuffer {};
Utils::PoolBufferWithTimedRetirement<FEXCore::X86Tables::DecodedInst*, 5000, 500> PoolObject;
@@ -117,6 +124,7 @@ private:
uint64_t ExecutableRangeEnd {};
bool ExecutableRangeWritable {};
bool HitNonExecutableRange {};
bool HitBadRelocation {};
const uint8_t* InstStream {};
IR::OpSize GetGPROpSize() const {
@@ -130,13 +138,11 @@ private:
FEXCore::X86Tables::DecodedInst* DecodeInst;
// This is for multiblock data tracking
bool SymbolAvailable {false};
uint64_t EntryPoint {};
uint64_t MaxCondBranchForward {};
uint64_t MaxCondBranchBackwards {~0ULL};
uint64_t SymbolMaxAddress {};
uint64_t SymbolMinAddress {~0ULL};
uint64_t SectionMaxAddress {~0ULL};
uint64_t SectionMinAddress {};
uint64_t NextBlockStartAddress {~0ULL};
DecodedBlockInformation BlockInfo;
@@ -145,6 +151,8 @@ private:
fextl::set<uint64_t> VisitedBlocks;
fextl::set<uint64_t>* ExternalBranches {nullptr};
const fextl::robin_map<uint32_t, GuestRelocationType>* Relocations {nullptr};
// ModRM rm decoding
using DecodeModRMPtr = void (FEXCore::Frontend::Decoder::*)(X86Tables::DecodedOperand* Operand, X86Tables::ModRMDecoded ModRM);
void DecodeModRM_16(X86Tables::DecodedOperand* Operand, X86Tables::ModRMDecoded ModRM);
@@ -435,12 +435,12 @@ struct OpHandlers<IR::OP_F80BCDSTORE> {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
X80SoftFloat Src1 = Src1q;
ScopedSoftFloatState State {FCW, Frame};
bool Negative = Src1.Sign;
bool Negative = Src1.Top.Sign;
Src1 = X80SoftFloat::FRNDINT(&State.State, Src1);
// Clear the Sign bit
Src1.Sign = 0;
Src1.Top.Sign = 0;
uint64_t Tmp = Src1.ToI64(&State.State);
X80SoftFloat Rv;
@@ -503,7 +503,7 @@ struct OpHandlers<IR::OP_F80BCDLOAD> {
X80SoftFloat Tmp;
Tmp = BCD;
Tmp.Sign = Negative;
Tmp.Top.Sign = Negative;
return Tmp;
}
};
@@ -2,14 +2,14 @@
#include "Interface/Core/Interpreter/Fallbacks/VectorFallbacks.h"
#include "Interface/IR/IR.h"
#ifdef _M_ARM_64
#ifdef ARCHITECTURE_arm64
#include <arm_neon.h>
#endif
#include <cstring>
namespace FEXCore::CPU {
#ifdef _M_ARM_64
#ifdef ARCHITECTURE_arm64
FEXCORE_PRESERVE_ALL_ATTR static int32_t GetImplicitLength(FEXCore::VectorRegType data, uint16_t control) {
const auto is_using_words = (control & 1) != 0;
+13 -6
View File
@@ -43,21 +43,28 @@ DEF_BINOP_WITH_CONSTANT(Ror, rorv, ror)
DEF_OP(Constant) {
auto Op = IROp->C<IR::IROp_Constant>();
auto Dst = GetReg(Node);
LoadConstant(ARMEmitter::Size::i64Bit, Dst, Op->Constant);
const auto PadType = [Pad = Op->Pad]() {
switch (Pad) {
case IR::ConstPad::NoPad: return CPU::Arm64Emitter::PadType::NOPAD;
case IR::ConstPad::DoPad: return CPU::Arm64Emitter::PadType::DOPAD;
default: return CPU::Arm64Emitter::PadType::AUTOPAD;
}
}();
LoadConstant(ARMEmitter::Size::i64Bit, Dst, Op->Constant, PadType, Op->MaxBytes);
}
DEF_OP(EntrypointOffset) {
auto Op = IROp->C<IR::IROp_EntrypointOffset>();
auto Constant = Entry + Op->Offset;
auto Dst = GetReg(Node);
uint64_t Mask = ~0ULL;
const auto OpSize = IROp->Size;
if (OpSize == IR::OpSize::i32Bit) {
Mask = 0xFFFF'FFFFULL;
}
LoadConstant(ARMEmitter::Size::i64Bit, Dst, Constant & Mask);
InsertGuestRIPMove(GetReg(Node), Constant & Mask);
}
DEF_OP(InlineConstant) {
@@ -267,7 +274,7 @@ DEF_OP(CmpPairZ) {
// Restore NzCV
if (CTX->HostFeatures.SupportsFlagM) {
rmif(TMP1, 0, 0xb /* NzCV */);
rmif(TMP1, 28, 0xb /* NzCV */);
} else {
cset(ARMEmitter::Size::i32Bit, TMP2, ARMEmitter::Condition::CC_EQ);
bfi(ARMEmitter::Size::i32Bit, TMP1, TMP2, 30 /* lsb: Z */, 1);
@@ -917,7 +924,7 @@ DEF_OP(Div) {
mov(EmitSize, TMP2, Lower);
mov(EmitSize, TMP3, Divisor);
ldr(TMP4, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.LDIVHandler));
ldr(TMP4, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.LDIVHandler));
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
blr(TMP4);
@@ -1000,7 +1007,7 @@ DEF_OP(UDiv) {
mov(EmitSize, TMP2, Lower);
mov(EmitSize, TMP3, Divisor);
ldr(TMP4, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.LUDIVHandler));
ldr(TMP4, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.LUDIVHandler));
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
blr(TMP4);
@@ -11,36 +11,33 @@ $end_info$
#include <FEXCore/Core/Thunks.h>
namespace FEXCore::CPU {
uint64_t Arm64JITCore::GetNamedSymbolLiteral(FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op) {
uint64_t GetNamedSymbolLiteral(FEXCore::Context::ContextImpl& CTX, FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op) {
switch (Op) {
case FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol::SYMBOL_LITERAL_EXITFUNCTION_LINKER:
return ThreadState->CurrentFrame->Pointers.Common.ExitFunctionLinker;
break;
default: ERROR_AND_DIE_FMT("Unknown named symbol literal: {}", static_cast<uint32_t>(Op)); break;
return CTX.Dispatcher->GetExitFunctionLinkerAddress();
default: ERROR_AND_DIE_FMT("Unknown named symbol literal: {}", static_cast<uint32_t>(Op));
}
return ~0ULL;
}
void Arm64JITCore::InsertNamedThunkRelocation(ARMEmitter::Register Reg, const IR::SHA256Sum& Sum) {
Relocation MoveABI {};
MoveABI.NamedThunkMove.Header.Type = FEXCore::CPU::RelocationTypes::RELOC_NAMED_THUNK_MOVE;
// Offset is the offset from the entrypoint of the block
auto CurrentCursor = GetCursorAddress<uint8_t*>();
MoveABI.NamedThunkMove.Offset = CurrentCursor - CodeData.BlockBegin;
MoveABI.NamedThunkMove.Header = {.Offset = GetCursorOffset(), .Type = FEXCore::CPU::RelocationTypes::RELOC_NAMED_THUNK_MOVE};
MoveABI.NamedThunkMove.Symbol = Sum;
MoveABI.NamedThunkMove.RegisterIndex = Reg.Idx();
uint64_t Pointer = reinterpret_cast<uint64_t>(EmitterCTX->ThunkHandler->LookupThunk(Sum));
LoadConstant(ARMEmitter::Size::i64Bit, Reg, Pointer, false);
// Pointers are required to fit within 48-bit VA space.
// TODO: Force 6-byte `MaxSize`, with zext extension to 64-bit. Current code not smart enough to handle negatives.
LoadConstant(ARMEmitter::Size::i64Bit, Reg, Pointer, FEXCore::CPU::Arm64Emitter::PadType::AUTOPAD);
Relocations.emplace_back(MoveABI);
}
Arm64JITCore::NamedSymbolLiteralPair Arm64JITCore::InsertNamedSymbolLiteral(FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op) {
uint64_t Pointer = GetNamedSymbolLiteral(Op);
uint64_t Pointer = GetNamedSymbolLiteral(*CTX, Op);
Arm64JITCore::NamedSymbolLiteralPair Lit {
NamedSymbolLiteralPair Lit {
.Lit = Pointer,
.MoveABI =
{
@@ -48,92 +45,76 @@ Arm64JITCore::NamedSymbolLiteralPair Arm64JITCore::InsertNamedSymbolLiteral(FEXC
{
.Header =
{
.Offset = 0, // Set by PlaceNamedSymbolLiteral
.Type = FEXCore::CPU::RelocationTypes::RELOC_NAMED_SYMBOL_LITERAL,
},
.Symbol = Op,
.Offset = 0,
},
},
};
return Lit;
}
void Arm64JITCore::PlaceNamedSymbolLiteral(NamedSymbolLiteralPair& Lit) {
// Offset is the offset from the entrypoint of the block
auto CurrentCursor = GetCursorAddress<uint8_t*>();
Lit.MoveABI.NamedSymbolLiteral.Offset = CurrentCursor - CodeData.BlockBegin;
void Arm64JITCore::PlaceNamedSymbolLiteral(NamedSymbolLiteralPair Lit) {
switch (Lit.MoveABI.Header.Type) {
case RelocationTypes::RELOC_NAMED_SYMBOL_LITERAL:
case RelocationTypes::RELOC_GUEST_RIP_LITERAL: {
Lit.MoveABI.Header.Offset = GetCursorOffset();
break;
}
default: ERROR_AND_DIE_FMT("Unknown relocation type for {}", __FUNCTION__);
}
BindOrRestart(&Lit.Loc);
dc64(Lit.Lit);
Relocations.emplace_back(Lit.MoveABI);
}
auto Arm64JITCore::InsertGuestRIPLiteral(uint64_t GuestRIP) -> NamedSymbolLiteralPair {
return {
.Lit = GuestRIP,
.MoveABI =
{
.GuestRIP = {.Header =
{
.Offset = 0, // Set by PlaceNamedSymbolLiteral
.Type = FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_LITERAL,
},
// NOTE: Cache serialization will subtract the guest binary base address later to produce consistency results
.GuestRIP = GuestRIP},
},
};
}
void Arm64JITCore::InsertGuestRIPMove(ARMEmitter::Register Reg, uint64_t Constant) {
Relocation MoveABI {};
MoveABI.GuestRIPMove.Header.Type = FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_MOVE;
// Offset is the offset from the entrypoint of the block
auto CurrentCursor = GetCursorAddress<uint8_t*>();
MoveABI.GuestRIPMove.Offset = CurrentCursor - CodeData.BlockBegin;
MoveABI.GuestRIPMove.GuestRIP = Constant;
MoveABI.GuestRIPMove.RegisterIndex = Reg.Idx();
MoveABI.GuestRIP.Header = {.Offset = GetCursorOffset(), .Type = FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_MOVE};
// NOTE: Cache serialization will subtract the guest binary base address later to produce consistency results
MoveABI.GuestRIP.GuestRIP = Constant;
MoveABI.GuestRIP.RegisterIndex = Reg.Idx();
LoadConstant(ARMEmitter::Size::i64Bit, Reg, Constant, false);
// Pointers are required to fit within 48-bit VA space.
// TODO: Force 6-byte `MaxSize`, with sign extension to 64-bit. Current code not smart enough to handle negatives.
// 48-bit sign extension works because x86-64 guests only receive 47-bit VA space, with 48-bit being reserved for kernel.
// Additional quirk, "canonical" 48-bit pointers on x86-64, sign extend the 48-bit as well (Which is why kernel pointers are negative).
LoadConstant(ARMEmitter::Size::i64Bit, Reg, Constant, FEXCore::CPU::Arm64Emitter::PadType::AUTOPAD);
Relocations.emplace_back(MoveABI);
}
bool Arm64JITCore::ApplyRelocations(uint64_t GuestEntry, std::span<std::byte> Code, std::span<const FEXCore::CPU::Relocation> Relocations) {
const auto OrigBase = GetBufferBase();
const auto OrigSize = GetBufferSize();
const auto OrigOffset = GetCursorOffset();
SetBuffer(reinterpret_cast<std::uint8_t*>(Code.data()), Code.size_bytes());
for (auto& Reloc : Relocations) {
switch (Reloc.Header.Type) {
case FEXCore::CPU::RelocationTypes::RELOC_NAMED_SYMBOL_LITERAL: {
uint64_t Pointer = GetNamedSymbolLiteral(Reloc.NamedSymbolLiteral.Symbol);
// Relocation occurs at the cursorEntry + offset relative to that cursor
SetCursorOffset(Reloc.NamedSymbolLiteral.Offset);
// Generate a literal so we can place it
dc64(Pointer);
break;
}
case FEXCore::CPU::RelocationTypes::RELOC_NAMED_THUNK_MOVE: {
uint64_t Pointer = reinterpret_cast<uint64_t>(EmitterCTX->ThunkHandler->LookupThunk(Reloc.NamedThunkMove.Symbol));
if (Pointer == ~0ULL) {
return false;
}
// Relocation occurs at the cursorEntry + offset relative to that cursor.
SetCursorOffset(Reloc.NamedThunkMove.Offset);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Register(Reloc.NamedThunkMove.RegisterIndex), Pointer, true);
break;
}
case FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_MOVE: {
// XXX: Reenable once the JIT Object Cache is upstream
// XXX: Should spin the relocation list, create a list of guest RIP moves, and ask for them all once, reduces lock contention.
uint64_t Pointer = ~0ULL; // EmitterCTX->JITObjectCache->FindRelocatedRIP(Reloc->GuestRIPMove.GuestRIP);
if (Pointer == ~0ULL) {
SetBuffer(OrigBase, OrigSize);
SetCursorOffset(OrigOffset);
return false;
}
// Relocation occurs at the cursorEntry + offset relative to that cursor.
SetCursorOffset(Reloc.GuestRIPMove.Offset);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Register(Reloc.GuestRIPMove.RegisterIndex), Pointer, true);
fextl::vector<FEXCore::CPU::Relocation> Arm64JITCore::TakeRelocations(uint64_t GuestBaseAddress) {
// Rebase relocations to library base address
for (auto& Relocation : Relocations) {
switch (Relocation.Header.Type) {
case FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_MOVE:
case FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_LITERAL: {
Relocation.GuestRIP.GuestRIP -= GuestBaseAddress;
break;
}
default:;
}
}
SetBuffer(OrigBase, OrigSize);
SetCursorOffset(OrigOffset);
return true;
}
fextl::vector<FEXCore::CPU::Relocation> Arm64JITCore::TakeRelocations() {
return std::move(Relocations);
}
@@ -138,26 +138,6 @@ DEF_OP(CAS) {
}
}
DEF_OP(AtomicXor) {
auto Op = IROp->C<IR::IROp_AtomicXor>();
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
auto MemSrc = GetReg(Op->Addr);
auto Src = GetReg(Op->Value);
if (CTX->HostFeatures.SupportsAtomics) {
steorl(SubEmitSize, Src, MemSrc);
} else {
ARMEmitter::BackwardLabel LoopTop;
(void)Bind(&LoopTop);
ldaxr(SubEmitSize, TMP2, MemSrc);
eor(EmitSize, TMP2, TMP2, Src);
stlxr(SubEmitSize, TMP2, TMP2, MemSrc);
(void)cbnz(EmitSize, TMP2, &LoopTop);
}
}
DEF_OP(AtomicSwap) {
auto Op = IROp->C<IR::IROp_AtomicSwap>();
const auto OpSize = IROp->Size;
@@ -349,7 +329,7 @@ DEF_OP(TelemetrySetValue) {
auto Op = IROp->C<IR::IROp_TelemetrySetValue>();
auto Src = GetReg(Op->Value);
ldr(TMP2, STATE_PTR(CpuStateFrame, Pointers.Common.TelemetryValueAddresses[Op->TelemetryValueIndex]));
ldr(TMP2, STATE_PTR_IDX(CpuStateFrame, Pointers.TelemetryValueAddresses, Op->TelemetryValueIndex));
// Cortex fuses cmp+cset.
cmp(ARMEmitter::Size::i32Bit, Src, 0);
+35 -21
View File
@@ -56,12 +56,12 @@ DEF_OP(ExitFunction) {
uint64_t NewRIP;
if (IsInlineConstant(Op->NewRIP, &NewRIP) || IsInlineEntrypointOffset(Op->NewRIP, &NewRIP)) {
#ifdef _M_ARM_64EC
#ifdef ARCHITECTURE_arm64ec
if (NewRIP < EC_CODE_BITMAP_MAX_ADDRESS && RtlIsEcCode(NewRIP)) {
str(REG_CALLRET_SP, STATE_PTR(CpuStateFrame, State.callret_sp));
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, StaticRegisters[X86State::REG_RSP], 0);
LoadConstant(ARMEmitter::Size::i64Bit, EC_CALL_CHECKER_PC_REG, NewRIP);
ldr(TMP2, STATE_PTR(CpuStateFrame, Pointers.Common.ExitFunctionEC));
InsertGuestRIPMove(EC_CALL_CHECKER_PC_REG, NewRIP);
ldr(TMP2, STATE_PTR(CpuStateFrame, Pointers.ExitFunctionEC));
br(TMP2);
} else {
#endif
@@ -150,16 +150,16 @@ DEF_OP(ExitFunction) {
ARMEmitter::ForwardLabel TFUnset;
ldrb(TMP1, STATE_PTR(CpuStateFrame, State.flags[X86State::RFLAG_TF_RAW_LOC]));
(void)cbz(ARMEmitter::Size::i32Bit, TMP1, &TFUnset);
LoadConstant(ARMEmitter::Size::i64Bit, TMP1, NewRIP);
InsertGuestRIPMove(TMP1, NewRIP);
str(TMP1, STATE, offsetof(FEXCore::Core::CpuStateFrame, State.rip));
ldr(TMP2, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.DispatcherLoopTop));
ldr(TMP2, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.DispatcherLoopTop));
blr(TMP2);
(void)Bind(&TFUnset);
}
EmitLinkedBranch(NewRIP, Op->Hint == IR::BranchHint::Call);
(void)Bind(&l_CallReturn);
#ifdef _M_ARM_64EC
#ifdef ARCHITECTURE_arm64ec
}
#endif
} else {
@@ -186,7 +186,7 @@ DEF_OP(ExitFunction) {
// Note: sub+cbnz used over cmp+br to preserve flags.
sub(TMP1, TMP1, RipReg.X());
(void)cbz(ARMEmitter::Size::i64Bit, TMP1, &SkipFullLookup);
ldr(TMP2, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.DispatcherLoopTop));
ldr(TMP2, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.DispatcherLoopTop));
str(RipReg.X(), STATE, offsetof(FEXCore::Core::CpuStateFrame, State.rip));
(void)Bind(&SkipFullLookup);
@@ -265,7 +265,10 @@ DEF_OP(Syscall) {
uint32_t GPRSpillMask = ~0U;
uint32_t FPRSpillMask = ~0U;
SpillStaticRegs(TMP1, true, GPRSpillMask, FPRSpillMask);
SpillStaticRegs(TMP1, {
.GPRSpillMask = GPRSpillMask,
.FPRSpillMask = FPRSpillMask,
});
// Now that we are spilled, store in the state that we are in a syscall
// Still without overwriting registers that matter
@@ -283,8 +286,8 @@ DEF_OP(Syscall) {
str(GetReg(Op->Header.Args[i]).X(), ARMEmitter::Reg::rsp, i * 8);
}
ldr(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.SyscallHandlerObj));
ldr(ARMEmitter::XReg::x3, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.SyscallHandlerFunc));
ldr(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.SyscallHandlerObj));
ldr(ARMEmitter::XReg::x3, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.SyscallHandlerFunc));
mov(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r1, STATE.R());
// SP supporting move
@@ -299,7 +302,12 @@ DEF_OP(Syscall) {
// Result is now in x0
// Fix the stack and any values that were stepped on
FillStaticRegs(true, GPRSpillMask, FPRSpillMask, ARMEmitter::Reg::r1, ARMEmitter::Reg::r2);
FillStaticRegs({
.OptionalReg = ARMEmitter::Reg::r1,
.OptionalReg2 = ARMEmitter::Reg::r2,
.GPRFillMask = GPRSpillMask,
.FPRFillMask = FPRSpillMask,
});
// Now the registers we've spilled are back in their original host registers
// We can safely claim we are no longer in a syscall
@@ -322,14 +330,16 @@ DEF_OP(Thunk) {
// X0: CTX
// X1: Args (from guest stack)
SpillStaticRegs(TMP1); // spill to ctx before ra64 spill
// spill to ctx before ra64 spill
SpillStaticRegs(TMP1, {
.NZCV = false,
});
PushDynamicRegs(TMP1);
mov(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, GetReg(Op->ArgPtr));
auto thunkFn = static_cast<Context::ContextImpl*>(ThreadState->CTX)->ThunkHandler->LookupThunk(Op->ThunkNameHash);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r2, (uintptr_t)thunkFn);
InsertNamedThunkRelocation(ARMEmitter::Reg::r2, Op->ThunkNameHash);
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<void, void*, void*>(ARMEmitter::Reg::r2);
} else {
@@ -338,7 +348,10 @@ DEF_OP(Thunk) {
PopDynamicRegs();
FillStaticRegs(); // load from ctx after ra64 refill
// load from ctx after ra64 refill
FillStaticRegs({
.NZCV = false,
});
}
DEF_OP(ValidateCode) {
@@ -398,9 +411,10 @@ DEF_OP(ThreadRemoveCodeEntry) {
// X1: RIP
mov(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, STATE.R());
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r1, Entry);
// TODO: Relocations don't seem to be wired up to this...?
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r1, Entry, CPU::Arm64Emitter::PadType::AUTOPAD);
ldr(ARMEmitter::XReg::x2, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.ThreadRemoveCodeEntryFromJIT));
ldr(ARMEmitter::XReg::x2, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.ThreadRemoveCodeEntryFromJIT));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<void, void*, void*>(ARMEmitter::Reg::r2);
} else {
@@ -425,8 +439,8 @@ DEF_OP(CPUID) {
// x0 = CPUID Handler
// x1 = CPUID Function
// x2 = CPUID Leaf
ldr(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.CPUIDObj));
ldr(ARMEmitter::XReg::x3, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.CPUIDFunction));
ldr(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.CPUIDObj));
ldr(ARMEmitter::XReg::x3, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.CPUIDFunction));
if (!TMP_ABIARGS) {
mov(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r1, TMP2);
@@ -466,8 +480,8 @@ DEF_OP(XGetBV) {
// x0 = CPUID Handler
// x1 = XCR Function
ldr(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.CPUIDObj));
ldr(ARMEmitter::XReg::x2, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.XCRFunction));
ldr(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.CPUIDObj));
ldr(ARMEmitter::XReg::x2, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.XCRFunction));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<uint64_t, void*, uint32_t>(ARMEmitter::Reg::r2);
} else {
+122 -103
View File
@@ -1,7 +1,7 @@
// SPDX-License-Identifier: MIT
/*
$info$
glossary: Splatter ~ a code generator backend that concaternates configurable macros instead of doing isel
glossary: Splatter ~ a code generator backend that concatenates configurable macros instead of doing isel
glossary: IR ~ Intermediate Representation, our high-level opcode representation, loosely modeling arm64
glossary: SSA ~ Single Static Assignment, a form of representing IR in memory
glossary: Basic Block ~ A block of instructions with no control flow, terminated by control flow
@@ -68,6 +68,10 @@ PrintValue(uint64_t Value) {
LogMan::Msg::DFmt("Value: 0x{:x}", Value);
}
static void PrintMsg(const char* Value) {
LogMan::Msg::DFmt("{}", Value);
}
static void PrintVectorValue(uint64_t Value, uint64_t ValueUpper) {
LogMan::Msg::DFmt("Value: 0x{:016x}'{:016x}", ValueUpper, Value);
}
@@ -133,8 +137,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
const auto Src1 = GetVReg(IROp->Args[0]);
fmov(VTMP1.S(), Src1.S());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -151,8 +155,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
const auto Src1 = GetVReg(IROp->Args[0]);
fmov(VTMP1.D(), Src1.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -176,8 +180,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
mov(ARMEmitter::Size::i32Bit, TMP2, Src1);
}
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -194,8 +198,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
const auto Src1 = GetVReg(IROp->Args[0]);
mov(VTMP1.Q(), Src1.Q());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -212,8 +216,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
const auto Src1 = GetVReg(IROp->Args[0]);
mov(VTMP1.Q(), Src1.Q());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -230,8 +234,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
const auto Src1 = GetVReg(IROp->Args[0]);
fmov(VTMP1.D(), Src1.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -254,8 +258,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
fmov(VTMP1.D(), Src1.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -276,8 +280,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
fmov(VTMP1.D(), Src1.D());
fmov(VTMP2.D(), Src2.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -294,8 +298,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
const auto Src1 = GetVReg(IROp->Args[0]);
mov(VTMP1.Q(), Src1.Q());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -312,8 +316,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
const auto Src1 = GetVReg(IROp->Args[0]);
mov(VTMP1.Q(), Src1.Q());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -330,8 +334,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
const auto Src1 = GetVReg(IROp->Args[0]);
mov(VTMP1.Q(), Src1.Q());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -351,8 +355,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
mov(VTMP1.Q(), Src1.Q());
mov(VTMP2.Q(), Src2.Q());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -369,8 +373,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
const auto Src1 = GetVReg(IROp->Args[0]);
mov(VTMP1.Q(), Src1.Q());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -394,8 +398,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
mov(VTMP1.Q(), Src1.Q());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -416,8 +420,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
mov(VTMP1.Q(), Src1.Q());
mov(VTMP2.Q(), Src2.Q());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -434,8 +438,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
// tmp2 (x1/x11): source 2
// tmp3 (x2/x12): source 3
const auto Op = IROp->C<IR::IROp_VPCMPESTRX>();
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP1, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
stp<ARMEmitter::IndexType::PRE>(TMP1, ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
@@ -476,8 +480,8 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
mov(VTMP2.Q(), Src2.Q());
movz(ARMEmitter::Size::i32Bit, TMP1, Control);
ldr(TMP2, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex].ABIHandler));
ldr(TMP4, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex].Func));
ldr(TMP2, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, ABIHandler));
ldr(TMP4, FALLBACK_HANDLER_OFFSET(Info.HandlerIndex, Func));
blr(TMP2);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
@@ -493,7 +497,7 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
}
}
static void DirectBlockDelinker(FEXCore::Core::CpuStateFrame* Frame, FEXCore::Context::ExitFunctionLinkData* Record, bool Call) {
static void DirectBlockDelinker(FEXCore::Context::ExitFunctionLinkData* Record, bool Call) {
uintptr_t JumpThunkStartAddress = reinterpret_cast<uintptr_t>(Record) - 0x10;
uintptr_t CallerAddress = JumpThunkStartAddress + Record->CallerOffset;
auto BranchOffset = JumpThunkStartAddress / 4 - CallerAddress / 4;
@@ -511,11 +515,12 @@ static void DirectBlockDelinker(FEXCore::Core::CpuStateFrame* Frame, FEXCore::Co
ARMEmitter::Emitter::ClearICache(reinterpret_cast<void*>(CallerAddress), 4);
}
static void IndirectBlockDelinker(FEXCore::Core::CpuStateFrame* Frame, FEXCore::Context::ExitFunctionLinkData* Record) {
static void IndirectBlockDelinker(FEXCore::Context::ExitFunctionLinkData* Record) {
uintptr_t JumpThunkStartAddress = reinterpret_cast<uintptr_t>(Record) - 0x10;
uint32_t BranchInst = 0;
ARMEmitter::Emitter BranchEmit(reinterpret_cast<uint8_t*>(&BranchInst), 4);
BranchEmit.b(0x8);
// Restore branch +2 instructions to jump to the linker block
BranchEmit.b(0x2);
std::atomic_ref<uint32_t>(*reinterpret_cast<uint32_t*>(JumpThunkStartAddress)).store(BranchInst, std::memory_order::relaxed);
ARMEmitter::Emitter::ClearICache(reinterpret_cast<void*>(JumpThunkStartAddress), 4);
@@ -532,7 +537,7 @@ uint64_t Arm64JITCore::ExitFunctionLink(FEXCore::Core::CpuStateFrame* Frame, FEX
if (TFSet) {
// If TF is set, the cache must be skipped as different code needs to be generated.
Frame->State.rip = GuestRip;
return Frame->Pointers.Common.DispatcherLoopTop;
return Frame->Pointers.DispatcherLoopTop;
} else {
{
// Guard the LookupCache lock with the code invalidation mutex, to avoid issues with forking
@@ -577,16 +582,11 @@ uint64_t Arm64JITCore::ExitFunctionLink(FEXCore::Core::CpuStateFrame* Frame, FEX
if (KnownCallMarkerInst == ExpectedKnownCallMarkerInst) {
BranchEmit.bl(BranchOffset);
Thread->LookupCache->AddBlockLink(
GuestRip, Record,
[](FEXCore::Core::CpuStateFrame* Frame, FEXCore::Context::ExitFunctionLinkData* Record) { DirectBlockDelinker(Frame, Record, true); }, lk);
GuestRip, Record, [](FEXCore::Context::ExitFunctionLinkData* Record) { DirectBlockDelinker(Record, true); }, lk);
} else {
BranchEmit.b(BranchOffset);
Thread->LookupCache->AddBlockLink(
GuestRip, Record,
[](FEXCore::Core::CpuStateFrame* Frame, FEXCore::Context::ExitFunctionLinkData* Record) {
DirectBlockDelinker(Frame, Record, false);
},
lk);
GuestRip, Record, [](FEXCore::Context::ExitFunctionLinkData* Record) { DirectBlockDelinker(Record, false); }, lk);
}
std::atomic_ref<uint32_t>(*reinterpret_cast<uint32_t*>(CallerAddress)).store(BranchInst, std::memory_order::relaxed);
@@ -594,7 +594,7 @@ uint64_t Arm64JITCore::ExitFunctionLink(FEXCore::Core::CpuStateFrame* Frame, FEX
} else {
// This case is common between calls and jumps as the thunk callsite can be left untouched.
std::atomic_ref<uint64_t>(Record->HostCode).store(HostCode, std::memory_order::seq_cst);
#ifdef _M_ARM_64
#ifdef ARCHITECTURE_arm64
// Make memory write visible to other threads reading the same location
asm volatile("dc cvau, %0; dsb ish" : : "r"(Record->HostCode) :);
#endif
@@ -636,36 +636,34 @@ Arm64JITCore::Arm64JITCore(FEXCore::Context::ContextImpl* ctx, FEXCore::Core::In
// Set up pointers that the JIT needs to load
// Common
auto& Common = ThreadState->CurrentFrame->Pointers.Common;
auto& Ptrs = ThreadState->CurrentFrame->Pointers;
Common.PrintValue = reinterpret_cast<uint64_t>(PrintValue);
Common.PrintVectorValue = reinterpret_cast<uint64_t>(PrintVectorValue);
Common.ThreadRemoveCodeEntryFromJIT = reinterpret_cast<uintptr_t>(&Context::ContextImpl::ThreadRemoveCodeEntryFromJit);
Common.MonoBackpatcherWrite = reinterpret_cast<uint64_t>(&Context::ContextImpl::MonoBackpatcherWrite);
Common.CPUIDObj = reinterpret_cast<uint64_t>(&CTX->CPUID);
Ptrs.PrintValue = reinterpret_cast<uint64_t>(PrintValue);
Ptrs.PrintVectorValue = reinterpret_cast<uint64_t>(PrintVectorValue);
Ptrs.PrintMsgValue = reinterpret_cast<uint64_t>(PrintMsg);
Ptrs.ThreadRemoveCodeEntryFromJIT = reinterpret_cast<uintptr_t>(&Context::ContextImpl::ThreadRemoveCodeEntryFromJit);
Ptrs.MonoBackpatcherWrite = reinterpret_cast<uint64_t>(&Context::ContextImpl::MonoBackpatcherWrite);
Ptrs.CPUIDObj = reinterpret_cast<uint64_t>(&CTX->CPUID);
{
FEXCore::Utils::MemberFunctionToPointerCast PMF(&FEXCore::CPUIDEmu::RunFunction);
Common.CPUIDFunction = PMF.GetConvertedPointer();
Ptrs.CPUIDFunction = PMF.GetConvertedPointer();
}
{
FEXCore::Utils::MemberFunctionToPointerCast PMF(&FEXCore::CPUIDEmu::RunXCRFunction);
Common.XCRFunction = PMF.GetConvertedPointer();
Ptrs.XCRFunction = PMF.GetConvertedPointer();
}
{
FEXCore::Utils::MemberFunctionToPointerCast PMF(&FEXCore::HLE::SyscallHandler::HandleSyscall);
Common.SyscallHandlerObj = reinterpret_cast<uint64_t>(CTX->SyscallHandler);
Common.SyscallHandlerFunc = PMF.GetVTableEntry(CTX->SyscallHandler);
Ptrs.SyscallHandlerObj = reinterpret_cast<uint64_t>(CTX->SyscallHandler);
Ptrs.SyscallHandlerFunc = PMF.GetVTableEntry(CTX->SyscallHandler);
}
Common.ExitFunctionLink = reinterpret_cast<uintptr_t>(&Arm64JITCore::ExitFunctionLink);
// Platform Specific
auto& AArch64 = ThreadState->CurrentFrame->Pointers.AArch64;
AArch64.LUDIV = reinterpret_cast<uint64_t>(LUDIV);
AArch64.LDIV = reinterpret_cast<uint64_t>(LDIV);
Ptrs.ExitFunctionLink = reinterpret_cast<uintptr_t>(&Arm64JITCore::ExitFunctionLink);
Ptrs.LUDIV = reinterpret_cast<uint64_t>(LUDIV);
Ptrs.LDIV = reinterpret_cast<uint64_t>(LDIV);
}
CurrentCodeBuffer = CodeBuffers.GetLatest();
@@ -684,7 +682,7 @@ void Arm64JITCore::ClearCache() {
auto lk = PrevCodeBuffer->LookupCache->AcquireWriteLock();
auto CodeBuffer = GetEmptyCodeBuffer();
SetBuffer(CodeBuffer->Ptr, CodeBuffer->Size);
SetBuffer(CodeBuffer->Ptr, CodeBuffer->AllocatedSize);
EmitDetectionString();
ThreadState->LookupCache->ChangeGuestToHostMapping(*PrevCodeBuffer, *CurrentCodeBuffer->LookupCache, lk);
@@ -764,7 +762,7 @@ void Arm64JITCore::EmitTFCheck() {
LoadConstant(ARMEmitter::Size::i64Bit, TMP1, Constant);
str(TMP1, STATE, offsetof(FEXCore::Core::CpuStateFrame, SynchronousFaultData));
ldr(TMP1, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.GuestSignal_SIGTRAP));
ldr(TMP1, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.GuestSignal_SIGTRAP));
br(TMP1);
(void)Bind(&l_TFBlocked);
@@ -782,7 +780,7 @@ void Arm64JITCore::EmitSuspendInterruptCheck() {
offsetof(FEXCore::Core::InternalThreadState, InterruptFaultPage) - offsetof(FEXCore::Core::InternalThreadState, BaseFrameState));
}
#ifdef _M_ARM_64EC
#ifdef ARCHITECTURE_arm64ec
static constexpr uint16_t SuspendMagic {0xCAFE};
ldr(TMP2.W(), STATE_PTR(CpuStateFrame, SuspendDoorbell));
@@ -820,19 +818,28 @@ void Arm64JITCore::EmitEntryPoint(ARMEmitter::BackwardLabel& HeaderLabel, bool C
CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size, bool SingleInst, const FEXCore::IR::IRListView* IR,
FEXCore::Core::DebugData* DebugData, bool CheckTF) {
FEXCORE_PROFILE_SCOPED("Arm64::CompileCode");
const auto PrevNumAllocations = Relocations.size();
this->Entry = Entry;
this->DebugData = DebugData;
this->IR = IR;
RequiresFarARM64Jumps = false;
SSANodeMultiplier = 24;
// Prepare restart via long jump in case branch encoding fails.
// This uses UncheckedLongJump since we don't implement std::longjmp in WoA setups
switch (static_cast<RestartOptions::Control>(FEXCore::UncheckedLongJump::SetJump(RestartControl.RestartJump))) {
switch (static_cast<RestartOptions::Control>(FEXCore::UncheckedLongJump::SetJump(ThreadState->RestartJump))) {
case RestartOptions::Control::Incoming:
// Nothing
break;
case RestartOptions::Control::EnableFarARM64Jumps: RequiresFarARM64Jumps = true; break;
default: ERROR_AND_DIE_FMT("Unhandled Arm64 restart condition!");
case RestartOptions::Control::NeedsLargerJITSpace:
// Get rid of the claimed buffer immediately, we can't fit in it at all.
TempAllocator.UnclaimBuffer();
SSANodeMultiplier *= 2;
break;
default: LOGMAN_MSG_A_FMT("Unhandled Arm64 restart condition!");
}
uint32_t SSACount = IR->GetSSACount();
@@ -840,16 +847,24 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
CallReturnTargets.clear();
PendingJumpThunks.clear();
JumpTargets.resize(IR->GetHeader()->BlockCount, {});
Relocations.resize(PrevNumAllocations, FEXCore::CPU::Relocation::Default()); // Discard any relocations generated from a previous attempt
CodeData.EntryPoints.clear();
// Fairly excessive buffer range to make sure we don't overflow
uint32_t BufferRange = 0x1000 + SSACount * 24;
// One page baseline, plus SSANodeMultipler bytes, plus another page for guard page.
const uint32_t DesiredBufferRange = AlignUp(FEXCore::Utils::FEX_PAGE_SIZE * 2 + SSACount * SSANodeMultiplier, FEXCore::Utils::FEX_PAGE_SIZE);
// JIT output is first written to a temporary buffer and later relocated to the CodeBuffer.
// This minimizes lock contention of CodeBufferWriteMutex.
auto TempCodeBuffer = TempAllocator.ReownOrClaimBuffer(BufferRange);
SetBuffer(TempCodeBuffer, BufferRange);
auto TempCodeBufferInfo = TempAllocator.ReownOrClaimBufferWithSize(DesiredBufferRange);
auto TempCodeBuffer = TempCodeBufferInfo.Ptr;
const uint32_t UsableBufferRange = TempCodeBufferInfo.Size - FEXCore::Utils::FEX_PAGE_SIZE;
SetBuffer(TempCodeBuffer, UsableBufferRange);
ThreadState->JITGuardPage = reinterpret_cast<uintptr_t>(TempCodeBuffer) + UsableBufferRange;
ThreadState->JITGuardOverflowArgument = FEXCore::ToUnderlying(RestartOptions::Control::NeedsLargerJITSpace);
CodeData.BlockBegin = GetCursorAddress<uint8_t*>();
@@ -979,22 +994,28 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
// This is a ExitFunctionLinkData struct
BindOrRestart(&l_ExitLink);
dc64(0); // HostCode
dc64(PendingJumpThunk.GuestRIP); // GuestRIP
dc64(PendingJumpThunk.CallerAddress - ThunkAddress); // CallerOffset
dc64(0); // HostCode
PlaceNamedSymbolLiteral(InsertGuestRIPLiteral(PendingJumpThunk.GuestRIP)); // GuestRIP
dc64(PendingJumpThunk.CallerAddress - ThunkAddress); // CallerOffset
}
BindOrRestart(&l_ExitLink);
dc64(ThreadState->CurrentFrame->Pointers.Common.ExitFunctionLinker);
PlaceNamedSymbolLiteral(InsertNamedSymbolLiteral(RelocNamedSymbolLiteral::NamedSymbol::SYMBOL_LITERAL_EXITFUNCTION_LINKER));
// CodeSize not including the header or tail data.
const uint64_t CodeOnlySize = GetCursorAddress<uint8_t*>() - CodeBegin;
// Add the JitCodeTail
// Add the JitCodeTail (written later)
Align(alignof(JITCodeTail));
auto JITBlockTailLocation = GetCursorAddress<uint8_t*>();
auto JITBlockTail = GetCursorAddress<JITCodeTail*>();
CursorIncrement(sizeof(JITCodeTail));
const auto JITBlockTailLocation = GetCursorAddress<uint8_t*>();
CodeHeader->OffsetToBlockTail = JITBlockTailLocation - CodeData.BlockBegin;
JITCodeTail JITBlockTail {
.RIP = Entry,
.GuestSize = Size,
.SpinLockFutex = 0,
.SingleInst = SingleInst,
};
// Entries that live after the JITCodeTail.
// These entries correlate JIT code regions with guest RIP regions.
@@ -1012,23 +1033,13 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
// FEXCore::Utils::vl64 GuestRIPOffset;
// };
auto JITRIPEntriesBegin = GetCursorAddress<uint8_t*>();
// Put the block's RIP entry in the tail.
// This will be used for RIP reconstruction in the future.
// TODO: This needs to be a data RIP relocation once code caching works.
// Current relocation code doesn't support this feature yet.
JITBlockTail->RIP = Entry;
JITBlockTail->GuestSize = Size;
JITBlockTail->SingleInst = SingleInst;
JITBlockTail->SpinLockFutex = 0;
const auto JITRIPEntriesBegin = JITBlockTailLocation + sizeof(JITBlockTail);
auto JITRIPEntriesLocation = JITRIPEntriesBegin;
{
// Store the RIP entries.
JITBlockTail->NumberOfRIPEntries = DebugData->GuestOpcodes.size();
JITBlockTail->OffsetToRIPEntries = JITRIPEntriesBegin - JITBlockTailLocation;
JITBlockTail.NumberOfRIPEntries = DebugData->GuestOpcodes.size();
JITBlockTail.OffsetToRIPEntries = JITRIPEntriesBegin - JITBlockTailLocation;
uintptr_t CurrentRIPOffset = 0;
uint64_t CurrentPCOffset = 0;
@@ -1044,14 +1055,20 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
}
}
CursorIncrement(JITRIPEntriesLocation - JITRIPEntriesBegin);
SetCursorOffset(JITRIPEntriesLocation - CodeData.BlockBegin);
Align();
CodeHeader->OffsetToBlockTail = JITBlockTailLocation - CodeData.BlockBegin;
CodeData.Size = GetCursorAddress<uint8_t*>() - CodeData.BlockBegin;
JITBlockTail->Size = CodeData.Size;
// Finalize and write block tail data
JITBlockTail.Size = CodeData.Size;
{
auto PrevCur = GetCursorOffset();
memcpy(JITBlockTailLocation, &JITBlockTail, sizeof(JITBlockTail));
SetCursorOffset(JITBlockTailLocation - CodeData.BlockBegin + offsetof(JITCodeTail, RIP));
PlaceNamedSymbolLiteral(InsertGuestRIPLiteral(JITBlockTail.RIP));
SetCursorOffset(PrevCur);
}
// Migrate the compile output from temporary storage to the actual CodeBuffer.
// This can block progress in other compiling threads, so the duration of the lock should be as small as possible.
@@ -1060,7 +1077,6 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
// Query size of generated code
const auto TempSize = GetCursorOffset();
LOGMAN_THROW_A_FMT(TempSize <= BufferRange, "Exceeded bounds of temporary buffer ({:#x} vs {:#x})", TempSize, BufferRange);
// Bring CodeBuffer up to date
{
@@ -1073,14 +1089,13 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
}
// NOTE: 16-byte alignment of the new cursor offset must be preserved for block linking records
SetBuffer(CurrentCodeBuffer->Ptr, CurrentCodeBuffer->Size);
SetCursorOffset(AlignUp(CodeBuffers.LatestOffset, 16));
if ((GetCursorOffset() + TempSize) > (CurrentCodeBuffer->Size - Utils::FEX_PAGE_SIZE)) {
SetBuffer(CurrentCodeBuffer->Ptr, CurrentCodeBuffer->AllocatedSize);
SetCursorOffset(CodeBuffers.LatestOffset);
Align16B();
if ((GetCursorOffset() + TempSize) > CurrentCodeBuffer->UsableSize()) {
CTX->ClearCodeCache(ThreadState);
}
Align16B();
CodeBuffers.LatestOffset = GetCursorOffset();
}
@@ -1092,6 +1107,10 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
}
CodeBegin += Delta;
for (std::size_t Idx = PrevNumAllocations; Idx != Relocations.size(); ++Idx) {
Relocations[Idx].Header.Offset += CodeBuffers.LatestOffset;
}
// Copy over CodeBuffer contents
memcpy(GetCursorAddress<uint8_t*>(), TempCodeBuffer, TempSize);
SetCursorOffset(CodeBuffers.LatestOffset + TempSize);
+41 -22
View File
@@ -68,10 +68,10 @@ private:
const bool HostSupportsAFP {};
struct RestartOptions {
FEXCore::UncheckedLongJump::JumpBuf RestartJump;
enum class Control : uint64_t {
Incoming = 0,
EnableFarARM64Jumps = 1,
NeedsLargerJITSpace = 2,
};
};
@@ -79,6 +79,8 @@ private:
// In the rare case when those assumptions are broken, FEX needs to safely restart the JIT.
RestartOptions RestartControl {};
bool RequiresFarARM64Jumps {};
// Default to 6 instructions per SSA node.
uint32_t SSANodeMultiplier {24};
ARMEmitter::BiDirectionalLabel* PendingTargetLabel {};
ARMEmitter::BiDirectionalLabel* PendingCallReturnTargetLabel {};
@@ -360,7 +362,7 @@ private:
// We can support this but currently unnecessary.
ERROR_AND_DIE_FMT("Tried to branch larger than 128MB away!");
FEXCore::UncheckedLongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
FEXCore::UncheckedLongJump::LongJump(ThreadState->RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
template<ARMEmitter::IsLabel T>
@@ -371,7 +373,7 @@ private:
// We can support this but currently unnecessary.
ERROR_AND_DIE_FMT("Tried to branch larger than 128MB away!");
FEXCore::UncheckedLongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
FEXCore::UncheckedLongJump::LongJump(ThreadState->RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
template<ARMEmitter::IsLabel T>
@@ -392,7 +394,7 @@ private:
return;
}
FEXCore::UncheckedLongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
FEXCore::UncheckedLongJump::LongJump(ThreadState->RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
template<ARMEmitter::IsLabel T>
@@ -413,7 +415,7 @@ private:
return;
}
FEXCore::UncheckedLongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
FEXCore::UncheckedLongJump::LongJump(ThreadState->RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
template<ARMEmitter::IsLabel T>
@@ -434,7 +436,7 @@ private:
return;
}
FEXCore::UncheckedLongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
FEXCore::UncheckedLongJump::LongJump(ThreadState->RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
template<ARMEmitter::IsLabel T>
@@ -455,7 +457,7 @@ private:
return;
}
FEXCore::UncheckedLongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
FEXCore::UncheckedLongJump::LongJump(ThreadState->RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
template<ARMEmitter::IsLabel T>
@@ -476,29 +478,37 @@ private:
return;
}
FEXCore::UncheckedLongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
FEXCore::UncheckedLongJump::LongJump(ThreadState->RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
template<ARMEmitter::IsLabel T>
void adr_OrRestart(ARMEmitter::Register rd, T* Label) {
if (RequiresFarARM64Jumps) {
if (LongAddressGen(rd, Label) == ARMEmitter::BranchEncodeSucceeded::Failure) {
ERROR_AND_DIE_FMT("Unable to encode long ADR.");
}
return;
}
if (adr(rd, Label) == ARMEmitter::BranchEncodeSucceeded::Success) {
return;
}
// We can support this but currently unnecessary.
ERROR_AND_DIE_FMT("Long ADR currently unsupported!");
FEXCore::UncheckedLongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
FEXCore::UncheckedLongJump::LongJump(ThreadState->RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
template<ARMEmitter::IsLabel T>
void adrp_OrRestart(ARMEmitter::Register rd, T* Label) {
if (RequiresFarARM64Jumps) {
if (LongAddressGen(rd, Label) == ARMEmitter::BranchEncodeSucceeded::Failure) {
ERROR_AND_DIE_FMT("Unable to encode long ADRP.");
}
return;
}
if (adrp(rd, Label) == ARMEmitter::BranchEncodeSucceeded::Success) {
return;
}
// We can support this but currently unnecessary.
ERROR_AND_DIE_FMT("Long ADRP currently unsupported!");
FEXCore::UncheckedLongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
FEXCore::UncheckedLongJump::LongJump(ThreadState->RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
template<ARMEmitter::IsLabel T>
@@ -513,7 +523,7 @@ private:
return;
}
FEXCore::UncheckedLongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
FEXCore::UncheckedLongJump::LongJump(ThreadState->RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
// This is purely a debugging aid for developers to see if they are in JIT code space when inspecting raw memory
@@ -526,8 +536,6 @@ private:
* @name Relocations
* @{ */
uint64_t GetNamedSymbolLiteral(FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op);
/**
* @brief A literal pair relocation object for named symbol literals
*/
@@ -564,19 +572,30 @@ private:
*/
NamedSymbolLiteralPair InsertNamedSymbolLiteral(FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op);
/**
* @brief Inserts a relocation for a constant value relative to the guest entrypoint
*
* @param Reg - The GPR to move the guest RIP in to
* @param Constant - The guest RIP that will be relocated
*/
NamedSymbolLiteralPair InsertGuestRIPLiteral(uint64_t GuestRIP);
/**
* @brief Place the named symbol literal relocation in memory
*
* @param Lit - Which literal to place
*/
void PlaceNamedSymbolLiteral(NamedSymbolLiteralPair& Lit);
void PlaceNamedSymbolLiteral(NamedSymbolLiteralPair Lit);
fextl::vector<FEXCore::CPU::Relocation> Relocations;
///< Relocation code loading
bool ApplyRelocations(uint64_t GuestEntry, std::span<std::byte> Code, std::span<const FEXCore::CPU::Relocation>);
fextl::vector<FEXCore::CPU::Relocation> TakeRelocations() override;
/**
* Returns any relocations generated since the last call to TakeRelocations.
*
* GuestBaseAddress must match the base virtual address to which the
* input x86 binary is mapped.
*/
fextl::vector<FEXCore::CPU::Relocation> TakeRelocations(uint64_t GuestBaseAddress) override;
/** @} */
+172 -81
View File
@@ -563,7 +563,25 @@ DEF_OP(LoadDF) {
auto Flag = X86State::RFLAG_DF_RAW_LOC;
// DF needs sign extension to turn 0x1/0xFF into 1/-1
ldrsb(Dst.X(), STATE, offsetof(FEXCore::Core::CPUState, flags[Flag]));
ldrsb(Dst.X(), STATE, ARRAY_OFFSETOF(FEXCore::Core::CPUState, flags, Flag));
}
DEF_OP(ContextClear) {
auto Op = IROp->C<IR::IROp_ContextClear>();
if (CTX->HostFeatures.SupportsCLZERO) {
// We can use CLZero directly when hardware supports it.
// Provides a fairly generous speed-up on Ampere1A hardware.
// TODO: When FEAT_MOPS hardware ships, test memset using MOPS.
for (size_t i = 0; i < Op->Size; i += 64) {
add(ARMEmitter::Size::i64Bit, TMP1, STATE.R(), Op->Offset + i);
dc(ARMEmitter::DataCacheOperation::ZVA, TMP1);
}
} else {
movi(ARMEmitter::SubRegSize::i64Bit, VTMP1.Q(), 0);
for (size_t i = 0; i < Op->Size; i += 32) {
stp<ARMEmitter::IndexType::OFFSET>(VTMP1.Q(), VTMP1.Q(), STATE.R(), Op->Offset + i);
}
}
}
ARMEmitter::ExtendedMemOperand Arm64JITCore::GenerateMemOperand(
@@ -1831,13 +1849,6 @@ DEF_OP(StoreMemTSO) {
}
DEF_OP(MemSet) {
// TODO: A future looking task would be to support this with ARM's MOPS instructions.
// The 8-bit non-atomic forward path directly matches ARM's SETP/SETM/SETE instruction,
// while the backward version needs some fixup to convert it to a forward direction.
//
// Assuming non-atomicity and non-faulting behaviour, this can accelerate this implementation.
// Additionally: This is commonly used as a memset to zero. If we know up-front with an inline constant
// that the value is zero, we can optimize any operation larger than 8-bit down to 8-bit to use the MOPS implementation.
const auto Op = IROp->C<IR::IROp_MemSet>();
const bool IsAtomic = CTX->IsMemcpyAtomicTSOEnabled();
@@ -1915,8 +1926,30 @@ DEF_OP(MemSet) {
ARMEmitter::SubRegSize::i8Bit;
auto EmitMemset = [&](int32_t Direction) {
const int32_t OpSize = Size;
const int32_t SizeDirection = Size * Direction;
const bool IsBackwards = Direction == -1;
// Sets the result to the final address written depending on
// whether or not the memset is forwards or backwards.
const auto MakeFinalAddress = [&] {
if (IsBackwards) {
switch (Size) {
case 1: sub(Dst.X(), MemReg.X(), Length.X()); break;
case 2:
case 4:
case 8: sub(Dst.X(), MemReg.X(), Length.X(), ARMEmitter::ShiftType::LSL, FEXCore::ilog2(Size)); break;
default: LOGMAN_MSG_A_FMT("Unhandled MemSet size: {}", Size); break;
}
} else {
switch (Size) {
case 1: add(Dst.X(), MemReg.X(), Length.X()); break;
case 2:
case 4:
case 8: add(Dst.X(), MemReg.X(), Length.X(), ARMEmitter::ShiftType::LSL, FEXCore::ilog2(Size)); break;
default: LOGMAN_MSG_A_FMT("Unhandled MemSet size: {}", Size); break;
}
}
};
ARMEmitter::BiDirectionalLabel AgainInternal {};
ARMEmitter::ForwardLabel DoneInternal {};
@@ -1925,12 +1958,56 @@ DEF_OP(MemSet) {
(void)cbz(ARMEmitter::Size::i64Bit, TMP1, &DoneInternal);
if (!IsAtomic) {
if (CTX->HostFeatures.SupportsMOPS) {
const bool Is8Bit = SubRegSize == ARMEmitter::SubRegSize::i8Bit;
// We can handle 8-bit memsets and any other size that happens
// to be using an inlined zero value (resulting in the use of ZR).
//
// NOTE:
// Strictly speaking, this can also be trivially expanded to handle other sizes
// that happen to use any value that could fit inside a byte if the need
// arises. This does increase branching and code generation, however, since
// we'd still need to emit the fallback in the event a value for a larger size
// falls outside the range of a byte instead of only generating the MOPS code.
if (Is8Bit || Value == ARMEmitter::Reg::zr) {
// If we're performing a non-byte-sized zeroing operation then we need to
// scale the counter accordingly. (e.g. a 64-bit memset of size 2 needs to
// be turned into an 8-bit memset of size 16)
if (!Is8Bit) {
lsl(ARMEmitter::Size::i64Bit, TMP1, TMP1, FEXCore::ToUnderlying(SubRegSize));
}
// If backwards, then we need to adjust the starting address because
// set{p, m, e} memset forwards, so we need to slide this bad boy
// back like: (address - count) + 1.
//
// This lets us offset the address such that we can treat a backwards
// memset as if it were a forwards one.
if (IsBackwards) {
sub(TMP2, TMP2, TMP1);
add(ARMEmitter::Size::i64Bit, TMP2, TMP2, 1);
}
// Unfortunately set operations fiddle with NZCV, so we need to preserve it.
mrs(TMP3, ARMEmitter::SystemRegister::NZCV);
setp(TMP2, TMP1, Value.X());
setm(TMP2, TMP1, Value.X());
sete(TMP2, TMP1, Value.X());
msr(ARMEmitter::SystemRegister::NZCV, TMP3);
MakeFinalAddress();
(void)Bind(&DoneInternal);
return;
}
}
ARMEmitter::ForwardLabel AgainInternal256Exit {};
ARMEmitter::BackwardLabel AgainInternal256 {};
ARMEmitter::ForwardLabel AgainInternal128Exit {};
ARMEmitter::BackwardLabel AgainInternal128 {};
if (Direction == -1) {
if (IsBackwards) {
sub(ARMEmitter::Size::i64Bit, TMP2, TMP2, 32 - Size);
}
@@ -1968,39 +2045,23 @@ DEF_OP(MemSet) {
add(ARMEmitter::Size::i64Bit, TMP1, TMP1, 32 / Size);
(void)cbz(ARMEmitter::Size::i64Bit, TMP1, &DoneInternal);
if (Direction == -1) {
if (IsBackwards) {
add(ARMEmitter::Size::i64Bit, TMP2, TMP2, 32 - Size);
}
}
(void)Bind(&AgainInternal);
if (IsAtomic) {
MemStoreTSO(Value, OpSize, SizeDirection);
MemStoreTSO(Value, Size, SizeDirection);
} else {
MemStore(Value, OpSize, SizeDirection);
MemStore(Value, Size, SizeDirection);
}
sub(ARMEmitter::Size::i64Bit, TMP1, TMP1, 1);
(void)cbnz(ARMEmitter::Size::i64Bit, TMP1, &AgainInternal);
(void)Bind(&DoneInternal);
if (SizeDirection >= 0) {
switch (OpSize) {
case 1: add(Dst.X(), MemReg.X(), Length.X()); break;
case 2: add(Dst.X(), MemReg.X(), Length.X(), ARMEmitter::ShiftType::LSL, 1); break;
case 4: add(Dst.X(), MemReg.X(), Length.X(), ARMEmitter::ShiftType::LSL, 2); break;
case 8: add(Dst.X(), MemReg.X(), Length.X(), ARMEmitter::ShiftType::LSL, 3); break;
default: LOGMAN_MSG_A_FMT("Unhandled {} size: {}", __func__, OpSize); break;
}
} else {
switch (OpSize) {
case 1: sub(Dst.X(), MemReg.X(), Length.X()); break;
case 2: sub(Dst.X(), MemReg.X(), Length.X(), ARMEmitter::ShiftType::LSL, 1); break;
case 4: sub(Dst.X(), MemReg.X(), Length.X(), ARMEmitter::ShiftType::LSL, 2); break;
case 8: sub(Dst.X(), MemReg.X(), Length.X(), ARMEmitter::ShiftType::LSL, 3); break;
default: LOGMAN_MSG_A_FMT("Unhandled {} size: {}", __func__, OpSize); break;
}
}
MakeFinalAddress();
};
if (DirectionIsInline) {
@@ -2023,10 +2084,6 @@ DEF_OP(MemSet) {
}
DEF_OP(MemCpy) {
// TODO: A future looking task would be to support this with ARM's MOPS instructions.
// The 8-bit non-atomic path directly matches ARM's CPYP/CPYM/CPYE instruction,
//
// Assuming non-atomicity and non-faulting behaviour, this can accelerate this implementation.
const auto Op = IROp->C<IR::IROp_MemCpy>();
const bool IsAtomic = CTX->IsMemcpyAtomicTSOEnabled();
@@ -2157,8 +2214,40 @@ DEF_OP(MemCpy) {
};
auto EmitMemcpy = [&](int32_t Direction) {
const int32_t OpSize = Size;
const int32_t SizeDirection = Size * Direction;
const bool IsBackwards = Direction == -1;
const auto FinalizeAddresses = [&] {
if (IsBackwards) {
switch (Size) {
case 1:
sub(Dst0.X(), TMP1, TMP3);
sub(Dst1.X(), TMP2, TMP3);
break;
case 2:
case 4:
case 8:
sub(Dst0.X(), TMP1, TMP3, ARMEmitter::ShiftType::LSL, FEXCore::ilog2(Size));
sub(Dst1.X(), TMP2, TMP3, ARMEmitter::ShiftType::LSL, FEXCore::ilog2(Size));
break;
default: LOGMAN_MSG_A_FMT("Unhandled MemCpy size: {}", Size); break;
}
} else {
switch (Size) {
case 1:
add(Dst0.X(), TMP1, TMP3);
add(Dst1.X(), TMP2, TMP3);
break;
case 2:
case 4:
case 8:
add(Dst0.X(), TMP1, TMP3, ARMEmitter::ShiftType::LSL, FEXCore::ilog2(Size));
add(Dst1.X(), TMP2, TMP3, ARMEmitter::ShiftType::LSL, FEXCore::ilog2(Size));
break;
default: LOGMAN_MSG_A_FMT("Unhandled MemCpy size: {}", Size); break;
}
}
};
ARMEmitter::BiDirectionalLabel AgainInternal {};
ARMEmitter::ForwardLabel DoneInternal {};
@@ -2167,6 +2256,48 @@ DEF_OP(MemCpy) {
(void)cbz(ARMEmitter::Size::i64Bit, TMP1, &DoneInternal);
if (!IsAtomic) {
if (CTX->HostFeatures.SupportsMOPS) {
// In the event we have an overlap (gross), we need to fall back
// to the non-mops copy handler. Since the overlap check needs to
// make use of NZCV, we need to save it. This can be avoided with
// ARMv9.6+'s FEAT_CMPBR, but alas, we don't have access to that right now.
//
// NOTE: That we need to temporarily trash TMP1 and restore it after the
// comparison.
ARMEmitter::ForwardLabel OverlapCase;
mrs(TMP4, ARMEmitter::SystemRegister::NZCV);
sub(ARMEmitter::Size::i64Bit, TMP1, TMP2, TMP3);
cmp(ARMEmitter::Size::i64Bit, TMP1, Length.X());
mov(TMP1, Length.X());
(void)bc(ARMEmitter::Condition::CC_LT, &OverlapCase);
// If doing something larger than a byte copy, then we need to scale
// the counter value accordingly to convert it to bytes.
if (Size > 1) {
lsl(ARMEmitter::Size::i64Bit, TMP1, TMP1, FEXCore::ilog2(Size));
}
// Adjust addresses so that we treat the backward copy as a forward copy
if (IsBackwards) {
sub(ARMEmitter::Size::i64Bit, TMP2, TMP2, TMP1);
sub(ARMEmitter::Size::i64Bit, TMP3, TMP3, TMP1);
add(ARMEmitter::Size::i64Bit, TMP2, TMP2, Size);
add(ARMEmitter::Size::i64Bit, TMP3, TMP3, Size);
}
// Unfortunately copy operations fiddle with NZCV, so we need to preserve it.
cpyfp(TMP2, TMP3, TMP1);
cpyfm(TMP2, TMP3, TMP1);
cpyfe(TMP2, TMP3, TMP1);
msr(ARMEmitter::SystemRegister::NZCV, TMP4);
(void)b(&DoneInternal);
// Turns out we overlap and need to fall back. Make sure to restore NZCV.
(void)Bind(&OverlapCase);
msr(ARMEmitter::SystemRegister::NZCV, TMP4);
}
ARMEmitter::ForwardLabel AbsPos {};
ARMEmitter::ForwardLabel AgainInternal256Exit {};
ARMEmitter::ForwardLabel AgainInternal128Exit {};
@@ -2180,7 +2311,7 @@ DEF_OP(MemCpy) {
sub(ARMEmitter::Size::i64Bit, TMP4, TMP4, 32);
(void)tbnz(TMP4, 63, &AgainInternal);
if (Direction == -1) {
if (IsBackwards) {
sub(ARMEmitter::Size::i64Bit, TMP2, TMP2, 32 - Size);
sub(ARMEmitter::Size::i64Bit, TMP3, TMP3, 32 - Size);
}
@@ -2215,7 +2346,7 @@ DEF_OP(MemCpy) {
add(ARMEmitter::Size::i64Bit, TMP1, TMP1, 32 / Size);
(void)cbz(ARMEmitter::Size::i64Bit, TMP1, &DoneInternal);
if (Direction == -1) {
if (IsBackwards) {
add(ARMEmitter::Size::i64Bit, TMP2, TMP2, 32 - Size);
add(ARMEmitter::Size::i64Bit, TMP3, TMP3, 32 - Size);
}
@@ -2223,9 +2354,9 @@ DEF_OP(MemCpy) {
(void)Bind(&AgainInternal);
if (IsAtomic) {
MemCpyTSO(OpSize, SizeDirection);
MemCpyTSO(Size, SizeDirection);
} else {
MemCpy(OpSize, SizeDirection);
MemCpy(Size, SizeDirection);
}
sub(ARMEmitter::Size::i64Bit, TMP1, TMP1, 1);
(void)cbnz(ARMEmitter::Size::i64Bit, TMP1, &AgainInternal);
@@ -2237,54 +2368,14 @@ DEF_OP(MemCpy) {
mov(TMP2, MemRegSrc.X());
mov(TMP3, Length.X());
if (SizeDirection >= 0) {
switch (OpSize) {
case 1:
add(Dst0.X(), TMP1, TMP3);
add(Dst1.X(), TMP2, TMP3);
break;
case 2:
add(Dst0.X(), TMP1, TMP3, ARMEmitter::ShiftType::LSL, 1);
add(Dst1.X(), TMP2, TMP3, ARMEmitter::ShiftType::LSL, 1);
break;
case 4:
add(Dst0.X(), TMP1, TMP3, ARMEmitter::ShiftType::LSL, 2);
add(Dst1.X(), TMP2, TMP3, ARMEmitter::ShiftType::LSL, 2);
break;
case 8:
add(Dst0.X(), TMP1, TMP3, ARMEmitter::ShiftType::LSL, 3);
add(Dst1.X(), TMP2, TMP3, ARMEmitter::ShiftType::LSL, 3);
break;
default: LOGMAN_MSG_A_FMT("Unhandled {} size: {}", __func__, OpSize); break;
}
} else {
switch (OpSize) {
case 1:
sub(Dst0.X(), TMP1, TMP3);
sub(Dst1.X(), TMP2, TMP3);
break;
case 2:
sub(Dst0.X(), TMP1, TMP3, ARMEmitter::ShiftType::LSL, 1);
sub(Dst1.X(), TMP2, TMP3, ARMEmitter::ShiftType::LSL, 1);
break;
case 4:
sub(Dst0.X(), TMP1, TMP3, ARMEmitter::ShiftType::LSL, 2);
sub(Dst1.X(), TMP2, TMP3, ARMEmitter::ShiftType::LSL, 2);
break;
case 8:
sub(Dst0.X(), TMP1, TMP3, ARMEmitter::ShiftType::LSL, 3);
sub(Dst1.X(), TMP2, TMP3, ARMEmitter::ShiftType::LSL, 3);
break;
default: LOGMAN_MSG_A_FMT("Unhandled {} size: {}", __func__, OpSize); break;
}
}
FinalizeAddresses();
};
if (DirectionIsInline) {
LOGMAN_THROW_A_FMT(DirectionConstant == 1 || DirectionConstant == -1, "unexpected direction");
EmitMemcpy(DirectionConstant);
} else {
// Emit forward direction memset then backward direction memset.
// Emit forward direction memcpy then backward direction memcpy.
for (int32_t Direction : {1, -1}) {
EmitMemcpy(Direction);
if (Direction == 1) {
+40 -12
View File
@@ -78,19 +78,19 @@ DEF_OP(Break) {
switch (Op->Reason.Signal) {
case Core::FAULT_SIGILL:
ldr(TMP1, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.GuestSignal_SIGILL));
ldr(TMP1, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.GuestSignal_SIGILL));
br(TMP1);
break;
case Core::FAULT_SIGTRAP:
ldr(TMP1, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.GuestSignal_SIGTRAP));
ldr(TMP1, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.GuestSignal_SIGTRAP));
br(TMP1);
break;
case Core::FAULT_SIGSEGV:
ldr(TMP1, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.GuestSignal_SIGSEGV));
ldr(TMP1, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.GuestSignal_SIGSEGV));
br(TMP1);
break;
default:
ldr(TMP1, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.GuestSignal_SIGTRAP));
ldr(TMP1, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.GuestSignal_SIGTRAP));
br(TMP1);
break;
}
@@ -189,11 +189,11 @@ DEF_OP(Print) {
if (IsGPR(Op->Value)) {
mov(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, GetReg(Op->Value));
ldr(ARMEmitter::XReg::x3, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.PrintValue));
ldr(ARMEmitter::XReg::x3, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.PrintValue));
} else {
fmov(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, GetVReg(Op->Value), false);
fmov(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r1, GetVReg(Op->Value), true);
ldr(ARMEmitter::XReg::x3, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.PrintVectorValue));
ldr(ARMEmitter::XReg::x3, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.PrintVectorValue));
}
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
@@ -210,6 +210,25 @@ DEF_OP(Print) {
PopDynamicRegs();
}
DEF_OP(PrintMsg) {
auto Op = IROp->C<IR::IROp_PrintMsg>();
PushDynamicRegs(TMP1);
SpillStaticRegs(TMP1);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, reinterpret_cast<uintptr_t>(Op->Value));
ldr(ARMEmitter::XReg::x1, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.PrintMsgValue));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<void, uint64_t>(ARMEmitter::Reg::r1);
} else {
blr(ARMEmitter::Reg::r1);
}
FillStaticRegs();
PopDynamicRegs();
}
DEF_OP(ProcessorID) {
if (CTX->HostFeatures.SupportsCPUIndexInTPIDRRO) {
mrs(GetReg(Node), ARMEmitter::SystemRegister::TPIDRRO_EL0);
@@ -227,7 +246,10 @@ DEF_OP(ProcessorID) {
// Ordering is incredibly important here
// We must spill any overlapping registers first THEN claim we are in a syscall without invalidating state at all
// Only spill the registers that intersect with our usage
SpillStaticRegs(TMP1, false, SpillMask);
SpillStaticRegs(TMP1, {
.GPRSpillMask = SpillMask,
.FPRs = false,
});
// Now that we are spilled, store in the state that we are in a syscall
// Still without overwriting registers that matter
@@ -241,7 +263,7 @@ DEF_OP(ProcessorID) {
sub(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, ARMEmitter::Reg::rsp, 16);
// Load the getcpu syscall number
#if defined(_M_X86_64)
#if defined(ARCHITECTURE_x86_64)
// Just to ensure the syscall number doesn't change if compiled for an x86_64 host.
constexpr auto GetCPUSyscallNum = 0xa8;
#else
@@ -264,7 +286,13 @@ DEF_OP(ProcessorID) {
// Now that we are done in the syscall we need to carefully peel back the state
// First unspill the registers from before
FillStaticRegs(false, SpillMask, ~0U, ARMEmitter::Reg::r8, ARMEmitter::Reg::r2);
FillStaticRegs({
.OptionalReg = ARMEmitter::Reg::r8,
.OptionalReg2 = ARMEmitter::Reg::r2,
.GPRFillMask = SpillMask,
.FPRs = false,
});
// Now the registers we've spilled are back in their original host registers
// We can safely claim we are no longer in a syscall
@@ -305,20 +333,20 @@ DEF_OP(MonoBackpatcherWrite) {
mov(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r3, TMP4);
}
#ifdef _M_ARM_64EC
#ifdef ARCHITECTURE_arm64ec
ldr(TMP2, ARMEmitter::XReg::x18, TEB_CPU_AREA_OFFSET);
LoadConstant(ARMEmitter::Size::i32Bit, TMP1, 1);
strb(TMP1.W(), TMP2, CPU_AREA_IN_SYSCALL_CALLBACK_OFFSET);
#endif
ldr(ARMEmitter::XReg::x4, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.MonoBackpatcherWrite));
ldr(ARMEmitter::XReg::x4, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.MonoBackpatcherWrite));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<void, void*, uint8_t, uint64_t, uint64_t>(ARMEmitter::Reg::r4);
} else {
blr(ARMEmitter::Reg::r4);
}
#ifdef _M_ARM_64EC
#ifdef ARCHITECTURE_arm64ec
ldr(TMP2, ARMEmitter::XReg::x18, TEB_CPU_AREA_OFFSET);
strb(ARMEmitter::WReg::zr, TMP2, CPU_AREA_IN_SYSCALL_CALLBACK_OFFSET);
#endif
+45 -24
View File
@@ -1,79 +1,100 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <FEXCore/IR/IR.h>
#include <FEXCore/Utils/CompilerDefs.h>
namespace FEXCore::Context {
class ContextImpl;
}
namespace FEXCore::CPU {
enum class RelocationTypes : uint8_t {
enum class RelocationTypes : uint32_t {
// 8 byte literal in memory for symbol
// Aligned to struct RelocNamedSymbolLiteral
RELOC_NAMED_SYMBOL_LITERAL,
// Fixed size named thunk move
// 4 instruction constant generation on AArch64
// 64-bit mov on x86-64
// 4 instruction constant generation
// Aligned to struct RelocNamedThunkMove
RELOC_NAMED_THUNK_MOVE,
// 8 byte literal (relative to binary base address)
RELOC_GUEST_RIP_LITERAL,
// Fixed size guest RIP move
// 4 instruction constant generation on AArch64
// 64-bit mov on x86-64
// Aligned to struct RelocGuestRIPMove
// 4 instruction constant generation
// Aligned to struct RelocGuestRIP
RELOC_GUEST_RIP_MOVE,
};
struct RelocationTypeHeader final {
struct FEX_PACKED RelocationHeader final {
// Offset to the relocated host code data
uint64_t Offset {};
RelocationTypes Type;
};
struct RelocNamedSymbolLiteral final {
enum class NamedSymbol : uint8_t {
enum class NamedSymbol : uint32_t {
///< Thread specific relocations
// JIT Literal pointers
SYMBOL_LITERAL_EXITFUNCTION_LINKER,
};
RelocationTypeHeader Header {};
RelocationHeader Header {};
NamedSymbol Symbol;
// Offset in to the code section to begin the relocation
uint64_t Offset {};
uint32_t Pad[8];
};
struct RelocNamedThunkMove final {
RelocationTypeHeader Header {};
RelocationHeader Header {};
// GPR index the constant is being moved to
uint8_t RegisterIndex;
uint32_t RegisterIndex;
// The thunk SHA256 hash
IR::SHA256Sum Symbol;
// Offset in to the code section to begin the relocation
uint64_t Offset {};
};
struct RelocGuestRIPMove final {
RelocationTypeHeader Header {};
struct RelocGuestRIP final {
RelocationHeader Header {};
// GPR index the constant is being moved to
// GPR index the constant is being moved to (for non-literal relocations)
uint8_t RegisterIndex;
// Offset in to the code section to begin the relocation
uint64_t Offset {};
char Pad[3];
// The unrelocated RIP that is being moved
// The base RIP (to be moved by the register for non-literal relocations).
// In a serialized code cache, this is relative to the binary base address.
uint64_t GuestRIP;
uint32_t pad2[6] {};
};
union Relocation {
RelocationTypeHeader Header {};
// Clang 16 Can't default-initialize this union
static Relocation Default() {
#if __clang_major__ < 17
Relocation Ret {.Header {}};
memset(&Ret, 0, sizeof(Ret));
return Ret;
#else
return {};
#endif
}
RelocationHeader Header {};
RelocNamedSymbolLiteral NamedSymbolLiteral;
// This makes our union of relocations at least 48 bytes
// It might be more efficient to not use a union
RelocNamedThunkMove NamedThunkMove;
RelocGuestRIPMove GuestRIPMove;
RelocGuestRIP GuestRIP;
};
uint64_t GetNamedSymbolLiteral(FEXCore::Context::ContextImpl&, RelocNamedSymbolLiteral::NamedSymbol);
} // namespace FEXCore::CPU
+135 -11
View File
@@ -977,7 +977,7 @@ DEF_OP(LoadNamedVectorConstant) {
}
// Load the pointer.
auto GenerateMemOperand = [this](IR::OpSize OpSize, uint32_t NamedConstant, ARMEmitter::Register Base) {
const auto ConstantOffset = offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.NamedVectorConstants[NamedConstant]);
const auto ConstantOffset = ARRAY_OFFSETOF(FEXCore::Core::CpuStateFrame, Pointers.NamedVectorConstants, NamedConstant);
if (ConstantOffset <= 255 || // Unscaled 9-bit signed
((ConstantOffset & (IR::OpSizeToSize(OpSize) - 1)) == 0 &&
@@ -985,13 +985,13 @@ DEF_OP(LoadNamedVectorConstant) {
return ARMEmitter::ExtendedMemOperand(Base.X(), ARMEmitter::IndexType::OFFSET, ConstantOffset);
}
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.Common.NamedVectorConstantPointers[NamedConstant]));
ldr(TMP1, STATE_PTR_IDX(CpuStateFrame, Pointers.NamedVectorConstantPointers, NamedConstant));
return ARMEmitter::ExtendedMemOperand(TMP1, ARMEmitter::IndexType::OFFSET, 0);
};
if (OpSize == IR::OpSize::i256Bit) {
// Handle SVE 32-byte variant upfront.
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.Common.NamedVectorConstantPointers[Op->Constant]));
ldr(TMP1, STATE_PTR_IDX(CpuStateFrame, Pointers.NamedVectorConstantPointers, Op->Constant));
ld1b<ARMEmitter::SubRegSize::i8Bit>(Dst.Z(), PRED_TMP_32B.Zeroing(), TMP1, 0);
return;
}
@@ -1013,7 +1013,7 @@ DEF_OP(LoadNamedVectorIndexedConstant) {
const auto Dst = GetVReg(Node);
// Load the pointer.
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.Common.IndexedNamedVectorConstantPointers[Op->Constant]));
ldr(TMP1, STATE_PTR_IDX(CpuStateFrame, Pointers.IndexedNamedVectorConstantPointers, Op->Constant));
switch (OpSize) {
case IR::OpSize::i8Bit: ldrb(Dst, TMP1, Op->Index); break;
@@ -1117,6 +1117,29 @@ DEF_OP(VAddP) {
}
}
DEF_OP(VOrn) {
const auto Op = IROp->C<IR::IROp_VOrn>();
const auto OpSize = IROp->Size;
const auto Is256Bit = OpSize == IR::OpSize::i256Bit;
const auto Is128Bit = OpSize == IR::OpSize::i128Bit;
LOGMAN_THROW_A_FMT(!Is256Bit || HostSupportsSVE256, "Need SVE256 support in order to use {} with 256-bit operation", __func__);
const auto Dst = GetVReg(Node);
const auto Vector1 = GetVReg(Op->Vector1);
const auto Vector2 = GetVReg(Op->Vector2);
if (HostSupportsSVE256 && Is256Bit) {
const auto Pred = PRED_TMP_32B.Merging();
not_(ARMEmitter::SubRegSize::i8Bit, VTMP1.Z(), Pred, Vector2.Z());
orr(Dst.Z(), Vector1.Z(), VTMP1.Z());
} else if (Is128Bit) {
orn(Dst.Q(), Vector1.Q(), Vector2.Q());
} else {
orn(Dst.D(), Vector1.D(), Vector2.D());
}
}
DEF_OP(VFAddV) {
const auto Op = IROp->C<IR::IROp_VFAddV>();
const auto OpSize = IROp->Size;
@@ -1411,8 +1434,8 @@ DEF_OP(VFMin) {
bif(Dst.Q(), Vector2.Q(), VTMP1.Q());
} else if (Dst == Vector2) {
// Destination is already Vector2, Invert arguments and insert Vector1 on false.
fcmgt(SubRegSize, VTMP1.Q(), Vector1.Q(), Vector2.Q());
bif(Dst.Q(), Vector1.Q(), VTMP1.Q());
fcmgt(SubRegSize, VTMP1.Q(), Vector2.Q(), Vector1.Q());
bit(Dst.Q(), Vector1.Q(), VTMP1.Q());
} else {
// Dst is not either source, need a move.
fcmgt(SubRegSize, VTMP1.Q(), Vector2.Q(), Vector1.Q());
@@ -1443,7 +1466,8 @@ DEF_OP(VFMax) {
const auto Mask = PRED_TMP_32B;
const auto ComparePred = ARMEmitter::PReg::p0;
fcmgt(SubRegSize, ComparePred, Mask.Zeroing(), Vector2.Z(), Vector1.Z());
fcmgt(SubRegSize, ComparePred, Mask.Zeroing(), Vector1.Z(), Vector2.Z());
not_(ComparePred, Mask.Zeroing(), ComparePred);
if (Dst == Vector1) {
// Trivial case where Vector1 is also the destination.
@@ -1465,17 +1489,17 @@ DEF_OP(VFMax) {
if (Dst == Vector1) {
// Destination is already Vector1, need to insert Vector2 on true.
fcmgt(SubRegSize, VTMP1.Q(), Vector2.Q(), Vector1.Q());
bit(Dst.Q(), Vector2.Q(), VTMP1.Q());
fcmgt(SubRegSize, VTMP1.Q(), Vector1.Q(), Vector2.Q());
bif(Dst.Q(), Vector2.Q(), VTMP1.Q());
} else if (Dst == Vector2) {
// Destination is already Vector2, Invert arguments and insert Vector1 on true.
fcmgt(SubRegSize, VTMP1.Q(), Vector1.Q(), Vector2.Q());
bit(Dst.Q(), Vector1.Q(), VTMP1.Q());
} else {
// Dst is not either source, need a move.
fcmgt(SubRegSize, VTMP1.Q(), Vector2.Q(), Vector1.Q());
fcmgt(SubRegSize, VTMP1.Q(), Vector1.Q(), Vector2.Q());
mov(Dst.Q(), Vector1.Q());
bit(Dst.Q(), Vector2.Q(), VTMP1.Q());
bif(Dst.Q(), Vector2.Q(), VTMP1.Q());
}
}
}
@@ -4588,4 +4612,104 @@ DEF_OP(VFCopySign) {
}
}
DEF_OP(F64SIN) {
const auto Op = IROp->C<IR::IROp_F64SIN>();
const auto Src = GetVReg(Op->Src);
const auto Dst = GetVReg(Node);
fmov(VTMP1.D(), Src.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.F64SinHandler));
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
fmov(Dst.D(), VTMP1.D());
}
DEF_OP(F64COS) {
const auto Op = IROp->C<IR::IROp_F64COS>();
const auto Src = GetVReg(Op->Src);
const auto Dst = GetVReg(Node);
fmov(VTMP1.D(), Src.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.F64CosHandler));
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
fmov(Dst.D(), VTMP1.D());
}
DEF_OP(F64TAN) {
const auto Op = IROp->C<IR::IROp_F64TAN>();
const auto Src = GetVReg(Op->Src);
const auto Dst = GetVReg(Node);
fmov(VTMP1.D(), Src.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.F64TanHandler));
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
fmov(Dst.D(), VTMP1.D());
}
// Src1=y(ST1), Src2=x(ST0). Marshal into VTMP1/VTMP2 and dispatch the shared handler.
DEF_OP(F64ATAN) {
const auto Op = IROp->C<IR::IROp_F64ATAN>();
const auto Src1 = GetVReg(Op->Src1);
const auto Src2 = GetVReg(Op->Src2);
const auto Dst = GetVReg(Node);
fmov(VTMP1.D(), Src1.D());
fmov(VTMP2.D(), Src2.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.F64AtanHandler));
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
fmov(Dst.D(), VTMP1.D());
}
// Src=x(ST0), Src2=y(ST1). Marshal into VTMP1/VTMP2 and dispatch the shared handler.
DEF_OP(F64FYL2X) {
const auto Op = IROp->C<IR::IROp_F64FYL2X>();
const auto Src = GetVReg(Op->Src);
const auto Src2 = GetVReg(Op->Src2);
const auto Dst = GetVReg(Node);
fmov(VTMP1.D(), Src.D());
fmov(VTMP2.D(), Src2.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.F64FYL2XHandler));
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
fmov(Dst.D(), VTMP1.D());
}
DEF_OP(F64SCALE) {
const auto Op = IROp->C<IR::IROp_F64SCALE>();
const auto Src1 = GetVReg(Op->Src1);
const auto Src2 = GetVReg(Op->Src2);
const auto Dst = GetVReg(Node);
fmov(VTMP1.D(), Src1.D());
fmov(VTMP2.D(), Src2.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.F64ScaleHandler));
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
fmov(Dst.D(), VTMP1.D());
}
DEF_OP(F64F2XM1) {
const auto Op = IROp->C<IR::IROp_F64F2XM1>();
const auto Src = GetVReg(Op->Src);
const auto Dst = GetVReg(Node);
fmov(VTMP1.D(), Src.D());
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.F64F2XM1Handler));
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
blr(TMP1);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
fmov(Dst.D(), VTMP1.D());
}
} // namespace FEXCore::CPU
@@ -41,6 +41,9 @@ LookupCache::LookupCache(FEXCore::Context::ContextImpl* CTX)
PagePointer = reinterpret_cast<uintptr_t>(FEXCore::Allocator::VirtualAlloc(TotalCacheSize, false, false));
LOGMAN_THROW_A_FMT(PagePointer != -1ULL, "Failed to allocate PagePointer");
// Disable THP on the Lookup cache.
FEXCore::Allocator::VirtualTHPControl(reinterpret_cast<const void*>(PagePointer), TotalCacheSize, FEXCore::Allocator::THPControl::Disable);
FEXCore::Allocator::VirtualName("FEXMem_Lookup", reinterpret_cast<void*>(PagePointer),
ctx->Config.VirtualMemSize / FEXCore::Utils::FEX_PAGE_SIZE * 8 + CODE_SIZE);
CTX->SyscallHandler->MarkOvercommitRange(PagePointer, TotalCacheSize);
@@ -75,7 +78,7 @@ LookupCache::~LookupCache() {
// These will get freed when their memory allocators are deallocated.
}
void LookupCache::ClearL2Cache(const FEXCore::LookupCacheReadLockToken& lk) {
void LookupCache::ClearL2Cache(const FEXCore::LookupCacheBaseLockToken& lk) {
// Clear out the page memory
// PagePointer and PageMemory are sequential with each other. Clear both at once.
FEXCore::Allocator::VirtualDontNeed(reinterpret_cast<void*>(PagePointer),
@@ -86,12 +89,12 @@ void LookupCache::ClearL2Cache(const FEXCore::LookupCacheReadLockToken& lk) {
void LookupCache::ClearThreadLocalCaches(const LookupCacheWriteLockToken&) {
// Clear L1 and L2 by clearing the full cache.
FEXCore::Allocator::VirtualDontNeed(reinterpret_cast<void*>(PagePointer), TotalCacheSize, false);
CachedCodePages.clear();
}
void LookupCache::ClearCache(const LookupCacheWriteLockToken& lk) {
// Clear L1 and L2 by clearing the full cache.
FEXCore::Allocator::VirtualDontNeed(reinterpret_cast<void*>(PagePointer), TotalCacheSize, false);
ClearThreadLocalCaches(lk);
Shared->ClearCache(lk);
}
+78 -36
View File
@@ -3,11 +3,12 @@
#include "Interface/Context/Context.h"
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/SHMStats.h>
#include "Utils/WritePriorityMutex.h"
#include <FEXCore/Utils/WritePriorityMutex.h>
#include <FEXCore/fextl/map.h>
#include <FEXCore/fextl/memory_resource.h>
#include <FEXCore/fextl/robin_map.h>
#include <FEXCore/fextl/robin_set.h>
#include <FEXCore/fextl/vector.h>
#include <FEXCore/fextl/memory_resource.h>
@@ -17,7 +18,13 @@
#include <mutex>
namespace FEXCore {
struct LookupCacheWriteLockToken {
struct LookupCacheBaseLockToken {
protected:
// Protected constructor - only derived classes can construct
LookupCacheBaseLockToken() = default;
};
struct LookupCacheWriteLockToken : public LookupCacheBaseLockToken {
private:
// Only constructible by GuestToHostMap
friend struct GuestToHostMap;
@@ -26,7 +33,7 @@ private:
std::lock_guard<FEXCore::Utils::WritePriorityMutex::Mutex> Lock;
};
struct LookupCacheReadLockToken {
struct LookupCacheReadLockToken : public LookupCacheBaseLockToken {
private:
// Only constructible by GuestToHostMap
friend struct GuestToHostMap;
@@ -74,48 +81,67 @@ struct GuestToHostMap {
fextl::unique_ptr<std::pmr::polymorphic_allocator<std::byte>> BlockLinks_pma;
BlockLinksMapType* BlockLinks;
fextl::robin_map<uint64_t, uint64_t> BlockList;
struct BlockEntry {
uint64_t HostCode;
fextl::vector<uint64_t> CodePages;
};
fextl::robin_map<uint64_t, BlockEntry> BlockList;
fextl::map<uint64_t, fextl::vector<uint64_t>> CodePages;
GuestToHostMap();
// Adds to Guest -> Host code mapping
void AddBlockMapping(uint64_t Address, void* HostCode, const LookupCacheWriteLockToken&) {
const BlockEntry& AddBlockMapping(uint64_t Address, const fextl::vector<uint64_t>& CodePages, void* HostCode, const LookupCacheWriteLockToken&) {
// This may replace an existing mapping
// NOTE: Generally no previous entry should exist, however there is one exception:
// If the backend updates the active thread's CodeBuffer, the new associated LookupCache
// may already contain the block address. Since is comparatively rare, we'll just leak
// one of the two blocks in this case.
BlockList[Address] = (uintptr_t)HostCode;
return BlockList.insert_or_assign(Address, BlockEntry {(uintptr_t)HostCode, CodePages}).first->second;
}
std::optional<uintptr_t> FindBlock(uint64_t Address, const LookupCacheReadLockToken&) {
const BlockEntry* FindBlock(uint64_t Address, const LookupCacheReadLockToken&) {
auto HostCode = BlockList.find(Address);
if (HostCode == BlockList.end()) {
return std::nullopt;
return nullptr;
}
return HostCode->second;
return &HostCode->second;
}
bool Erase(FEXCore::Core::CpuStateFrame* Frame, uint64_t Address, const LookupCacheWriteLockToken&) {
bool Erase(uint64_t Address, const LookupCacheWriteLockToken&) {
// Sever any links to this block
auto lower = BlockLinks->lower_bound({Address, nullptr});
auto upper = BlockLinks->upper_bound({Address, reinterpret_cast<FEXCore::Context::ExitFunctionLinkData*>(UINTPTR_MAX)});
for (auto it = lower; it != upper; it = BlockLinks->erase(it)) {
it->second(Frame, it->first.HostLink);
it->second(it->first.HostLink);
}
// Remove from BlockList
return BlockList.erase(Address) != 0;
}
void InvalidateRange(uint64_t Start, uint64_t Length) {
auto lk = AcquireWriteLock();
auto lower = CodePages.lower_bound(Start >> 12);
auto upper = CodePages.upper_bound((Start + Length - 1) >> 12);
for (auto it = lower; it != upper; it++) {
for (const auto& Entry : it->second) {
Erase(Entry, lk);
}
}
CodePages.erase(lower, upper);
}
void AddBlockLink(uint64_t GuestDestination, FEXCore::Context::ExitFunctionLinkData* HostLink,
const FEXCore::Context::BlockDelinkerFunc& delinker, const LookupCacheWriteLockToken&) {
BlockLinks->insert({{GuestDestination, HostLink}, delinker});
}
bool AddBlockExecutableRange(const fextl::set<uint64_t>& Addresses, uint64_t Start, uint64_t Length, const LookupCacheWriteLockToken&) {
bool AddBlockExecutableRange(const std::ranges::input_range auto& Addresses, uint64_t Start, uint64_t Length, const LookupCacheWriteLockToken&) {
bool rv = false;
for (auto CurrentPage = Start >> 12, EndPage = (Start + Length - 1) >> 12; CurrentPage <= EndPage; CurrentPage++) {
@@ -185,10 +211,10 @@ public:
if (!HostPtr) {
// Try L3
auto HostCode = Shared->FindBlock(Address, lk);
if (HostCode) {
CacheBlockMapping(Address, HostCode.value(), lk);
HostPtr = HostCode.value();
auto Entry = Shared->FindBlock(Address, lk);
if (Entry) {
CacheBlockMapping(Address, *Entry, false, lk);
HostPtr = Entry->HostCode;
}
}
}
@@ -259,32 +285,25 @@ public:
}
// Adds to Guest -> Host code mapping
void AddBlockMapping(FEXCore::Core::InternalThreadState* Thread, uint64_t Address, void* HostCode) {
void AddBlockMapping(FEXCore::Core::InternalThreadState* Thread, uint64_t Address, const fextl::vector<uint64_t>& CodePages, void* HostCode) {
std::optional<FEXCore::SHMStats::AccumulationBlock<uint64_t>> LockTime(
Thread->ThreadStats ? &Thread->ThreadStats->AccumulatedCacheWriteLockTime : nullptr);
auto lk = Shared->AcquireWriteLock();
LockTime.reset();
Shared->AddBlockMapping(Address, HostCode, lk);
const auto& Entry = Shared->AddBlockMapping(Address, CodePages, HostCode, lk);
// There is no need to update L1 or L2, they will get updated on first lookup
// However, adding to L1 here increases performance
auto& L1Entry = reinterpret_cast<LookupCacheEntry*>(L1Pointer)[Address & L1PointerMask];
L1Entry.GuestCode = Address;
L1Entry.HostCode = (uintptr_t)HostCode;
CacheBlockMapping(Address, Entry, true, lk);
}
// NOTE: It's the caller's responsibility to call Erase() for all other
// GuestToHostMaps that share the same LookupCache. Otherwise, the
// L1/L2 caches will contain stale references to deallocated memory.
bool Erase(FEXCore::Core::CpuStateFrame* Frame, uint64_t Address, const LookupCacheWriteLockToken& lk) {
bool ErasedAny = Shared->Erase(Frame, Address, lk);
// Invalidates L1/L2 for a given guest block
void InvalidateCache(uint64_t Address, const LookupCacheWriteLockToken& lk) {
// Do L1
auto& L1Entry = reinterpret_cast<LookupCacheEntry*>(L1Pointer)[Address & L1PointerMask];
if (L1Entry.GuestCode == Address) {
L1Entry.GuestCode = 0;
ErasedAny = true;
// Leave L1Entry.HostCode as is, so that concurrent lookups won't read a null pointer
// This is a soft guarantee for cross thread invalidation, as atomics are not used
// and it hasn't been thoroughly tested
@@ -300,7 +319,7 @@ public:
uint64_t LocalPagePointer = Pointers[Address];
if (!LocalPagePointer) {
// Page for this code didn't even exist, nothing to do
return ErasedAny;
return;
}
// Page exists, just set the offset to zero
@@ -308,7 +327,23 @@ public:
BlockPointers[PageOffset].GuestCode = 0;
BlockPointers[PageOffset].HostCode = 0;
}
return true;
}
// Invalidates all L1/L2 entries for all guest block that intersect the given range
bool InvalidateCacheRange(uint64_t Start, uint64_t Length) {
auto lk = Shared->AcquireWriteLock();
auto lower = CachedCodePages.lower_bound(Start >> 12);
auto upper = CachedCodePages.upper_bound((Start + Length - 1) >> 12);
for (auto it = lower; it != upper; it++) {
for (const auto& Entry : it->second) {
InvalidateCache(Entry, lk);
}
}
bool ret = upper != lower;
CachedCodePages.erase(lower, upper);
return ret;
}
void AddBlockLink(uint64_t GuestDestination, FEXCore::Context::ExitFunctionLinkData* HostLink,
@@ -317,7 +352,7 @@ public:
}
void ClearCache(const LookupCacheWriteLockToken&);
void ClearL2Cache(const LookupCacheReadLockToken&);
void ClearL2Cache(const LookupCacheBaseLockToken&);
void ClearThreadLocalCaches(const LookupCacheWriteLockToken&);
uintptr_t GetL1Pointer() const {
@@ -345,13 +380,17 @@ public:
}
private:
void CacheBlockMapping(uint64_t Address, uintptr_t HostCode, const LookupCacheReadLockToken& lk) {
void CacheBlockMapping(uint64_t Address, const GuestToHostMap::BlockEntry& Entry, bool L1Only, const LookupCacheBaseLockToken& lk) {
for (const auto& CodePage : Entry.CodePages) {
CachedCodePages[CodePage >> 12].insert(Address);
}
// Do L1
auto& L1Entry = reinterpret_cast<LookupCacheEntry*>(L1Pointer)[Address & L1PointerMask];
L1Entry.GuestCode = Address;
L1Entry.HostCode = HostCode;
L1Entry.HostCode = Entry.HostCode;
if (!DisableL2Cache()) {
if (!DisableL2Cache() && !L1Only) {
// Do ful map
auto FullAddress = Address;
Address = Address & (VirtualMemSize - 1);
@@ -368,7 +407,7 @@ private:
if (!NewPageBacking) {
// Couldn't allocate, clear L2 and retry
ClearL2Cache(lk);
CacheBlockMapping(Address, HostCode, lk);
CacheBlockMapping(FullAddress, Entry, false, lk);
return;
}
Pointers[Address] = NewPageBacking;
@@ -380,7 +419,7 @@ private:
// This silently replaces existing mappings
BlockPointers[PageOffset].GuestCode = FullAddress;
BlockPointers[PageOffset].HostCode = HostCode;
BlockPointers[PageOffset].HostCode = Entry.HostCode;
}
}
@@ -398,6 +437,9 @@ private:
return PageMemory + NewBase;
}
// Maps from a page index to all blocks in the page that have at some point been fetched into L1/L2
fextl::map<uint64_t, fextl::robin_set<uint64_t>> CachedCodePages;
uintptr_t PagePointer;
uintptr_t PageMemory;
uintptr_t L1Pointer;
+162 -115
View File
@@ -514,18 +514,17 @@ void OpDispatchBuilder::CALLOp(OpcodeArgs) {
BlockSetRIP = true;
// Call instruction only uses up to 32-bit signed displacement
int64_t TargetOffset = Op->Src[0].Literal();
const int64_t TargetOffset = Op->Src[0].Literal();
auto ConstantPC = GetRelocatedPC(Op);
const auto ConstantPC = GetRelocatedPC(Op);
// Push the return address.
Push(GPRSize, ConstantPC);
const uint64_t NextRIP = Op->PC + Op->InstSize;
uint64_t TargetRIP = NextRIP + TargetOffset;
if (NextRIP != TargetRIP) {
if (TargetOffset != 0) {
// Store the RIP
const uint64_t NextRIP = Op->PC + Op->InstSize;
ExitRelocatedPC(Op, TargetOffset, BranchHint::Call, ConstantPC, [&]() {
auto CallReturnJumpTarget = JumpTargets.find(NextRIP);
if (CallReturnJumpTarget != JumpTargets.end() && CallReturnJumpTarget->second.IsEntryPoint) {
@@ -1040,6 +1039,34 @@ void OpDispatchBuilder::TESTOp(OpcodeArgs, uint32_t SrcIndex) {
InvalidateAF();
}
void OpDispatchBuilder::ARPLOp(OpcodeArgs) {
// ARPL r/m16, r16
// If the RPL field in the destination selector is less privileged than the
// RPL field in the source selector, then adjust destination RPL to match
// source RPL and set ZF=1. Otherwise ZF=0 and destination is unchanged.
//
// Only ZF is modified by ARPL.
constexpr auto Size = OpSize::i16Bit;
Ref Dest = LoadSourceGPR_WithOpSize(Op, Op->Dest, Size, Op->Flags, {.AllowUpperGarbage = true});
Ref Src = LoadSourceGPR_WithOpSize(Op, Op->Src[0], Size, Op->Flags, {.AllowUpperGarbage = true});
// RPL is the low two bits of the selector.
Ref DestRPL = _Bfe(OpSize::i32Bit, 2, 0, Dest);
Ref SrcRPL = _Bfe(OpSize::i32Bit, 2, 0, Src);
// NeedUpdate is 1 when DestRPL < SrcRPL, else 0.
Ref NeedUpdate = _Select(OpSize::i32Bit, OpSize::i32Bit, CondClass::ULT, DestRPL, SrcRPL, Constant(1), Constant(0));
SetRFLAG<FEXCore::X86State::RFLAG_ZF_RAW_LOC>(NeedUpdate);
// Compute adjusted destination selector: (Dest & ~3) | SrcRPL.
auto NewDest = _Bfxil(OpSize::i32Bit, 2, 0, Dest, SrcRPL);
// Conditionally select updated selector based on NeedUpdate.
Ref FinalDest = _Select(OpSize::i32Bit, OpSize::i32Bit, CondClass::NEQ, NeedUpdate, Constant(0), NewDest, Dest);
StoreResultGPR_WithOpSize(Op, Op->Dest, FinalDest, Size);
}
void OpDispatchBuilder::MOVSXDOp(OpcodeArgs) {
// This instruction is a bit special
// if SrcSize == 2
@@ -1326,35 +1353,12 @@ void OpDispatchBuilder::MOVSegOp(OpcodeArgs, bool ToSeg) {
}
void OpDispatchBuilder::MOVOffsetOp(OpcodeArgs) {
auto GenMemSrcFromOp = [&](size_t StartingSource) -> AddressMode {
const uint64_t Lower = Op->Src[StartingSource].Literal();
const uint64_t Upper = Op->Src[StartingSource + 1].Literal();
const uint64_t Combined = (Upper << 32) | Lower;
const auto GPRSize = GetGPROpSize();
AddressMode A {
.Segment = GetSegment(Op->Flags),
.Offset = static_cast<int64_t>(Combined),
.AddrSize = (Op->Flags & X86Tables::DecodeFlags::FLAG_ADDRESS_SIZE) != 0 ? (GPRSize >> 1) : GPRSize,
.NonTSO = false,
};
return A;
};
switch (Op->OP) {
case 0xA0:
case 0xA1: {
// Source is memory(literal)
// Dest is GPR
Ref Src {};
if (Op->Src[0].Data.Literal.Size <= 4) {
Src = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.ForceLoad = true});
} else {
const auto OpSize = OpSizeFromSrc(Op);
auto A = GenMemSrcFromOp(0);
Src = _LoadMemGPRAutoTSO(OpSize, A, OpSize::i8Bit);
}
auto Src = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.ForceLoad = true});
StoreResultGPR(Op, Op->Dest, Src);
break;
}
@@ -1366,13 +1370,7 @@ void OpDispatchBuilder::MOVOffsetOp(OpcodeArgs) {
// This one is a bit special since the destination is a literal
// So the destination gets stored in Src[1]
if (Op->Src[1].Data.Literal.Size <= 4) {
StoreResultGPR(Op, Op->Src[1], Src);
} else {
const auto OpSize = OpSizeFromSrc(Op);
auto A = GenMemSrcFromOp(1);
_StoreMemGPRAutoTSO(OpSize, A, Src, OpSize::i8Bit);
}
StoreResultGPR(Op, Op->Src[1], Src);
break;
}
}
@@ -1397,7 +1395,7 @@ void OpDispatchBuilder::CPUIDOp(OpcodeArgs) {
StoreGPRRegister(X86State::REG_RDX, RDX);
}
uint32_t OpDispatchBuilder::LoadConstantShift(X86Tables::DecodedOp Op, bool Is1Bit) {
uint32_t OpDispatchBuilder::GetConstantShift(X86Tables::DecodedOp Op, bool Is1Bit) {
if (Is1Bit) {
return 1;
} else {
@@ -1432,7 +1430,7 @@ void OpDispatchBuilder::SHLOp(OpcodeArgs) {
void OpDispatchBuilder::SHLImmediateOp(OpcodeArgs, bool SHL1Bit) {
Ref Dest = LoadSourceGPR(Op, Op->Dest, Op->Flags, {.AllowUpperGarbage = true});
uint64_t Shift = LoadConstantShift(Op, SHL1Bit);
uint64_t Shift = GetConstantShift(Op, SHL1Bit);
const auto Size = GetSrcBitSize(Op);
Ref Result = _Lshl(Size == 64 ? OpSize::i64Bit : OpSize::i32Bit, Dest, Constant(Shift));
@@ -1455,7 +1453,7 @@ void OpDispatchBuilder::SHRImmediateOp(OpcodeArgs, bool SHR1Bit) {
const auto Size = GetSrcBitSize(Op);
auto Dest = LoadSourceGPR(Op, Op->Dest, Op->Flags, {.AllowUpperGarbage = Size >= 32});
uint64_t Shift = LoadConstantShift(Op, SHR1Bit);
uint64_t Shift = GetConstantShift(Op, SHR1Bit);
auto ALUOp = _Lshr(Size == 64 ? OpSize::i64Bit : OpSize::i32Bit, Dest, Constant(Shift));
CalculateFlags_ShiftRightImmediate(OpSizeFromSrc(Op), ALUOp, Dest, Shift);
@@ -1507,7 +1505,7 @@ void OpDispatchBuilder::SHLDOp(OpcodeArgs) {
}
void OpDispatchBuilder::SHLDImmediateOp(OpcodeArgs) {
uint64_t Shift = LoadConstantShift(Op, false);
uint64_t Shift = GetConstantShift(Op, false);
const auto Size = GetSrcBitSize(Op);
Ref Src = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.AllowUpperGarbage = Size >= 32});
@@ -1575,7 +1573,7 @@ void OpDispatchBuilder::SHRDImmediateOp(OpcodeArgs) {
Ref Src = LoadSourceGPR(Op, Op->Src[0], Op->Flags);
Ref Dest = LoadSourceGPR(Op, Op->Dest, Op->Flags);
uint64_t Shift = LoadConstantShift(Op, false);
uint64_t Shift = GetConstantShift(Op, false);
const auto Size = GetSrcBitSize(Op);
if (Shift != 0) {
@@ -1616,7 +1614,7 @@ void OpDispatchBuilder::ASHROp(OpcodeArgs, bool Immediate, bool SHR1Bit) {
}
if (Immediate) {
uint64_t Shift = LoadConstantShift(Op, SHR1Bit);
uint64_t Shift = GetConstantShift(Op, SHR1Bit);
Ref Result = _Ashr(OpSize, Dest, Constant(Shift));
CalculateFlags_SignShiftRightImmediate(OpSizeFromSrc(Op), Result, Dest, Shift);
@@ -1644,7 +1642,7 @@ void OpDispatchBuilder::RotateOp(OpcodeArgs, bool Left, bool IsImmediate, bool I
ArithRef UnmaskedSrc;
if (Is1Bit || IsImmediate) {
UnmaskedConst = LoadConstantShift(Op, Is1Bit);
UnmaskedConst = GetConstantShift(Op, Is1Bit);
UnmaskedSrc = ARef(UnmaskedConst);
} else {
UnmaskedSrc = ARef(LoadSourceGPR(Op, Op->Src[1], Op->Flags, {.AllowUpperGarbage = true}));
@@ -2645,7 +2643,10 @@ void OpDispatchBuilder::IMULOp(OpcodeArgs) {
}
// 64-bit special cased to save a move
Ref Result = Size < OpSize::i64Bit ? _Mul(OpSize::i64Bit, Src1, Src2) : nullptr;
Ref Result {};
if (Size < OpSize::i64Bit) {
Result = _Mul(OpSize::i64Bit, Src1, Src2);
}
Ref ResultHigh {};
if (Size == OpSize::i8Bit) {
// Result is stored in AX
@@ -2750,7 +2751,8 @@ void OpDispatchBuilder::NOTOp(OpcodeArgs) {
if (DestIsLockedMem(Op)) {
HandledLock = true;
Ref DestMem = MakeSegmentAddress(Op, Op->Dest);
_AtomicXor(Size, MaskConst, DestMem);
// Result unused
_AtomicFetchXor(Size, MaskConst, DestMem);
} else if (!Op->Dest.IsGPR()) {
// GPR version plays fast and loose with sizes, be safe for memory tho.
Ref Src = LoadSourceGPR(Op, Op->Dest, Op->Flags);
@@ -3067,7 +3069,6 @@ void OpDispatchBuilder::SMSWOp(OpcodeArgs) {
(0U << 2) | ///< EM - Emulation
(1U << 1) | ///< MP - Monitor Coprocessor
(1U << 0)); ///< PE - Protection Enabled
const auto OpAddr = X86Tables::DecodeFlags::GetOpAddr(Op->Flags, 0);
if (Is64BitMode) {
DstSize = OpAddr == X86Tables::DecodeFlags::FLAG_OPERAND_SIZE_LAST ? OpSize::i16Bit :
@@ -3199,7 +3200,7 @@ void OpDispatchBuilder::DECOp(OpcodeArgs) {
void OpDispatchBuilder::STOSOp(OpcodeArgs) {
if (Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_ADDRESS_SIZE) {
LogMan::Msg::EFmt("Can't handle adddress size");
LogMan::Msg::EFmt("STOSOp: Can't handle address size override (OP: 0x{:04X}, Flags: 0x{:08X})", Op->OP, Op->Flags);
DecodeFailure = true;
return;
}
@@ -3244,7 +3245,7 @@ void OpDispatchBuilder::STOSOp(OpcodeArgs) {
void OpDispatchBuilder::MOVSOp(OpcodeArgs) {
if (Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_ADDRESS_SIZE) {
LogMan::Msg::EFmt("Can't handle adddress size");
LogMan::Msg::EFmt("MOVSOp: Can't handle address size override (OP: 0x{:04X}, Flags: 0x{:08X})", Op->OP, Op->Flags);
DecodeFailure = true;
return;
}
@@ -3297,45 +3298,57 @@ void OpDispatchBuilder::MOVSOp(OpcodeArgs) {
_StoreMem(RegClass::GPR, Size, Src, RDI, Invalid(), OpSize::i8Bit, MemOffsetType::SXTX, 1);
}
auto PtrDir = LoadDir(IR::OpSizeToSize(Size));
RSI = Add(OpSize::i64Bit, RSI, PtrDir);
RDI = Add(OpSize::i64Bit, RDI, PtrDir);
RSI = OffsetByDir(RSI, IR::OpSizeToSize(Size));
RDI = OffsetByDir(RDI, IR::OpSizeToSize(Size));
StoreGPRRegister(X86State::REG_RSI, RSI);
StoreGPRRegister(X86State::REG_RDI, RDI);
}
}
IR::OpSize OpDispatchBuilder::GetStringOpSize(X86Tables::DecodedOp Op) const {
LOGMAN_THROW_A_FMT(Is64BitMode || !(Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_ADDRESS_SIZE), "Invalid modifier on 32bit address");
return !Is64BitMode || (Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_ADDRESS_SIZE) ? OpSize::i32Bit : OpSize::i64Bit;
}
void OpDispatchBuilder::CMPSOp(OpcodeArgs) {
if (Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_ADDRESS_SIZE) {
LogMan::Msg::EFmt("Can't handle adddress size");
if (!Is64BitMode && (Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_ADDRESS_SIZE)) {
LogMan::Msg::EFmt("CMPSOp: Address size override (0x67) not supported in 32-bit mode (OP: 0x{:04X}).", Op->OP);
DecodeFailure = true;
return;
}
const auto Size = OpSizeFromSrc(Op);
OpSize AddrSize = GetStringOpSize(Op);
bool Repeat = Op->Flags & (FEXCore::X86Tables::DecodeFlags::FLAG_REPNE_PREFIX | FEXCore::X86Tables::DecodeFlags::FLAG_REP_PREFIX);
if (!Repeat) {
// Default DS prefix
Ref Dest_RSI = MakeSegmentAddress(X86State::REG_RSI, Op->Flags, X86Tables::DecodeFlags::FLAG_DS_PREFIX);
// Only ES prefix
Ref Dest_RDI = MakeSegmentAddress(X86State::REG_RDI, 0, X86Tables::DecodeFlags::FLAG_ES_PREFIX, true);
Ref Src_RSI = LoadGPRRegister(X86State::REG_RSI, AddrSize);
Ref Src_RDI = LoadGPRRegister(X86State::REG_RDI, AddrSize);
Ref Dest_RSI = AppendSegmentOffset(Src_RSI, Op->Flags, X86Tables::DecodeFlags::FLAG_DS_PREFIX);
Ref Dest_RDI = AppendSegmentOffset(Src_RDI, 0, X86Tables::DecodeFlags::FLAG_ES_PREFIX, true);
auto Src1 = _LoadMemGPRAutoTSO(Size, Dest_RDI, Size);
auto Src2 = _LoadMemGPRAutoTSO(Size, Dest_RSI, Size);
CalculateFlags_SUB(OpSizeFromSrc(Op), Src2, Src1);
auto PtrDir = LoadDir(IR::OpSizeToSize(Size));
Dest_RDI = OffsetByDir(Src_RDI, IR::OpSizeToSize(Size));
if (Is64BitMode && AddrSize == OpSize::i32Bit) {
Dest_RDI = _Bfe(OpSize::i64Bit, 32, 0, Dest_RDI);
StoreGPRRegister(X86State::REG_RDI, Dest_RDI);
} else {
StoreGPRRegister(X86State::REG_RDI, Dest_RDI, AddrSize);
}
// Offset the pointer
Dest_RDI = Add(OpSize::i64Bit, Dest_RDI, PtrDir);
StoreGPRRegister(X86State::REG_RDI, Dest_RDI);
// Offset second pointer
Dest_RSI = Add(OpSize::i64Bit, Dest_RSI, PtrDir);
StoreGPRRegister(X86State::REG_RSI, Dest_RSI);
Dest_RSI = OffsetByDir(Src_RSI, IR::OpSizeToSize(Size));
if (Is64BitMode && AddrSize == OpSize::i32Bit) {
Dest_RSI = _Bfe(OpSize::i64Bit, 32, 0, Dest_RSI);
StoreGPRRegister(X86State::REG_RSI, Dest_RSI);
} else {
StoreGPRRegister(X86State::REG_RSI, Dest_RSI, AddrSize);
}
} else {
// Calculate flags early.
CalculateDeferredFlags();
@@ -3351,7 +3364,7 @@ void OpDispatchBuilder::CMPSOp(OpcodeArgs) {
SetCurrentCodeBlock(BeforeLoop);
StartNewBlock();
ForeachDirection([this, Op, Size, REPE](int32_t PtrDir) {
ForeachDirection([this, Op, Size, AddrSize, REPE](int32_t PtrDir) {
IRPair<IROp_CondJump> InnerJump;
auto JumpIntoLoop = Jump();
@@ -3363,10 +3376,11 @@ void OpDispatchBuilder::CMPSOp(OpcodeArgs) {
// Working loop
{
// Default DS prefix
Ref Dest_RSI = MakeSegmentAddress(X86State::REG_RSI, Op->Flags, X86Tables::DecodeFlags::FLAG_DS_PREFIX);
// Only ES prefix
Ref Dest_RDI = MakeSegmentAddress(X86State::REG_RDI, 0, X86Tables::DecodeFlags::FLAG_ES_PREFIX, true);
Ref Src_RSI = LoadGPRRegister(X86State::REG_RSI, AddrSize);
Ref Src_RDI = LoadGPRRegister(X86State::REG_RDI, AddrSize);
Ref Dest_RSI = AppendSegmentOffset(Src_RSI, Op->Flags, X86Tables::DecodeFlags::FLAG_DS_PREFIX);
Ref Dest_RDI = AppendSegmentOffset(Src_RDI, 0, X86Tables::DecodeFlags::FLAG_ES_PREFIX, true);
auto Src1 = _LoadMemGPRAutoTSO(Size, Dest_RDI, Size);
auto Src2 = _LoadMemGPR(Size, Dest_RSI, Size);
@@ -3383,13 +3397,21 @@ void OpDispatchBuilder::CMPSOp(OpcodeArgs) {
// Store the counter since we don't have phis
StoreGPRRegister(X86State::REG_RCX, TailCounter);
// Offset the pointer
Dest_RDI = Add(OpSize::i64Bit, Dest_RDI, PtrDir * static_cast<int32_t>(IR::OpSizeToSize(Size)));
StoreGPRRegister(X86State::REG_RDI, Dest_RDI);
Dest_RDI = Add(AddrSize, Src_RDI, PtrDir * static_cast<int32_t>(IR::OpSizeToSize(Size)));
if (Is64BitMode && AddrSize == OpSize::i32Bit) {
Dest_RDI = _Bfe(OpSize::i64Bit, 32, 0, Dest_RDI);
StoreGPRRegister(X86State::REG_RDI, Dest_RDI);
} else {
StoreGPRRegister(X86State::REG_RDI, Dest_RDI, AddrSize);
}
// Offset second pointer
Dest_RSI = Add(OpSize::i64Bit, Dest_RSI, PtrDir * static_cast<int32_t>(IR::OpSizeToSize(Size)));
StoreGPRRegister(X86State::REG_RSI, Dest_RSI);
Dest_RSI = Add(AddrSize, Src_RSI, PtrDir * static_cast<int32_t>(IR::OpSizeToSize(Size)));
if (Is64BitMode && AddrSize == OpSize::i32Bit) {
Dest_RSI = _Bfe(OpSize::i64Bit, 32, 0, Dest_RSI);
StoreGPRRegister(X86State::REG_RSI, Dest_RSI);
} else {
StoreGPRRegister(X86State::REG_RSI, Dest_RSI, AddrSize);
}
// If TailCounter != 0, compare sources.
// If TailCounter == 0, set ZF iff that would break.
@@ -3428,7 +3450,7 @@ void OpDispatchBuilder::CMPSOp(OpcodeArgs) {
void OpDispatchBuilder::LODSOp(OpcodeArgs) {
if (Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_ADDRESS_SIZE) {
LogMan::Msg::EFmt("Can't handle adddress size");
LogMan::Msg::EFmt("LODSOp: Can't handle address size override (OP: 0x{:04X}, Flags: 0x{:08X})", Op->OP, Op->Flags);
DecodeFailure = true;
return;
}
@@ -3510,31 +3532,37 @@ void OpDispatchBuilder::LODSOp(OpcodeArgs) {
}
void OpDispatchBuilder::SCASOp(OpcodeArgs) {
if (Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_ADDRESS_SIZE) {
LogMan::Msg::EFmt("Can't handle adddress size");
if (!Is64BitMode && (Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_ADDRESS_SIZE)) {
LogMan::Msg::EFmt("SCASOp: Address size override (0x67) not supported in 32-bit mode (OP: 0x{:04X}).", Op->OP);
DecodeFailure = true;
return;
}
const auto Size = OpSizeFromSrc(Op);
OpSize AddrSize = GetStringOpSize(Op);
const bool Repeat = (Op->Flags & (FEXCore::X86Tables::DecodeFlags::FLAG_REPNE_PREFIX | FEXCore::X86Tables::DecodeFlags::FLAG_REP_PREFIX)) != 0;
if (!Repeat) {
Ref Dest_RDI = MakeSegmentAddress(X86State::REG_RDI, 0, X86Tables::DecodeFlags::FLAG_ES_PREFIX, true);
Ref Src_RDI = LoadGPRRegister(X86State::REG_RDI, AddrSize);
Ref Dest_RDI = AppendSegmentOffset(Src_RDI, 0, X86Tables::DecodeFlags::FLAG_ES_PREFIX, true);
auto Src1 = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.AllowUpperGarbage = true});
auto Src2 = _LoadMemGPRAutoTSO(Size, Dest_RDI, Size);
CalculateFlags_SUB(OpSizeFromSrc(Op), Src1, Src2);
// Offset the pointer
Ref TailDest_RDI = LoadGPRRegister(X86State::REG_RDI);
StoreGPRRegister(X86State::REG_RDI, OffsetByDir(TailDest_RDI, IR::OpSizeToSize(Size)));
Ref TailDest_RDI = OffsetByDir(Src_RDI, IR::OpSizeToSize(Size));
if (Is64BitMode && AddrSize == OpSize::i32Bit) {
TailDest_RDI = _Bfe(OpSize::i64Bit, 32, 0, TailDest_RDI);
StoreGPRRegister(X86State::REG_RDI, TailDest_RDI);
} else {
StoreGPRRegister(X86State::REG_RDI, TailDest_RDI, AddrSize);
}
} else {
// Calculate flags early. because end of block
CalculateDeferredFlags();
ForeachDirection([this, Op, Size](int32_t Dir) {
ForeachDirection([this, Op, Size, AddrSize](int32_t Dir) {
bool REPE = Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_REP_PREFIX;
auto JumpStart = Jump();
@@ -3558,7 +3586,8 @@ void OpDispatchBuilder::SCASOp(OpcodeArgs) {
// Working loop
{
Ref Dest_RDI = MakeSegmentAddress(X86State::REG_RDI, 0, X86Tables::DecodeFlags::FLAG_ES_PREFIX, true);
Ref Src_RDI = LoadGPRRegister(X86State::REG_RDI, AddrSize);
Ref Dest_RDI = AppendSegmentOffset(Src_RDI, 0, X86Tables::DecodeFlags::FLAG_ES_PREFIX, true);
auto Src1 = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.AllowUpperGarbage = true});
auto Src2 = _LoadMemGPRAutoTSO(Size, Dest_RDI, Size);
@@ -3569,7 +3598,7 @@ void OpDispatchBuilder::SCASOp(OpcodeArgs) {
CalculateDeferredFlags();
Ref TailCounter = LoadGPRRegister(X86State::REG_RCX);
Ref TailDest_RDI = LoadGPRRegister(X86State::REG_RDI);
Ref Src_RDI_Tail = LoadGPRRegister(X86State::REG_RDI, AddrSize);
// Decrement counter
TailCounter = Sub(OpSize::i64Bit, TailCounter, 1);
@@ -3577,9 +3606,13 @@ void OpDispatchBuilder::SCASOp(OpcodeArgs) {
// Store the counter since we don't have phis
StoreGPRRegister(X86State::REG_RCX, TailCounter);
// Offset the pointer
TailDest_RDI = Add(OpSize::i64Bit, TailDest_RDI, Dir * static_cast<int32_t>(IR::OpSizeToSize(Size)));
StoreGPRRegister(X86State::REG_RDI, TailDest_RDI);
Ref TailDest_RDI = Add(AddrSize, Src_RDI_Tail, Dir * static_cast<int32_t>(IR::OpSizeToSize(Size)));
if (Is64BitMode && AddrSize == OpSize::i32Bit) {
TailDest_RDI = _Bfe(OpSize::i64Bit, 32, 0, TailDest_RDI);
StoreGPRRegister(X86State::REG_RDI, TailDest_RDI);
} else {
StoreGPRRegister(X86State::REG_RDI, TailDest_RDI, AddrSize);
}
CalculateDeferredFlags();
InternalCondJump = CondJumpNZCV(REPE ? CondClass::EQ : CondClass::NEQ);
@@ -4099,6 +4132,8 @@ void OpDispatchBuilder::CheckLegacySegmentRead(Ref NewNode, uint32_t SegmentReg)
// Will set the telemetry value if NewNode is != 0
_TelemetrySetValue(NewNode, TelemIndex);
// Telemetry will dirty flags, and user code does not expect LoadSource to clobber flags, fix that up here as this is an edge case.
CalculateDeferredFlags();
#endif
}
@@ -4137,6 +4172,8 @@ void OpDispatchBuilder::CheckLegacySegmentWrite(Ref NewNode, uint32_t SegmentReg
// Will set the telemetry value if NewNode is != 0
_TelemetrySetValue(NewNode, TelemIndex);
// Telemetry will dirty flags, and user code does not expect LoadSource to clobber flags, fix that up here as this is an edge case.
CalculateDeferredFlags();
#endif
}
@@ -4202,18 +4239,26 @@ AddressMode OpDispatchBuilder::DecodeAddress(const X86Tables::DecodedOp& Op, con
} else if (Operand.IsGPRDirect()) {
A.Base = LoadGPRRegister(Operand.Data.GPR.GPR, GPRSize);
A.NonTSO |= IsNonTSOReg(AccessType, Operand.Data.GPR.GPR);
} else if (Operand.IsGPRIndirect()) {
} else if (Operand.IsGPRIndirect() || Operand.IsGPRIndirectRelocation()) {
A.Base = LoadGPRRegister(Operand.Data.GPRIndirect.GPR, GPRSize);
A.Offset = Operand.Data.GPRIndirect.Displacement;
if (Operand.IsGPRIndirectRelocation()) {
A.Base = Add(GPRSize, _EntrypointOffset(GPRSize, Operand.Data.GPRIndirect.Displacement), A.Base);
} else {
A.Offset = static_cast<int32_t>(Operand.Data.GPRIndirect.Displacement);
}
A.NonTSO |= IsNonTSOReg(AccessType, Operand.Data.GPRIndirect.GPR);
} else if (Operand.IsRIPRelative()) {
} else if (Operand.IsRIPRelative() || Operand.IsRIPRelativeRelocation()) {
if (Is64BitMode) {
A.Base = GetRelocatedPC(Op, Operand.Data.RIPLiteral.Value.s);
A.Base = GetRelocatedPC(Op, static_cast<int32_t>(Operand.Data.RIPLiteral.Value));
} else {
// 32bit this isn't RIP relative but instead absolute
A.Offset = Operand.Data.RIPLiteral.Value.u;
if (Operand.IsRIPRelativeRelocation()) {
A.Base = _EntrypointOffset(GPRSize, Operand.Data.RIPLiteral.Value);
} else {
A.Offset = Operand.Data.RIPLiteral.Value;
}
}
} else if (Operand.IsSIB()) {
} else if (Operand.IsSIB() || Operand.IsSIBRelocation()) {
const bool IsVSIB = IsLoad && ((Op->Flags & X86Tables::DecodeFlags::FLAG_VSIB_BYTE) != 0);
if (Operand.Data.SIB.Base != FEXCore::X86State::REG_INVALID) {
@@ -4234,8 +4279,20 @@ AddressMode OpDispatchBuilder::DecodeAddress(const X86Tables::DecodedOp& Op, con
A.IndexScale = Operand.Data.SIB.Scale;
}
A.Offset = Operand.Data.SIB.Offset;
if (Operand.IsSIBRelocation()) {
auto EPOffset = _EntrypointOffset(GPRSize, Operand.Data.SIB.Offset);
if (A.Base) {
A.Base = Add(GPRSize, EPOffset, A.Base);
} else {
A.Base = EPOffset;
}
} else {
A.Offset = static_cast<int32_t>(Operand.Data.SIB.Offset);
}
A.NonTSO |= IsNonTSOReg(AccessType, Operand.Data.SIB.Base) || IsNonTSOReg(AccessType, Operand.Data.SIB.Index);
} else if (Operand.IsLiteralRelocation()) {
A.Base = _EntrypointOffset(GPRSize, Operand.Data.LiteralRelocation.EntrypointOffset);
} else {
LOGMAN_MSG_A_FMT("Unknown Src Type: {}\n", Operand.Type);
}
@@ -4433,8 +4490,6 @@ void OpDispatchBuilder::StoreResult(RegClass Class, X86Tables::DecodedOp Op, Ref
OpDispatchBuilder::OpDispatchBuilder(FEXCore::Context::ContextImpl* ctx)
: IREmitter {ctx->OpDispatcherAllocator, ctx->HostFeatures.SupportsTSOImm9}
, CTX {ctx} {
ResetWorkingList();
if (CTX->HostFeatures.SupportsAVX && CTX->HostFeatures.SupportsSVE256) {
SaveAVXStateFunc = &OpDispatchBuilder::SaveAVXState;
RestoreAVXStateFunc = &OpDispatchBuilder::RestoreAVXState;
@@ -4447,7 +4502,8 @@ OpDispatchBuilder::OpDispatchBuilder(FEXCore::Context::ContextImpl* ctx)
}
void OpDispatchBuilder::ResetWorkingList() {
IREmitter::ResetWorkingList();
IREmitter::ReownOrClaimBuffer();
JumpTargets.clear();
BlockSetRIP = false;
DecodeFailure = false;
@@ -4468,20 +4524,6 @@ void OpDispatchBuilder::MOVGPROp(OpcodeArgs, uint32_t SrcIndex) {
StoreResultGPR(Op, Src, OpSize::i8Bit);
}
void OpDispatchBuilder::MOVGPRImmediate(OpcodeArgs) {
Ref Src {};
if (Op->Src[0].Data.Literal.Size <= 4) {
Src = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.Align = OpSize::i8Bit, .AllowUpperGarbage = true});
} else {
// 8-byte literal is special cased.
const uint64_t Lower = Op->Src[0].Literal();
const uint64_t Upper = Op->Src[1].Literal();
const uint64_t Combined = (Upper << 32) | Lower;
Src = _Constant(Combined);
}
StoreResultGPR(Op, Src, OpSize::i8Bit);
}
void OpDispatchBuilder::MOVGPRNTOp(OpcodeArgs) {
Ref Src = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.Align = OpSize::i8Bit});
StoreResultGPR(Op, Src, OpSize::i8Bit, MemoryAccessType::STREAM);
@@ -4626,7 +4668,7 @@ void OpDispatchBuilder::INTOp(OpcodeArgs) {
}
#endif
#ifdef _M_ARM_64EC
#ifdef ARCHITECTURE_arm64ec
// This is used when QueryPerformanceCounter is called on recent Windows versions, it causes CNTVCT to be written into RAX.
constexpr uint8_t GET_CNTVCT_LITERAL = 0x81;
if (Literal == GET_CNTVCT_LITERAL) {
@@ -4847,6 +4889,11 @@ void OpDispatchBuilder::CLZeroOp(OpcodeArgs) {
}
void OpDispatchBuilder::Prefetch(OpcodeArgs, bool ForStore, bool Stream, uint8_t Level) {
if (Op->Src[0].IsGPR()) {
// NOP instance.
return;
}
Ref DestMem = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.LoadData = false});
_Prefetch(ForStore, Stream, Level, DestMem, Invalid(), MemOffsetType::SXTX, 1);
}
@@ -27,6 +27,42 @@
#include <xxhash.h>
namespace FEXCore::IR {
enum class VectorCompareType {
// SSE comparisons.
EQ_OQ = 0,
LT_OS = 1,
LE_OS = 2,
UNORD_Q = 3,
NEQ_UQ = 4,
NLT_US = 5,
NLE_US = 6,
ORD_Q = 7,
// AVX-only comparisons.
EQ_UQ = 8,
NGE_US = 9,
NGT_US = 10,
FALSE_OQ = 11,
NEQ_OQ = 12,
GE_OS = 13,
GT_OS = 14,
TRUE_UQ = 15,
EQ_OS = 16,
LT_OQ = 17,
LE_OQ = 18,
UNORD_S = 19,
NEQ_US = 20,
NLT_UQ = 21,
NLE_UQ = 22,
ORD_S = 23,
EQ_US = 24,
NGE_UQ = 25,
NGT_UQ = 26,
FALSE_OS = 27,
NEQ_OS = 28,
GE_OQ = 29,
GT_OQ = 30,
TRUE_US = 31,
};
enum class MemoryAccessType {
// Choose TSO or Non-TSO depending on access type
@@ -269,7 +305,9 @@ public:
OpDispatchBuilder(FEXCore::Context::ContextImpl* ctx);
// Should only be called at the start of IR Emission.
void ResetWorkingList();
void ResetDecodeFailure() {
NeedsBlockEnd = DecodeFailure = false;
}
@@ -319,7 +357,6 @@ public:
void UnhandledOp(OpcodeArgs);
void MOVGPROp(OpcodeArgs, uint32_t SrcIndex);
void MOVGPRImmediate(OpcodeArgs);
void MOVGPRNTOp(OpcodeArgs);
void MOVVectorAlignedOp(OpcodeArgs);
void MOVVectorUnalignedOp(OpcodeArgs);
@@ -357,6 +394,7 @@ public:
void CALLFARIndirectOp(OpcodeArgs);
void RETFARIndirectOp(OpcodeArgs);
void TESTOp(OpcodeArgs, uint32_t SrcIndex);
void ARPLOp(OpcodeArgs);
void MOVSXDOp(OpcodeArgs);
void MOVSXOp(OpcodeArgs);
void MOVZXOp(OpcodeArgs);
@@ -373,7 +411,7 @@ public:
void CMOVOp(OpcodeArgs);
void CPUIDOp(OpcodeArgs);
void XGetBVOp(OpcodeArgs);
uint32_t LoadConstantShift(X86Tables::DecodedOp Op, bool Is1Bit);
uint32_t GetConstantShift(X86Tables::DecodedOp Op, bool Is1Bit);
void SHLOp(OpcodeArgs);
void SHLImmediateOp(OpcodeArgs, bool SHL1Bit);
void SHROp(OpcodeArgs);
@@ -1544,7 +1582,7 @@ private:
[[nodiscard]]
static bool IsOperandMem(const X86Tables::DecodedOperand& Operand, bool Load) {
// Literals are immediates as sources but memory addresses as destinations.
return !(Load && Operand.IsLiteral()) && !Operand.IsGPR();
return !(Load && (Operand.IsLiteral() || Operand.IsLiteralRelocation())) && !Operand.IsGPR();
}
[[nodiscard]]
@@ -1630,7 +1668,7 @@ private:
[[nodiscard]]
static uint32_t GPROffset(X86State::X86Reg reg) {
LOGMAN_THROW_A_FMT(reg <= X86State::X86Reg::REG_R15, "Invalid reg used");
return static_cast<uint32_t>(offsetof(Core::CPUState, gregs[static_cast<size_t>(reg)]));
return static_cast<uint32_t>(ARRAY_OFFSETOF(Core::CPUState, gregs, reg));
}
[[nodiscard]]
@@ -1655,6 +1693,9 @@ private:
return IR::SizeToOpSize(GetSrcSize(Op));
}
[[nodiscard]]
IR::OpSize GetStringOpSize(X86Tables::DecodedOp Op) const;
// Set flag tracking to prepare for an operation that directly writes NZCV.
void HandleNZCVWrite() {
CachedNZCV = nullptr;
@@ -1846,15 +1887,15 @@ private:
// For DF, we need to transform 0/1 into 1/-1
StoreDF(_SubShift(OpSize::i64Bit, Constant(1), Value, ShiftType::LSL, 1));
} else if (BitOffset == FEXCore::X86State::RFLAG_TF_RAW_LOC) {
auto PackedTF = _LoadContextGPR(OpSize::i8Bit, offsetof(FEXCore::Core::CPUState, flags[BitOffset]));
auto PackedTF = _LoadContextGPR(OpSize::i8Bit, ARRAY_OFFSETOF(FEXCore::Core::CPUState, flags, BitOffset));
// An exception should still be raised after an instruction that unsets TF, leave the unblocked bit set but unset
// the TF bit to cause such behaviour. The handling code at the start of the next block will then unset the
// unblocked bit before raising the exception.
auto NewPackedTF =
_Select(OpSize::i64Bit, OpSize::i64Bit, CondClass::EQ, Value, Constant(0), _And(OpSize::i32Bit, PackedTF, Constant(~1)), Constant(1));
_StoreContextGPR(OpSize::i8Bit, NewPackedTF, offsetof(FEXCore::Core::CPUState, flags[BitOffset]));
_StoreContextGPR(OpSize::i8Bit, NewPackedTF, ARRAY_OFFSETOF(FEXCore::Core::CPUState, flags, BitOffset));
} else {
_StoreContextGPR(OpSize::i8Bit, Value, offsetof(FEXCore::Core::CPUState, flags[BitOffset]));
_StoreContextGPR(OpSize::i8Bit, Value, ARRAY_OFFSETOF(FEXCore::Core::CPUState, flags, BitOffset));
}
}
@@ -1909,8 +1950,8 @@ private:
[[nodiscard]]
static uint32_t CacheIndexToContextOffset(int Index) {
switch (Index) {
case MM0Index ... MM7Index: return offsetof(FEXCore::Core::CPUState, mm[Index - MM0Index]);
case AVXHigh0Index ... AVXHigh15Index: return offsetof(FEXCore::Core::CPUState, avx_high[Index - AVXHigh0Index][0]);
case MM0Index ... MM7Index: return ARRAY_OFFSETOF(FEXCore::Core::CPUState, mm, Index - MM0Index);
case AVXHigh0Index ... AVXHigh15Index: return ARRAY_OFFSETOF(FEXCore::Core::CPUState, avx_high, Index - AVXHigh0Index);
default: return ~0U;
}
}
@@ -2067,6 +2108,12 @@ private:
RegCache.Written |= Bit;
}
void InvalidateHighAVXRegisters() {
for (size_t i = 0; i < 16; ++i) {
InvalidateReg(AVXHigh0Index + i);
}
}
void StoreRegister(uint8_t Reg, bool FPR, Ref Value) {
StoreContext(Reg + (FPR ? FPR0Index : GPR0Index), Value);
}
@@ -2104,7 +2151,7 @@ private:
// Recover the sign bit, it is the logical DF value
return _Lshr(OpSize::i64Bit, LoadDF(), Constant(63));
} else {
return _LoadContextGPR(OpSize::i8Bit, offsetof(Core::CPUState, flags[BitOffset]));
return _LoadContextGPR(OpSize::i8Bit, ARRAY_OFFSETOF(Core::CPUState, flags, BitOffset));
}
}
@@ -52,7 +52,8 @@ OpDispatchBuilder::RefPair OpDispatchBuilder::AVX128_LoadSource_WithOpSize(
OpDispatchBuilder::RefVSIB
OpDispatchBuilder::AVX128_LoadVSIB(const X86Tables::DecodedOp& Op, const X86Tables::DecodedOperand& Operand, uint32_t Flags, bool NeedsHigh) {
const bool IsVSIB = (Op->Flags & X86Tables::DecodeFlags::FLAG_VSIB_BYTE) != 0;
LOGMAN_THROW_A_FMT(Operand.IsSIB() && IsVSIB, "Trying to load VSIB for something that isn't the correct type!");
LOGMAN_THROW_A_FMT((Operand.IsSIB() || Operand.IsSIBRelocation()) && IsVSIB, "Trying to load VSIB for something that isn't the correct "
"type!");
// VSIB is a very special case which has a ton of encoded data.
// Get it in a format we can reason about.
@@ -64,13 +65,25 @@ OpDispatchBuilder::AVX128_LoadVSIB(const X86Tables::DecodedOp& Op, const X86Tabl
"Base must be a GPR.");
const auto Index_XMM_gpr = Index_gpr - X86State::REG_XMM_0;
return {
OpDispatchBuilder::RefVSIB A {
.Low = AVX128_LoadXMMRegister(Index_XMM_gpr, false),
.High = NeedsHigh ? AVX128_LoadXMMRegister(Index_XMM_gpr, true) : Invalid(),
.BaseAddr = Base_gpr != FEXCore::X86State::REG_INVALID ? LoadGPRRegister(Base_gpr, OpSize::i64Bit, 0, false) : nullptr,
.Displacement = Operand.Data.SIB.Offset,
.Scale = Operand.Data.SIB.Scale,
};
if (Operand.IsSIBRelocation()) {
auto EPOffset = _EntrypointOffset(OpSize::i64Bit, Operand.Data.SIB.Offset);
if (A.BaseAddr) {
A.BaseAddr = Add(OpSize::i64Bit, EPOffset, A.BaseAddr);
} else {
A.BaseAddr = EPOffset;
}
} else {
A.Displacement = static_cast<int32_t>(Operand.Data.SIB.Offset);
}
return A;
}
void OpDispatchBuilder::AVX128_StoreResult_WithOpSize(FEXCore::X86Tables::DecodedOp Op, const FEXCore::X86Tables::DecodedOperand& Operand,
@@ -327,16 +340,12 @@ void OpDispatchBuilder::AVX128_VZERO(OpcodeArgs) {
AVX128_StoreXMMRegister(i, ZeroVector, false);
}
// More efficient for non-SRA upper-halves to use a cached constant and store directly.
for (uint32_t i = 0; i < NumRegs; i++) {
AVX128_StoreXMMRegister(i, ZeroVector, true);
}
InvalidateHighAVXRegisters();
_ContextClear(offsetof(FEXCore::Core::CPUState, avx_high), sizeof(FEXCore::Core::CPUState::avx_high[0]) * NumRegs);
} else {
// Likewise, VZEROUPPER will only ever zero only up to the first 16 registers
const auto ZeroVector = LoadZeroVector(OpSize::i128Bit);
for (uint32_t i = 0; i < NumRegs; i++) {
AVX128_StoreXMMRegister(i, ZeroVector, true);
}
InvalidateHighAVXRegisters();
_ContextClear(offsetof(FEXCore::Core::CPUState, avx_high), sizeof(FEXCore::Core::CPUState::avx_high[0]) * NumRegs);
}
}
@@ -663,10 +672,10 @@ void OpDispatchBuilder::AVX128_VFCMP(OpcodeArgs, IR::OpSize ElementSize) {
struct {
FEXCore::X86Tables::DecodedOp Op;
uint8_t CompType {};
uint32_t CompType {};
} Capture {
.Op = Op,
.CompType = CompType,
.CompType = CompType & 0b11111u,
};
AVX128_VectorBinaryImpl(Op, OpSizeFromSrc(Op), ElementSize, [this, &Capture](IR::OpSize _ElementSize, Ref Src1, Ref Src2) {
@@ -692,7 +701,7 @@ void OpDispatchBuilder::AVX128_InsertScalarFCMP(OpcodeArgs, IR::OpSize ElementSi
const uint8_t CompType = Op->Src[2].Literal();
RefPair Result {};
Result.Low = InsertScalarFCMPOpImpl(OpSize::i128Bit, OpSize::i128Bit, ElementSize, Src1.Low, Src2.Low, CompType, false);
Result.Low = InsertScalarFCMPOpImpl(OpSize::i128Bit, OpSize::i128Bit, ElementSize, Src1.Low, Src2.Low, CompType & 0b11111, false);
Result.High = LoadZeroVector(OpSize::i128Bit);
AVX128_StoreResult_WithOpSize(Op, Op->Dest, Result);
}
@@ -53,7 +53,7 @@ constexpr inline DispatchTableEntry OpDispatch_BaseOpTable[] = {
{0xAA, 2, &OpDispatchBuilder::STOSOp},
{0xAC, 2, &OpDispatchBuilder::LODSOp},
{0xAE, 2, &OpDispatchBuilder::SCASOp},
{0xB0, 16, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVGPRImmediate>},
{0xB0, 16, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVGPROp, 0>},
{0xC2, 2, &OpDispatchBuilder::RETOp},
{0xC8, 1, &OpDispatchBuilder::EnterOp},
{0xC9, 1, &OpDispatchBuilder::LEAVEOp},
@@ -552,7 +552,7 @@ void OpDispatchBuilder::AVXInsertScalarRound(OpcodeArgs) {
const uint64_t Mode = Op->Src[2].Literal();
const auto DstSize = GetGuestVectorLength();
Ref Result = InsertScalarRoundImpl(Op, DstSize, ElementSize, Op->Dest, Op->Src[0], Mode, true);
Ref Result = InsertScalarRoundImpl(Op, DstSize, ElementSize, Op->Src[0], Op->Src[1], Mode, true);
StoreResultFPR_WithOpSize(Op, Op->Dest, Result, DstSize);
}
@@ -562,31 +562,90 @@ template void OpDispatchBuilder::AVXInsertScalarRound<OpSize::i64Bit>(OpcodeArgs
Ref OpDispatchBuilder::InsertScalarFCMPOpImpl(OpSize Size, IR::OpSize OpDstSize, IR::OpSize ElementSize, Ref Src1, Ref Src2,
uint8_t CompType, bool ZeroUpperBits) {
switch (CompType & 7) {
case 0x0: // EQ
return _VFCMPScalarInsert(Size, ElementSize, Src1, Src2, FloatCompareOp::EQ, ZeroUpperBits);
case 0x1: // LT, GT(Swapped operand)
return _VFCMPScalarInsert(Size, ElementSize, Src1, Src2, FloatCompareOp::LT, ZeroUpperBits);
case 0x2: // LE, GE(Swapped operand)
return _VFCMPScalarInsert(Size, ElementSize, Src1, Src2, FloatCompareOp::LE, ZeroUpperBits);
case 0x3: // Unordered
return _VFCMPScalarInsert(Size, ElementSize, Src1, Src2, FloatCompareOp::UNO, ZeroUpperBits);
case 0x4: // NEQ
return _VFCMPScalarInsert(Size, ElementSize, Src1, Src2, FloatCompareOp::NEQ, ZeroUpperBits);
case 0x5: { // NLT, NGT(Swapped operand)
switch (static_cast<VectorCompareType>(CompType)) {
case VectorCompareType::EQ_OQ:
case VectorCompareType::EQ_OS: return _VFCMPScalarInsert(Size, ElementSize, Src1, Src2, FloatCompareOp::EQ, ZeroUpperBits);
case VectorCompareType::LT_OS: // GT(Swapped operand)
case VectorCompareType::LT_OQ: return _VFCMPScalarInsert(Size, ElementSize, Src1, Src2, FloatCompareOp::LT, ZeroUpperBits);
case VectorCompareType::LE_OS: // GE(Swapped operand)
case VectorCompareType::LE_OQ: return _VFCMPScalarInsert(Size, ElementSize, Src1, Src2, FloatCompareOp::LE, ZeroUpperBits);
case VectorCompareType::UNORD_Q:
case VectorCompareType::UNORD_S: return _VFCMPScalarInsert(Size, ElementSize, Src1, Src2, FloatCompareOp::UNO, ZeroUpperBits);
case VectorCompareType::NEQ_UQ:
case VectorCompareType::NEQ_US: return _VFCMPScalarInsert(Size, ElementSize, Src1, Src2, FloatCompareOp::NEQ, ZeroUpperBits);
case VectorCompareType::NLT_US: // NGT(Swapped operand)
case VectorCompareType::NLT_UQ: {
Ref Result = _VFCMPLT(ElementSize, ElementSize, Src1, Src2);
Result = _VNot(ElementSize, ElementSize, Result);
// Insert the lower bits
return _VInsElement(OpDstSize, ElementSize, 0, 0, Src1, Result);
}
case 0x6: { // NLE, NGE(Swapped operand)
case VectorCompareType::NLE_US: // NGE(Swapped operand)
case VectorCompareType::NLE_UQ: {
Ref Result = _VFCMPLE(ElementSize, ElementSize, Src1, Src2);
Result = _VNot(ElementSize, ElementSize, Result);
// Insert the lower bits
return _VInsElement(OpDstSize, ElementSize, 0, 0, Src1, Result);
}
case 0x7: // Ordered
return _VFCMPScalarInsert(Size, ElementSize, Src1, Src2, FloatCompareOp::ORD, ZeroUpperBits);
case VectorCompareType::ORD_Q:
case VectorCompareType::ORD_S: return _VFCMPScalarInsert(Size, ElementSize, Src1, Src2, FloatCompareOp::ORD, ZeroUpperBits);
case VectorCompareType::NGT_UQ:
case VectorCompareType::NGT_US: {
Ref Result = _VFCMPLT(ElementSize, ElementSize, Src2, Src1);
Result = _VNot(ElementSize, ElementSize, Result);
// Insert the lower bits
return _VInsElement(OpDstSize, ElementSize, 0, 0, Src1, Result);
}
case VectorCompareType::NGE_UQ:
case VectorCompareType::NGE_US: {
Ref Result = _VFCMPLE(ElementSize, ElementSize, Src2, Src1);
Result = _VNot(ElementSize, ElementSize, Result);
// Insert the lower bits
return _VInsElement(OpDstSize, ElementSize, 0, 0, Src1, Result);
}
case VectorCompareType::GT_OQ:
case VectorCompareType::GT_OS: {
Ref Result = _VFCMPLT(ElementSize, ElementSize, Src2, Src1);
// Insert the lower bits
return _VInsElement(OpDstSize, ElementSize, 0, 0, Src1, Result);
}
case VectorCompareType::GE_OQ:
case VectorCompareType::GE_OS: {
Ref Result = _VFCMPLE(ElementSize, ElementSize, Src2, Src1);
// Insert the lower bits
return _VInsElement(OpDstSize, ElementSize, 0, 0, Src1, Result);
}
case VectorCompareType::EQ_UQ:
case VectorCompareType::EQ_US: {
// If either of the sources are unordered, then returns true.
Ref Src1_U = _VFCMPEQ(Size, ElementSize, Src1, Src1);
Ref Src2_U = _VFCMPEQ(Size, ElementSize, Src2, Src2);
auto Ordered = _VAnd(Size, ElementSize, Src1_U, Src2_U);
Ref Compare_Ordered = _VFCMPEQ(Size, ElementSize, Src1, Src2);
Ref Result = _VOrn(Size, ElementSize, Compare_Ordered, Ordered);
// Insert the lower bits
return _VInsElement(OpDstSize, ElementSize, 0, 0, Src1, Result);
}
case VectorCompareType::NEQ_OQ:
case VectorCompareType::NEQ_OS: {
// If either of the sources are unordered, then returns false.
Ref Src1_U = _VFCMPEQ(Size, ElementSize, Src1, Src1);
Ref Src2_U = _VFCMPEQ(Size, ElementSize, Src2, Src2);
Ref Compare_Ordered = _VFCMPEQ(Size, ElementSize, Src1, Src2);
Ref Result = _VAndn(Size, ElementSize, Src1_U, Compare_Ordered);
Result = _VAnd(Size, ElementSize, Result, Src2_U);
// Insert the lower bits
return _VInsElement(OpDstSize, ElementSize, 0, 0, Src1, Result);
}
case VectorCompareType::FALSE_OQ:
case VectorCompareType::FALSE_OS: return _VInsElement(OpDstSize, ElementSize, 0, 0, Src1, LoadZeroVector(OpSize::i128Bit));
case VectorCompareType::TRUE_UQ:
case VectorCompareType::TRUE_US:
return _VInsElement(OpDstSize, ElementSize, 0, 0, Src1, _VectorImm(OpSize::i128Bit, OpSize::i8Bit, -1, 0));
}
FEX_UNREACHABLE;
}
@@ -600,7 +659,7 @@ void OpDispatchBuilder::InsertScalarFCMPOp(OpcodeArgs) {
Ref Src1 = LoadSourceFPR_WithOpSize(Op, Op->Dest, DstSize, Op->Flags);
Ref Src2 = LoadSourceFPR_WithOpSize(Op, Op->Src[0], SrcSize, Op->Flags, {.AllowUpperGarbage = true});
Ref Result = InsertScalarFCMPOpImpl(DstSize, OpSizeFromDst(Op), ElementSize, Src1, Src2, CompType, false);
Ref Result = InsertScalarFCMPOpImpl(DstSize, OpSizeFromDst(Op), ElementSize, Src1, Src2, CompType & 0b111, false);
StoreResultFPR_WithOpSize(Op, Op->Dest, Result, DstSize);
}
@@ -619,7 +678,7 @@ void OpDispatchBuilder::AVXInsertScalarFCMPOp(OpcodeArgs) {
Ref Src1 = LoadSourceFPR_WithOpSize(Op, Op->Src[0], DstSize, Op->Flags);
Ref Src2 = LoadSourceFPR_WithOpSize(Op, Op->Src[1], SrcSize, Op->Flags, {.AllowUpperGarbage = true});
Ref Result = InsertScalarFCMPOpImpl(DstSize, OpSizeFromDst(Op), ElementSize, Src1, Src2, CompType, true);
Ref Result = InsertScalarFCMPOpImpl(DstSize, OpSizeFromDst(Op), ElementSize, Src1, Src2, CompType & 0b11111, true);
StoreResultFPR_WithOpSize(Op, Op->Dest, Result, DstSize);
}
@@ -952,10 +1011,522 @@ Ref OpDispatchBuilder::PShufWLane(IR::OpSize Size, FEXCore::IR::IndexNamedVector
}
void OpDispatchBuilder::PSHUFW8ByteOp(OpcodeArgs) {
uint16_t Shuffle = Op->Src[1].Data.Literal.Value;
uint8_t Shuffle = Op->Src[1].Data.Literal.Value;
const auto Size = OpSizeFromSrc(Op);
const auto TBLIndex = FEXCore::IR::INDEXED_NAMED_VECTOR_PSHUFLW;
Ref Src = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
Ref Dest = PShufWLane(Size, FEXCore::IR::INDEXED_NAMED_VECTOR_PSHUFLW, true, Src, Shuffle);
// Single MMX 64-bit shuffle. Shuffle selector can fit full selection.
Ref Dest {};
switch (Shuffle) {
// Single-instruction shuffle operations.
case 0b00'00'00'00: Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 0); break;
case 0b00'10'01'00: Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 0, Src, Src); break;
case 0b01'00'01'00: Dest = _VZip(OpSize::i64Bit, OpSize::i32Bit, Src, Src); break;
case 0b00'11'10'01: Dest = _VExtr(OpSize::i64Bit, OpSize::i16Bit, Src, Src, 1); break;
case 0b01'00'11'10: Dest = _VRev64(OpSize::i64Bit, OpSize::i32Bit, Src); break;
case 0b01'01'01'01: Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 1); break;
case 0b01'10'01'00: Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 1, Src, Src); break;
case 0b10'10'01'00: Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 2, Src, Src); break;
case 0b10'10'10'10: Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 2); break;
case 0b11'00'01'00: Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 0, Src, Src); break;
case 0b11'01'01'00: Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 1, Src, Src); break;
case 0b11'10'00'00: Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 0, Src, Src); break;
case 0b11'10'01'00: Dest = Src; break;
case 0b11'10'01'01: Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 1, Src, Src); break;
case 0b11'10'01'10: Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 2, Src, Src); break;
case 0b11'10'01'11: Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 3, Src, Src); break;
case 0b11'10'10'00: Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 2, Src, Src); break;
case 0b11'10'11'00: Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 3, Src, Src); break;
case 0b11'10'11'10: Dest = _VDupElement(OpSize::i64Bit, OpSize::i32Bit, Src, 1); break;
case 0b11'11'01'00: Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 3, Src, Src); break;
case 0b11'11'11'11: Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 3); break;
// Two instruction shuffle operations.
case 0b00'00'00'01:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 0);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 1, Dest, Src);
break;
case 0b00'00'00'10:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 0);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 2, Dest, Src);
break;
case 0b00'00'00'11:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 0);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 3, Dest, Src);
break;
case 0b00'00'01'00:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 0);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 1, Dest, Src);
break;
case 0b00'00'10'00:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 0);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 2, Dest, Src);
break;
case 0b00'00'11'00:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 0);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 3, Dest, Src);
break;
case 0b00'00'11'10:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 0);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i32Bit, 0, 1, Dest, Src);
break;
case 0b00'01'00'00:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 0);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 1, Dest, Src);
break;
case 0b00'01'00'01:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i32Bit, Src, 0);
Dest = _VExtr(OpSize::i64Bit, OpSize::i16Bit, Dest, Dest, 1);
break;
case 0b00'01'00'11:
Dest = _VZip(OpSize::i64Bit, OpSize::i32Bit, Src, Src);
Dest = _VExtr(OpSize::i64Bit, OpSize::i8Bit, Dest, Src, 6);
break;
case 0b00'01'01'00:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 1, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 0, Dest, Src);
break;
case 0b00'01'01'01:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 1);
Dest = _VExtr(OpSize::i64Bit, OpSize::i16Bit, Src, Dest, 1);
break;
case 0b00'10'00'00:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 0);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 2, Dest, Src);
break;
case 0b00'10'00'10:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 0, Src, Src);
Dest = _VDupElement(OpSize::i64Bit, OpSize::i32Bit, Dest, 1);
break;
case 0b00'10'01'01:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 1, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 0, Dest, Src);
break;
case 0b00'10'01'10:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 2, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 0, Dest, Src);
break;
case 0b00'10'01'11:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 3, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 0, Dest, Src);
break;
case 0b00'10'10'00:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 2, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 0, Dest, Src);
break;
case 0b00'10'10'10:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 2);
Dest = _VExtr(OpSize::i64Bit, OpSize::i16Bit, Src, Dest, 1);
break;
case 0b00'10'11'00:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 3, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 0, Dest, Src);
break;
case 0b00'10'11'10:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i32Bit, Src, 1);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 0, Dest, Src);
break;
case 0b00'11'00'00:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 0);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 3, Dest, Src);
break;
case 0b00'11'01'00:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 3, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 0, Dest, Src);
break;
case 0b00'11'10'11:
Dest = _VZip2(OpSize::i64Bit, OpSize::i32Bit, Src, Src);
Dest = _VExtr(OpSize::i64Bit, OpSize::i16Bit, Src, Dest, 1);
break;
case 0b00'11'11'11:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 3);
Dest = _VExtr(OpSize::i64Bit, OpSize::i16Bit, Src, Dest, 1);
break;
case 0b01'00'00'00:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 0);
Dest = _VZip(OpSize::i64Bit, OpSize::i32Bit, Dest, Src);
break;
case 0b01'00'00'10:
Dest = _VRev64(OpSize::i64Bit, OpSize::i32Bit, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 0, Dest, Src);
break;
case 0b01'00'01'01:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 1);
Dest = _VZip(OpSize::i64Bit, OpSize::i32Bit, Dest, Src);
break;
case 0b01'00'01'10:
Dest = _VZip(OpSize::i64Bit, OpSize::i32Bit, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 2, Dest, Src);
break;
case 0b01'00'01'11:
Dest = _VZip(OpSize::i64Bit, OpSize::i32Bit, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 3, Dest, Src);
break;
case 0b01'00'10'00:
Dest = _VZip(OpSize::i64Bit, OpSize::i32Bit, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 2, Dest, Src);
break;
case 0b01'00'10'01:
Dest = _VExtr(OpSize::i64Bit, OpSize::i8Bit, Src, Src, 2);
Dest = _VZip(OpSize::i64Bit, OpSize::i32Bit, Dest, Src);
break;
case 0b01'00'10'10:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 2);
Dest = _VZip(OpSize::i64Bit, OpSize::i32Bit, Dest, Src);
break;
case 0b01'00'11'00:
Dest = _VZip(OpSize::i64Bit, OpSize::i32Bit, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 3, Dest, Src);
break;
case 0b01'00'11'01:
Dest = _VRev64(OpSize::i64Bit, OpSize::i32Bit, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 1, Dest, Src);
break;
case 0b01'00'11'11:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 3);
Dest = _VZip(OpSize::i64Bit, OpSize::i32Bit, Dest, Src);
break;
case 0b01'01'00'01:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 1);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 0, Dest, Src);
break;
case 0b01'01'01'00:
Dest = _VZip(OpSize::i64Bit, OpSize::i32Bit, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 1, Dest, Src);
break;
case 0b01'01'01'10:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 1);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 2, Dest, Src);
break;
case 0b01'01'01'11:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 1);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 3, Dest, Src);
break;
case 0b01'01'10'01:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 1);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 2, Dest, Src);
break;
case 0b01'01'11'01:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 1);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 3, Dest, Src);
break;
case 0b01'01'11'10:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 1);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i32Bit, 0, 1, Dest, Src);
break;
case 0b01'10'00'00:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 0, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 1, Dest, Src);
break;
case 0b01'10'01'01:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 1);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 2, Dest, Src);
break;
case 0b01'10'01'10:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 2, Src, Src);
Dest = _VZip(OpSize::i64Bit, OpSize::i32Bit, Dest, Dest);
break;
case 0b01'10'01'11:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 3, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 1, Dest, Src);
break;
case 0b01'10'10'00:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 2, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 1, Dest, Src);
break;
case 0b01'10'10'10:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 2);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 1, Dest, Src);
break;
case 0b01'10'11'00:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 3, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 1, Dest, Src);
break;
case 0b01'10'11'10:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i32Bit, Src, 1);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 1, Dest, Src);
break;
case 0b01'11'01'00:
Dest = _VZip(OpSize::i64Bit, OpSize::i32Bit, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 3, Dest, Src);
break;
case 0b01'11'01'01:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 1);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 3, Dest, Src);
break;
case 0b01'11'01'11:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 3, Src, Src);
Dest = _VZip(OpSize::i64Bit, OpSize::i32Bit, Dest, Dest);
break;
case 0b01'11'10'01:
Dest = _VExtr(OpSize::i64Bit, OpSize::i8Bit, Src, Src, 2);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 0, Dest, Dest);
break;
case 0b01'11'11'10:
Dest = _VRev64(OpSize::i64Bit, OpSize::i32Bit, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 1, Dest, Dest);
break;
case 0b01'11'11'11:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 3);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 1, Dest, Src);
break;
case 0b10'00'00'00:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 0);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 2, Dest, Src);
break;
case 0b10'00'01'00:
Dest = _VZip(OpSize::i64Bit, OpSize::i32Bit, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 2, Dest, Src);
break;
case 0b10'00'10'00:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 2, Src, Src);
Dest = _VZip(OpSize::i64Bit, OpSize::i32Bit, Dest, Dest);
break;
case 0b10'00'10'10:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 2);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 0, Dest, Src);
break;
case 0b10'00'11'10:
Dest = _VRev64(OpSize::i64Bit, OpSize::i32Bit, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 0, Dest, Dest);
break;
case 0b10'01'00'00:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 0);
Dest = _VExtr(OpSize::i64Bit, OpSize::i8Bit, Src, Dest, 6);
break;
case 0b10'01'00'01:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 1);
Dest = _VExtr(OpSize::i64Bit, OpSize::i8Bit, Src, Dest, 6);
break;
case 0b10'01'00'10:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 2);
Dest = _VExtr(OpSize::i64Bit, OpSize::i8Bit, Src, Dest, 6);
break;
case 0b10'01'00'11:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 3);
Dest = _VExtr(OpSize::i64Bit, OpSize::i8Bit, Src, Dest, 6);
break;
case 0b10'01'01'00:
Dest = _VExtr(OpSize::i64Bit, OpSize::i8Bit, Src, Src, 6);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i32Bit, 0, 0, Dest, Src);
break;
case 0b10'01'01'01:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 1);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 2, Dest, Src);
break;
case 0b10'01'10'01:
Dest = _VExtr(OpSize::i64Bit, OpSize::i8Bit, Src, Src, 2);
Dest = _VZip(OpSize::i64Bit, OpSize::i32Bit, Dest, Dest);
break;
case 0b10'01'10'10:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 2);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 1, Dest, Src);
break;
case 0b10'01'11'10:
Dest = _VExtr(OpSize::i64Bit, OpSize::i8Bit, Src, Src, 6);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i32Bit, 0, 1, Dest, Src);
break;
case 0b10'10'00'00:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 0, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 2, Dest, Src);
break;
case 0b10'10'00'10:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 2);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 0, Dest, Src);
break;
case 0b10'10'01'01:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 1, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 2, Dest, Src);
break;
case 0b10'10'01'10:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 2);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 1, Dest, Src);
break;
case 0b10'10'01'11:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 3, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 2, Dest, Src);
break;
case 0b10'10'10'00:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 2);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 0, Dest, Src);
break;
case 0b10'10'10'01:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 2);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 1, Dest, Src);
break;
case 0b10'10'10'11:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 2);
Dest = _VExtr(OpSize::i64Bit, OpSize::i8Bit, Dest, Src, 6);
break;
case 0b10'10'11'00:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 3, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 2, Dest, Src);
break;
case 0b10'10'11'10:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 2);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 3, Dest, Src);
break;
case 0b10'11'01'00:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 3, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 2, Dest, Src);
break;
case 0b10'11'10'10:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 2);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 3, Dest, Src);
break;
case 0b10'11'10'11:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i32Bit, Src, 1);
Dest = _VExtr(OpSize::i64Bit, OpSize::i16Bit, Dest, Dest, 1);
break;
case 0b10'11'11'11:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 3);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 2, Dest, Src);
break;
case 0b11'00'00'00:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 0);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 3, Dest, Src);
break;
case 0b11'00'01'01:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 1, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 0, Dest, Src);
break;
case 0b11'00'01'10:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 2, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 0, Dest, Src);
break;
case 0b11'00'01'11:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 3, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 0, Dest, Src);
break;
case 0b11'00'10'00:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 2, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 0, Dest, Src);
break;
case 0b11'00'11'00:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 3, Src, Src);
Dest = _VZip(OpSize::i64Bit, OpSize::i32Bit, Dest, Dest);
break;
case 0b11'00'11'10:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i32Bit, Src, 1);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 0, Dest, Src);
break;
case 0b11'00'11'11:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 3);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 0, Dest, Src);
break;
case 0b11'01'00'00:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 0, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 1, Dest, Src);
break;
case 0b11'01'01'01:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 1);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 3, Dest, Src);
break;
case 0b11'01'01'10:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 2, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 1, Dest, Src);
break;
case 0b11'01'01'11:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 3, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 1, Dest, Src);
break;
case 0b11'01'10'00:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 2, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 1, Dest, Src);
break;
case 0b11'01'11'00:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 3, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 1, Dest, Src);
break;
case 0b11'01'11'01:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 1, Src, Src);
Dest = _VDupElement(OpSize::i64Bit, OpSize::i32Bit, Dest, 1);
break;
case 0b11'01'11'10:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i32Bit, Src, 1);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 1, Dest, Src);
break;
case 0b11'01'11'11:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 3);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 1, Dest, Src);
break;
case 0b11'10'00'01:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 1, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 0, Dest, Src);
break;
case 0b11'10'00'10:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i32Bit, Src, 1);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 0, Dest, Src);
break;
case 0b11'10'00'11:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 3, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 0, Dest, Src);
break;
case 0b11'10'10'01:
Dest = _VExtr(OpSize::i64Bit, OpSize::i8Bit, Src, Src, 2);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i32Bit, 1, 1, Dest, Src);
break;
case 0b11'10'10'10:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 2);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 3, 3, Dest, Src);
break;
case 0b11'10'10'11:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 3, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 2, Dest, Src);
break;
case 0b11'10'11'01:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i32Bit, Src, 1);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 1, Dest, Src);
break;
case 0b11'10'11'11:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 3);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 2, Dest, Src);
break;
case 0b11'11'00'00:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 0, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 3, Dest, Src);
break;
case 0b11'11'00'11:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 3);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 0, Dest, Src);
break;
case 0b11'11'01'01:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 1, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 3, Dest, Src);
break;
case 0b11'11'01'10:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 2, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 3, Dest, Src);
break;
case 0b11'11'01'11:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 3);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 1, Dest, Src);
break;
case 0b11'11'10'00:
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 2, Src, Src);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 2, 3, Dest, Src);
break;
case 0b11'11'10'11:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 3);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 1, 2, Dest, Src);
break;
case 0b11'11'11'00:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 3);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 0, Dest, Src);
break;
case 0b11'11'11'01:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 3);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 1, Dest, Src);
break;
case 0b11'11'11'10:
Dest = _VDupElement(OpSize::i64Bit, OpSize::i16Bit, Src, 3);
Dest = _VInsElement(OpSize::i64Bit, OpSize::i16Bit, 0, 2, Dest, Src);
break;
default:
auto LookupIndexes = LoadAndCacheIndexedNamedVectorConstant(Size, TBLIndex, Shuffle * 16);
Dest = _VTBL1(Size, Src, LookupIndexes);
break;
}
StoreResultFPR(Op, Dest);
}
@@ -1288,6 +1859,11 @@ Ref OpDispatchBuilder::SHUFOpImpl(OpcodeArgs, IR::OpSize DstSize, IR::OpSize Ele
Shuffle >>= ShiftAmount;
}
} else {
if (Src1 == Src2 && Shuffle == 0) {
// TODO: We can optimize significantly more shuffles when we know the sources match.
// Special case broadcast element 0.
return _VDupElement(DstSize, ElementSize, Src1, Shuffle & SelectionMask);
}
if (ElementSize == OpSize::i32Bit) {
// We can shuffle optimally in a lot of cases.
// TODO: We can optimize more of these cases.
@@ -2424,26 +3000,67 @@ void OpDispatchBuilder::MOVBetweenGPR_FPR(OpcodeArgs, VectorOpType VectorType) {
}
Ref OpDispatchBuilder::VFCMPOpImpl(OpSize Size, IR::OpSize ElementSize, Ref Src1, Ref Src2, uint8_t CompType) {
Ref Result {};
switch (CompType & 0x7) {
case 0x0: // EQ
return _VFCMPEQ(Size, ElementSize, Src1, Src2);
case 0x1: // LT, GT(Swapped operand)
return _VFCMPLT(Size, ElementSize, Src1, Src2);
case 0x2: // LE, GE(Swapped operand)
return _VFCMPLE(Size, ElementSize, Src1, Src2);
case 0x3: // Unordered
return _VFCMPUNO(Size, ElementSize, Src1, Src2);
case 0x4: // NEQ
return _VFCMPNEQ(Size, ElementSize, Src1, Src2);
case 0x5: // NLT, NGT(Swapped operand)
Result = _VFCMPLT(Size, ElementSize, Src1, Src2);
switch (static_cast<VectorCompareType>(CompType)) {
case VectorCompareType::EQ_OQ:
case VectorCompareType::EQ_OS: return _VFCMPEQ(Size, ElementSize, Src1, Src2);
case VectorCompareType::LT_OS: // GT(Swapped operand)
case VectorCompareType::LT_OQ: return _VFCMPLT(Size, ElementSize, Src1, Src2);
case VectorCompareType::LE_OS: // GE(Swapped operand)
case VectorCompareType::LE_OQ: return _VFCMPLE(Size, ElementSize, Src1, Src2);
case VectorCompareType::UNORD_Q:
case VectorCompareType::UNORD_S: return _VFCMPUNO(Size, ElementSize, Src1, Src2);
case VectorCompareType::NEQ_UQ:
case VectorCompareType::NEQ_US: return _VFCMPNEQ(Size, ElementSize, Src1, Src2);
case VectorCompareType::NLT_US: // NGT(Swapped operand)
case VectorCompareType::NLT_UQ: {
Ref Result = _VFCMPLT(Size, ElementSize, Src1, Src2);
return _VNot(Size, ElementSize, Result);
case 0x6: // NLE, NGE(Swapped operand)
Result = _VFCMPLE(Size, ElementSize, Src1, Src2);
}
case VectorCompareType::NLE_US: // NGE(Swapped operand)
case VectorCompareType::NLE_UQ: {
Ref Result = _VFCMPLE(Size, ElementSize, Src1, Src2);
return _VNot(Size, ElementSize, Result);
case 0x7: // Ordered
return _VFCMPORD(Size, ElementSize, Src1, Src2);
}
case VectorCompareType::ORD_Q:
case VectorCompareType::ORD_S: return _VFCMPORD(Size, ElementSize, Src1, Src2);
case VectorCompareType::NGT_UQ:
case VectorCompareType::NGT_US: {
Ref Result = _VFCMPLT(Size, ElementSize, Src2, Src1);
return _VNot(Size, ElementSize, Result);
}
case VectorCompareType::NGE_UQ:
case VectorCompareType::NGE_US: {
Ref Result = _VFCMPLE(Size, ElementSize, Src2, Src1);
return _VNot(Size, ElementSize, Result);
}
case VectorCompareType::GT_OQ:
case VectorCompareType::GT_OS: return _VFCMPLT(Size, ElementSize, Src2, Src1);
case VectorCompareType::GE_OQ:
case VectorCompareType::GE_OS: return _VFCMPLE(Size, ElementSize, Src2, Src1);
case VectorCompareType::EQ_UQ:
case VectorCompareType::EQ_US: {
// If either of the sources are unordered, then returns true.
Ref Src1_U = _VFCMPEQ(Size, ElementSize, Src1, Src1);
Ref Src2_U = _VFCMPEQ(Size, ElementSize, Src2, Src2);
auto Ordered = _VAnd(Size, ElementSize, Src1_U, Src2_U);
Ref Compare_Ordered = _VFCMPEQ(Size, ElementSize, Src1, Src2);
return _VOrn(Size, ElementSize, Compare_Ordered, Ordered);
}
case VectorCompareType::NEQ_OQ:
case VectorCompareType::NEQ_OS: {
// If either of the sources are unordered, then returns false.
Ref Src1_U = _VFCMPEQ(Size, ElementSize, Src1, Src1);
Ref Src2_U = _VFCMPEQ(Size, ElementSize, Src2, Src2);
Ref Compare_Ordered = _VFCMPEQ(Size, ElementSize, Src1, Src2);
Ref Result = _VAndn(Size, ElementSize, Src1_U, Compare_Ordered);
return _VAnd(Size, ElementSize, Result, Src2_U);
}
case VectorCompareType::FALSE_OQ:
case VectorCompareType::FALSE_OS: return LoadZeroVector(Size);
case VectorCompareType::TRUE_UQ:
case VectorCompareType::TRUE_US: return _VectorImm(Size, OpSize::i8Bit, -1, 0);
}
FEX_UNREACHABLE;
}
@@ -2459,7 +3076,7 @@ void OpDispatchBuilder::VFCMPOp(OpcodeArgs) {
Ref Dest = LoadSourceFPR_WithOpSize(Op, Op->Dest, DstSize, Op->Flags);
const uint8_t CompType = Op->Src[1].Data.Literal.Value;
Ref Result = VFCMPOpImpl(OpSizeFromSrc(Op), ElementSize, Dest, Src, CompType);
Ref Result = VFCMPOpImpl(OpSizeFromSrc(Op), ElementSize, Dest, Src, CompType & 0b111);
StoreResultFPR(Op, Result);
}
@@ -2477,7 +3094,7 @@ void OpDispatchBuilder::AVXVFCMPOp(OpcodeArgs) {
Ref Src1 = LoadSourceFPR_WithOpSize(Op, Op->Src[0], DstSize, Op->Flags);
Ref Src2 = LoadSourceFPR_WithOpSize(Op, Op->Src[1], SrcSize, Op->Flags);
Ref Result = VFCMPOpImpl(OpSizeFromSrc(Op), ElementSize, Src1, Src2, CompType);
Ref Result = VFCMPOpImpl(OpSizeFromSrc(Op), ElementSize, Src1, Src2, CompType & 0b11111);
StoreResultFPR(Op, Result);
}
@@ -2656,7 +3273,7 @@ void OpDispatchBuilder::SaveSSEState(Ref MemBase) {
void OpDispatchBuilder::SaveMXCSRState(Ref MemBase) {
// Store MXCSR and the mask for all bits.
_StoreMemPairGPR(OpSize::i32Bit, GetMXCSR(), Constant(0xFFFF), MemBase, 24);
_StoreMemPairGPR(OpSize::i32Bit, GetMXCSR(), Constant(0xFFC0), MemBase, 24);
}
void OpDispatchBuilder::SaveAVXState(Ref MemBase) {
@@ -5017,7 +5634,8 @@ void OpDispatchBuilder::VFMAddSubImpl(OpcodeArgs, bool AddSub, uint8_t Src1Idx,
OpDispatchBuilder::RefVSIB OpDispatchBuilder::LoadVSIB(const X86Tables::DecodedOp& Op, const X86Tables::DecodedOperand& Operand, uint32_t Flags) {
const bool IsVSIB = (Op->Flags & X86Tables::DecodeFlags::FLAG_VSIB_BYTE) != 0;
LOGMAN_THROW_A_FMT(Operand.IsSIB() && IsVSIB, "Trying to load VSIB for something that isn't the correct type!");
LOGMAN_THROW_A_FMT((Operand.IsSIB() || Operand.IsSIBRelocation()) && IsVSIB, "Trying to load VSIB for something that isn't the correct "
"type!");
// VSIB is a very special case which has a ton of encoded data.
// Get it in a format we can reason about.
@@ -5029,12 +5647,24 @@ OpDispatchBuilder::RefVSIB OpDispatchBuilder::LoadVSIB(const X86Tables::DecodedO
"Base must be a GPR.");
const auto Index_XMM_gpr = Index_gpr - X86State::REG_XMM_0;
return {
OpDispatchBuilder::RefVSIB A {
.Low = LoadXMMRegister(Index_XMM_gpr),
.BaseAddr = Base_gpr != FEXCore::X86State::REG_INVALID ? LoadGPRRegister(Base_gpr, OpSize::i64Bit, 0, false) : nullptr,
.Displacement = Operand.Data.SIB.Offset,
.Scale = Operand.Data.SIB.Scale,
};
if (Operand.IsSIBRelocation()) {
auto EPOffset = _EntrypointOffset(OpSize::i64Bit, Operand.Data.SIB.Offset);
if (A.BaseAddr) {
A.BaseAddr = Add(OpSize::i64Bit, EPOffset, A.BaseAddr);
} else {
A.BaseAddr = EPOffset;
}
} else {
A.Displacement = static_cast<int32_t>(Operand.Data.SIB.Offset);
}
return A;
}
template<OpSize AddrElementSize>
@@ -60,15 +60,13 @@ void OpDispatchBuilder::SetX87Top(Ref Value) {
// Float LoaD operation with memory operand
void OpDispatchBuilder::FLD(OpcodeArgs, IR::OpSize Width) {
const auto ReadWidth = (Width == OpSize::f80Bit) ? OpSize::i128Bit : Width;
Ref Data = LoadSourceFPR_WithOpSize(Op, Op->Src[0], Width, Op->Flags);
Ref ConvertedData = Data;
// Convert to 80bit float
if (Width == OpSize::i32Bit || Width == OpSize::i64Bit) {
ConvertedData = _F80CVTTo(Data, ReadWidth);
ConvertedData = _F80CVTTo(Data, Width);
}
_PushStack(ConvertedData, Data, ReadWidth);
_PushStack(ConvertedData, Data, Width);
}
// Float LoaD operation with memory operand
@@ -80,7 +78,7 @@ void OpDispatchBuilder::FBLD(OpcodeArgs) {
// Read from memory
Ref Data = LoadSourceFPR_WithOpSize(Op, Op->Src[0], OpSize::f80Bit, Op->Flags);
Ref ConvertedData = _F80BCDLoad(Data);
_PushStack(ConvertedData, Data, OpSize::i128Bit);
_PushStack(ConvertedData, Invalid(), OpSize::iInvalid);
}
void OpDispatchBuilder::FBSTP(OpcodeArgs) {
@@ -92,7 +90,7 @@ void OpDispatchBuilder::FBSTP(OpcodeArgs) {
void OpDispatchBuilder::FLD_Const(OpcodeArgs, NamedVectorConstant K) {
// Update TOP
Ref Data = LoadAndCacheNamedVectorConstant(OpSize::i128Bit, K);
_PushStack(Data, Data, OpSize::i128Bit);
_PushStack(Data, Data, OpSize::f80Bit);
}
void OpDispatchBuilder::FILD(OpcodeArgs) {
@@ -123,11 +121,12 @@ void OpDispatchBuilder::FILD(OpcodeArgs) {
auto upper = _Or(OpSize::i64Bit, sign, zeroed_exponent);
Ref ConvertedData = _VLoadTwoGPRs(shifted, upper);
_PushStack(ConvertedData, Invalid(), ReadWidth);
_PushStack(ConvertedData, Invalid(), OpSize::iInvalid);
}
void OpDispatchBuilder::FST(OpcodeArgs, IR::OpSize Width) {
const auto SourceSize = ReducedPrecisionMode ? OpSize::i64Bit : OpSize::i128Bit;
LOGMAN_THROW_A_FMT(Width == OpSize::i32Bit || Width == OpSize::i64Bit || Width == OpSize::f80Bit, "Invalid store width for FST");
const auto SourceSize = ReducedPrecisionMode ? OpSize::i64Bit : OpSize::f80Bit;
AddressMode A = DecodeAddress(Op, Op->Dest, MemoryAccessType::DEFAULT, false);
A = SelectAddressMode(this, A, GetGPROpSize(), CTX->HostFeatures.SupportsTSOImm9, false, false, Width);
@@ -164,7 +163,7 @@ void OpDispatchBuilder::FIST(OpcodeArgs, bool Truncate) {
// Check for NaN/Infinity: exponent = 0x7fff
SaveNZCV();
_TestNZ(OpSize::i64Bit, Exponent, Constant(0x7fff));
SubWithFlags(OpSize::i64Bit, Exponent, 0x7fff);
Ref IsSpecial = _NZCVSelect01(CondClass::EQ);
// For overflow detection, check if exponent indicates a value >= 2^15
@@ -179,7 +178,8 @@ void OpDispatchBuilder::FIST(OpcodeArgs, bool Truncate) {
Data = _F80CVTInt(Size, Data, Truncate);
StoreResultGPR_WithOpSize(Op, Op->Dest, Data, Size, OpSize::i8Bit);
StoreResultGPR_WithOpSize(Op, Op->Dest, Data, Size, OpSize::i8Bit,
CTX->IsVectorAtomicTSOEnabled() ? MemoryAccessType::DEFAULT : MemoryAccessType::NONTSO);
if ((Op->TableInfo->Flags & X86Tables::InstFlags::FLAGS_POP) != 0) {
_PopStackDestroy();
@@ -877,8 +877,8 @@ void OpDispatchBuilder::X87FXTRACT(OpcodeArgs) {
_PopStackDestroy();
auto Exp = _F80XTRACT_EXP(Top);
auto Sig = _F80XTRACT_SIG(Top);
_PushStack(Exp, Invalid(), OpSize::f80Bit);
_PushStack(Sig, Invalid(), OpSize::f80Bit);
_PushStack(Exp, Invalid(), OpSize::iInvalid);
_PushStack(Sig, Invalid(), OpSize::iInvalid);
}
} // namespace FEXCore::IR
@@ -59,7 +59,6 @@ void OpDispatchBuilder::X87FLDCWF64(OpcodeArgs) {
// F64 ops
// Float load op with memory operand
void OpDispatchBuilder::FLDF64(OpcodeArgs, IR::OpSize Width) {
const auto ReadWidth = (Width == OpSize::f80Bit) ? OpSize::i128Bit : Width;
Ref Data = LoadSourceFPR_WithOpSize(Op, Op->Src[0], Width, Op->Flags);
// Convert to 64bit float
Ref ConvertedData = Data;
@@ -68,7 +67,7 @@ void OpDispatchBuilder::FLDF64(OpcodeArgs, IR::OpSize Width) {
} else if (Width == OpSize::f80Bit) {
ConvertedData = _F80CVT(OpSize::i64Bit, Data);
}
_PushStack(ConvertedData, Data, ReadWidth);
_PushStack(ConvertedData, Data, Width);
}
void OpDispatchBuilder::FBLDF64(OpcodeArgs) {
@@ -76,7 +75,7 @@ void OpDispatchBuilder::FBLDF64(OpcodeArgs) {
Ref Data = LoadSourceFPR_WithOpSize(Op, Op->Src[0], OpSize::f80Bit, Op->Flags);
Ref ConvertedData = _F80BCDLoad(Data);
ConvertedData = _F80CVT(OpSize::i64Bit, ConvertedData);
_PushStack(ConvertedData, Data, OpSize::i64Bit);
_PushStack(ConvertedData, Invalid(), OpSize::iInvalid);
}
void OpDispatchBuilder::FBSTPF64(OpcodeArgs) {
@@ -100,19 +99,31 @@ void OpDispatchBuilder::FILDF64(OpcodeArgs) {
Data = _Sbfe(OpSize::i64Bit, IR::OpSizeAsBits(ReadWidth), 0, Data);
}
auto ConvertedData = _Float_FromGPR_S(OpSize::i64Bit, ReadWidth == OpSize::i32Bit ? OpSize::i32Bit : OpSize::i64Bit, Data);
_PushStack(ConvertedData, Invalid(), ReadWidth);
_PushStack(ConvertedData, Invalid(), OpSize::iInvalid);
}
void OpDispatchBuilder::FISTF64(OpcodeArgs, bool Truncate) {
const auto Size = OpSizeFromSrc(Op);
Ref data = _ReadStackValue(0);
if (Truncate) {
data = _Float_ToGPR_ZS(Size == OpSize::i32Bit ? OpSize::i32Bit : OpSize::i64Bit, OpSize::i64Bit, data);
bool CanUseFloatReg = Size == OpSize::i64Bit;
if (CanUseFloatReg) {
// If possible, it's faster to keep the data in an FPR than doing a GPR transfer.
if (Truncate) {
data = _Vector_FToZS(OpSize::i128Bit, OpSize::i64Bit, data);
} else {
data = _Vector_FToS(OpSize::i128Bit, OpSize::i64Bit, data);
}
StoreResultFPR_WithOpSize(Op, Op->Dest, data, OpSize::i64Bit, OpSize::i8Bit);
} else {
data = _Float_ToGPR_S(Size == OpSize::i32Bit ? OpSize::i32Bit : OpSize::i64Bit, OpSize::i64Bit, data);
if (Truncate) {
data = _Float_ToGPR_ZS(Size == OpSize::i32Bit ? OpSize::i32Bit : OpSize::i64Bit, OpSize::i64Bit, data);
} else {
data = _Float_ToGPR_S(Size == OpSize::i32Bit ? OpSize::i32Bit : OpSize::i64Bit, OpSize::i64Bit, data);
}
StoreResultGPR_WithOpSize(Op, Op->Dest, data, Size, OpSize::i8Bit,
CTX->IsVectorAtomicTSOEnabled() ? MemoryAccessType::DEFAULT : MemoryAccessType::NONTSO);
}
StoreResultGPR_WithOpSize(Op, Op->Dest, data, Size, OpSize::i8Bit);
if ((Op->TableInfo->Flags & X86Tables::InstFlags::FLAGS_POP) != 0) {
_PopStackDestroy();
@@ -371,6 +382,8 @@ void OpDispatchBuilder::X87FXTRACTF64(OpcodeArgs) {
// Split node into SIG and EXP while handling the special zero case.
// i.e. if val == 0.0, then sig = 0.0, exp = -inf
// if val == -0.0, then sig = -0.0, exp = -inf
// if val is +/-Inf, then sig = val, exp = +inf
// if val is NaN, then sig = val, exp = val
// otherwise we just extract the 64-bit sig and exp as normal.
Ref Node = _ReadStackValue(0);
@@ -380,6 +393,11 @@ void OpDispatchBuilder::X87FXTRACTF64(OpcodeArgs) {
Ref ExpZV = _VCastFromGPR(OpSize::i64Bit, OpSize::i64Bit, Constant(0xfff0'0000'0000'0000UL));
Ref SigZV = Node;
// Inf/NaN case
Ref ExpInfOnlyV = _VCastFromGPR(OpSize::i64Bit, OpSize::i64Bit, Constant(0x7ff0'0000'0000'0000UL));
Ref ExpNanV = Node;
Ref SigInfV = Node;
// non zero case
Ref ExpNZ = _Bfe(OpSize::i64Bit, 11, 52, Gpr);
ExpNZ = Sub(OpSize::i64Bit, ExpNZ, Constant(1023));
@@ -389,15 +407,27 @@ void OpDispatchBuilder::X87FXTRACTF64(OpcodeArgs) {
SigNZ = _Or(OpSize::i64Bit, SigNZ, Constant(0x3ff0'0000'0000'0000LL));
Ref SigNZV = _VCastFromGPR(OpSize::i64Bit, OpSize::i64Bit, SigNZ);
// Comparison and select to push onto stack
SaveNZCV();
// Mantissa non-zero => NaN (exp result = input); else Inf (exp result = +Inf)
Ref Mantissa = _And(OpSize::i64Bit, Gpr, Constant(0x000f'ffff'ffff'ffffULL));
_TestNZ(OpSize::i64Bit, Mantissa, Constant(~0ULL));
Ref ExpInfV = _NZCVSelectV(OpSize::i64Bit, CondClass::EQ, ExpInfOnlyV, ExpNanV);
// Biased exponent == 0x7ff => Inf/NaN path, else non-zero-case.
Ref BiasedExp = _Bfe(OpSize::i64Bit, 11, 52, Gpr);
SubWithFlags(OpSize::i64Bit, BiasedExp, 0x7ff);
Ref ExpNZOrInf = _NZCVSelectV(OpSize::i64Bit, CondClass::EQ, ExpInfV, ExpNZV);
Ref SigNZOrInf = _NZCVSelectV(OpSize::i64Bit, CondClass::EQ, SigInfV, SigNZV);
// Zero folds on top.
_TestNZ(OpSize::i64Bit, Gpr, Constant(0x7fff'ffff'ffff'ffffUL));
Ref Sig = _NZCVSelectV(OpSize::i64Bit, CondClass::EQ, SigZV, SigNZV);
Ref Exp = _NZCVSelectV(OpSize::i64Bit, CondClass::EQ, ExpZV, ExpNZV);
Ref Sig = _NZCVSelectV(OpSize::i64Bit, CondClass::EQ, SigZV, SigNZOrInf);
Ref Exp = _NZCVSelectV(OpSize::i64Bit, CondClass::EQ, ExpZV, ExpNZOrInf);
_PopStackDestroy();
_PushStack(Exp, Invalid(), OpSize::i64Bit);
_PushStack(Sig, Invalid(), OpSize::i64Bit);
_PushStack(Exp, Invalid(), OpSize::iInvalid);
_PushStack(Sig, Invalid(), OpSize::iInvalid);
}
} // namespace FEXCore::IR
@@ -1,89 +0,0 @@
// SPDX-License-Identifier: MIT
/*
$info$
tags: glue|x86-guest-code
desc: Guest-side assembly helpers used by the backends
$end_info$
*/
#include "Interface/Core/X86HelperGen.h"
#include "FEXCore/Utils/AllocatorHooks.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/Utils/Allocator.h>
#include <FEXHeaderUtils/Syscalls.h>
#include <cstdint>
#include <cstring>
namespace FEXCore {
constexpr size_t CODE_SIZE = 0x1000;
X86GeneratedCode::X86GeneratedCode() {
#ifdef _WIN32
// No need to allocate anything in this config.
#else
// Allocate a page for our emulated guest
CodePtr = AllocateGuestCodeSpace(CODE_SIZE);
constexpr std::array<uint8_t, 2> SignalReturnCode = {
0x0F, 0x3E, // CALLBACKRET FEX Instruction
};
CallbackReturn = reinterpret_cast<uint64_t>(CodePtr);
memcpy(reinterpret_cast<void*>(CallbackReturn), SignalReturnCode.data(), SignalReturnCode.size());
mprotect(CodePtr, CODE_SIZE, PROT_READ);
#endif
}
X86GeneratedCode::~X86GeneratedCode() {
#ifndef _WIN32
FEXCore::Allocator::VirtualFree(CodePtr, CODE_SIZE);
#endif
}
void* X86GeneratedCode::AllocateGuestCodeSpace(size_t Size) {
#ifndef _WIN32
FEX_CONFIG_OPT(Is64BitMode, IS64BIT_MODE);
if (Is64BitMode()) {
// 64bit mode can have its sigret handler anywhere
auto Result = FEXCore::Allocator::VirtualAlloc(Size);
FEXCore::Allocator::VirtualName("FEXMem_Misc", reinterpret_cast<void*>(Result), Size);
return Result;
}
// First 64bit page
constexpr uintptr_t LOCATION_MAX = 0x1'0000'0000;
// 32bit mode
// We need to have the sigret handler in the lower 32bits of memory space
// Scan top down and try to allocate a location
for (size_t Location = 0xFFFF'E000; Location != 0x0; Location -= 0x1000) {
void* Ptr = ::mmap(reinterpret_cast<void*>(Location), Size, PROT_READ | PROT_WRITE, MAP_FIXED_NOREPLACE | MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
if (Ptr != MAP_FAILED && reinterpret_cast<uintptr_t>(Ptr) >= LOCATION_MAX) {
// Failed to map in the lower 32bits
// Try again
// Can happen in the case that host kernel ignores MAP_FIXED_NOREPLACE
::munmap(Ptr, Size);
continue;
}
if (Ptr != MAP_FAILED) {
return Ptr;
}
}
// Can't do anything about this
// Here's hoping the application doesn't use signals
return MAP_FAILED;
#else
return nullptr;
#endif
}
} // namespace FEXCore
@@ -1,25 +0,0 @@
// SPDX-License-Identifier: MIT
/*
$info$
tags: glue|x86-guest-code
$end_info$
*/
#pragma once
#include <stddef.h>
#include <stdint.h>
namespace FEXCore {
class X86GeneratedCode final {
public:
X86GeneratedCode();
~X86GeneratedCode();
uint64_t CallbackReturn {};
private:
void* CodePtr {};
void* AllocateGuestCodeSpace(size_t Size);
};
} // namespace FEXCore
@@ -124,7 +124,7 @@ constexpr std::array<X86InstInfo[2], ENTRY_MAX> Primary_ArchSelect_LUT = {{
},
// ENTRY_63
{
{"ARPL", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
{"ARPL", TYPE_INST, GenFlagsSameSize(SIZE_16BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, { .OpDispatch = &IR::OpDispatchBuilder::ARPLOp } },
{"MOVSXD", TYPE_INST, GenFlagsDstSize(SIZE_64BIT) | FLAGS_MODRM, 0, { .OpDispatch = &IR::OpDispatchBuilder::MOVSXDOp } },
},
// ENTRY_9A
@@ -436,4 +436,3 @@ const std::array<X86InstInfo, MAX_PRIMARY_TABLE_SIZE> BaseOps = []() consteval {
}();
}
@@ -50,7 +50,7 @@ constexpr std::array<X86InstInfo[2], ENTRY_MAX> SecondGroup_ArchSelect_LUT = {{
},
}};
constexpr auto SecondInstGroupOps = []() consteval {
constexpr std::array<X86InstInfo, MAX_INST_SECOND_GROUP_TABLE_SIZE> SecondInstGroupOps = []() consteval {
std::array<X86InstInfo, MAX_INST_SECOND_GROUP_TABLE_SIZE> Table{};
constexpr U16U8InfoStruct SecondaryExtensionOpTable[] = {
// GROUP 1
@@ -402,37 +402,37 @@ constexpr auto SecondInstGroupOps = []() consteval {
// GROUP 16
// AMD documentation claims again that this entire group is n/a to prefix
// Tooling once again fails to disassemble oens with the prefix. Disable until proven otherwise
{OPD(TYPE_GROUP_16, PF_NONE, 0), 1, X86InstInfo{"PREFETCH NTA", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_NONE, 1), 1, X86InstInfo{"PREFETCH T0", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_NONE, 2), 1, X86InstInfo{"PREFETCH T1", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_NONE, 3), 1, X86InstInfo{"PREFETCH T2", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_NONE, 0), 1, X86InstInfo{"PREFETCH NTA", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_NONE, 1), 1, X86InstInfo{"PREFETCH T0", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_NONE, 2), 1, X86InstInfo{"PREFETCH T1", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_NONE, 3), 1, X86InstInfo{"PREFETCH T2", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_NONE, 4), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_NONE, 5), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_NONE, 6), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_NONE, 7), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 0), 1, X86InstInfo{"PREFETCH NTA", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 1), 1, X86InstInfo{"PREFETCH T0", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 2), 1, X86InstInfo{"PREFETCH T1", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 3), 1, X86InstInfo{"PREFETCH T2", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 0), 1, X86InstInfo{"PREFETCH NTA", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 1), 1, X86InstInfo{"PREFETCH T0", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 2), 1, X86InstInfo{"PREFETCH T1", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 3), 1, X86InstInfo{"PREFETCH T2", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 4), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 5), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 6), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F3, 7), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_66, 0), 1, X86InstInfo{"PREFETCH NTA", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_66, 1), 1, X86InstInfo{"PREFETCH T0", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_66, 2), 1, X86InstInfo{"PREFETCH T1", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_66, 3), 1, X86InstInfo{"PREFETCH T2", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_66, 0), 1, X86InstInfo{"PREFETCH NTA", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_66, 1), 1, X86InstInfo{"PREFETCH T0", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_66, 2), 1, X86InstInfo{"PREFETCH T1", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_66, 3), 1, X86InstInfo{"PREFETCH T2", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_66, 4), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_66, 5), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_66, 6), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_66, 7), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F2, 0), 1, X86InstInfo{"PREFETCH NTA", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_F2, 1), 1, X86InstInfo{"PREFETCH T0", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_F2, 2), 1, X86InstInfo{"PREFETCH T1", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_F2, 3), 1, X86InstInfo{"PREFETCH T2", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY, 0}},
{OPD(TYPE_GROUP_16, PF_F2, 0), 1, X86InstInfo{"PREFETCH NTA", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F2, 1), 1, X86InstInfo{"PREFETCH T0", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F2, 2), 1, X86InstInfo{"PREFETCH T1", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F2, 3), 1, X86InstInfo{"PREFETCH T2", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F2, 4), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F2, 5), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
{OPD(TYPE_GROUP_16, PF_F2, 6), 1, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM, 0}},
@@ -31,19 +31,19 @@ constexpr std::array<X86InstInfo[2], ENTRY_MAX> Secondary_ArchSelect_LUT = {{
},
{
{"PUSH FS", TYPE_INST, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_FS_PREFIX> } },
{"PUSH FS", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_FS_PREFIX> } },
{"PUSH FS", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_FS_PREFIX> } },
},
{
{"POP FS", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_DEF) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::POPSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_FS_PREFIX> } },
{"POP FS", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_64BIT) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::POPSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_FS_PREFIX> } },
{"POP FS", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_64BITDEF) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::POPSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_FS_PREFIX> } },
},
{
{"PUSH GS", TYPE_INST, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_GS_PREFIX> } },
{"PUSH GS", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_GS_PREFIX> } },
{"PUSH GS", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_GS_PREFIX> } },
},
{
{"POP GS", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_DEF) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::POPSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_GS_PREFIX> } },
{"POP GS", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_64BIT) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::POPSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_GS_PREFIX> } },
{"POP GS", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_64BITDEF) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::POPSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_GS_PREFIX> } },
},
}};
@@ -24,92 +24,98 @@ namespace FEXCore::X86Tables {
struct X86InstInfo;
namespace DecodeFlags {
constexpr uint32_t FLAG_OPERAND_SIZE = (1 << 0);
constexpr uint32_t FLAG_ADDRESS_SIZE = (1 << 1);
constexpr uint32_t FLAG_LOCK = (1 << 2);
constexpr uint32_t FLAG_LEGACY_PREFIX = (1 << 3);
constexpr uint32_t FLAG_REX_PREFIX = (1 << 4);
constexpr uint32_t FLAG_VSIB_BYTE = (1 << 5);
constexpr uint32_t FLAG_OPTION_AVX_W = (1 << 6);
constexpr uint32_t FLAG_REX_WIDENING = (1 << 7);
constexpr uint32_t FLAG_REX_XGPR_B = (1 << 8);
constexpr uint32_t FLAG_REX_XGPR_X = (1 << 9);
constexpr uint32_t FLAG_REX_XGPR_R = (1 << 10);
constexpr uint32_t FLAG_NO_PREFIX = (0b000 << 11);
constexpr uint32_t FLAG_ES_PREFIX = (0b001 << 11);
constexpr uint32_t FLAG_CS_PREFIX = (0b010 << 11);
constexpr uint32_t FLAG_SS_PREFIX = (0b011 << 11);
constexpr uint32_t FLAG_DS_PREFIX = (0b100 << 11);
constexpr uint32_t FLAG_FS_PREFIX = (0b101 << 11);
constexpr uint32_t FLAG_GS_PREFIX = (0b110 << 11);
constexpr uint32_t FLAG_SEGMENTS = (0b111 << 11);
constexpr uint32_t FLAG_FORCE_TSO = (1 << 14);
constexpr uint32_t FLAG_DECODED_MODRM = (1 << 15);
constexpr uint32_t FLAG_DECODED_SIB = (1 << 16);
constexpr uint32_t FLAG_REP_PREFIX = (1 << 17);
constexpr uint32_t FLAG_REPNE_PREFIX = (1 << 18);
// Size flags
constexpr uint32_t FLAG_SIZE_DST_OFF = 19;
constexpr uint32_t FLAG_SIZE_SRC_OFF = FLAG_SIZE_DST_OFF + 3;
constexpr uint32_t SIZE_MASK = 0b111;
constexpr uint32_t SIZE_DEF = 0b000; // This should be invalid past decoding
constexpr uint32_t SIZE_8BIT = 0b001;
constexpr uint32_t SIZE_16BIT = 0b010;
constexpr uint32_t SIZE_32BIT = 0b011;
constexpr uint32_t SIZE_64BIT = 0b100;
constexpr uint32_t SIZE_128BIT = 0b101;
constexpr uint32_t SIZE_256BIT = 0b110;
constexpr uint32_t FLAG_OPERAND_SIZE = (1 << 0);
constexpr uint32_t FLAG_ADDRESS_SIZE = (1 << 1);
constexpr uint32_t FLAG_LOCK = (1 << 2);
constexpr uint32_t FLAG_LEGACY_PREFIX = (1 << 3);
constexpr uint32_t FLAG_REX_PREFIX = (1 << 4);
constexpr uint32_t FLAG_VSIB_BYTE = (1 << 5);
constexpr uint32_t FLAG_OPTION_AVX_W = (1 << 6);
constexpr uint32_t FLAG_REX_WIDENING = (1 << 7);
constexpr uint32_t FLAG_REX_XGPR_B = (1 << 8);
constexpr uint32_t FLAG_REX_XGPR_X = (1 << 9);
constexpr uint32_t FLAG_REX_XGPR_R = (1 << 10);
constexpr uint32_t FLAG_NO_PREFIX = (0b000 << 11);
constexpr uint32_t FLAG_ES_PREFIX = (0b001 << 11);
constexpr uint32_t FLAG_CS_PREFIX = (0b010 << 11);
constexpr uint32_t FLAG_SS_PREFIX = (0b011 << 11);
constexpr uint32_t FLAG_DS_PREFIX = (0b100 << 11);
constexpr uint32_t FLAG_FS_PREFIX = (0b101 << 11);
constexpr uint32_t FLAG_GS_PREFIX = (0b110 << 11);
constexpr uint32_t FLAG_SEGMENTS = (0b111 << 11);
constexpr uint32_t FLAG_FORCE_TSO = (1 << 14);
constexpr uint32_t FLAG_DECODED_MODRM = (1 << 15);
constexpr uint32_t FLAG_DECODED_SIB = (1 << 16);
constexpr uint32_t FLAG_REP_PREFIX = (1 << 17);
constexpr uint32_t FLAG_REPNE_PREFIX = (1 << 18);
// Size flags
constexpr uint32_t FLAG_SIZE_DST_OFF = 19;
constexpr uint32_t FLAG_SIZE_SRC_OFF = FLAG_SIZE_DST_OFF + 3;
constexpr uint32_t SIZE_MASK = 0b111;
constexpr uint32_t SIZE_DEF = 0b000; // This should be invalid past decoding
constexpr uint32_t SIZE_8BIT = 0b001;
constexpr uint32_t SIZE_16BIT = 0b010;
constexpr uint32_t SIZE_32BIT = 0b011;
constexpr uint32_t SIZE_64BIT = 0b100;
constexpr uint32_t SIZE_128BIT = 0b101;
constexpr uint32_t SIZE_256BIT = 0b110;
constexpr uint32_t FLAG_OPADDR_OFF = (FLAG_SIZE_SRC_OFF + 3);
constexpr uint32_t FLAG_OPADDR_STACKSIZE = 4; // Two level deep stack
constexpr uint32_t FLAG_OPADDR_FLAG_SIZE = 2;
constexpr uint32_t FLAG_OPADDR_MASK = (((1 << FLAG_OPADDR_STACKSIZE) - 1) << FLAG_OPADDR_OFF);
constexpr uint32_t FLAG_OPADDR_OFF = (FLAG_SIZE_SRC_OFF + 3);
constexpr uint32_t FLAG_OPADDR_STACKSIZE = 4; // Two level deep stack
constexpr uint32_t FLAG_OPADDR_FLAG_SIZE = 2;
constexpr uint32_t FLAG_OPADDR_MASK = (((1 << FLAG_OPADDR_STACKSIZE) - 1) << FLAG_OPADDR_OFF);
// 00 = NONE
constexpr uint32_t FLAG_OPERAND_SIZE_LAST = 0b01;
constexpr uint32_t FLAG_WIDENING_SIZE_LAST = 0b10;
// 00 = NONE
constexpr uint32_t FLAG_OPERAND_SIZE_LAST = 0b01;
constexpr uint32_t FLAG_WIDENING_SIZE_LAST = 0b10;
constexpr uint32_t GetSizeDstFlags(uint32_t Flags) { return (Flags >> FLAG_SIZE_DST_OFF) & SIZE_MASK; }
constexpr uint32_t GetSizeSrcFlags(uint32_t Flags) { return (Flags >> FLAG_SIZE_SRC_OFF) & SIZE_MASK; }
constexpr uint32_t GenSizeDstSize(uint32_t Size) { return Size << FLAG_SIZE_DST_OFF; }
constexpr uint32_t GenSizeSrcSize(uint32_t Size) { return Size << FLAG_SIZE_SRC_OFF; }
constexpr uint32_t GetOpAddr(uint32_t Flags, uint32_t Index) {
return (((Flags & FLAG_OPADDR_MASK) >> FLAG_OPADDR_OFF) >> (Index * 2)) & ((1 << FLAG_OPADDR_FLAG_SIZE) - 1);
}
inline void PushOpAddr(uint32_t *Flags, uint32_t Flag) {
uint32_t TmpFlags = *Flags;
uint32_t BottomOfStack = ((TmpFlags & FLAG_OPADDR_MASK) >> FLAG_OPADDR_OFF) & ((1 << FLAG_OPADDR_FLAG_SIZE) - 1);
TmpFlags &= ~(FLAG_OPADDR_MASK);
TmpFlags |=
(BottomOfStack << (FLAG_OPADDR_OFF + FLAG_OPADDR_FLAG_SIZE)) |
(Flag << FLAG_OPADDR_OFF);
*Flags = TmpFlags;
}
inline void PopOpAddrIf(uint32_t *Flags, uint32_t Flag) {
uint32_t TmpFlags = *Flags;
uint32_t BottomOfStack = ((TmpFlags & FLAG_OPADDR_MASK) >> FLAG_OPADDR_OFF) & ((1 << FLAG_OPADDR_FLAG_SIZE) - 1);
// Only pop the stack if the bottom flag is the one we care about
// Necessary for escape prefixes that overlap regular prefixes
if (BottomOfStack != Flag) {
return;
constexpr uint32_t GetSizeDstFlags(uint32_t Flags) {
return (Flags >> FLAG_SIZE_DST_OFF) & SIZE_MASK;
}
constexpr uint32_t GetSizeSrcFlags(uint32_t Flags) {
return (Flags >> FLAG_SIZE_SRC_OFF) & SIZE_MASK;
}
uint32_t TopOfStack = ((TmpFlags & FLAG_OPADDR_MASK) >> (FLAG_OPADDR_OFF + FLAG_OPADDR_FLAG_SIZE)) & ((1 << FLAG_OPADDR_FLAG_SIZE) - 1);
constexpr uint32_t GenSizeDstSize(uint32_t Size) {
return Size << FLAG_SIZE_DST_OFF;
}
constexpr uint32_t GenSizeSrcSize(uint32_t Size) {
return Size << FLAG_SIZE_SRC_OFF;
}
TmpFlags &= ~(FLAG_OPADDR_MASK);
TmpFlags |= (TopOfStack << FLAG_OPADDR_OFF);
constexpr uint32_t GetOpAddr(uint32_t Flags, uint32_t Index) {
return (((Flags & FLAG_OPADDR_MASK) >> FLAG_OPADDR_OFF) >> (Index * 2)) & ((1 << FLAG_OPADDR_FLAG_SIZE) - 1);
}
*Flags = TmpFlags;
}
inline void PushOpAddr(uint32_t* Flags, uint32_t Flag) {
uint32_t TmpFlags = *Flags;
uint32_t BottomOfStack = ((TmpFlags & FLAG_OPADDR_MASK) >> FLAG_OPADDR_OFF) & ((1 << FLAG_OPADDR_FLAG_SIZE) - 1);
}
TmpFlags &= ~(FLAG_OPADDR_MASK);
TmpFlags |= (BottomOfStack << (FLAG_OPADDR_OFF + FLAG_OPADDR_FLAG_SIZE)) | (Flag << FLAG_OPADDR_OFF);
*Flags = TmpFlags;
}
inline void PopOpAddrIf(uint32_t* Flags, uint32_t Flag) {
uint32_t TmpFlags = *Flags;
uint32_t BottomOfStack = ((TmpFlags & FLAG_OPADDR_MASK) >> FLAG_OPADDR_OFF) & ((1 << FLAG_OPADDR_FLAG_SIZE) - 1);
// Only pop the stack if the bottom flag is the one we care about
// Necessary for escape prefixes that overlap regular prefixes
if (BottomOfStack != Flag) {
return;
}
uint32_t TopOfStack = ((TmpFlags & FLAG_OPADDR_MASK) >> (FLAG_OPADDR_OFF + FLAG_OPADDR_FLAG_SIZE)) & ((1 << FLAG_OPADDR_FLAG_SIZE) - 1);
TmpFlags &= ~(FLAG_OPADDR_MASK);
TmpFlags |= (TopOfStack << FLAG_OPADDR_OFF);
*Flags = TmpFlags;
}
} // namespace DecodeFlags
struct DecodedOperand {
enum class OpType : uint8_t {
@@ -117,9 +123,13 @@ struct DecodedOperand {
GPR,
GPRDirect,
GPRIndirect,
GPRIndirectRelocation,
RIPRelative,
RIPRelativeRelocation,
Literal,
LiteralRelocation,
SIB,
SIBRelocation
};
bool IsNone() const {
@@ -134,20 +144,30 @@ struct DecodedOperand {
bool IsGPRIndirect() const {
return Type == OpType::GPRIndirect;
}
bool IsGPRIndirectRelocation() const {
return Type == OpType::GPRIndirectRelocation;
}
bool IsRIPRelative() const {
return Type == OpType::RIPRelative;
}
bool IsRIPRelativeRelocation() const {
return Type == OpType::RIPRelativeRelocation;
}
bool IsLiteral() const {
return Type == OpType::Literal;
}
bool IsLiteralRelocation() const {
return Type == OpType::LiteralRelocation;
}
bool IsSIB() const {
return Type == OpType::SIB;
}
bool IsSIBRelocation() const {
return Type == OpType::SIBRelocation;
}
uint64_t Literal() const {
LOGMAN_THROW_A_FMT(IsLiteral(), "Precondition: must be a literal");
if (Data.Literal.SignExtend) {
return static_cast<int64_t>(static_cast<int32_t>(Data.Literal.Value));
}
return Data.Literal.Value;
}
@@ -159,30 +179,29 @@ struct DecodedOperand {
} GPR;
struct {
int32_t Displacement;
int64_t Displacement;
uint8_t GPR;
} GPRIndirect;
} GPRIndirect; // Shared with GPRIndirectRelocation
struct {
union {
int32_t s;
uint32_t u;
} Value;
} RIPLiteral;
int64_t Value;
} RIPLiteral; // Shared with RIPLiteralRelocation
struct LiteralType {
uint32_t Value;
uint8_t Size : 7 ;
bool SignExtend : 1;
auto operator<=>(const LiteralType&) const = default;
uint64_t Value;
uint8_t Size;
} Literal;
struct {
int32_t Offset;
int64_t EntrypointOffset;
} LiteralRelocation;
struct {
int64_t Offset;
uint8_t Scale;
uint8_t Index; // ~0 invalid
uint8_t Base; // ~0 invalid
} SIB;
uint8_t Base; // ~0 invalid
} SIB; // Shared with SIBRelocation
};
TypeUnion Data;
@@ -196,7 +215,7 @@ struct DecodedInst {
DecodedOperand Src[3];
// Constains the dispatcher handler pointer
X86InstInfo const* TableInfo;
const X86InstInfo* TableInfo;
uint32_t Flags;
uint16_t OP;
@@ -205,21 +224,22 @@ struct DecodedInst {
uint8_t ModRM;
uint8_t SIB;
uint8_t InstSize;
int8_t REXIndex;
};
union ModRMDecoded {
uint8_t Hex{};
uint8_t Hex {};
struct {
uint8_t rm : 3;
uint8_t rm : 3;
uint8_t reg : 3;
uint8_t mod : 2;
};
};
union SIBDecoded {
uint8_t Hex{};
uint8_t Hex {};
struct {
uint8_t base : 3;
uint8_t base : 3;
uint8_t index : 3;
uint8_t scale : 2;
};
@@ -290,123 +310,136 @@ enum InstType {
namespace InstFlags {
using InstFlagType = uint64_t;
using InstFlagType = uint64_t;
constexpr InstFlagType FLAGS_NONE = 0;
// The secondary Opcode Map uses prefix bytes to overlay more instruction
// But some instructions need to ignore this overlay and consume these prefixes.
constexpr InstFlagType FLAGS_NO_OVERLAY = (1ULL << 0);
// Some instructions partially ignore overlay
// Ignore OpSize (0x66) in this case
constexpr InstFlagType FLAGS_NO_OVERLAY66 = (1ULL << 1);
constexpr InstFlagType FLAGS_DEBUG_MEM_ACCESS = (1ULL << 2);
// Only SEXT if the instruction is operating in 64bit operand size
constexpr InstFlagType FLAGS_SRC_SEXT64BIT = (1ULL << 3);
constexpr InstFlagType FLAGS_BLOCK_END = (1ULL << 4);
constexpr InstFlagType FLAGS_SETS_RIP = (1ULL << 5);
constexpr InstFlagType FLAGS_NONE = 0;
// The secondary Opcode Map uses prefix bytes to overlay more instruction
// But some instructions need to ignore this overlay and consume these prefixes.
constexpr InstFlagType FLAGS_NO_OVERLAY = (1ULL << 0);
// Some instructions partially ignore overlay
// Ignore OpSize (0x66) in this case
constexpr InstFlagType FLAGS_NO_OVERLAY66 = (1ULL << 1);
constexpr InstFlagType FLAGS_DEBUG_MEM_ACCESS = (1ULL << 2);
// Only SEXT if the instruction is operating in 64bit operand size
constexpr InstFlagType FLAGS_SRC_SEXT64BIT = (1ULL << 3);
constexpr InstFlagType FLAGS_BLOCK_END = (1ULL << 4);
constexpr InstFlagType FLAGS_SETS_RIP = (1ULL << 5);
constexpr InstFlagType FLAGS_DISPLACE_SIZE_MUL_2 = (1ULL << 6);
constexpr InstFlagType FLAGS_DISPLACE_SIZE_DIV_2 = (1ULL << 7);
constexpr InstFlagType FLAGS_SRC_SEXT = (1ULL << 8);
constexpr InstFlagType FLAGS_MEM_OFFSET = (1ULL << 9);
constexpr InstFlagType FLAGS_DISPLACE_SIZE_MUL_2 = (1ULL << 6);
constexpr InstFlagType FLAGS_DISPLACE_SIZE_DIV_2 = (1ULL << 7);
constexpr InstFlagType FLAGS_SRC_SEXT = (1ULL << 8);
constexpr InstFlagType FLAGS_MEM_OFFSET = (1ULL << 9);
// Enables XMM based subflags
// Current reserved range for this SF is [10, 15]
constexpr InstFlagType FLAGS_XMM_FLAGS = (1ULL << 10);
// Enables XMM based subflags
// Current reserved range for this SF is [10, 15]
constexpr InstFlagType FLAGS_XMM_FLAGS = (1ULL << 10);
// X87 flags aliased to XMM flags selection
// Allows X87 instruction table that is abusing the flag for 64BIT selection to work
constexpr InstFlagType FLAGS_X87_FLAGS = (1ULL << 10);
// X87 flags aliased to XMM flags selection
// Allows X87 instruction table that is abusing the flag for 64BIT selection to work
constexpr InstFlagType FLAGS_X87_FLAGS = (1ULL << 10);
// Non-XMM subflags
constexpr InstFlagType FLAGS_SF_DST_RAX = (1ULL << 11);
constexpr InstFlagType FLAGS_SF_DST_RDX = (1ULL << 12);
constexpr InstFlagType FLAGS_SF_SRC_RAX = (1ULL << 13);
constexpr InstFlagType FLAGS_SF_SRC_RCX = (1ULL << 14);
constexpr InstFlagType FLAGS_SF_REX_IN_BYTE = (1ULL << 15);
constexpr InstFlagType FLAGS_SF_DST_RAX = (1ULL << 11);
constexpr InstFlagType FLAGS_SF_DST_RDX = (1ULL << 12);
constexpr InstFlagType FLAGS_SF_SRC_RAX = (1ULL << 13);
constexpr InstFlagType FLAGS_SF_SRC_RCX = (1ULL << 14);
constexpr InstFlagType FLAGS_SF_REX_IN_BYTE = (1ULL << 15);
// XMM subflags
constexpr InstFlagType FLAGS_SF_UNUSED = (1ULL << 11); // No assigned behavior yet
constexpr InstFlagType FLAGS_SF_DST_GPR = (1ULL << 12);
constexpr InstFlagType FLAGS_SF_SRC_GPR = (1ULL << 13);
constexpr InstFlagType FLAGS_SF_MMX_DST = (1ULL << 14);
constexpr InstFlagType FLAGS_SF_MMX_SRC = (1ULL << 15);
constexpr InstFlagType FLAGS_SF_MMX = FLAGS_SF_MMX_DST | FLAGS_SF_MMX_SRC;
constexpr InstFlagType FLAGS_SF_UNUSED = (1ULL << 11); // No assigned behavior yet
constexpr InstFlagType FLAGS_SF_DST_GPR = (1ULL << 12);
constexpr InstFlagType FLAGS_SF_SRC_GPR = (1ULL << 13);
constexpr InstFlagType FLAGS_SF_MMX_DST = (1ULL << 14);
constexpr InstFlagType FLAGS_SF_MMX_SRC = (1ULL << 15);
constexpr InstFlagType FLAGS_SF_MMX = FLAGS_SF_MMX_DST | FLAGS_SF_MMX_SRC;
// Enables MODRM specific subflags
// Current reserved range for this SF is [14, 17]
constexpr InstFlagType FLAGS_MODRM = (1ULL << 16);
// Enables MODRM specific subflags
// Current reserved range for this SF is [14, 17]
constexpr InstFlagType FLAGS_MODRM = (1ULL << 16);
// With ModRM SF flag enabled
// Direction of ModRM. Dst ^ Src
// Set means destination is rm bits
// Unset means src is rm bits
constexpr InstFlagType FLAGS_SF_MOD_DST = (1ULL << 17);
constexpr InstFlagType FLAGS_SF_MOD_DST = (1ULL << 17);
// If the instruction is restricted to mem or reg only
// 0b00 = Regular ModRM support
// 0b01 = Memory accesses only
// 0b10 = Register accesses only
// 0b11 = <Reserved>
constexpr InstFlagType FLAGS_SF_MOD_MEM_ONLY = (1ULL << 18);
constexpr InstFlagType FLAGS_SF_MOD_REG_ONLY = (1ULL << 19);
constexpr InstFlagType FLAGS_SF_MOD_MEM_ONLY = (1ULL << 18);
constexpr InstFlagType FLAGS_SF_MOD_REG_ONLY = (1ULL << 19);
constexpr InstFlagType FLAGS_SF_MOD_ZERO_REG = (1ULL << 20);
constexpr InstFlagType FLAGS_SF_MOD_ZERO_REG = (1ULL << 20);
// x87
constexpr InstFlagType FLAGS_POP = (1ULL << 21);
// x87
constexpr InstFlagType FLAGS_POP = (1ULL << 21);
// Whether or not the instruction has a VEX prefix for the dest, first, or second source.
constexpr InstFlagType FLAGS_VEX_SRC_MASK = (0b11ULL << 22);
constexpr InstFlagType FLAGS_VEX_NO_OPERAND = (0b00ULL << 22);
constexpr InstFlagType FLAGS_VEX_DST = (0b01ULL << 22);
constexpr InstFlagType FLAGS_VEX_1ST_SRC = (0b10ULL << 22);
constexpr InstFlagType FLAGS_VEX_2ND_SRC = (0b11ULL << 22);
// Whether or not the instruction has a VSIB byte
constexpr InstFlagType FLAGS_VEX_VSIB = (1ULL << 24);
constexpr InstFlagType FLAGS_VEX_L_IGNORE = (1ULL << 25);
constexpr InstFlagType FLAGS_VEX_L_0 = (1ULL << 26);
constexpr InstFlagType FLAGS_VEX_L_1 = (1ULL << 27);
// Whether or not the instruction has a VEX prefix for the dest, first, or second source.
constexpr InstFlagType FLAGS_VEX_SRC_MASK = (0b11ULL << 22);
constexpr InstFlagType FLAGS_VEX_NO_OPERAND = (0b00ULL << 22);
constexpr InstFlagType FLAGS_VEX_DST = (0b01ULL << 22);
constexpr InstFlagType FLAGS_VEX_1ST_SRC = (0b10ULL << 22);
constexpr InstFlagType FLAGS_VEX_2ND_SRC = (0b11ULL << 22);
// Whether or not the instruction has a VSIB byte
constexpr InstFlagType FLAGS_VEX_VSIB = (1ULL << 24);
constexpr InstFlagType FLAGS_VEX_L_IGNORE = (1ULL << 25);
constexpr InstFlagType FLAGS_VEX_L_0 = (1ULL << 26);
constexpr InstFlagType FLAGS_VEX_L_1 = (1ULL << 27);
constexpr InstFlagType FLAGS_REX_W_0 = (1ULL << 28);
constexpr InstFlagType FLAGS_REX_W_1 = (1ULL << 29);
constexpr InstFlagType FLAGS_REX_W_0 = (1ULL << 28);
constexpr InstFlagType FLAGS_REX_W_1 = (1ULL << 29);
constexpr InstFlagType FLAGS_CALL = (1ULL << 30);
constexpr InstFlagType FLAGS_CALL = (1ULL << 30);
constexpr InstFlagType FLAGS_SIZE_DST_OFF = 58;
constexpr InstFlagType FLAGS_SIZE_SRC_OFF = FLAGS_SIZE_DST_OFF + 3;
constexpr InstFlagType FLAGS_SIZE_DST_OFF = 58;
constexpr InstFlagType FLAGS_SIZE_SRC_OFF = FLAGS_SIZE_DST_OFF + 3;
constexpr InstFlagType SIZE_MASK = 0b111;
constexpr InstFlagType SIZE_DEF = 0b000;
constexpr InstFlagType SIZE_8BIT = 0b001;
constexpr InstFlagType SIZE_16BIT = 0b010;
constexpr InstFlagType SIZE_32BIT = 0b011;
constexpr InstFlagType SIZE_64BIT = 0b100;
constexpr InstFlagType SIZE_128BIT = 0b101;
constexpr InstFlagType SIZE_256BIT = 0b110;
constexpr InstFlagType SIZE_64BITDEF = 0b111; // Default mode is 64bit instead of typical 32bit
constexpr InstFlagType SIZE_MASK = 0b111;
constexpr InstFlagType SIZE_DEF = 0b000;
constexpr InstFlagType SIZE_8BIT = 0b001;
constexpr InstFlagType SIZE_16BIT = 0b010;
constexpr InstFlagType SIZE_32BIT = 0b011;
constexpr InstFlagType SIZE_64BIT = 0b100;
constexpr InstFlagType SIZE_128BIT = 0b101;
constexpr InstFlagType SIZE_256BIT = 0b110;
constexpr InstFlagType SIZE_64BITDEF = 0b111; // Default mode is 64bit instead of typical 32bit
#ifndef _WIN32
constexpr uint32_t DEFAULT_SYSCALL_FLAGS = FLAGS_NO_OVERLAY;
#else
// Syscall ends a block on WIN32 because the instruction can update the CPU's RIP.
// Syscall ends a block on WIN32 because the instruction can update the CPU's RIP.
constexpr uint32_t DEFAULT_SYSCALL_FLAGS = FLAGS_NO_OVERLAY | FLAGS_BLOCK_END;
#endif
constexpr InstFlagType GetSizeDstFlags(InstFlagType Flags) { return (Flags >> FLAGS_SIZE_DST_OFF) & SIZE_MASK; }
constexpr InstFlagType GetSizeSrcFlags(InstFlagType Flags) { return (Flags >> FLAGS_SIZE_SRC_OFF) & SIZE_MASK; }
constexpr InstFlagType GetSizeDstFlags(InstFlagType Flags) {
return (Flags >> FLAGS_SIZE_DST_OFF) & SIZE_MASK;
}
constexpr InstFlagType GetSizeSrcFlags(InstFlagType Flags) {
return (Flags >> FLAGS_SIZE_SRC_OFF) & SIZE_MASK;
}
constexpr InstFlagType GenFlagsDstSize(InstFlagType Size) { return Size << FLAGS_SIZE_DST_OFF; }
constexpr InstFlagType GenFlagsSrcSize(InstFlagType Size) { return Size << FLAGS_SIZE_SRC_OFF; }
constexpr InstFlagType GenFlagsSameSize(InstFlagType Size) { return (Size << FLAGS_SIZE_DST_OFF) | (Size << FLAGS_SIZE_SRC_OFF); }
constexpr InstFlagType GenFlagsSizes(InstFlagType Dest, InstFlagType Src) { return (Dest << FLAGS_SIZE_DST_OFF) | (Src << FLAGS_SIZE_SRC_OFF); }
constexpr InstFlagType GenFlagsDstSize(InstFlagType Size) {
return Size << FLAGS_SIZE_DST_OFF;
}
constexpr InstFlagType GenFlagsSrcSize(InstFlagType Size) {
return Size << FLAGS_SIZE_SRC_OFF;
}
constexpr InstFlagType GenFlagsSameSize(InstFlagType Size) {
return (Size << FLAGS_SIZE_DST_OFF) | (Size << FLAGS_SIZE_SRC_OFF);
}
constexpr InstFlagType GenFlagsSizes(InstFlagType Dest, InstFlagType Src) {
return (Dest << FLAGS_SIZE_DST_OFF) | (Src << FLAGS_SIZE_SRC_OFF);
}
// If it has an xmm subflag
#define HAS_XMM_SUBFLAG(x, flag) (((x) & (FEXCore::X86Tables::InstFlags::FLAGS_XMM_FLAGS | (flag))) == (FEXCore::X86Tables::InstFlags::FLAGS_XMM_FLAGS | (flag)))
#define HAS_XMM_SUBFLAG(x, flag) \
(((x) & (FEXCore::X86Tables::InstFlags::FLAGS_XMM_FLAGS | (flag))) == (FEXCore::X86Tables::InstFlags::FLAGS_XMM_FLAGS | (flag)))
// If it has non-xmm subflag
#define HAS_NON_XMM_SUBFLAG(x, flag) (((x) & (FEXCore::X86Tables::InstFlags::FLAGS_XMM_FLAGS | (flag))) == (flag))
}
} // namespace InstFlags
constexpr uint8_t OpToIndex(uint8_t Op) {
switch (Op) {
@@ -436,32 +469,30 @@ constexpr uint8_t OpToIndex(uint8_t Op) {
return 0;
}
using DecodedOp = DecodedInst const*;
using DecodedOp = const DecodedInst*;
using OpDispatchPtr = void (IR::OpDispatchBuilder::*)(DecodedOp);
union OpDispatchPtrWrapper {
OpDispatchPtr OpDispatch;
const struct X86InstInfo *Indirect;
const struct X86InstInfo* Indirect;
};
struct X86InstInfo {
char const *Name;
const char* Name;
InstType Type;
InstFlags::InstFlagType Flags; ///< Must be larger than InstFlags enum
uint8_t MoreBytes;
OpDispatchPtrWrapper OpcodeDispatcher;
bool operator==(const X86InstInfo &b) const {
if (strcmp(Name, b.Name) != 0 ||
Type != b.Type ||
Flags != b.Flags ||
MoreBytes != b.MoreBytes)
bool operator==(const X86InstInfo& b) const {
if (strcmp(Name, b.Name) != 0 || Type != b.Type || Flags != b.Flags || MoreBytes != b.MoreBytes) {
return false;
}
// We don't care if the opcode dispatcher differs
return true;
}
bool operator!=(const X86InstInfo &b) const {
bool operator!=(const X86InstInfo& b) const {
return !operator==(b);
}
};
@@ -513,7 +544,7 @@ extern const std::array<X86InstInfo, MAX_VEX_GROUP_TABLE_SIZE> VEXTableGroupOps;
extern const std::array<X86InstInfo, MAX_VEX_TABLE_SIZE> VEXTableOps_AVX128;
extern const std::array<X86InstInfo, MAX_VEX_GROUP_TABLE_SIZE> VEXTableGroupOps_AVX128;
template <typename OpcodeType>
template<typename OpcodeType>
struct X86TablesInfoStruct {
OpcodeType first;
uint8_t second;
@@ -523,11 +554,11 @@ using U8U8InfoStruct = X86TablesInfoStruct<uint8_t>;
using U16U8InfoStruct = X86TablesInfoStruct<uint16_t>;
template<typename OpcodeType>
constexpr static inline void GenerateTable(X86InstInfo *FinalTable, X86TablesInfoStruct<OpcodeType> const *LocalTable, size_t TableSize) {
constexpr static inline void GenerateTable(X86InstInfo* FinalTable, const X86TablesInfoStruct<OpcodeType>* LocalTable, size_t TableSize) {
for (size_t j = 0; j < TableSize; ++j) {
X86TablesInfoStruct<OpcodeType> const &Op = LocalTable[j];
const X86TablesInfoStruct<OpcodeType>& Op = LocalTable[j];
auto OpNum = Op.first;
X86InstInfo const &Info = Op.Info;
const X86InstInfo& Info = Op.Info;
for (uint32_t i = 0; i < Op.second; ++i) {
if (FinalTable[OpNum + i].Type != TYPE_UNKNOWN) {
LOGMAN_MSG_A_FMT("Duplicate Entry {}->{}", FinalTable[OpNum + i].Name, Info.Name);
@@ -541,19 +572,19 @@ constexpr static inline void GenerateTable(X86InstInfo *FinalTable, X86TablesInf
};
template<typename OpcodeType>
constexpr static inline void GenerateTableWithCopy(X86InstInfo *FinalTable, X86TablesInfoStruct<OpcodeType> const *LocalTable, size_t TableSize, const X86InstInfo *OtherLocal) {
constexpr static inline void GenerateTableWithCopy(X86InstInfo* FinalTable, const X86TablesInfoStruct<OpcodeType>* LocalTable,
size_t TableSize, const X86InstInfo* OtherLocal) {
for (size_t j = 0; j < TableSize; ++j) {
X86TablesInfoStruct<OpcodeType> const &Op = LocalTable[j];
const X86TablesInfoStruct<OpcodeType>& Op = LocalTable[j];
auto OpNum = Op.first;
X86InstInfo const &Info = Op.Info;
const X86InstInfo& Info = Op.Info;
for (uint32_t i = 0; i < Op.second; ++i) {
if (FinalTable[OpNum + i].Type != TYPE_UNKNOWN) {
LOGMAN_MSG_A_FMT("Duplicate Entry {}->{}", FinalTable[OpNum + i].Name, Info.Name);
}
if (Info.Type == TYPE_COPY_OTHER) {
FinalTable[OpNum + i] = OtherLocal[OpNum + i];
}
else {
} else {
FinalTable[OpNum + i] = Info;
}
}
@@ -561,11 +592,11 @@ constexpr static inline void GenerateTableWithCopy(X86InstInfo *FinalTable, X86T
};
template<typename OpcodeType>
constexpr static inline void GenerateX87Table(X86InstInfo *FinalTable, X86TablesInfoStruct<OpcodeType> const *LocalTable, size_t TableSize) {
constexpr static inline void GenerateX87Table(X86InstInfo* FinalTable, const X86TablesInfoStruct<OpcodeType>* LocalTable, size_t TableSize) {
for (size_t j = 0; j < TableSize; ++j) {
X86TablesInfoStruct<OpcodeType> const &Op = LocalTable[j];
const X86TablesInfoStruct<OpcodeType>& Op = LocalTable[j];
auto OpNum = Op.first;
X86InstInfo const &Info = Op.Info;
const X86InstInfo& Info = Op.Info;
for (uint32_t i = 0; i < Op.second; ++i) {
if (FinalTable[OpNum + i].Type != TYPE_UNKNOWN) {
LOGMAN_MSG_A_FMT("Duplicate Entry {}->{}", FinalTable[OpNum + i].Name, Info.Name);
@@ -573,8 +604,7 @@ constexpr static inline void GenerateX87Table(X86InstInfo *FinalTable, X86Tables
if ((OpNum & 0b11'000'000) == 0b11'000'000) {
// If the mod field is 0b11 then it is a regular op
FinalTable[OpNum + i] = Info;
}
else {
} else {
// If the mod field is !0b11 then this instruction is duplicated through the whole mod [0b00, 0b10] range
// and the modrm.rm space because that is used part of the instruction encoding
if ((OpNum & 0b11'000'000) != 0) {
@@ -17,7 +17,7 @@ using namespace IR;
// All OPDReg versions need it
#define OPDReg(op, reg) ((1 << 15) | ((op - 0xD8) << 8) | (reg << 3))
#define OPD(op, modrmop) (((op - 0xD8) << 8) | modrmop)
constexpr std::array<DispatchTableEntry, 133> X87F64OpTable = {{
constexpr std::array<DispatchTableEntry, 140> X87F64OpTable = {{
{OPDReg(0xD8, 0) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FADDF64, OpSize::i32Bit, false, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xD8, 1) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FMULF64, OpSize::i32Bit, false, OpDispatchBuilder::OpResult::RES_ST0>},
@@ -194,6 +194,10 @@ constexpr std::array<DispatchTableEntry, 133> X87F64OpTable = {{
{OPD(0xDC, 0xC0), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FADDF64, OpSize::f80Bit, false, OpDispatchBuilder::OpResult::RES_STI>},
{OPD(0xDC, 0xC8), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FMULF64, OpSize::f80Bit, false, OpDispatchBuilder::OpResult::RES_STI>},
{OPD(0xDC, 0xD0), 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FCOMIF64, OpSize::f80Bit, false, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, false>},
{OPD(0xDC, 0xD8), 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FCOMIF64, OpSize::f80Bit, false, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, false>},
{OPD(0xDC, 0xE0), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FSUBF64, OpSize::f80Bit, false, true, OpDispatchBuilder::OpResult::RES_STI>},
{OPD(0xDC, 0xE8), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FSUBF64, OpSize::f80Bit, false, false, OpDispatchBuilder::OpResult::RES_STI>},
{OPD(0xDC, 0xF0), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FDIVF64, OpSize::f80Bit, false, true, OpDispatchBuilder::OpResult::RES_STI>},
@@ -215,6 +219,7 @@ constexpr std::array<DispatchTableEntry, 133> X87F64OpTable = {{
{OPDReg(0xDD, 7) | 0x00, 8, &OpDispatchBuilder::X87FNSTSW},
{OPD(0xDD, 0xC0), 8, &OpDispatchBuilder::X87FFREE},
{OPD(0xDD, 0xC8), 8, &OpDispatchBuilder::FXCH},
{OPD(0xDD, 0xD0), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FSTToStack>}, // register-register from regular X87
{OPD(0xDD, 0xD8), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FSTToStack>}, //^
@@ -247,6 +252,8 @@ constexpr std::array<DispatchTableEntry, 133> X87F64OpTable = {{
{OPD(0xDE, 0xC0), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FADDF64, OpSize::f80Bit, false, OpDispatchBuilder::OpResult::RES_STI>},
{OPD(0xDE, 0xC8), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FMULF64, OpSize::f80Bit, false, OpDispatchBuilder::OpResult::RES_STI>},
{OPD(0xDE, 0xD0), 8,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FCOMIF64, OpSize::f80Bit, false, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, false>},
{OPD(0xDE, 0xD9), 1,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FCOMIF64, OpSize::f80Bit, false, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, true>},
{OPD(0xDE, 0xE0), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FSUBF64, OpSize::f80Bit, false, true, OpDispatchBuilder::OpResult::RES_STI>},
@@ -273,6 +280,9 @@ constexpr std::array<DispatchTableEntry, 133> X87F64OpTable = {{
// XXX: This should also set the x87 tag bits to empty
// We don't support this currently, so just pop the stack
{OPD(0xDF, 0xC0), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::X87ModifySTP, true>},
{OPD(0xDF, 0xC8), 8, &OpDispatchBuilder::FXCH},
{OPD(0xDF, 0xD0), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FSTToStack>},
{OPD(0xDF, 0xD8), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FSTToStack>},
{OPD(0xDF, 0xE0), 8, &OpDispatchBuilder::X87FNSTSW},
{OPD(0xDF, 0xE8), 8,
@@ -281,7 +291,7 @@ constexpr std::array<DispatchTableEntry, 133> X87F64OpTable = {{
&OpDispatchBuilder::Bind<&OpDispatchBuilder::FCOMIF64, OpSize::f80Bit, false, OpDispatchBuilder::FCOMIFlags::FLAGS_RFLAGS, false>},
}};
constexpr std::array<DispatchTableEntry, 133> X87F80OpTable = {{
constexpr std::array<DispatchTableEntry, 140> X87F80OpTable = {{
{OPDReg(0xD8, 0) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FADD, OpSize::i32Bit, false, OpDispatchBuilder::OpResult::RES_ST0>},
{OPDReg(0xD8, 1) | 0x00, 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FMUL, OpSize::i32Bit, false, OpDispatchBuilder::OpResult::RES_ST0>},
@@ -453,6 +463,8 @@ constexpr std::array<DispatchTableEntry, 133> X87F80OpTable = {{
{OPD(0xDC, 0xC0), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FADD, OpSize::f80Bit, false, OpDispatchBuilder::OpResult::RES_STI>},
{OPD(0xDC, 0xC8), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FMUL, OpSize::f80Bit, false, OpDispatchBuilder::OpResult::RES_STI>},
{OPD(0xDC, 0xD0), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FCOMI, OpSize::f80Bit, false, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, false>},
{OPD(0xDC, 0xD8), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FCOMI, OpSize::f80Bit, false, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, false>},
{OPD(0xDC, 0xE0), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FSUB, OpSize::f80Bit, false, true, OpDispatchBuilder::OpResult::RES_STI>},
{OPD(0xDC, 0xE8), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FSUB, OpSize::f80Bit, false, false, OpDispatchBuilder::OpResult::RES_STI>},
{OPD(0xDC, 0xF0), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FDIV, OpSize::f80Bit, false, true, OpDispatchBuilder::OpResult::RES_STI>},
@@ -474,6 +486,7 @@ constexpr std::array<DispatchTableEntry, 133> X87F80OpTable = {{
{OPDReg(0xDD, 7) | 0x00, 8, &OpDispatchBuilder::X87FNSTSW},
{OPD(0xDD, 0xC0), 8, &OpDispatchBuilder::X87FFREE},
{OPD(0xDD, 0xC8), 8, &OpDispatchBuilder::FXCH},
{OPD(0xDD, 0xD0), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FSTToStack>},
{OPD(0xDD, 0xD8), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FSTToStack>},
@@ -502,6 +515,7 @@ constexpr std::array<DispatchTableEntry, 133> X87F80OpTable = {{
{OPD(0xDE, 0xC0), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FADD, OpSize::f80Bit, false, OpDispatchBuilder::OpResult::RES_STI>},
{OPD(0xDE, 0xC8), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FMUL, OpSize::f80Bit, false, OpDispatchBuilder::OpResult::RES_STI>},
{OPD(0xDE, 0xD0), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FCOMI, OpSize::f80Bit, false, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, false>},
{OPD(0xDE, 0xD9), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FCOMI, OpSize::f80Bit, false, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, true>},
{OPD(0xDE, 0xE0), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FSUB, OpSize::f80Bit, false, true, OpDispatchBuilder::OpResult::RES_STI>},
{OPD(0xDE, 0xE8), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FSUB, OpSize::f80Bit, false, false, OpDispatchBuilder::OpResult::RES_STI>},
@@ -527,6 +541,9 @@ constexpr std::array<DispatchTableEntry, 133> X87F80OpTable = {{
// XXX: This should also set the x87 tag bits to empty
// We don't support this currently, so just pop the stack
{OPD(0xDF, 0xC0), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::X87ModifySTP, true>},
{OPD(0xDF, 0xC8), 8, &OpDispatchBuilder::FXCH},
{OPD(0xDF, 0xD0), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FSTToStack>},
{OPD(0xDF, 0xD8), 8, &OpDispatchBuilder::Bind<&OpDispatchBuilder::FSTToStack>},
{OPD(0xDF, 0xE0), 8, &OpDispatchBuilder::X87FNSTSW},
{OPD(0xDF, 0xE8), 8,
@@ -688,9 +705,9 @@ auto GenerateX87TableLambda = [](const auto DispatchTable) consteval {
// / 1
{OPD(0xDC, 0xC8), 8, X86InstInfo{"FMUL", TYPE_X87, FLAGS_NONE, 0}},
// / 2
{OPD(0xDC, 0xD0), 8, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(0xDC, 0xD0), 8, X86InstInfo{"FCOM", TYPE_X87, FLAGS_X87_FLAGS, 0}},
// / 3
{OPD(0xDC, 0xD8), 8, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(0xDC, 0xD8), 8, X86InstInfo{"FCOMP", TYPE_X87, FLAGS_X87_FLAGS | FLAGS_POP, 0}},
// / 4
{OPD(0xDC, 0xE0), 8, X86InstInfo{"FSUBR", TYPE_X87, FLAGS_NONE, 0}},
// / 5
@@ -711,7 +728,7 @@ auto GenerateX87TableLambda = [](const auto DispatchTable) consteval {
// / 0
{OPD(0xDD, 0xC0), 8, X86InstInfo{"FFREE", TYPE_X87, FLAGS_NONE, 0}},
// / 1
{OPD(0xDD, 0xC8), 8, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(0xDD, 0xC8), 8, X86InstInfo{"FXCH", TYPE_X87, FLAGS_NONE, 0}},
// / 2
{OPD(0xDD, 0xD0), 8, X86InstInfo{"FST", TYPE_INST, FLAGS_SF_MOD_DST, 0}},
// / 3
@@ -738,7 +755,7 @@ auto GenerateX87TableLambda = [](const auto DispatchTable) consteval {
// / 1
{OPD(0xDE, 0xC8), 8, X86InstInfo{"FMULP", TYPE_X87, FLAGS_POP, 0}},
// / 2
{OPD(0xDE, 0xD0), 8, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(0xDE, 0xD0), 8, X86InstInfo{"FCOMP", TYPE_X87, FLAGS_X87_FLAGS | FLAGS_POP, 0}},
// / 3
{OPD(0xDE, 0xD8), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(0xDE, 0xD9), 1, X86InstInfo{"FCOMPP", TYPE_X87, FLAGS_POP, 0}},
@@ -771,11 +788,11 @@ auto GenerateX87TableLambda = [](const auto DispatchTable) consteval {
// Almost all x86 CPUs implement this, and it is expected to be around
{OPD(0xDF, 0xC0), 8, X86InstInfo{"FFREEP", TYPE_X87, FLAGS_POP, 0}},
// / 1
{OPD(0xDF, 0xC8), 8, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(0xDF, 0xC8), 8, X86InstInfo{"FXCH", TYPE_X87, FLAGS_NONE, 0}},
// / 2
{OPD(0xDF, 0xD0), 8, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(0xDF, 0xD0), 8, X86InstInfo{"FSTP", TYPE_X87, FLAGS_SF_MOD_DST | FLAGS_POP, 0}},
// / 3
{OPD(0xDF, 0xD8), 8, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
{OPD(0xDF, 0xD8), 8, X86InstInfo{"FSTP", TYPE_X87, FLAGS_SF_MOD_DST | FLAGS_POP, 0}},
// / 4
{OPD(0xDF, 0xE0), 1, X86InstInfo{"FNSTSW", TYPE_INST, GenFlagsSameSize(SIZE_16BIT) | FLAGS_SF_DST_RAX, 0}},
{OPD(0xDF, 0xE1), 7, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0}},
+2 -2
View File
@@ -42,7 +42,7 @@ void __attribute__((noinline)) __jit_debug_register_code() {
namespace FEXCore {
void GDBJITRegister(FEXCore::ExecutableFileInfo& Entry, uintptr_t VAFileStart, uint64_t GuestRIP, uintptr_t HostEntry,
void GDBJITRegister(const FEXCore::ExecutableFileInfo& Entry, uintptr_t VAFileStart, uint64_t GuestRIP, uintptr_t HostEntry,
FEXCore::Core::DebugData& DebugData) {
auto map = Entry.SourcecodeMap.get();
@@ -113,7 +113,7 @@ void GDBJITRegister(FEXCore::ExecutableFileInfo& Entry, uintptr_t VAFileStart, u
} // namespace FEXCore
#else
namespace FEXCore {
void GDBJITRegister(FEXCore::ExecutableFileInfo&, uintptr_t, uint64_t, uintptr_t, FEXCore::Core::DebugData&) {
void GDBJITRegister(const FEXCore::ExecutableFileInfo&, uintptr_t, uint64_t, uintptr_t, FEXCore::Core::DebugData&) {
ERROR_AND_DIE_FMT("GDBSymbols support not compiled in");
}
} // namespace FEXCore
+1 -1
View File
@@ -4,5 +4,5 @@
#include <Interface/Core/JIT/DebugData.h>
namespace FEXCore {
void GDBJITRegister(FEXCore::ExecutableFileInfo&, uintptr_t VAFileStart, uint64_t GuestRIP, uintptr_t HostEntry, FEXCore::Core::DebugData&);
void GDBJITRegister(const FEXCore::ExecutableFileInfo&, uintptr_t VAFileStart, uint64_t GuestRIP, uintptr_t HostEntry, FEXCore::Core::DebugData&);
}
+1 -18
View File
@@ -60,24 +60,7 @@ struct NodeID final {
Value = 0;
}
[[nodiscard]] friend constexpr bool operator==(NodeID, NodeID) noexcept = default;
[[nodiscard]]
friend constexpr bool operator<(NodeID lhs, NodeID rhs) noexcept {
return lhs.Value < rhs.Value;
}
[[nodiscard]]
friend constexpr bool operator>(NodeID lhs, NodeID rhs) noexcept {
return operator<(rhs, lhs);
}
[[nodiscard]]
friend constexpr bool operator<=(NodeID lhs, NodeID rhs) noexcept {
return !operator>(lhs, rhs);
}
[[nodiscard]]
friend constexpr bool operator>=(NodeID lhs, NodeID rhs) noexcept {
return !operator<(lhs, rhs);
}
[[nodiscard]] constexpr auto operator<=>(const NodeID&) const noexcept = default;
friend std::ostream& operator<<(std::ostream& out, NodeID ID) {
out << ID.Value;
+47 -21
View File
@@ -105,6 +105,11 @@
"PosInfinity = 2,",
"TowardsZero = 3, /* Truncate */",
"Host = 4,"
],
"class ConstPad : uint8_t": [
"NoPad = 0,",
"DoPad = 1,",
"AutoPad = 2,"
]
},
"Defines": [
@@ -131,6 +136,7 @@
"u16": "uint16_t",
"u32": "uint32_t",
"u64": "uint64_t",
"c_str": "const char*",
"OpSize": "FEXCore::IR::OpSize",
"SSA": "OrderedNode*",
"GPR": "OrderedNode*",
@@ -142,6 +148,7 @@
"MemOffsetType": "MemOffsetType",
"BreakDefinition": "BreakDefinition",
"RoundType": "RoundMode",
"ConstPad": "ConstPad",
"FloatCompareOp": "FloatCompareOp",
"NamedVectorConstant": "FEXCore::IR::NamedVectorConstant",
"IndexNamedVectorConstant": "FEXCore::IR::IndexNamedVectorConstant",
@@ -234,6 +241,12 @@
"Desc": ["Debug operation that prints an SSA value to the console",
"May only print 64bits of the value"]
},
"PrintMsg c_str:$Value": {
"HasSideEffects": true,
"Desc": ["Debug operation that prints an string to the console.",
"This is for debug only! Will break code caching!"
]
},
"GPR = AllocateGPR i1:$ForPair": {
"Desc": ["Silly pseudo-instruction to allocate a register for a future destination",
"Note: if an instruction uses allocated destinations-as-sources,",
@@ -482,7 +495,7 @@
"HasSideEffects": true,
"Desc": ["Spills an SSA value to memory",
"Spill slots are register allocated and has live ranges calculated to handle slot calculation",
"```diff\n- !Don't use this op. It is for RA to handle spilling and filling!\n```"
"!Don't use this op. It is for RA to handle spilling and filling!"
],
"EmitValidation": [
"WalkFindRegClass($Value) == $Class"
@@ -492,7 +505,7 @@
"SSA = FillRegister OpSize:#Size, OpSize:#ElementSize, u32:$Slot, RegisterClass:$Class": {
"Desc": ["Fills a register from a spill slot",
"Spill slots are register allocated and has live ranges calculated to handle slot calculation",
"```diff\n- !Don't use this op. It is for RA to handle spilling and filling!\n```"
"!Don't use this op. It is for RA to handle spilling and filling!"
],
"DestSize": "Size",
"ElementSize": "ElementSize"
@@ -771,6 +784,17 @@
"RegisterSize == FEXCore::IR::OpSize::i128Bit || RegisterSize == FEXCore::IR::OpSize::i256Bit",
"Offset % IR::OpSizeToSize(RegisterSize) == 0"
]
},
"ContextClear u32:$Offset, u32:$Size": {
"Desc": [
"Clears a region of the context by CLZero size",
"Both the offset and size alignment need to be by CLZero size"
],
"HasSideEffects": true,
"EmitValidation": [
"Offset % 64 == 0",
"Size % 64 == 0"
]
}
},
"Atomic": {
@@ -804,16 +828,6 @@
"Size == FEXCore::IR::OpSize::i32Bit || Size == FEXCore::IR::OpSize::i64Bit"
]
},
"AtomicXor OpSize:#Size, GPR:$Value, GPR:$Addr": {
"HasSideEffects": true,
"Desc": ["Atomic integer xor",
"IR layout must match Fetch-variant, otherwise DCE IR optimization breaks!"
],
"DestSize": "Size",
"EmitValidation": [
"Size == FEXCore::IR::OpSize::i8Bit || Size == FEXCore::IR::OpSize::i16Bit || Size == FEXCore::IR::OpSize::i32Bit || Size == FEXCore::IR::OpSize::i64Bit"
]
},
"GPR = AtomicSwap OpSize:#Size, GPR:$Value, GPR:$Addr": {
"HasSideEffects": true,
"Desc": ["Atomic integer swap"
@@ -940,11 +954,15 @@
]
},
"GPR = Constant i64:$Constant": {
"GPR = Constant i64:$Constant, ConstPad:$Pad{IR::ConstPad::NoPad}, i32:$MaxBytes{0}": {
"Desc": ["Generates a 64bit constant inside of a GPR",
"Unsupported to create a constant in FPR"
],
"DestSize": "OpSize::i64Bit"
"DestSize": "OpSize::i64Bit",
"EmitValidation": [
"MaxBytes >= 0 && MaxBytes <= 8 && (MaxBytes & 1) == 0",
"MaxBytes == 0 || (Constant >> (MaxBytes * 8)) == 0"
]
},
"InlineConstant i64:$Constant": {
@@ -2151,6 +2169,14 @@
]
},
"FPR = VOrn OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$Vector1, FPR:$Vector2": {
"DestSize": "RegisterSize",
"ElementSize": "ElementSize",
"EmitValidation": [
"RegisterSize == FEXCore::IR::OpSize::i256Bit || RegisterSize == FEXCore::IR::OpSize::i128Bit || RegisterSize == FEXCore::IR::OpSize::i64Bit"
]
},
"FPR = VOr OpSize:#RegisterSize, OpSize:#ElementSize, FPR:$Vector1, FPR:$Vector2": {
"DestSize": "RegisterSize",
"ElementSize": "ElementSize",
@@ -2732,7 +2758,7 @@
"F64": {
"FPR = F64ATAN FPR:$Src1, FPR:$Src2": {
"DestSize": "OpSize::i64Bit",
"JITDispatch": false
"JITDispatch": true
},
"FPR = F64FPREM FPR:$Src1, FPR:$Src2": {
"DestSize": "OpSize::i64Bit",
@@ -2744,27 +2770,27 @@
},
"FPR = F64SCALE FPR:$Src1, FPR:$Src2": {
"DestSize": "OpSize::i64Bit",
"JITDispatch": false
"JITDispatch": true
},
"FPR = F64F2XM1 FPR:$Src": {
"DestSize": "OpSize::i64Bit",
"JITDispatch": false
"JITDispatch": true
},
"FPR = F64FYL2X FPR:$Src, FPR:$Src2": {
"DestSize": "OpSize::i64Bit",
"JITDispatch": false
"JITDispatch": true
},
"FPR = F64TAN FPR:$Src": {
"DestSize": "OpSize::i64Bit",
"JITDispatch": false
"JITDispatch": true
},
"FPR = F64SIN FPR:$Src": {
"DestSize": "OpSize::i64Bit",
"JITDispatch": false
"JITDispatch": true
},
"FPR = F64COS FPR:$Src": {
"DestSize": "OpSize::i64Bit",
"JITDispatch": false
"JITDispatch": true
},
"FPR:$Sin, FPR:$Cos = F64SINCOS FPR:$Src": {
"DestSize": "OpSize::i64Bit",
+18 -4
View File
@@ -38,6 +38,10 @@ static void PrintArg(fextl::stringstream* out, const IRListView*, uint64_t Arg)
*out << fextl::fmt::format("#{:#x}", Arg);
}
static void PrintArg(fextl::stringstream* out, const IRListView*, const char* const Arg) {
*out << fextl::fmt::format("'{}'", Arg);
}
static void PrintArg(fextl::stringstream* out, const IRListView*, CondClass Arg) {
if (Arg == CondClass::AL) {
*out << "ALWAYS";
@@ -149,6 +153,17 @@ static void PrintArg(fextl::stringstream* out, const IRListView*, RoundMode Arg)
}();
}
static void PrintArg(fextl::stringstream* out, const IRListView*, ConstPad Arg) {
*out << [Arg] {
switch (Arg) {
case ConstPad::NoPad: return "NoPad";
case ConstPad::DoPad: return "DoPad";
case ConstPad::AutoPad: return "AutoPad";
}
return "<Unknown ConstPad Type>";
}();
}
static void PrintArg(fextl::stringstream* out, const IRListView*, NamedVectorConstant Arg) {
*out << [Arg] {
// clang-format off
@@ -345,16 +360,15 @@ void Dump(fextl::stringstream* out, const IRListView* IR) {
++CurrentIndent;
AddIndent();
*out << fextl::fmt::format("(%0) IRHeader %{}, #{:#x}, #{}, #{}\n", HeaderOp->Blocks.ID(), HeaderOp->OriginalRIP, HeaderOp->BlockCount,
HeaderOp->NumHostInstructions);
*out << fextl::fmt::format("(%0) IRHeader %{}, #{:#x}, #{}, #{}\n", HeaderOp->Blocks.ID(), +HeaderOp->OriginalRIP, +HeaderOp->BlockCount,
+HeaderOp->NumHostInstructions);
for (auto [BlockNode, BlockHeader] : IR->GetBlocks()) {
{
auto BlockIROp = BlockHeader->C<FEXCore::IR::IROp_CodeBlock>();
AddIndent();
*out << "(%" << IR->GetID(BlockNode) << ") "
<< "CodeBlock ";
*out << "(%" << IR->GetID(BlockNode) << ") " << "CodeBlock ";
*out << "%" << BlockIROp->Begin.ID() << ", ";
*out << "%" << BlockIROp->Last.ID() << std::endl;
+23 -10
View File
@@ -21,15 +21,15 @@ class IREmitter {
public:
IREmitter(FEXCore::Utils::IntrusivePooledAllocator& ThreadAllocator, bool SupportsTSOImm9)
: DualListData {ThreadAllocator, 8 * 1024 * 1024}
, SupportsTSOImm9(SupportsTSOImm9) {
ReownOrClaimBuffer();
ResetWorkingList();
}
, SupportsTSOImm9(SupportsTSOImm9) {}
virtual ~IREmitter() = default;
void ReownOrClaimBuffer() {
DualListData.ReownOrClaimBuffer();
// Reset the working list on new buffer.
ResetWorkingList();
}
void DelayedDisownBuffer() {
@@ -39,7 +39,6 @@ public:
IRListView ViewIR() {
return IRListView(&DualListData);
}
void ResetWorkingList();
/**
* @name IR allocation routines
@@ -239,22 +238,33 @@ public:
DEF_ADDSUB(AddWithFlags)
DEF_ADDSUB(SubWithFlags)
int64_t Constants[32];
struct ConstantData {
int64_t Value;
ConstPad Pad;
int32_t MaxBytes;
[[nodiscard]] auto operator<=>(const ConstantData&) const noexcept = default;
};
ConstantData Constants[32];
Ref ConstantRefs[32];
uint32_t NrConstants;
Ref Constant(int64_t Value) {
Ref Constant(int64_t Value, ConstPad Pad = IR::ConstPad::NoPad, int32_t MaxBytes = 0) {
const ConstantData Data {
.Value = Value,
.Pad = Pad,
.MaxBytes = MaxBytes,
};
// Search for the constant in the pool.
for (unsigned i = 0; i < std::min(NrConstants, 32u); ++i) {
if (Constants[i] == Value) {
if (Constants[i] == Data) {
return ConstantRefs[i];
}
}
// Otherwise, materialize a fresh constant and pool it.
Ref R = _Constant(Value);
Ref R = _Constant(Value, Pad, MaxBytes);
unsigned i = (NrConstants++) & 31;
Constants[i] = Value;
Constants[i] = Data;
ConstantRefs[i] = R;
return R;
}
@@ -501,6 +511,9 @@ protected:
fextl::vector<Ref> CodeBlocks;
uint64_t Entry {};
bool SupportsTSOImm9 {};
private:
void ResetWorkingList();
};
} // namespace FEXCore::IR
Loaded 100 of 517 files, more files were not shown because too many files have changed in this diff. Show more