Compare commits

..
Author SHA1 Message Date
Billy Laws 9fe5eb1979 JIT: Restore behaviour of emitting interrupt checks at every block entry
This is needed to handle suspend in infinite loops that occur as a
result of block-size constraints or indirect jumps. Fixes grow home.
2025-10-29 00:35:46 +00:00
Ryan Houdek f414c92963 Code view 2025-10-28 23:53:15 +00:00
Ryan Houdek 90c59e37cb unittests/ASM: Adds test for too large branch objects 2025-10-28 23:53:15 +00:00
Ryan Houdek b7c7789a01 FEXCore/JIT: Supports restarting JIT in case of encoding failure
ARM64 branches have fairly small relative distances they can encode.
These can be +-1MB, or even +-32KB. The largest relative branch is
+-128MB, which we already set as an upper limit of our block JIT cache
size.

We have for a long time just compiled these without checking with the
expectation that things just happen to work. We didn't hit the asserts
so it was relatively low priority. Apparently now with Steam and a
MaxInst limit of 5000, we are now hitting an assert where we are
encoding too large of a range.

Implement support for long jumping from anywhere in the JIT for when a
long jump tries to be encoded and fails, allowing us to restart the JIT
at any moment. This is implemented as a long jump when this singular
feature could have gotten away with some sort of invasive check and
early exit path for two reasons. For one, that would be even more
invasive, effectively doing try-catch logic manually. And two, the next
step is supporting JIT buffer overflow for when our block size heuristic
fails.

This next step will mandate longjump on SIGSEGV (with cooperative
interaction with the frontend) from effectively /anywhere/ in the JIT.
One of the design goals of the CodeEmitter is that every code emission
function doesn't do a size remaining check to allow the compiler to do
some very effective optimization of emitting code blocks to memory (and
it works!).

But we lose the ability to sanely size check. When writing the emitter I
knew we were going to need to write this cooperative guard page handler,
and we're finally at a point where it needs to be done. This will be in
the next PR although.
2025-10-28 23:53:15 +00:00
Ryan Houdek f653c5e0c0 FEXCore/JIT: Ignore local encoding limit checks
These are guaranteed not to hit encoding distance limits, so we can
ignore the returns.
2025-10-28 23:53:15 +00:00
Ryan Houdek 65fff73959 FEXCore/Dispatcher: Check encoding errors 2025-10-28 23:53:15 +00:00
Ryan Houdek 93b7c513d8 FEXCore/VectorRegType: Trivial header fix 2025-10-28 23:53:15 +00:00
Ryan Houdek f1d14c6325 Linux/BPFEmitter: Explicitly ignored encoding bool
We know these won't encode in errors.
2025-10-28 23:53:15 +00:00
Ryan Houdek 8223c6ac36 unittests/Emitter: Explicitly ignore encoding bool
We know these won't encode in errors.
2025-10-28 23:53:15 +00:00
Ryan Houdek e17677580d CodeEmitter: Return bool if Label instructions can't be encoded
Programming error if they aren't checked, as they will encode
incorrectly if they are too large for their respective instructions.
2025-10-28 23:53:15 +00:00
Ryan Houdek 150bf7b30c FEXCore: Moves longjump implementation from FEX frontend
This will be getting used by FEXCore in a bit.
2025-10-28 23:53:15 +00:00
Ryan Houdek 674efc69c4 FEX: Print a log when kernel unaligned atomics are used 2025-10-28 23:53:15 +00:00
Billy Laws 38049c5281 Windows: Enable downstream kernel-side unaligned atomic handling 2025-10-28 23:53:15 +00:00
Billy Laws e3627349a1 FEXLoader: Enable downstream kernel-side unaligned atomic handling 2025-10-28 23:53:15 +00:00
Billy Laws 7c207080a4 Windows: Support new two-stage invalidation model 2025-10-28 23:53:12 +00:00
Billy Laws eeee5b53ca Linux: Support new two-stage invalidation model 2025-10-28 23:53:12 +00:00
Billy Laws 47619063c2 LookupCache: Introduce two-pass code invalidation model
Shared code buffer support introduced the concept of having a single
GuestToHostMaps shared across many threads. In the common case all
threads will share one however if e.g. a resize recently occured and
specific thread is yet to compile any code with the new codebuffer it
will still use the old GuestToHostMap. The current invalidation
approach handles this by repeatedly calling erase for every single
thread's GuestToHostMap, even if it is repeated. An accumulator is used
to ensure when two threads share a map, the L1/L2 cache entries in the
second thread will still be invalidated even if the the iteration for
the first thread removed them from the map.

Unfortunately this is incredibly slow in cases with many threads, as
a significant number of redundant map lookups and L1/L2 cache erasures
on threads that never even observed a given block can occur. Solve this
by introducing a two-pass model:
- First, all active codebuffers (and their associated GuestToHostMaps)
  have their entries invalidated for the given range, these codebuffers
  are tracked internally within FEXCore. It is at this point that delinking
  callbacks are ran.
- Second, each thread will have its caches invalidated. But rather than
  naively invalidating the L1/L2 caches for every invalidated block for
  every thread, threads now track on their own what specific entries
  have been potentially fetched into their L1/L2 caches. This is
  aided by GuestToHostMap now tracking the pages each block touches. (an
  inverse CodePages so to speak).
2025-10-28 23:53:12 +00:00
Billy Laws cb7076cbab FEXCore: Keep a list of weak refs to all allocated codebuffers
We currently rely on the frontend to keep track of threads and then
iterate over all threads to perform per-codebuffer operations. However
as codebuffers are shared between many threads (the common case is a
single code buffer across all) this ends up being inefficient. Introduce
a list of codebuffers to solve that (new codebuffers are very rare, so a
vector is plenty fine here for erasing invalid weak refs).
2025-10-28 23:53:12 +00:00
Billy Laws 8dde79826e LookupCache: Drop unused state frame argument for delinker cbs 2025-10-28 23:53:12 +00:00
213 changed files with 3454 additions and 7382 deletions

No files matched your search

-79
View File
@@ -1,79 +0,0 @@
name: steamrt4 build
on:
push:
branches:
- main
pull_request:
branches:
- main
env:
BUILD_TYPE: Release
CC: clang
CXX: clang++
jobs:
steamrt4_build:
runs-on: ${{ matrix.arch }}
strategy:
matrix:
arch: [[self-hosted, ARM64, distrobox]]
fail-fast: false
steps:
- uses: actions/checkout@v3
- name: Set runner label
run: echo "runner_label=${{ matrix.arch[1] }}" >> $GITHUB_ENV
- name : submodule checkout
run: |
git submodule sync --recursive
git submodule update --init --depth 1
- name: Clean Build Environment
run: |
rm -Rf ${{runner.workspace}}/build
cmake -E make_directory ${{runner.workspace}}/build
# Setup everything required.
- name : distrobox setup
run: |
distrobox create -Y -i registry.gitlab.steamos.cloud/steamrt/steamrt4/sdk/arm64:4.0.20251117.183306 steamrt4 || true
distrobox upgrade steamrt4
distrobox enter --name steamrt4 -- sudo apt-get install -y \
git cmake ninja-build ccache \
lld clang \
libclang-dev llvm-dev \
libstdc++-14-dev-i386-cross libgcc-14-dev-i386-cross \
libstdc++-14-dev-amd64-cross libgcc-14-dev-amd64-cross
- name: Create Build Environment
run: distrobox enter --name steamrt4 -- cmake -E make_directory ${{runner.workspace}}/build
- name: Configure CMake
shell: bash
working-directory: ${{runner.workspace}}/build
run: distrobox enter --name steamrt4 -- cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DBUILD_STEAM_SUPPORT=True -DENABLE_LTO=True -DENABLE_ASSERTIONS=False -DBUILD_THUNKS=True -DBUILD_FEXCONFIG=False -DBUILD_TESTING=False -DENABLE_CLANG_THUNKS=True -DUSE_LINKER=lld -DCMAKE_INSTALL_PREFIX=/usr
- name: Build
working-directory: ${{runner.workspace}}/build
shell: bash
run: distrobox enter --name steamrt4 -- cmake --build . --config $BUILD_TYPE
- name: install
working-directory: ${{runner.workspace}}/build
shell: bash
env:
DESTDIR: ${{runner.workspace}}/install
run: distrobox enter --name steamrt4 -- cmake --build . --config $BUILD_TYPE -t install
- name: Upload libraries
uses: 'actions/upload-artifact@v4'
timeout-minutes: 1
with:
overwrite: true
name: steamrt4_steampipe_depot
path: ${{runner.workspace}}/install/*
retention-days: 1
compression-level: 9
+2 -14
View File
@@ -33,7 +33,6 @@ option(ENABLE_FEXCORE_PROFILER "Enables use of the FEXCore timeline profiling ca
set (FEXCORE_PROFILER_BACKEND "gpuvis" CACHE STRING "Set which backend to use for the FEXCore profiler (gpuvis, tracy)")
option(ENABLE_GLIBC_ALLOCATOR_HOOK_FAULT "Enables glibc memory allocation hooking with fault for CI testing")
option(USE_PDB_DEBUGINFO "Builds debug info in PDB format" FALSE)
option(BUILD_STEAM_SUPPORT "Builds FEX for integration into Steam" FALSE)
set (X86_32_TOOLCHAIN_FILE "${CMAKE_CURRENT_SOURCE_DIR}/Data/CMake/toolchain_x86_32.cmake" CACHE FILEPATH "Toolchain file for the (cross-)compiler targeting i686")
set (X86_64_TOOLCHAIN_FILE "${CMAKE_CURRENT_SOURCE_DIR}/Data/CMake/toolchain_x86_64.cmake" CACHE FILEPATH "Toolchain file for the (cross-)compiler targeting x86_64")
@@ -65,10 +64,6 @@ if (NOT MINGW_BUILD)
endif()
endif()
if (BUILD_STEAM_SUPPORT)
add_definitions(-DFEX_STEAM_SUPPORT=1)
endif()
if (ENABLE_FEXCORE_PROFILER)
add_definitions(-DENABLE_FEXCORE_PROFILER=1)
string(TOUPPER "${FEXCORE_PROFILER_BACKEND}" FEXCORE_PROFILER_BACKEND)
@@ -481,16 +476,13 @@ add_subdirectory(FEXHeaderUtils/)
add_subdirectory(CodeEmitter/)
add_subdirectory(FEXCore/)
if (_M_ARM_64 AND NOT MINGW_BUILD AND NOT BUILD_STEAM_SUPPORT)
if (_M_ARM_64 AND NOT MINGW_BUILD)
# Binfmt_misc files must be installed prior to Source/ installs
add_subdirectory(Data/binfmts/)
endif()
add_subdirectory(Source/)
if (NOT BUILD_STEAM_SUPPORT)
add_subdirectory(Data/AppConfig/)
endif()
add_subdirectory(Data/AppConfig/)
# Install the ThunksDB file
file(GLOB CONFIG_SOURCES CONFIGURE_DEPENDS ${CMAKE_CURRENT_SOURCE_DIR}/Data/*.json)
@@ -590,10 +582,6 @@ if (BUILD_THUNKS)
add_dependencies(uninstall uninstall_guest-libs-32)
endif()
if (BUILD_STEAM_SUPPORT)
add_subdirectory(Source/Steam/)
endif()
set(FEX_VERSION_MAJOR "0")
set(FEX_VERSION_MINOR "0")
set(FEX_VERSION_PATCH "0")
+13 -21
View File
@@ -38,7 +38,9 @@ public:
[[nodiscard]] BranchEncodeSucceeded adr(ARMEmitter::Register rd, const BackwardLabel* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
if (IsADRRange(Imm)) {
LOGMAN_THROW_A_FMT(IsADRRange(Imm), "Unscaled offset too large");
if (IsADRRange(Imm)) [[likely]] {
constexpr uint32_t Op = 0b0001'0000 << 24;
DataProcessing_PCRel_Imm(Op, rd, Imm);
return BranchEncodeSucceeded::Success;
@@ -71,8 +73,9 @@ public:
[[nodiscard]] BranchEncodeSucceeded adrp(ARMEmitter::Register rd, const BackwardLabel* Label) {
int64_t Imm = reinterpret_cast<int64_t>(Label->Location) - (GetCursorAddress<int64_t>() & ~0xFFFLL);
LOGMAN_THROW_A_FMT(IsADRPRange(Imm) && IsADRPAligned(Imm), "Unscaled offset too large");
if (IsADRPRange(Imm) && IsADRPAligned(Imm)) {
if (IsADRPRange(Imm) && IsADRPAligned(Imm)) [[likely]] {
constexpr uint32_t Op = 0b1001'0000 << 24;
DataProcessing_PCRel_Imm(Op, rd, Imm);
return BranchEncodeSucceeded::Success;
@@ -100,22 +103,16 @@ public:
}
[[nodiscard]] BranchEncodeSucceeded LongAddressGen(ARMEmitter::Register rd, const BackwardLabel* Label) {
const auto SLocation = reinterpret_cast<int64_t>(Label->Location);
const auto ULocation = std::bit_cast<uint64_t>(SLocation);
const int64_t Imm = SLocation - (GetCursorAddress<int64_t>());
const auto UImm = std::bit_cast<uint64_t>(Imm);
int64_t Imm = reinterpret_cast<int64_t>(Label->Location) - (GetCursorAddress<int64_t>());
if (IsADRRange(Imm)) {
// If the range is in ADR range then we can just use ADR.
return adr(rd, Label);
}
if (IsADRPRange(Imm)) {
const int64_t ADRPImm = (SLocation & ~0xFFFLL) - (GetCursorAddress<int64_t>() & ~0xFFFLL);
} else if (IsADRPRange(Imm)) {
int64_t ADRPImm = (reinterpret_cast<int64_t>(Label->Location) & ~0xFFFLL) - (GetCursorAddress<int64_t>() & ~0xFFFLL);
// If the range is in the ADRP range then we can use ADRP.
const bool NeedsOffset = !IsADRPAligned(ULocation);
const uint64_t AlignedOffset = ULocation & 0xFFFULL;
bool NeedsOffset = !IsADRPAligned(reinterpret_cast<uint64_t>(Label->Location));
uint64_t AlignedOffset = reinterpret_cast<uint64_t>(Label->Location) & 0xFFFULL;
// First emit ADRP
adrp(rd, ADRPImm >> 12);
@@ -128,19 +125,14 @@ public:
return BranchEncodeSucceeded::Success;
}
// Stinky path, we need to load the address as a sequence of movz+movk+movk
movz(ARMEmitter::Size::i64Bit, rd, (UImm >> 32) & 0xFFFF, 32);
movk(ARMEmitter::Size::i64Bit, rd, (UImm >> 16) & 0xFFFF, 16);
movk(ARMEmitter::Size::i64Bit, rd, UImm & 0xFFFF);
return BranchEncodeSucceeded::Success;
// Can't encode.
return BranchEncodeSucceeded::Failure;
}
[[nodiscard]] BranchEncodeSucceeded LongAddressGen(ARMEmitter::Register rd, ForwardLabel* Label) {
AddLocationToLabel(Label, ForwardLabel::Reference {.Location = GetCursorAddress<uint8_t*>(), .Type = ForwardLabel::InstType::LONG_ADDRESS_GEN});
// Emit a register index and two nops. These will be backpatched.
// Emit a register index and a nop. These will be backpatched.
dc32(rd.Idx());
nop();
nop();
// Forward label doesn't know if it can encode until Bind.
return BranchEncodeSucceeded::Success;
+9 -8
View File
@@ -22,7 +22,7 @@ public:
}
[[nodiscard]] BranchEncodeSucceeded b(ARMEmitter::Condition Cond, const BackwardLabel* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
if (Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0)) {
if (Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0)) [[likely]] {
constexpr uint32_t Op = 0b0101'010 << 25;
Branch_Conditional(Op, 0, 0, Cond, Imm >> 2);
return BranchEncodeSucceeded::Success;
@@ -55,7 +55,7 @@ public:
}
[[nodiscard]] BranchEncodeSucceeded bc(ARMEmitter::Condition Cond, const BackwardLabel* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
if (Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0)) {
if (Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0)) [[likely]] {
constexpr uint32_t Op = 0b0101'010 << 25;
Branch_Conditional(Op, 0, 1, Cond, Imm >> 2);
return BranchEncodeSucceeded::Success;
@@ -116,7 +116,7 @@ public:
}
[[nodiscard]] BranchEncodeSucceeded b(const BackwardLabel* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
if (Imm >= -134217728 && Imm <= 134217724 && ((Imm & 0b11) == 0)) {
if (Imm >= -134217728 && Imm <= 134217724 && ((Imm & 0b11) == 0)) [[likely]] {
constexpr uint32_t Op = 0b0001'01 << 26;
UnconditionalBranch(Op, Imm >> 2);
return BranchEncodeSucceeded::Success;
@@ -151,7 +151,7 @@ public:
[[nodiscard]] BranchEncodeSucceeded bl(const BackwardLabel* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
if (Imm >= -134217728 && Imm <= 134217724 && ((Imm & 0b11) == 0)) {
if (Imm >= -134217728 && Imm <= 134217724 && ((Imm & 0b11) == 0)) [[likely]] {
constexpr uint32_t Op = 0b1001'01 << 26;
UnconditionalBranch(Op, Imm >> 2);
@@ -189,7 +189,7 @@ public:
[[nodiscard]] BranchEncodeSucceeded cbz(ARMEmitter::Size s, ARMEmitter::Register rt, const BackwardLabel* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
if (Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0)) {
if (Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0)) [[likely]] {
constexpr uint32_t Op = 0b0011'0100 << 24;
CompareAndBranch(Op, s, rt, Imm >> 2);
return BranchEncodeSucceeded::Success;
@@ -227,7 +227,7 @@ public:
[[nodiscard]] BranchEncodeSucceeded cbnz(ARMEmitter::Size s, ARMEmitter::Register rt, const BackwardLabel* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
if (Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0)) {
if (Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0)) [[likely]] {
constexpr uint32_t Op = 0b0011'0101 << 24;
CompareAndBranch(Op, s, rt, Imm >> 2);
return BranchEncodeSucceeded::Success;
@@ -265,7 +265,7 @@ public:
[[nodiscard]] BranchEncodeSucceeded tbz(ARMEmitter::Register rt, uint32_t Bit, const BackwardLabel* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
if (Imm >= -32768 && Imm <= 32764 && ((Imm & 0b11) == 0)) {
if (Imm >= -32768 && Imm <= 32764 && ((Imm & 0b11) == 0)) [[likely]] {
constexpr uint32_t Op = 0b0011'0110 << 24;
TestAndBranch(Op, rt, Bit, Imm >> 2);
return BranchEncodeSucceeded::Success;
@@ -301,8 +301,9 @@ public:
}
[[nodiscard]] BranchEncodeSucceeded tbnz(ARMEmitter::Register rt, uint32_t Bit, const BackwardLabel* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(Imm >= -32768 && Imm <= 32764 && ((Imm & 0b11) == 0), "Unscaled offset too large");
if (Imm >= -32768 && Imm <= 32764 && ((Imm & 0b11) == 0)) {
if (Imm >= -32768 && Imm <= 32764 && ((Imm & 0b11) == 0)) [[likely]] {
constexpr uint32_t Op = 0b0011'0111 << 24;
TestAndBranch(Op, rt, Bit, Imm >> 2);
return BranchEncodeSucceeded::Success;
+21 -27
View File
@@ -662,7 +662,7 @@ public:
case ForwardLabel::InstType::ADR: {
uint32_t* Instruction = reinterpret_cast<uint32_t*>(Label->Location);
int64_t Imm = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(Instruction);
if (!IsADRRange(Imm)) {
if (!IsADRRange(Imm)) [[unlikely]] {
// Can't bind.
return false;
}
@@ -678,7 +678,7 @@ public:
uint32_t* Instruction = reinterpret_cast<uint32_t*>(Label->Location);
int64_t Imm = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(Instruction);
if (!(IsADRPRange(Imm) && IsADRPAligned(Imm))) {
if (!(IsADRPRange(Imm) && IsADRPAligned(Imm))) [[unlikely]] {
// Can't bind.
return false;
}
@@ -695,7 +695,7 @@ public:
case ForwardLabel::InstType::B: {
uint32_t* Instruction = reinterpret_cast<uint32_t*>(Label->Location);
int64_t Imm = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(Instruction);
if (!(Imm >= -134217728 && Imm <= 134217724 && ((Imm & 0b11) == 0))) {
if (!(Imm >= -134217728 && Imm <= 134217724 && ((Imm & 0b11) == 0))) [[unlikely]] {
// Can't bind.
return false;
}
@@ -711,7 +711,7 @@ public:
case ForwardLabel::InstType::TEST_BRANCH: {
uint32_t* Instruction = reinterpret_cast<uint32_t*>(Label->Location);
int64_t Imm = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(Instruction);
if (!(Imm >= -32768 && Imm <= 32764 && ((Imm & 0b11) == 0))) {
if (!(Imm >= -32768 && Imm <= 32764 && ((Imm & 0b11) == 0))) [[unlikely]] {
// Can't bind.
return false;
}
@@ -728,7 +728,7 @@ public:
case ForwardLabel::InstType::RELATIVE_LOAD: {
uint32_t* Instruction = reinterpret_cast<uint32_t*>(Label->Location);
int64_t Imm = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(Instruction);
if (!(Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0))) {
if (!(Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0))) [[unlikely]] {
// Can't bind.
return false;
}
@@ -741,44 +741,38 @@ public:
break;
}
case ForwardLabel::InstType::LONG_ADDRESS_GEN: {
const auto* Instructions = reinterpret_cast<uint32_t*>(Label->Location);
const auto ImmInstOne = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(&Instructions[0]);
const auto ImmInstTwo = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(&Instructions[1]);
const auto ImmInstThree = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(&Instructions[2]);
const auto OriginalOffset = GetCursorOffset();
uint32_t* Instructions = reinterpret_cast<uint32_t*>(Label->Location);
int64_t ImmInstOne = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(&Instructions[0]);
int64_t ImmInstTwo = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(&Instructions[1]);
auto OriginalOffset = GetCursorOffset();
const auto InstOffset = GetCursorOffsetFromAddress(Instructions);
auto InstOffset = GetCursorOffsetFromAddress(Instructions);
SetCursorOffset(InstOffset);
// We encoded the destination register in to the first instruction space.
// Read it back.
ARMEmitter::Register DestReg(Instructions[0]);
if (IsADRRange(ImmInstThree)) {
// If within ADR range from the third instruction, then we can emit NOP+NOP+ADR
if (IsADRRange(ImmInstTwo)) {
// If within ADR range from the second instruction, then we can emit NOP+ADR
nop();
nop();
adr(DestReg, static_cast<uint32_t>(ImmInstThree) & 0x7FFF);
} else if (IsADRPRange(ImmInstTwo)) {
adr(DestReg, static_cast<uint32_t>(ImmInstTwo) & 0x7FFF);
} else if (IsADRPRange(ImmInstOne)) {
// If within ADRP range from the first instruction, then we are /definitely/ in range for the second instruction.
// First check if we are in non-offset range for second instruction.
if (IsADRPAligned(reinterpret_cast<uint64_t>(CurrentAddress))) {
// We can emit nop + nop + adrp
nop();
nop();
adrp(DestReg, static_cast<uint32_t>(ImmInstThree >> 12) & 0x7FFF);
} else {
// Not aligned, need nop + adrp + add
// We can emit nop + adrp
nop();
adrp(DestReg, static_cast<uint32_t>(ImmInstTwo >> 12) & 0x7FFF);
add(ARMEmitter::Size::i64Bit, DestReg, DestReg, ImmInstTwo & 0xFFF);
} else {
// Not aligned, need adrp + add
adrp(DestReg, static_cast<uint32_t>(ImmInstOne >> 12) & 0x7FFF);
add(ARMEmitter::Size::i64Bit, DestReg, DestReg, ImmInstOne & 0xFFF);
}
} else {
// Stinky path, we need to emit a movz+movk+movk sequence.
movz(ARMEmitter::Size::i64Bit, DestReg, uint32_t(ImmInstOne >> 32) & 0x7FFF, 32);
movk(ARMEmitter::Size::i64Bit, DestReg, uint32_t(ImmInstOne >> 16) & 0xFFFF, 16);
movk(ARMEmitter::Size::i64Bit, DestReg, uint32_t(ImmInstOne) & 0xFFFF);
LOGMAN_MSG_A_FMT("Unscaled offset is too large");
FEX_UNREACHABLE;
}
SetCursorOffset(OriginalOffset);
+1 -1
+1 -1
+1 -1
+3 -5
View File
@@ -74,11 +74,9 @@ add_compile_options($<$<COMPILE_LANGUAGE:CXX>:-fno-strict-aliasing> $<$<COMPILE_
add_subdirectory(Source/)
if (NOT BUILD_STEAM_SUPPORT)
install (DIRECTORY include/FEXCore ${CMAKE_BINARY_DIR}/include/FEXCore
DESTINATION include
COMPONENT Development)
endif()
install (DIRECTORY include/FEXCore ${CMAKE_BINARY_DIR}/include/FEXCore
DESTINATION include
COMPONENT Development)
if (BUILD_TESTING)
add_subdirectory(unittests/)
-9
View File
@@ -200,15 +200,6 @@ def print_man_environment_tail():
],
"''", True)
print_man_env_option(
"APP_CACHE_LOCATION",
[
"Allows the user to override where FEX stores and loads cache files",
"By default FEX will look in $XDG_CACHE_HOME/fex-emu/ or $HOME/.cache/fex-emu/",
"This will override the full path, trailing forward-slash is expected to exist",
],
"''", True)
def print_man_header():
header ='''.Dd {0}
.Dt FEX
+4 -5
View File
@@ -31,6 +31,7 @@ set (SRCS
Interface/Core/OpcodeDispatcher/X87.cpp
Interface/Core/OpcodeDispatcher/X87F64.cpp
Interface/Core/OpcodeDispatcher.cpp
Interface/Core/X86HelperGen.cpp
Interface/Core/ArchHelpers/Arm64Emitter.cpp
Interface/Core/Dispatcher/Dispatcher.cpp
Interface/Core/Interpreter/Fallbacks/InterpreterFallbacks.cpp
@@ -201,10 +202,8 @@ add_custom_target(CONFIG_INC
DEPENDS "${OUTPUT_MAN_NAME}"
DEPENDS "${OUTPUT_MAN_NAME_COMPRESS}")
if (NOT BUILD_STEAM_SUPPORT)
# Install the compressed man page
install(FILES ${OUTPUT_MAN_NAME_COMPRESS} COMPONENT Runtime DESTINATION ${MAN_DIR}/man1)
endif()
# Install the compressed man page
install(FILES ${OUTPUT_MAN_NAME_COMPRESS} COMPONENT Runtime DESTINATION ${MAN_DIR}/man1)
# Add in diagnostic colours if the option is available.
# Ninja code generator will kill colours if this isn't here
@@ -291,7 +290,7 @@ AddObject(${PROJECT_NAME}_object OBJECT)
AddLibrary(${PROJECT_NAME} STATIC)
AddLibrary(${PROJECT_NAME}_shared SHARED)
if (NOT MINGW_BUILD AND NOT BUILD_STEAM_SUPPORT)
if (NOT MINGW_BUILD)
install(TARGETS ${PROJECT_NAME}_shared
LIBRARY
DESTINATION ${CMAKE_INSTALL_LIBDIR}
+14 -19
View File
@@ -30,14 +30,14 @@ class Context;
}
namespace FEXCore::Config {
namespace detail {
namespace DefaultValues {
#define P(x) x
#define OPT_BASE(type, group, enum, json, default) const P(type) P(enum) = P(default);
#define OPT_STR(group, enum, json, default) const std::string_view P(enum) = P(default);
#define OPT_STRARRAY(group, enum, json, default) OPT_STR(group, enum, json, default)
#define OPT_STRENUM(group, enum, json, default) const uint64_t P(enum) = FEXCore::ToUnderlying(P(default));
#include <FEXCore/Config/ConfigValues.inl>
} // namespace detail
} // namespace DefaultValues
enum Paths {
PATH_DATA_DIR_LOCAL = 0,
@@ -134,7 +134,7 @@ public:
void Load();
template<typename T>
requires (!std::is_same_v<fextl::string, T> && !std::is_same_v<StringArrayType, T>)
requires (!std::is_same_v<fextl::string, T> && !std::is_same_v<DefaultValues::Type::StringArrayType, T>)
std::optional<T> GetConv(ConfigOption Option) {
const auto it = OptionMap.find(Option);
if (it == OptionMap.end()) {
@@ -142,7 +142,7 @@ public:
}
const auto& Value = it->second;
LOGMAN_THROW_A_FMT(!std::holds_alternative<StringArrayType>(Value), "Tried to get config of invalid type!");
LOGMAN_THROW_A_FMT(!std::holds_alternative<DefaultValues::Type::StringArrayType>(Value), "Tried to get config of invalid type!");
if (std::holds_alternative<T>(Value)) [[likely]] {
return std::get<T>(Value);
@@ -165,7 +165,7 @@ public:
private:
void MergeConfigMap(const LayerOptions& Options);
void MergeEnvironmentVariables(const ConfigOption& Option, const StringArrayType& Value);
void MergeEnvironmentVariables(const ConfigOption& Option, const DefaultValues::Type::StringArrayType& Value);
};
void MetaLayer::Load() {
@@ -181,7 +181,7 @@ void MetaLayer::Load() {
}
void MetaLayer::MergeEnvironmentVariables(const ConfigOption& Option, const StringArrayType& Value) {
void MetaLayer::MergeEnvironmentVariables(const ConfigOption& Option, const DefaultValues::Type::StringArrayType& Value) {
// Environment variables need a bit of additional work
// We want to merge the arrays rather than overwrite entirely
auto MetaEnvironment = OptionMap.find(Option);
@@ -193,7 +193,7 @@ void MetaLayer::MergeEnvironmentVariables(const ConfigOption& Option, const Stri
// If an environment variable exists in both current meta and in the incoming layer then the meta layer value is overwritten
fextl::unordered_map<fextl::string, fextl::string> LookupMap;
const auto AddToMap = [&LookupMap](const StringArrayType& Value) {
const auto AddToMap = [&LookupMap](const DefaultValues::Type::StringArrayType& Value) {
for (const auto& EnvVar : Value) {
const auto ItEq = EnvVar.find_first_of('=');
if (ItEq == fextl::string::npos) {
@@ -209,7 +209,7 @@ void MetaLayer::MergeEnvironmentVariables(const ConfigOption& Option, const Stri
}
};
AddToMap(std::get<StringArrayType>(MetaEnvironment->second));
AddToMap(std::get<DefaultValues::Type::StringArrayType>(MetaEnvironment->second));
AddToMap(Value);
// Now with the two layers merged in the map
@@ -225,8 +225,8 @@ void MetaLayer::MergeConfigMap(const LayerOptions& Options) {
// Insert this layer's options, overlaying previous options that exist here
for (auto& it : Options) {
if (it.first == FEXCore::Config::ConfigOption::CONFIG_ENV || it.first == FEXCore::Config::ConfigOption::CONFIG_HOSTENV) {
LOGMAN_THROW_A_FMT(std::holds_alternative<StringArrayType>(it.second), "Tried to get config of invalid type!");
MergeEnvironmentVariables(it.first, std::get<StringArrayType>(it.second));
LOGMAN_THROW_A_FMT(std::holds_alternative<DefaultValues::Type::StringArrayType>(it.second), "Tried to get config of invalid type!");
MergeEnvironmentVariables(it.first, std::get<DefaultValues::Type::StringArrayType>(it.second));
} else {
OptionMap.insert_or_assign(it.first, it.second);
}
@@ -423,7 +423,7 @@ bool Exists(ConfigOption Option) {
return Meta->OptionExists(Option);
}
std::optional<StringArrayType*> All(ConfigOption Option) {
std::optional<DefaultValues::Type::StringArrayType*> All(ConfigOption Option) {
return Meta->All(Option);
}
@@ -436,12 +436,6 @@ std::optional<T> GetConv(ConfigOption Option) {
return Meta->GetConv<T>(Option);
}
template std::optional<bool> GetConv(ConfigOption Option);
template std::optional<uint8_t> GetConv(ConfigOption Option);
template std::optional<int32_t> GetConv(ConfigOption Option);
template std::optional<uint32_t> GetConv(ConfigOption Option);
template std::optional<uint64_t> GetConv(ConfigOption Option);
void Set(ConfigOption Option, std::string_view Data) {
Meta->Set(Option, Data);
}
@@ -497,12 +491,13 @@ template Value<uint8_t>::Value(FEXCore::Config::ConfigOption _Option, uint8_t De
template Value<uint64_t>::Value(FEXCore::Config::ConfigOption _Option, uint64_t Default);
template<typename T>
void Value<T>::GetListIfExists(FEXCore::Config::ConfigOption Option, StringArrayType* List) {
void Value<T>::GetListIfExists(FEXCore::Config::ConfigOption Option, DefaultValues::Type::StringArrayType* List) {
auto Value = FEXCore::Config::All(Option);
List->clear();
if (Value) {
*List = **Value;
}
}
template void Value<StringArrayType>::GetListIfExists(FEXCore::Config::ConfigOption Option, StringArrayType* List);
template void Value<DefaultValues::Type::StringArrayType>::GetListIfExists(FEXCore::Config::ConfigOption Option,
DefaultValues::Type::StringArrayType* List);
} // namespace FEXCore::Config
@@ -16,13 +16,6 @@
"Maximum number of instruction to store in a block"
]
},
"EnableCodeCachingWIP": {
"Type": "bool",
"Default": "false",
"Desc": [
"Enable the code caching subsystem"
]
},
"HostFeatures": {
"Type": "strenum",
"Default": "FEXCore::Config::HostFeatures::OFF",
@@ -101,13 +94,6 @@
"Desc": [
"Scales the cycle counter on systems that have low frequencies."
]
},
"CPUFeatureRegisters": {
"Type": "str",
"Default": "",
"Desc": [
"Allows overriding cpu feature flags for manual testing"
]
}
},
"Emulation": {
@@ -443,13 +429,6 @@
"This is required to ensure a split-lock doesn't tear inside the process"
]
},
"KernelUnalignedAtomicBackpatching": {
"Type": "bool",
"Default": "true",
"Desc": [
"When the kernel unaligned atomic handler is enabled, use backpatching to reduce kernel context switches."
]
},
"VolatileMetadata": {
"Type": "bool",
"Default": "true",
+4 -35
View File
@@ -4,6 +4,7 @@
#include "Common/JitSymbols.h"
#include "Interface/Core/CPUBackend.h"
#include "Interface/Core/CPUID.h"
#include "Interface/Core/X86HelperGen.h"
#include <Interface/IR/IntrusiveIRList.h>
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/Context.h>
@@ -72,34 +73,12 @@ public:
ContextImpl& CTX;
bool IsGeneratingCache = false;
uint64_t ComputeCodeMapId(std::string_view Filename, int FD) override;
void LoadData(Core::InternalThreadState&, std::byte* MappedCacheFile, const ExecutableFileSectionInfo&) override;
bool SaveData(Core::InternalThreadState&, int TargetFD, const ExecutableFileSectionInfo&, uint64_t SerializedBaseAddress) override;
void InitiateCacheGeneration() override {
IsGeneratingCache = true;
}
/**
* Applies a set of FEX relocations to the given code section.
*
* FEX relocations describe runtime-dependencies of FEX-generated code.
* When loading a code cache, they are used to move cached code to the
* dynamically chosen base address of the guest binary.
*
* Conversely, relocations are applied in reverse when writing code caches
* to ensure consistency across generation runs.
*
* Note that FEX relocations are unrelated to ELF/PE relocations.
*
* @param GuestDelta Guest address offset to apply to RIP-relative data
* @param ForStorage True for serializing data (producing deterministic output); false for de-serializing it (resolving dynamic symbols)
*
* @return Returns true on success
*/
[[nodiscard]]
bool ApplyCodeRelocations(uint64_t GuestDelta, std::span<std::byte> Code, std::span<const CPU::Relocation> Relocations, bool ForStorage);
};
class ContextImpl final : public FEXCore::Context::Context, public CPU::CodeBufferManager {
@@ -176,17 +155,7 @@ public:
return CodeCache;
}
void SetCodeMapWriter(fextl::unique_ptr<CodeMapWriter> Writer) override {
CodeMapWriter = std::move(Writer);
}
void FlushAndCloseCodeMap() override {
if (CodeMapWriter) {
CodeMapWriter.reset();
}
}
void OnCodeBufferAllocated(const std::shared_ptr<CPU::CodeBuffer>&) override;
void OnCodeBufferAllocated(const std::shared_ptr<CPU::CodeBuffer> &) override;
void ClearCodeCache(FEXCore::Core::InternalThreadState* Thread, bool NewCodeBuffer = true) override;
void InvalidateCodeBuffersCodeRange(uint64_t Start, uint64_t Length) override;
void InvalidateThreadCachedCodeRange(FEXCore::Core::InternalThreadState* Thread, uint64_t Start, uint64_t Length) override;
@@ -256,9 +225,9 @@ public:
FEXCore::ThunkHandler* ThunkHandler {};
fextl::unique_ptr<FEXCore::CPU::Dispatcher> Dispatcher;
CodeCache CodeCache;
fextl::unique_ptr<CodeMapWriter> CodeMapWriter;
SignalDelegator* SignalDelegation {};
X86GeneratedCode X86CodeGen;
ContextImpl(const FEXCore::HostFeatures& Features);
@@ -298,7 +267,7 @@ public:
FEXCore::Utils::PooledAllocatorVirtual OpDispatcherAllocator {"FEXMem_OpDispatcher"};
FEXCore::Utils::PooledAllocatorVirtual FrontendAllocator {"FEXMem_Frontend"};
FEXCore::Utils::PooledAllocatorVirtualWithGuard CPUBackendAllocator {"FEXMem_CPUBackend"};
FEXCore::Utils::PooledAllocatorVirtual CPUBackendAllocator {"FEXMem_CPUBackend"};
// If Atomic-based TSO emulation is enabled or not.
bool IsAtomicTSOEnabled() const {
@@ -105,12 +105,9 @@ constexpr ARMEmitter::PRegister PRED_TMP_32B = ARMEmitter::PReg::p7;
// This class contains common emitter utility functions that can
// be used by both Arm64 JIT and ARM64 Dispatcher
class Arm64Emitter : public ARMEmitter::Emitter {
public:
protected:
Arm64Emitter(FEXCore::Context::ContextImpl* ctx, void* EmissionPtr = nullptr, size_t size = 0);
void LoadConstant(ARMEmitter::Size s, ARMEmitter::Register Reg, uint64_t Constant, bool NOPPad = false);
protected:
FEXCore::Context::ContextImpl* EmitterCTX;
std::span<const ARMEmitter::Register> StaticRegisters {};
@@ -120,6 +117,8 @@ protected:
std::span<const ARMEmitter::VRegister> GeneralFPRegisters {};
uint32_t PairRegisters = 0;
void LoadConstant(ARMEmitter::Size s, ARMEmitter::Register Reg, uint64_t Constant, bool NOPPad = false);
void FillSpecialRegs(ARMEmitter::Register TmpReg, ARMEmitter::Register TmpReg2, bool SetFIZ, bool SetPredRegs);
// Correlate an ARM register back to an x86 register index.
+1 -1
View File
@@ -161,7 +161,7 @@ namespace CPU {
virtual CompiledCode CompileCode(uint64_t Entry, uint64_t Size, bool SingleInst, const FEXCore::IR::IRListView* IR,
FEXCore::Core::DebugData* DebugData, bool CheckTF) = 0;
virtual fextl::vector<FEXCore::CPU::Relocation> TakeRelocations(uint64_t GuestBaseAddress) = 0;
virtual fextl::vector<FEXCore::CPU::Relocation> TakeRelocations() = 0;
virtual void ClearCache() {}
+8 -11
View File
@@ -14,7 +14,6 @@ $end_info$
#include <FEXCore/Core/CPUID.h>
#include <FEXCore/Core/HostFeatures.h>
#include <FEXCore/Utils/FileLoading.h>
#include <FEXCore/Utils/MathUtils.h>
#include <FEXCore/fextl/string.h>
#include <FEXHeaderUtils/Syscalls.h>
@@ -89,7 +88,6 @@ namespace ProductNames {
static const char ARM_Blizzard_M2Pro[] = "Apple Blizzard (M2 Pro)";
static const char ARM_Avalanche_M2Max[] = "Apple Avalanche (M2 Max)";
static const char ARM_Blizzard_M2Max[] = "Apple Blizzard (M2 Max)";
static const char ARM_AppleSilicon[] = "Apple Silicon";
static const char ARM_ORYON_1[] = "Oryon-1";
static const char ARM_Ampere_1[] = "AmpereOne";
@@ -190,7 +188,6 @@ void CPUIDEmu::SetupHostHybridFlag() {
{0x61, 0x029, 1, ProductNames::ARM_Firestorm_M1Max}, // Apple Firestorm (M1 Max)
{0x61, 0x025, 1, ProductNames::ARM_Firestorm_M1Pro}, // Apple Firestorm (M1 Pro)
{0x61, 0x023, 1, ProductNames::ARM_Firestorm_M1}, // Apple Firestorm (M1)
{0x61, 0, 1, ProductNames::ARM_AppleSilicon}, // QEmu Apple Silicon
{0x41, 0xd8c, 1, ProductNames::ARM_C1Ultra}, // C1-Ultra
{0x41, 0xd90, 1, ProductNames::ARM_C1Premium}, // C1-Premium
@@ -444,10 +441,10 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_01h(uint32_t Leaf) const {
Res.eax = FAMILY_IDENTIFIER;
Res.ebx = 0 | // Brand index
(8 << 8) | // Cache line size in bytes
(Cores << 16) | // Number of addressable IDs for the logical cores in the physical CPU
(GetCPUID() << 24); // Local APIC ID
Res.ebx = 0 | // Brand index
(8 << 8) | // Cache line size in bytes
(Cores << 16) | // Number of addressable IDs for the logical cores in the physical CPU
(0 << 24); // Local APIC ID
Res.ecx = (1 << 0) | // SSE3
(CTX->HostFeatures.SupportsPMULL_128Bit << 1) | // PCLMULQDQ
@@ -510,7 +507,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_01h(uint32_t Leaf) const {
(1 << 25) | // SSE
(1 << 26) | // SSE2
(0 << 27) | // Self Snoop
(0 << 28) | // (HTT) Max APIC IDs reserved field is valid
(1 << 28) | // Max APIC IDs reserved field is valid
(1 << 29) | // Thermal monitor
(0 << 30) | // Reserved
(0 << 31); // Pending break enable
@@ -1097,9 +1094,9 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0008h(uint32_t Leaf) con
(CTX->HostFeatures.SupportsCLZERO << 0); // CLZERO support
uint32_t CoreCount = Cores - 1;
Res.ecx = (0 << 16) | // PerfTscSize: Performance timestamp count size
(std::bit_ceil(Cores) << 12) | // ApicIdSize: Number of bits in ApicID
(CoreCount << 0); // Count count subtract one
Res.ecx = (0 << 16) | // PerfTscSize: Performance timestamp count size
((uint32_t)std::log2(CoreCount + 1) << 12) | // ApicIdSize: Number of bits in ApicID
(CoreCount << 0); // Count count subtract one
return Res;
}
+1 -1
View File
@@ -277,7 +277,7 @@ private:
// 0: Highest function parameter and ID
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 1: Processor info
{SupportsConstant::NONCONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 2: Cache and TLB info
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 3: Serial Number(previously), now reserved
+1 -346
View File
@@ -1,213 +1,12 @@
// SPDX-License-Identifier: MIT
#include "Utils/SpinWaitLock.h"
#include <Interface/Context/Context.h>
#include <Interface/Core/ArchHelpers/Arm64Emitter.h>
#include <Interface/Core/JIT/Relocations.h>
#include <Interface/Core/LookupCache.h>
#include <FEXCore/Core/Thunks.h>
#include <FEXCore/HLE/SourcecodeResolver.h>
#include <FEXHeaderUtils/Filesystem.h>
#include <git_version.h>
#include <xxhash.h>
#include <fstream>
namespace FEXCore {
#if __clang_major__ < 16
ExecutableFileInfo::ExecutableFileInfo(fextl::unique_ptr<HLE::SourcecodeMap> Map, uint64_t FileId, fextl::string Filename)
: SourcecodeMap(std::move(Map))
, FileId(FileId)
, Filename(Filename) {}
#endif
ExecutableFileInfo::~ExecutableFileInfo() = default;
fextl::string CodeMap::GetBaseFilename(const ExecutableFileInfo& MainExecutable, bool AddNombSuffix) {
auto FileId = MainExecutable.FileId;
std::string_view base_filename = FHU::Filesystem::GetFilename(std::string_view {MainExecutable.Filename});
if (FileId != 0xffff'ffff'ffff'ffff) {
return fextl::fmt::format("{}-{:016x}{}", base_filename, MainExecutable.FileId, AddNombSuffix ? "-nomb" : "");
}
return "";
}
fextl::map<CodeMapFileId, CodeMap::ParsedContents> CodeMap::ParseCodeMap(std::ifstream& File) {
fextl::map<CodeMapFileId, CodeMap::ParsedContents> Ret;
while (true) {
Entry Entry;
File.read(reinterpret_cast<char*>(&Entry), sizeof(Entry));
if (!File) {
break;
}
if (Entry.FileId == LoadExternalLibrary.FileId && Entry.BlockOffset == LoadExternalLibrary.BlockOffset) {
ExternalLibraryInfo Info;
File.read(reinterpret_cast<char*>(&Info), sizeof(Info));
fextl::string Filename;
std::getline(File, Filename, '\0');
// Align to 4-byte boundary
char Null[4];
File.read(Null, AlignUp(Filename.size() + 1, 4) - Filename.size() - 1);
if (!File) {
break;
}
Ret[Info.ExternalFileId].Filename = std::move(Filename);
} else if (Entry.FileId == SetExecutableFileId {}.Marker.FileId && Entry.BlockOffset == SetExecutableFileId {}.Marker.BlockOffset) {
CodeMapFileId ExecutableFileId;
File.read(reinterpret_cast<char*>(&ExecutableFileId), sizeof(ExecutableFileId));
if (!File) {
break;
}
Ret[ExecutableFileId].IsExecutable = true;
} else {
if (!Ret.contains(Entry.FileId)) {
LogMan::Msg::EFmt("Code map referenced unknown file id {:016x}", Entry.FileId);
} else {
Ret[Entry.FileId].Blocks.insert(Entry.BlockOffset);
}
}
if (!File) {
break;
}
}
return Ret;
}
CodeMapWriter::CodeMapWriter(CodeMapOpener& Opener, bool OpenEagerly)
: Buffer(4096)
, FileOpener(Opener) {
if (OpenEagerly) {
CodeMapFD = FileOpener.OpenCodeMapFile();
}
}
CodeMapWriter::~CodeMapWriter() {
if (CodeMapFD.value_or(-1) != -1) {
Flush(BufferOffset);
close(*CodeMapFD);
}
}
bool CodeMapWriter::IsWriteEnabled(const ExecutableFileSectionInfo& Section) {
if (CodeMapFD == -1) {
return false;
}
// PV libraries can't yet be read by FEXServer, so skip dumping them
if (Section.FileInfo.Filename.starts_with("/run/pressure-vessel")) {
return false;
}
if (CodeMapFD) {
return true;
}
// Acquire mutex and re-check CodeMapFD to avoid race conditions
auto lk = std::unique_lock {Mutex};
if (!CodeMapFD) {
CodeMapFD = FileOpener.OpenCodeMapFile();
}
return CodeMapFD != -1;
}
void CodeMapWriter::Flush(size_t Offset) {
// Acquire exclusive lock and flush circular buffer
std::unique_lock Lock {Mutex};
Flush(Offset, Lock);
}
void CodeMapWriter::Flush(size_t Offset, std::unique_lock<std::shared_mutex>&) {
write(*CodeMapFD, Buffer.data(), Offset);
BufferOffset = 0;
}
void CodeMapWriter::AppendBlock(const FEXCore::ExecutableFileSectionInfo& SectionInfo, uint64_t BlockEntry) {
if (!IsWriteEnabled(SectionInfo)) {
return;
}
BlockEntry -= SectionInfo.FileStartVA;
if (BlockEntry > std::numeric_limits<uint32_t>::max()) {
ERROR_AND_DIE_FMT("Cannot write code map");
}
// Register new library if not already known
bool NewLibraryLoad = false;
{
// Check prior registration with shared lock
std::shared_lock Lock {Mutex};
NewLibraryLoad = !KnownFileIds.contains(SectionInfo.FileInfo.FileId);
}
if (NewLibraryLoad) {
// Register to map with exclusive lock
std::unique_lock Lock {Mutex};
NewLibraryLoad &= KnownFileIds.insert(SectionInfo.FileInfo.FileId).second;
}
if (NewLibraryLoad) {
// Add entry to code map
AppendLibraryLoad(SectionInfo.FileInfo);
}
// Register the actual code block
CodeMap::Entry DataEntry {SectionInfo.FileInfo.FileId, static_cast<uint32_t>(BlockEntry)};
AppendData(std::as_bytes(std::span {&DataEntry, 1}));
}
void CodeMapWriter::AppendLibraryLoad(const FEXCore::ExecutableFileInfo& FileInfo) {
// See CodeMap::ExternalLibraryInfo
auto ExternalFileId = FileInfo.FileId;
auto TotalSize = AlignUp(sizeof(CodeMap::LoadExternalLibrary) + sizeof(ExternalFileId) + FileInfo.Filename.size() + 1, 4);
const auto Data = reinterpret_cast<char*>(alloca(TotalSize));
auto WritePtr = std::copy_n(reinterpret_cast<const char*>(&CodeMap::LoadExternalLibrary), sizeof(CodeMap::LoadExternalLibrary), Data);
WritePtr = std::copy_n(reinterpret_cast<const char*>(&ExternalFileId), sizeof(ExternalFileId), WritePtr);
WritePtr = std::copy(FileInfo.Filename.begin(), FileInfo.Filename.end(), WritePtr);
std::fill(WritePtr, Data + TotalSize, 0);
AppendData(std::as_bytes(std::span {Data, TotalSize}));
}
void CodeMapWriter::AppendSetMainExecutable(const FEXCore::ExecutableFileInfo& FileInfo) {
CodeMap::SetExecutableFileId Data {.ExecutableFileId = FileInfo.FileId};
AppendData(std::span {reinterpret_cast<const std::byte*>(&Data), sizeof(Data)});
}
void CodeMapWriter::AppendData(std::span<const std::byte> Data) {
std::shared_lock Lock {Mutex};
auto Offset = BufferOffset.fetch_add(Data.size_bytes());
if (Offset + Data.size_bytes() > Buffer.size()) {
// Acquire exclusive lock and flush the buffer.
// Under heavy pressure, multiple threads may observe an exhausted buffer simultaneously.
// The thread with the last in-bounds Offset is responsible for flushing the buffer.
Lock.unlock();
bool IsResponsibleForFlush = false;
{
std::unique_lock ExclusiveLock {Mutex};
IsResponsibleForFlush = (Offset <= Buffer.size());
if (IsResponsibleForFlush) {
Flush(Offset, ExclusiveLock);
}
}
if (!IsResponsibleForFlush) {
// Wait for the buffer to be flushed on the responsible thread
Utils::SpinWaitLock::WaitPred<std::less_equal<>, size_t>(reinterpret_cast<size_t*>(&BufferOffset), Buffer.size());
}
AppendData(Data);
return;
}
memcpy(&Buffer.at(Offset), Data.data(), Data.size_bytes());
}
} // namespace FEXCore
namespace FEXCore::Context {
@@ -216,156 +15,12 @@ CodeCache::CodeCache(ContextImpl& CTX_)
: CTX(CTX_) {}
CodeCache::~CodeCache() = default;
uint64_t CodeCache::ComputeCodeMapId(std::string_view Filename, int FD) {
if (Filename.empty()) {
return 0xffff'ffff'ffff'ffff;
}
// For now, we just use the file path as an identifier.
// TODO: Ensure the hash is unique enough to distinguish executables while remaining independent of the installation location
return XXH3_64bits(Filename.data(), Filename.size());
}
struct CodeCacheHeader {
char Magic[4] = {'F', 'X', 'C', 'C'};
uint32_t FormatVersion = 1;
char FEXVersion[8] = {};
uint32_t NumBlocks;
uint32_t NumCodePages;
uint32_t CodeBufferSize;
uint32_t NumRelocations;
uint64_t SerializedBaseAddress;
// TODO: Consider including information from LookupCache.BlockLinks
};
void CodeCache::LoadData(Core::InternalThreadState& Thread, std::byte* MappedCacheFile, const ExecutableFileSectionInfo& GuestRIPLookup) {
// TODO
}
template<typename T>
static constexpr auto IsOrderedContainer(const T&) -> std::false_type;
template<typename... T>
static constexpr auto IsOrderedContainer(const std::map<T...>&) -> std::true_type;
template<typename... T>
static constexpr auto IsOrderedContainer(const std::set<T...>&) -> std::true_type;
bool CodeCache::SaveData(Core::InternalThreadState& Thread, int fd, const ExecutableFileSectionInfo& SourceBinary, uint64_t SerializedBaseAddress) {
auto CodeBuffer = CTX.GetLatest();
auto& LookupCache = *Thread.LookupCache->Shared;
auto Relocations = Thread.CPUBackend->TakeRelocations(SourceBinary.FileStartVA);
// Write file header
CodeCacheHeader header;
memcpy(&header.FEXVersion[0], GIT_SHORT_HASH, strlen(GIT_SHORT_HASH));
header.NumBlocks = LookupCache.BlockList.size();
header.NumCodePages = LookupCache.CodePages.size();
header.CodeBufferSize = CTX.LatestOffset;
header.NumRelocations = Relocations.size();
header.SerializedBaseAddress = SerializedBaseAddress;
::write(fd, &header, sizeof(header));
// Dump guest<->host block mappings
{
// Cache contents must be deterministic, so copy the unordered block list and then sort by key
static_assert(!decltype(IsOrderedContainer(LookupCache.BlockList))::value, "Already deterministic; drop temporary container");
fextl::vector<std::pair<uint64_t, const GuestToHostMap::BlockEntry*>> BlockList;
BlockList.reserve(LookupCache.BlockList.size());
for (auto& [Guest, BlockEntry] : LookupCache.BlockList) {
static_assert(sizeof(Guest) == 8, "Breaking change in code cache data layout");
BlockList.emplace_back(Guest, &BlockEntry);
}
std::ranges::sort(BlockList);
for (auto [Guest, Host] : BlockList) {
static_assert(sizeof(Host->HostCode) == 8, "Breaking change in code cache data layout");
static_assert(sizeof(Host->CodePages[0]) == 8, "Breaking change in code cache data layout");
Guest -= SourceBinary.FileStartVA;
::write(fd, &Guest, sizeof(Guest));
uint64_t HostCode = Host->HostCode - reinterpret_cast<uintptr_t>(CodeBuffer->Ptr);
::write(fd, &HostCode, sizeof(HostCode));
uint64_t NumCodePages = Host->CodePages.size();
::write(fd, &NumCodePages, sizeof(NumCodePages));
LOGMAN_THROW_A_FMT(std::ranges::is_sorted(Host->CodePages), "Code pages aren't sorted");
for (auto CodePage : Host->CodePages) {
CodePage -= SourceBinary.FileStartVA;
::write(fd, &CodePage, sizeof(CodePage));
}
}
}
// Dump relocations
static_assert(sizeof(Relocations[0]) == 48, "Breaking change in code cache data layout");
::write(fd, Relocations.data(), Relocations.size() * sizeof(Relocations[0]));
// Pad to next page in file so that the CodeBuffer can be mmap'ed into process on load
char Zero[64] {};
auto Off = lseek(fd, 0, SEEK_CUR);
while (Off != AlignUp(Off, Utils::FEX_PAGE_SIZE)) {
auto BytesToWrite = std::min(AlignUp(Off, Utils::FEX_PAGE_SIZE) - Off, sizeof(Zero));
::write(fd, Zero, BytesToWrite);
Off += BytesToWrite;
}
// Dump the host code (relocated for position-independent serialization)
std::vector CodeBufferData(reinterpret_cast<std::byte*>(CodeBuffer->Ptr), reinterpret_cast<std::byte*>(CodeBuffer->Ptr) + CTX.LatestOffset);
if (!ApplyCodeRelocations(SerializedBaseAddress, CodeBufferData, Relocations, true)) {
LOGMAN_THROW_A_FMT(false, "Failed to apply code relocations");
return false;
}
::write(fd, CodeBufferData.data(), CodeBufferData.size());
// Dump code pages
static_assert(decltype(IsOrderedContainer(LookupCache.CodePages))::value, "Non-deterministic data source");
for (auto& [Page, Entrypoints] : LookupCache.CodePages) {
static_assert(sizeof(Page) == 8, "Breaking change in code cache data layout");
::write(fd, &Page, sizeof(Page));
uint64_t NumEntrypoints = Entrypoints.size();
::write(fd, &NumEntrypoints, sizeof(NumEntrypoints));
::write(fd, Entrypoints.data(), Entrypoints.size() * sizeof(Entrypoints[0]));
}
return true;
}
bool CodeCache::ApplyCodeRelocations(uint64_t GuestEntry, std::span<std::byte> Code,
std::span<const FEXCore::CPU::Relocation> EntryRelocations, bool ForStorage) {
CPU::Arm64Emitter Emitter(&CTX, Code.data(), Code.size_bytes());
for (size_t j = 0; j < EntryRelocations.size(); ++j) {
const FEXCore::CPU::Relocation& Reloc = EntryRelocations[j];
Emitter.SetCursorOffset(Reloc.Header.Offset);
switch (Reloc.Header.Type) {
case FEXCore::CPU::RelocationTypes::RELOC_NAMED_SYMBOL_LITERAL: {
// Generate a literal so we can place it
uint64_t Pointer = ForStorage ? 0 : GetNamedSymbolLiteral(CTX, Reloc.NamedSymbolLiteral.Symbol);
Emitter.dc64(Pointer);
break;
}
case FEXCore::CPU::RelocationTypes::RELOC_NAMED_THUNK_MOVE: {
uint64_t Pointer = ForStorage ? 0 : reinterpret_cast<uint64_t>(CTX.ThunkHandler->LookupThunk(Reloc.NamedThunkMove.Symbol));
if (Pointer == ~0ULL) {
return false;
}
Emitter.LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Register(Reloc.NamedThunkMove.RegisterIndex), Pointer, true);
break;
}
case FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_LITERAL: {
Emitter.dc64(GuestEntry + Reloc.GuestRIP.GuestRIP);
break;
}
case FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_MOVE: {
uint64_t Pointer = Reloc.GuestRIP.GuestRIP + GuestEntry;
Emitter.LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Register(Reloc.GuestRIP.RegisterIndex), Pointer, true);
break;
}
default: ERROR_AND_DIE_FMT("Unknown relocation type {}", ToUnderlying(Reloc.Header.Type));
}
}
// TODO
return true;
}
+6 -18
View File
@@ -437,10 +437,6 @@ void ContextImpl::UnlockAfterFork(FEXCore::Core::InternalThreadState* LiveThread
Profiler::PostForkAction(Child);
if (Child) {
if (CodeMapWriter) {
CodeMapWriter->ResetAfterFork();
}
CodeInvalidationMutex.StealAndDropActiveLocks();
if (Config.StrictInProcessSplitLocks) {
StrictSplitLockMutex = 0;
@@ -469,7 +465,7 @@ void ContextImpl::OnCodeBufferAllocated(const fextl::shared_ptr<CPU::CodeBuffer>
}
{
std::scoped_lock lk {CodeBufferListLock};
std::scoped_lock lk{CodeBufferListLock};
CodeBufferList.emplace_back(Buffer);
}
}
@@ -724,8 +720,7 @@ ContextImpl::CompileCodeResult ContextImpl::CompileCode(FEXCore::Core::InternalT
if (SourcecodeResolver && Config.GDBSymbols()) {
auto MappedSection = SyscallHandler->LookupExecutableFileSection(*Thread, GuestRIP);
if (MappedSection) {
MappedSection->FileInfo.SourcecodeMap =
SourcecodeResolver->GenerateMap(MappedSection->FileInfo.Filename, CodeMap::GetBaseFilename(MappedSection->FileInfo, false));
MappedSection->FileInfo.SourcecodeMap = SourcecodeResolver->GenerateMap(MappedSection->FileInfo.Filename, MappedSection->FileInfo.FileId);
}
}
@@ -864,13 +859,6 @@ uintptr_t ContextImpl::CompileBlock(FEXCore::Core::CpuStateFrame* Frame, uint64_
Thread->LookupCache->AddBlockMapping(Thread, GuestAddr, CodePages, HostAddr);
}
if (CodeMapWriter) {
auto Region = SyscallHandler->LookupExecutableFileSection(*Thread, GuestRIP);
if (Region && Region->FileStartVA != 0) {
CodeMapWriter->AppendBlock(*Region, GuestRIP);
}
}
return (uintptr_t)CodePtr;
}
@@ -898,11 +886,11 @@ uintptr_t ContextImpl::CompileSingleStep(FEXCore::Core::CpuStateFrame* Frame, ui
void ContextImpl::InvalidateCodeBuffersCodeRange(uint64_t Start, uint64_t Length) {
FEXCORE_PROFILE_SCOPED("InvalidateCodeBuffersCodeRange");
LOGMAN_THROW_A_FMT(CodeInvalidationMutex.try_lock() == false, "CodeInvalidationMutex needs to be unique_locked here");
LogMan::Throw::AFmt(CodeInvalidationMutex.try_lock() == false, "CodeInvalidationMutex needs to be unique_locked here");
std::scoped_lock lk {CodeBufferListLock};
auto it = CodeBufferList.begin();
while (it != CodeBufferList.end()) {
if (auto Strong = it->lock()) {
if (auto Strong = it->lock(); Strong) {
Strong->LookupCache->InvalidateRange(Start, Length);
it++;
} else {
@@ -912,9 +900,9 @@ void ContextImpl::InvalidateCodeBuffersCodeRange(uint64_t Start, uint64_t Length
}
void ContextImpl::InvalidateThreadCachedCodeRange(FEXCore::Core::InternalThreadState* Thread, uint64_t Start, uint64_t Length) {
LOGMAN_THROW_A_FMT(CodeInvalidationMutex.try_lock() == false, "CodeInvalidationMutex needs to be unique_locked here");
LogMan::Throw::AFmt(CodeInvalidationMutex.try_lock() == false, "CodeInvalidationMutex needs to be unique_locked here");
// Ensures now-modified mappings aren't cached as being in their previous non-executable state.
// Ensures now-modified mappings aren't cached as being in their previous non-executable state.
// Accessing FrontendDecoder is safe as the thread's code invalidation mutex must be locked here.
Thread->FrontendDecoder->ResetExecutableRangeCache();
@@ -5,6 +5,7 @@
#include "Interface/Core/CPUBackend.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Core/LookupCache.h"
#include "Interface/Core/X86HelperGen.h"
#include "Utils/MemberFunctionToPointer.h"
#include <FEXCore/Config/Config.h>
@@ -97,7 +98,7 @@ void Dispatcher::EmitDispatcher() {
ARMEmitter::BiDirectionalLabel LoopTop {};
#ifdef _M_ARM_64EC
(void)b(&LoopTop);
b(&LoopTop);
AbsoluteLoopTopAddressEnterECFillSRA = GetCursorAddress<uint64_t>();
ldr(STATE, EC_ENTRY_CPUAREA_REG, CPU_AREA_EMULATOR_DATA_OFFSET);
@@ -105,10 +106,10 @@ void Dispatcher::EmitDispatcher() {
ldr(RipReg, STATE_PTR(CpuStateFrame, State.rip));
// Force a single instruction block if ENTRY_FILL_SRA_SINGLE_INST_REG is nonzero entering the JIT, used for inline SMC handling.
(void)cbnz(ARMEmitter::Size::i32Bit, ENTRY_FILL_SRA_SINGLE_INST_REG, &CompileSingleStep);
cbnz(ARMEmitter::Size::i32Bit, ENTRY_FILL_SRA_SINGLE_INST_REG, &CompileSingleStep);
// Enter JIT
(void)b(&LoopTop);
b(&LoopTop);
AbsoluteLoopTopAddressEnterEC = GetCursorAddress<uint64_t>();
// Load ThreadState and write the target PC there
@@ -129,7 +130,7 @@ void Dispatcher::EmitDispatcher() {
ldp<ARMEmitter::IndexType::OFFSET>(TMP1, TMP2, REG_CALLRET_SP);
// EC_CALL_CHECKER_PC_REG is REG_PF which isn't touched by any of the above
sub(ARMEmitter::Size::i64Bit, TMP1, EC_CALL_CHECKER_PC_REG, TMP1);
(void)cbnz(ARMEmitter::Size::i64Bit, TMP1, &LoopTop);
cbnz(ARMEmitter::Size::i64Bit, TMP1, &LoopTop);
// If the entry at the TOS is for the target address, pop it and return to the JIT code
add(ARMEmitter::Size::i64Bit, REG_CALLRET_SP, REG_CALLRET_SP, 0x10);
@@ -488,7 +489,7 @@ void Dispatcher::EmitDispatcher() {
// Now push the callback return trampoline to the guest stack
// Guest will be misaligned because calling a thunk won't correct the guest's stack once we call the callback from the host
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, CTX->SignalDelegation->GetThunkCallbackRET());
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, CTX->X86CodeGen.CallbackReturn);
ldr(ARMEmitter::XReg::x2, STATE_PTR(CpuStateFrame, State.gregs[X86State::REG_RSP]));
sub(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r2, ARMEmitter::Reg::r2, CTX->Config.Is64BitMode ? 16 : 12);
@@ -51,10 +51,6 @@ public:
}
#endif
uint64_t GetExitFunctionLinkerAddress() const {
return ExitFunctionLinkerAddress;
}
SignalDelegatorConfig MakeSignalDelegatorConfig() const;
protected:
@@ -9,6 +9,7 @@ $end_info$
#include "Interface/Context/Context.h"
#include "Interface/Core/Frontend.h"
#include "Interface/Core/X86Tables/X86Tables.h"
#include "Interface/Core/X86HelperGen.h"
#include "Interface/Core/LookupCache.h"
#include <array>
@@ -89,6 +90,11 @@ Decoder::Decoder(FEXCore::Core::InternalThreadState* Thread)
}
bool Decoder::CheckRangeExecutable(uint64_t Address, uint64_t Size) {
// Treat FEX-internal X86 callbacks as always executable
if (EntryPoint == CTX->X86CodeGen.CallbackReturn) {
return true;
}
while (Address < ExecutableRangeBase || Address + Size > ExecutableRangeEnd) {
auto RangeInfo = CTX->SyscallHandler->QueryGuestExecutableRange(Thread, Address);
ExecutableRangeBase = RangeInfo.Base;
+2 -1
View File
@@ -50,13 +50,14 @@ DEF_OP(EntrypointOffset) {
auto Op = IROp->C<IR::IROp_EntrypointOffset>();
auto Constant = Entry + Op->Offset;
auto Dst = GetReg(Node);
uint64_t Mask = ~0ULL;
const auto OpSize = IROp->Size;
if (OpSize == IR::OpSize::i32Bit) {
Mask = 0xFFFF'FFFFULL;
}
InsertGuestRIPMove(GetReg(Node), Constant & Mask);
LoadConstant(ARMEmitter::Size::i64Bit, Dst, Constant & Mask);
}
DEF_OP(InlineConstant) {
@@ -11,18 +11,23 @@ $end_info$
#include <FEXCore/Core/Thunks.h>
namespace FEXCore::CPU {
uint64_t GetNamedSymbolLiteral(FEXCore::Context::ContextImpl& CTX, FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op) {
uint64_t Arm64JITCore::GetNamedSymbolLiteral(FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op) {
switch (Op) {
case FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol::SYMBOL_LITERAL_EXITFUNCTION_LINKER:
return CTX.Dispatcher->GetExitFunctionLinkerAddress();
default: ERROR_AND_DIE_FMT("Unknown named symbol literal: {}", static_cast<uint32_t>(Op));
return ThreadState->CurrentFrame->Pointers.Common.ExitFunctionLinker;
break;
default: ERROR_AND_DIE_FMT("Unknown named symbol literal: {}", static_cast<uint32_t>(Op)); break;
}
return ~0ULL;
}
void Arm64JITCore::InsertNamedThunkRelocation(ARMEmitter::Register Reg, const IR::SHA256Sum& Sum) {
Relocation MoveABI {};
MoveABI.NamedThunkMove.Header = {.Offset = GetCursorOffset(), .Type = FEXCore::CPU::RelocationTypes::RELOC_NAMED_THUNK_MOVE};
MoveABI.NamedThunkMove.Header.Type = FEXCore::CPU::RelocationTypes::RELOC_NAMED_THUNK_MOVE;
// Offset is the offset from the entrypoint of the block
auto CurrentCursor = GetCursorAddress<uint8_t*>();
MoveABI.NamedThunkMove.Offset = CurrentCursor - CodeData.BlockBegin;
MoveABI.NamedThunkMove.Symbol = Sum;
MoveABI.NamedThunkMove.RegisterIndex = Reg.Idx();
@@ -33,9 +38,9 @@ void Arm64JITCore::InsertNamedThunkRelocation(ARMEmitter::Register Reg, const IR
}
Arm64JITCore::NamedSymbolLiteralPair Arm64JITCore::InsertNamedSymbolLiteral(FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op) {
uint64_t Pointer = GetNamedSymbolLiteral(*CTX, Op);
uint64_t Pointer = GetNamedSymbolLiteral(Op);
NamedSymbolLiteralPair Lit {
Arm64JITCore::NamedSymbolLiteralPair Lit {
.Lit = Pointer,
.MoveABI =
{
@@ -43,72 +48,92 @@ Arm64JITCore::NamedSymbolLiteralPair Arm64JITCore::InsertNamedSymbolLiteral(FEXC
{
.Header =
{
.Offset = 0, // Set by PlaceNamedSymbolLiteral
.Type = FEXCore::CPU::RelocationTypes::RELOC_NAMED_SYMBOL_LITERAL,
},
.Symbol = Op,
.Offset = 0,
},
},
};
return Lit;
}
void Arm64JITCore::PlaceNamedSymbolLiteral(NamedSymbolLiteralPair Lit) {
switch (Lit.MoveABI.Header.Type) {
case RelocationTypes::RELOC_NAMED_SYMBOL_LITERAL:
case RelocationTypes::RELOC_GUEST_RIP_LITERAL: {
Lit.MoveABI.Header.Offset = GetCursorOffset();
break;
}
default: ERROR_AND_DIE_FMT("Unknown relocation type for {}", __FUNCTION__);
}
void Arm64JITCore::PlaceNamedSymbolLiteral(NamedSymbolLiteralPair& Lit) {
// Offset is the offset from the entrypoint of the block
auto CurrentCursor = GetCursorAddress<uint8_t*>();
Lit.MoveABI.NamedSymbolLiteral.Offset = CurrentCursor - CodeData.BlockBegin;
BindOrRestart(&Lit.Loc);
dc64(Lit.Lit);
Relocations.emplace_back(Lit.MoveABI);
}
auto Arm64JITCore::InsertGuestRIPLiteral(uint64_t GuestRIP) -> NamedSymbolLiteralPair {
return {
.Lit = GuestRIP,
.MoveABI =
{
.GuestRIP = {.Header =
{
.Offset = 0, // Set by PlaceNamedSymbolLiteral
.Type = FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_LITERAL,
},
// NOTE: Cache serialization will subtract the guest binary base address later to produce consistency results
.GuestRIP = GuestRIP},
},
};
}
void Arm64JITCore::InsertGuestRIPMove(ARMEmitter::Register Reg, uint64_t Constant) {
Relocation MoveABI {};
MoveABI.GuestRIP.Header = {.Offset = GetCursorOffset(), .Type = FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_MOVE};
// NOTE: Cache serialization will subtract the guest binary base address later to produce consistency results
MoveABI.GuestRIP.GuestRIP = Constant;
MoveABI.GuestRIP.RegisterIndex = Reg.Idx();
MoveABI.GuestRIPMove.Header.Type = FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_MOVE;
// Offset is the offset from the entrypoint of the block
auto CurrentCursor = GetCursorAddress<uint8_t*>();
MoveABI.GuestRIPMove.Offset = CurrentCursor - CodeData.BlockBegin;
MoveABI.GuestRIPMove.GuestRIP = Constant;
MoveABI.GuestRIPMove.RegisterIndex = Reg.Idx();
LoadConstant(ARMEmitter::Size::i64Bit, Reg, Constant, false);
Relocations.emplace_back(MoveABI);
}
fextl::vector<FEXCore::CPU::Relocation> Arm64JITCore::TakeRelocations(uint64_t GuestBaseAddress) {
// Rebase relocations to library base address
for (auto& Relocation : Relocations) {
switch (Relocation.Header.Type) {
case FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_MOVE:
case FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_LITERAL: {
Relocation.GuestRIP.GuestRIP -= GuestBaseAddress;
bool Arm64JITCore::ApplyRelocations(uint64_t GuestEntry, std::span<std::byte> Code, std::span<const FEXCore::CPU::Relocation> Relocations) {
const auto OrigBase = GetBufferBase();
const auto OrigSize = GetBufferSize();
const auto OrigOffset = GetCursorOffset();
SetBuffer(reinterpret_cast<std::uint8_t*>(Code.data()), Code.size_bytes());
for (auto& Reloc : Relocations) {
switch (Reloc.Header.Type) {
case FEXCore::CPU::RelocationTypes::RELOC_NAMED_SYMBOL_LITERAL: {
uint64_t Pointer = GetNamedSymbolLiteral(Reloc.NamedSymbolLiteral.Symbol);
// Relocation occurs at the cursorEntry + offset relative to that cursor
SetCursorOffset(Reloc.NamedSymbolLiteral.Offset);
// Generate a literal so we can place it
dc64(Pointer);
break;
}
case FEXCore::CPU::RelocationTypes::RELOC_NAMED_THUNK_MOVE: {
uint64_t Pointer = reinterpret_cast<uint64_t>(EmitterCTX->ThunkHandler->LookupThunk(Reloc.NamedThunkMove.Symbol));
if (Pointer == ~0ULL) {
return false;
}
// Relocation occurs at the cursorEntry + offset relative to that cursor.
SetCursorOffset(Reloc.NamedThunkMove.Offset);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Register(Reloc.NamedThunkMove.RegisterIndex), Pointer, true);
break;
}
case FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_MOVE: {
// XXX: Reenable once the JIT Object Cache is upstream
// XXX: Should spin the relocation list, create a list of guest RIP moves, and ask for them all once, reduces lock contention.
uint64_t Pointer = ~0ULL; // EmitterCTX->JITObjectCache->FindRelocatedRIP(Reloc->GuestRIPMove.GuestRIP);
if (Pointer == ~0ULL) {
SetBuffer(OrigBase, OrigSize);
SetCursorOffset(OrigOffset);
return false;
}
// Relocation occurs at the cursorEntry + offset relative to that cursor.
SetCursorOffset(Reloc.GuestRIPMove.Offset);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Register(Reloc.GuestRIPMove.RegisterIndex), Pointer, true);
break;
}
default:;
}
}
SetBuffer(OrigBase, OrigSize);
SetCursorOffset(OrigOffset);
return true;
}
fextl::vector<FEXCore::CPU::Relocation> Arm64JITCore::TakeRelocations() {
return std::move(Relocations);
}
@@ -138,6 +138,26 @@ DEF_OP(CAS) {
}
}
DEF_OP(AtomicXor) {
auto Op = IROp->C<IR::IROp_AtomicXor>();
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
auto MemSrc = GetReg(Op->Addr);
auto Src = GetReg(Op->Value);
if (CTX->HostFeatures.SupportsAtomics) {
steorl(SubEmitSize, Src, MemSrc);
} else {
ARMEmitter::BackwardLabel LoopTop;
(void)Bind(&LoopTop);
ldaxr(SubEmitSize, TMP2, MemSrc);
eor(EmitSize, TMP2, TMP2, Src);
stlxr(SubEmitSize, TMP2, TMP2, MemSrc);
(void)cbnz(EmitSize, TMP2, &LoopTop);
}
}
DEF_OP(AtomicSwap) {
auto Op = IROp->C<IR::IROp_AtomicSwap>();
const auto OpSize = IROp->Size;
@@ -328,7 +328,8 @@ DEF_OP(Thunk) {
mov(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, GetReg(Op->ArgPtr));
InsertNamedThunkRelocation(ARMEmitter::Reg::r2, Op->ThunkNameHash);
auto thunkFn = static_cast<Context::ContextImpl*>(ThreadState->CTX)->ThunkHandler->LookupThunk(Op->ThunkNameHash);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r2, (uintptr_t)thunkFn);
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<void, void*, void*>(ARMEmitter::Reg::r2);
} else {
+40 -59
View File
@@ -515,8 +515,7 @@ static void IndirectBlockDelinker(FEXCore::Context::ExitFunctionLinkData* Record
uintptr_t JumpThunkStartAddress = reinterpret_cast<uintptr_t>(Record) - 0x10;
uint32_t BranchInst = 0;
ARMEmitter::Emitter BranchEmit(reinterpret_cast<uint8_t*>(&BranchInst), 4);
// Restore branch +2 instructions to jump to the linker block
BranchEmit.b(0x2);
BranchEmit.b(0x8);
std::atomic_ref<uint32_t>(*reinterpret_cast<uint32_t*>(JumpThunkStartAddress)).store(BranchInst, std::memory_order::relaxed);
ARMEmitter::Emitter::ClearICache(reinterpret_cast<void*>(JumpThunkStartAddress), 4);
@@ -578,11 +577,16 @@ uint64_t Arm64JITCore::ExitFunctionLink(FEXCore::Core::CpuStateFrame* Frame, FEX
if (KnownCallMarkerInst == ExpectedKnownCallMarkerInst) {
BranchEmit.bl(BranchOffset);
Thread->LookupCache->AddBlockLink(
GuestRip, Record, [](FEXCore::Context::ExitFunctionLinkData* Record) { DirectBlockDelinker(Record, true); }, lk);
GuestRip, Record,
[](FEXCore::Context::ExitFunctionLinkData* Record) { DirectBlockDelinker(Record, true); }, lk);
} else {
BranchEmit.b(BranchOffset);
Thread->LookupCache->AddBlockLink(
GuestRip, Record, [](FEXCore::Context::ExitFunctionLinkData* Record) { DirectBlockDelinker(Record, false); }, lk);
GuestRip, Record,
[](FEXCore::Context::ExitFunctionLinkData* Record) {
DirectBlockDelinker(Record, false);
},
lk);
}
std::atomic_ref<uint32_t>(*reinterpret_cast<uint32_t*>(CallerAddress)).store(BranchInst, std::memory_order::relaxed);
@@ -783,7 +787,7 @@ void Arm64JITCore::EmitSuspendInterruptCheck() {
ldr(TMP2.W(), STATE_PTR(CpuStateFrame, SuspendDoorbell));
ARMEmitter::ForwardLabel l_NoSuspend;
(void)cbz(ARMEmitter::Size::i32Bit, TMP2, &l_NoSuspend);
cbz(ARMEmitter::Size::i32Bit, TMP2, &l_NoSuspend);
brk(SuspendMagic);
(void)Bind(&l_NoSuspend);
#endif
@@ -816,28 +820,17 @@ void Arm64JITCore::EmitEntryPoint(ARMEmitter::BackwardLabel& HeaderLabel, bool C
CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size, bool SingleInst, const FEXCore::IR::IRListView* IR,
FEXCore::Core::DebugData* DebugData, bool CheckTF) {
FEXCORE_PROFILE_SCOPED("Arm64::CompileCode");
const auto PrevNumAllocations = Relocations.size();
this->Entry = Entry;
this->DebugData = DebugData;
this->IR = IR;
RequiresFarARM64Jumps = false;
SSANodeMultiplier = 24;
// Prepare restart via long jump in case branch encoding fails.
// This uses UncheckedLongJump since we don't implement std::longjmp in WoA setups
switch (static_cast<RestartOptions::Control>(FEXCore::UncheckedLongJump::SetJump(ThreadState->RestartJump))) {
switch (static_cast<RestartOptions::Control>(FEXCore::LongJump::SetJump(RestartControl.RestartJump))) {
case RestartOptions::Control::Incoming:
// Nothing
break;
case RestartOptions::Control::EnableFarARM64Jumps: RequiresFarARM64Jumps = true; break;
case RestartOptions::Control::NeedsLargerJITSpace:
// Get rid of the claimed buffer immediately, we can't fit in it at all.
TempAllocator.UnclaimBuffer();
SSANodeMultiplier *= 2;
break;
default: LOGMAN_MSG_A_FMT("Unhandled Arm64 restart condition!");
default: ERROR_AND_DIE_FMT("Unhandled Arm64 restart condition!");
}
uint32_t SSACount = IR->GetSSACount();
@@ -849,19 +842,12 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
CodeData.EntryPoints.clear();
// Fairly excessive buffer range to make sure we don't overflow
// One page baseline, plus SSANodeMultipler bytes, plus another page for guard page.
const uint32_t DesiredBufferRange = AlignUp(FEXCore::Utils::FEX_PAGE_SIZE * 2 + SSACount * SSANodeMultiplier, FEXCore::Utils::FEX_PAGE_SIZE);
uint32_t BufferRange = 0x1000 + SSACount * 24;
// JIT output is first written to a temporary buffer and later relocated to the CodeBuffer.
// This minimizes lock contention of CodeBufferWriteMutex.
auto TempCodeBufferInfo = TempAllocator.ReownOrClaimBufferWithSize(DesiredBufferRange);
auto TempCodeBuffer = TempCodeBufferInfo.Ptr;
const uint32_t UsableBufferRange = TempCodeBufferInfo.Size - FEXCore::Utils::FEX_PAGE_SIZE;
SetBuffer(TempCodeBuffer, UsableBufferRange);
ThreadState->JITGuardPage = reinterpret_cast<uintptr_t>(TempCodeBuffer) + UsableBufferRange;
ThreadState->JITGuardOverflowArgument = FEXCore::ToUnderlying(RestartOptions::Control::NeedsLargerJITSpace);
auto TempCodeBuffer = TempAllocator.ReownOrClaimBuffer(BufferRange);
SetBuffer(TempCodeBuffer, BufferRange);
CodeData.BlockBegin = GetCursorAddress<uint8_t*>();
@@ -991,28 +977,22 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
// This is a ExitFunctionLinkData struct
BindOrRestart(&l_ExitLink);
dc64(0); // HostCode
PlaceNamedSymbolLiteral(InsertGuestRIPLiteral(PendingJumpThunk.GuestRIP)); // GuestRIP
dc64(PendingJumpThunk.CallerAddress - ThunkAddress); // CallerOffset
dc64(0); // HostCode
dc64(PendingJumpThunk.GuestRIP); // GuestRIP
dc64(PendingJumpThunk.CallerAddress - ThunkAddress); // CallerOffset
}
BindOrRestart(&l_ExitLink);
PlaceNamedSymbolLiteral(InsertNamedSymbolLiteral(RelocNamedSymbolLiteral::NamedSymbol::SYMBOL_LITERAL_EXITFUNCTION_LINKER));
dc64(ThreadState->CurrentFrame->Pointers.Common.ExitFunctionLinker);
// CodeSize not including the header or tail data.
const uint64_t CodeOnlySize = GetCursorAddress<uint8_t*>() - CodeBegin;
// Add the JitCodeTail (written later)
// Add the JitCodeTail
Align(alignof(JITCodeTail));
const auto JITBlockTailLocation = GetCursorAddress<uint8_t*>();
CodeHeader->OffsetToBlockTail = JITBlockTailLocation - CodeData.BlockBegin;
JITCodeTail JITBlockTail {
.RIP = Entry,
.GuestSize = Size,
.SpinLockFutex = 0,
.SingleInst = SingleInst,
};
auto JITBlockTailLocation = GetCursorAddress<uint8_t*>();
auto JITBlockTail = GetCursorAddress<JITCodeTail*>();
CursorIncrement(sizeof(JITCodeTail));
// Entries that live after the JITCodeTail.
// These entries correlate JIT code regions with guest RIP regions.
@@ -1030,13 +1010,23 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
// FEXCore::Utils::vl64 GuestRIPOffset;
// };
const auto JITRIPEntriesBegin = JITBlockTailLocation + sizeof(JITBlockTail);
auto JITRIPEntriesBegin = GetCursorAddress<uint8_t*>();
// Put the block's RIP entry in the tail.
// This will be used for RIP reconstruction in the future.
// TODO: This needs to be a data RIP relocation once code caching works.
// Current relocation code doesn't support this feature yet.
JITBlockTail->RIP = Entry;
JITBlockTail->GuestSize = Size;
JITBlockTail->SingleInst = SingleInst;
JITBlockTail->SpinLockFutex = 0;
auto JITRIPEntriesLocation = JITRIPEntriesBegin;
{
// Store the RIP entries.
JITBlockTail.NumberOfRIPEntries = DebugData->GuestOpcodes.size();
JITBlockTail.OffsetToRIPEntries = JITRIPEntriesBegin - JITBlockTailLocation;
JITBlockTail->NumberOfRIPEntries = DebugData->GuestOpcodes.size();
JITBlockTail->OffsetToRIPEntries = JITRIPEntriesBegin - JITBlockTailLocation;
uintptr_t CurrentRIPOffset = 0;
uint64_t CurrentPCOffset = 0;
@@ -1052,20 +1042,14 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
}
}
SetCursorOffset(JITRIPEntriesLocation - CodeData.BlockBegin);
CursorIncrement(JITRIPEntriesLocation - JITRIPEntriesBegin);
Align();
CodeHeader->OffsetToBlockTail = JITBlockTailLocation - CodeData.BlockBegin;
CodeData.Size = GetCursorAddress<uint8_t*>() - CodeData.BlockBegin;
// Finalize and write block tail data
JITBlockTail.Size = CodeData.Size;
{
auto PrevCur = GetCursorOffset();
memcpy(JITBlockTailLocation, &JITBlockTail, sizeof(JITBlockTail));
SetCursorOffset(JITBlockTailLocation - CodeData.BlockBegin + offsetof(JITCodeTail, RIP));
PlaceNamedSymbolLiteral(InsertGuestRIPLiteral(JITBlockTail.RIP));
SetCursorOffset(PrevCur);
}
JITBlockTail->Size = CodeData.Size;
// Migrate the compile output from temporary storage to the actual CodeBuffer.
// This can block progress in other compiling threads, so the duration of the lock should be as small as possible.
@@ -1074,6 +1058,7 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
// Query size of generated code
const auto TempSize = GetCursorOffset();
LOGMAN_THROW_A_FMT(TempSize <= BufferRange, "Exceeded bounds of temporary buffer ({:#x} vs {:#x})", TempSize, BufferRange);
// Bring CodeBuffer up to date
{
@@ -1105,10 +1090,6 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
}
CodeBegin += Delta;
for (std::size_t Idx = PrevNumAllocations; Idx != Relocations.size(); ++Idx) {
Relocations[Idx].Header.Offset += CodeBuffers.LatestOffset;
}
// Copy over CodeBuffer contents
memcpy(GetCursorAddress<uint8_t*>(), TempCodeBuffer, TempSize);
SetCursorOffset(CodeBuffers.LatestOffset + TempSize);
+22 -41
View File
@@ -68,10 +68,10 @@ private:
const bool HostSupportsAFP {};
struct RestartOptions {
FEXCore::LongJump::JumpBuf RestartJump;
enum class Control : uint64_t {
Incoming = 0,
EnableFarARM64Jumps = 1,
NeedsLargerJITSpace = 2,
};
};
@@ -79,8 +79,6 @@ private:
// In the rare case when those assumptions are broken, FEX needs to safely restart the JIT.
RestartOptions RestartControl {};
bool RequiresFarARM64Jumps {};
// Default to 6 instructions per SSA node.
uint32_t SSANodeMultiplier {24};
ARMEmitter::BiDirectionalLabel* PendingTargetLabel {};
ARMEmitter::BiDirectionalLabel* PendingCallReturnTargetLabel {};
@@ -362,7 +360,7 @@ private:
// We can support this but currently unnecessary.
ERROR_AND_DIE_FMT("Tried to branch larger than 128MB away!");
FEXCore::UncheckedLongJump::LongJump(ThreadState->RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
FEXCore::LongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
template<ARMEmitter::IsLabel T>
@@ -373,7 +371,7 @@ private:
// We can support this but currently unnecessary.
ERROR_AND_DIE_FMT("Tried to branch larger than 128MB away!");
FEXCore::UncheckedLongJump::LongJump(ThreadState->RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
FEXCore::LongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
template<ARMEmitter::IsLabel T>
@@ -394,7 +392,7 @@ private:
return;
}
FEXCore::UncheckedLongJump::LongJump(ThreadState->RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
FEXCore::LongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
template<ARMEmitter::IsLabel T>
@@ -415,7 +413,7 @@ private:
return;
}
FEXCore::UncheckedLongJump::LongJump(ThreadState->RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
FEXCore::LongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
template<ARMEmitter::IsLabel T>
@@ -436,7 +434,7 @@ private:
return;
}
FEXCore::UncheckedLongJump::LongJump(ThreadState->RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
FEXCore::LongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
template<ARMEmitter::IsLabel T>
@@ -457,7 +455,7 @@ private:
return;
}
FEXCore::UncheckedLongJump::LongJump(ThreadState->RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
FEXCore::LongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
template<ARMEmitter::IsLabel T>
@@ -478,37 +476,29 @@ private:
return;
}
FEXCore::UncheckedLongJump::LongJump(ThreadState->RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
FEXCore::LongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
template<ARMEmitter::IsLabel T>
void adr_OrRestart(ARMEmitter::Register rd, T* Label) {
if (RequiresFarARM64Jumps) {
if (LongAddressGen(rd, Label) == ARMEmitter::BranchEncodeSucceeded::Failure) {
ERROR_AND_DIE_FMT("Unable to encode long ADR.");
}
return;
}
if (adr(rd, Label) == ARMEmitter::BranchEncodeSucceeded::Success) {
return;
}
FEXCore::UncheckedLongJump::LongJump(ThreadState->RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
// We can support this but currently unnecessary.
ERROR_AND_DIE_FMT("Long ADR currently unsupported!");
FEXCore::LongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
template<ARMEmitter::IsLabel T>
void adrp_OrRestart(ARMEmitter::Register rd, T* Label) {
if (RequiresFarARM64Jumps) {
if (LongAddressGen(rd, Label) == ARMEmitter::BranchEncodeSucceeded::Failure) {
ERROR_AND_DIE_FMT("Unable to encode long ADRP.");
}
return;
}
if (adrp(rd, Label) == ARMEmitter::BranchEncodeSucceeded::Success) {
return;
}
FEXCore::UncheckedLongJump::LongJump(ThreadState->RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
// We can support this but currently unnecessary.
ERROR_AND_DIE_FMT("Long ADRP currently unsupported!");
FEXCore::LongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
template<ARMEmitter::IsLabel T>
@@ -523,7 +513,7 @@ private:
return;
}
FEXCore::UncheckedLongJump::LongJump(ThreadState->RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
FEXCore::LongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
// This is purely a debugging aid for developers to see if they are in JIT code space when inspecting raw memory
@@ -536,6 +526,8 @@ private:
* @name Relocations
* @{ */
uint64_t GetNamedSymbolLiteral(FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op);
/**
* @brief A literal pair relocation object for named symbol literals
*/
@@ -572,30 +564,19 @@ private:
*/
NamedSymbolLiteralPair InsertNamedSymbolLiteral(FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op);
/**
* @brief Inserts a relocation for a constant value relative to the guest entrypoint
*
* @param Reg - The GPR to move the guest RIP in to
* @param Constant - The guest RIP that will be relocated
*/
NamedSymbolLiteralPair InsertGuestRIPLiteral(uint64_t GuestRIP);
/**
* @brief Place the named symbol literal relocation in memory
*
* @param Lit - Which literal to place
*/
void PlaceNamedSymbolLiteral(NamedSymbolLiteralPair Lit);
void PlaceNamedSymbolLiteral(NamedSymbolLiteralPair& Lit);
fextl::vector<FEXCore::CPU::Relocation> Relocations;
/**
* Returns any relocations generated since the last call to TakeRelocations.
*
* GuestBaseAddress must match the base virtual address to which the
* input x86 binary is mapped.
*/
fextl::vector<FEXCore::CPU::Relocation> TakeRelocations(uint64_t GuestBaseAddress) override;
///< Relocation code loading
bool ApplyRelocations(uint64_t GuestEntry, std::span<std::byte> Code, std::span<const FEXCore::CPU::Relocation>);
fextl::vector<FEXCore::CPU::Relocation> TakeRelocations() override;
/** @} */
+24 -34
View File
@@ -1,89 +1,79 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <FEXCore/IR/IR.h>
#include <FEXCore/Utils/CompilerDefs.h>
namespace FEXCore::Context {
class ContextImpl;
}
namespace FEXCore::CPU {
enum class RelocationTypes : uint32_t {
enum class RelocationTypes : uint8_t {
// 8 byte literal in memory for symbol
// Aligned to struct RelocNamedSymbolLiteral
RELOC_NAMED_SYMBOL_LITERAL,
// Fixed size named thunk move
// 4 instruction constant generation
// 4 instruction constant generation on AArch64
// 64-bit mov on x86-64
// Aligned to struct RelocNamedThunkMove
RELOC_NAMED_THUNK_MOVE,
// 8 byte literal (relative to binary base address)
RELOC_GUEST_RIP_LITERAL,
// Fixed size guest RIP move
// 4 instruction constant generation
// Aligned to struct RelocGuestRIP
// 4 instruction constant generation on AArch64
// 64-bit mov on x86-64
// Aligned to struct RelocGuestRIPMove
RELOC_GUEST_RIP_MOVE,
};
struct FEX_PACKED RelocationHeader final {
// Offset to the relocated host code data
uint64_t Offset {};
struct RelocationTypeHeader final {
RelocationTypes Type;
};
struct RelocNamedSymbolLiteral final {
enum class NamedSymbol : uint32_t {
enum class NamedSymbol : uint8_t {
///< Thread specific relocations
// JIT Literal pointers
SYMBOL_LITERAL_EXITFUNCTION_LINKER,
};
RelocationHeader Header {};
RelocationTypeHeader Header {};
NamedSymbol Symbol;
uint32_t Pad[8];
// Offset in to the code section to begin the relocation
uint64_t Offset {};
};
struct RelocNamedThunkMove final {
RelocationHeader Header {};
RelocationTypeHeader Header {};
// GPR index the constant is being moved to
uint32_t RegisterIndex;
uint8_t RegisterIndex;
// The thunk SHA256 hash
IR::SHA256Sum Symbol;
// Offset in to the code section to begin the relocation
uint64_t Offset {};
};
struct RelocGuestRIP final {
RelocationHeader Header {};
struct RelocGuestRIPMove final {
RelocationTypeHeader Header {};
// GPR index the constant is being moved to (for non-literal relocations)
// GPR index the constant is being moved to
uint8_t RegisterIndex;
char Pad[3];
// Offset in to the code section to begin the relocation
uint64_t Offset {};
// The base RIP (to be moved by the register for non-literal relocations).
// In a serialized code cache, this is relative to the binary base address.
// The unrelocated RIP that is being moved
uint64_t GuestRIP;
uint32_t pad2[6] {};
};
union Relocation {
RelocationHeader Header {};
RelocationTypeHeader Header {};
RelocNamedSymbolLiteral NamedSymbolLiteral;
// This makes our union of relocations at least 48 bytes
// It might be more efficient to not use a union
RelocNamedThunkMove NamedThunkMove;
RelocGuestRIP GuestRIP;
RelocGuestRIPMove GuestRIPMove;
};
uint64_t GetNamedSymbolLiteral(FEXCore::Context::ContextImpl&, RelocNamedSymbolLiteral::NamedSymbol);
} // namespace FEXCore::CPU
@@ -39,8 +39,6 @@ LookupCache::LookupCache(FEXCore::Context::ContextImpl* CTX)
// We need one pointer per page of virtual memory
// At 64GB of virtual memory this will allocate 128MB of virtual memory space
PagePointer = reinterpret_cast<uintptr_t>(FEXCore::Allocator::VirtualAlloc(TotalCacheSize, false, false));
LOGMAN_THROW_A_FMT(PagePointer != -1ULL, "Failed to allocate PagePointer");
FEXCore::Allocator::VirtualName("FEXMem_Lookup", reinterpret_cast<void*>(PagePointer),
ctx->Config.VirtualMemSize / FEXCore::Utils::FEX_PAGE_SIZE * 8 + CODE_SIZE);
CTX->SyscallHandler->MarkOvercommitRange(PagePointer, TotalCacheSize);
@@ -51,11 +49,14 @@ LookupCache::LookupCache(FEXCore::Context::ContextImpl* CTX)
// We currently limit to 128MB of real memory for caching for the total cache size.
// Can end up being inefficient if we compile a small number of blocks per page
PageMemory = PagePointer + ctx->Config.VirtualMemSize / FEXCore::Utils::FEX_PAGE_SIZE * 8;
LOGMAN_THROW_A_FMT(PageMemory != -1ULL, "Failed to allocate page memory");
// L1 Cache
L1Pointer = PageMemory + CODE_SIZE;
FEXCore::Allocator::VirtualName("FEXMem_Lookup_L1", reinterpret_cast<void*>(L1Pointer), MAX_L1_SIZE);
LOGMAN_THROW_A_FMT(L1Pointer != -1ULL, "Failed to allocate L1Pointer");
VirtualMemSize = ctx->Config.VirtualMemSize;
if (DynamicL1Cache()) {
@@ -75,7 +76,7 @@ LookupCache::~LookupCache() {
// These will get freed when their memory allocators are deallocated.
}
void LookupCache::ClearL2Cache(const FEXCore::LookupCacheBaseLockToken& lk) {
void LookupCache::ClearL2Cache(const FEXCore::LookupCacheWriteLockToken& lk) {
// Clear out the page memory
// PagePointer and PageMemory are sequential with each other. Clear both at once.
FEXCore::Allocator::VirtualDontNeed(reinterpret_cast<void*>(PagePointer),
+12 -34
View File
@@ -3,13 +3,11 @@
#include "Interface/Context/Context.h"
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/SHMStats.h>
#include "Utils/WritePriorityMutex.h"
#include <FEXCore/fextl/map.h>
#include <FEXCore/fextl/memory_resource.h>
#include <FEXCore/fextl/robin_map.h>
#include <FEXCore/fextl/robin_set.h>
#include <FEXCore/fextl/vector.h>
#include <FEXCore/fextl/unordered_set.h>
#include <FEXCore/fextl/memory_resource.h>
#include <cstdint>
@@ -18,41 +16,22 @@
#include <mutex>
namespace FEXCore {
struct LookupCacheBaseLockToken {
protected:
// Protected constructor - only derived classes can construct
LookupCacheBaseLockToken() = default;
};
struct LookupCacheWriteLockToken : public LookupCacheBaseLockToken {
struct LookupCacheWriteLockToken {
private:
// Only constructible by GuestToHostMap
friend struct GuestToHostMap;
LookupCacheWriteLockToken(FEXCore::Utils::WritePriorityMutex::Mutex& Mutex)
LookupCacheWriteLockToken(std::mutex& Mutex)
: Lock {Mutex} {}
std::lock_guard<FEXCore::Utils::WritePriorityMutex::Mutex> Lock;
};
struct LookupCacheReadLockToken : public LookupCacheBaseLockToken {
private:
// Only constructible by GuestToHostMap
friend struct GuestToHostMap;
LookupCacheReadLockToken(FEXCore::Utils::WritePriorityMutex::Mutex& Mutex)
: Lock {Mutex} {}
std::shared_lock<FEXCore::Utils::WritePriorityMutex::Mutex> Lock;
std::lock_guard<std::mutex> Lock;
};
struct GuestToHostMap {
FEXCore::Utils::WritePriorityMutex::Mutex Lock {};
std::mutex WriteLock;
[[nodiscard]]
LookupCacheWriteLockToken AcquireWriteLock() {
return LookupCacheWriteLockToken {Lock};
}
[[nodiscard]]
LookupCacheReadLockToken AcquireReadLock() {
return LookupCacheReadLockToken {Lock};
return LookupCacheWriteLockToken {WriteLock};
}
struct BlockLinkTag {
@@ -102,7 +81,7 @@ struct GuestToHostMap {
return BlockList.insert_or_assign(Address, BlockEntry {(uintptr_t)HostCode, CodePages}).first->second;
}
const BlockEntry* FindBlock(uint64_t Address, const LookupCacheReadLockToken&) {
const BlockEntry* FindBlock(uint64_t Address, const LookupCacheWriteLockToken&) {
auto HostCode = BlockList.find(Address);
if (HostCode == BlockList.end()) {
return nullptr;
@@ -185,7 +164,7 @@ public:
{
std::optional<FEXCore::SHMStats::AccumulationBlock<uint64_t>> LockTime(
Thread->ThreadStats ? &Thread->ThreadStats->AccumulatedCacheReadLockTime : nullptr);
auto lk = Shared->AcquireReadLock();
auto lk = Shared->AcquireWriteLock();
LockTime.reset();
if (!DisableL2Cache()) {
@@ -341,9 +320,8 @@ public:
InvalidateCache(Entry, lk);
}
}
bool ret = upper != lower;
CachedCodePages.erase(lower, upper);
return ret;
return upper != lower;
}
void AddBlockLink(uint64_t GuestDestination, FEXCore::Context::ExitFunctionLinkData* HostLink,
@@ -352,7 +330,7 @@ public:
}
void ClearCache(const LookupCacheWriteLockToken&);
void ClearL2Cache(const LookupCacheBaseLockToken&);
void ClearL2Cache(const LookupCacheWriteLockToken&);
void ClearThreadLocalCaches(const LookupCacheWriteLockToken&);
uintptr_t GetL1Pointer() const {
@@ -380,7 +358,7 @@ public:
}
private:
void CacheBlockMapping(uint64_t Address, const GuestToHostMap::BlockEntry& Entry, bool L1Only, const LookupCacheBaseLockToken& lk) {
void CacheBlockMapping(uint64_t Address, const GuestToHostMap::BlockEntry& Entry, bool L1Only, const LookupCacheWriteLockToken& lk) {
for (const auto& CodePage : Entry.CodePages) {
CachedCodePages[CodePage >> 12].insert(Address);
}
@@ -438,7 +416,7 @@ private:
}
// Maps from a page index to all blocks in the page that have at some point been fetched into L1/L2
fextl::map<uint64_t, fextl::robin_set<uint64_t>> CachedCodePages;
fextl::map<uint64_t, fextl::unordered_set<uint64_t>> CachedCodePages;
uintptr_t PagePointer;
uintptr_t PageMemory;
@@ -514,17 +514,18 @@ void OpDispatchBuilder::CALLOp(OpcodeArgs) {
BlockSetRIP = true;
// Call instruction only uses up to 32-bit signed displacement
const int64_t TargetOffset = Op->Src[0].Literal();
int64_t TargetOffset = Op->Src[0].Literal();
const auto ConstantPC = GetRelocatedPC(Op);
auto ConstantPC = GetRelocatedPC(Op);
// Push the return address.
Push(GPRSize, ConstantPC);
if (TargetOffset != 0) {
// Store the RIP
const uint64_t NextRIP = Op->PC + Op->InstSize;
const uint64_t NextRIP = Op->PC + Op->InstSize;
uint64_t TargetRIP = NextRIP + TargetOffset;
if (NextRIP != TargetRIP) {
// Store the RIP
ExitRelocatedPC(Op, TargetOffset, BranchHint::Call, ConstantPC, [&]() {
auto CallReturnJumpTarget = JumpTargets.find(NextRIP);
if (CallReturnJumpTarget != JumpTargets.end() && CallReturnJumpTarget->second.IsEntryPoint) {
@@ -2749,8 +2750,7 @@ void OpDispatchBuilder::NOTOp(OpcodeArgs) {
if (DestIsLockedMem(Op)) {
HandledLock = true;
Ref DestMem = MakeSegmentAddress(Op, Op->Dest);
// Result unused
_AtomicFetchXor(Size, MaskConst, DestMem);
_AtomicXor(Size, MaskConst, DestMem);
} else if (!Op->Dest.IsGPR()) {
// GPR version plays fast and loose with sizes, be safe for memory tho.
Ref Src = LoadSourceGPR(Op, Op->Dest, Op->Flags);
@@ -3199,7 +3199,7 @@ void OpDispatchBuilder::DECOp(OpcodeArgs) {
void OpDispatchBuilder::STOSOp(OpcodeArgs) {
if (Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_ADDRESS_SIZE) {
LogMan::Msg::EFmt("STOSOp: Can't handle address size override (OP: 0x{:04X}, Flags: 0x{:08X})", Op->OP, Op->Flags);
LogMan::Msg::EFmt("Can't handle adddress size");
DecodeFailure = true;
return;
}
@@ -3244,7 +3244,7 @@ void OpDispatchBuilder::STOSOp(OpcodeArgs) {
void OpDispatchBuilder::MOVSOp(OpcodeArgs) {
if (Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_ADDRESS_SIZE) {
LogMan::Msg::EFmt("MOVSOp: Can't handle address size override (OP: 0x{:04X}, Flags: 0x{:08X})", Op->OP, Op->Flags);
LogMan::Msg::EFmt("Can't handle adddress size");
DecodeFailure = true;
return;
}
@@ -3297,57 +3297,45 @@ void OpDispatchBuilder::MOVSOp(OpcodeArgs) {
_StoreMem(RegClass::GPR, Size, Src, RDI, Invalid(), OpSize::i8Bit, MemOffsetType::SXTX, 1);
}
RSI = OffsetByDir(RSI, IR::OpSizeToSize(Size));
RDI = OffsetByDir(RDI, IR::OpSizeToSize(Size));
auto PtrDir = LoadDir(IR::OpSizeToSize(Size));
RSI = Add(OpSize::i64Bit, RSI, PtrDir);
RDI = Add(OpSize::i64Bit, RDI, PtrDir);
StoreGPRRegister(X86State::REG_RSI, RSI);
StoreGPRRegister(X86State::REG_RDI, RDI);
}
}
IR::OpSize OpDispatchBuilder::GetStringOpSize(X86Tables::DecodedOp Op) const {
LOGMAN_THROW_A_FMT(Is64BitMode || !(Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_ADDRESS_SIZE), "Invalid modifier on 32bit address");
return !Is64BitMode || (Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_ADDRESS_SIZE) ? OpSize::i32Bit : OpSize::i64Bit;
}
void OpDispatchBuilder::CMPSOp(OpcodeArgs) {
if (!Is64BitMode && (Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_ADDRESS_SIZE)) {
LogMan::Msg::EFmt("CMPSOp: Address size override (0x67) not supported in 32-bit mode (OP: 0x{:04X}).", Op->OP);
if (Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_ADDRESS_SIZE) {
LogMan::Msg::EFmt("Can't handle adddress size");
DecodeFailure = true;
return;
}
const auto Size = OpSizeFromSrc(Op);
OpSize AddrSize = GetStringOpSize(Op);
bool Repeat = Op->Flags & (FEXCore::X86Tables::DecodeFlags::FLAG_REPNE_PREFIX | FEXCore::X86Tables::DecodeFlags::FLAG_REP_PREFIX);
if (!Repeat) {
Ref Src_RSI = LoadGPRRegister(X86State::REG_RSI, AddrSize);
Ref Src_RDI = LoadGPRRegister(X86State::REG_RDI, AddrSize);
Ref Dest_RSI = AppendSegmentOffset(Src_RSI, Op->Flags, X86Tables::DecodeFlags::FLAG_DS_PREFIX);
Ref Dest_RDI = AppendSegmentOffset(Src_RDI, 0, X86Tables::DecodeFlags::FLAG_ES_PREFIX, true);
// Default DS prefix
Ref Dest_RSI = MakeSegmentAddress(X86State::REG_RSI, Op->Flags, X86Tables::DecodeFlags::FLAG_DS_PREFIX);
// Only ES prefix
Ref Dest_RDI = MakeSegmentAddress(X86State::REG_RDI, 0, X86Tables::DecodeFlags::FLAG_ES_PREFIX, true);
auto Src1 = _LoadMemGPRAutoTSO(Size, Dest_RDI, Size);
auto Src2 = _LoadMemGPRAutoTSO(Size, Dest_RSI, Size);
CalculateFlags_SUB(OpSizeFromSrc(Op), Src2, Src1);
Dest_RDI = OffsetByDir(Src_RDI, IR::OpSizeToSize(Size));
if (Is64BitMode && AddrSize == OpSize::i32Bit) {
Dest_RDI = _Bfe(OpSize::i64Bit, 32, 0, Dest_RDI);
StoreGPRRegister(X86State::REG_RDI, Dest_RDI);
} else {
StoreGPRRegister(X86State::REG_RDI, Dest_RDI, AddrSize);
}
auto PtrDir = LoadDir(IR::OpSizeToSize(Size));
Dest_RSI = OffsetByDir(Src_RSI, IR::OpSizeToSize(Size));
if (Is64BitMode && AddrSize == OpSize::i32Bit) {
Dest_RSI = _Bfe(OpSize::i64Bit, 32, 0, Dest_RSI);
StoreGPRRegister(X86State::REG_RSI, Dest_RSI);
} else {
StoreGPRRegister(X86State::REG_RSI, Dest_RSI, AddrSize);
}
// Offset the pointer
Dest_RDI = Add(OpSize::i64Bit, Dest_RDI, PtrDir);
StoreGPRRegister(X86State::REG_RDI, Dest_RDI);
// Offset second pointer
Dest_RSI = Add(OpSize::i64Bit, Dest_RSI, PtrDir);
StoreGPRRegister(X86State::REG_RSI, Dest_RSI);
} else {
// Calculate flags early.
CalculateDeferredFlags();
@@ -3363,7 +3351,7 @@ void OpDispatchBuilder::CMPSOp(OpcodeArgs) {
SetCurrentCodeBlock(BeforeLoop);
StartNewBlock();
ForeachDirection([this, Op, Size, AddrSize, REPE](int32_t PtrDir) {
ForeachDirection([this, Op, Size, REPE](int32_t PtrDir) {
IRPair<IROp_CondJump> InnerJump;
auto JumpIntoLoop = Jump();
@@ -3375,11 +3363,10 @@ void OpDispatchBuilder::CMPSOp(OpcodeArgs) {
// Working loop
{
Ref Src_RSI = LoadGPRRegister(X86State::REG_RSI, AddrSize);
Ref Src_RDI = LoadGPRRegister(X86State::REG_RDI, AddrSize);
Ref Dest_RSI = AppendSegmentOffset(Src_RSI, Op->Flags, X86Tables::DecodeFlags::FLAG_DS_PREFIX);
Ref Dest_RDI = AppendSegmentOffset(Src_RDI, 0, X86Tables::DecodeFlags::FLAG_ES_PREFIX, true);
// Default DS prefix
Ref Dest_RSI = MakeSegmentAddress(X86State::REG_RSI, Op->Flags, X86Tables::DecodeFlags::FLAG_DS_PREFIX);
// Only ES prefix
Ref Dest_RDI = MakeSegmentAddress(X86State::REG_RDI, 0, X86Tables::DecodeFlags::FLAG_ES_PREFIX, true);
auto Src1 = _LoadMemGPRAutoTSO(Size, Dest_RDI, Size);
auto Src2 = _LoadMemGPR(Size, Dest_RSI, Size);
@@ -3396,21 +3383,13 @@ void OpDispatchBuilder::CMPSOp(OpcodeArgs) {
// Store the counter since we don't have phis
StoreGPRRegister(X86State::REG_RCX, TailCounter);
Dest_RDI = Add(AddrSize, Src_RDI, PtrDir * static_cast<int32_t>(IR::OpSizeToSize(Size)));
if (Is64BitMode && AddrSize == OpSize::i32Bit) {
Dest_RDI = _Bfe(OpSize::i64Bit, 32, 0, Dest_RDI);
StoreGPRRegister(X86State::REG_RDI, Dest_RDI);
} else {
StoreGPRRegister(X86State::REG_RDI, Dest_RDI, AddrSize);
}
// Offset the pointer
Dest_RDI = Add(OpSize::i64Bit, Dest_RDI, PtrDir * static_cast<int32_t>(IR::OpSizeToSize(Size)));
StoreGPRRegister(X86State::REG_RDI, Dest_RDI);
Dest_RSI = Add(AddrSize, Src_RSI, PtrDir * static_cast<int32_t>(IR::OpSizeToSize(Size)));
if (Is64BitMode && AddrSize == OpSize::i32Bit) {
Dest_RSI = _Bfe(OpSize::i64Bit, 32, 0, Dest_RSI);
StoreGPRRegister(X86State::REG_RSI, Dest_RSI);
} else {
StoreGPRRegister(X86State::REG_RSI, Dest_RSI, AddrSize);
}
// Offset second pointer
Dest_RSI = Add(OpSize::i64Bit, Dest_RSI, PtrDir * static_cast<int32_t>(IR::OpSizeToSize(Size)));
StoreGPRRegister(X86State::REG_RSI, Dest_RSI);
// If TailCounter != 0, compare sources.
// If TailCounter == 0, set ZF iff that would break.
@@ -3449,7 +3428,7 @@ void OpDispatchBuilder::CMPSOp(OpcodeArgs) {
void OpDispatchBuilder::LODSOp(OpcodeArgs) {
if (Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_ADDRESS_SIZE) {
LogMan::Msg::EFmt("LODSOp: Can't handle address size override (OP: 0x{:04X}, Flags: 0x{:08X})", Op->OP, Op->Flags);
LogMan::Msg::EFmt("Can't handle adddress size");
DecodeFailure = true;
return;
}
@@ -3531,37 +3510,31 @@ void OpDispatchBuilder::LODSOp(OpcodeArgs) {
}
void OpDispatchBuilder::SCASOp(OpcodeArgs) {
if (!Is64BitMode && (Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_ADDRESS_SIZE)) {
LogMan::Msg::EFmt("SCASOp: Address size override (0x67) not supported in 32-bit mode (OP: 0x{:04X}).", Op->OP);
if (Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_ADDRESS_SIZE) {
LogMan::Msg::EFmt("Can't handle adddress size");
DecodeFailure = true;
return;
}
const auto Size = OpSizeFromSrc(Op);
OpSize AddrSize = GetStringOpSize(Op);
const bool Repeat = (Op->Flags & (FEXCore::X86Tables::DecodeFlags::FLAG_REPNE_PREFIX | FEXCore::X86Tables::DecodeFlags::FLAG_REP_PREFIX)) != 0;
if (!Repeat) {
Ref Src_RDI = LoadGPRRegister(X86State::REG_RDI, AddrSize);
Ref Dest_RDI = AppendSegmentOffset(Src_RDI, 0, X86Tables::DecodeFlags::FLAG_ES_PREFIX, true);
Ref Dest_RDI = MakeSegmentAddress(X86State::REG_RDI, 0, X86Tables::DecodeFlags::FLAG_ES_PREFIX, true);
auto Src1 = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.AllowUpperGarbage = true});
auto Src2 = _LoadMemGPRAutoTSO(Size, Dest_RDI, Size);
CalculateFlags_SUB(OpSizeFromSrc(Op), Src1, Src2);
Ref TailDest_RDI = OffsetByDir(Src_RDI, IR::OpSizeToSize(Size));
if (Is64BitMode && AddrSize == OpSize::i32Bit) {
TailDest_RDI = _Bfe(OpSize::i64Bit, 32, 0, TailDest_RDI);
StoreGPRRegister(X86State::REG_RDI, TailDest_RDI);
} else {
StoreGPRRegister(X86State::REG_RDI, TailDest_RDI, AddrSize);
}
// Offset the pointer
Ref TailDest_RDI = LoadGPRRegister(X86State::REG_RDI);
StoreGPRRegister(X86State::REG_RDI, OffsetByDir(TailDest_RDI, IR::OpSizeToSize(Size)));
} else {
// Calculate flags early. because end of block
CalculateDeferredFlags();
ForeachDirection([this, Op, Size, AddrSize](int32_t Dir) {
ForeachDirection([this, Op, Size](int32_t Dir) {
bool REPE = Op->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_REP_PREFIX;
auto JumpStart = Jump();
@@ -3585,8 +3558,7 @@ void OpDispatchBuilder::SCASOp(OpcodeArgs) {
// Working loop
{
Ref Src_RDI = LoadGPRRegister(X86State::REG_RDI, AddrSize);
Ref Dest_RDI = AppendSegmentOffset(Src_RDI, 0, X86Tables::DecodeFlags::FLAG_ES_PREFIX, true);
Ref Dest_RDI = MakeSegmentAddress(X86State::REG_RDI, 0, X86Tables::DecodeFlags::FLAG_ES_PREFIX, true);
auto Src1 = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.AllowUpperGarbage = true});
auto Src2 = _LoadMemGPRAutoTSO(Size, Dest_RDI, Size);
@@ -3597,7 +3569,7 @@ void OpDispatchBuilder::SCASOp(OpcodeArgs) {
CalculateDeferredFlags();
Ref TailCounter = LoadGPRRegister(X86State::REG_RCX);
Ref Src_RDI_Tail = LoadGPRRegister(X86State::REG_RDI, AddrSize);
Ref TailDest_RDI = LoadGPRRegister(X86State::REG_RDI);
// Decrement counter
TailCounter = Sub(OpSize::i64Bit, TailCounter, 1);
@@ -3605,13 +3577,9 @@ void OpDispatchBuilder::SCASOp(OpcodeArgs) {
// Store the counter since we don't have phis
StoreGPRRegister(X86State::REG_RCX, TailCounter);
Ref TailDest_RDI = Add(AddrSize, Src_RDI_Tail, Dir * static_cast<int32_t>(IR::OpSizeToSize(Size)));
if (Is64BitMode && AddrSize == OpSize::i32Bit) {
TailDest_RDI = _Bfe(OpSize::i64Bit, 32, 0, TailDest_RDI);
StoreGPRRegister(X86State::REG_RDI, TailDest_RDI);
} else {
StoreGPRRegister(X86State::REG_RDI, TailDest_RDI, AddrSize);
}
// Offset the pointer
TailDest_RDI = Add(OpSize::i64Bit, TailDest_RDI, Dir * static_cast<int32_t>(IR::OpSizeToSize(Size)));
StoreGPRRegister(X86State::REG_RDI, TailDest_RDI);
CalculateDeferredFlags();
InternalCondJump = CondJumpNZCV(REPE ? CondClass::EQ : CondClass::NEQ);
@@ -1655,9 +1655,6 @@ private:
return IR::SizeToOpSize(GetSrcSize(Op));
}
[[nodiscard]]
IR::OpSize GetStringOpSize(X86Tables::DecodedOp Op) const;
// Set flag tracking to prepare for an operation that directly writes NZCV.
void HandleNZCVWrite() {
CachedNZCV = nullptr;
@@ -552,7 +552,7 @@ void OpDispatchBuilder::AVXInsertScalarRound(OpcodeArgs) {
const uint64_t Mode = Op->Src[2].Literal();
const auto DstSize = GetGuestVectorLength();
Ref Result = InsertScalarRoundImpl(Op, DstSize, ElementSize, Op->Src[0], Op->Src[1], Mode, true);
Ref Result = InsertScalarRoundImpl(Op, DstSize, ElementSize, Op->Dest, Op->Src[0], Mode, true);
StoreResultFPR_WithOpSize(Op, Op->Dest, Result, DstSize);
}
@@ -17,6 +17,7 @@ $end_info$
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/FPState.h>
#include <cmath>
#include <stddef.h>
#include <stdint.h>
@@ -60,13 +61,15 @@ void OpDispatchBuilder::SetX87Top(Ref Value) {
// Float LoaD operation with memory operand
void OpDispatchBuilder::FLD(OpcodeArgs, IR::OpSize Width) {
const auto ReadWidth = (Width == OpSize::f80Bit) ? OpSize::i128Bit : Width;
Ref Data = LoadSourceFPR_WithOpSize(Op, Op->Src[0], Width, Op->Flags);
Ref ConvertedData = Data;
// Convert to 80bit float
if (Width == OpSize::i32Bit || Width == OpSize::i64Bit) {
ConvertedData = _F80CVTTo(Data, Width);
ConvertedData = _F80CVTTo(Data, ReadWidth);
}
_PushStack(ConvertedData, Data, Width);
_PushStack(ConvertedData, Data, ReadWidth, true);
}
// Float LoaD operation with memory operand
@@ -78,7 +81,7 @@ void OpDispatchBuilder::FBLD(OpcodeArgs) {
// Read from memory
Ref Data = LoadSourceFPR_WithOpSize(Op, Op->Src[0], OpSize::f80Bit, Op->Flags);
Ref ConvertedData = _F80BCDLoad(Data);
_PushStack(ConvertedData, Invalid(), OpSize::iInvalid);
_PushStack(ConvertedData, Data, OpSize::i128Bit, true);
}
void OpDispatchBuilder::FBSTP(OpcodeArgs) {
@@ -90,7 +93,7 @@ void OpDispatchBuilder::FBSTP(OpcodeArgs) {
void OpDispatchBuilder::FLD_Const(OpcodeArgs, NamedVectorConstant K) {
// Update TOP
Ref Data = LoadAndCacheNamedVectorConstant(OpSize::i128Bit, K);
_PushStack(Data, Data, OpSize::f80Bit);
_PushStack(Data, Data, OpSize::i128Bit, true);
}
void OpDispatchBuilder::FILD(OpcodeArgs) {
@@ -121,16 +124,15 @@ void OpDispatchBuilder::FILD(OpcodeArgs) {
auto upper = _Or(OpSize::i64Bit, sign, zeroed_exponent);
Ref ConvertedData = _VLoadTwoGPRs(shifted, upper);
_PushStack(ConvertedData, Invalid(), OpSize::iInvalid);
_PushStack(ConvertedData, Data, ReadWidth, false);
}
void OpDispatchBuilder::FST(OpcodeArgs, IR::OpSize Width) {
LOGMAN_THROW_A_FMT(Width == OpSize::i32Bit || Width == OpSize::i64Bit || Width == OpSize::f80Bit, "Invalid store width for FST");
const auto SourceSize = ReducedPrecisionMode ? OpSize::i64Bit : OpSize::f80Bit;
const auto SourceSize = ReducedPrecisionMode ? OpSize::i64Bit : OpSize::i128Bit;
AddressMode A = DecodeAddress(Op, Op->Dest, MemoryAccessType::DEFAULT, false);
A = SelectAddressMode(this, A, GetGPROpSize(), CTX->HostFeatures.SupportsTSOImm9, false, false, Width);
_StoreStackMem(SourceSize, Width, A.Base, A.Index, OpSize::iInvalid, A.IndexType, A.IndexScale);
_StoreStackMem(SourceSize, Width, A.Base, A.Index, OpSize::iInvalid, A.IndexType, A.IndexScale, /*Float=*/true);
if (Op->TableInfo->Flags & X86Tables::InstFlags::FLAGS_POP) {
_PopStackDestroy();
@@ -876,8 +878,8 @@ void OpDispatchBuilder::X87FXTRACT(OpcodeArgs) {
_PopStackDestroy();
auto Exp = _F80XTRACT_EXP(Top);
auto Sig = _F80XTRACT_SIG(Top);
_PushStack(Exp, Invalid(), OpSize::iInvalid);
_PushStack(Sig, Invalid(), OpSize::iInvalid);
_PushStack(Exp, Exp, OpSize::f80Bit, true);
_PushStack(Sig, Sig, OpSize::f80Bit, true);
}
} // namespace FEXCore::IR
@@ -59,6 +59,7 @@ void OpDispatchBuilder::X87FLDCWF64(OpcodeArgs) {
// F64 ops
// Float load op with memory operand
void OpDispatchBuilder::FLDF64(OpcodeArgs, IR::OpSize Width) {
const auto ReadWidth = (Width == OpSize::f80Bit) ? OpSize::i128Bit : Width;
Ref Data = LoadSourceFPR_WithOpSize(Op, Op->Src[0], Width, Op->Flags);
// Convert to 64bit float
Ref ConvertedData = Data;
@@ -67,7 +68,7 @@ void OpDispatchBuilder::FLDF64(OpcodeArgs, IR::OpSize Width) {
} else if (Width == OpSize::f80Bit) {
ConvertedData = _F80CVT(OpSize::i64Bit, Data);
}
_PushStack(ConvertedData, Data, Width);
_PushStack(ConvertedData, Data, ReadWidth, true);
}
void OpDispatchBuilder::FBLDF64(OpcodeArgs) {
@@ -75,7 +76,7 @@ void OpDispatchBuilder::FBLDF64(OpcodeArgs) {
Ref Data = LoadSourceFPR_WithOpSize(Op, Op->Src[0], OpSize::f80Bit, Op->Flags);
Ref ConvertedData = _F80BCDLoad(Data);
ConvertedData = _F80CVT(OpSize::i64Bit, ConvertedData);
_PushStack(ConvertedData, Invalid(), OpSize::iInvalid);
_PushStack(ConvertedData, Data, OpSize::i64Bit, true);
}
void OpDispatchBuilder::FBSTPF64(OpcodeArgs) {
@@ -87,7 +88,7 @@ void OpDispatchBuilder::FBSTPF64(OpcodeArgs) {
void OpDispatchBuilder::FLDF64_Const(OpcodeArgs, uint64_t Num) {
auto Data = _VCastFromGPR(OpSize::i64Bit, OpSize::i64Bit, Constant(Num));
_PushStack(Data, Data, OpSize::i64Bit);
_PushStack(Data, Data, OpSize::i64Bit, true);
}
void OpDispatchBuilder::FILDF64(OpcodeArgs) {
@@ -99,7 +100,7 @@ void OpDispatchBuilder::FILDF64(OpcodeArgs) {
Data = _Sbfe(OpSize::i64Bit, IR::OpSizeAsBits(ReadWidth), 0, Data);
}
auto ConvertedData = _Float_FromGPR_S(OpSize::i64Bit, ReadWidth == OpSize::i32Bit ? OpSize::i32Bit : OpSize::i64Bit, Data);
_PushStack(ConvertedData, Invalid(), OpSize::iInvalid);
_PushStack(ConvertedData, Data, ReadWidth, false);
}
void OpDispatchBuilder::FISTF64(OpcodeArgs, bool Truncate) {
@@ -396,7 +397,7 @@ void OpDispatchBuilder::X87FXTRACTF64(OpcodeArgs) {
Ref Exp = _NZCVSelectV(OpSize::i64Bit, CondClass::EQ, ExpZV, ExpNZV);
_PopStackDestroy();
_PushStack(Exp, Invalid(), OpSize::iInvalid);
_PushStack(Sig, Invalid(), OpSize::iInvalid);
_PushStack(Exp, Exp, OpSize::i64Bit, true);
_PushStack(Sig, Sig, OpSize::i64Bit, true);
}
} // namespace FEXCore::IR
@@ -0,0 +1,89 @@
// SPDX-License-Identifier: MIT
/*
$info$
tags: glue|x86-guest-code
desc: Guest-side assembly helpers used by the backends
$end_info$
*/
#include "Interface/Core/X86HelperGen.h"
#include "FEXCore/Utils/AllocatorHooks.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/Utils/Allocator.h>
#include <FEXHeaderUtils/Syscalls.h>
#include <cstdint>
#include <cstring>
namespace FEXCore {
constexpr size_t CODE_SIZE = 0x1000;
X86GeneratedCode::X86GeneratedCode() {
#ifdef _WIN32
// No need to allocate anything in this config.
#else
// Allocate a page for our emulated guest
CodePtr = AllocateGuestCodeSpace(CODE_SIZE);
constexpr std::array<uint8_t, 2> SignalReturnCode = {
0x0F, 0x3E, // CALLBACKRET FEX Instruction
};
CallbackReturn = reinterpret_cast<uint64_t>(CodePtr);
memcpy(reinterpret_cast<void*>(CallbackReturn), SignalReturnCode.data(), SignalReturnCode.size());
mprotect(CodePtr, CODE_SIZE, PROT_READ);
#endif
}
X86GeneratedCode::~X86GeneratedCode() {
#ifndef _WIN32
FEXCore::Allocator::VirtualFree(CodePtr, CODE_SIZE);
#endif
}
void* X86GeneratedCode::AllocateGuestCodeSpace(size_t Size) {
#ifndef _WIN32
FEX_CONFIG_OPT(Is64BitMode, IS64BIT_MODE);
if (Is64BitMode()) {
// 64bit mode can have its sigret handler anywhere
auto Result = FEXCore::Allocator::VirtualAlloc(Size);
FEXCore::Allocator::VirtualName("FEXMem_Misc", reinterpret_cast<void*>(Result), Size);
return Result;
}
// First 64bit page
constexpr uintptr_t LOCATION_MAX = 0x1'0000'0000;
// 32bit mode
// We need to have the sigret handler in the lower 32bits of memory space
// Scan top down and try to allocate a location
for (size_t Location = 0xFFFF'E000; Location != 0x0; Location -= 0x1000) {
void* Ptr = ::mmap(reinterpret_cast<void*>(Location), Size, PROT_READ | PROT_WRITE, MAP_FIXED_NOREPLACE | MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
if (Ptr != MAP_FAILED && reinterpret_cast<uintptr_t>(Ptr) >= LOCATION_MAX) {
// Failed to map in the lower 32bits
// Try again
// Can happen in the case that host kernel ignores MAP_FIXED_NOREPLACE
::munmap(Ptr, Size);
continue;
}
if (Ptr != MAP_FAILED) {
return Ptr;
}
}
// Can't do anything about this
// Here's hoping the application doesn't use signals
return MAP_FAILED;
#else
return nullptr;
#endif
}
} // namespace FEXCore
@@ -0,0 +1,25 @@
// SPDX-License-Identifier: MIT
/*
$info$
tags: glue|x86-guest-code
$end_info$
*/
#pragma once
#include <stddef.h>
#include <stdint.h>
namespace FEXCore {
class X86GeneratedCode final {
public:
X86GeneratedCode();
~X86GeneratedCode();
uint64_t CallbackReturn {};
private:
void* CodePtr {};
void* AllocateGuestCodeSpace(size_t Size);
};
} // namespace FEXCore
+18 -1
View File
@@ -60,7 +60,24 @@ struct NodeID final {
Value = 0;
}
[[nodiscard]] constexpr auto operator<=>(const NodeID&) const noexcept = default;
[[nodiscard]] friend constexpr bool operator==(NodeID, NodeID) noexcept = default;
[[nodiscard]]
friend constexpr bool operator<(NodeID lhs, NodeID rhs) noexcept {
return lhs.Value < rhs.Value;
}
[[nodiscard]]
friend constexpr bool operator>(NodeID lhs, NodeID rhs) noexcept {
return operator<(rhs, lhs);
}
[[nodiscard]]
friend constexpr bool operator<=(NodeID lhs, NodeID rhs) noexcept {
return !operator>(lhs, rhs);
}
[[nodiscard]]
friend constexpr bool operator>=(NodeID lhs, NodeID rhs) noexcept {
return !operator<(lhs, rhs);
}
friend std::ostream& operator<<(std::ostream& out, NodeID ID) {
out << ID.Value;
+20 -5
View File
@@ -804,6 +804,16 @@
"Size == FEXCore::IR::OpSize::i32Bit || Size == FEXCore::IR::OpSize::i64Bit"
]
},
"AtomicXor OpSize:#Size, GPR:$Value, GPR:$Addr": {
"HasSideEffects": true,
"Desc": ["Atomic integer xor",
"IR layout must match Fetch-variant, otherwise DCE IR optimization breaks!"
],
"DestSize": "Size",
"EmitValidation": [
"Size == FEXCore::IR::OpSize::i8Bit || Size == FEXCore::IR::OpSize::i16Bit || Size == FEXCore::IR::OpSize::i32Bit || Size == FEXCore::IR::OpSize::i64Bit"
]
},
"GPR = AtomicSwap OpSize:#Size, GPR:$Value, GPR:$Addr": {
"HasSideEffects": true,
"Desc": ["Atomic integer swap"
@@ -2808,13 +2818,17 @@
"X87": true,
"HasSideEffects": true
},
"PushStack FPR:$X80Src, FPR:$OriginalValue, OpSize:$LoadSize": {
"PushStack FPR:$X80Src, SSA:$OriginalValue, OpSize:$LoadSize, i1:$Float": {
"Desc": [
"Pushes the provided X80Src source on to the x87 stack.",
"Tracks OriginalValue as the original value of X80Src. OriginalValue can be Invalid() in which case no tracking is done.",
"Tracks OriginalValue as the original value of X80Src.",
"Opsize is 128bit for F80 values, 64-bit for low precision.",
"LoadSize the original load size, i.e. of size of OriginalValue.",
"Float: 80-bit, 64-bit, 32-bit"
"Float: 80-bit, 64-bit, 32-bit",
"Int: 64-bit, 32-bit, 16-bit"
],
"EmitValidation": [
"WalkFindRegClass($OriginalValue) == RegClass::FPR || WalkFindRegClass($OriginalValue) == RegClass::GPR"
],
"HasSideEffects": true,
"X87": true
@@ -2826,12 +2840,13 @@
"HasSideEffects": true,
"X87": true
},
"StoreStackMem OpSize:$SourceSize, OpSize:$StoreSize, GPR:$Addr, GPR:$Offset, OpSize:$Align, MemOffsetType:$OffsetType, u8:$OffsetScale": {
"StoreStackMem OpSize:$SourceSize, OpSize:$StoreSize, GPR:$Addr, GPR:$Offset, OpSize:$Align, MemOffsetType:$OffsetType, u8:$OffsetScale, i1:$Float": {
"Desc": [
"Takes the top value off the x87 stack and stores it to memory.",
"SourceSize is 128bit for F80 values, 64-bit for low precision.",
"StoreSize is the store size for conversion:",
"Float: 80-bit, 64-bit, or 32-bit"
"Float: 80-bit, 64-bit, or 32-bit",
"Int: 64-bit, 32-bit, 16-bit"
],
"HasSideEffects": true,
"X87": true
+2 -1
View File
@@ -353,7 +353,8 @@ void Dump(fextl::stringstream* out, const IRListView* IR) {
auto BlockIROp = BlockHeader->C<FEXCore::IR::IROp_CodeBlock>();
AddIndent();
*out << "(%" << IR->GetID(BlockNode) << ") " << "CodeBlock ";
*out << "(%" << IR->GetID(BlockNode) << ") "
<< "CodeBlock ";
*out << "%" << BlockIROp->Begin.ID() << ", ";
*out << "%" << BlockIROp->Last.ID() << std::endl;
@@ -6,6 +6,7 @@
#include "Interface/IR/PassManager.h"
#include "FEXCore/IR/IR.h"
#include "FEXCore/Utils/Profiler.h"
#include "FEXCore/Utils/MathUtils.h"
#include "FEXCore/Core/HostFeatures.h"
#include "Interface/Core/Addressing.h"
@@ -65,7 +66,7 @@ public:
int8_t TopOffset = 0;
FixedSizeStack()
: buffer(FixedSizeStack::size, {StackSlot::UNUSED, T::Invalid}) {}
: buffer(FixedSizeStack::size, {StackSlot::UNUSED, T()}) {}
void push(const T& Value) {
rotate();
@@ -84,7 +85,7 @@ public:
}
void pop() {
buffer.front() = {StackSlot::INVALID, T::Invalid};
buffer.front() = {StackSlot::INVALID, T()};
rotate(false);
}
@@ -102,7 +103,7 @@ public:
void clear() {
for (auto& Elem : buffer) {
Elem = {StackSlot::UNUSED, T::Invalid};
Elem = {StackSlot::UNUSED, T()};
}
TopOffset = 0;
}
@@ -170,8 +171,13 @@ private:
// Helpers
Ref RotateRight8(uint32_t V, Ref Amount);
void F80SplitStore_Helper(const IROp_StoreStackMem* Op, Ref StackNode, Ref AddrNode, Ref Offset, OpSize Align, MemOffsetType OffsetType,
uint8_t OffsetScale) {
void F80SplitStore_Helper(const IROp_StoreStackMem* Op, Ref StackNode) {
Ref AddrNode = IR->GetNode(Op->Addr);
Ref Offset = IR->GetNode(Op->Offset);
OpSize Align = Op->Align;
MemOffsetType OffsetType = Op->OffsetType;
uint8_t OffsetScale = Op->OffsetScale;
IREmit->_StoreMemFPR(OpSize::i64Bit, StackNode, AddrNode, Offset, Align, OffsetType, OffsetScale);
auto Upper = IREmit->_VExtractToGPR(OpSize::i128Bit, OpSize::i64Bit, StackNode, 1);
@@ -186,24 +192,7 @@ private:
IREmit->_StoreMemGPR(OpSize::i16Bit, Upper, A.Base, A.Index, OpSize::i64Bit, MemOffsetType::SXTX, A.IndexScale);
}
void Store80BitToMem(const IROp_StoreStackMem* Op, Ref StackNode, Ref AddrNode, Ref Offset, OpSize Align, MemOffsetType OffsetType,
uint8_t OffsetScale) {
if (Features.SupportsSVE128 || Features.SupportsSVE256) {
AddressMode A {.Base = AddrNode,
.Index = Op->Offset.IsInvalid() ? nullptr : Offset,
.IndexType = MemOffsetType::SXTX,
.IndexScale = OffsetScale,
.AddrSize = OpSize::i64Bit};
AddrNode = LoadEffectiveAddress(IREmit, A, GPROpSize, false);
IREmit->_StoreMemX87SVEOptPredicate(OpSize::i128Bit, OpSize::i16Bit, StackNode, AddrNode);
} else {
F80SplitStore_Helper(Op, StackNode, AddrNode, Offset, Align, OffsetType, OffsetScale);
}
}
void StoreStackMem_Helper(const IROp_StoreStackMem* Op, Ref StackNode) {
LOGMAN_THROW_A_FMT(!ReducedPrecisionMode, "Full precision mode expected.");
Ref AddrNode = IR->GetNode(Op->Addr);
Ref Offset = IR->GetNode(Op->Offset);
OpSize Align = Op->Align;
@@ -220,7 +209,17 @@ private:
}
case OpSize::f80Bit: {
Store80BitToMem(Op, StackNode, AddrNode, Offset, Align, OffsetType, OffsetScale);
if (Features.SupportsSVE128 || Features.SupportsSVE256) {
AddressMode A {.Base = AddrNode,
.Index = Op->Offset.IsInvalid() ? nullptr : Offset,
.IndexType = MemOffsetType::SXTX,
.IndexScale = OffsetScale,
.AddrSize = OpSize::i64Bit};
AddrNode = LoadEffectiveAddress(IREmit, A, GPROpSize, false);
IREmit->_StoreMemX87SVEOptPredicate(OpSize::i128Bit, OpSize::i16Bit, StackNode, AddrNode);
} else { // 80bit requires split-store
F80SplitStore_Helper(Op, StackNode);
}
break;
}
default: ERROR_AND_DIE_FMT("Unsupported x87 size");
@@ -230,8 +229,6 @@ private:
// Performs a store to memory from a value the stack passed in as StackNode.
// This is the version dealing with the reduced precision case.
void StoreStackMem_Reduced_Helper(const IROp_StoreStackMem* Op, Ref StackNode) {
LOGMAN_THROW_A_FMT(ReducedPrecisionMode, "Reduced precision mode expected.");
Ref AddrNode = IR->GetNode(Op->Addr);
Ref Offset = IR->GetNode(Op->Offset);
OpSize Align = Op->Align;
@@ -248,9 +245,10 @@ private:
break;
}
// 80bit requires split-store
case OpSize::f80Bit: {
StackNode = IREmit->_F80CVTTo(StackNode, OpSize::i64Bit);
Store80BitToMem(Op, StackNode, AddrNode, Offset, Align, OffsetType, OffsetScale);
F80SplitStore_Helper(Op, StackNode);
break;
}
default: ERROR_AND_DIE_FMT("Unsupported x87 size");
@@ -292,24 +290,23 @@ private:
void Reset();
struct StackMemberInfo {
StackMemberInfo() = delete;
StackMemberInfo() {}
StackMemberInfo(Ref Data)
: StackDataNode(Data) {}
StackMemberInfo(Ref Data, Ref Source, OpSize Size)
StackMemberInfo(Ref Data, Ref Source, OpSize Size, bool Float)
: StackDataNode(Data)
, Source({Size, Source}) {}
, Source({Size, Source})
, InterpretAsFloat(Float) {}
Ref StackDataNode {}; // Reference to the data in the Stack.
// This is the source data node in the stack format, possibly converted to 64/80 bits.
struct StackMemberData final {
OpSize Size;
Ref Node;
};
static const StackMemberInfo Invalid;
// Tuple is only valid if we have information about the Source of the Stack Data Node.
// In it's valid then OpSize is the original source size and Ref is the original source node.
std::optional<StackMemberData> Source {};
bool InterpretAsFloat {false}; // True if this is a floating point value, false if integer
};
// StackData, TopCache need to be always properly set to ensure
@@ -362,8 +359,6 @@ private:
IRListView* IR = nullptr;
};
inline const X87StackOptimization::StackMemberInfo X87StackOptimization::StackMemberInfo::Invalid {nullptr};
inline void X87StackOptimization::InvalidateCaches() {
InvalidateCachedRegs();
ConstantPool.fill(nullptr);
@@ -733,7 +728,6 @@ void X87StackOptimization::Run(IREmitter* Emit) {
// The optimization should run per-block
Reset();
IREmit->SetCurrentCodeBlock(BlockNode);
for (auto [CodeNode, IROp] : CurrentIR.GetCode(BlockNode)) {
if (!LoweredX87(IROp->Op)) {
continue;
@@ -933,13 +927,8 @@ void X87StackOptimization::Run(IREmitter* Emit) {
StoreStackValueAtOffset_Slow(SourceNode);
} else {
auto* SourceNode = CurrentIR.GetNode(Op->X80Src);
if (Op->OriginalValue.IsInvalid()) {
// No original value to track - just push the converted data
StackData.push(StackMemberInfo {SourceNode});
} else {
auto* OriginalNode = CurrentIR.GetNode(Op->OriginalValue);
StackData.push(StackMemberInfo {SourceNode, OriginalNode, Op->LoadSize});
}
auto* OriginalNode = CurrentIR.GetNode(Op->OriginalValue);
StackData.push(StackMemberInfo {SourceNode, OriginalNode, Op->LoadSize, Op->Float});
}
break;
}
@@ -1004,16 +993,9 @@ void X87StackOptimization::Run(IREmitter* Emit) {
// str w2, [x1]
// or similar. As long as the source size and dest size are one and the same.
// This will avoid any conversions between source and stack element size and conversion back.
OpSize StoreSize = Op->StoreSize;
LOGMAN_THROW_A_FMT(Op->StoreSize == OpSize::i32Bit || Op->StoreSize == OpSize::i64Bit || Op->StoreSize == OpSize::f80Bit,
"Invalid store size in x87 store stack mem");
if (!SlowPath && Value->Source && Value->Source->Size == StoreSize) {
Ref SourceValue = Value->Source->Node;
if (Op->StoreSize == OpSize::f80Bit) {
Store80BitToMem(Op, SourceValue, AddrNode, Offset, Align, OffsetType, OffsetScale);
} else {
IREmit->_StoreMemFPR(StoreSize, SourceValue, AddrNode, Offset, Align, OffsetType, OffsetScale);
}
if (!SlowPath && Value->Source && Value->Source->Size == Op->StoreSize && Value->InterpretAsFloat) {
const auto ClassType = Value->InterpretAsFloat ? RegClass::FPR : RegClass::GPR;
IREmit->_StoreMem(ClassType, Op->StoreSize, Value->Source->Node, AddrNode, Offset, Align, OffsetType, OffsetScale);
break;
}
@@ -1053,26 +1035,11 @@ void X87StackOptimization::Run(IREmitter* Emit) {
case OP_F80STACKXCHANGE: {
const auto* Op = IROp->C<IROp_F80StackXchange>();
auto Offset = Op->SrcStack;
Ref ValueTop = LoadStackValue();
Ref ValueOffset = LoadStackValue(Offset);
if (Offset == 0) {
// No-op
break;
}
const auto [ValidTop, StackMemberTop] = StackData.top(0);
const auto [ValidOffset, StackMemberOffset] = StackData.top(Offset);
if (ValidTop != StackSlot::VALID || ValidOffset != StackSlot::VALID) {
// Slow path: do actual memory operations
Ref ValueTop = LoadStackValue();
Ref ValueOffset = LoadStackValue(Offset);
StoreStackValue(ValueOffset);
StoreStackValue(ValueTop, Offset);
} else {
// Fast path: swap complete StackMemberInfo preserving Source metadata
StackData.setTop(StackMemberOffset, 0);
StackData.setTop(StackMemberTop, Offset);
}
StoreStackValue(ValueOffset);
StoreStackValue(ValueTop, Offset);
break;
}
+4 -1
View File
@@ -140,7 +140,10 @@ void ClearHooks() {
FEXCore::Allocator::mmap = ::mmap;
FEXCore::Allocator::munmap = ::munmap;
Alloc::OSAllocator::ReleaseAllocatorWorkaround(std::move(Alloc64));
// XXX: This is currently a leak.
// We can't work around this yet until static initializers that allocate memory are completely removed from our codebase
// Luckily we only remove this on process shutdown, so the kernel will do the cleanup for us
Alloc64.release();
}
#pragma GCC diagnostic pop
@@ -207,7 +207,7 @@ OSAllocator_64Bit::LiveVMARegion* OSAllocator_64Bit::FindLiveRegionForAddress(ui
uintptr_t RegionBegin = (*it)->SlabInfo->Base;
uintptr_t RegionEnd = RegionBegin + (*it)->SlabInfo->RegionSize;
if (Addr >= RegionBegin && AddrEnd < RegionEnd) {
if (Addr >= RegionBegin && Addr < RegionEnd) {
LiveRegion = *it;
// Leave our loop
break;
@@ -405,18 +405,14 @@ again:
// Mark the pages as used
uintptr_t RegionBegin = LiveRegion->SlabInfo->Base;
uintptr_t MappedBegin = (AllocatedOffset - RegionBegin) >> FEXCore::Utils::FEX_PAGE_SHIFT;
size_t PagesSet {};
for (size_t i = 0; i < NumberOfPages; ++i) {
PagesSet += LiveRegion->UsedPages.TestAndSet(MappedBegin + i) == false;
LiveRegion->UsedPages.Set(MappedBegin + i);
}
// Change our last allocation region
LiveRegion->LastPageAllocation = MappedBegin + NumberOfPages;
LiveRegion->FreeSpace -= PagesSet * FEXCore::Utils::FEX_PAGE_SIZE;
LOGMAN_THROW_A_FMT(LiveRegion->FreeSpace <= LiveRegion->SlabInfo->RegionSize,
"Corrupt LiveRegion free space! 0x{:x} > 0x{:x}. After allocating 0x{:x} (0x{:x} overlapped)", LiveRegion->FreeSpace,
LiveRegion->SlabInfo->RegionSize, length, PagesSet);
LiveRegion->FreeSpace -= length;
}
if (!AllocatedOffset) {
+6 -20
View File
@@ -27,11 +27,6 @@ struct FlexBitSet final {
Memory[Element / MinimumSizeBits] &= ~(1ULL << (Element % MinimumSizeBits));
return Value;
}
bool TestAndSet(size_t Element) {
bool Value = Get(Element);
Memory[Element / MinimumSizeBits] |= (1ULL << (Element % MinimumSizeBits));
return Value;
}
void Set(size_t Element) {
Memory[Element / MinimumSizeBits] |= (1ULL << (Element % MinimumSizeBits));
}
@@ -75,17 +70,12 @@ struct FlexBitSet final {
template<bool WantUnset>
BitsetScanResults BackwardScanForRange(size_t BeginningElement, size_t ElementCount, size_t MinimumElement) {
bool FoundHole {};
// Final element to iterate to.
const size_t FinalElement = MinimumElement + ElementCount - 1;
for (size_t CurrentPage = BeginningElement; CurrentPage >= FinalElement;) {
for (size_t CurrentPage = BeginningElement; CurrentPage >= (MinimumElement + ElementCount);) {
size_t Remaining = ElementCount;
LOGMAN_THROW_A_FMT(CurrentPage <= BeginningElement && CurrentPage >= FinalElement, "BackwardScanForRange: Scanning less than "
"available range");
LOGMAN_THROW_A_FMT(Remaining <= CurrentPage, "Scanning less than available range");
while (Remaining) {
if (this->Get(CurrentPage - Remaining + 1) == WantUnset) {
if (this->Get(CurrentPage - Remaining) == WantUnset) {
// Has an intersecting range
break;
}
@@ -102,7 +92,7 @@ struct FlexBitSet final {
CurrentPage -= Remaining;
} else {
// We have a slab range
return BitsetScanResults {CurrentPage - ElementCount + 1, FoundHole};
return BitsetScanResults {CurrentPage - ElementCount, FoundHole};
}
}
@@ -118,15 +108,11 @@ struct FlexBitSet final {
BitsetScanResults ForwardScanForRange(size_t BeginningElement, size_t ElementCount, size_t ElementsInSet) {
bool FoundHole {};
// Final element to iterate to.
const size_t FinalElement = ElementsInSet - ElementCount + 1;
for (size_t CurrentElement = BeginningElement; CurrentElement <= FinalElement;) {
for (size_t CurrentElement = BeginningElement; CurrentElement < (ElementsInSet - ElementCount);) {
// If we have enough free space, check if we have enough free pages that are contiguous
size_t Remaining = ElementCount;
LOGMAN_THROW_A_FMT(CurrentElement >= BeginningElement && CurrentElement <= FinalElement, "ForwardScanForRange: Scanning less than "
"available range");
LOGMAN_THROW_A_FMT((CurrentElement + Remaining - 1) < ElementsInSet, "Scanning less than available range");
while (Remaining) {
if (this->Get(CurrentElement + Remaining - 1) == WantUnset) {
@@ -53,12 +53,4 @@ public:
namespace Alloc::OSAllocator {
fextl::unique_ptr<Alloc::HostAllocator> Create64BitAllocator();
fextl::unique_ptr<Alloc::HostAllocator> Create64BitAllocatorWithRegions(fextl::vector<FEXCore::Allocator::MemoryRegion>& Regions);
static inline void ReleaseAllocatorWorkaround(fextl::unique_ptr<Alloc::HostAllocator> Allocator) {
// XXX: This is currently a leak.
// We can't work around this yet until static initializers that allocate memory are completely removed from our codebase
// The allocator is also intrusively allocated, so the unique_ptr tries to double free the HostAllocator object.
// Luckily we only remove this on process shutdown, so the kernel will do the cleanup for us
Allocator.release();
}
} // namespace Alloc::OSAllocator
+4 -32
View File
@@ -1,10 +1,7 @@
// SPDX-License-Identifier: MIT
#include <FEXCore/Utils/LongJump.h>
#include <FEXCore/Utils/LogManager.h>
#include <cstring>
namespace FEXCore::UncheckedLongJump {
namespace FEXCore::LongJump {
#if defined(_M_ARM_64)
[[nodiscard]]
FEX_DEFAULT_VISIBILITY FEX_NAKED uint64_t SetJump(JumpBuf& Buffer) {
@@ -35,7 +32,7 @@ FEX_DEFAULT_VISIBILITY FEX_NAKED uint64_t SetJump(JumpBuf& Buffer) {
}
[[noreturn]]
FEX_DEFAULT_VISIBILITY FEX_NAKED void LongJump(const JumpBuf& Buffer, uint64_t Value) {
FEX_DEFAULT_VISIBILITY FEX_NAKED void LongJump(JumpBuf& Buffer, uint64_t Value) {
__asm volatile(R"(
// x0 contains the jumpbuffer
ldp x19, x20, [x0, #( 0 * 8)];
@@ -61,27 +58,6 @@ FEX_DEFAULT_VISIBILITY FEX_NAKED void LongJump(const JumpBuf& Buffer, uint64_t V
)" ::
: "memory");
}
FEX_DEFAULT_VISIBILITY void ManuallyLoadJumpBuf(const JumpBuf& Buffer, uint64_t Value, uint64_t* GPRs, __uint128_t* FPRs, uint64_t* PC) {
// First 12 values are registers [x19,x30].
memcpy(&GPRs[19], &Buffer.Registers[0], sizeof(uint64_t) * 12);
// Next 8 values are [D8,D15]
// Retain upper 64-bits of the register, only modifying lower 64-bits.
for (size_t i = 0; i < 8; ++i) {
memcpy(&FPRs[8 + i], &Buffer.Registers[12 + i], sizeof(uint64_t));
}
// Last value is stack pointer
memcpy(&GPRs[31], &Buffer.Registers[20], sizeof(uint64_t));
// Load the expected value in to X0
GPRs[0] = Value;
// Load the PC with the current LR.
*PC = GPRs[30];
}
#else
[[nodiscard]]
FEX_DEFAULT_VISIBILITY FEX_NAKED uint64_t SetJump(JumpBuf& Buffer) {
@@ -110,7 +86,7 @@ FEX_DEFAULT_VISIBILITY FEX_NAKED uint64_t SetJump(JumpBuf& Buffer) {
}
[[noreturn]]
FEX_DEFAULT_VISIBILITY FEX_NAKED void LongJump(const JumpBuf& Buffer, uint64_t Value) {
FEX_DEFAULT_VISIBILITY FEX_NAKED void LongJump(JumpBuf& Buffer, uint64_t Value) {
__asm volatile(R"(
.intel_syntax noprefix;
// rdi contains the jumpbuffer
@@ -139,9 +115,5 @@ FEX_DEFAULT_VISIBILITY FEX_NAKED void LongJump(const JumpBuf& Buffer, uint64_t V
: "memory");
}
FEX_DEFAULT_VISIBILITY void ManuallyLoadJumpBuf(JumpBuf& Buffer, uint64_t Value, uint64_t* GPRs, __uint128_t* FPRs, uint64_t* PC) {
LOGMAN_MSG_A_FMT("This is unimplemented on x86-64");
}
#endif
} // namespace FEXCore::UncheckedLongJump
} // namespace FEXCore::LongJump
+22 -44
View File
@@ -1,14 +1,9 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <atomic>
#include <chrono>
#include <mutex>
#include <type_traits>
#include <FEXCore/fextl/functional.h>
#include <FEXCore/Utils/EnumUtils.h>
namespace FEXCore::Utils::SpinWaitLock {
/**
* @brief This provides routines to implement implement an "efficient spin-loop" using ARM's WFE and exclusive monitor interfaces.
@@ -128,21 +123,30 @@ static inline uint64_t WFELoadAtomic(uint64_t* Futex) {
return Result;
}
template<typename Pred, typename T>
static inline void WaitPred(T* Futex, T ComparisonValue) {
template<typename T, typename TT = T>
static inline void Wait(T* Futex, TT ExpectedValue) {
auto AtomicFutex = std::atomic_ref<T>(*Futex);
T Result = AtomicFutex.load();
while (!Pred {}(Result, ComparisonValue)) {
// Early exit if possible.
if (Result == ExpectedValue) {
return;
}
do {
Result = LoadExclusive(Futex);
if (Pred {}(Result, ComparisonValue)) {
if (Result == ExpectedValue) {
return;
}
Result = WFELoadAtomic(Futex);
}
} while (Result != ExpectedValue);
}
template void Wait<uint8_t>(uint8_t*, uint8_t);
template void Wait<uint16_t>(uint16_t*, uint16_t);
template void Wait<uint32_t>(uint32_t*, uint32_t);
template void Wait<uint64_t>(uint64_t*, uint64_t);
template<typename T, typename TT>
static inline bool Wait(T* Futex, TT ExpectedValue, const std::chrono::nanoseconds& Timeout) {
auto AtomicFutex = std::atomic_ref<T>(*Futex);
@@ -180,36 +184,20 @@ template bool Wait<uint16_t>(uint16_t*, uint16_t, const std::chrono::nanoseconds
template bool Wait<uint32_t>(uint32_t*, uint32_t, const std::chrono::nanoseconds&);
template bool Wait<uint64_t>(uint64_t*, uint64_t, const std::chrono::nanoseconds&);
template<typename T>
static inline T OneShotWFEBitComparison(T* Futex, T Mask, T Comp) {
#else
template<typename T, typename TT>
static inline void Wait(T* Futex, TT ExpectedValue) {
auto AtomicFutex = std::atomic_ref<T>(*Futex);
T Result = AtomicFutex.load();
// Early exit if possible.
if ((Result & Mask) == Comp) {
return Result;
if (Result == ExpectedValue) {
return;
}
Result = LoadExclusive(Futex);
if ((Result & Mask) == Comp) {
return Result;
}
// Waits for write and returns result.
Result = WFELoadAtomic(Futex);
return Result;
}
#else
template<typename Pred, typename T>
static inline void WaitPred(T* Futex, T ComparisonValue) {
auto AtomicFutex = std::atomic_ref<T>(*Futex);
T Result = AtomicFutex.load();
while (!Pred {}(Result, ComparisonValue)) {
do {
Result = AtomicFutex.load();
}
} while (Result != ExpectedValue);
}
template<typename T, typename TT>
@@ -240,16 +228,6 @@ static inline bool Wait(T* Futex, TT ExpectedValue, const std::chrono::nanosecon
}
#endif
template<typename T, typename TT = T>
static inline void Wait(T* Futex, TT ExpectedValue) {
WaitPred<std::equal_to<>, T>(Futex, ExpectedValue);
}
template void Wait<uint8_t>(uint8_t*, uint8_t);
template void Wait<uint16_t>(uint16_t*, uint16_t);
template void Wait<uint32_t>(uint32_t*, uint32_t);
template void Wait<uint64_t>(uint64_t*, uint64_t);
template<typename T>
static inline void lock(T* Futex) {
auto AtomicFutex = std::atomic_ref<T>(*Futex);
-381
View File
@@ -1,381 +0,0 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <atomic>
#include <cstdint>
#if !defined(_WIN32)
#include <linux/futex.h> /* Definition of FUTEX_* constants */
#include <sys/syscall.h> /* Definition of SYS_* constants */
#include <unistd.h>
#else
#include <synchapi.h>
#endif
#include <FEXCore/Utils/LogManager.h>
#include "Utils/SpinWaitLock.h"
namespace FEXCore::Utils::WritePriorityMutex {
// A custom mutex that prioritizes exclusive locks.
// In highly contested scenarios, this can help minimize overall contention time.
//
// Features:
// - Up to 32767 pending exclusive locks ("writers")
// - Up to 32767 pending shared_locks ("readers")
// - Low-overhead waiting via WFE with a fallback to futex on timeout
// - Direct writer->reader hand-off and vice-versa to further reduce overhead
//
// Trade-offs:
// - No guaranteed order of wake-ups besides prioritizing writers
// - No support for recursive locking
// - We can't use FUTEX_LOCK_PI to enable priority inheritance
class Mutex final {
public:
Mutex() = default;
// Move-only type
Mutex(const Mutex&) = delete;
Mutex& operator=(const Mutex&) = delete;
Mutex(Mutex&& rhs) = delete;
Mutex& operator=(Mutex&&) = delete;
void lock() {
// Try a non-blocking lock first.
if (try_lock()) {
return;
}
// Try a quick WFE write-lock.
if (Attempt_WFE_WriteLock()) {
return;
}
// Still couldn't get it. Start waiting.
auto AtomicFutex = std::atomic_ref<uint32_t>(Futex);
uint32_t Expected {};
uint32_t Desired {};
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
Expected = AtomicFutex.load(std::memory_order_relaxed);
do {
// Increment the number of write waiters.
Desired = Expected + WRITE_WAITER_INCREMENT;
LOGMAN_THROW_A_FMT((Desired & WRITE_WAITER_COUNT_MASK) != 0, "Overflow in write-waiters!");
} while (AtomicFutex.compare_exchange_strong(Expected, Desired, std::memory_order_acq_rel, std::memory_order_acquire) == false);
#else
Expected = AtomicFutex.fetch_add(WRITE_WAITER_INCREMENT);
Desired = Expected + WRITE_WAITER_INCREMENT;
#endif
// Thread added to waiter list.
Expected = Desired;
while (true) {
bool Sleep = false;
do {
if ((Expected & WRITE_OWNED_BIT) == 0 && (Expected & READ_OWNER_COUNT_MASK) == 0) {
// If not write-owned, and no read-owners, try to acquire.
LOGMAN_THROW_A_FMT((Expected & WRITE_WAITER_COUNT_MASK) != 0, "Underflow in write-waiters!");
// Add write-owned bit.
Desired = Expected | WRITE_OWNED_BIT;
// Remove ourselves from the wait list.
Desired -= WRITE_WAITER_INCREMENT;
Sleep = false;
} else {
// Already write-owned or read-locked. Go to sleep.
Desired = Expected;
Sleep = true;
break;
}
} while (AtomicFutex.compare_exchange_strong(Expected, Desired, std::memory_order_acq_rel, std::memory_order_acquire) == false);
if (!Sleep) {
// Acquired early.
LOGMAN_THROW_A_FMT((Desired & WRITE_OWNED_BIT) == WRITE_OWNED_BIT, "Somehow acquired a write-lock without it being set!");
return;
}
FutexWaitForWriteAvailable(Desired);
Expected = AtomicFutex.load(std::memory_order_relaxed);
}
}
void lock_shared() {
// Try an uncontended lock first.
if (try_lock_shared()) {
return;
}
// Try a quick WFE read-lock.
if (Attempt_WFE_ReadLock()) {
return;
}
auto AtomicFutex = std::atomic_ref<uint32_t>(Futex);
uint32_t Expected = AtomicFutex.load(std::memory_order_relaxed);
uint32_t Desired {};
while (true) {
bool Sleep = false;
do {
if ((Expected & WRITE_OWNED_BIT) == 0 && (Expected & WRITE_WAITER_COUNT_MASK) == 0) {
// If no write-owner and no write-waiting, try and acquire.
Desired = Expected + READ_OWNER_INCREMENT;
LOGMAN_THROW_A_FMT((Desired & READ_OWNER_COUNT_MASK) != 0, "Overflow in read-owners!");
Sleep = false;
} else {
// Waiting for lock to become available. Add to waiters.
Desired = Expected | READ_WAITER_BIT;
Sleep = true;
}
} while (AtomicFutex.compare_exchange_strong(Expected, Desired, std::memory_order_acq_rel, std::memory_order_acquire) == false);
if (!Sleep) {
// Acquired early.
LOGMAN_THROW_A_FMT((Desired & WRITE_OWNED_BIT) != WRITE_OWNED_BIT, "Somehow read-locked and got a write lock!");
return;
}
FutexWaitForReadAvailable(Desired);
Expected = AtomicFutex.load(std::memory_order_relaxed);
}
}
void unlock() {
auto AtomicFutex = std::atomic_ref<uint32_t>(Futex);
uint32_t Expected = AtomicFutex.load(std::memory_order_relaxed);
uint32_t Desired {};
do {
LOGMAN_THROW_A_FMT((Expected & WRITE_OWNED_BIT) == WRITE_OWNED_BIT, "Trying to write-unlock something not write-locked!");
// Remove the exclusive lock bit.
Desired = Expected & ~WRITE_OWNED_BIT;
// If no more writers, then make sure to clear the read-waiters bit as well.
if ((Desired & WRITE_WAITER_COUNT_MASK) == 0) {
Desired &= ~READ_WAITER_BIT;
}
} while (AtomicFutex.compare_exchange_strong(Expected, Desired, std::memory_order_acq_rel, std::memory_order_acquire) == false);
// If success, then `Expected` has old value. Containing `READ_WAITER_BIT` which was just masked off, and also `WRITE_WAITER_COUNT_MASK`.
if ((Expected & WRITE_WAITER_COUNT_MASK)) {
// Handle write-write handoff.
FutexWakeWriter();
} else if ((Expected & READ_WAITER_BIT)) {
// Handle write-reader handoff.
FutexWakeReaders();
}
}
void unlock_shared() {
auto AtomicFutex = std::atomic_ref<uint32_t>(Futex);
uint32_t Desired {};
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
uint32_t Expected = AtomicFutex.load(std::memory_order_relaxed);
do {
LOGMAN_THROW_A_FMT((Expected & WRITE_OWNED_BIT) != WRITE_OWNED_BIT, "Trying to read-unlock something write-locked!");
LOGMAN_THROW_A_FMT((Expected & READ_OWNER_COUNT_MASK) != 0, "Trying to read-unlock something not read-locked!");
// Decrement the shared counter.
Desired = Expected - READ_OWNER_INCREMENT;
} while (AtomicFutex.compare_exchange_strong(Expected, Desired, std::memory_order_acq_rel, std::memory_order_acquire) == false);
#else
Desired = AtomicFutex.fetch_sub(READ_OWNER_INCREMENT) - READ_OWNER_INCREMENT;
#endif
// Handle read->write handoff if there are any waiting writers, and no readers left.
if ((Desired & WRITE_WAITER_COUNT_MASK) && (Desired & READ_OWNER_COUNT_MASK) == 0) {
FutexWakeWriter();
}
}
bool try_lock() {
auto AtomicFutex = std::atomic_ref<uint32_t>(Futex);
uint32_t Expected = 0;
// Try and grab the owned bit.
uint32_t Desired = WRITE_OWNED_BIT;
// try to CAS immediately.
return AtomicFutex.compare_exchange_strong(Expected, Desired, std::memory_order_acq_rel, std::memory_order_acquire);
}
// Can race with other threads trying to lock shared!
bool try_lock_shared() {
auto AtomicFutex = std::atomic_ref<uint32_t>(Futex);
uint32_t Expected = AtomicFutex.load(std::memory_order_relaxed);
// Exclusively owned or has a list of waiting owners. Can't pass.
if ((Expected & WRITE_OWNED_BIT) || (Expected & WRITE_WAITER_COUNT_MASK)) {
return false;
}
// Try to add reader.
uint32_t Desired = Expected + READ_OWNER_INCREMENT;
LOGMAN_THROW_A_FMT((Desired & READ_OWNER_COUNT_MASK) != 0, "Overflow in read-owners!");
// Uncontended mutex check
return AtomicFutex.compare_exchange_strong(Expected, Desired, std::memory_order_acq_rel, std::memory_order_acquire);
}
#if !defined(_WIN32)
// Initialize the internal mutex object to its default initializer state.
// Should only ever be used in the child process when a Linux fork() has occured.
void StealAndDropActiveLocks() {
Futex = 0;
}
#endif
private:
#if !defined(_WIN32)
void FutexWaitForWriteAvailable(uint32_t Expected) {
::syscall(SYS_futex, &Futex, FUTEX_PRIVATE_FLAG | FUTEX_WAIT_BITSET, Expected, nullptr, nullptr, FUTEX_BITSET_WAIT_WRITERS);
}
// Read-lock waiting for writers to drain out.
void FutexWaitForReadAvailable(uint32_t Expected) {
::syscall(SYS_futex, &Futex, FUTEX_PRIVATE_FLAG | FUTEX_WAIT_BITSET, Expected, nullptr, nullptr, FUTEX_BITSET_WAIT_READERS);
}
// Read-Lock or Write-lock unlocked, wake one writer.
// - Read->Write handoff.
// - Write->Write handoff.
void FutexWakeWriter() {
::syscall(SYS_futex, &Futex, FUTEX_PRIVATE_FLAG | FUTEX_WAKE_BITSET, 1, nullptr, nullptr, FUTEX_BITSET_WAIT_WRITERS);
}
// Write-lock unlocked, wake read-locks waiting.
void FutexWakeReaders() {
// Wake all readers.
::syscall(SYS_futex, &Futex, FUTEX_PRIVATE_FLAG | FUTEX_WAKE_BITSET, INT_MAX, nullptr, nullptr, FUTEX_BITSET_WAIT_READERS);
}
#else
// Writers wait for the full 32-bit futex.
void FutexWaitForWriteAvailable(uint32_t Expected) {
WaitOnAddress(&Futex, &Expected, sizeof(Futex), INFINITE);
}
// Readers wait for Futex bits [31:16] to be zero.
void FutexWaitForReadAvailable(uint32_t Expected) {
auto ReadWaiterAddress = reinterpret_cast<uint8_t*>(&Futex) + 2;
uint16_t smol_Expected = Expected >> 16;
WaitOnAddress(ReadWaiterAddress, &smol_Expected, sizeof(smol_Expected), INFINITE);
}
void FutexWakeWriter() {
WakeByAddressSingle(&Futex);
}
void FutexWakeReaders() {
auto ReadWaiterAddress = reinterpret_cast<uint8_t*>(&Futex) + 2;
WakeByAddressAll(ReadWaiterAddress);
}
#endif
// Reuse the SpinWaitLock WFE implementations for read/write lock acquiring with WFE.
// Can't reuse the spin-lock directly as some bit-representations are different.
// WFE-write-lock is less likely to occur the more read-lock threads are participating. Can still occur so good to try.
// WFE-read-lock is actually quite likely to succeed.
// Return: true if the lock was acquired.
bool Attempt_WFE_WriteLock() {
#ifdef _M_ARM_64
const auto Begin = FEXCore::Utils::SpinWaitLock::GetCycleCounter();
auto Now = Begin;
const auto Duration = FEXCore::Utils::SpinWaitLock::CycleCounterFrequency / CYCLECOUNT_DIVISOR;
auto AtomicFutex = std::atomic_ref<uint32_t>(Futex);
uint32_t Expected = AtomicFutex.load(std::memory_order_relaxed);
while ((Now - Begin) < Duration) {
if (Expected == 0) {
// Try and grab the owned bit.
uint32_t Desired = WRITE_OWNED_BIT;
if (AtomicFutex.compare_exchange_strong(Expected, Desired, std::memory_order_acq_rel, std::memory_order_acquire)) {
return true;
}
}
// One-shot attempt to wait for mask to be zero.
Expected = FEXCore::Utils::SpinWaitLock::OneShotWFEBitComparison(&Futex, ~0U, 0U);
Now = FEXCore::Utils::SpinWaitLock::GetCycleCounter();
}
#endif
return false;
}
// Return: true if the lock was acquired.
bool Attempt_WFE_ReadLock() {
#ifdef _M_ARM_64
// Spin on a WFE for a short-amount of time, waiting for write-owned and writer-count to be zero.
// - Attempt to acquire read-lock at that point.
// - Don't add read-waiters bit on failure, return false.
const auto Begin = FEXCore::Utils::SpinWaitLock::GetCycleCounter();
auto Now = Begin;
const auto Duration = FEXCore::Utils::SpinWaitLock::CycleCounterFrequency / CYCLECOUNT_DIVISOR;
auto AtomicFutex = std::atomic_ref<uint32_t>(Futex);
uint32_t Expected = AtomicFutex.load(std::memory_order_relaxed);
uint32_t Desired {};
while ((Now - Begin) < Duration) {
if ((Expected & WRITE_OWNED_BIT) == 0 && (Expected & WRITE_WAITER_COUNT_MASK) == 0) {
// If no write-owner and no write-waiting, try and acquire.
Desired = Expected + READ_OWNER_INCREMENT;
LOGMAN_THROW_A_FMT((Desired & READ_OWNER_COUNT_MASK) != 0, "Overflow in read-owners!");
if (AtomicFutex.compare_exchange_strong(Expected, Desired, std::memory_order_acq_rel, std::memory_order_acquire)) {
return true;
}
}
// One-shot attempt to wait for mask to be zero.
Expected = FEXCore::Utils::SpinWaitLock::OneShotWFEBitComparison(&Futex, WRITE_OWNED_BIT | WRITE_WAITER_COUNT_MASK, 0U);
Now = FEXCore::Utils::SpinWaitLock::GetCycleCounter();
}
#endif
return false;
}
constexpr static uint32_t WRITE_OWNED_BIT = 1U << 31;
constexpr static uint32_t READ_WAITER_BIT = 1U << 15;
constexpr static uint32_t WRITE_WAITER_OFFSET = 16;
constexpr static uint32_t WRITE_WAITER_INCREMENT = 1U << WRITE_WAITER_OFFSET;
constexpr static uint32_t READ_OWNER_INCREMENT = 1;
// Count masks
constexpr static uint32_t WRITE_WAITER_COUNT_MASK = 0x7FFFU << WRITE_WAITER_OFFSET;
constexpr static uint32_t READ_OWNER_COUNT_MASK = 0x7FFFU;
// Independent futex bit-set masks.
// Wait for readers to drain.
constexpr static uint32_t FUTEX_BITSET_WAIT_READERS = 1U << 0;
// Wait for writers to drain.
constexpr static uint32_t FUTEX_BITSET_WAIT_WRITERS = 1U << 1;
// Only spin on WFE for 0.01ms (10k ns).
constexpr static uint64_t CYCLECOUNT_DIVISOR = 1'000'000'000ULL / 10'000U;
// Layout:
// Bits[31]: Write-lock bit.
// Bits[30:16]: Write-waiter count.
// Bits[15]: Read-waiter bit.
// Bits[14:0]: Read-owner count.
uint32_t Futex {};
};
} // namespace FEXCore::Utils::WritePriorityMutex
+34 -64
View File
@@ -103,25 +103,28 @@ static inline std::optional<fextl::string> EnumParser(const ArrayPairType& EnumP
return fextl::fmt::format("{}", EnumMask);
}
using StringArrayType = fextl::list<fextl::string>;
namespace detail {
template<ConfigOption Option>
struct ConfigOptionInfo;
#define DEFINE_METAINFO(type, enum, default) \
template<> \
struct ConfigOptionInfo<ConfigOption::CONFIG_##enum> { \
using Type = type; \
static auto Default() { \
extern default; \
return enum; \
} \
};
#define OPT_BASE(type, group, enum, json, default) DEFINE_METAINFO(type, enum, const type enum)
#define OPT_STR(group, enum, json, default) DEFINE_METAINFO(fextl::string, enum, const std::string_view enum)
#define OPT_STRARRAY(group, enum, json, default) DEFINE_METAINFO(StringArrayType, enum, const std::string_view enum)
namespace DefaultValues {
#define P(x) x
#define OPT_BASE(type, group, enum, json, default) extern const P(type) P(enum);
#define OPT_STR(group, enum, json, default) extern const std::string_view P(enum);
#define OPT_STRARRAY(group, enum, json, default) OPT_STR(group, enum, json, default)
#include <FEXCore/Config/ConfigValues.inl>
} // namespace detail
namespace Type {
using StringArrayType = fextl::list<fextl::string>;
#define OPT_BASE(type, group, enum, json, default) using P(enum) = P(type);
#define OPT_STR(group, enum, json, default) using P(enum) = fextl::string;
#define OPT_STRARRAY(group, enum, json, default) using P(enum) = StringArrayType;
#include <FEXCore/Config/ConfigValues.inl>
} // namespace Type
#define FEX_CONFIG_OPT(name, enum) \
FEXCore::Config::Value<FEXCore::Config::DefaultValues::Type::enum> name { \
FEXCore::Config::CONFIG_##enum, \
FEXCore::Config::DefaultValues::enum \
}
#undef P
} // namespace DefaultValues
FEX_DEFAULT_VISIBILITY void SetDataDirectory(std::string_view Path, bool Global);
FEX_DEFAULT_VISIBILITY void SetConfigDirectory(const std::string_view Path, bool Global);
@@ -132,7 +135,8 @@ FEX_DEFAULT_VISIBILITY const fextl::string& GetConfigDirectory(bool Global);
FEX_DEFAULT_VISIBILITY const fextl::string& GetConfigFileLocation(bool Global = false);
FEX_DEFAULT_VISIBILITY fextl::string GetApplicationConfig(const std::string_view Program, bool Global);
using LayerValue = std::variant< fextl::string, StringArrayType, uint8_t, int8_t, uint16_t, int16_t, uint32_t, int32_t, uint64_t, int64_t, bool >;
using LayerValue =
std::variant< fextl::string, DefaultValues::Type::StringArrayType, uint8_t, int8_t, uint16_t, int16_t, uint32_t, int32_t, uint64_t, int64_t, bool >;
using LayerOptions = fextl::unordered_map<ConfigOption, LayerValue>;
@@ -147,16 +151,16 @@ public:
return OptionMap.find(Option) != OptionMap.end();
}
std::optional<StringArrayType*> All(ConfigOption Option) {
std::optional<DefaultValues::Type::StringArrayType*> All(ConfigOption Option) {
const auto it = OptionMap.find(Option);
if (it == OptionMap.end()) {
return std::nullopt;
}
auto& Value = it->second;
LOGMAN_THROW_A_FMT(std::holds_alternative<StringArrayType>(Value), "Tried to get config of invalid type!");
LOGMAN_THROW_A_FMT(std::holds_alternative<DefaultValues::Type::StringArrayType>(Value), "Tried to get config of invalid type!");
return &std::get<StringArrayType>(Value);
return &std::get<DefaultValues::Type::StringArrayType>(Value);
}
std::optional<fextl::string*> Get(ConfigOption Option) {
@@ -197,12 +201,12 @@ public:
auto it = OptionMap.find(Option);
if (it == OptionMap.end()) {
// If the option didn't exist as a StringArrayType yet, emplace it.
it = OptionMap.emplace(Option, StringArrayType {}).first;
it = OptionMap.emplace(Option, DefaultValues::Type::StringArrayType {}).first;
}
auto& Value = it->second;
LOGMAN_THROW_A_FMT(std::holds_alternative<StringArrayType>(Value), "Tried to get config of invalid type!");
std::get<StringArrayType>(Value).emplace_back(Data);
LOGMAN_THROW_A_FMT(std::holds_alternative<DefaultValues::Type::StringArrayType>(Value), "Tried to get config of invalid type!");
std::get<DefaultValues::Type::StringArrayType>(Value).emplace_back(Data);
}
void Erase(ConfigOption Option) {
@@ -232,9 +236,7 @@ FEX_DEFAULT_VISIBILITY fextl::string FindContainerPrefix();
FEX_DEFAULT_VISIBILITY void AddLayer(fextl::unique_ptr<FEXCore::Config::Layer> _Layer);
FEX_DEFAULT_VISIBILITY bool Exists(ConfigOption Option);
FEX_DEFAULT_VISIBILITY std::optional<StringArrayType*> All(ConfigOption Option);
template<typename T>
FEX_DEFAULT_VISIBILITY std::optional<T> GetConv(ConfigOption Option);
FEX_DEFAULT_VISIBILITY std::optional<DefaultValues::Type::StringArrayType*> All(ConfigOption Option);
FEX_DEFAULT_VISIBILITY std::optional<fextl::string*> Get(ConfigOption Option);
FEX_DEFAULT_VISIBILITY void Set(ConfigOption Option, std::string_view Data);
FEX_DEFAULT_VISIBILITY void Erase(ConfigOption Option);
@@ -269,18 +271,18 @@ public:
return ValueData;
}
Value(T Value) requires (!std::is_same_v<T, StringArrayType>)
Value(T Value) requires (!std::is_same_v<T, DefaultValues::Type::StringArrayType>)
{
ValueData = std::move(Value);
}
// Array value types.
Value(FEXCore::Config::ConfigOption Option, std::string_view) requires (std::is_same_v<T, StringArrayType>)
Value(FEXCore::Config::ConfigOption Option, std::string_view) requires (std::is_same_v<T, DefaultValues::Type::StringArrayType>)
{
GetListIfExists(Option, &ValueData);
}
StringArrayType& All() requires (std::is_same_v<T, StringArrayType>)
DefaultValues::Type::StringArrayType& All() requires (std::is_same_v<T, DefaultValues::Type::StringArrayType>)
{
return ValueData;
}
@@ -291,38 +293,6 @@ private:
static T GetIfExists(FEXCore::Config::ConfigOption Option, T Default);
static T GetIfExists(FEXCore::Config::ConfigOption Option, std::string_view Default);
static void GetListIfExists(FEXCore::Config::ConfigOption Option, StringArrayType* List);
static void GetListIfExists(FEXCore::Config::ConfigOption Option, DefaultValues::Type::StringArrayType* List);
};
/**
* Wrapper around Value that automatically picks the default for the given ConfigOption
*/
template<ConfigOption Option>
struct FEX_DEFAULT_VISIBILITY Getter : public Value<typename detail::ConfigOptionInfo<Option>::Type> {
using OptionInfo = detail::ConfigOptionInfo<Option>;
Getter()
: Value<typename OptionInfo::Type> {Option, OptionInfo::Default()} {}
};
/**
* Helper for reading a config value with caching.
*
* Typically this is used to declare class members so that the value is read
* on construction of the parent.
*/
#define FEX_CONFIG_OPT(name, enum) FEXCore::Config::Getter<FEXCore::Config::ConfigOption::CONFIG_##enum> name {}
#define OPT_BASE(type, group, enum, json, default) \
/** \
* Helper for reading a config value. \
* \
* In contrast to FEX_CONFIG_OPT, this can be used in arbitrary expressions, \
* at the expense of not caching the value. Use Getter instead if the value \
* is read frequently. \
*/ \
inline auto Get_##enum() { \
return Getter<FEXCore::Config::ConfigOption::CONFIG_##enum> {}; \
}
#include <FEXCore/Config/ConfigValues.inl>
} // namespace FEXCore::Config
+1 -136
View File
@@ -1,20 +1,10 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <FEXCore/fextl/functional.h>
#include <FEXCore/fextl/map.h>
#include <FEXCore/fextl/memory.h>
#include <FEXCore/fextl/set.h>
#include <FEXCore/fextl/string.h>
#include <FEXCore/fextl/vector.h>
#include <atomic>
#include <cstdint>
#include <mutex>
#include <optional>
#include <shared_mutex>
#include <span>
#include <unistd.h>
namespace FEXCore {
@@ -30,14 +20,8 @@ namespace HLE {
struct ExecutableFileInfo {
~ExecutableFileInfo();
#if __clang_major__ < 16
// Workaround for broken aggregate-initialization with std::piecewise_construct
ExecutableFileInfo(fextl::unique_ptr<HLE::SourcecodeMap>, uint64_t, fextl::string);
ExecutableFileInfo() = default;
#endif
fextl::unique_ptr<HLE::SourcecodeMap> SourcecodeMap;
uint64_t FileId = 0;
fextl::string FileId;
fextl::string Filename;
};
@@ -49,129 +33,10 @@ struct ExecutableFileSectionInfo {
uintptr_t FileStartVA;
};
using CodeMapFileId = uint64_t;
/**
* Code maps capture information required for offline code cache generation
* and are written to disk during execution of FEX.
*
* Almost all CodeMap data will be an Entry that indicates blocks to be
* compiled for cache generation. The reserved value `LoadExternalLibrary`
* indicates that an instance of ExternalLibraryInfo follows (the entry data
* itself should be skipped in that case).
*/
struct CodeMap {
// Describes the location of an entry block compiled during execution
struct FEX_PACKED Entry {
CodeMapFileId FileId;
uint32_t BlockOffset;
};
// Describes an external library referenced during execution
struct ExternalLibraryInfo {
CodeMapFileId ExternalFileId;
// null-terminated file path; EITHER relative to the main executable OR an absolute path OR starting with a magic identifier:
// - WINE/: Path to Wine/Proton installation
// - WINEPREFIX/: Path to Wine/Proton prefix
// - SLR/: Path to Steam Linux Runtime
// At runtime, FEX will always dump absolute paths
char Path[];
// Followed by padding to a 4 byte boundary
};
// Followed by ExternalLibraryInfo
static constexpr Entry LoadExternalLibrary = {0xffff'ffff'ffff'ffff, 0xffff'ffff};
struct FEX_PACKED SetExecutableFileId {
Entry Marker = {0xffff'ffff'ffff'ffff, 0xffff'fffe};
CodeMapFileId ExecutableFileId;
};
struct ParsedContents {
fextl::string Filename;
fextl::set<uint64_t> Blocks;
bool IsExecutable = false;
};
// Follows scheme fileid[-nomb]
// The nomb ("no multiblock") suffix signifies that the code map is for use without multiblock, only.
static fextl::string GetBaseFilename(const ExecutableFileInfo& MainExecutable, bool AddNombSuffix);
static fextl::map<CodeMapFileId, ParsedContents> ParseCodeMap(std::ifstream& File);
};
struct CodeMapOpener {
virtual ~CodeMapOpener() = default;
virtual int OpenCodeMapFile() = 0;
};
class CodeMapWriter {
public:
CodeMapWriter(CodeMapOpener&, bool OpenEagerly = false);
~CodeMapWriter();
// Checks if writing is enabled. Calls to this functions may also be interpreted as signals that writes are about to happen
bool IsWriteEnabled(const ExecutableFileSectionInfo&);
void ResetAfterFork() {
if (CodeMapFD.value_or(-1) != -1) {
close(CodeMapFD.value());
CodeMapFD.reset();
}
BufferOffset = 0;
KnownFileIds.clear();
}
bool IsBackingFD(int FD) const {
if (FD == CodeMapFD) {
LogMan::Msg::DFmt("Hiding directory entry for code map FD");
return true;
}
return false;
}
void AppendBlock(const FEXCore::ExecutableFileSectionInfo&, uint64_t Entry);
void AppendLibraryLoad(const FEXCore::ExecutableFileInfo&);
void AppendSetMainExecutable(const FEXCore::ExecutableFileInfo&);
// Thread-safely commit any pending data to disk
void Flush(size_t Offset);
private:
// Queues data into an internal ring buffer.
// Call Flush() to commit the data to disk.
void AppendData(std::span<const std::byte> Data);
// Commit given data range to disk
void Flush(size_t Offset, std::unique_lock<std::shared_mutex>&);
std::shared_mutex Mutex;
fextl::vector<std::byte> Buffer;
std::atomic<size_t> BufferOffset {0};
fextl::set<CodeMapFileId> KnownFileIds;
// std::nullopt: We haven't requested a CodeMapFD yet
// value is -1: We requested a CodeMapFD but FEXServer told us not to write any data
// other values: Code map writing is active
std::optional<int> CodeMapFD;
CodeMapOpener& FileOpener;
};
class AbstractCodeCache {
public:
virtual ~AbstractCodeCache() = default;
/**
* Computes a unique identifier for the referenced binary file to be used for
* generating the code map.
* This identifier is independent of FEX build/runtime configuration and
* stable across FEX updates.
*/
virtual uint64_t ComputeCodeMapId(std::string_view Filename, int FD) = 0;
/**
* Loads a code cache from mapped memory and appends it to the current Core state.
* TODO: Optionally recompiles all contained code blocks at runtime for validation.
-2
View File
@@ -136,8 +136,6 @@ public:
FEX_DEFAULT_VISIBILITY virtual FEXCore::CPUID::FunctionResults RunCPUIDFunctionName(uint32_t Function, uint32_t Leaf, uint32_t CPU) = 0;
virtual AbstractCodeCache& GetCodeCache() = 0;
virtual void SetCodeMapWriter(fextl::unique_ptr<CodeMapWriter>) = 0;
virtual void FlushAndCloseCodeMap() = 0;
FEX_DEFAULT_VISIBILITY virtual void ClearCodeCache(FEXCore::Core::InternalThreadState* Thread, bool NewCodeBuffer = true) = 0;
FEX_DEFAULT_VISIBILITY virtual void InvalidateCodeBuffersCodeRange(uint64_t Start, uint64_t Length) = 0;
+5
View File
@@ -353,6 +353,7 @@ struct JITPointers {
uint64_t ExitFunctionLinker {};
uint64_t ThreadStopHandlerSpillSRA {};
uint64_t ThreadPauseHandlerSpillSRA {};
uint64_t UnimplementedInstructionHandler {};
uint64_t GuestSignal_SIGILL {};
uint64_t GuestSignal_SIGTRAP {};
uint64_t GuestSignal_SIGSEGV {};
@@ -370,6 +371,8 @@ struct JITPointers {
// Process specific
uint64_t LUDIV {};
uint64_t LDIV {};
uint64_t LUREM {};
uint64_t LREM {};
// Thread Specific
@@ -378,6 +381,8 @@ struct JITPointers {
* @{ */
uint64_t LUDIVHandler {};
uint64_t LDIVHandler {};
uint64_t LUREMHandler {};
uint64_t LREMHandler {};
/** @} */
} AArch64;
@@ -67,10 +67,6 @@ public:
return Config;
}
virtual uintptr_t GetThunkCallbackRET() const {
return 0;
}
protected:
SignalDelegatorConfig Config;
};
@@ -4,7 +4,6 @@
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Utils/AllocatorHooks.h>
#include <FEXCore/Utils/TypeDefines.h>
#include <FEXCore/Utils/LongJump.h>
#include <FEXCore/fextl/memory.h>
#include <FEXCore/fextl/vector.h>
@@ -119,10 +118,6 @@ struct alignas(FEXCore::Utils::FEX_PAGE_SIZE) InternalThreadState : public FEXCo
// The low address of the call-ret stack allocation (not including guard pages)
void* CallRetStackBase {};
uintptr_t JITGuardPage {};
uint64_t JITGuardOverflowArgument {};
FEXCore::UncheckedLongJump::JumpBuf RestartJump;
// BaseFrameState should always be at the end, directly before the interrupt fault page
alignas(16) FEXCore::Core::CpuStateFrame BaseFrameState {};
+9 -2
View File
@@ -3,7 +3,6 @@
#include <FEXCore/Utils/EnumOperators.h>
#include <compare>
#include <cstdint>
#include <cstring>
@@ -77,7 +76,15 @@ enum IndexNamedVectorConstant : uint8_t {
struct SHA256Sum final {
uint8_t data[32];
[[nodiscard]] auto operator<=>(const SHA256Sum&) const noexcept = default;
[[nodiscard]]
bool operator<(const SHA256Sum& rhs) const {
return memcmp(data, rhs.data, sizeof(data)) < 0;
}
[[nodiscard]]
bool operator==(const SHA256Sum& rhs) const {
return memcmp(data, rhs.data, sizeof(data)) == 0;
}
};
typedef void ThunkedFunction(void* ArgsRv);
+4 -6
View File
@@ -5,9 +5,8 @@
#include <cstdint>
// Reimplementation of longjmp without glibc fortification checks.
// This is useful when false positives need to be avoided or when using
// a libc implementation that does not implement std::longjmp.
namespace FEXCore::UncheckedLongJump {
// This is useful to avoid false positives reported by glibc.
namespace FEXCore::LongJump {
// JumpBuf definition needs to be public because the frontend needs to understand it.
#if defined(_M_ARM_64)
struct JumpBuf {
@@ -34,6 +33,5 @@ struct JumpBuf {
#endif
[[nodiscard]] FEX_DEFAULT_VISIBILITY uint64_t SetJump(JumpBuf& Buffer);
[[noreturn]] FEX_DEFAULT_VISIBILITY void LongJump(const JumpBuf& Buffer, uint64_t Value);
FEX_DEFAULT_VISIBILITY void ManuallyLoadJumpBuf(const JumpBuf& Buffer, uint64_t Value, uint64_t* GPRs, __uint128_t* FPRs, uint64_t* PC);
} // namespace FEXCore::UncheckedLongJump
[[noreturn]] FEX_DEFAULT_VISIBILITY void LongJump(JumpBuf& Buffer, uint64_t Value);
} // namespace FEXCore::LongJump
-1
View File
@@ -1,7 +1,6 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <atomic>
#include <cstddef>
#include <cstdint>
#ifdef _M_X86_64
@@ -28,19 +28,4 @@ inline fextl::string Trim(fextl::string String, std::string_view TrimTokens = "
return RightTrim(LeftTrim(std::move(String), TrimTokens), TrimTokens);
}
inline fextl::string& ReplaceAllInPlace(fextl::string& Str, std::string_view Token, std::string_view New) {
const auto OriginalTokenSize = Token.size();
const auto NewTokenSize = New.size();
size_t TokenPos {};
auto TokenIter = Str.find(Token, TokenPos);
while (TokenIter != Str.npos) {
Str.replace(TokenIter, OriginalTokenSize, New);
TokenPos += NewTokenSize;
TokenIter = Str.find(Token, TokenPos);
}
return Str;
}
} // namespace FEXCore::StringUtils
@@ -3,8 +3,6 @@
#include <FEXCore/Utils/Allocator.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/MathUtils.h>
#include <FEXCore/Utils/TypeDefines.h>
#include <FEXCore/fextl/list.h>
#include <atomic>
@@ -39,6 +37,12 @@ namespace FEXCore::Utils {
*/
class IntrusivePooledAllocator {
public:
template<typename T>
struct AllocationInfo {
T Ptr;
size_t Size;
};
struct MemoryBuffer;
/**
* @brief Container for tracking the buffers
@@ -399,42 +403,6 @@ private:
const char* Name {};
};
/**
* @brief Thread pool allocator that allocates and frees objects that uses mmap, with a guard page.
*
* The last page of the size provided has the guard.
*/
class PooledAllocatorVirtualWithGuard final : public IntrusivePooledAllocator {
public:
PooledAllocatorVirtualWithGuard() = default;
PooledAllocatorVirtualWithGuard(const char* Name)
: Name {Name} {}
virtual ~PooledAllocatorVirtualWithGuard() {
FreeAllBuffers();
}
private:
void* Alloc(size_t Size) override {
auto Ptr = FEXCore::Allocator::VirtualAlloc(Size);
uintptr_t LastPageAddr = AlignDown(reinterpret_cast<uintptr_t>(Ptr) + Size - 1, FEXCore::Utils::FEX_PAGE_SIZE);
if (!FEXCore::Allocator::VirtualProtect(reinterpret_cast<void*>(LastPageAddr), FEXCore::Utils::FEX_PAGE_SIZE,
FEXCore::Allocator::ProtectOptions::None)) {
LogMan::Msg::EFmt("Failed to mprotect last page of code buffer.");
}
if (Name) {
FEXCore::Allocator::VirtualName(Name, Ptr, Size);
}
return Ptr;
}
void Free(void* Ptr, size_t Size) override {
FEXCore::Allocator::VirtualFree(Ptr, Size);
}
const char* Name {};
};
/**
* @brief Wrapper around the pool allocator for delayed pool reclaiming
*
@@ -492,11 +460,6 @@ public:
UnclaimBuffer();
}
struct AllocationInfo {
Type Ptr;
size_t Size;
};
/**
* @brief Return the owned buffer or allocate another one from the `Allocator`
*
@@ -505,9 +468,9 @@ public:
*
* @param NewSize Optional new size for managed data
*
* @return A usable pointer of type `Type` and the size of the backing store.
* @return object of type `Type` allocated within the selected buffer
*/
AllocationInfo ReownOrClaimBufferWithSize(std::optional<size_t> NewSize = std::nullopt) {
Type ReownOrClaimBuffer(std::optional<size_t> NewSize = std::nullopt) {
// Check if we can cheaply re-own a previous buffer
std::optional Buffer =
IntrusivePooledAllocator::IsClientBufferOwned(ClientOwnedFlag) ? Info : ThreadAllocator.TryToReownBuffer(Info, Size, &ClientOwnedFlag);
@@ -530,14 +493,7 @@ public:
// Leaving this here for future excavation that will definitely occur here
// memset((*Info)->Ptr, 0, Size);
return {
.Ptr = reinterpret_cast<Type>((*Info)->Ptr),
.Size = (*Info)->Size,
};
}
Type ReownOrClaimBuffer(std::optional<size_t> NewSize = std::nullopt) {
return ReownOrClaimBufferWithSize(NewSize).Ptr;
return reinterpret_cast<Type>((*Info)->Ptr);
}
/**
-10
View File
@@ -1,10 +0,0 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <FEXCore/fextl/allocator.h>
#include <tsl/robin_set.h>
namespace fextl {
template<class Key, class Hash = std::hash<Key>, class KeyEqual = std::equal_to<Key>, class Allocator = fextl::FEXAlloc<Key>>
using robin_set = tsl::robin_set<Key, Hash, KeyEqual, Allocator>;
}
-66
View File
@@ -1,66 +0,0 @@
// SPDX-License-Identifier: MIT
#include <catch2/catch_test_macros.hpp>
#include <catch2/generators/catch_generators_range.hpp>
#include "Utils/Allocator/HostAllocator.h"
#include <FEXCore/Utils/Allocator.h>
#include <sys/mman.h>
template<typename T>
bool HasSyscallError(T Result) {
constexpr uint64_t MAX_ERRNO = 0xFFFF'FFFF'FFFF'0001ULL;
return reinterpret_cast<uint64_t>(Result) >= MAX_ERRNO;
}
TEST_CASE("Allocator - Fixed replacement") {
const auto RegionSize = 128 * 1024 * 1024;
fextl::vector<FEXCore::Allocator::MemoryRegion> MemoryRegions {};
for (size_t i = 0; i < 2; ++i) {
auto Ptr = mmap(nullptr, RegionSize, PROT_NONE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
MemoryRegions.emplace_back(FEXCore::Allocator::MemoryRegion {
.Ptr = Ptr,
.Size = RegionSize,
});
}
auto Allocator = Alloc::OSAllocator::Create64BitAllocatorWithRegions(MemoryRegions);
auto Base = Allocator->Mmap(nullptr, 4096, PROT_NONE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
REQUIRE(!HasSyscallError(Base));
// Allocate perfectly overlapping pages. Allocate as many pages as the region.
// FEX had a bug where the allocator could run out of memory with MAP_FIXED.
for (size_t i = 0; i < (RegionSize / 4096); ++i) {
auto NewBase = Allocator->Mmap(Base, 4096, PROT_NONE, MAP_FIXED | MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
REQUIRE(Base == NewBase);
}
Alloc::OSAllocator::ReleaseAllocatorWorkaround(std::move(Allocator));
}
TEST_CASE("Allocator - Non-Fit") {
const auto RegionSize = 128 * 1024 * 1024;
fextl::vector<FEXCore::Allocator::MemoryRegion> MemoryRegions {};
for (size_t i = 0; i < 2; ++i) {
auto Ptr = mmap(nullptr, RegionSize, PROT_NONE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
MemoryRegions.emplace_back(FEXCore::Allocator::MemoryRegion {
.Ptr = Ptr,
.Size = RegionSize,
});
}
auto Allocator = Alloc::OSAllocator::Create64BitAllocatorWithRegions(MemoryRegions);
auto Base = Allocator->Mmap(nullptr, RegionSize / 4, PROT_NONE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
REQUIRE(!HasSyscallError(Base));
// Try to allocate within the whole VMA size minus a small amount.
// FEX had a bug where if the allocation fit within a VMA region, it would try and allocate past the end without checking.
// Only occurred when `MAP_FIXED` was used.
auto NewBase = Allocator->Mmap(Base, RegionSize - (4096 * 64), PROT_NONE, MAP_FIXED | MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
// Must either fit in the VMA region, or fail.
// - If it matches previous allocation, then it fit in the VMA region.
// - This can happen if FEX's allocator gains support for VMA merging.
// - If it errors, then it doesn't fit in the VMA region.
REQUIRE((NewBase == Base || HasSyscallError(NewBase)));
Alloc::OSAllocator::ReleaseAllocatorWorkaround(std::move(Allocator));
}
-23
View File
@@ -3,7 +3,6 @@
#include <catch2/generators/catch_generators_range.hpp>
#include "Utils/Allocator/FlexBitSet.h"
#include <sys/mman.h>
TEST_CASE("FlexBitSet - Sizing") {
// Ensure that FlexBitSet sizing is correct.
@@ -41,25 +40,3 @@ TEST_CASE("FlexBitSet - Sizing") {
CHECK(FEXCore::FlexBitSet<uint32_t>::SizeInBits(sizeof(uint32_t) * 8) == sizeof(uint32_t) * 8);
CHECK(FEXCore::FlexBitSet<uint64_t>::SizeInBits(sizeof(uint64_t) * 8) == sizeof(uint64_t) * 8);
}
TEST_CASE("FlexBitSet - Limit") {
// Ensure that the FlexBitSet doesn't read past the limits, and returns correct indexes.
const auto Size = 4096 * 3;
auto Ptr = mmap(nullptr, Size, PROT_NONE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
auto PtrMiddle = reinterpret_cast<void*>(reinterpret_cast<uintptr_t>(Ptr) + 4096);
REQUIRE(mprotect(PtrMiddle, 4096, PROT_READ | PROT_WRITE) != -1);
using ElementType = uint8_t;
const size_t NumElements = 4096 * 8;
auto FlexBit = reinterpret_cast<FEXCore::FlexBitSet<ElementType>*>(PtrMiddle);
for (size_t i = 0; i < NumElements; ++i) {
auto Result = FlexBit->ForwardScanForRange<true>(i, 1, NumElements);
CHECK(Result.FoundElement == i);
}
for (size_t i = 0; i < NumElements; ++i) {
auto Result = FlexBit->BackwardScanForRange<true>(i, 1, 0);
CHECK(Result.FoundElement == i);
}
}
+16 -22
View File
@@ -95,15 +95,14 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: ALU: PC relative") {
CHECK(DisassembleEncoding(1) == 0x10fffffe);
}
{
// Will generate nop + nop + adr.
// Will generate nop + adr.
ForwardLabel Label;
(void)LongAddressGen(Reg::r30, &Label);
(void)Bind(&Label);
dc32(0);
CHECK(DisassembleEncoding(0) == 0xd503201f);
CHECK(DisassembleEncoding(0) == 0xd503201f);
CHECK(DisassembleEncoding(2) == 0x1000003e);
CHECK(DisassembleEncoding(1) == 0x1000003e);
}
{
// Will generate adr.
@@ -116,15 +115,14 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: ALU: PC relative") {
}
{
// Will generate nop + nop + adr.
// Will generate nop + adr.
BiDirectionalLabel Label;
(void)LongAddressGen(Reg::r30, &Label);
(void)Bind(&Label);
dc32(0);
CHECK(DisassembleEncoding(0) == 0xd503201f);
CHECK(DisassembleEncoding(1) == 0xd503201f);
CHECK(DisassembleEncoding(2) == 0x1000003e);
CHECK(DisassembleEncoding(1) == 0x1000003e);
}
{
@@ -145,12 +143,12 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: ALU: PC relative") {
}
{
// Will generate nop + nop + adrp.
// Will generate nop + adrp.
ForwardLabel Label;
(void)LongAddressGen(Reg::r30, &Label);
// Move label 1MB away, plus a page, and then aligned to a page.
for (size_t i = 0; i < ((1 * 1024 * 1024 + 4096) / 4 - 3); ++i) {
for (size_t i = 0; i < ((1 * 1024 * 1024 + 4096) / 4 - 2); ++i) {
nop();
}
@@ -158,12 +156,11 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: ALU: PC relative") {
dc32(0);
CHECK(DisassembleEncoding(0) == 0xd503201f);
CHECK(DisassembleEncoding(1) == 0xd503201f);
CHECK(DisassembleEncoding(2) == 0x9000081e);
CHECK(DisassembleEncoding(1) == 0x9000081e);
}
{
// Will generate nop + adrp + add.
// Will generate adrp + add.
ForwardLabel Label;
(void)LongAddressGen(Reg::r30, &Label);
@@ -175,9 +172,8 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: ALU: PC relative") {
(void)Bind(&Label);
dc32(0);
CHECK(DisassembleEncoding(0) == 0xd503201f);
CHECK(DisassembleEncoding(1) == 0xb000081e);
CHECK(DisassembleEncoding(2) == 0x910013de);
CHECK(DisassembleEncoding(0) == 0xb000081e);
CHECK(DisassembleEncoding(1) == 0x910013de);
}
@@ -199,12 +195,12 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: ALU: PC relative") {
}
{
// Will generate nop + nop + adrp.
// Will generate nop + adrp.
BiDirectionalLabel Label;
(void)LongAddressGen(Reg::r30, &Label);
// Move label 1MB away, plus a page, and then aligned to a page.
for (size_t i = 0; i < ((1 * 1024 * 1024 + 4096) / 4 - 3); ++i) {
for (size_t i = 0; i < ((1 * 1024 * 1024 + 4096) / 4 - 2); ++i) {
nop();
}
@@ -212,12 +208,11 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: ALU: PC relative") {
dc32(0);
CHECK(DisassembleEncoding(0) == 0xd503201f);
CHECK(DisassembleEncoding(1) == 0xd503201f);
CHECK(DisassembleEncoding(2) == 0x9000081e);
CHECK(DisassembleEncoding(1) == 0x9000081e);
}
{
// Will generate nop + adrp + add.
// Will generate adrp + add.
BiDirectionalLabel Label;
(void)LongAddressGen(Reg::r30, &Label);
@@ -229,9 +224,8 @@ TEST_CASE_METHOD(TestDisassembler, "Emitter: ALU: PC relative") {
(void)Bind(&Label);
dc32(0);
CHECK(DisassembleEncoding(0) == 0xd503201f);
CHECK(DisassembleEncoding(1) == 0xb000081e);
CHECK(DisassembleEncoding(2) == 0x910013de);
CHECK(DisassembleEncoding(0) == 0xb000081e);
CHECK(DisassembleEncoding(1) == 0x910013de);
}
}
TEST_CASE_METHOD(TestDisassembler, "Emitter: ALU: Add/subtract immediate") {
+2 -1
View File
@@ -1,5 +1,6 @@
#!/usr/bin/python3
import xxhash
import hashlib
import sys
import os
import shutil
@@ -187,5 +188,5 @@ def main():
return 0
if __name__ == "__main__":
# execute only if run as a script
# execute only if run as a script
sys.exit(main())
+3
View File
@@ -1,5 +1,8 @@
#!/usr/bin/python3
import os
import subprocess
import sys
import tempfile
import platform
def ListContainsRequired(Features, RequiredFeatures):
+2 -2
View File
@@ -4,7 +4,7 @@ from clang.cindex import CursorKind
from clang.cindex import TypeKind
from clang.cindex import TranslationUnit
import sys
from dataclasses import dataclass
from dataclasses import dataclass, field
import subprocess
import logging
logger = logging.getLogger()
@@ -548,5 +548,5 @@ def main():
PrintFunctionDecls()
if __name__ == "__main__":
# execute only if run as a script
# execute only if run as a script
sys.exit(main())
+3 -3
View File
@@ -9,19 +9,19 @@ for fileid in ~/.fex-emu/aotir/*.path; do
else
args="$args --no-abilocalflags"
fi
if [ "${fileid: -7 : 1}" == "T" ]; then
args="$args --tsoenabled"
else
args="$args --no-tsoenabled"
fi
if [ "${fileid: -8 : 1}" == "S" ]; then
args="$args --smc=full"
else
args="$args --smc=mman"
fi
if [ -f "${fileid%.path}.aotir" ]; then
echo "`basename $fileid` has already been generated"
else
+2 -2
View File
@@ -1,5 +1,5 @@
#!/usr/bin/python3
from dataclasses import dataclass
from dataclasses import dataclass, field
import math
import sys
import logging
@@ -282,5 +282,5 @@ def main():
ExportCommonSyscallDefines()
if __name__ == "__main__":
# execute only if run as a script
# execute only if run as a script
sys.exit(main())
+21 -26
View File
@@ -2,6 +2,7 @@
import os
import subprocess
import sys
import tempfile
import re
_Arch = None
@@ -202,26 +203,10 @@ def UpdatePPA():
return DidUpdate
def InstallPackages(PackagesToInstall):
DidInstall = False
try:
CmdResult = subprocess.call(["sudo", "apt-get", "-y", "install"] + PackagesToInstall)
DidInstall = CmdResult == 0
except KeyboardInterrupt:
print ("Keyboard interrupt")
DidInstall = False
pass
if DidInstall:
print("Packages updated")
else:
print("Packages failed to update")
return DidInstall
def CheckAndInstallPackageUpdates(PackagesToInstall, InstallIfNotFound=False):
def CheckAndInstallPackageUpdates():
PackagesToInstall = GetPackagesToInstall()
for Package in PackagesToInstall[:]:
UpgradableStatus = subprocess.check_output(["apt", "list", "--upgradable", Package], stderr=None).decode("utf-8")
UpgradableStatus = subprocess.check_output(["apt", "list", "--upgradable", Package]).decode("utf-8")
Found = False
for Line in UpgradableStatus.split("\n"):
# If the package exists to be upgraded then it will appear in this list
@@ -236,14 +221,28 @@ def CheckAndInstallPackageUpdates(PackagesToInstall, InstallIfNotFound=False):
if Package in Line and "upgradable" in Line:
Found = True
if InstallIfNotFound == False and Found == False:
if Found == False:
PackagesToInstall.remove(Package)
if len(PackagesToInstall) > 0:
print ("Found updates for packages: {}".format(PackagesToInstall))
print ("This bit may ask for your password")
return InstallPackages(PackagesToInstall)
DidInstall = False
try:
CmdResult = subprocess.call(["sudo", "apt-get", "-y", "install"] + PackagesToInstall)
DidInstall = CmdResult == 0
except KeyboardInterrupt:
print ("Keyboard interrupt")
DidInstall = False
pass
if DidInstall:
print("Packages updated")
else:
print("Packages failed to update")
return DidInstall
return True
@@ -355,14 +354,10 @@ def main():
if not UpdatePPA():
print ("apt sources failed to update. Not continuing")
ExitWithStatus(-1)
if not CheckAndInstallPackageUpdates(GetPackagesToInstall()):
if not CheckAndInstallPackageUpdates():
print ("apt packages failed to update. Not continuing")
ExitWithStatus(-1)
else:
if not CheckAndInstallPackageUpdates(["software-properties-common"], True):
print ("software-properties-common package failed to update. Not continuing")
ExitWithStatus(-1)
if not InstallPPA():
print ("PPA failed to install. Not continuing")
ExitWithStatus(-1)
+2 -2
View File
@@ -1,6 +1,6 @@
#!/usr/bin/python3
import base64
from dataclasses import dataclass
from dataclasses import dataclass, field
from enum import Flag
import json
import struct
@@ -254,5 +254,5 @@ def main():
return 0
if __name__ == "__main__":
# execute only if run as a script
# execute only if run as a script
sys.exit(main())
+2 -2
View File
@@ -4,7 +4,7 @@ from clang.cindex import CursorKind
from clang.cindex import TypeKind
from clang.cindex import TranslationUnit
import sys
from dataclasses import dataclass
from dataclasses import dataclass, field
import subprocess
import logging
logger = logging.getLogger()
@@ -774,5 +774,5 @@ def main():
return Result
if __name__ == "__main__":
# execute only if run as a script
# execute only if run as a script
sys.exit(main())
+4
View File
@@ -1,9 +1,13 @@
#!/usr/bin/python3
from enum import Flag
import json
import os
import struct
import sys
import glob
from threading import Thread
import subprocess
import time
import multiprocessing
from shutil import which
+1 -1
View File
@@ -76,6 +76,6 @@ def main():
return 0
if __name__ == "__main__":
# execute only if run as a script
# execute only if run as a script
sys.exit(main())
+1 -2
View File
@@ -1,6 +1,7 @@
#!/usr/bin/python3
import re
import sys
import subprocess
try:
from packaging.version import Version as version_check
except:
@@ -80,8 +81,6 @@ BigCoreIDs = {
[ ["apple-a13", "0.0"], # If we aren't on 12.0+
["apple-a14", "12.0"], # Only exists in 12.0+
],
# QEmu HVF 10.2+
tuple([0x61, 0]): "apple-a13", # Can't determine variant, choose lowest.
}
LittleCoreIDs = {
+1
View File
@@ -1,6 +1,7 @@
#!/bin/env python3
import sys
import fileinput
import re
# Handles the following formats:
+2 -2
View File
@@ -4,8 +4,8 @@ import os
import sys
import subprocess
# Check if FEX indicates support for AVX
def DoesFEXSupportAVX(mode):
# Check if FEX indicates support for AVX
fex_interpreter_path = os.path.dirname(sys.argv[7]) + "/FEX"
args = list()
@@ -22,8 +22,8 @@ def DoesFEXSupportAVX(mode):
return 'avx' in flags and 'avx2' in flags
return False
# Check if the test itself requires AVX
def TestRequiresAVXSupport():
# Check if the test itself requires AVX
exe_path = sys.argv[len(sys.argv) - 1]
json_path = os.path.dirname(os.path.dirname(exe_path)) + '/requirements/' + os.path.basename(exe_path) + '.json'
+3
View File
@@ -1,3 +1,6 @@
from enum import Flag
import json
import struct
import sys
from json_config_parse import parse_json
+3
View File
@@ -1,3 +1,6 @@
from enum import Flag
import json
import struct
import sys
from json_config_parse import parse_json
+3
View File
@@ -3,6 +3,9 @@
# Save current directory
DIR=$(pwd)
# Get the absolute path to the Scripts directory
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
# Parse arguments
CHANGED_ONLY=false
TARGET_DIR=""
+2 -1
View File
@@ -1,6 +1,7 @@
#!/usr/bin/python3
import sys
import subprocess
import os.path
from os import path
from shutil import which
@@ -76,4 +77,4 @@ if (is_known_failure):
sys.exit(1)
else:
# Just return the result code if we don't have this test as a known failure
sys.exit(ResultCode)
sys.exit(ResultCode);
-3
View File
@@ -14,12 +14,10 @@
#include <span>
#include <utility>
#include <vector>
#include <unistd.h>
#include <FEXCore/fextl/functional.h>
#include <FEXCore/fextl/map.h>
#include <FEXCore/fextl/vector.h>
#include <FEXCore/Utils/LogManager.h>
namespace fasio {
@@ -370,7 +368,6 @@ std::size_t read(AsyncReadStream& Stream, mutable_buffer Buffers, error& ec) {
auto BytesRead = Stream.read_some(Buffers, ec);
TotalBytesRead += BytesRead;
if (Buffers.FD) {
LOGMAN_THROW_A_FMT(**Buffers.FD != -1, "Receiver requested a file descriptor but none was sent");
(void)Buffers.consume_fd();
}
Buffers += BytesRead;
+2 -4
View File
@@ -160,10 +160,8 @@ private:
struct cmsghdr* cmsg = CMSG_FIRSTHDR(&msg);
if (Buffers.FD &&
(cmsg == nullptr || cmsg->cmsg_len != CMSG_LEN(sizeof(int)) || cmsg->cmsg_level != SOL_SOCKET || cmsg->cmsg_type != SCM_RIGHTS)) {
// Not a failure since some data was read for the main message
**Buffers.FD = -1;
ec = error::success;
return BytesRead;
ec = error::invalid;
return 0;
}
if (Buffers.FD) {
+2 -34
View File
@@ -79,8 +79,8 @@ static char* SaveLayerToJSON(char* JsonBuffer, const FEXCore::Config::Layer* Lay
}
if (std::holds_alternative<fextl::string>(it.second)) {
JsonBuffer = json_str(JsonBuffer, Name.data(), std::get<fextl::string>(it.second).c_str());
} else if (std::holds_alternative<FEXCore::Config::StringArrayType>(it.second)) {
for (auto& var : std::get<FEXCore::Config::StringArrayType>(it.second)) {
} else if (std::holds_alternative<FEXCore::Config::DefaultValues::Type::StringArrayType>(it.second)) {
for (auto& var : std::get<FEXCore::Config::DefaultValues::Type::StringArrayType>(it.second)) {
JsonBuffer = json_str(JsonBuffer, Name.data(), var.c_str());
}
} else {
@@ -542,13 +542,6 @@ const char* GetHomeDirectory() {
#endif
fextl::string GetDataDirectory(bool Global, const PortableInformation& PortableInfo) {
#ifdef FEX_STEAM_SUPPORT
const char* SteamDataPath = getenv("STEAM_COMPAT_DATA_PATH");
if (SteamDataPath) {
return fextl::fmt::format("{}/fex-emu/", SteamDataPath);
}
#endif
const char* DataOverride = getenv("FEX_APP_DATA_LOCATION");
if (PortableInfo.IsPortable && (Global || !DataOverride)) {
@@ -573,13 +566,6 @@ fextl::string GetDataDirectory(bool Global, const PortableInformation& PortableI
}
fextl::string GetConfigDirectory(bool Global, const PortableInformation& PortableInfo) {
#ifdef FEX_STEAM_SUPPORT
const char* SteamDataPath = getenv("STEAM_COMPAT_DATA_PATH");
if (SteamDataPath) {
return fextl::fmt::format("{}/fex-emu/", SteamDataPath);
}
#endif
const char* ConfigOverride = getenv("FEX_APP_CONFIG_LOCATION");
if (PortableInfo.IsPortable && (Global || !ConfigOverride)) {
return fextl::fmt::format("{}/fex-emu/", PortableInfo.InterpreterPath);
@@ -616,24 +602,6 @@ fextl::string GetConfigDirectory(bool Global, const PortableInformation& Portabl
return ConfigDir;
}
fextl::string GetCacheDirectory() {
const char* CacheOverride = getenv("FEX_APP_CACHE_LOCATION");
if (CacheOverride) {
return CacheOverride;
}
#ifdef FEX_STEAM_SUPPORT
const char* SteamDataPath = getenv("STEAM_COMPAT_SHADER_PATH");
if (SteamDataPath) {
return fextl::fmt::format("{}/fex-emu/", SteamDataPath);
}
#endif
const char* HomeDir = GetHomeDirectory();
const char* CacheXDG = getenv("XDG_CACHE_HOME");
return (CacheXDG ? fextl::string {CacheXDG} : (fextl::string {HomeDir} + "/.cache")) + "/fex-emu/";
}
fextl::string GetConfigFileLocation(bool Global, const PortableInformation& PortableInfo) {
return GetConfigDirectory(Global, PortableInfo) + "Config.json";
}
-1
View File
@@ -62,7 +62,6 @@ const char* GetHomeDirectory();
fextl::string GetDataDirectory(const PortableInformation& PortableInfo);
fextl::string GetConfigDirectory(bool Global, const PortableInformation& PortableInfo);
fextl::string GetConfigFileLocation(bool Global, const PortableInformation& PortableInfo);
fextl::string GetCacheDirectory();
void InitializeConfigs(const PortableInformation& PortableInfo);
+102 -164
View File
@@ -1,7 +1,6 @@
// SPDX-License-Identifier: MIT
#include "Common/AsyncNet.h"
#include "Common/Config.h"
#include "FDUtils.h"
#include "Common/FEXServerClient.h"
#include <FEXCore/Utils/CompilerDefs.h>
@@ -60,8 +59,8 @@ int RequestPIDFDPacket(int ServerSocket, PacketType Type) {
fasio::mutable_buffer ResBuffer {std::as_writable_bytes(std::span {&Res, 1})};
int NewFD = -1;
ResBuffer.FD = &NewFD;
auto BytesRead = Socket.read_some(ResBuffer, ec);
if (ec != fasio::error::success || BytesRead != sizeof(Res) || Res.Header.Type != PacketType::TYPE_SUCCESS) {
read(Socket, ResBuffer, ec);
if (ec != fasio::error::success || Res.Header.Type != PacketType::TYPE_SUCCESS) {
return -1;
}
@@ -138,21 +137,15 @@ fextl::string GetServerSocketName() {
}
fextl::string GetServerSocketPath() {
fextl::string name {};
#ifndef FEX_STEAM_SUPPORT
FEX_CONFIG_OPT(ServerSocketPath, SERVERSOCKETPATH);
name = ServerSocketPath();
auto name = ServerSocketPath();
if (name.starts_with("/")) {
return name;
}
auto Folder = GetTempFolder();
#else
// Under Steam the FEXServer's socket is a game-specific directory.
auto Folder = GetServerLockFolder();
#endif
if (name.empty()) {
return fextl::fmt::format("{}/{}.FEXServer.Socket", Folder, ::getuid());
@@ -166,31 +159,25 @@ int GetServerFD() {
}
int ConnectToServer(ConnectionOption ConnectionOption) {
int SocketFD {-1};
size_t SizeOfAddr {};
struct sockaddr_un addr {};
size_t SizeOfSocketString {};
auto ServerSocketName = GetServerSocketName();
// Create the initial unix socket
SocketFD = socket(AF_UNIX, SOCK_STREAM | SOCK_CLOEXEC, 0);
int SocketFD = socket(AF_UNIX, SOCK_STREAM | SOCK_CLOEXEC, 0);
if (SocketFD == -1) {
LogMan::Msg::EFmt("Couldn't open AF_UNIX socket {}", errno);
return -1;
}
// Steam doesn't get to connect to global sockets.
#ifndef FEX_STEAM_SUPPORT
auto ServerSocketName = GetServerSocketName();
// AF_UNIX has a special feature for named socket paths.
// If the name of the socket begins with `\0` then it is an "abstract" socket address.
// The entirety of the name is used as a path to a socket that doesn't have any filesystem backing.
struct sockaddr_un addr {};
addr.sun_family = AF_UNIX;
SizeOfSocketString = std::min(ServerSocketName.size() + 1, sizeof(addr.sun_path) - 1);
size_t SizeOfSocketString = std::min(ServerSocketName.size() + 1, sizeof(addr.sun_path) - 1);
addr.sun_path[0] = 0; // Abstract AF_UNIX sockets start with \0
strncpy(addr.sun_path + 1, ServerSocketName.data(), SizeOfSocketString);
// Include final null character.
SizeOfAddr = sizeof(addr.sun_family) + SizeOfSocketString;
size_t SizeOfAddr = sizeof(addr.sun_family) + SizeOfSocketString;
if (connect(SocketFD, reinterpret_cast<struct sockaddr*>(&addr), SizeOfAddr) == -1) {
if (ConnectionOption == ConnectionOption::Default || errno != ECONNREFUSED) {
@@ -199,13 +186,11 @@ int ConnectToServer(ConnectionOption ConnectionOption) {
} else {
return SocketFD;
}
#endif
// Try again with a path-based socket, since abstract sockets will fail if we have been
// placed in a new netns as part of a sandbox.
auto ServerSocketPath = GetServerSocketPath();
addr.sun_family = AF_UNIX;
SizeOfSocketString = std::min(ServerSocketPath.size(), sizeof(addr.sun_path) - 1);
strncpy(addr.sun_path, ServerSocketPath.data(), SizeOfSocketString);
SizeOfAddr = sizeof(addr.sun_family) + SizeOfSocketString;
@@ -239,125 +224,110 @@ bool SetupClient(std::string_view InterpreterPath) {
return true;
}
int StartServer(std::string_view InterpreterPath, int watch_fd) {
int LocalServerFD {-1};
// Couldn't connect to the server. Start one
int ConnectToAndStartServer(std::string_view InterpreterPath) {
int ServerFD = ConnectToServer(ConnectionOption::NoPrintConnectionError);
if (ServerFD == -1) {
// Couldn't connect to the server. Start one
// Open some pipes for letting us know when the server is ready
int fds[2] {};
if (pipe2(fds, 0) != 0) {
LogMan::Msg::EFmt("Couldn't open pipe");
return -1;
}
// Extract directory from InterpreterPath
fextl::string InterpreterDir {InterpreterPath};
size_t LastSlash = InterpreterDir.rfind('/');
if (LastSlash != fextl::string::npos) {
InterpreterDir = InterpreterDir.substr(0, LastSlash);
}
fextl::string FEXServerPath = fextl::fmt::format("{}/FEXServer", InterpreterDir);
// Check if a local FEXServer next to FEX exists
// If it does then it takes priority over the installed one
if (!FHU::Filesystem::Exists(FEXServerPath)) {
FEXServerPath = "FEXServer";
}
// Set-up our SIGCHLD handler to ignore the signal.
// This is early in the initialization stage so no handlers have been installed.
//
// We want to ignore the signal so that if FEXServer starts in daemon mode, it
// doesn't leave a zombie process around waiting for something to get the result.
struct sigaction action {};
action.sa_handler = SIG_IGN;
sigaction(SIGCHLD, &action, &action);
pid_t pid = fork();
if (pid == 0) {
// Child
close(fds[0]); // Close read end of pipe
const char* argv[6];
auto pipe_string = fextl::fmt::format("{}", fds[1]);
auto watch_fd_string = fextl::fmt::format("{}", watch_fd);
size_t arg_count {};
argv[arg_count++] = FEXServerPath.c_str();
argv[arg_count++] = "--wait_pipe";
argv[arg_count++] = pipe_string.c_str();
if (watch_fd != -1) {
argv[arg_count++] = "--watch_fd";
argv[arg_count++] = watch_fd_string.c_str();
}
argv[arg_count++] = nullptr;
if (execvp(argv[0], (char* const*)argv) == -1) {
// Let the parent know that we couldn't execute for some reason
uint64_t error {1};
write(fds[1], &error, sizeof(error));
// Give a hopefully helpful error message for users
LogMan::Msg::EFmt("Couldn't execute: {}", argv[0]);
LogMan::Msg::EFmt("This means the squashFS rootfs won't be mounted.");
LogMan::Msg::EFmt("Expect errors!");
// Destroy this fork
exit(1);
}
FEX_UNREACHABLE;
} else {
// Parent
// Wait for the child to exit so we can check if it is mounted or not
close(fds[1]); // Close write end of the pipe
// Wait for a message from FEXServer
pollfd PollFD;
PollFD.fd = fds[0];
PollFD.events = POLLIN | POLLOUT | POLLRDHUP | POLLERR | POLLHUP | POLLNVAL;
// Wait for a result on the pipe that isn't EINTR
while (poll(&PollFD, 1, -1) == -1 && errno == EINTR)
;
// Check if child signaled an error
uint64_t error = 0;
ssize_t bytes_read = read(fds[0], &error, sizeof(error));
close(fds[0]);
if (bytes_read > 0 && error != 0) {
// Open some pipes for letting us know when the server is ready
int fds[2] {};
if (pipe2(fds, 0) != 0) {
LogMan::Msg::EFmt("Couldn't open pipe");
return -1;
}
for (size_t i = 0; i < 5; ++i) {
LocalServerFD = ConnectToServer(ConnectionOption::Default);
// Extract directory from InterpreterPath
fextl::string InterpreterDir {InterpreterPath};
size_t LastSlash = InterpreterDir.rfind('/');
if (LastSlash != fextl::string::npos) {
InterpreterDir = InterpreterDir.substr(0, LastSlash);
}
if (LocalServerFD != -1) {
break;
fextl::string FEXServerPath = fextl::fmt::format("{}/FEXServer", InterpreterDir);
// Check if a local FEXServer next to FEX exists
// If it does then it takes priority over the installed one
if (!FHU::Filesystem::Exists(FEXServerPath)) {
FEXServerPath = "FEXServer";
}
// Set-up our SIGCHLD handler to ignore the signal.
// This is early in the initialization stage so no handlers have been installed.
//
// We want to ignore the signal so that if FEXServer starts in daemon mode, it
// doesn't leave a zombie process around waiting for something to get the result.
struct sigaction action {};
action.sa_handler = SIG_IGN;
sigaction(SIGCHLD, &action, &action);
pid_t pid = fork();
if (pid == 0) {
// Child
close(fds[0]); // Close read end of pipe
const char* argv[4];
auto pipe_string = fextl::fmt::format("{}", fds[1]);
argv[0] = FEXServerPath.c_str();
argv[1] = "--wait_pipe";
argv[2] = pipe_string.c_str();
argv[3] = nullptr;
if (execvp(argv[0], (char* const*)argv) == -1) {
// Let the parent know that we couldn't execute for some reason
uint64_t error {1};
write(fds[1], &error, sizeof(error));
// Give a hopefully helpful error message for users
LogMan::Msg::EFmt("Couldn't execute: {}", argv[0]);
LogMan::Msg::EFmt("This means the squashFS rootfs won't be mounted.");
LogMan::Msg::EFmt("Expect errors!");
// Destroy this fork
exit(1);
}
std::this_thread::sleep_for(std::chrono::seconds(1));
FEX_UNREACHABLE;
} else {
// Parent
// Wait for the child to exit so we can check if it is mounted or not
close(fds[1]); // Close write end of the pipe
// Wait for a message from FEXServer
pollfd PollFD;
PollFD.fd = fds[0];
PollFD.events = POLLIN | POLLOUT | POLLRDHUP | POLLERR | POLLHUP | POLLNVAL;
// Wait for a result on the pipe that isn't EINTR
while (poll(&PollFD, 1, -1) == -1 && errno == EINTR)
;
// Check if child signaled an error
uint64_t error = 0;
ssize_t bytes_read = read(fds[0], &error, sizeof(error));
close(fds[0]);
if (bytes_read > 0 && error != 0) {
return -1;
}
for (size_t i = 0; i < 5; ++i) {
ServerFD = ConnectToServer(ConnectionOption::Default);
if (ServerFD != -1) {
break;
}
std::this_thread::sleep_for(std::chrono::seconds(1));
}
if (ServerFD == -1) {
// Still couldn't connect to the socket.
LogMan::Msg::EFmt("Couldn't connect to FEXServer socket {} after launching the process", GetServerSocketName());
}
}
if (LocalServerFD == -1) {
// Still couldn't connect to the socket.
LogMan::Msg::EFmt("Couldn't connect to FEXServer socket after launching the process");
}
// Restore the original SIGCHLD handler if it existed.
sigaction(SIGCHLD, &action, nullptr);
}
// Restore the original SIGCHLD handler if it existed.
sigaction(SIGCHLD, &action, nullptr);
return LocalServerFD;
}
int ConnectToAndStartServer(std::string_view InterpreterPath) {
int LocalServerFD = ConnectToServer(ConnectionOption::NoPrintConnectionError);
if (LocalServerFD == -1) {
LocalServerFD = StartServer(InterpreterPath);
}
return LocalServerFD;
return ServerFD;
}
/**
@@ -405,38 +375,6 @@ int RequestPIDFD(int ServerSocket) {
return RequestPIDFDPacket(ServerSocket, PacketType::TYPE_GET_PID_FD);
}
int RequestCodeMapFD(int ServerSocket, int ProgramFD, bool HasMultiblock) {
fasio::tcp_socket Socket {ServerSocket};
FEXServerRequestPacket Req {
.Header {
.Type = HasMultiblock ? PacketType::TYPE_QUERY_CODE_MAP : PacketType::TYPE_QUERY_CODE_MAP_NO_MULTIBLOCK,
},
};
// Send request
fasio::error ec;
{
fasio::mutable_buffer WriteBuffer {std::as_writable_bytes(std::span {&Req, 1})};
WriteBuffer.FD = &ProgramFD;
write(Socket, WriteBuffer, ec);
if (ec != fasio::error::success) {
return -1;
}
}
// Wait for success response and log FD
FEXServerResultPacket Res {};
fasio::mutable_buffer ResBuffer {std::as_writable_bytes(std::span {&Res, 1})};
int NewFD = -1;
ResBuffer.FD = &NewFD;
read(Socket, ResBuffer, ec);
if (ec != fasio::error::success || Res.Header.Type != PacketType::TYPE_SUCCESS) {
return -1;
}
return NewFD;
}
/** @} */
/**
-19
View File
@@ -19,8 +19,6 @@ enum class PacketType {
TYPE_GET_LOG_FD,
TYPE_GET_ROOTFS_PATH,
TYPE_GET_PID_FD,
TYPE_QUERY_CODE_MAP,
TYPE_QUERY_CODE_MAP_NO_MULTIBLOCK,
// Result only
TYPE_SUCCESS,
@@ -67,13 +65,6 @@ int GetServerFD();
bool SetupClient(std::string_view InterpreterPath);
/**
* @brief Start a FEXServer instance if possible
*
* @return socket FD for communicating with server
*/
int StartServer(std::string_view InterpreterPath, int watch_fd = -1);
/**
* @brief Connect to and start a FEXServer instance if required
*
@@ -122,16 +113,6 @@ fextl::string RequestRootFSPath(int ServerSocket);
*/
int RequestPIDFD(int ServerSocket);
/**
* @brief Request FEXServer to create a new code map for disk cache population
*
* @param ServerSocket - Socket to the server
* @param ProgramFD - FD for program binary
*
* @return FD to write code map to
*/
int RequestCodeMapFD(int ServerSocket, int ProgramFD, bool HasMultiblock);
/** @} */
/**
+5 -9
View File
@@ -2,9 +2,9 @@
#pragma once
#include <FEXCore/Utils/TypeDefines.h>
#include <FEXCore/fextl/vector.h>
#include <cstdint>
#include <optional>
#include <span>
#include <elf.h>
@@ -15,16 +15,12 @@ namespace FEXCore {
* Infers the base virtual address from a file mapping (as described by parameters to a single
* call to mmap()).
*
* Usually the base address can uniquely be inferred, but in edge cases multiple possible
* candidates are returned.
*
* The file offset of any given mapping need not match its virtual address offset from the base
* mapping (file offset = 0). Instead, this function searches the corresponding ELF program headers
* for an entry that generated the given file mapping.
*/
inline fextl::vector<uint64_t>
inline std::optional<uint64_t>
InferMappingBaseAddress(std::span<const Elf64_Phdr> ProgramHeaders, uint64_t Addr, uint64_t Size, uint64_t FileOffset, int AccessFlags) {
fextl::vector<uint64_t> Ret;
for (auto& phdr : ProgramHeaders) {
if (phdr.p_type != PT_LOAD) {
// Skip headers that don't trigger memory mappings
@@ -40,11 +36,11 @@ InferMappingBaseAddress(std::span<const Elf64_Phdr> ProgramHeaders, uint64_t Add
if (FileOffset >= SegmentStartOffset && FileOffset < SegmentStartOffset + phdr.p_filesz &&
(FileOffset & Utils::FEX_PAGE_MASK) == (phdr.p_offset & Utils::FEX_PAGE_MASK)) {
// Compute VA offset relative to the base mapping
Ret.push_back(Addr - (phdr.p_vaddr - (phdr.p_offset & 0xfff)) + (ProgramHeaders[0].p_vaddr - (ProgramHeaders[0].p_offset & 0xfff)) -
(FileOffset - SegmentStartOffset));
return Addr - (phdr.p_vaddr - (phdr.p_offset & 0xfff)) + (ProgramHeaders[0].p_vaddr - (ProgramHeaders[0].p_offset & 0xfff)) -
(FileOffset - SegmentStartOffset);
}
}
return Ret;
return std::nullopt;
}
} // namespace FEXCore
+68 -192
View File
@@ -7,9 +7,6 @@
#include <FEXCore/Utils/FileLoading.h>
#include <FEXCore/Utils/StringUtils.h>
#include <range/v3/view/split.hpp>
#include <range/v3/view/transform.hpp>
#ifdef _M_X86_64
#include "Common/X86Features.h"
#endif
@@ -38,22 +35,6 @@ void FillMIDRInformationViaLinux(FEXCore::HostFeatures* Features) {
#endif
}
#if defined(_M_ARM_64) && !defined(VIXL_SIMULATOR)
__attribute__((naked)) static uint64_t ReadSVEVectorLengthInBits() {
///< Can't use rdvl instruction directly because compilers will complain that sve/sme is required.
__asm(R"(
.word 0x04bf5100 // rdvl x0, #8
ret;
)");
}
#else
[[maybe_unused]]
static int ReadSVEVectorLengthInBits() {
// Return unsupported
return 0;
}
#endif
#ifdef _M_ARM_64
#define GetSysReg(name, reg) \
static uint64_t Get_##name() { \
@@ -72,7 +53,6 @@ GetSysReg(MMFR2_EL1, ID_AA64MMFR2_EL1);
GetSysReg(ZFR0_EL1, s3_0_c0_c4_4); // Can't request by name
GetSysReg(MMFR1_EL1, ID_AA64MMFR1_EL1);
GetSysReg(ISAR2_EL1, ID_AA64ISAR2_EL1);
GetSysReg(DCZID_EL0, DCZID_EL0);
class CPUFeaturesFromID final : public FEX::CPUFeatures {
public:
@@ -86,17 +66,12 @@ public:
MMFR2.SetReg(Get_MMFR2_EL1());
MMFR1.SetReg(Get_MMFR1_EL1());
ISAR2.SetReg(Get_ISAR2_EL1());
DCZID.SetReg(Get_DCZID_EL0());
if (PFR0.SupportsSVE()) {
// Can only query if SVE is supported.
ZFR0.SetReg(Get_ZFR0_EL1());
}
FillFeatureFlags();
if (Supports(CPUFeatures::Feature::SVE2)) {
SVEVL.SetReg(ReadSVEVectorLengthInBits());
}
}
};
@@ -105,78 +80,6 @@ FEX::CPUFeatures GetCPUFeaturesFromIDRegisters() {
}
#endif
class CPUFeaturesFromConfig final : public FEX::CPUFeatures {
public:
CPUFeaturesFromConfig(std::string_view Config) {
auto to_string_view = [](auto rng) {
return std::string_view(&*rng.begin(), ranges::distance(rng));
};
for (auto Option : ranges::views::split(Config, ',') | ranges::views::transform(to_string_view)) {
auto OptionData = ranges::views::split(Option, '=') | ranges::views::transform(to_string_view);
auto OptionDataBegin = ranges::begin(OptionData);
auto OptionDataEnd = ranges::end(OptionData);
if (OptionDataBegin == OptionDataEnd) {
continue;
}
auto Key = *OptionDataBegin;
if (Key.empty()) {
continue;
}
++OptionDataBegin;
if (OptionDataBegin == OptionDataEnd) {
continue;
}
auto Value = *OptionDataBegin;
uint64_t ValueHex {};
char* str_end {};
ValueHex = std::strtoull(Value.data(), &str_end, 16);
if (str_end == Value.data()) {
LogMan::Msg::EFmt("Couldn't parse '{}={}'\n", Key, Value);
continue;
}
if (Key == "isar0") {
ISAR0.SetReg(ValueHex);
} else if (Key == "isar1") {
ISAR1.SetReg(ValueHex);
} else if (Key == "isar2") {
ISAR2.SetReg(ValueHex);
} else if (Key == "pfr0") {
PFR0.SetReg(ValueHex);
} else if (Key == "pfr1") {
PFR1.SetReg(ValueHex);
} else if (Key == "midr") {
MIDR.SetReg(ValueHex);
} else if (Key == "mmfr0") {
MMFR0.SetReg(ValueHex);
} else if (Key == "mmfr1") {
MMFR1.SetReg(ValueHex);
} else if (Key == "mmfr2") {
MMFR2.SetReg(ValueHex);
} else if (Key == "zfr0") {
ZFR0.SetReg(ValueHex);
} else if (Key == "dczid") {
DCZID.SetReg(ValueHex);
} else if (Key == "svevl") {
SVEVL.SetReg(ValueHex);
} else {
LogMan::Msg::EFmt("Unknown Key: {}", Key);
}
}
FillFeatureFlags();
}
};
FEX::CPUFeatures GetCPUFeaturesFromConfig(std::string_view Config) {
return CPUFeaturesFromConfig {Config};
}
class CPUFeaturesAll final : public FEX::CPUFeatures {
public:
CPUFeaturesAll() {
@@ -184,9 +87,6 @@ public:
for (uint32_t i = 0; i < FEXCore::ToUnderlying(FEX::CPUFeatures::Feature::MAX); ++i) {
SetFeature(FEX::CPUFeatures::Feature {i});
}
// Report unsupported for DCZVA
DCZID.SetReg(0b1'0000);
}
};
@@ -446,7 +346,21 @@ void FEX::CPUFeatures::FillFeatureFlags() {
}
}
// Data Zero Prohibited flag
// 0b0 = ZVA/GVA/GZVA permitted
// 0b1 = ZVA/GVA/GZVA prohibited
[[maybe_unused]] constexpr uint32_t DCZID_DZP_MASK = 0b1'0000;
// Log2 of the blocksize in 32-bit words
[[maybe_unused]] constexpr uint32_t DCZID_BS_MASK = 0b0'1111;
#ifdef _M_ARM_64
[[maybe_unused]]
static uint32_t GetDCZID() {
uint64_t Result {};
__asm("mrs %[Res], DCZID_EL0" : [Res] "=r"(Result));
return Result;
}
static uint32_t GetFPCR() {
uint64_t Result {};
__asm("mrs %[Res], FPCR" : [Res] "=r"(Result));
@@ -457,6 +371,27 @@ static void SetFPCR(uint64_t Value) {
__asm("msr FPCR, %[Value]" ::[Value] "r"(Value));
}
#ifndef VIXL_SIMULATOR
__attribute__((naked)) static uint64_t ReadSVEVectorLengthInBits() {
///< Can't use rdvl instruction directly because compilers will complain that sve/sme is required.
__asm(R"(
.word 0x04bf5100 // rdvl x0, #8
ret;
)");
}
#endif
#else
[[maybe_unused]]
static uint32_t GetDCZID() {
// Return unsupported
return DCZID_DZP_MASK;
}
[[maybe_unused]]
static int ReadSVEVectorLengthInBits() {
// Return unsupported
return 0;
}
#endif
static void OverrideFeatures(FEXCore::HostFeatures* Features, uint64_t ForceSVEWidth) {
@@ -525,76 +460,9 @@ static void OverrideFeatures(FEXCore::HostFeatures* Features, uint64_t ForceSVEW
Features->SupportsSVE256 = ForceSVEWidth && ForceSVEWidth >= 256;
}
static void HandleErrata(FEXCore::HostFeatures* HostFeatures, uint64_t MIDR) {
constexpr uint32_t Implementer_ARM = 0x41;
constexpr uint32_t PartNum_V2 = 0xd4f;
constexpr uint32_t PartNum_V3 = 0xd84;
constexpr uint32_t PartNum_V3AE = 0xd83;
constexpr uint32_t PartNum_X3 = 0xd4e;
constexpr uint32_t PartNum_X4 = 0xd82;
constexpr uint32_t PartNum_X925 = 0xd85;
constexpr uint32_t PartNum_C1Ultra = 0xd8c;
constexpr uint32_t PartNum_C1Premium = 0xd90;
FEXCore::HostFeatures FetchHostFeatures(FEX::CPUFeatures& Features, bool SupportsCacheMaintenanceOps, uint64_t CTR, uint64_t MIDR) {
FEXCore::HostFeatures HostFeatures;
constexpr uint32_t Implementer_QCOM = 0x51;
constexpr uint32_t PartNum_Oryon1 = 0x001;
auto GetMIDRImplementer = [](uint32_t MIDR) -> uint32_t {
return (MIDR >> 24) & 0xFF;
};
auto GetMIDRPartNum = [](uint32_t MIDR) -> uint32_t {
return (MIDR >> 4) & 0xFFF;
};
const uint32_t MIDR_Implementer = GetMIDRImplementer(MIDR);
const uint32_t MIDR_PartNum = GetMIDRPartNum(MIDR);
#ifdef _M_ARM_64
if (MIDR_Implementer == Implementer_QCOM && MIDR_PartNum == PartNum_Oryon1) {
// Work around an errata in Qualcomm's Oryon.
// While this CPU implements the RAND extension:
// - The RNDR register works.
// - The RNDRRS register will never read a random number. (Always return failure)
// This is contrary to x86 RNG behaviour where it allows spurious failure with RDSEED, but guarantees eventual success.
// This manifested itself on Linux when an x86 processor failed to guarantee forward progress and boot of services would infinite
// loop. Just disable this extension if this CPU is detected.
HostFeatures->SupportsRAND = false;
}
#endif
// The LDAPUR instruction suffers from significant performance issues on many ARM implementations. This is
// listed in the official Cortex errata list as follows:
//
// 3877900
// LDAPUR, LDAPURB, LDAPURH instructions have stricter memory ordering than required
//
// LDAPUR instructions execute with full Load-Acquire ordering instead of the relaxed ordering described
// in the LDAPUR pseudocode. This might cause significant performance degradation in workloads that do
// not require this stricter memory ordering. Note that this erratum only affects the unscaled versions of
// LDAPUR (LDAPUR, LDAPURB, LDAPURH), and not LDAPR (LDAPR, LDAPRB, LDAPRH).
//
// The list of cores to disable its use on was taken from the following LLVM PR that accomplishes the same
// thing: https://github.com/llvm/llvm-project/pull/124274
for (uint32_t CoreIndex = 0; CoreIndex < HostFeatures->CPUMIDRs.size(); CoreIndex++) {
const uint32_t CoreMIDR = HostFeatures->CPUMIDRs[CoreIndex];
const uint32_t Core_MIDR_Implementer = GetMIDRImplementer(CoreMIDR);
const uint32_t Core_MIDR_PartNum = GetMIDRPartNum(CoreMIDR);
bool IgnoreLRCPC2 = (Core_MIDR_Implementer == Implementer_ARM) &&
((Core_MIDR_PartNum == PartNum_V2) || (Core_MIDR_PartNum == PartNum_V3) || (Core_MIDR_PartNum == PartNum_X3) ||
(Core_MIDR_PartNum == PartNum_X4) || (Core_MIDR_PartNum == PartNum_X925) || (Core_MIDR_PartNum == PartNum_V3AE) ||
(Core_MIDR_PartNum == PartNum_C1Ultra) || (Core_MIDR_PartNum == PartNum_C1Premium));
if (IgnoreLRCPC2) {
HostFeatures->SupportsTSOImm9 = false;
break;
}
}
}
void FetchHostFeatures(FEX::CPUFeatures& Features, FEXCore::HostFeatures& HostFeatures, bool SupportsCacheMaintenanceOps, uint64_t CTR,
uint64_t MIDR) {
FEX_CONFIG_OPT(ForceSVEWidth, FORCESVEWIDTH);
FEX_CONFIG_OPT(Is64BitMode, IS64BIT_MODE);
@@ -627,7 +495,7 @@ void FetchHostFeatures(FEX::CPUFeatures& Features, FEXCore::HostFeatures& HostFe
HostFeatures.SupportsSVE256 = ForceSVEWidth() ? ForceSVEWidth() >= 256 : true;
#else
HostFeatures.SupportsSVE128 = Features.Supports(CPUFeatures::Feature::SVE2);
HostFeatures.SupportsSVE256 = Features.Supports(CPUFeatures::Feature::SVE2) && Features.GetSVEVectorLengthInBits() >= 256;
HostFeatures.SupportsSVE256 = Features.Supports(CPUFeatures::Feature::SVE2) && ReadSVEVectorLengthInBits() >= 256;
#endif
HostFeatures.SupportsAVX = true;
@@ -664,6 +532,23 @@ void FetchHostFeatures(FEX::CPUFeatures& Features, FEXCore::HostFeatures& HostFe
// Set FPCR back to original just in case anything changed
SetFPCR(OriginalFPCR);
if (HostFeatures.SupportsRAND) {
constexpr uint32_t Implementer_QCOM = 0x51;
constexpr uint32_t PartNum_Oryon1 = 0x001;
const uint32_t MIDR_Implementer = (MIDR >> 24) & 0xFF;
const uint32_t MIDR_PartNum = (MIDR >> 4) & 0xFFF;
if (MIDR_Implementer == Implementer_QCOM && MIDR_PartNum == PartNum_Oryon1) {
// Work around an errata in Qualcomm's Oryon.
// While this CPU implements the RAND extension:
// - The RNDR register works.
// - The RNDRRS register will never read a random number. (Always return failure)
// This is contrary to x86 RNG behaviour where it allows spurious failure with RDSEED, but guarantees eventual success.
// This manifested itself on Linux when an x86 processor failed to guarantee forward progress and boot of services would infinite
// loop. Just disable this extension if this CPU is detected.
HostFeatures.SupportsRAND = false;
}
}
#endif
#ifdef VIXL_SIMULATOR
@@ -675,17 +560,16 @@ void FetchHostFeatures(FEX::CPUFeatures& Features, FEXCore::HostFeatures& HostFe
HostFeatures.SupportsSHA = true;
HostFeatures.SupportsPMULL_128Bit = true;
HostFeatures.SupportsAES256 = true;
// Simulator doesn't support these
HostFeatures.SupportsRPRES = false;
HostFeatures.SupportsAFP = false;
#else
// Check if we can support cacheline clears
if (Features.GetDCZID().SupportsDCZVA()) {
uint32_t DCZID = GetDCZID();
if ((DCZID & DCZID_DZP_MASK) == 0) {
uint32_t DCZID_Log2 = DCZID & DCZID_BS_MASK;
uint32_t DCZID_Bytes = (1 << DCZID_Log2) * sizeof(uint32_t);
// If the DC ZVA size matches the emulated cache line size
// This means we can use the instruction
constexpr static uint64_t CACHELINE_SIZE = 64;
HostFeatures.SupportsCLZERO = Features.GetDCZID().BlockSizeInBytes() == CACHELINE_SIZE;
HostFeatures.SupportsCLZERO = DCZID_Bytes == CACHELINE_SIZE;
}
#endif
@@ -714,28 +598,21 @@ void FetchHostFeatures(FEX::CPUFeatures& Features, FEXCore::HostFeatures& HostFe
#endif
HostFeatures.SupportsPreserveAllABI = FEX_HAS_PRESERVE_ALL_ATTR;
HandleErrata(&HostFeatures, MIDR);
OverrideFeatures(&HostFeatures, ForceSVEWidth());
return HostFeatures;
}
FEXCore::HostFeatures FetchHostFeatures() {
FEX_CONFIG_OPT(CPUFeatureRegisters, CPUFEATUREREGISTERS);
CPUFeatures Features {};
if (!CPUFeatureRegisters().empty()) {
Features = GetCPUFeaturesFromConfig(CPUFeatureRegisters());
} else {
#ifdef _M_X86_64
Features = CPUFeaturesAll {};
CPUFeatures Features = CPUFeaturesAll {};
// Vixl simulator doesn't support AFP.
Features.RemoveFeature(CPUFeatures::Feature::AFP);
// Vixl simulator doesn't support RPRES.
Features.RemoveFeature(CPUFeatures::Feature::RPRES);
// Vixl simulator doesn't support AFP.
Features.RemoveFeature(CPUFeatures::Feature::AFP);
// Vixl simulator doesn't support RPRES.
Features.RemoveFeature(CPUFeatures::Feature::RPRES);
#else
Features = GetCPUFeaturesFromIDRegisters();
CPUFeatures Features = GetCPUFeaturesFromIDRegisters();
#endif
}
uint64_t CTR = 0;
uint64_t MIDR = 0;
@@ -746,9 +623,8 @@ FEXCore::HostFeatures FetchHostFeatures() {
__asm volatile("mrs %[midr], midr_el1" : [midr] "=r"(MIDR));
#endif
FEXCore::HostFeatures HostFeatures = {};
auto HostFeatures = FetchHostFeatures(Features, true, CTR, MIDR);
FillMIDRInformationViaLinux(&HostFeatures);
FetchHostFeatures(Features, HostFeatures, true, CTR, MIDR);
HostFeatures.SupportsCPUIndexInTPIDRRO = false;
return HostFeatures;
+30 -69
View File
@@ -8,23 +8,6 @@
namespace FEX {
class CPUFeatures {
public:
class FeatureReg {
public:
void SetReg(uint64_t _Reg) {
Reg = _Reg;
}
uint64_t Get() const {
return Reg;
}
protected:
// All feature flag fields are 4-bits.
uint64_t GetField(uint64_t Offset) const {
return (Reg >> Offset) & 0b1111;
}
uint64_t Reg {};
};
enum class Feature : uint32_t {
// ISAR0
AES,
@@ -117,25 +100,23 @@ public:
MAX,
};
class DCZIDReg final : public FeatureReg {
public:
bool SupportsDCZVA() const {
return (Reg & DCZID_DZP_MASK) == 0;
}
static_assert(FEXCore::ToUnderlying(Feature::MAX) < 128);
static_assert((FEXCore::ToUnderlying(Feature::MAX) / (sizeof(uint64_t) * 8)) == 1);
uint32_t BlockSizeInBytes() const {
uint32_t DCZID_Log2 = Reg & DCZID_BS_MASK;
return (1 << DCZID_Log2) * sizeof(uint32_t);
}
bool Supports(Feature feat) const {
const size_t DWordSelect = FEXCore::ToUnderlying(feat) / (sizeof(uint64_t) * 8);
const size_t BitSelect = FEXCore::ToUnderlying(feat) - (DWordSelect * (sizeof(uint64_t) * 8));
return (FeatureBits[DWordSelect] >> BitSelect) & 1;
}
private:
// Data Zero Prohibited flag
// 0b0 = ZVA/GVA/GZVA permitted
// 0b1 = ZVA/GVA/GZVA prohibited
[[maybe_unused]] constexpr static uint32_t DCZID_DZP_MASK = 0b1'0000;
// Log2 of the blocksize in 32-bit words
[[maybe_unused]] constexpr static uint32_t DCZID_BS_MASK = 0b0'1111;
};
void RemoveFeature(Feature feat) {
const size_t DWordSelect = FEXCore::ToUnderlying(feat) / (sizeof(uint64_t) * 8);
const size_t BitSelect = FEXCore::ToUnderlying(feat) - (DWordSelect * (sizeof(uint64_t) * 8));
FeatureBits[DWordSelect] &= ~(1ULL << BitSelect);
}
protected:
void FillFeatureFlags();
// This list is informed by Linux kernel's `Documentation/arch/arm64/cpu-feature-registers.rst`
enum class FeatureRegType {
@@ -151,11 +132,24 @@ public:
ISAR2_EL1,
};
class FeatureReg {
public:
void SetReg(uint64_t _Reg) {
Reg = _Reg;
}
protected:
// All feature flag fields are 4-bits.
uint64_t GetField(uint64_t Offset) const {
return (Reg >> Offset) & 0b1111;
}
uint64_t Reg {};
};
#define FIELD_FETCHER(feature, field, minimum_field) \
bool Supports##feature() const { \
return GetField(field) >= minimum_field; \
}
class ISAR0Reg final : public FeatureReg {
public:
FIELD_FETCHER(AES, AES, 0b0001);
@@ -607,11 +601,8 @@ public:
ATS1A = 15 * 4,
};
};
class SVEVLReg final : public FeatureReg {};
#undef FIELD_FETCHER
ISAR0Reg ISAR0;
PFR0Reg PFR0;
PFR1Reg PFR1;
@@ -622,34 +613,6 @@ public:
MMFR2Reg MMFR2;
MMFR1Reg MMFR1;
ISAR2Reg ISAR2;
DCZIDReg DCZID;
SVEVLReg SVEVL;
static_assert(FEXCore::ToUnderlying(Feature::MAX) < 128);
static_assert((FEXCore::ToUnderlying(Feature::MAX) / (sizeof(uint64_t) * 8)) == 1);
bool Supports(Feature feat) const {
const size_t DWordSelect = FEXCore::ToUnderlying(feat) / (sizeof(uint64_t) * 8);
const size_t BitSelect = FEXCore::ToUnderlying(feat) - (DWordSelect * (sizeof(uint64_t) * 8));
return (FeatureBits[DWordSelect] >> BitSelect) & 1;
}
void RemoveFeature(Feature feat) {
const size_t DWordSelect = FEXCore::ToUnderlying(feat) / (sizeof(uint64_t) * 8);
const size_t BitSelect = FEXCore::ToUnderlying(feat) - (DWordSelect * (sizeof(uint64_t) * 8));
FeatureBits[DWordSelect] &= ~(1ULL << BitSelect);
}
const DCZIDReg& GetDCZID() const {
return DCZID;
}
uint64_t GetSVEVectorLengthInBits() const {
return SVEVL.Get();
}
protected:
void FillFeatureFlags();
uint64_t FeatureBits[(FEXCore::ToUnderlying(Feature::MAX) / (sizeof(uint64_t) * 8)) + 1] {};
@@ -662,8 +625,6 @@ protected:
void FillMIDRInformationViaLinux(FEXCore::HostFeatures* Features);
void FetchHostFeatures(FEX::CPUFeatures& Features, FEXCore::HostFeatures& HostFeatures, bool SupportsCacheMaintenanceOps, uint64_t CTR,
uint64_t MIDR);
FEXCore::HostFeatures FetchHostFeatures(FEX::CPUFeatures& Features, bool SupportsCacheMaintenanceOps, uint64_t CTR, uint64_t MIDR);
FEXCore::HostFeatures FetchHostFeatures();
FEX::CPUFeatures GetCPUFeaturesFromIDRegisters();
} // namespace FEX
-33
View File
@@ -1,33 +0,0 @@
add_executable(FEXCompatTool
CompatTool.cpp)
target_link_libraries(FEXCompatTool
PRIVATE
FEXCore Common CommonTools JemallocLibs)
install(TARGETS FEXCompatTool
RUNTIME
DESTINATION /
COMPONENT Runtime
)
add_executable(FEXServerManager
ServerManager.cpp)
target_link_libraries(FEXServerManager
PRIVATE
FEXCore Common CommonTools JemallocLibs)
install(TARGETS FEXServerManager
RUNTIME
DESTINATION bin
COMPONENT Runtime
)
# Description json gets installed into root of depot
install(FILES emulator.json
DESTINATION /
COMPONENT Runtime)
install(FILES ConfigTemplate.json
DESTINATION /
COMPONENT Runtime)
-197
View File
@@ -1,197 +0,0 @@
// SPDX-License-Identifier: MIT
/*
$info$
tags: Bin|FEXCompatTool
desc: Used for launching games from Steam
$end_info$
*/
#include "PortabilityInfo.h"
#include "Common/Config.h"
#include "FEXCore/Utils/FileLoading.h"
#include "FEXCore/Utils/StringUtils.h"
#include "FEXHeaderUtils/Filesystem.h"
#include <stdlib.h>
#include <tiny-json.h>
fextl::string GenerateSteamConfigTemplate(const FEX::Config::PortableInformation& PortableInfo) {
const auto ConfigTemplatePath = PortableInfo.InterpreterPath + "ConfigTemplate.json";
if (!FHU::Filesystem::Exists(ConfigTemplatePath)) {
return {};
}
fextl::string Data;
if (!FEXCore::FileLoading::LoadFile(Data, ConfigTemplatePath)) {
return {};
}
// Try and find a mount point.
fextl::string MountPoint {};
const char* RuntimeDir = getenv("XDG_RUNTIME_DIR");
if (RuntimeDir) {
MountPoint = fextl::fmt::format("{}/fexrootfs/", RuntimeDir);
} else {
const auto UserDirectory = fextl::fmt::format("/run/user/{}", geteuid());
if (FHU::Filesystem::Exists(UserDirectory)) {
MountPoint = fextl::fmt::format("{}/fexrootfs/", UserDirectory);
} else {
const char* CacheDir = getenv("XDG_CACHE_HOME");
if (CacheDir) {
MountPoint = fextl::fmt::format("{}/fexrootfs/", CacheDir);
} else {
// We tried really hard to find a mount path.
MountPoint = "~/.cache/fexrootfs/";
}
}
}
// Update the @FEX_COMPAT_TOOL@ config to point to the root of the depot.
FEXCore::StringUtils::ReplaceAllInPlace(Data, "@FEX_COMPAT_TOOL@", PortableInfo.InterpreterPath);
// TODO: This path is getting phased out.
FEXCore::StringUtils::ReplaceAllInPlace(Data, "@FEX_ROOTFS_PATH@", MountPoint);
// Save the json.
const auto ConfigPath = FEX::Config::GetConfigDirectory(false, PortableInfo);
const auto ConfigLocation = ConfigPath + "Config.json";
if (!FHU::Filesystem::CreateDirectories(ConfigPath)) {
return {};
}
auto File = FEXCore::File::File(ConfigLocation.c_str(),
FEXCore::File::FileModes::WRITE | FEXCore::File::FileModes::CREATE | FEXCore::File::FileModes::TRUNCATE);
if (!File.IsValid()) {
return {};
}
File.Write(Data.data(), Data.size());
return ConfigPath;
}
fextl::string GenerateSteamAppConfig(const FEX::Config::PortableInformation& PortableInfo) {
const auto user_config = getenv("FEX_APP_CONFIG");
if (user_config) {
// If user supplied config then don't use Steam config.
return {};
}
// Current supported Steam options.
struct SteamOptions {
bool TSO = true;
bool Multiblock = true;
bool Thunks_GL = false;
bool Thunks_Vulkan = false;
bool EnableLogging = false;
};
SteamOptions Options {};
// Game overrides.
const auto steam_fex_tso = getenv("STEAM_FEX_TSOENABLED");
if (steam_fex_tso) {
Options.TSO = std::strtoull(steam_fex_tso, nullptr, 0) != 0;
}
const auto steam_fex_multiblock = getenv("STEAM_FEX_MULTIBLOCK");
if (steam_fex_multiblock) {
Options.Multiblock = std::strtoull(steam_fex_multiblock, nullptr, 0) != 0;
}
const auto steam_fex_logging = getenv("STEAM_FEX_LOG");
if (steam_fex_logging) {
Options.EnableLogging = std::strtoull(steam_fex_logging, nullptr, 0) != 0;
}
// UI overrides.
const auto steam_fex_compat = getenv("STEAM_COMPAT_FEX_CONFIG");
if (steam_fex_compat) {
const auto steam_fex_compat_view = std::string_view(steam_fex_compat);
if (steam_fex_compat_view.find("TSOEnabled:1") != steam_fex_compat_view.npos) {
Options.TSO = true;
}
if (steam_fex_compat_view.find("Multiblock:1") != steam_fex_compat_view.npos) {
Options.Multiblock = true;
}
if (steam_fex_compat_view.find("ThunksDB_GL:1") != steam_fex_compat_view.npos) {
Options.Thunks_GL = true;
}
if (steam_fex_compat_view.find("ThunksDB_Vulkan:1") != steam_fex_compat_view.npos) {
Options.Thunks_Vulkan = true;
}
}
// Create the json.
char Buffer[4096];
char* Dest {};
Dest = json_objOpen(Buffer, nullptr);
{
Dest = json_objOpen(Dest, "Config");
Dest = json_str(Dest, "TSOEnabled", Options.TSO ? "1" : "0");
Dest = json_str(Dest, "Multiblock", Options.Multiblock ? "1" : "0");
Dest = json_str(Dest, "SilentLog", Options.EnableLogging ? "0" : "1");
if (Options.EnableLogging) {
Dest = json_str(Dest, "OutputLog", "server");
}
Dest = json_objClose(Dest);
}
{
Dest = json_objOpen(Dest, "ThunksDB");
Dest = json_str(Dest, "GL", Options.Thunks_GL ? "1" : "0");
Dest = json_str(Dest, "Vulkan", Options.Thunks_Vulkan ? "1" : "0");
Dest = json_objClose(Dest);
}
Dest = json_objClose(Dest);
json_end(Dest);
// Save the json.
const auto ConfigPath = FEX::Config::GetConfigDirectory(false, PortableInfo);
const auto ConfigLocation = ConfigPath + "app_config.json";
if (!FHU::Filesystem::CreateDirectories(ConfigPath)) {
return {};
}
auto File = FEXCore::File::File(ConfigLocation.c_str(),
FEXCore::File::FileModes::WRITE | FEXCore::File::FileModes::CREATE | FEXCore::File::FileModes::TRUNCATE);
if (!File.IsValid()) {
return {};
}
File.Write(Buffer, strlen(Buffer));
return ConfigLocation;
}
int main(int argc, const char** argv) {
const auto PortableInfo = FEX::ReadPortabilityInformation();
const auto TemplateConfigPath = GenerateSteamConfigTemplate(PortableInfo);
const auto AppConfigPath = GenerateSteamAppConfig(PortableInfo);
if (!TemplateConfigPath.empty()) {
setenv("FEX_APP_CONFIG_LOCATION", TemplateConfigPath.c_str(), true);
}
if (!AppConfigPath.empty()) {
setenv("FEX_APP_CONFIG", AppConfigPath.c_str(), true);
}
const auto FEXInterpreterPath = PortableInfo.InterpreterPath + "usr/bin/FEX";
// Due to no arguments for this application, just replace argv[0] and execve again.
argv[0] = FEXInterpreterPath.c_str();
execv(FEXInterpreterPath.c_str(), const_cast<char* const*>(argv));
// Save errno as it can change after calling `perror`.
const auto saved_errno = errno;
perror(argv[0]);
if (saved_errno == ENOENT) {
return 127;
}
return 126;
}
-9
View File
@@ -1,9 +0,0 @@
{
"Config": {
"X87ReducedPrecision": "1",
"RootFS": "@FEX_ROOTFS_PATH@/",
"ThunkHostLibs": "@FEX_COMPAT_TOOL@/usr/lib/aarch64-linux-gnu/fex-emu/HostThunks",
"ThunkGuestLibs": "@FEX_COMPAT_TOOL@/usr/share/fex-emu/GuestThunks",
"ProfileStats": "1"
}
}
-96
View File
@@ -1,96 +0,0 @@
// SPDX-License-Identifier: MIT
#include "PortabilityInfo.h"
#include "Common/FEXServerClient.h"
#include <cstdio>
#include <unistd.h>
#include <poll.h>
void MsgHandler(LogMan::DebugLevels Level, const char* Message) {
const auto Style = fmt::text_style {};
const auto Output = fextl::fmt::format("{} {}\n", fmt::styled(LogMan::DebugLevelStr(Level), Style), Message);
write(STDERR_FILENO, Output.c_str(), Output.size());
fsync(STDERR_FILENO);
}
void AssertHandler(const char* Message) {
return MsgHandler(LogMan::ASSERT, Message);
}
void SignalPVToContinue() {
// Tell pressure-vessel that the startup was a success.
const auto ReadyMsg = "READY=1\n";
write(STDOUT_FILENO, ReadyMsg, strlen(ReadyMsg));
// pressure-vessel is waiting for EOF on STDOUT from this process to ensure it can run FEX processes.
// dup2 atomically replaces stdout with a copy of stderr to achieve this.
dup2(STDERR_FILENO, STDOUT_FILENO);
}
struct PipesType {
int read_pipe {-1};
int write_pipe {-1};
};
PipesType get_pipe() {
PipesType pipes {};
pipe(&pipes.read_pipe);
return pipes;
}
int main(int argc, const char** argv, char** const envp) {
LogMan::Throw::InstallHandler(AssertHandler);
LogMan::Msg::InstallHandler(MsgHandler);
const auto PortableInfo = FEX::ReadPortabilityInformation();
FEX::Config::LoadConfig({}, envp, PortableInfo);
// Reload the meta layer
FEXCore::Config::ReloadMetaLayer();
auto pipes = get_pipe();
// Set the write side to close on exec.
fcntl(pipes.write_pipe, F_SETFD, FD_CLOEXEC);
// Give the read end of the pipe to FEXServer.
auto ServerFD = FEXServerClient::StartServer(PortableInfo.InterpreterPath, pipes.read_pipe);
if (ServerFD == -1) {
perror("Couldn't start FEXServer");
return 126;
}
// FEXServer is now running. Tell PV to continue.
SignalPVToContinue();
// Don't need the read pipe anymore.
close(pipes.read_pipe);
pipes.read_pipe = -1;
// Now that the server is started and watching our pipe, we can close the returned FD, as it'll stay open as long as the pipe is open.
close(ServerFD);
ServerFD = -1;
// stdin will be a pipe, so wait until that FD is closed.
while (true) {
pollfd p {
.fd = STDIN_FILENO,
.events = POLLRDHUP,
.revents = 0,
};
int events = poll(&p, 1, -1);
if (events == -1 && errno == EINTR) {
continue;
}
if (events > 0 && (p.revents & (POLLRDHUP | POLLERR | POLLHUP | POLLNVAL))) {
// Error or pressure-vessel hung-up.
break;
}
}
// Terminating will clean-up.
return 0;
}
-13
View File
@@ -1,13 +0,0 @@
{
"emulator_v0": {
"argv": "./usr/bin/FEX",
"environment": { "FEX_PORTABLE": "1" },
"container_argv": "./usr/bin/FEX",
"container_environment": { "FEX_ROOTFS": "" },
"main_argv": "./FEXCompatTool",
"server_argv": "./usr/bin/FEXServerManager",
"emulated_architectures": ["x86_64-linux-gnu", "i386-linux-gnu"],
"required_architectures": ["aarch64-linux-gnu"],
"required_libraries": ["libc.so.6", "libstdc++.so.6"]
}
}
+1 -5
View File
@@ -14,14 +14,10 @@ if (NOT MINGW_BUILD)
add_subdirectory(FEXGDBReader/)
endif()
if (NOT BUILD_STEAM_SUPPORT)
add_subdirectory(FEXRootFSFetcher/)
endif()
add_subdirectory(FEXRootFSFetcher/)
add_subdirectory(FEXGetConfig/)
add_subdirectory(FEXServer/)
add_subdirectory(FEXBash/)
add_subdirectory(FEXOfflineCompiler/)
add_subdirectory(CodeSizeValidation/)
add_subdirectory(LinuxEmulation/)
-1
View File
@@ -5,7 +5,6 @@
#include <FEXCore/fextl/vector.h>
#include <cstdint>
#include <unistd.h>
namespace FEX {
+1 -5
View File
@@ -78,10 +78,6 @@ void ConfigModel::Reload() {
if (!LoadedConfig->OptionExists(Option.first)) {
continue;
}
if (std::holds_alternative<fextl::list<fextl::string>>(Option.second)) {
// Omit string lists from the model since they require special handling
continue;
}
auto& [Name, TypeId] = ConfigToNameLookup.find(Option.first)->second;
auto Item = new QStandardItem(QString::fromStdString(Name));
@@ -126,7 +122,7 @@ void ConfigModel::setStringList(const QString& Name, const QStringList& Values)
const auto& Option = NameToConfigLookup.at(Name.toStdString());
LoadedConfig->Erase(Option);
for (auto& Value : Values) {
LoadedConfig->AppendStrArrayValue(Option, Value.toStdString().c_str());
LoadedConfig->Set(Option, Value.toStdString().c_str());
}
Reload();
}
Loaded 100 of 213 files, more files were not shown because too many files have changed in this diff. Show more