Compare commits

..
Author SHA1 Message Date
Billy Laws 9fe5eb1979 JIT: Restore behaviour of emitting interrupt checks at every block entry
This is needed to handle suspend in infinite loops that occur as a
result of block-size constraints or indirect jumps. Fixes grow home.
2025-10-29 00:35:46 +00:00
Ryan Houdek f414c92963 Code view 2025-10-28 23:53:15 +00:00
Ryan Houdek 90c59e37cb unittests/ASM: Adds test for too large branch objects 2025-10-28 23:53:15 +00:00
Ryan Houdek b7c7789a01 FEXCore/JIT: Supports restarting JIT in case of encoding failure
ARM64 branches have fairly small relative distances they can encode.
These can be +-1MB, or even +-32KB. The largest relative branch is
+-128MB, which we already set as an upper limit of our block JIT cache
size.

We have for a long time just compiled these without checking with the
expectation that things just happen to work. We didn't hit the asserts
so it was relatively low priority. Apparently now with Steam and a
MaxInst limit of 5000, we are now hitting an assert where we are
encoding too large of a range.

Implement support for long jumping from anywhere in the JIT for when a
long jump tries to be encoded and fails, allowing us to restart the JIT
at any moment. This is implemented as a long jump when this singular
feature could have gotten away with some sort of invasive check and
early exit path for two reasons. For one, that would be even more
invasive, effectively doing try-catch logic manually. And two, the next
step is supporting JIT buffer overflow for when our block size heuristic
fails.

This next step will mandate longjump on SIGSEGV (with cooperative
interaction with the frontend) from effectively /anywhere/ in the JIT.
One of the design goals of the CodeEmitter is that every code emission
function doesn't do a size remaining check to allow the compiler to do
some very effective optimization of emitting code blocks to memory (and
it works!).

But we lose the ability to sanely size check. When writing the emitter I
knew we were going to need to write this cooperative guard page handler,
and we're finally at a point where it needs to be done. This will be in
the next PR although.
2025-10-28 23:53:15 +00:00
Ryan Houdek f653c5e0c0 FEXCore/JIT: Ignore local encoding limit checks
These are guaranteed not to hit encoding distance limits, so we can
ignore the returns.
2025-10-28 23:53:15 +00:00
Ryan Houdek 65fff73959 FEXCore/Dispatcher: Check encoding errors 2025-10-28 23:53:15 +00:00
Ryan Houdek 93b7c513d8 FEXCore/VectorRegType: Trivial header fix 2025-10-28 23:53:15 +00:00
Ryan Houdek f1d14c6325 Linux/BPFEmitter: Explicitly ignored encoding bool
We know these won't encode in errors.
2025-10-28 23:53:15 +00:00
Ryan Houdek 8223c6ac36 unittests/Emitter: Explicitly ignore encoding bool
We know these won't encode in errors.
2025-10-28 23:53:15 +00:00
Ryan Houdek e17677580d CodeEmitter: Return bool if Label instructions can't be encoded
Programming error if they aren't checked, as they will encode
incorrectly if they are too large for their respective instructions.
2025-10-28 23:53:15 +00:00
Ryan Houdek 150bf7b30c FEXCore: Moves longjump implementation from FEX frontend
This will be getting used by FEXCore in a bit.
2025-10-28 23:53:15 +00:00
Ryan Houdek 674efc69c4 FEX: Print a log when kernel unaligned atomics are used 2025-10-28 23:53:15 +00:00
Billy Laws 38049c5281 Windows: Enable downstream kernel-side unaligned atomic handling 2025-10-28 23:53:15 +00:00
Billy Laws e3627349a1 FEXLoader: Enable downstream kernel-side unaligned atomic handling 2025-10-28 23:53:15 +00:00
Billy Laws 7c207080a4 Windows: Support new two-stage invalidation model 2025-10-28 23:53:12 +00:00
Billy Laws eeee5b53ca Linux: Support new two-stage invalidation model 2025-10-28 23:53:12 +00:00
Billy Laws 47619063c2 LookupCache: Introduce two-pass code invalidation model
Shared code buffer support introduced the concept of having a single
GuestToHostMaps shared across many threads. In the common case all
threads will share one however if e.g. a resize recently occured and
specific thread is yet to compile any code with the new codebuffer it
will still use the old GuestToHostMap. The current invalidation
approach handles this by repeatedly calling erase for every single
thread's GuestToHostMap, even if it is repeated. An accumulator is used
to ensure when two threads share a map, the L1/L2 cache entries in the
second thread will still be invalidated even if the the iteration for
the first thread removed them from the map.

Unfortunately this is incredibly slow in cases with many threads, as
a significant number of redundant map lookups and L1/L2 cache erasures
on threads that never even observed a given block can occur. Solve this
by introducing a two-pass model:
- First, all active codebuffers (and their associated GuestToHostMaps)
  have their entries invalidated for the given range, these codebuffers
  are tracked internally within FEXCore. It is at this point that delinking
  callbacks are ran.
- Second, each thread will have its caches invalidated. But rather than
  naively invalidating the L1/L2 caches for every invalidated block for
  every thread, threads now track on their own what specific entries
  have been potentially fetched into their L1/L2 caches. This is
  aided by GuestToHostMap now tracking the pages each block touches. (an
  inverse CodePages so to speak).
2025-10-28 23:53:12 +00:00
Billy Laws cb7076cbab FEXCore: Keep a list of weak refs to all allocated codebuffers
We currently rely on the frontend to keep track of threads and then
iterate over all threads to perform per-codebuffer operations. However
as codebuffers are shared between many threads (the common case is a
single code buffer across all) this ends up being inefficient. Introduce
a list of codebuffers to solve that (new codebuffers are very rare, so a
vector is plenty fine here for erasing invalid weak refs).
2025-10-28 23:53:12 +00:00
Billy Laws 8dde79826e LookupCache: Drop unused state frame argument for delinker cbs 2025-10-28 23:53:12 +00:00
68 changed files with 2493 additions and 2848 deletions

No files matched your search

+2 -2
View File
@@ -40,7 +40,7 @@ public:
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(IsADRRange(Imm), "Unscaled offset too large");
if (IsADRRange(Imm)) {
if (IsADRRange(Imm)) [[likely]] {
constexpr uint32_t Op = 0b0001'0000 << 24;
DataProcessing_PCRel_Imm(Op, rd, Imm);
return BranchEncodeSucceeded::Success;
@@ -75,7 +75,7 @@ public:
int64_t Imm = reinterpret_cast<int64_t>(Label->Location) - (GetCursorAddress<int64_t>() & ~0xFFFLL);
LOGMAN_THROW_A_FMT(IsADRPRange(Imm) && IsADRPAligned(Imm), "Unscaled offset too large");
if (IsADRPRange(Imm) && IsADRPAligned(Imm)) {
if (IsADRPRange(Imm) && IsADRPAligned(Imm)) [[likely]] {
constexpr uint32_t Op = 0b1001'0000 << 24;
DataProcessing_PCRel_Imm(Op, rd, Imm);
return BranchEncodeSucceeded::Success;
+8 -8
View File
@@ -22,7 +22,7 @@ public:
}
[[nodiscard]] BranchEncodeSucceeded b(ARMEmitter::Condition Cond, const BackwardLabel* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
if (Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0)) {
if (Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0)) [[likely]] {
constexpr uint32_t Op = 0b0101'010 << 25;
Branch_Conditional(Op, 0, 0, Cond, Imm >> 2);
return BranchEncodeSucceeded::Success;
@@ -55,7 +55,7 @@ public:
}
[[nodiscard]] BranchEncodeSucceeded bc(ARMEmitter::Condition Cond, const BackwardLabel* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
if (Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0)) {
if (Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0)) [[likely]] {
constexpr uint32_t Op = 0b0101'010 << 25;
Branch_Conditional(Op, 0, 1, Cond, Imm >> 2);
return BranchEncodeSucceeded::Success;
@@ -116,7 +116,7 @@ public:
}
[[nodiscard]] BranchEncodeSucceeded b(const BackwardLabel* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
if (Imm >= -134217728 && Imm <= 134217724 && ((Imm & 0b11) == 0)) {
if (Imm >= -134217728 && Imm <= 134217724 && ((Imm & 0b11) == 0)) [[likely]] {
constexpr uint32_t Op = 0b0001'01 << 26;
UnconditionalBranch(Op, Imm >> 2);
return BranchEncodeSucceeded::Success;
@@ -151,7 +151,7 @@ public:
[[nodiscard]] BranchEncodeSucceeded bl(const BackwardLabel* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
if (Imm >= -134217728 && Imm <= 134217724 && ((Imm & 0b11) == 0)) {
if (Imm >= -134217728 && Imm <= 134217724 && ((Imm & 0b11) == 0)) [[likely]] {
constexpr uint32_t Op = 0b1001'01 << 26;
UnconditionalBranch(Op, Imm >> 2);
@@ -189,7 +189,7 @@ public:
[[nodiscard]] BranchEncodeSucceeded cbz(ARMEmitter::Size s, ARMEmitter::Register rt, const BackwardLabel* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
if (Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0)) {
if (Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0)) [[likely]] {
constexpr uint32_t Op = 0b0011'0100 << 24;
CompareAndBranch(Op, s, rt, Imm >> 2);
return BranchEncodeSucceeded::Success;
@@ -227,7 +227,7 @@ public:
[[nodiscard]] BranchEncodeSucceeded cbnz(ARMEmitter::Size s, ARMEmitter::Register rt, const BackwardLabel* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
if (Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0)) {
if (Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0)) [[likely]] {
constexpr uint32_t Op = 0b0011'0101 << 24;
CompareAndBranch(Op, s, rt, Imm >> 2);
return BranchEncodeSucceeded::Success;
@@ -265,7 +265,7 @@ public:
[[nodiscard]] BranchEncodeSucceeded tbz(ARMEmitter::Register rt, uint32_t Bit, const BackwardLabel* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
if (Imm >= -32768 && Imm <= 32764 && ((Imm & 0b11) == 0)) {
if (Imm >= -32768 && Imm <= 32764 && ((Imm & 0b11) == 0)) [[likely]] {
constexpr uint32_t Op = 0b0011'0110 << 24;
TestAndBranch(Op, rt, Bit, Imm >> 2);
return BranchEncodeSucceeded::Success;
@@ -303,7 +303,7 @@ public:
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(Imm >= -32768 && Imm <= 32764 && ((Imm & 0b11) == 0), "Unscaled offset too large");
if (Imm >= -32768 && Imm <= 32764 && ((Imm & 0b11) == 0)) {
if (Imm >= -32768 && Imm <= 32764 && ((Imm & 0b11) == 0)) [[likely]] {
constexpr uint32_t Op = 0b0011'0111 << 24;
TestAndBranch(Op, rt, Bit, Imm >> 2);
return BranchEncodeSucceeded::Success;
+5 -5
View File
@@ -662,7 +662,7 @@ public:
case ForwardLabel::InstType::ADR: {
uint32_t* Instruction = reinterpret_cast<uint32_t*>(Label->Location);
int64_t Imm = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(Instruction);
if (!IsADRRange(Imm)) {
if (!IsADRRange(Imm)) [[unlikely]] {
// Can't bind.
return false;
}
@@ -678,7 +678,7 @@ public:
uint32_t* Instruction = reinterpret_cast<uint32_t*>(Label->Location);
int64_t Imm = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(Instruction);
if (!(IsADRPRange(Imm) && IsADRPAligned(Imm))) {
if (!(IsADRPRange(Imm) && IsADRPAligned(Imm))) [[unlikely]] {
// Can't bind.
return false;
}
@@ -695,7 +695,7 @@ public:
case ForwardLabel::InstType::B: {
uint32_t* Instruction = reinterpret_cast<uint32_t*>(Label->Location);
int64_t Imm = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(Instruction);
if (!(Imm >= -134217728 && Imm <= 134217724 && ((Imm & 0b11) == 0))) {
if (!(Imm >= -134217728 && Imm <= 134217724 && ((Imm & 0b11) == 0))) [[unlikely]] {
// Can't bind.
return false;
}
@@ -711,7 +711,7 @@ public:
case ForwardLabel::InstType::TEST_BRANCH: {
uint32_t* Instruction = reinterpret_cast<uint32_t*>(Label->Location);
int64_t Imm = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(Instruction);
if (!(Imm >= -32768 && Imm <= 32764 && ((Imm & 0b11) == 0))) {
if (!(Imm >= -32768 && Imm <= 32764 && ((Imm & 0b11) == 0))) [[unlikely]] {
// Can't bind.
return false;
}
@@ -728,7 +728,7 @@ public:
case ForwardLabel::InstType::RELATIVE_LOAD: {
uint32_t* Instruction = reinterpret_cast<uint32_t*>(Label->Location);
int64_t Imm = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(Instruction);
if (!(Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0))) {
if (!(Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0))) [[unlikely]] {
// Can't bind.
return false;
}
+7 -6
View File
@@ -62,7 +62,7 @@ struct CustomIRResult {
, Data(Data) {}
};
using BlockDelinkerFunc = void (*)(FEXCore::Core::CpuStateFrame* Frame, FEXCore::Context::ExitFunctionLinkData* Record);
using BlockDelinkerFunc = void (*)(FEXCore::Context::ExitFunctionLinkData* Record);
constexpr uint32_t TSC_SCALE_MAXIMUM = 1'000'000'000; ///< 1Ghz
class CodeCache : public AbstractCodeCache {
@@ -155,10 +155,10 @@ public:
return CodeCache;
}
void OnCodeBufferAllocated(CPU::CodeBuffer&) override;
void OnCodeBufferAllocated(const std::shared_ptr<CPU::CodeBuffer> &) override;
void ClearCodeCache(FEXCore::Core::InternalThreadState* Thread, bool NewCodeBuffer = true) override;
void InvalidateGuestCodeRange(FEXCore::Core::InternalThreadState* Thread, InvalidatedEntryAccumulator& Accumulator, uint64_t Start,
uint64_t Length) override;
void InvalidateCodeBuffersCodeRange(uint64_t Start, uint64_t Length) override;
void InvalidateThreadCachedCodeRange(FEXCore::Core::InternalThreadState* Thread, uint64_t Start, uint64_t Length) override;
FEXCore::ForkableSharedMutex& GetCodeInvalidationMutex() override {
return CodeInvalidationMutex;
}
@@ -231,8 +231,6 @@ public:
ContextImpl(const FEXCore::HostFeatures& Features);
static bool ThreadRemoveCodeEntry(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP, const FEXCore::LookupCacheWriteLockToken& lk);
static void ThreadRemoveCodeEntryFromJit(FEXCore::Core::CpuStateFrame* Frame, uint64_t GuestRIP);
// This is used as a replacement for the SMC writes in the mono callsite backpatcher that avoids atomic operations
@@ -348,5 +346,8 @@ private:
bool MonoDetected = false;
std::atomic<uint64_t> MonoBackpatcherBlock;
std::mutex CodeBufferListLock;
fextl::vector<std::weak_ptr<CPU::CodeBuffer>> CodeBufferList;
};
} // namespace FEXCore::Context
+1 -1
View File
@@ -400,7 +400,7 @@ namespace CPU {
Latest = Buffer;
LatestOffset = 0;
OnCodeBufferAllocated(*Buffer);
OnCodeBufferAllocated(Buffer);
return Buffer;
}
+1 -1
View File
@@ -81,7 +81,7 @@ namespace CPU {
// Protects writes to the latest CodeBuffer and changes to LatestOffset
FEXCore::ForkableUniqueMutex CodeBufferWriteMutex;
virtual void OnCodeBufferAllocated(CodeBuffer&) {};
virtual void OnCodeBufferAllocated(const std::shared_ptr<CodeBuffer>&) {};
private:
fextl::shared_ptr<CodeBuffer> Latest;
-2
View File
@@ -88,7 +88,6 @@ namespace ProductNames {
static const char ARM_Blizzard_M2Pro[] = "Apple Blizzard (M2 Pro)";
static const char ARM_Avalanche_M2Max[] = "Apple Avalanche (M2 Max)";
static const char ARM_Blizzard_M2Max[] = "Apple Blizzard (M2 Max)";
static const char ARM_AppleSilicon[] = "Apple Silicon";
static const char ARM_ORYON_1[] = "Oryon-1";
static const char ARM_Ampere_1[] = "AmpereOne";
@@ -189,7 +188,6 @@ void CPUIDEmu::SetupHostHybridFlag() {
{0x61, 0x029, 1, ProductNames::ARM_Firestorm_M1Max}, // Apple Firestorm (M1 Max)
{0x61, 0x025, 1, ProductNames::ARM_Firestorm_M1Pro}, // Apple Firestorm (M1 Pro)
{0x61, 0x023, 1, ProductNames::ARM_Firestorm_M1}, // Apple Firestorm (M1)
{0x61, 0, 1, ProductNames::ARM_AppleSilicon}, // QEmu Apple Silicon
{0x41, 0xd8c, 1, ProductNames::ARM_C1Ultra}, // C1-Ultra
{0x41, 0xd90, 1, ProductNames::ARM_C1Premium}, // C1-Premium
+35 -38
View File
@@ -459,9 +459,14 @@ void ContextImpl::LockBeforeFork(FEXCore::Core::InternalThreadState* Thread) {
}
#endif
void ContextImpl::OnCodeBufferAllocated(CPU::CodeBuffer& Buffer) {
void ContextImpl::OnCodeBufferAllocated(const fextl::shared_ptr<CPU::CodeBuffer>& Buffer) {
if (Config.GlobalJITNaming()) {
Symbols.RegisterJITSpace(Buffer.Ptr, Buffer.Size);
Symbols.RegisterJITSpace(Buffer->Ptr, Buffer->Size);
}
{
std::scoped_lock lk{CodeBufferListLock};
CodeBufferList.emplace_back(Buffer);
}
}
@@ -833,10 +838,14 @@ uintptr_t ContextImpl::CompileBlock(FEXCore::Core::CpuStateFrame* Frame, uint64_
Thread->CPUBackend->ClearRelocations();
}
fextl::vector<uint64_t> CodePages;
if (NeedsAddGuestCodeRanges) {
// Track in the guest to host map all entrypoints for all pages the compiled block touches, if any page didn't previously
// contain code, inform the frontend so it can setup SMC detection.
auto BlockInfo = Thread->FrontendDecoder->GetDecodedBlockInfo();
CodePages.reserve(BlockInfo->CodePages.size());
CodePages.insert(CodePages.end(), BlockInfo->CodePages.begin(), BlockInfo->CodePages.end());
for (auto CodePage : BlockInfo->CodePages) {
if (Thread->LookupCache->AddBlockExecutableRange(Thread, BlockInfo->EntryPoints, CodePage, FEXCore::Utils::FEX_PAGE_SIZE)) {
SyscallHandler->MarkGuestExecutableRange(Thread, CodePage, FEXCore::Utils::FEX_PAGE_SIZE);
@@ -845,8 +854,9 @@ uintptr_t ContextImpl::CompileBlock(FEXCore::Core::CpuStateFrame* Frame, uint64_
}
// Insert to lookup cache
for (auto [GuestAddr, HostAddr] : CompiledCode.EntryPoints) {
Thread->LookupCache->AddBlockMapping(Thread, GuestAddr, HostAddr);
Thread->LookupCache->AddBlockMapping(Thread, GuestAddr, CodePages, HostAddr);
}
return (uintptr_t)CodePtr;
@@ -873,50 +883,37 @@ uintptr_t ContextImpl::CompileSingleStep(FEXCore::Core::CpuStateFrame* Frame, ui
return (uintptr_t)CodePtr;
}
static void InvalidateGuestThreadCodeRange(FEXCore::Core::InternalThreadState* Thread, InvalidatedEntryAccumulator& Accumulator,
uint64_t Start, uint64_t Length) {
// Ensures now-modified mappings aren't cached as being in their previous non-executable state.
void ContextImpl::InvalidateCodeBuffersCodeRange(uint64_t Start, uint64_t Length) {
FEXCORE_PROFILE_SCOPED("InvalidateCodeBuffersCodeRange");
LogMan::Throw::AFmt(CodeInvalidationMutex.try_lock() == false, "CodeInvalidationMutex needs to be unique_locked here");
std::scoped_lock lk {CodeBufferListLock};
auto it = CodeBufferList.begin();
while (it != CodeBufferList.end()) {
if (auto Strong = it->lock(); Strong) {
Strong->LookupCache->InvalidateRange(Start, Length);
it++;
} else {
it = CodeBufferList.erase(it);
}
}
}
void ContextImpl::InvalidateThreadCachedCodeRange(FEXCore::Core::InternalThreadState* Thread, uint64_t Start, uint64_t Length) {
LogMan::Throw::AFmt(CodeInvalidationMutex.try_lock() == false, "CodeInvalidationMutex needs to be unique_locked here");
// Ensures now-modified mappings aren't cached as being in their previous non-executable state.
// Accessing FrontendDecoder is safe as the thread's code invalidation mutex must be locked here.
Thread->FrontendDecoder->ResetExecutableRangeCache();
auto lk = Thread->LookupCache->AcquireWriteLock();
auto& CodePages = Thread->LookupCache->Shared->CodePages;
if (Thread->LookupCache->InvalidateCacheRange(Start, Length)) {
FEXCORE_PROFILE_SCOPED("InvalidateCallRet");
auto lower = CodePages.lower_bound(Start >> 12);
auto upper = CodePages.upper_bound((Start + Length - 1) >> 12);
for (auto it = lower; it != upper; it++) {
Accumulator.emplace_back(std::move(it->second));
}
bool InvalidatedAnyEntries = false;
for (const auto& PageEntries : Accumulator) {
for (const auto& Entry : PageEntries) {
if (ContextImpl::ThreadRemoveCodeEntry(Thread, Entry, lk)) {
InvalidatedAnyEntries = true;
}
}
}
if (InvalidatedAnyEntries) {
// This may cause access violations in the thread on Windows as zeroing is not atomic, this is handled by the frontend
Allocator::VirtualDontNeed(Thread->CallRetStackBase, FEXCore::Core::InternalThreadState::CALLRET_STACK_SIZE);
}
}
void ContextImpl::InvalidateGuestCodeRange(FEXCore::Core::InternalThreadState* Thread, InvalidatedEntryAccumulator& Accumulator,
uint64_t Start, uint64_t Length) {
InvalidateGuestThreadCodeRange(Thread, Accumulator, Start, Length);
}
bool ContextImpl::ThreadRemoveCodeEntry(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP,
const FEXCore::LookupCacheWriteLockToken& lk) {
LogMan::Throw::AFmt(static_cast<ContextImpl*>(Thread->CTX)->CodeInvalidationMutex.try_lock() == false, "CodeInvalidationMutex needs to "
"be unique_locked here");
return Thread->LookupCache->Erase(Thread->CurrentFrame, GuestRIP, lk);
}
void ContextImpl::ThreadRemoveCodeEntryFromJit(FEXCore::Core::CpuStateFrame* Frame, uint64_t GuestRIP) {
static_cast<ContextImpl*>(Frame->Thread->CTX)->SyscallHandler->InvalidateGuestCodeRange(Frame->Thread, GuestRIP, 1);
}
@@ -98,7 +98,7 @@ void Dispatcher::EmitDispatcher() {
ARMEmitter::BiDirectionalLabel LoopTop {};
#ifdef _M_ARM_64EC
(void)b(&LoopTop);
b(&LoopTop);
AbsoluteLoopTopAddressEnterECFillSRA = GetCursorAddress<uint64_t>();
ldr(STATE, EC_ENTRY_CPUAREA_REG, CPU_AREA_EMULATOR_DATA_OFFSET);
@@ -106,10 +106,10 @@ void Dispatcher::EmitDispatcher() {
ldr(RipReg, STATE_PTR(CpuStateFrame, State.rip));
// Force a single instruction block if ENTRY_FILL_SRA_SINGLE_INST_REG is nonzero entering the JIT, used for inline SMC handling.
(void)cbnz(ARMEmitter::Size::i32Bit, ENTRY_FILL_SRA_SINGLE_INST_REG, &CompileSingleStep);
cbnz(ARMEmitter::Size::i32Bit, ENTRY_FILL_SRA_SINGLE_INST_REG, &CompileSingleStep);
// Enter JIT
(void)b(&LoopTop);
b(&LoopTop);
AbsoluteLoopTopAddressEnterEC = GetCursorAddress<uint64_t>();
// Load ThreadState and write the target PC there
@@ -130,7 +130,7 @@ void Dispatcher::EmitDispatcher() {
ldp<ARMEmitter::IndexType::OFFSET>(TMP1, TMP2, REG_CALLRET_SP);
// EC_CALL_CHECKER_PC_REG is REG_PF which isn't touched by any of the above
sub(ARMEmitter::Size::i64Bit, TMP1, EC_CALL_CHECKER_PC_REG, TMP1);
(void)cbnz(ARMEmitter::Size::i64Bit, TMP1, &LoopTop);
cbnz(ARMEmitter::Size::i64Bit, TMP1, &LoopTop);
// If the entry at the TOS is for the target address, pop it and return to the JIT code
add(ARMEmitter::Size::i64Bit, REG_CALLRET_SP, REG_CALLRET_SP, 0x10);
+7 -9
View File
@@ -493,7 +493,7 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
}
}
static void DirectBlockDelinker(FEXCore::Core::CpuStateFrame* Frame, FEXCore::Context::ExitFunctionLinkData* Record, bool Call) {
static void DirectBlockDelinker(FEXCore::Context::ExitFunctionLinkData* Record, bool Call) {
uintptr_t JumpThunkStartAddress = reinterpret_cast<uintptr_t>(Record) - 0x10;
uintptr_t CallerAddress = JumpThunkStartAddress + Record->CallerOffset;
auto BranchOffset = JumpThunkStartAddress / 4 - CallerAddress / 4;
@@ -511,7 +511,7 @@ static void DirectBlockDelinker(FEXCore::Core::CpuStateFrame* Frame, FEXCore::Co
ARMEmitter::Emitter::ClearICache(reinterpret_cast<void*>(CallerAddress), 4);
}
static void IndirectBlockDelinker(FEXCore::Core::CpuStateFrame* Frame, FEXCore::Context::ExitFunctionLinkData* Record) {
static void IndirectBlockDelinker(FEXCore::Context::ExitFunctionLinkData* Record) {
uintptr_t JumpThunkStartAddress = reinterpret_cast<uintptr_t>(Record) - 0x10;
uint32_t BranchInst = 0;
ARMEmitter::Emitter BranchEmit(reinterpret_cast<uint8_t*>(&BranchInst), 4);
@@ -578,13 +578,13 @@ uint64_t Arm64JITCore::ExitFunctionLink(FEXCore::Core::CpuStateFrame* Frame, FEX
BranchEmit.bl(BranchOffset);
Thread->LookupCache->AddBlockLink(
GuestRip, Record,
[](FEXCore::Core::CpuStateFrame* Frame, FEXCore::Context::ExitFunctionLinkData* Record) { DirectBlockDelinker(Frame, Record, true); }, lk);
[](FEXCore::Context::ExitFunctionLinkData* Record) { DirectBlockDelinker(Record, true); }, lk);
} else {
BranchEmit.b(BranchOffset);
Thread->LookupCache->AddBlockLink(
GuestRip, Record,
[](FEXCore::Core::CpuStateFrame* Frame, FEXCore::Context::ExitFunctionLinkData* Record) {
DirectBlockDelinker(Frame, Record, false);
[](FEXCore::Context::ExitFunctionLinkData* Record) {
DirectBlockDelinker(Record, false);
},
lk);
}
@@ -787,7 +787,7 @@ void Arm64JITCore::EmitSuspendInterruptCheck() {
ldr(TMP2.W(), STATE_PTR(CpuStateFrame, SuspendDoorbell));
ARMEmitter::ForwardLabel l_NoSuspend;
(void)cbz(ARMEmitter::Size::i32Bit, TMP2, &l_NoSuspend);
cbz(ARMEmitter::Size::i32Bit, TMP2, &l_NoSuspend);
brk(SuspendMagic);
(void)Bind(&l_NoSuspend);
#endif
@@ -825,9 +825,7 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
this->IR = IR;
RequiresFarARM64Jumps = false;
// Prepare restart via long jump in case branch encoding fails.
// This uses UncheckedLongJump since we don't implement std::longjmp in WoA setups
switch (static_cast<RestartOptions::Control>(FEXCore::UncheckedLongJump::SetJump(RestartControl.RestartJump))) {
switch (static_cast<RestartOptions::Control>(FEXCore::LongJump::SetJump(RestartControl.RestartJump))) {
case RestartOptions::Control::Incoming:
// Nothing
break;
+11 -11
View File
@@ -68,7 +68,7 @@ private:
const bool HostSupportsAFP {};
struct RestartOptions {
FEXCore::UncheckedLongJump::JumpBuf RestartJump;
FEXCore::LongJump::JumpBuf RestartJump;
enum class Control : uint64_t {
Incoming = 0,
EnableFarARM64Jumps = 1,
@@ -360,7 +360,7 @@ private:
// We can support this but currently unnecessary.
ERROR_AND_DIE_FMT("Tried to branch larger than 128MB away!");
FEXCore::UncheckedLongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
FEXCore::LongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
template<ARMEmitter::IsLabel T>
@@ -371,7 +371,7 @@ private:
// We can support this but currently unnecessary.
ERROR_AND_DIE_FMT("Tried to branch larger than 128MB away!");
FEXCore::UncheckedLongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
FEXCore::LongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
template<ARMEmitter::IsLabel T>
@@ -392,7 +392,7 @@ private:
return;
}
FEXCore::UncheckedLongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
FEXCore::LongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
template<ARMEmitter::IsLabel T>
@@ -413,7 +413,7 @@ private:
return;
}
FEXCore::UncheckedLongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
FEXCore::LongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
template<ARMEmitter::IsLabel T>
@@ -434,7 +434,7 @@ private:
return;
}
FEXCore::UncheckedLongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
FEXCore::LongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
template<ARMEmitter::IsLabel T>
@@ -455,7 +455,7 @@ private:
return;
}
FEXCore::UncheckedLongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
FEXCore::LongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
template<ARMEmitter::IsLabel T>
@@ -476,7 +476,7 @@ private:
return;
}
FEXCore::UncheckedLongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
FEXCore::LongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
template<ARMEmitter::IsLabel T>
@@ -487,7 +487,7 @@ private:
// We can support this but currently unnecessary.
ERROR_AND_DIE_FMT("Long ADR currently unsupported!");
FEXCore::UncheckedLongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
FEXCore::LongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
template<ARMEmitter::IsLabel T>
@@ -498,7 +498,7 @@ private:
// We can support this but currently unnecessary.
ERROR_AND_DIE_FMT("Long ADRP currently unsupported!");
FEXCore::UncheckedLongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
FEXCore::LongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
template<ARMEmitter::IsLabel T>
@@ -513,7 +513,7 @@ private:
return;
}
FEXCore::UncheckedLongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
FEXCore::LongJump::LongJump(RestartControl.RestartJump, FEXCore::ToUnderlying(RestartOptions::Control::EnableFarARM64Jumps));
}
// This is purely a debugging aid for developers to see if they are in JIT code space when inspecting raw memory
@@ -39,8 +39,6 @@ LookupCache::LookupCache(FEXCore::Context::ContextImpl* CTX)
// We need one pointer per page of virtual memory
// At 64GB of virtual memory this will allocate 128MB of virtual memory space
PagePointer = reinterpret_cast<uintptr_t>(FEXCore::Allocator::VirtualAlloc(TotalCacheSize, false, false));
LOGMAN_THROW_A_FMT(PagePointer != -1ULL, "Failed to allocate PagePointer");
FEXCore::Allocator::VirtualName("FEXMem_Lookup", reinterpret_cast<void*>(PagePointer),
ctx->Config.VirtualMemSize / FEXCore::Utils::FEX_PAGE_SIZE * 8 + CODE_SIZE);
CTX->SyscallHandler->MarkOvercommitRange(PagePointer, TotalCacheSize);
@@ -51,11 +49,14 @@ LookupCache::LookupCache(FEXCore::Context::ContextImpl* CTX)
// We currently limit to 128MB of real memory for caching for the total cache size.
// Can end up being inefficient if we compile a small number of blocks per page
PageMemory = PagePointer + ctx->Config.VirtualMemSize / FEXCore::Utils::FEX_PAGE_SIZE * 8;
LOGMAN_THROW_A_FMT(PageMemory != -1ULL, "Failed to allocate page memory");
// L1 Cache
L1Pointer = PageMemory + CODE_SIZE;
FEXCore::Allocator::VirtualName("FEXMem_Lookup_L1", reinterpret_cast<void*>(L1Pointer), MAX_L1_SIZE);
LOGMAN_THROW_A_FMT(L1Pointer != -1ULL, "Failed to allocate L1Pointer");
VirtualMemSize = ctx->Config.VirtualMemSize;
if (DynamicL1Cache()) {
@@ -75,7 +76,7 @@ LookupCache::~LookupCache() {
// These will get freed when their memory allocators are deallocated.
}
void LookupCache::ClearL2Cache(const FEXCore::LookupCacheReadLockToken& lk) {
void LookupCache::ClearL2Cache(const FEXCore::LookupCacheWriteLockToken& lk) {
// Clear out the page memory
// PagePointer and PageMemory are sequential with each other. Clear both at once.
FEXCore::Allocator::VirtualDontNeed(reinterpret_cast<void*>(PagePointer),
@@ -86,12 +87,12 @@ void LookupCache::ClearL2Cache(const FEXCore::LookupCacheReadLockToken& lk) {
void LookupCache::ClearThreadLocalCaches(const LookupCacheWriteLockToken&) {
// Clear L1 and L2 by clearing the full cache.
FEXCore::Allocator::VirtualDontNeed(reinterpret_cast<void*>(PagePointer), TotalCacheSize, false);
CachedCodePages.clear();
}
void LookupCache::ClearCache(const LookupCacheWriteLockToken& lk) {
// Clear L1 and L2 by clearing the full cache.
FEXCore::Allocator::VirtualDontNeed(reinterpret_cast<void*>(PagePointer), TotalCacheSize, false);
ClearThreadLocalCaches(lk);
Shared->ClearCache(lk);
}
+73 -53
View File
@@ -3,12 +3,11 @@
#include "Interface/Context/Context.h"
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/SHMStats.h>
#include "Utils/WritePriorityMutex.h"
#include <FEXCore/fextl/map.h>
#include <FEXCore/fextl/memory_resource.h>
#include <FEXCore/fextl/robin_map.h>
#include <FEXCore/fextl/vector.h>
#include <FEXCore/fextl/unordered_set.h>
#include <FEXCore/fextl/memory_resource.h>
#include <cstdint>
@@ -17,35 +16,22 @@
#include <mutex>
namespace FEXCore {
struct LookupCacheWriteLockToken {
private:
// Only constructible by GuestToHostMap
friend struct GuestToHostMap;
LookupCacheWriteLockToken(FEXCore::Utils::WritePriorityMutex::Mutex& Mutex)
LookupCacheWriteLockToken(std::mutex& Mutex)
: Lock {Mutex} {}
std::lock_guard<FEXCore::Utils::WritePriorityMutex::Mutex> Lock;
};
struct LookupCacheReadLockToken {
private:
// Only constructible by GuestToHostMap
friend struct GuestToHostMap;
LookupCacheReadLockToken(FEXCore::Utils::WritePriorityMutex::Mutex& Mutex)
: Lock {Mutex} {}
std::shared_lock<FEXCore::Utils::WritePriorityMutex::Mutex> Lock;
std::lock_guard<std::mutex> Lock;
};
struct GuestToHostMap {
FEXCore::Utils::WritePriorityMutex::Mutex Lock {};
std::mutex WriteLock;
[[nodiscard]]
LookupCacheWriteLockToken AcquireWriteLock() {
return LookupCacheWriteLockToken {Lock};
}
[[nodiscard]]
LookupCacheReadLockToken AcquireReadLock() {
return LookupCacheReadLockToken {Lock};
return LookupCacheWriteLockToken {WriteLock};
}
struct BlockLinkTag {
@@ -74,42 +60,61 @@ struct GuestToHostMap {
fextl::unique_ptr<std::pmr::polymorphic_allocator<std::byte>> BlockLinks_pma;
BlockLinksMapType* BlockLinks;
fextl::robin_map<uint64_t, uint64_t> BlockList;
struct BlockEntry {
uint64_t HostCode;
fextl::vector<uint64_t> CodePages;
};
fextl::robin_map<uint64_t, BlockEntry> BlockList;
fextl::map<uint64_t, fextl::vector<uint64_t>> CodePages;
GuestToHostMap();
// Adds to Guest -> Host code mapping
void AddBlockMapping(uint64_t Address, void* HostCode, const LookupCacheWriteLockToken&) {
const BlockEntry& AddBlockMapping(uint64_t Address, const fextl::vector<uint64_t>& CodePages, void* HostCode, const LookupCacheWriteLockToken&) {
// This may replace an existing mapping
// NOTE: Generally no previous entry should exist, however there is one exception:
// If the backend updates the active thread's CodeBuffer, the new associated LookupCache
// may already contain the block address. Since is comparatively rare, we'll just leak
// one of the two blocks in this case.
BlockList[Address] = (uintptr_t)HostCode;
return BlockList.insert_or_assign(Address, BlockEntry {(uintptr_t)HostCode, CodePages}).first->second;
}
std::optional<uintptr_t> FindBlock(uint64_t Address, const LookupCacheReadLockToken&) {
const BlockEntry* FindBlock(uint64_t Address, const LookupCacheWriteLockToken&) {
auto HostCode = BlockList.find(Address);
if (HostCode == BlockList.end()) {
return std::nullopt;
return nullptr;
}
return HostCode->second;
return &HostCode->second;
}
bool Erase(FEXCore::Core::CpuStateFrame* Frame, uint64_t Address, const LookupCacheWriteLockToken&) {
bool Erase(uint64_t Address, const LookupCacheWriteLockToken&) {
// Sever any links to this block
auto lower = BlockLinks->lower_bound({Address, nullptr});
auto upper = BlockLinks->upper_bound({Address, reinterpret_cast<FEXCore::Context::ExitFunctionLinkData*>(UINTPTR_MAX)});
for (auto it = lower; it != upper; it = BlockLinks->erase(it)) {
it->second(Frame, it->first.HostLink);
it->second(it->first.HostLink);
}
// Remove from BlockList
return BlockList.erase(Address) != 0;
}
void InvalidateRange(uint64_t Start, uint64_t Length) {
auto lk = AcquireWriteLock();
auto lower = CodePages.lower_bound(Start >> 12);
auto upper = CodePages.upper_bound((Start + Length - 1) >> 12);
for (auto it = lower; it != upper; it++) {
for (const auto& Entry : it->second) {
Erase(Entry, lk);
}
}
CodePages.erase(lower, upper);
}
void AddBlockLink(uint64_t GuestDestination, FEXCore::Context::ExitFunctionLinkData* HostLink,
const FEXCore::Context::BlockDelinkerFunc& delinker, const LookupCacheWriteLockToken&) {
BlockLinks->insert({{GuestDestination, HostLink}, delinker});
@@ -159,7 +164,7 @@ public:
{
std::optional<FEXCore::SHMStats::AccumulationBlock<uint64_t>> LockTime(
Thread->ThreadStats ? &Thread->ThreadStats->AccumulatedCacheReadLockTime : nullptr);
auto lk = Shared->AcquireReadLock();
auto lk = Shared->AcquireWriteLock();
LockTime.reset();
if (!DisableL2Cache()) {
@@ -185,10 +190,10 @@ public:
if (!HostPtr) {
// Try L3
auto HostCode = Shared->FindBlock(Address, lk);
if (HostCode) {
CacheBlockMapping(Address, HostCode.value(), lk);
HostPtr = HostCode.value();
auto Entry = Shared->FindBlock(Address, lk);
if (Entry) {
CacheBlockMapping(Address, *Entry, false, lk);
HostPtr = Entry->HostCode;
}
}
}
@@ -259,32 +264,25 @@ public:
}
// Adds to Guest -> Host code mapping
void AddBlockMapping(FEXCore::Core::InternalThreadState* Thread, uint64_t Address, void* HostCode) {
void AddBlockMapping(FEXCore::Core::InternalThreadState* Thread, uint64_t Address, const fextl::vector<uint64_t>& CodePages, void* HostCode) {
std::optional<FEXCore::SHMStats::AccumulationBlock<uint64_t>> LockTime(
Thread->ThreadStats ? &Thread->ThreadStats->AccumulatedCacheWriteLockTime : nullptr);
auto lk = Shared->AcquireWriteLock();
LockTime.reset();
Shared->AddBlockMapping(Address, HostCode, lk);
const auto& Entry = Shared->AddBlockMapping(Address, CodePages, HostCode, lk);
// There is no need to update L1 or L2, they will get updated on first lookup
// However, adding to L1 here increases performance
auto& L1Entry = reinterpret_cast<LookupCacheEntry*>(L1Pointer)[Address & L1PointerMask];
L1Entry.GuestCode = Address;
L1Entry.HostCode = (uintptr_t)HostCode;
CacheBlockMapping(Address, Entry, true, lk);
}
// NOTE: It's the caller's responsibility to call Erase() for all other
// GuestToHostMaps that share the same LookupCache. Otherwise, the
// L1/L2 caches will contain stale references to deallocated memory.
bool Erase(FEXCore::Core::CpuStateFrame* Frame, uint64_t Address, const LookupCacheWriteLockToken& lk) {
bool ErasedAny = Shared->Erase(Frame, Address, lk);
// Invalidates L1/L2 for a given guest block
void InvalidateCache(uint64_t Address, const LookupCacheWriteLockToken& lk) {
// Do L1
auto& L1Entry = reinterpret_cast<LookupCacheEntry*>(L1Pointer)[Address & L1PointerMask];
if (L1Entry.GuestCode == Address) {
L1Entry.GuestCode = 0;
ErasedAny = true;
// Leave L1Entry.HostCode as is, so that concurrent lookups won't read a null pointer
// This is a soft guarantee for cross thread invalidation, as atomics are not used
// and it hasn't been thoroughly tested
@@ -300,7 +298,7 @@ public:
uint64_t LocalPagePointer = Pointers[Address];
if (!LocalPagePointer) {
// Page for this code didn't even exist, nothing to do
return ErasedAny;
return;
}
// Page exists, just set the offset to zero
@@ -308,7 +306,22 @@ public:
BlockPointers[PageOffset].GuestCode = 0;
BlockPointers[PageOffset].HostCode = 0;
}
return true;
}
// Invalidates all L1/L2 entries for all guest block that intersect the given range
bool InvalidateCacheRange(uint64_t Start, uint64_t Length) {
auto lk = Shared->AcquireWriteLock();
auto lower = CachedCodePages.lower_bound(Start >> 12);
auto upper = CachedCodePages.upper_bound((Start + Length - 1) >> 12);
for (auto it = lower; it != upper; it++) {
for (const auto& Entry : it->second) {
InvalidateCache(Entry, lk);
}
}
CachedCodePages.erase(lower, upper);
return upper != lower;
}
void AddBlockLink(uint64_t GuestDestination, FEXCore::Context::ExitFunctionLinkData* HostLink,
@@ -317,7 +330,7 @@ public:
}
void ClearCache(const LookupCacheWriteLockToken&);
void ClearL2Cache(const LookupCacheReadLockToken&);
void ClearL2Cache(const LookupCacheWriteLockToken&);
void ClearThreadLocalCaches(const LookupCacheWriteLockToken&);
uintptr_t GetL1Pointer() const {
@@ -345,13 +358,17 @@ public:
}
private:
void CacheBlockMapping(uint64_t Address, uintptr_t HostCode, const LookupCacheReadLockToken& lk) {
void CacheBlockMapping(uint64_t Address, const GuestToHostMap::BlockEntry& Entry, bool L1Only, const LookupCacheWriteLockToken& lk) {
for (const auto& CodePage : Entry.CodePages) {
CachedCodePages[CodePage >> 12].insert(Address);
}
// Do L1
auto& L1Entry = reinterpret_cast<LookupCacheEntry*>(L1Pointer)[Address & L1PointerMask];
L1Entry.GuestCode = Address;
L1Entry.HostCode = HostCode;
L1Entry.HostCode = Entry.HostCode;
if (!DisableL2Cache()) {
if (!DisableL2Cache() && !L1Only) {
// Do ful map
auto FullAddress = Address;
Address = Address & (VirtualMemSize - 1);
@@ -368,7 +385,7 @@ private:
if (!NewPageBacking) {
// Couldn't allocate, clear L2 and retry
ClearL2Cache(lk);
CacheBlockMapping(Address, HostCode, lk);
CacheBlockMapping(Address, Entry, false, lk);
return;
}
Pointers[Address] = NewPageBacking;
@@ -380,7 +397,7 @@ private:
// This silently replaces existing mappings
BlockPointers[PageOffset].GuestCode = FullAddress;
BlockPointers[PageOffset].HostCode = HostCode;
BlockPointers[PageOffset].HostCode = Entry.HostCode;
}
}
@@ -398,6 +415,9 @@ private:
return PageMemory + NewBase;
}
// Maps from a page index to all blocks in the page that have at some point been fetched into L1/L2
fextl::map<uint64_t, fextl::unordered_set<uint64_t>> CachedCodePages;
uintptr_t PagePointer;
uintptr_t PageMemory;
uintptr_t L1Pointer;
@@ -17,6 +17,7 @@ $end_info$
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/FPState.h>
#include <cmath>
#include <stddef.h>
#include <stdint.h>
@@ -68,7 +69,7 @@ void OpDispatchBuilder::FLD(OpcodeArgs, IR::OpSize Width) {
if (Width == OpSize::i32Bit || Width == OpSize::i64Bit) {
ConvertedData = _F80CVTTo(Data, ReadWidth);
}
_PushStack(ConvertedData, Data, ReadWidth);
_PushStack(ConvertedData, Data, ReadWidth, true);
}
// Float LoaD operation with memory operand
@@ -80,7 +81,7 @@ void OpDispatchBuilder::FBLD(OpcodeArgs) {
// Read from memory
Ref Data = LoadSourceFPR_WithOpSize(Op, Op->Src[0], OpSize::f80Bit, Op->Flags);
Ref ConvertedData = _F80BCDLoad(Data);
_PushStack(ConvertedData, Data, OpSize::i128Bit);
_PushStack(ConvertedData, Data, OpSize::i128Bit, true);
}
void OpDispatchBuilder::FBSTP(OpcodeArgs) {
@@ -92,7 +93,7 @@ void OpDispatchBuilder::FBSTP(OpcodeArgs) {
void OpDispatchBuilder::FLD_Const(OpcodeArgs, NamedVectorConstant K) {
// Update TOP
Ref Data = LoadAndCacheNamedVectorConstant(OpSize::i128Bit, K);
_PushStack(Data, Data, OpSize::i128Bit);
_PushStack(Data, Data, OpSize::i128Bit, true);
}
void OpDispatchBuilder::FILD(OpcodeArgs) {
@@ -123,7 +124,7 @@ void OpDispatchBuilder::FILD(OpcodeArgs) {
auto upper = _Or(OpSize::i64Bit, sign, zeroed_exponent);
Ref ConvertedData = _VLoadTwoGPRs(shifted, upper);
_PushStack(ConvertedData, Invalid(), ReadWidth);
_PushStack(ConvertedData, Data, ReadWidth, false);
}
void OpDispatchBuilder::FST(OpcodeArgs, IR::OpSize Width) {
@@ -131,7 +132,7 @@ void OpDispatchBuilder::FST(OpcodeArgs, IR::OpSize Width) {
AddressMode A = DecodeAddress(Op, Op->Dest, MemoryAccessType::DEFAULT, false);
A = SelectAddressMode(this, A, GetGPROpSize(), CTX->HostFeatures.SupportsTSOImm9, false, false, Width);
_StoreStackMem(SourceSize, Width, A.Base, A.Index, OpSize::iInvalid, A.IndexType, A.IndexScale);
_StoreStackMem(SourceSize, Width, A.Base, A.Index, OpSize::iInvalid, A.IndexType, A.IndexScale, /*Float=*/true);
if (Op->TableInfo->Flags & X86Tables::InstFlags::FLAGS_POP) {
_PopStackDestroy();
@@ -877,8 +878,8 @@ void OpDispatchBuilder::X87FXTRACT(OpcodeArgs) {
_PopStackDestroy();
auto Exp = _F80XTRACT_EXP(Top);
auto Sig = _F80XTRACT_SIG(Top);
_PushStack(Exp, Invalid(), OpSize::f80Bit);
_PushStack(Sig, Invalid(), OpSize::f80Bit);
_PushStack(Exp, Exp, OpSize::f80Bit, true);
_PushStack(Sig, Sig, OpSize::f80Bit, true);
}
} // namespace FEXCore::IR
@@ -68,7 +68,7 @@ void OpDispatchBuilder::FLDF64(OpcodeArgs, IR::OpSize Width) {
} else if (Width == OpSize::f80Bit) {
ConvertedData = _F80CVT(OpSize::i64Bit, Data);
}
_PushStack(ConvertedData, Data, ReadWidth);
_PushStack(ConvertedData, Data, ReadWidth, true);
}
void OpDispatchBuilder::FBLDF64(OpcodeArgs) {
@@ -76,7 +76,7 @@ void OpDispatchBuilder::FBLDF64(OpcodeArgs) {
Ref Data = LoadSourceFPR_WithOpSize(Op, Op->Src[0], OpSize::f80Bit, Op->Flags);
Ref ConvertedData = _F80BCDLoad(Data);
ConvertedData = _F80CVT(OpSize::i64Bit, ConvertedData);
_PushStack(ConvertedData, Data, OpSize::i64Bit);
_PushStack(ConvertedData, Data, OpSize::i64Bit, true);
}
void OpDispatchBuilder::FBSTPF64(OpcodeArgs) {
@@ -88,7 +88,7 @@ void OpDispatchBuilder::FBSTPF64(OpcodeArgs) {
void OpDispatchBuilder::FLDF64_Const(OpcodeArgs, uint64_t Num) {
auto Data = _VCastFromGPR(OpSize::i64Bit, OpSize::i64Bit, Constant(Num));
_PushStack(Data, Data, OpSize::i64Bit);
_PushStack(Data, Data, OpSize::i64Bit, true);
}
void OpDispatchBuilder::FILDF64(OpcodeArgs) {
@@ -100,7 +100,7 @@ void OpDispatchBuilder::FILDF64(OpcodeArgs) {
Data = _Sbfe(OpSize::i64Bit, IR::OpSizeAsBits(ReadWidth), 0, Data);
}
auto ConvertedData = _Float_FromGPR_S(OpSize::i64Bit, ReadWidth == OpSize::i32Bit ? OpSize::i32Bit : OpSize::i64Bit, Data);
_PushStack(ConvertedData, Invalid(), ReadWidth);
_PushStack(ConvertedData, Data, ReadWidth, false);
}
void OpDispatchBuilder::FISTF64(OpcodeArgs, bool Truncate) {
@@ -397,7 +397,7 @@ void OpDispatchBuilder::X87FXTRACTF64(OpcodeArgs) {
Ref Exp = _NZCVSelectV(OpSize::i64Bit, CondClass::EQ, ExpZV, ExpNZV);
_PopStackDestroy();
_PushStack(Exp, Invalid(), OpSize::i64Bit);
_PushStack(Sig, Invalid(), OpSize::i64Bit);
_PushStack(Exp, Exp, OpSize::i64Bit, true);
_PushStack(Sig, Sig, OpSize::i64Bit, true);
}
} // namespace FEXCore::IR
+10 -5
View File
@@ -2818,13 +2818,17 @@
"X87": true,
"HasSideEffects": true
},
"PushStack FPR:$X80Src, FPR:$OriginalValue, OpSize:$LoadSize": {
"PushStack FPR:$X80Src, SSA:$OriginalValue, OpSize:$LoadSize, i1:$Float": {
"Desc": [
"Pushes the provided X80Src source on to the x87 stack.",
"Tracks OriginalValue as the original value of X80Src. OriginalValue can be Invalid() in which case no tracking is done.",
"Tracks OriginalValue as the original value of X80Src.",
"Opsize is 128bit for F80 values, 64-bit for low precision.",
"LoadSize the original load size, i.e. of size of OriginalValue.",
"Float: 80-bit, 64-bit, 32-bit"
"Float: 80-bit, 64-bit, 32-bit",
"Int: 64-bit, 32-bit, 16-bit"
],
"EmitValidation": [
"WalkFindRegClass($OriginalValue) == RegClass::FPR || WalkFindRegClass($OriginalValue) == RegClass::GPR"
],
"HasSideEffects": true,
"X87": true
@@ -2836,12 +2840,13 @@
"HasSideEffects": true,
"X87": true
},
"StoreStackMem OpSize:$SourceSize, OpSize:$StoreSize, GPR:$Addr, GPR:$Offset, OpSize:$Align, MemOffsetType:$OffsetType, u8:$OffsetScale": {
"StoreStackMem OpSize:$SourceSize, OpSize:$StoreSize, GPR:$Addr, GPR:$Offset, OpSize:$Align, MemOffsetType:$OffsetType, u8:$OffsetScale, i1:$Float": {
"Desc": [
"Takes the top value off the x87 stack and stores it to memory.",
"SourceSize is 128bit for F80 values, 64-bit for low precision.",
"StoreSize is the store size for conversion:",
"Float: 80-bit, 64-bit, or 32-bit"
"Float: 80-bit, 64-bit, or 32-bit",
"Int: 64-bit, 32-bit, 16-bit"
],
"HasSideEffects": true,
"X87": true
@@ -6,6 +6,7 @@
#include "Interface/IR/PassManager.h"
#include "FEXCore/IR/IR.h"
#include "FEXCore/Utils/Profiler.h"
#include "FEXCore/Utils/MathUtils.h"
#include "FEXCore/Core/HostFeatures.h"
#include "Interface/Core/Addressing.h"
@@ -292,9 +293,10 @@ private:
StackMemberInfo() {}
StackMemberInfo(Ref Data)
: StackDataNode(Data) {}
StackMemberInfo(Ref Data, Ref Source, OpSize Size)
StackMemberInfo(Ref Data, Ref Source, OpSize Size, bool Float)
: StackDataNode(Data)
, Source({Size, Source}) {}
, Source({Size, Source})
, InterpretAsFloat(Float) {}
Ref StackDataNode {}; // Reference to the data in the Stack.
// This is the source data node in the stack format, possibly converted to 64/80 bits.
struct StackMemberData final {
@@ -304,6 +306,7 @@ private:
// Tuple is only valid if we have information about the Source of the Stack Data Node.
// In it's valid then OpSize is the original source size and Ref is the original source node.
std::optional<StackMemberData> Source {};
bool InterpretAsFloat {false}; // True if this is a floating point value, false if integer
};
// StackData, TopCache need to be always properly set to ensure
@@ -924,13 +927,8 @@ void X87StackOptimization::Run(IREmitter* Emit) {
StoreStackValueAtOffset_Slow(SourceNode);
} else {
auto* SourceNode = CurrentIR.GetNode(Op->X80Src);
if (Op->OriginalValue.IsInvalid()) {
// No original value to track - just push the converted data
StackData.push(StackMemberInfo {SourceNode});
} else {
auto* OriginalNode = CurrentIR.GetNode(Op->OriginalValue);
StackData.push(StackMemberInfo {SourceNode, OriginalNode, Op->LoadSize});
}
auto* OriginalNode = CurrentIR.GetNode(Op->OriginalValue);
StackData.push(StackMemberInfo {SourceNode, OriginalNode, Op->LoadSize, Op->Float});
}
break;
}
@@ -995,8 +993,9 @@ void X87StackOptimization::Run(IREmitter* Emit) {
// str w2, [x1]
// or similar. As long as the source size and dest size are one and the same.
// This will avoid any conversions between source and stack element size and conversion back.
if (!SlowPath && Value->Source && Value->Source->Size == Op->StoreSize) {
IREmit->_StoreMemFPR(Op->StoreSize, Value->Source->Node, AddrNode, Offset, Align, OffsetType, OffsetScale);
if (!SlowPath && Value->Source && Value->Source->Size == Op->StoreSize && Value->InterpretAsFloat) {
const auto ClassType = Value->InterpretAsFloat ? RegClass::FPR : RegClass::GPR;
IREmit->_StoreMem(ClassType, Op->StoreSize, Value->Source->Node, AddrNode, Offset, Align, OffsetType, OffsetScale);
break;
}
@@ -1036,26 +1035,11 @@ void X87StackOptimization::Run(IREmitter* Emit) {
case OP_F80STACKXCHANGE: {
const auto* Op = IROp->C<IROp_F80StackXchange>();
auto Offset = Op->SrcStack;
Ref ValueTop = LoadStackValue();
Ref ValueOffset = LoadStackValue(Offset);
if (Offset == 0) {
// No-op
break;
}
const auto [ValidTop, StackMemberTop] = StackData.top(0);
const auto [ValidOffset, StackMemberOffset] = StackData.top(Offset);
if (ValidTop != StackSlot::VALID || ValidOffset != StackSlot::VALID) {
// Slow path: do actual memory operations
Ref ValueTop = LoadStackValue();
Ref ValueOffset = LoadStackValue(Offset);
StoreStackValue(ValueOffset);
StoreStackValue(ValueTop, Offset);
} else {
// Fast path: swap complete StackMemberInfo preserving Source metadata
StackData.setTop(StackMemberOffset, 0);
StackData.setTop(StackMemberTop, Offset);
}
StoreStackValue(ValueOffset);
StoreStackValue(ValueTop, Offset);
break;
}
+4 -1
View File
@@ -140,7 +140,10 @@ void ClearHooks() {
FEXCore::Allocator::mmap = ::mmap;
FEXCore::Allocator::munmap = ::munmap;
Alloc::OSAllocator::ReleaseAllocatorWorkaround(std::move(Alloc64));
// XXX: This is currently a leak.
// We can't work around this yet until static initializers that allocate memory are completely removed from our codebase
// Luckily we only remove this on process shutdown, so the kernel will do the cleanup for us
Alloc64.release();
}
#pragma GCC diagnostic pop
@@ -207,7 +207,7 @@ OSAllocator_64Bit::LiveVMARegion* OSAllocator_64Bit::FindLiveRegionForAddress(ui
uintptr_t RegionBegin = (*it)->SlabInfo->Base;
uintptr_t RegionEnd = RegionBegin + (*it)->SlabInfo->RegionSize;
if (Addr >= RegionBegin && AddrEnd < RegionEnd) {
if (Addr >= RegionBegin && Addr < RegionEnd) {
LiveRegion = *it;
// Leave our loop
break;
@@ -405,18 +405,14 @@ again:
// Mark the pages as used
uintptr_t RegionBegin = LiveRegion->SlabInfo->Base;
uintptr_t MappedBegin = (AllocatedOffset - RegionBegin) >> FEXCore::Utils::FEX_PAGE_SHIFT;
size_t PagesSet {};
for (size_t i = 0; i < NumberOfPages; ++i) {
PagesSet += LiveRegion->UsedPages.TestAndSet(MappedBegin + i) == false;
LiveRegion->UsedPages.Set(MappedBegin + i);
}
// Change our last allocation region
LiveRegion->LastPageAllocation = MappedBegin + NumberOfPages;
LiveRegion->FreeSpace -= PagesSet * FEXCore::Utils::FEX_PAGE_SIZE;
LOGMAN_THROW_A_FMT(LiveRegion->FreeSpace <= LiveRegion->SlabInfo->RegionSize,
"Corrupt LiveRegion free space! 0x{:x} > 0x{:x}. After allocating 0x{:x} (0x{:x} overlapped)", LiveRegion->FreeSpace,
LiveRegion->SlabInfo->RegionSize, length, PagesSet);
LiveRegion->FreeSpace -= length;
}
if (!AllocatedOffset) {
+6 -20
View File
@@ -27,11 +27,6 @@ struct FlexBitSet final {
Memory[Element / MinimumSizeBits] &= ~(1ULL << (Element % MinimumSizeBits));
return Value;
}
bool TestAndSet(size_t Element) {
bool Value = Get(Element);
Memory[Element / MinimumSizeBits] |= (1ULL << (Element % MinimumSizeBits));
return Value;
}
void Set(size_t Element) {
Memory[Element / MinimumSizeBits] |= (1ULL << (Element % MinimumSizeBits));
}
@@ -75,17 +70,12 @@ struct FlexBitSet final {
template<bool WantUnset>
BitsetScanResults BackwardScanForRange(size_t BeginningElement, size_t ElementCount, size_t MinimumElement) {
bool FoundHole {};
// Final element to iterate to.
const size_t FinalElement = MinimumElement + ElementCount - 1;
for (size_t CurrentPage = BeginningElement; CurrentPage >= FinalElement;) {
for (size_t CurrentPage = BeginningElement; CurrentPage >= (MinimumElement + ElementCount);) {
size_t Remaining = ElementCount;
LOGMAN_THROW_A_FMT(CurrentPage <= BeginningElement && CurrentPage >= FinalElement, "BackwardScanForRange: Scanning less than "
"available range");
LOGMAN_THROW_A_FMT(Remaining <= CurrentPage, "Scanning less than available range");
while (Remaining) {
if (this->Get(CurrentPage - Remaining + 1) == WantUnset) {
if (this->Get(CurrentPage - Remaining) == WantUnset) {
// Has an intersecting range
break;
}
@@ -102,7 +92,7 @@ struct FlexBitSet final {
CurrentPage -= Remaining;
} else {
// We have a slab range
return BitsetScanResults {CurrentPage - ElementCount + 1, FoundHole};
return BitsetScanResults {CurrentPage - ElementCount, FoundHole};
}
}
@@ -118,15 +108,11 @@ struct FlexBitSet final {
BitsetScanResults ForwardScanForRange(size_t BeginningElement, size_t ElementCount, size_t ElementsInSet) {
bool FoundHole {};
// Final element to iterate to.
const size_t FinalElement = ElementsInSet - ElementCount + 1;
for (size_t CurrentElement = BeginningElement; CurrentElement <= FinalElement;) {
for (size_t CurrentElement = BeginningElement; CurrentElement < (ElementsInSet - ElementCount);) {
// If we have enough free space, check if we have enough free pages that are contiguous
size_t Remaining = ElementCount;
LOGMAN_THROW_A_FMT(CurrentElement >= BeginningElement && CurrentElement <= FinalElement, "ForwardScanForRange: Scanning less than "
"available range");
LOGMAN_THROW_A_FMT((CurrentElement + Remaining - 1) < ElementsInSet, "Scanning less than available range");
while (Remaining) {
if (this->Get(CurrentElement + Remaining - 1) == WantUnset) {
@@ -53,12 +53,4 @@ public:
namespace Alloc::OSAllocator {
fextl::unique_ptr<Alloc::HostAllocator> Create64BitAllocator();
fextl::unique_ptr<Alloc::HostAllocator> Create64BitAllocatorWithRegions(fextl::vector<FEXCore::Allocator::MemoryRegion>& Regions);
static inline void ReleaseAllocatorWorkaround(fextl::unique_ptr<Alloc::HostAllocator> Allocator) {
// XXX: This is currently a leak.
// We can't work around this yet until static initializers that allocate memory are completely removed from our codebase
// The allocator is also intrusively allocated, so the unique_ptr tries to double free the HostAllocator object.
// Luckily we only remove this on process shutdown, so the kernel will do the cleanup for us
Allocator.release();
}
} // namespace Alloc::OSAllocator
+2 -2
View File
@@ -1,7 +1,7 @@
// SPDX-License-Identifier: MIT
#include <FEXCore/Utils/LongJump.h>
namespace FEXCore::UncheckedLongJump {
namespace FEXCore::LongJump {
#if defined(_M_ARM_64)
[[nodiscard]]
FEX_DEFAULT_VISIBILITY FEX_NAKED uint64_t SetJump(JumpBuf& Buffer) {
@@ -116,4 +116,4 @@ FEX_DEFAULT_VISIBILITY FEX_NAKED void LongJump(JumpBuf& Buffer, uint64_t Value)
}
#endif
} // namespace FEXCore::UncheckedLongJump
} // namespace FEXCore::LongJump
-22
View File
@@ -1,6 +1,4 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <atomic>
#include <chrono>
#include <mutex>
@@ -186,26 +184,6 @@ template bool Wait<uint16_t>(uint16_t*, uint16_t, const std::chrono::nanoseconds
template bool Wait<uint32_t>(uint32_t*, uint32_t, const std::chrono::nanoseconds&);
template bool Wait<uint64_t>(uint64_t*, uint64_t, const std::chrono::nanoseconds&);
template<typename T>
static inline T OneShotWFEBitComparison(T* Futex, T Mask, T Comp) {
auto AtomicFutex = std::atomic_ref<T>(*Futex);
T Result = AtomicFutex.load();
// Early exit if possible.
if ((Result & Mask) == Comp) {
return Result;
}
Result = LoadExclusive(Futex);
if ((Result & Mask) == Comp) {
return Result;
}
// Waits for write and returns result.
Result = WFELoadAtomic(Futex);
return Result;
}
#else
template<typename T, typename TT>
static inline void Wait(T* Futex, TT ExpectedValue) {
-384
View File
@@ -1,384 +0,0 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <atomic>
#include <cstdint>
#if !defined(_WIN32)
#include <linux/futex.h> /* Definition of FUTEX_* constants */
#include <sys/syscall.h> /* Definition of SYS_* constants */
#include <unistd.h>
#else
#include <synchapi.h>
#endif
#include <FEXCore/Utils/LogManager.h>
#include "Utils/SpinWaitLock.h"
namespace FEXCore::Utils::WritePriorityMutex {
#if !defined(_WIN32)
// A custom mutex that prioritizes exclusive locks.
// In highly contested scenarios, this can help minimize overall contention time.
//
// Features:
// - Up to 32767 pending exclusive locks ("writers")
// - Up to 32767 pending shared_locks ("readers")
// - Low-overhead waiting via WFE with a fallback to futex on timeout
// - Direct writer->reader hand-off and vice-versa to further reduce overhead
//
// Trade-offs:
// - No guaranteed order of wake-ups besides prioritizing writers
// - No support for recursive locking
// - We can't use FUTEX_LOCK_PI to enable priority inheritance
class Mutex final {
public:
Mutex() = default;
// Move-only type
Mutex(const Mutex&) = delete;
Mutex& operator=(const Mutex&) = delete;
Mutex(Mutex&& rhs) = delete;
Mutex& operator=(Mutex&&) = delete;
void lock() {
// Try a non-blocking lock first.
if (try_lock()) {
return;
}
// Try a quick WFE write-lock.
if (Attempt_WFE_WriteLock()) {
return;
}
// Still couldn't get it. Start waiting.
auto AtomicFutex = std::atomic_ref<uint32_t>(Futex);
uint32_t Expected {};
uint32_t Desired {};
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
Expected = AtomicFutex.load(std::memory_order_relaxed);
do {
// Increment the number of write waiters.
Desired = Expected + WRITE_WAITER_INCREMENT;
LOGMAN_THROW_A_FMT((Desired & WRITE_WAITER_COUNT_MASK) != 0, "Overflow in write-waiters!");
} while (AtomicFutex.compare_exchange_strong(Expected, Desired, std::memory_order_acq_rel, std::memory_order_acquire) == false);
#else
Expected = AtomicFutex.fetch_add(WRITE_WAITER_INCREMENT);
Desired = Expected + WRITE_WAITER_INCREMENT;
#endif
// Thread added to waiter list.
Expected = Desired;
while (true) {
bool Sleep = false;
do {
if ((Expected & WRITE_OWNED_BIT) == 0 && (Expected & READ_OWNER_COUNT_MASK) == 0) {
// If not write-owned, and no read-owners, try to acquire.
LOGMAN_THROW_A_FMT((Expected & WRITE_WAITER_COUNT_MASK) != 0, "Underflow in write-waiters!");
// Add write-owned bit.
Desired = Expected | WRITE_OWNED_BIT;
// Remove ourselves from the wait list.
Desired -= WRITE_WAITER_INCREMENT;
Sleep = false;
} else {
// Already write-owned or read-locked. Go to sleep.
Desired = Expected;
Sleep = true;
break;
}
} while (AtomicFutex.compare_exchange_strong(Expected, Desired, std::memory_order_acq_rel, std::memory_order_acquire) == false);
if (!Sleep) {
// Acquired early.
LOGMAN_THROW_A_FMT((Desired & WRITE_OWNED_BIT) == WRITE_OWNED_BIT, "Somehow acquired a write-lock without it being set!");
return;
}
FutexWaitForWriteAvailable(Desired);
Expected = AtomicFutex.load(std::memory_order_relaxed);
}
}
void lock_shared() {
// Try an uncontended lock first.
if (try_lock_shared()) {
return;
}
// Try a quick WFE read-lock.
if (Attempt_WFE_ReadLock()) {
return;
}
auto AtomicFutex = std::atomic_ref<uint32_t>(Futex);
uint32_t Expected = AtomicFutex.load(std::memory_order_relaxed);
uint32_t Desired {};
while (true) {
bool Sleep = false;
do {
if ((Expected & WRITE_OWNED_BIT) == 0 && (Expected & WRITE_WAITER_COUNT_MASK) == 0) {
// If no write-owner and no write-waiting, try and acquire.
Desired = Expected + READ_OWNER_INCREMENT;
LOGMAN_THROW_A_FMT((Desired & READ_OWNER_COUNT_MASK) != 0, "Overflow in read-owners!");
Sleep = false;
} else {
// Waiting for lock to become available. Add to waiters.
Desired = Expected | READ_WAITER_BIT;
Sleep = true;
}
} while (AtomicFutex.compare_exchange_strong(Expected, Desired, std::memory_order_acq_rel, std::memory_order_acquire) == false);
if (!Sleep) {
// Acquired early.
LOGMAN_THROW_A_FMT((Desired & WRITE_OWNED_BIT) != WRITE_OWNED_BIT, "Somehow read-locked and got a write lock!");
return;
}
FutexWaitForReadAvailable(Desired);
Expected = AtomicFutex.load(std::memory_order_relaxed);
}
}
void unlock() {
auto AtomicFutex = std::atomic_ref<uint32_t>(Futex);
uint32_t Expected = AtomicFutex.load(std::memory_order_relaxed);
uint32_t Desired {};
do {
LOGMAN_THROW_A_FMT((Expected & WRITE_OWNED_BIT) == WRITE_OWNED_BIT, "Trying to write-unlock something not write-locked!");
// Remove the exclusive lock bit.
Desired = Expected & ~WRITE_OWNED_BIT;
// If no more writers, then make sure to clear the read-waiters bit as well.
if ((Desired & WRITE_WAITER_COUNT_MASK) == 0) {
Desired &= ~READ_WAITER_BIT;
}
} while (AtomicFutex.compare_exchange_strong(Expected, Desired, std::memory_order_acq_rel, std::memory_order_acquire) == false);
// If success, then `Expected` has old value. Containing `READ_WAITER_BIT` which was just masked off, and also `WRITE_WAITER_COUNT_MASK`.
if ((Expected & WRITE_WAITER_COUNT_MASK)) {
// Handle write-write handoff.
FutexWakeWriter();
} else if ((Expected & READ_WAITER_BIT)) {
// Handle write-reader handoff.
FutexWakeReaders();
}
}
void unlock_shared() {
auto AtomicFutex = std::atomic_ref<uint32_t>(Futex);
uint32_t Desired {};
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
uint32_t Expected = AtomicFutex.load(std::memory_order_relaxed);
do {
LOGMAN_THROW_A_FMT((Expected & WRITE_OWNED_BIT) != WRITE_OWNED_BIT, "Trying to read-unlock something write-locked!");
LOGMAN_THROW_A_FMT((Expected & READ_OWNER_COUNT_MASK) != 0, "Trying to read-unlock something not read-locked!");
// Decrement the shared counter.
Desired = Expected - READ_OWNER_INCREMENT;
} while (AtomicFutex.compare_exchange_strong(Expected, Desired, std::memory_order_acq_rel, std::memory_order_acquire) == false);
#else
Desired = AtomicFutex.fetch_sub(READ_OWNER_INCREMENT) - READ_OWNER_INCREMENT;
#endif
// Handle read->write handoff if there are any waiting writers, and no readers left.
if ((Desired & WRITE_WAITER_COUNT_MASK) && (Desired & READ_OWNER_COUNT_MASK) == 0) {
FutexWakeWriter();
}
}
bool try_lock() {
auto AtomicFutex = std::atomic_ref<uint32_t>(Futex);
uint32_t Expected = 0;
// Try and grab the owned bit.
uint32_t Desired = WRITE_OWNED_BIT;
// try to CAS immediately.
return AtomicFutex.compare_exchange_strong(Expected, Desired, std::memory_order_acq_rel, std::memory_order_acquire);
}
// Can race with other threads trying to lock shared!
bool try_lock_shared() {
auto AtomicFutex = std::atomic_ref<uint32_t>(Futex);
uint32_t Expected = AtomicFutex.load(std::memory_order_relaxed);
// Exclusively owned or has a list of waiting owners. Can't pass.
if ((Expected & WRITE_OWNED_BIT) || (Expected & WRITE_WAITER_COUNT_MASK)) {
return false;
}
// Try to add reader.
uint32_t Desired = Expected + READ_OWNER_INCREMENT;
LOGMAN_THROW_A_FMT((Desired & READ_OWNER_COUNT_MASK) != 0, "Overflow in read-owners!");
// Uncontended mutex check
return AtomicFutex.compare_exchange_strong(Expected, Desired, std::memory_order_acq_rel, std::memory_order_acquire);
}
private:
void FutexWaitForWriteAvailable(uint32_t Expected) {
::syscall(SYS_futex, &Futex, FUTEX_PRIVATE_FLAG | FUTEX_WAIT_BITSET, Expected, nullptr, nullptr, FUTEX_BITSET_WAIT_WRITERS);
}
// Read-lock waiting for writers to drain out.
void FutexWaitForReadAvailable(uint32_t Expected) {
::syscall(SYS_futex, &Futex, FUTEX_PRIVATE_FLAG | FUTEX_WAIT_BITSET, Expected, nullptr, nullptr, FUTEX_BITSET_WAIT_READERS);
}
// Read-Lock or Write-lock unlocked, wake one writer.
// - Read->Write handoff.
// - Write->Write handoff.
void FutexWakeWriter() {
::syscall(SYS_futex, &Futex, FUTEX_PRIVATE_FLAG | FUTEX_WAKE_BITSET, 1, nullptr, nullptr, FUTEX_BITSET_WAIT_WRITERS);
}
// Write-lock unlocked, wake read-locks waiting.
void FutexWakeReaders() {
// Wake all readers.
::syscall(SYS_futex, &Futex, FUTEX_PRIVATE_FLAG | FUTEX_WAKE_BITSET, INT_MAX, nullptr, nullptr, FUTEX_BITSET_WAIT_READERS);
}
// Reuse the SpinWaitLock WFE implementations for read/write lock acquiring with WFE.
// Can't reuse the spin-lock directly as some bit-representations are different.
// WFE-write-lock is less likely to occur the more read-lock threads are participating. Can still occur so good to try.
// WFE-read-lock is actually quite likely to succeed.
// Return: true if the lock was acquired.
bool Attempt_WFE_WriteLock() {
#ifdef _M_ARM_64
const auto Begin = FEXCore::Utils::SpinWaitLock::GetCycleCounter();
auto Now = Begin;
const auto Duration = FEXCore::Utils::SpinWaitLock::CycleCounterFrequency / CYCLECOUNT_DIVISOR;
auto AtomicFutex = std::atomic_ref<uint32_t>(Futex);
uint32_t Expected = AtomicFutex.load(std::memory_order_relaxed);
while ((Now - Begin) < Duration) {
if (Expected == 0) {
// Try and grab the owned bit.
uint32_t Desired = WRITE_OWNED_BIT;
if (AtomicFutex.compare_exchange_strong(Expected, Desired, std::memory_order_acq_rel, std::memory_order_acquire)) {
return true;
}
}
// One-shot attempt to wait for mask to be zero.
Expected = FEXCore::Utils::SpinWaitLock::OneShotWFEBitComparison(&Futex, ~0U, 0U);
Now = FEXCore::Utils::SpinWaitLock::GetCycleCounter();
}
#endif
return false;
}
// Return: true if the lock was acquired.
bool Attempt_WFE_ReadLock() {
#ifdef _M_ARM_64
// Spin on a WFE for a short-amount of time, waiting for write-owned and writer-count to be zero.
// - Attempt to acquire read-lock at that point.
// - Don't add read-waiters bit on failure, return false.
const auto Begin = FEXCore::Utils::SpinWaitLock::GetCycleCounter();
auto Now = Begin;
const auto Duration = FEXCore::Utils::SpinWaitLock::CycleCounterFrequency / CYCLECOUNT_DIVISOR;
auto AtomicFutex = std::atomic_ref<uint32_t>(Futex);
uint32_t Expected = AtomicFutex.load(std::memory_order_relaxed);
uint32_t Desired {};
while ((Now - Begin) < Duration) {
if ((Expected & WRITE_OWNED_BIT) == 0 && (Expected & WRITE_WAITER_COUNT_MASK) == 0) {
// If no write-owner and no write-waiting, try and acquire.
Desired = Expected + READ_OWNER_INCREMENT;
LOGMAN_THROW_A_FMT((Desired & READ_OWNER_COUNT_MASK) != 0, "Overflow in read-owners!");
if (AtomicFutex.compare_exchange_strong(Expected, Desired, std::memory_order_acq_rel, std::memory_order_acquire)) {
return true;
}
}
// One-shot attempt to wait for mask to be zero.
Expected = FEXCore::Utils::SpinWaitLock::OneShotWFEBitComparison(&Futex, WRITE_OWNED_BIT | WRITE_WAITER_COUNT_MASK, 0U);
Now = FEXCore::Utils::SpinWaitLock::GetCycleCounter();
}
#endif
return false;
}
constexpr static uint32_t WRITE_OWNED_BIT = 1U << 31;
constexpr static uint32_t READ_WAITER_BIT = 1U << 15;
constexpr static uint32_t WRITE_WAITER_OFFSET = 16;
constexpr static uint32_t WRITE_WAITER_INCREMENT = 1U << WRITE_WAITER_OFFSET;
constexpr static uint32_t READ_OWNER_INCREMENT = 1;
// Count masks
constexpr static uint32_t WRITE_WAITER_COUNT_MASK = 0x7FFFU << WRITE_WAITER_OFFSET;
constexpr static uint32_t READ_OWNER_COUNT_MASK = 0x7FFFU;
// Independent futex bit-set masks.
// Wait for readers to drain.
constexpr static uint32_t FUTEX_BITSET_WAIT_READERS = 1U << 0;
// Wait for writers to drain.
constexpr static uint32_t FUTEX_BITSET_WAIT_WRITERS = 1U << 1;
// Only spin on WFE for 0.01ms (10k ns).
constexpr static uint64_t CYCLECOUNT_DIVISOR = 1'000'000'000ULL / 10'000U;
// Layout:
// Bits[31]: Write-lock bit.
// Bits[30:16]: Write-waiter count.
// Bits[15]: Read-waiter bit.
// Bits[14:0]: Read-owner count.
uint32_t Futex {};
};
#else
// SRWLocks are already write-priority locks in WINE and Windows. Use them to avoid lock stampeding.
class Mutex final {
public:
void lock() {
AcquireSRWLockExclusive(&Futex);
}
void lock_shared() {
AcquireSRWLockShared(&Futex);
}
void unlock() {
ReleaseSRWLockExclusive(&Futex);
}
void unlock_shared() {
ReleaseSRWLockShared(&Futex);
}
bool try_lock() {
return TryAcquireSRWLockExclusive(&Futex);
}
bool try_lock_shared() {
return TryAcquireSRWLockShared(&Futex);
}
private:
SRWLOCK Futex = SRWLOCK_INIT;
};
#endif
} // namespace FEXCore::Utils::WritePriorityMutex
+3 -5
View File
@@ -42,9 +42,6 @@ enum OperatingMode {
using CodeRangeInvalidationFn = std::function<void(uint64_t start, uint64_t Length)>;
// Nested vector of guest block entrypoints
using InvalidatedEntryAccumulator = fextl::vector<fextl::vector<uint64_t>>;
using CustomIREntrypointHandler = std::function<void(uintptr_t Entrypoint, IR::IREmitter*)>;
using ExitHandler = std::function<void(Core::InternalThreadState* Thread)>;
@@ -141,8 +138,9 @@ public:
virtual AbstractCodeCache& GetCodeCache() = 0;
FEX_DEFAULT_VISIBILITY virtual void ClearCodeCache(FEXCore::Core::InternalThreadState* Thread, bool NewCodeBuffer = true) = 0;
FEX_DEFAULT_VISIBILITY virtual void InvalidateGuestCodeRange(
FEXCore::Core::InternalThreadState* Thread, InvalidatedEntryAccumulator& Accumulator, uint64_t Start, uint64_t Length) = 0;
FEX_DEFAULT_VISIBILITY virtual void InvalidateCodeBuffersCodeRange(uint64_t Start, uint64_t Length) = 0;
FEX_DEFAULT_VISIBILITY virtual void
InvalidateThreadCachedCodeRange(FEXCore::Core::InternalThreadState* Thread, uint64_t Start, uint64_t Length) = 0;
FEX_DEFAULT_VISIBILITY virtual FEXCore::ForkableSharedMutex& GetCodeInvalidationMutex() = 0;
FEX_DEFAULT_VISIBILITY virtual void
+5
View File
@@ -353,6 +353,7 @@ struct JITPointers {
uint64_t ExitFunctionLinker {};
uint64_t ThreadStopHandlerSpillSRA {};
uint64_t ThreadPauseHandlerSpillSRA {};
uint64_t UnimplementedInstructionHandler {};
uint64_t GuestSignal_SIGILL {};
uint64_t GuestSignal_SIGTRAP {};
uint64_t GuestSignal_SIGSEGV {};
@@ -370,6 +371,8 @@ struct JITPointers {
// Process specific
uint64_t LUDIV {};
uint64_t LDIV {};
uint64_t LUREM {};
uint64_t LREM {};
// Thread Specific
@@ -378,6 +381,8 @@ struct JITPointers {
* @{ */
uint64_t LUDIVHandler {};
uint64_t LDIVHandler {};
uint64_t LUREMHandler {};
uint64_t LREMHandler {};
/** @} */
} AArch64;
+3 -4
View File
@@ -5,9 +5,8 @@
#include <cstdint>
// Reimplementation of longjmp without glibc fortification checks.
// This is useful when false positives need to be avoided or when using
// a libc implementation that does not implement std::longjmp.
namespace FEXCore::UncheckedLongJump {
// This is useful to avoid false positives reported by glibc.
namespace FEXCore::LongJump {
// JumpBuf definition needs to be public because the frontend needs to understand it.
#if defined(_M_ARM_64)
struct JumpBuf {
@@ -35,4 +34,4 @@ struct JumpBuf {
[[nodiscard]] FEX_DEFAULT_VISIBILITY uint64_t SetJump(JumpBuf& Buffer);
[[noreturn]] FEX_DEFAULT_VISIBILITY void LongJump(JumpBuf& Buffer, uint64_t Value);
} // namespace FEXCore::UncheckedLongJump
} // namespace FEXCore::LongJump
-66
View File
@@ -1,66 +0,0 @@
// SPDX-License-Identifier: MIT
#include <catch2/catch_test_macros.hpp>
#include <catch2/generators/catch_generators_range.hpp>
#include "Utils/Allocator/HostAllocator.h"
#include <FEXCore/Utils/Allocator.h>
#include <sys/mman.h>
template<typename T>
bool HasSyscallError(T Result) {
constexpr uint64_t MAX_ERRNO = 0xFFFF'FFFF'FFFF'0001ULL;
return reinterpret_cast<uint64_t>(Result) >= MAX_ERRNO;
}
TEST_CASE("Allocator - Fixed replacement") {
const auto RegionSize = 128 * 1024 * 1024;
fextl::vector<FEXCore::Allocator::MemoryRegion> MemoryRegions {};
for (size_t i = 0; i < 2; ++i) {
auto Ptr = mmap(nullptr, RegionSize, PROT_NONE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
MemoryRegions.emplace_back(FEXCore::Allocator::MemoryRegion {
.Ptr = Ptr,
.Size = RegionSize,
});
}
auto Allocator = Alloc::OSAllocator::Create64BitAllocatorWithRegions(MemoryRegions);
auto Base = Allocator->Mmap(nullptr, 4096, PROT_NONE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
REQUIRE(!HasSyscallError(Base));
// Allocate perfectly overlapping pages. Allocate as many pages as the region.
// FEX had a bug where the allocator could run out of memory with MAP_FIXED.
for (size_t i = 0; i < (RegionSize / 4096); ++i) {
auto NewBase = Allocator->Mmap(Base, 4096, PROT_NONE, MAP_FIXED | MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
REQUIRE(Base == NewBase);
}
Alloc::OSAllocator::ReleaseAllocatorWorkaround(std::move(Allocator));
}
TEST_CASE("Allocator - Non-Fit") {
const auto RegionSize = 128 * 1024 * 1024;
fextl::vector<FEXCore::Allocator::MemoryRegion> MemoryRegions {};
for (size_t i = 0; i < 2; ++i) {
auto Ptr = mmap(nullptr, RegionSize, PROT_NONE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
MemoryRegions.emplace_back(FEXCore::Allocator::MemoryRegion {
.Ptr = Ptr,
.Size = RegionSize,
});
}
auto Allocator = Alloc::OSAllocator::Create64BitAllocatorWithRegions(MemoryRegions);
auto Base = Allocator->Mmap(nullptr, RegionSize / 4, PROT_NONE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
REQUIRE(!HasSyscallError(Base));
// Try to allocate within the whole VMA size minus a small amount.
// FEX had a bug where if the allocation fit within a VMA region, it would try and allocate past the end without checking.
// Only occurred when `MAP_FIXED` was used.
auto NewBase = Allocator->Mmap(Base, RegionSize - (4096 * 64), PROT_NONE, MAP_FIXED | MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
// Must either fit in the VMA region, or fail.
// - If it matches previous allocation, then it fit in the VMA region.
// - This can happen if FEX's allocator gains support for VMA merging.
// - If it errors, then it doesn't fit in the VMA region.
REQUIRE((NewBase == Base || HasSyscallError(NewBase)));
Alloc::OSAllocator::ReleaseAllocatorWorkaround(std::move(Allocator));
}
-23
View File
@@ -3,7 +3,6 @@
#include <catch2/generators/catch_generators_range.hpp>
#include "Utils/Allocator/FlexBitSet.h"
#include <sys/mman.h>
TEST_CASE("FlexBitSet - Sizing") {
// Ensure that FlexBitSet sizing is correct.
@@ -41,25 +40,3 @@ TEST_CASE("FlexBitSet - Sizing") {
CHECK(FEXCore::FlexBitSet<uint32_t>::SizeInBits(sizeof(uint32_t) * 8) == sizeof(uint32_t) * 8);
CHECK(FEXCore::FlexBitSet<uint64_t>::SizeInBits(sizeof(uint64_t) * 8) == sizeof(uint64_t) * 8);
}
TEST_CASE("FlexBitSet - Limit") {
// Ensure that the FlexBitSet doesn't read past the limits, and returns correct indexes.
const auto Size = 4096 * 3;
auto Ptr = mmap(nullptr, Size, PROT_NONE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
auto PtrMiddle = reinterpret_cast<void*>(reinterpret_cast<uintptr_t>(Ptr) + 4096);
REQUIRE(mprotect(PtrMiddle, 4096, PROT_READ | PROT_WRITE) != -1);
using ElementType = uint8_t;
const size_t NumElements = 4096 * 8;
auto FlexBit = reinterpret_cast<FEXCore::FlexBitSet<ElementType>*>(PtrMiddle);
for (size_t i = 0; i < NumElements; ++i) {
auto Result = FlexBit->ForwardScanForRange<true>(i, 1, NumElements);
CHECK(Result.FoundElement == i);
}
for (size_t i = 0; i < NumElements; ++i) {
auto Result = FlexBit->BackwardScanForRange<true>(i, 1, 0);
CHECK(Result.FoundElement == i);
}
}
-2
View File
@@ -81,8 +81,6 @@ BigCoreIDs = {
[ ["apple-a13", "0.0"], # If we aren't on 12.0+
["apple-a14", "12.0"], # Only exists in 12.0+
],
# QEmu HVF 10.2+
tuple([0x61, 0]): "apple-a13", # Can't determine variant, choose lowest.
}
LittleCoreIDs = {
+2 -2
View File
@@ -52,8 +52,8 @@ public:
{
auto CodeInvalidationlk = FEXCore::GuardSignalDeferringSection(CTX->GetCodeInvalidationMutex(), Thread);
FEXCore::Context::InvalidatedEntryAccumulator Accumulator;
CTX->InvalidateGuestCodeRange(Thread, Accumulator, reinterpret_cast<uint64_t>(CodeStart), MAX_CODE_SIZE);
CTX->InvalidateCodeBuffersCodeRange(reinterpret_cast<uint64_t>(CodeStart), MAX_CODE_SIZE);
CTX->InvalidateThreadCachedCodeRange(Thread, reinterpret_cast<uint64_t>(CodeStart), MAX_CODE_SIZE);
}
ClearStats();
+5
View File
@@ -710,6 +710,11 @@ ApplicationWindow {
config: "X87ReducedPrecision"
}
ConfigCheckBox {
text: qsTr("Unsafe local flags optimization")
config: "ABILocalFlags"
}
ConfigCheckBox {
text: qsTr("Disable JIT optimization passes")
config: "O0"
@@ -222,6 +222,28 @@ void CheckForGCS() {
}
} // namespace FEX::GCS
namespace FEX::UnalignedAtomic {
void SetupKernelUnalignedAtomics() {
#ifndef PR_ARM64_SET_UNALIGN_ATOMIC
#define PR_ARM64_SET_UNALIGN_ATOMIC 0x46455849
#define PR_ARM64_UNALIGN_ATOMIC_EMULATE (1UL << 0)
#define PR_ARM64_UNALIGN_ATOMIC_BACKPATCH (1UL << 1)
#define PR_ARM64_UNALIGN_ATOMIC_STRICT_SPLIT_LOCKS (1UL << 2)
#endif
// Interfaces with downstream FEX kernel patches to control unaligned atomic handling
FEX_CONFIG_OPT(ParanoidTSO, PARANOIDTSO);
FEX_CONFIG_OPT(StrictInProcessSplitLocks, STRICTINPROCESSSPLITLOCKS);
uint64_t Flags = (StrictInProcessSplitLocks() ? PR_ARM64_UNALIGN_ATOMIC_STRICT_SPLIT_LOCKS : 0) |
(ParanoidTSO() ? 0 : PR_ARM64_UNALIGN_ATOMIC_BACKPATCH) | PR_ARM64_UNALIGN_ATOMIC_EMULATE;
if (prctl(PR_ARM64_SET_UNALIGN_ATOMIC, Flags, 0, 0, 0) != -1) {
LogMan::Msg::IFmt("FEX: Kernel unaligned atomics enabled!");
}
}
} // namespace FEX::UnalignedAtomic
/**
* @brief Get an FD from an environment variable and then unset the environment variable.
*
@@ -458,6 +480,7 @@ int main(int argc, char** argv, char** const envp) {
// Setup TSO hardware emulation immediately after initializing the context.
FEX::TSO::SetupTSOEmulation(CTX.get());
FEX::UnalignedAtomic::SetupKernelUnalignedAtomics();
if (!Loader.Is64BitMode()) {
// Tell the kernel we want to use the compat input syscalls even though we're
@@ -1,51 +0,0 @@
// SPDX-License-Identifier: MIT
#include "ArchHelpers/MContext.h"
namespace FEX::ArchHelpers::Context {
#ifdef _M_ARM_64
std::string_view GetESRName(uint64_t ESR) {
switch ((ESR & ESR1_EC) >> 26) {
case 0b000'000: return "Unknown";
case 0b000'001: return "Trapped WF*";
case 0b000'011: return "Trapped MCR/MRC";
case 0b000'100: return "Trapped MCRR/MRRC";
case 0b000'101: return "Trapped MCR/MRC (coproc==0b1110)";
case 0b000'110: return "Trapped LDC/STC";
case 0b000'111: return "Trapped SME;SVE,ASIMD,FP";
case 0b001'010: return "Trapped non-covered instruction";
case 0b001'100: return "Trapped MRRC (coproc==0b1110)";
case 0b001'101: return "Branch target exception";
case 0b001'110: return "Illegal Execution State";
case 0b010'001: return "AArch32 SVC";
case 0b010'100: return "Trapped MSRR/MRRS/System instruction";
case 0b010'101: return "AArch64 SVC";
case 0b011'000: return "Trapped MSR/MRS/System instruction";
case 0b011'001: return "Trapped SVE from ZEN";
case 0b011'011: return "TSTART Exception";
case 0b011'100: return "PAC Exception";
case 0b011'101: return "Trapped SME from SMEN";
case 0b100'000: return "Instruction abort";
case 0b100'001: return "Instruction abort w/o change to exception level";
case 0b100'010: return "PC Alignment fault";
case 0b100'100: return "Data abort";
case 0b100'101: return "Data abort w/o change to exception level";
case 0b100'110: return "SP Alignment fault";
case 0b100'111: return "Memory operation exception";
case 0b101'000: return "AArch32 Trapped FP Exception";
case 0b101'100: return "AArch64 Trapped FP Exception";
case 0b101'101: return "GCS exception";
case 0b101'111: return "SError exception";
case 0b110'000: return "BP Exception";
case 0b110'001: return "BP Exception w/o change to exception level";
case 0b110'010: return "Software step Exception";
case 0b110'011: return "Software step Exception w/o change to exception level";
case 0b110'100: return "Watchpoint Exception";
case 0b110'101: return "Watchpoit Exception w/o change to exception level";
case 0b111'000: return "AArch32 BKPT";
case 0b111'100: return "AArch64 BRK";
case 0b111'101: return "Profiling Exception";
default: return "Reserved";
}
}
#endif
} // namespace FEX::ArchHelpers::Context
@@ -203,12 +203,9 @@ constexpr static uint64_t ESR1_DataAbort_Level_EL2 = 0b01;
constexpr static uint64_t ESR1_DataAbort_Level_EL1 = 0b10;
constexpr static uint64_t ESR1_DataAbort_Level_EL0 = 0b11;
std::string_view GetESRName(uint64_t ESR);
static inline uint32_t GetProtectFlags(void* ucontext) {
uint64_t ESR = GetArmESR(ucontext);
LOGMAN_THROW_A_FMT((ESR & ESR1_EC) == ESR1_EC_DataAbort, "Unknown ESR1 EC type: 0x{:x} != 0x{:x}. Received '{}'", ESR & ESR1_EC,
ESR1_EC_DataAbort, GetESRName(ESR));
LOGMAN_THROW_A_FMT((ESR & ESR1_EC) == ESR1_EC_DataAbort, "Unknown ESR1 EC type: 0x{:x} != 0x{:x}", ESR & ESR1_EC, ESR1_EC_DataAbort);
uint32_t ProtectFlags {};
if ((ESR & ESR1_DataAbort_Level) == ESR1_DataAbort_Level_EL0) {
@@ -3,7 +3,6 @@ add_compile_options(-fno-operator-names)
set (SRCS
VDSO_Emulation.cpp
Thunks.cpp
ArchHelpers/MContext.cpp
GdbServer/Info.cpp
LinuxSyscalls/GdbServer.cpp
LinuxSyscalls/EmulatedFiles/EmulatedFiles.cpp
@@ -208,10 +208,9 @@ public:
// Thread object isn't valid very early in frontend's initialization.
// To be more optimal the frontend should provide this code with a valid Thread object earlier.
auto CodeInvalidationlk = GuardSignalDeferringSectionWithFallback(CTX->GetCodeInvalidationMutex(), CallingThread);
FEXCore::Context::InvalidatedEntryAccumulator Accumulator;
CTX->InvalidateCodeBuffersCodeRange(Start, Length);
for (auto& Thread : Threads) {
CTX->InvalidateGuestCodeRange(Thread->Thread, Accumulator, Start, Length);
CTX->InvalidateThreadCachedCodeRange(Thread->Thread, Start, Length);
}
}
@@ -223,10 +222,9 @@ public:
// Thread object isn't valid very early in frontend's initialization.
// To be more optimal the frontend should provide this code with a valid Thread object earlier.
auto CodeInvalidationlk = GuardSignalDeferringSectionWithFallback(CTX->GetCodeInvalidationMutex(), CallingThread);
FEXCore::Context::InvalidatedEntryAccumulator Accumulator;
CTX->InvalidateCodeBuffersCodeRange(Start, Length);
for (auto& Thread : Threads) {
CTX->InvalidateGuestCodeRange(Thread->Thread, Accumulator, Start, Length);
CTX->InvalidateThreadCachedCodeRange(Thread->Thread, Start, Length);
}
// Callback while holding the locks.
@@ -261,7 +261,7 @@ namespace PThreads {
return STracker;
}
void SetupLongJump(FEXCore::UncheckedLongJump::JumpBuf* exit_resolver) {
void SetupLongJump(FEXCore::LongJump::JumpBuf* exit_resolver) {
_exit_resolver = exit_resolver;
}
@@ -269,7 +269,7 @@ namespace PThreads {
void LongJumpExit(FEX::HLE::ThreadStateObject* ThreadObject, uint32_t Status) {
this->Status = Status;
this->ThreadObject = ThreadObject;
FEXCore::UncheckedLongJump::LongJump(*_exit_resolver, 1);
FEXCore::LongJump::LongJump(*_exit_resolver, 1);
FEX_UNREACHABLE;
}
@@ -288,9 +288,9 @@ namespace PThreads {
void* UserArg;
void* Stack {};
// Use FEXCore's UncheckedLongJump to avoid fortification checks.
// Use FEXCore's LongJump to avoid fortification checks.
// This avoids a false positive since glibc does not understand stack pivots.
FEXCore::UncheckedLongJump::JumpBuf* _exit_resolver {};
FEXCore::LongJump::JumpBuf* _exit_resolver {};
FEX::HLE::ThreadStateObject* ThreadObject {};
uint32_t Status {};
};
@@ -301,11 +301,11 @@ namespace PThreads {
PThread* Thread {reinterpret_cast<PThread*>(Ptr)};
StackBase = Thread->GetPivotStack();
STracker = Thread->GetStackTracker();
FEXCore::UncheckedLongJump::JumpBuf exit_resolver {};
FEXCore::LongJump::JumpBuf exit_resolver {};
bool LongJumpExit {};
if (FEXCore::UncheckedLongJump::SetJump(exit_resolver) == 0) {
if (FEXCore::LongJump::SetJump(exit_resolver) == 0) {
Thread->SetupLongJump(&exit_resolver);
// Run the user function.
// `Thread` object is dead after this function returns.
@@ -14,6 +14,3 @@ _BASIC_META(DRM_IOCTL_AMDGPU_WAIT_FENCES)
_BASIC_META(DRM_IOCTL_AMDGPU_VM)
_BASIC_META(DRM_IOCTL_AMDGPU_FENCE_TO_HANDLE)
_BASIC_META(DRM_IOCTL_AMDGPU_SCHED)
_BASIC_META(DRM_IOCTL_AMDGPU_USERQ)
_BASIC_META(DRM_IOCTL_AMDGPU_USERQ_SIGNAL)
_BASIC_META(DRM_IOCTL_AMDGPU_USERQ_WAIT)
@@ -1,11 +0,0 @@
_BASIC_META(DRM_IOCTL_ASAHI_GET_PARAMS)
_BASIC_META(DRM_IOCTL_ASAHI_GET_TIME)
_BASIC_META(DRM_IOCTL_ASAHI_VM_CREATE)
_BASIC_META(DRM_IOCTL_ASAHI_VM_DESTROY)
_BASIC_META(DRM_IOCTL_ASAHI_VM_BIND)
_BASIC_META(DRM_IOCTL_ASAHI_GEM_CREATE)
_BASIC_META(DRM_IOCTL_ASAHI_GEM_MMAP_OFFSET)
_BASIC_META(DRM_IOCTL_ASAHI_GEM_BIND_OBJECT)
_BASIC_META(DRM_IOCTL_ASAHI_QUEUE_CREATE)
_BASIC_META(DRM_IOCTL_ASAHI_QUEUE_DESTROY)
_BASIC_META(DRM_IOCTL_ASAHI_SUBMIT)
@@ -15,12 +15,10 @@ extern "C" {
#include "fex-drm/drm_mode.h"
#include "fex-drm/i915_drm.h"
#include "fex-drm/amdgpu_drm.h"
#include "fex-drm/asahi_drm.h"
#include "fex-drm/lima_drm.h"
#include "fex-drm/panfrost_drm.h"
#include "fex-drm/msm_drm.h"
#include "fex-drm/nouveau_drm.h"
#include "fex-drm/nova_drm.h"
#include "fex-drm/radeon_drm.h"
#include "fex-drm/vc4_drm.h"
#include "fex-drm/v3d_drm.h"
@@ -1274,13 +1272,11 @@ namespace V3D {
#include "LinuxSyscalls/x32/Ioctl/drm.inl"
#include "LinuxSyscalls/x32/Ioctl/amdgpu_drm.inl"
#include "LinuxSyscalls/x32/Ioctl/asahi_drm.inl"
#include "LinuxSyscalls/x32/Ioctl/msm_drm.inl"
#include "LinuxSyscalls/x32/Ioctl/i915_drm.inl"
#include "LinuxSyscalls/x32/Ioctl/lima_drm.inl"
#include "LinuxSyscalls/x32/Ioctl/panfrost_drm.inl"
#include "LinuxSyscalls/x32/Ioctl/nouveau_drm.inl"
#include "LinuxSyscalls/x32/Ioctl/nova_drm.inl"
#include "LinuxSyscalls/x32/Ioctl/radeon_drm.inl"
#include "LinuxSyscalls/x32/Ioctl/vc4_drm.inl"
#include "LinuxSyscalls/x32/Ioctl/v3d_drm.inl"
@@ -10,4 +10,4 @@ _BASIC_META(DRM_IOCTL_MSM_GEM_MADVISE)
_BASIC_META(DRM_IOCTL_MSM_SUBMITQUEUE_NEW)
_BASIC_META(DRM_IOCTL_MSM_SUBMITQUEUE_CLOSE)
_BASIC_META(DRM_IOCTL_MSM_SUBMITQUEUE_QUERY)
_BASIC_META(DRM_IOCTL_MSM_VM_BIND)
@@ -1,3 +0,0 @@
_BASIC_META(DRM_IOCTL_NOVA_GETPARAM)
_BASIC_META(DRM_IOCTL_NOVA_GEM_CREATE)
_BASIC_META(DRM_IOCTL_NOVA_GEM_INFO)
@@ -7,4 +7,3 @@ _BASIC_META(DRM_IOCTL_PANFROST_GET_BO_OFFSET)
_BASIC_META(DRM_IOCTL_PANFROST_MADVISE)
_BASIC_META(DRM_IOCTL_PANFROST_PERFCNT_ENABLE)
_BASIC_META(DRM_IOCTL_PANFROST_PERFCNT_DUMP)
_BASIC_META(DRM_IOCTL_PANFROST_SET_LABEL_BO)
@@ -11,5 +11,3 @@ _BASIC_META(DRM_IOCTL_PANTHOR_GROUP_SUBMIT)
_BASIC_META(DRM_IOCTL_PANTHOR_GROUP_GET_STATE)
_BASIC_META(DRM_IOCTL_PANTHOR_TILER_HEAP_CREATE)
_BASIC_META(DRM_IOCTL_PANTHOR_TILER_HEAP_DESTROY)
_BASIC_META(DRM_IOCTL_PANTHOR_BO_SET_LABEL)
_BASIC_META(DRM_IOCTL_PANTHOR_SET_USER_MMIO_OFFSET)
@@ -11,4 +11,3 @@ _BASIC_META(DRM_IOCTL_V3D_PERFMON_DESTROY)
_BASIC_META(DRM_IOCTL_V3D_PERFMON_GET_VALUES)
_BASIC_META(DRM_IOCTL_V3D_SUBMIT_CPU)
_BASIC_META(DRM_IOCTL_V3D_PERFMON_GET_COUNTER)
_BASIC_META(DRM_IOCTL_V3D_PERFMON_SET_GLOBAL)
+2 -9
View File
@@ -87,20 +87,13 @@ bool FindWineFEXApplication(int64_t PID, std::string_view exe, const std::vector
}
// Wine was found, scan the mapped files to see if anything mapped "libarm64ecfex.dll" or "libwow64fex.dll"
std::error_code ec {};
auto dir_iter = std::filesystem::directory_iterator(fmt::format("/proc/{}/map_files", PID), ec);
// If error reading symlink then skip.
if (ec) {
return false;
}
for (const auto& Entry : dir_iter) {
for (const auto& Entry : std::filesystem::directory_iterator(fmt::format("/proc/{}/map_files", PID))) {
// If not a symlink then skip.
if (!Entry.is_symlink()) {
continue;
}
std::error_code ec {};
const auto symlink_path = std::filesystem::read_symlink(Entry.path(), ec);
// If error reading symlink then skip.
if (ec) {
+1 -2
View File
@@ -639,7 +639,6 @@ NTSTATUS ProcessInit() {
SignalDelegator = fextl::make_unique<FEX::DummyHandlers::DummySignalDelegator>();
SyscallHandler = fextl::make_unique<Exception::ECSyscallHandler>();
Exception::HandlerConfig.emplace();
const auto NtDll = GetModuleHandle("ntdll.dll");
const bool IsWine = !!GetProcAddress(NtDll, "wine_get_version");
@@ -653,7 +652,7 @@ NTSTATUS ProcessInit() {
CTX->SetSignalDelegator(SignalDelegator.get());
CTX->SetSyscallHandler(SyscallHandler.get());
CTX->InitCore();
Exception::HandlerConfig.emplace(*CTX);
InvalidationTracker.emplace(*CTX, Threads);
HandleImageMap(NtDllBase);
+14 -27
View File
@@ -56,11 +56,7 @@ void InvalidationTracker::HandleMemoryProtectionNotification(uint64_t Address, u
if (NeedsInvalidate) {
// IntervalsLock cannot be held during invalidation
std::scoped_lock Lock(CTX.GetCodeInvalidationMutex());
FEXCore::Context::InvalidatedEntryAccumulator Accumulator;
for (auto Thread : Threads) {
CTX.InvalidateGuestCodeRange(Thread.second, Accumulator, AlignedBase, AlignedSize);
}
InvalidateIntervalInternal(AlignedBase, AlignedSize);
}
}
@@ -115,13 +111,8 @@ InvalidationTracker::InvalidateContainingSectionResult InvalidationTracker::Inva
reinterpret_cast<uint64_t>(Info.AllocationBase) == SectionBase) {
SectionSize += Info.RegionSize;
}
{
std::scoped_lock Lock(CTX.GetCodeInvalidationMutex());
FEXCore::Context::InvalidatedEntryAccumulator Accumulator;
for (auto Thread : Threads) {
CTX.InvalidateGuestCodeRange(Thread.second, Accumulator, SectionBase, SectionSize);
}
}
InvalidateIntervalInternal(SectionBase, SectionSize);
if (Free) {
std::unique_lock Lock(IntervalsLock);
@@ -141,13 +132,7 @@ void InvalidationTracker::InvalidateAlignedInterval(uint64_t Address, uint64_t S
const auto AlignedBase = Address & FEXCore::Utils::FEX_PAGE_MASK;
const auto AlignedSize = std::max(Size, (Address - AlignedBase + Size + FEXCore::Utils::FEX_PAGE_SIZE - 1) & FEXCore::Utils::FEX_PAGE_MASK);
{
std::scoped_lock Lock(CTX.GetCodeInvalidationMutex());
FEXCore::Context::InvalidatedEntryAccumulator Accumulator;
for (auto Thread : Threads) {
CTX.InvalidateGuestCodeRange(Thread.second, Accumulator, AlignedBase, AlignedSize);
}
}
InvalidateIntervalInternal(AlignedBase, AlignedSize);
if (Free) {
std::unique_lock Lock(IntervalsLock);
@@ -197,14 +182,8 @@ bool InvalidationTracker::HandleRWXAccessViolation(FEXCore::Core::InternalThread
}(FaultAddress);
if (NeedsInvalidate) {
{
// IntervalsLock cannot be held during invalidation
std::scoped_lock Lock(CTX.GetCodeInvalidationMutex());
FEXCore::Context::InvalidatedEntryAccumulator Accumulator;
for (auto Thread : Threads) {
CTX.InvalidateGuestCodeRange(Thread.second, Accumulator, FaultAddress & FEXCore::Utils::FEX_PAGE_MASK, FEXCore::Utils::FEX_PAGE_SIZE);
}
}
// IntervalsLock cannot be held during invalidation
InvalidateIntervalInternal(FaultAddress & FEXCore::Utils::FEX_PAGE_MASK, FEXCore::Utils::FEX_PAGE_SIZE);
DetectMonoBackpatcherBlock(Thread, HostPc);
return true;
}
@@ -274,4 +253,12 @@ void InvalidationTracker::DisableSMCDetection() {
} while (Query.Size);
}
void InvalidationTracker::InvalidateIntervalInternal(uint64_t Address, uint64_t Size) {
std::scoped_lock Lock(CTX.GetCodeInvalidationMutex());
CTX.InvalidateCodeBuffersCodeRange(Address, Size);
for (auto Thread : Threads) {
CTX.InvalidateThreadCachedCodeRange(Thread.second, Address, Size);
}
}
} // namespace FEX::Windows
@@ -4,6 +4,7 @@
#include <FEXCore/Utils/IntervalList.h>
#include <FEXCore/HLE/SyscallHandler.h>
#include <mutex>
#include <shared_mutex>
#include <unordered_map>
#include <string_view>
@@ -37,6 +38,8 @@ public:
private:
void DetectMonoBackpatcherBlock(FEXCore::Core::InternalThreadState* Thread, uint64_t HostPC);
void DisableSMCDetection();
void InvalidateIntervalInternal(uint64_t Address, uint64_t Size);
FEXCore::IntervalList<uint64_t> XIntervals;
FEXCore::IntervalList<uint64_t> RWXIntervals;
+19 -1
View File
@@ -1,18 +1,34 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <FEXCore/Core/Context.h>
#include <FEXCore/Config/Config.h>
#include <FEXCore/Utils/ArchHelpers/Arm64.h>
namespace FEX::Windows {
class TSOHandlerConfig final {
public:
TSOHandlerConfig() {
TSOHandlerConfig(FEXCore::Context::Context& CTX) {
if (HalfBarrierTSOEnabled()) {
UnalignedHandlerType = FEXCore::ArchHelpers::Arm64::UnalignedHandlerType::HalfBarrier;
} else {
UnalignedHandlerType = FEXCore::ArchHelpers::Arm64::UnalignedHandlerType::NonAtomic;
}
if (TSOEnabled()) {
BOOL Enable = TRUE;
NTSTATUS Status = NtSetInformationProcess(NtCurrentProcess(), ProcessFexHardwareTso, &Enable, sizeof(Enable));
if (Status == STATUS_SUCCESS) {
CTX.SetHardwareTSOSupport(true);
}
}
uint64_t Flags = (StrictInProcessSplitLocks() ? FEX_UNALIGN_ATOMIC_STRICT_SPLIT_LOCKS : 0) |
FEX_UNALIGN_ATOMIC_BACKPATCH | FEX_UNALIGN_ATOMIC_EMULATE;
if (NtSetInformationProcess(NtCurrentProcess(), ProcessFexUnalignAtomic, &Flags, sizeof(Flags)) == STATUS_SUCCESS) {
LogMan::Msg::IFmt("FEX: Kernel unaligned atomics enabled!");
}
}
FEXCore::ArchHelpers::Arm64::UnalignedHandlerType GetUnalignedHandlerType() const {
@@ -20,7 +36,9 @@ public:
}
private:
FEX_CONFIG_OPT(TSOEnabled, TSOENABLED);
FEX_CONFIG_OPT(HalfBarrierTSOEnabled, HALFBARRIERTSOENABLED);
FEX_CONFIG_OPT(StrictInProcessSplitLocks, STRICTINPROCESSSPLITLOCKS);
FEXCore::ArchHelpers::Arm64::UnalignedHandlerType UnalignedHandlerType {FEXCore::ArchHelpers::Arm64::UnalignedHandlerType::HalfBarrier};
};
-12
View File
@@ -28,18 +28,6 @@ void ReleaseSRWLockExclusive(PSRWLOCK SRWLock) {
RtlReleaseSRWLockExclusive(SRWLock);
}
void AcquireSRWLockShared(PSRWLOCK SRWLock) {
RtlAcquireSRWLockShared(SRWLock);
}
void ReleaseSRWLockShared(PSRWLOCK SRWLock) {
RtlReleaseSRWLockShared(SRWLock);
}
DLLEXPORT_FUNC(BOOLEAN, TryAcquireSRWLockShared, (PSRWLOCK SRWLock)) {
return RtlTryAcquireSRWLockShared(SRWLock);
}
DLLEXPORT_FUNC(BOOLEAN, TryAcquireSRWLockExclusive, (PSRWLOCK SRWLock)) {
return RtlTryAcquireSRWLockExclusive(SRWLock);
}
+1 -4
View File
@@ -522,22 +522,19 @@ void BTCpuProcessInit() {
SignalDelegator = fextl::make_unique<FEX::DummyHandlers::DummySignalDelegator>();
SyscallHandler = fextl::make_unique<WowSyscallHandler>();
Context::HandlerConfig.emplace();
const auto NtDll = GetModuleHandle("ntdll.dll");
const bool IsWine = !!GetProcAddress(NtDll, "wine_get_version");
OvercommitTracker.emplace(IsWine);
{
auto HostFeatures = FEX::Windows::CPUFeatures::FetchHostFeatures(IsWine);
// AVX is unsupported for WOW64
HostFeatures.SupportsAVX = false;
CTX = FEXCore::Context::Context::CreateNewContext(HostFeatures);
}
CTX->SetSignalDelegator(SignalDelegator.get());
CTX->SetSyscallHandler(SyscallHandler.get());
CTX->InitCore();
Context::HandlerConfig.emplace(*CTX);
InvalidationTracker.emplace(*CTX, Threads);
auto NtDllX86 = reinterpret_cast<SYSTEM_DLL_INIT_BLOCK*>(GetProcAddress(NtDll, "LdrSystemDllInitBlock"))->ntdll_handle;
+6 -3
View File
@@ -455,6 +455,12 @@ typedef enum _MEMORY_INFORMATION_CLASS {
#define SystemEmulationBasicInformation (SYSTEM_INFORMATION_CLASS)62
#define ProcessFexHardwareTso (PROCESSINFOCLASS)2000
#define ProcessFexUnalignAtomic (PROCESSINFOCLASS)2001
// These match the prctl flag values
#define FEX_UNALIGN_ATOMIC_EMULATE (1ULL << 0)
#define FEX_UNALIGN_ATOMIC_BACKPATCH (1ULL << 1)
#define FEX_UNALIGN_ATOMIC_STRICT_SPLIT_LOCKS (1ULL << 2)
typedef enum _KEY_VALUE_INFORMATION_CLASS {
KeyValueBasicInformation,
@@ -540,9 +546,6 @@ NTSTATUS WINAPI RtlWow64SetThreadContext(HANDLE, const WOW64_CONTEXT*);
void WINAPI Wow64ProcessPendingCrossProcessItems(void);
NTSTATUS WINAPI Wow64SystemServiceEx(UINT, UINT*);
NTSTATUS WINAPI RtlWow64SuspendThread(HANDLE, ULONG*);
void WINAPI RtlAcquireSRWLockShared(RTL_SRWLOCK*);
void WINAPI RtlReleaseSRWLockShared(RTL_SRWLOCK*);
BOOLEAN WINAPI RtlTryAcquireSRWLockShared(RTL_SRWLOCK*);
#ifdef __cplusplus
}
+4 -4
View File
@@ -80,10 +80,10 @@ static void* malloc_wrapper(size_t size) {
}
static void OnInit() {
fexfn_pack_GL_SetGuestMalloc((uintptr_t)malloc_wrapper, (uintptr_t)CallbackUnpack<decltype(malloc_wrapper)>::Unpack);
fexfn_pack_GL_SetGuestXSync((uintptr_t)XSync, (uintptr_t)CallbackUnpack<decltype(XSync)>::Unpack);
fexfn_pack_GL_SetGuestXGetVisualInfo((uintptr_t)XGetVisualInfo, (uintptr_t)CallbackUnpack<decltype(XGetVisualInfo)>::Unpack);
fexfn_pack_GL_SetGuestXDisplayString((uintptr_t)XDisplayString, (uintptr_t)CallbackUnpack<decltype(XDisplayString)>::Unpack);
fexfn_pack_SetGuestMalloc((uintptr_t)malloc_wrapper, (uintptr_t)CallbackUnpack<decltype(malloc_wrapper)>::Unpack);
fexfn_pack_SetGuestXSync((uintptr_t)XSync, (uintptr_t)CallbackUnpack<decltype(XSync)>::Unpack);
fexfn_pack_SetGuestXGetVisualInfo((uintptr_t)XGetVisualInfo, (uintptr_t)CallbackUnpack<decltype(XGetVisualInfo)>::Unpack);
fexfn_pack_SetGuestXDisplayString((uintptr_t)XDisplayString, (uintptr_t)CallbackUnpack<decltype(XDisplayString)>::Unpack);
}
// libGL.so must pull in libX11.so as a dependency. Referencing some libX11
+4 -4
View File
@@ -50,19 +50,19 @@ host_layout<_XDisplay*>::~host_layout() {
// Functions returning _XDisplay* should be handled explicitly via ptr_passthrough
guest_layout<_XDisplay*> to_guest(host_layout<_XDisplay*>) = delete;
static void fexfn_impl_libGL_GL_SetGuestMalloc(uintptr_t GuestTarget, uintptr_t GuestUnpacker) {
static void fexfn_impl_libGL_SetGuestMalloc(uintptr_t GuestTarget, uintptr_t GuestUnpacker) {
MakeHostTrampolineForGuestFunctionAt(GuestTarget, GuestUnpacker, &GuestMalloc);
}
static void fexfn_impl_libGL_GL_SetGuestXGetVisualInfo(uintptr_t GuestTarget, uintptr_t GuestUnpacker) {
static void fexfn_impl_libGL_SetGuestXGetVisualInfo(uintptr_t GuestTarget, uintptr_t GuestUnpacker) {
MakeHostTrampolineForGuestFunctionAt(GuestTarget, GuestUnpacker, &x11_manager.GuestXGetVisualInfo);
}
static void fexfn_impl_libGL_GL_SetGuestXSync(uintptr_t GuestTarget, uintptr_t GuestUnpacker) {
static void fexfn_impl_libGL_SetGuestXSync(uintptr_t GuestTarget, uintptr_t GuestUnpacker) {
MakeHostTrampolineForGuestFunctionAt(GuestTarget, GuestUnpacker, &x11_manager.GuestXSync);
}
static void fexfn_impl_libGL_GL_SetGuestXDisplayString(uintptr_t GuestTarget, uintptr_t GuestUnpacker) {
static void fexfn_impl_libGL_SetGuestXDisplayString(uintptr_t GuestTarget, uintptr_t GuestUnpacker) {
MakeHostTrampolineForGuestFunctionAt(GuestTarget, GuestUnpacker, &x11_manager.GuestXDisplayString);
}
+8 -8
View File
@@ -22,18 +22,18 @@ template<>
struct fex_gen_config<glXGetProcAddress> : fexgen::custom_host_impl, fexgen::custom_guest_entrypoint, fexgen::returns_guest_pointer {};
// internal use
void GL_SetGuestMalloc(uintptr_t, uintptr_t);
void GL_SetGuestXSync(uintptr_t, uintptr_t);
void GL_SetGuestXGetVisualInfo(uintptr_t, uintptr_t);
void GL_SetGuestXDisplayString(uintptr_t, uintptr_t);
void SetGuestMalloc(uintptr_t, uintptr_t);
void SetGuestXSync(uintptr_t, uintptr_t);
void SetGuestXGetVisualInfo(uintptr_t, uintptr_t);
void SetGuestXDisplayString(uintptr_t, uintptr_t);
template<>
struct fex_gen_config<GL_SetGuestMalloc> : fexgen::custom_guest_entrypoint, fexgen::custom_host_impl {};
struct fex_gen_config<SetGuestMalloc> : fexgen::custom_guest_entrypoint, fexgen::custom_host_impl {};
template<>
struct fex_gen_config<GL_SetGuestXSync> : fexgen::custom_guest_entrypoint, fexgen::custom_host_impl {};
struct fex_gen_config<SetGuestXSync> : fexgen::custom_guest_entrypoint, fexgen::custom_host_impl {};
template<>
struct fex_gen_config<GL_SetGuestXGetVisualInfo> : fexgen::custom_guest_entrypoint, fexgen::custom_host_impl {};
struct fex_gen_config<SetGuestXGetVisualInfo> : fexgen::custom_guest_entrypoint, fexgen::custom_host_impl {};
template<>
struct fex_gen_config<GL_SetGuestXDisplayString> : fexgen::custom_guest_entrypoint, fexgen::custom_host_impl {};
struct fex_gen_config<SetGuestXDisplayString> : fexgen::custom_guest_entrypoint, fexgen::custom_host_impl {};
template<typename>
struct fex_gen_type {};
+3 -3
View File
@@ -87,9 +87,9 @@ PFN_vkVoidFunction vkGetInstanceProcAddr(VkInstance a_0, const char* a_1) {
void OnInit() {
// TODO: Load libX11 on-demand instead
void* libx11 = dlopen("libX11.so.6", RTLD_LAZY);
fexfn_pack_Vulkan_SetGuestXSync((uintptr_t)dlsym(libx11, "XSync"), (uintptr_t)CallbackUnpack<decltype(XSync)>::Unpack);
fexfn_pack_Vulkan_SetGuestXGetVisualInfo((uintptr_t)dlsym(libx11, "XGetVisualInfo"), (uintptr_t)CallbackUnpack<decltype(XGetVisualInfo)>::Unpack);
fexfn_pack_Vulkan_SetGuestXDisplayString((uintptr_t)dlsym(libx11, "XDisplayString"), (uintptr_t)CallbackUnpack<decltype(XDisplayString)>::Unpack);
fexfn_pack_SetGuestXSync((uintptr_t)dlsym(libx11, "XSync"), (uintptr_t)CallbackUnpack<decltype(XSync)>::Unpack);
fexfn_pack_SetGuestXGetVisualInfo((uintptr_t)dlsym(libx11, "XGetVisualInfo"), (uintptr_t)CallbackUnpack<decltype(XGetVisualInfo)>::Unpack);
fexfn_pack_SetGuestXDisplayString((uintptr_t)dlsym(libx11, "XDisplayString"), (uintptr_t)CallbackUnpack<decltype(XDisplayString)>::Unpack);
}
LOAD_LIB_INIT(libvulkan, OnInit)
+3 -3
View File
@@ -63,15 +63,15 @@ static void DoSetupWithInstance(VkInstance instance) {
static X11Manager x11_manager;
static void fexfn_impl_libvulkan_Vulkan_SetGuestXGetVisualInfo(uintptr_t GuestTarget, uintptr_t GuestUnpacker) {
static void fexfn_impl_libvulkan_SetGuestXGetVisualInfo(uintptr_t GuestTarget, uintptr_t GuestUnpacker) {
MakeHostTrampolineForGuestFunctionAt(GuestTarget, GuestUnpacker, &x11_manager.GuestXGetVisualInfo);
}
static void fexfn_impl_libvulkan_Vulkan_SetGuestXSync(uintptr_t GuestTarget, uintptr_t GuestUnpacker) {
static void fexfn_impl_libvulkan_SetGuestXSync(uintptr_t GuestTarget, uintptr_t GuestUnpacker) {
MakeHostTrampolineForGuestFunctionAt(GuestTarget, GuestUnpacker, &x11_manager.GuestXSync);
}
static void fexfn_impl_libvulkan_Vulkan_SetGuestXDisplayString(uintptr_t GuestTarget, uintptr_t GuestUnpacker) {
static void fexfn_impl_libvulkan_SetGuestXDisplayString(uintptr_t GuestTarget, uintptr_t GuestUnpacker) {
MakeHostTrampolineForGuestFunctionAt(GuestTarget, GuestUnpacker, &x11_manager.GuestXDisplayString);
}
+6 -6
View File
@@ -29,15 +29,15 @@ template<typename>
struct fex_gen_type {};
// internal use
void Vulkan_SetGuestXSync(uintptr_t, uintptr_t);
void Vulkan_SetGuestXGetVisualInfo(uintptr_t, uintptr_t);
void Vulkan_SetGuestXDisplayString(uintptr_t, uintptr_t);
void SetGuestXSync(uintptr_t, uintptr_t);
void SetGuestXGetVisualInfo(uintptr_t, uintptr_t);
void SetGuestXDisplayString(uintptr_t, uintptr_t);
template<>
struct fex_gen_config<Vulkan_SetGuestXSync> : fexgen::custom_guest_entrypoint, fexgen::custom_host_impl {};
struct fex_gen_config<SetGuestXSync> : fexgen::custom_guest_entrypoint, fexgen::custom_host_impl {};
template<>
struct fex_gen_config<Vulkan_SetGuestXGetVisualInfo> : fexgen::custom_guest_entrypoint, fexgen::custom_host_impl {};
struct fex_gen_config<SetGuestXGetVisualInfo> : fexgen::custom_guest_entrypoint, fexgen::custom_host_impl {};
template<>
struct fex_gen_config<Vulkan_SetGuestXDisplayString> : fexgen::custom_guest_entrypoint, fexgen::custom_host_impl {};
struct fex_gen_config<SetGuestXDisplayString> : fexgen::custom_guest_entrypoint, fexgen::custom_host_impl {};
// So-called "dispatchable" handles are represented as opaque pointers.
// In addition to marking them as such, API functions that create these objects
+1 -1
View File
@@ -1,4 +1,4 @@
# FEX-2511
# FEX-2510
## FEXCore
See [FEXCore/Readme.md](../FEXCore/Readme.md) for more details
+30 -30
View File
@@ -19,7 +19,7 @@
"str x20, [x28, #24]",
"mov w1, #0x401",
"str x1, [x28, #1488]",
"ldr x0, [x28, #2872]",
"ldr x0, [x28, #2880]",
"br x0"
]
},
@@ -33,7 +33,7 @@
"str x20, [x28, #24]",
"mov w1, #0x401",
"str x1, [x28, #1488]",
"ldr x0, [x28, #2872]",
"ldr x0, [x28, #2880]",
"br x0"
]
},
@@ -47,7 +47,7 @@
"str x20, [x28, #24]",
"mov w1, #0x401",
"str x1, [x28, #1488]",
"ldr x0, [x28, #2872]",
"ldr x0, [x28, #2880]",
"br x0"
]
},
@@ -61,7 +61,7 @@
"str x20, [x28, #24]",
"mov w1, #0x401",
"str x1, [x28, #1488]",
"ldr x0, [x28, #2872]",
"ldr x0, [x28, #2880]",
"br x0"
]
},
@@ -75,7 +75,7 @@
"str x20, [x28, #24]",
"mov w1, #0x401",
"str x1, [x28, #1488]",
"ldr x0, [x28, #2872]",
"ldr x0, [x28, #2880]",
"br x0"
]
},
@@ -89,7 +89,7 @@
"str x20, [x28, #24]",
"mov w1, #0x401",
"str x1, [x28, #1488]",
"ldr x0, [x28, #2872]",
"ldr x0, [x28, #2880]",
"br x0"
]
},
@@ -103,7 +103,7 @@
"str x20, [x28, #24]",
"mov w1, #0x401",
"str x1, [x28, #1488]",
"ldr x0, [x28, #2872]",
"ldr x0, [x28, #2880]",
"br x0"
]
},
@@ -117,7 +117,7 @@
"str x20, [x28, #24]",
"mov w1, #0x401",
"str x1, [x28, #1488]",
"ldr x0, [x28, #2872]",
"ldr x0, [x28, #2880]",
"br x0"
]
},
@@ -131,7 +131,7 @@
"str x20, [x28, #24]",
"mov w1, #0x401",
"str x1, [x28, #1488]",
"ldr x0, [x28, #2872]",
"ldr x0, [x28, #2880]",
"br x0"
]
},
@@ -145,7 +145,7 @@
"str x20, [x28, #24]",
"mov w1, #0x401",
"str x1, [x28, #1488]",
"ldr x0, [x28, #2872]",
"ldr x0, [x28, #2880]",
"br x0"
]
},
@@ -159,7 +159,7 @@
"str x20, [x28, #24]",
"mov w1, #0x401",
"str x1, [x28, #1488]",
"ldr x0, [x28, #2872]",
"ldr x0, [x28, #2880]",
"br x0"
]
},
@@ -173,7 +173,7 @@
"str x20, [x28, #24]",
"mov w1, #0x401",
"str x1, [x28, #1488]",
"ldr x0, [x28, #2872]",
"ldr x0, [x28, #2880]",
"br x0"
]
},
@@ -187,7 +187,7 @@
"str x20, [x28, #24]",
"mov w1, #0x401",
"str x1, [x28, #1488]",
"ldr x0, [x28, #2872]",
"ldr x0, [x28, #2880]",
"br x0"
]
},
@@ -201,7 +201,7 @@
"str x20, [x28, #24]",
"mov w1, #0x401",
"str x1, [x28, #1488]",
"ldr x0, [x28, #2872]",
"ldr x0, [x28, #2880]",
"br x0"
]
},
@@ -215,7 +215,7 @@
"str x20, [x28, #24]",
"mov w1, #0x401",
"str x1, [x28, #1488]",
"ldr x0, [x28, #2872]",
"ldr x0, [x28, #2880]",
"br x0"
]
},
@@ -229,7 +229,7 @@
"str x20, [x28, #24]",
"mov w1, #0x401",
"str x1, [x28, #1488]",
"ldr x0, [x28, #2872]",
"ldr x0, [x28, #2880]",
"br x0"
]
},
@@ -243,7 +243,7 @@
"str x20, [x28, #24]",
"mov w1, #0x401",
"str x1, [x28, #1488]",
"ldr x0, [x28, #2872]",
"ldr x0, [x28, #2880]",
"br x0"
]
},
@@ -257,7 +257,7 @@
"str x20, [x28, #24]",
"mov w1, #0x401",
"str x1, [x28, #1488]",
"ldr x0, [x28, #2872]",
"ldr x0, [x28, #2880]",
"br x0"
]
},
@@ -271,7 +271,7 @@
"str x20, [x28, #24]",
"mov w1, #0x401",
"str x1, [x28, #1488]",
"ldr x0, [x28, #2872]",
"ldr x0, [x28, #2880]",
"br x0"
]
},
@@ -285,7 +285,7 @@
"str x20, [x28, #24]",
"mov w1, #0x401",
"str x1, [x28, #1488]",
"ldr x0, [x28, #2872]",
"ldr x0, [x28, #2880]",
"br x0"
]
},
@@ -299,7 +299,7 @@
"str x20, [x28, #24]",
"mov w1, #0x401",
"str x1, [x28, #1488]",
"ldr x0, [x28, #2872]",
"ldr x0, [x28, #2880]",
"br x0"
]
},
@@ -313,7 +313,7 @@
"str x20, [x28, #24]",
"mov w1, #0x401",
"str x1, [x28, #1488]",
"ldr x0, [x28, #2872]",
"ldr x0, [x28, #2880]",
"br x0"
]
},
@@ -327,7 +327,7 @@
"str x20, [x28, #24]",
"mov w1, #0x401",
"str x1, [x28, #1488]",
"ldr x0, [x28, #2872]",
"ldr x0, [x28, #2880]",
"br x0"
]
},
@@ -341,7 +341,7 @@
"str x20, [x28, #24]",
"mov w1, #0x401",
"str x1, [x28, #1488]",
"ldr x0, [x28, #2872]",
"ldr x0, [x28, #2880]",
"br x0"
]
},
@@ -355,7 +355,7 @@
"str x20, [x28, #24]",
"mov w1, #0x401",
"str x1, [x28, #1488]",
"ldr x0, [x28, #2872]",
"ldr x0, [x28, #2880]",
"br x0"
]
},
@@ -369,7 +369,7 @@
"str x20, [x28, #24]",
"mov w1, #0x401",
"str x1, [x28, #1488]",
"ldr x0, [x28, #2872]",
"ldr x0, [x28, #2880]",
"br x0"
]
},
@@ -383,7 +383,7 @@
"str x20, [x28, #24]",
"mov w1, #0x401",
"str x1, [x28, #1488]",
"ldr x0, [x28, #2872]",
"ldr x0, [x28, #2880]",
"br x0"
]
},
@@ -397,7 +397,7 @@
"str x20, [x28, #24]",
"mov w1, #0x401",
"str x1, [x28, #1488]",
"ldr x0, [x28, #2872]",
"ldr x0, [x28, #2880]",
"br x0"
]
},
@@ -411,7 +411,7 @@
"str x20, [x28, #24]",
"mov w1, #0x401",
"str x1, [x28, #1488]",
"ldr x0, [x28, #2872]",
"ldr x0, [x28, #2880]",
"br x0"
]
},
@@ -425,7 +425,7 @@
"str x20, [x28, #24]",
"mov w1, #0x401",
"str x1, [x28, #1488]",
"ldr x0, [x28, #2872]",
"ldr x0, [x28, #2880]",
"br x0"
]
}
File diff suppressed because it is too large. Load diff
File diff suppressed because it is too large. Load diff
@@ -2085,220 +2085,220 @@
"fcvt s2, d2",
"stur s2, [x9, #-68]",
"ldur s2, [x9, #-68]",
"fcvt d4, s2",
"ldur s5, [x9, #-72]",
"fcvt d2, s2",
"ldur s4, [x9, #-72]",
"fcvt d4, s4",
"fadd d4, d2, d4",
"fcvt s4, d4",
"stur s4, [x9, #-72]",
"ldur s4, [x9, #-72]",
"fcvt d4, s4",
"ldur s5, [x9, #-80]",
"fcvt d5, s5",
"fadd d5, d4, d5",
"fcvt s5, d5",
"stur s5, [x9, #-72]",
"ldur s5, [x9, #-72]",
"fcvt d5, s5",
"ldur s6, [x9, #-80]",
"fcvt d6, s6",
"fadd d5, d5, d6",
"fcvt s5, d5",
"stur s5, [x9, #-80]",
"ldur s5, [x9, #-72]",
"fcvt d5, s5",
"ldur s6, [x9, #-76]",
"fcvt d6, s6",
"fadd d5, d5, d6",
"fcvt s5, d5",
"stur s5, [x9, #-72]",
"fadd d4, d4, d5",
"fcvt s4, d4",
"stur s4, [x9, #-80]",
"ldur s4, [x9, #-72]",
"fcvt d4, s4",
"ldur s5, [x9, #-76]",
"fcvt d5, s5",
"fadd d5, d4, d5",
"fcvt s5, d5",
"stur s5, [x9, #-76]",
"ldur s5, [x9, #-188]",
"fcvt d5, s5",
"ldur s6, [x9, #-192]",
"fcvt d6, s6",
"fadd d5, d5, d6",
"fcvt s5, d5",
"stur s5, [x9, #-64]",
"fadd d4, d4, d5",
"fcvt s4, d4",
"stur s4, [x9, #-72]",
"ldur s4, [x9, #-76]",
"fcvt d4, s4",
"fadd d4, d2, d4",
"fcvt s4, d4",
"stur s4, [x9, #-76]",
"ldur s4, [x9, #-188]",
"fcvt d4, s4",
"ldur s5, [x9, #-192]",
"fcvt d5, s5",
"ldur s6, [x9, #-188]",
"fcvt d6, s6",
"fsub d5, d5, d6",
"fmul d5, d5, d3",
"fcvt s5, d5",
"stur s5, [x9, #-60]",
"ldur s5, [x9, #-180]",
"fadd d4, d4, d5",
"fcvt s4, d4",
"stur s4, [x9, #-64]",
"ldur s4, [x9, #-192]",
"fcvt d4, s4",
"ldur s5, [x9, #-188]",
"fcvt d5, s5",
"ldur s6, [x9, #-184]",
"fcvt d6, s6",
"fadd d5, d5, d6",
"fcvt s5, d5",
"stur s5, [x9, #-56]",
"ldur s5, [x9, #-180]",
"fsub d4, d4, d5",
"fmul d4, d4, d3",
"fcvt s4, d4",
"stur s4, [x9, #-60]",
"ldur s4, [x9, #-180]",
"fcvt d4, s4",
"ldur s5, [x9, #-184]",
"fcvt d5, s5",
"ldur s6, [x9, #-184]",
"fcvt d6, s6",
"fsub d5, d5, d6",
"fmul d5, d5, d3",
"fcvt s5, d5",
"stur s5, [x9, #-52]",
"ldur s5, [x9, #-56]",
"fadd d4, d4, d5",
"fcvt s4, d4",
"stur s4, [x9, #-56]",
"ldur s4, [x9, #-180]",
"fcvt d4, s4",
"ldur s5, [x9, #-184]",
"fcvt d5, s5",
"ldur s6, [x9, #-52]",
"fcvt d6, s6",
"fadd d5, d5, d6",
"fcvt s5, d5",
"stur s5, [x9, #-56]",
"ldur s5, [x9, #-172]",
"fsub d4, d4, d5",
"fmul d4, d4, d3",
"fcvt s4, d4",
"stur s4, [x9, #-52]",
"ldur s4, [x9, #-56]",
"fcvt d4, s4",
"ldur s5, [x9, #-52]",
"fcvt d5, s5",
"ldur s6, [x9, #-176]",
"fcvt d6, s6",
"fadd d5, d5, d6",
"fcvt s5, d5",
"stur s5, [x9, #-48]",
"fadd d4, d4, d5",
"fcvt s4, d4",
"stur s4, [x9, #-56]",
"ldur s4, [x9, #-172]",
"fcvt d4, s4",
"ldur s5, [x9, #-176]",
"fcvt d5, s5",
"ldur s6, [x9, #-172]",
"fcvt d6, s6",
"fsub d5, d5, d6",
"fmul d5, d5, d3",
"fcvt s5, d5",
"stur s5, [x9, #-44]",
"ldur s5, [x9, #-164]",
"fadd d4, d4, d5",
"fcvt s4, d4",
"stur s4, [x9, #-48]",
"ldur s4, [x9, #-176]",
"fcvt d4, s4",
"ldur s5, [x9, #-172]",
"fcvt d5, s5",
"ldur s6, [x9, #-168]",
"fcvt d6, s6",
"fadd d5, d5, d6",
"fcvt s5, d5",
"stur s5, [x9, #-40]",
"ldur s5, [x9, #-164]",
"fsub d4, d4, d5",
"fmul d4, d4, d3",
"fcvt s4, d4",
"stur s4, [x9, #-44]",
"ldur s4, [x9, #-164]",
"fcvt d4, s4",
"ldur s5, [x9, #-168]",
"fcvt d5, s5",
"ldur s6, [x9, #-168]",
"fcvt d6, s6",
"fsub d5, d5, d6",
"fmul d5, d5, d3",
"fcvt s5, d5",
"stur s5, [x9, #-36]",
"ldur s5, [x9, #-40]",
"fadd d4, d4, d5",
"fcvt s4, d4",
"stur s4, [x9, #-40]",
"ldur s4, [x9, #-164]",
"fcvt d4, s4",
"ldur s5, [x9, #-168]",
"fcvt d5, s5",
"ldur s6, [x9, #-36]",
"fcvt d6, s6",
"fadd d5, d5, d6",
"fsub d4, d4, d5",
"fmul d4, d4, d3",
"fcvt s4, d4",
"stur s4, [x9, #-36]",
"ldur s4, [x9, #-40]",
"fcvt d4, s4",
"ldur s5, [x9, #-36]",
"fcvt d5, s5",
"fadd d4, d4, d5",
"strb wzr, [x28, #1049]",
"fcvt s5, d5",
"stur s5, [x9, #-40]",
"ldur s5, [x9, #-48]",
"fcvt d5, s5",
"ldur s7, [x9, #-40]",
"fcvt d7, s7",
"fadd d5, d5, d7",
"fcvt s5, d5",
"stur s5, [x9, #-48]",
"ldur s5, [x9, #-44]",
"fcvt d5, s5",
"ldur s7, [x9, #-40]",
"fcvt d7, s7",
"fadd d5, d5, d7",
"fcvt s5, d5",
"stur s5, [x9, #-40]",
"ldur s5, [x9, #-44]",
"fcvt d5, s5",
"fadd d5, d5, d6",
"fcvt s5, d5",
"stur s5, [x9, #-44]",
"ldur s5, [x9, #-160]",
"fcvt d5, s5",
"ldur s7, [x9, #-156]",
"fcvt d7, s7",
"fadd d5, d5, d7",
"fcvt s5, d5",
"stur s5, [x9, #-32]",
"ldur s5, [x9, #-160]",
"fcvt d5, s5",
"ldur s7, [x9, #-156]",
"fcvt d7, s7",
"fsub d5, d5, d7",
"fmul d5, d5, d3",
"fcvt s5, d5",
"stur s5, [x9, #-28]",
"ldur s5, [x9, #-152]",
"fcvt d5, s5",
"ldur s7, [x9, #-148]",
"fcvt d7, s7",
"fadd d5, d5, d7",
"fcvt s5, d5",
"stur s5, [x9, #-24]",
"ldur s5, [x9, #-148]",
"fcvt d5, s5",
"ldur s7, [x9, #-152]",
"fcvt d7, s7",
"fsub d5, d5, d7",
"fmul d5, d5, d3",
"fcvt s5, d5",
"stur s5, [x9, #-20]",
"ldur s5, [x9, #-24]",
"fcvt d5, s5",
"ldur s7, [x9, #-20]",
"fcvt d7, s7",
"fadd d5, d5, d7",
"fcvt s5, d5",
"stur s5, [x9, #-24]",
"ldur s5, [x9, #-144]",
"fcvt d5, s5",
"ldur s7, [x9, #-140]",
"fcvt d7, s7",
"fadd d5, d5, d7",
"fcvt s5, d5",
"stur s5, [x9, #-16]",
"fcvt s4, d4",
"stur s4, [x9, #-40]",
"ldur s4, [x9, #-48]",
"fcvt d4, s4",
"ldur s6, [x9, #-40]",
"fcvt d6, s6",
"fadd d4, d4, d6",
"fcvt s4, d4",
"stur s4, [x9, #-48]",
"ldur s4, [x9, #-44]",
"fcvt d4, s4",
"ldur s6, [x9, #-40]",
"fcvt d6, s6",
"fadd d4, d4, d6",
"fcvt s4, d4",
"stur s4, [x9, #-40]",
"ldur s4, [x9, #-44]",
"fcvt d4, s4",
"fadd d4, d4, d5",
"fcvt s4, d4",
"stur s4, [x9, #-44]",
"ldur s4, [x9, #-160]",
"fcvt d4, s4",
"ldur s6, [x9, #-156]",
"fcvt d6, s6",
"fadd d4, d4, d6",
"fcvt s4, d4",
"stur s4, [x9, #-32]",
"ldur s4, [x9, #-160]",
"fcvt d4, s4",
"ldur s6, [x9, #-156]",
"fcvt d6, s6",
"fsub d4, d4, d6",
"fmul d4, d4, d3",
"fcvt s4, d4",
"stur s4, [x9, #-28]",
"ldur s4, [x9, #-152]",
"fcvt d4, s4",
"ldur s6, [x9, #-148]",
"fcvt d6, s6",
"fadd d4, d4, d6",
"fcvt s4, d4",
"stur s4, [x9, #-24]",
"ldur s4, [x9, #-148]",
"fcvt d4, s4",
"ldur s6, [x9, #-152]",
"fcvt d6, s6",
"fsub d4, d4, d6",
"fmul d4, d4, d3",
"fcvt s4, d4",
"stur s4, [x9, #-20]",
"ldur s4, [x9, #-24]",
"fcvt d4, s4",
"ldur s6, [x9, #-20]",
"fcvt d6, s6",
"fadd d4, d4, d6",
"fcvt s4, d4",
"stur s4, [x9, #-24]",
"ldur s4, [x9, #-144]",
"fcvt d4, s4",
"ldur s6, [x9, #-140]",
"fcvt d6, s6",
"fadd d4, d4, d6",
"fcvt s4, d4",
"stur s4, [x9, #-16]",
"ldr w4, [x9, #8]",
"ldur s5, [x9, #-144]",
"fcvt d5, s5",
"ldur s4, [x9, #-144]",
"fcvt d4, s4",
"ldr w7, [x9, #12]",
"ldur s7, [x9, #-140]",
"fcvt d7, s7",
"fsub d5, d5, d7",
"fmul d5, d5, d3",
"fcvt s5, d5",
"stur s5, [x9, #-12]",
"ldur s5, [x9, #-136]",
"fcvt d5, s5",
"ldur s7, [x9, #-132]",
"fcvt d7, s7",
"fadd d5, d5, d7",
"fcvt s5, d5",
"stur s5, [x9, #-8]",
"ldur s5, [x9, #-132]",
"fcvt d5, s5",
"ldur s7, [x9, #-136]",
"fcvt d7, s7",
"fsub d5, d5, d7",
"fmul d3, d3, d5",
"ldur s6, [x9, #-140]",
"fcvt d6, s6",
"fsub d4, d4, d6",
"fmul d4, d4, d3",
"fcvt s4, d4",
"stur s4, [x9, #-12]",
"ldur s4, [x9, #-136]",
"fcvt d4, s4",
"ldur s6, [x9, #-132]",
"fcvt d6, s6",
"fadd d4, d4, d6",
"fcvt s4, d4",
"stur s4, [x9, #-8]",
"ldur s4, [x9, #-132]",
"fcvt d4, s4",
"ldur s6, [x9, #-136]",
"fcvt d6, s6",
"fsub d4, d4, d6",
"fmul d3, d3, d4",
"strb wzr, [x28, #1049]",
"fcvt s3, d3",
"stur s3, [x9, #-4]",
"ldur s3, [x9, #-8]",
"fcvt d3, s3",
"ldur s5, [x9, #-4]",
"fcvt d7, s5",
"fadd d3, d3, d7",
"ldur s4, [x9, #-4]",
"fcvt d4, s4",
"fadd d3, d3, d4",
"strb wzr, [x28, #1049]",
"fcvt s3, d3",
"stur s3, [x9, #-8]",
"ldur s3, [x9, #-16]",
"fcvt d3, s3",
"ldur s8, [x9, #-8]",
"fcvt d8, s8",
"fadd d3, d3, d8",
"ldur s6, [x9, #-8]",
"fcvt d6, s6",
"fadd d3, d3, d6",
"fcvt s3, d3",
"stur s3, [x9, #-16]",
"ldur s3, [x9, #-12]",
"fcvt d3, s3",
"ldur s8, [x9, #-8]",
"fcvt d8, s8",
"fadd d3, d3, d8",
"ldur s6, [x9, #-8]",
"fcvt d6, s6",
"fadd d3, d3, d6",
"fcvt s3, d3",
"stur s3, [x9, #-8]",
"ldur s3, [x9, #-12]",
"fcvt d3, s3",
"fadd d3, d3, d7",
"fadd d3, d3, d4",
"fcvt s3, d3",
"stur s3, [x9, #-12]",
"ldur s3, [x9, #-128]",
@@ -2321,66 +2321,67 @@
"str s3, [x7, #768]",
"ldur s3, [x9, #-80]",
"fcvt d3, s3",
"ldur s8, [x9, #-96]",
"fcvt d8, s8",
"fadd d3, d3, d8",
"ldur s6, [x9, #-96]",
"fcvt d6, s6",
"fadd d3, d3, d6",
"fcvt s3, d3",
"stur s3, [x9, #-96]",
"ldur s3, [x9, #-96]",
"str s3, [x4, #896]",
"ldur s3, [x9, #-80]",
"fcvt d3, s3",
"ldur s8, [x9, #-88]",
"fcvt d8, s8",
"fadd d3, d3, d8",
"ldur s6, [x9, #-88]",
"fcvt d6, s6",
"fadd d3, d3, d6",
"fcvt s3, d3",
"stur s3, [x9, #-80]",
"ldur s3, [x9, #-80]",
"str s3, [x4, #640]",
"ldur s3, [x9, #-72]",
"fcvt d3, s3",
"ldur s8, [x9, #-88]",
"fcvt d8, s8",
"fadd d3, d3, d8",
"ldur s6, [x9, #-88]",
"fcvt d6, s6",
"fadd d3, d3, d6",
"fcvt s3, d3",
"stur s3, [x9, #-88]",
"ldur s3, [x9, #-88]",
"str s3, [x4, #384]",
"ldur s3, [x9, #-72]",
"fcvt d3, s3",
"ldur s8, [x9, #-92]",
"fcvt d8, s8",
"fadd d3, d3, d8",
"ldur s6, [x9, #-92]",
"fcvt d6, s6",
"fadd d3, d3, d6",
"fcvt s3, d3",
"stur s3, [x9, #-72]",
"ldur s3, [x9, #-72]",
"str s3, [x4, #128]",
"ldur s3, [x9, #-76]",
"fcvt d3, s3",
"ldur s8, [x9, #-92]",
"fcvt d8, s8",
"fadd d3, d3, d8",
"ldur s6, [x9, #-92]",
"fcvt d6, s6",
"fadd d3, d3, d6",
"fcvt s3, d3",
"stur s3, [x9, #-92]",
"ldur s3, [x9, #-92]",
"str s3, [x7, #128]",
"ldur s3, [x9, #-76]",
"fcvt d3, s3",
"ldur s8, [x9, #-84]",
"fcvt d8, s8",
"fadd d3, d3, d8",
"ldur s6, [x9, #-84]",
"fcvt d6, s6",
"fadd d3, d3, d6",
"fcvt s3, d3",
"stur s3, [x9, #-76]",
"ldur s3, [x9, #-76]",
"str s3, [x7, #384]",
"ldur s3, [x9, #-84]",
"fcvt d3, s3",
"fadd d3, d4, d3",
"fadd d3, d2, d3",
"fcvt s3, d3",
"stur s3, [x9, #-84]",
"ldur s3, [x9, #-84]",
"str s3, [x7, #640]",
"strb wzr, [x28, #1049]",
"fcvt s2, d2",
"str s2, [x7, #896]",
"ldur s2, [x9, #-32]",
"fcvt d2, s2",
@@ -2510,7 +2511,7 @@
"str s2, [x7, #448]",
"ldur s2, [x9, #-20]",
"fcvt d2, s2",
"fadd d2, d2, d7",
"fadd d2, d2, d4",
"fcvt s2, d2",
"stur s2, [x9, #-20]",
"ldur s2, [x9, #-52]",
@@ -2522,14 +2523,15 @@
"str s2, [x7, #576]",
"ldur s2, [x9, #-20]",
"fcvt d2, s2",
"fadd d2, d6, d2",
"fadd d2, d5, d2",
"fcvt s2, d2",
"str s2, [x7, #704]",
"fadd d2, d6, d7",
"fadd d2, d5, d4",
"strb wzr, [x28, #1049]",
"fcvt s2, d2",
"str s2, [x7, #832]",
"str s5, [x7, #960]",
"fcvt s2, d4",
"str s2, [x7, #960]",
"mov x8, x9",
"ldp w9, w20, [x8], #8",
"ldrb w21, [x28, #1051]",
@@ -2542,7 +2544,7 @@
"strb w21, [x28, #1202]"
],
"x86InstructionCount": 809,
"ExpectedInstructionCount": 1712
"ExpectedInstructionCount": 1714
}
}
}
@@ -17,7 +17,7 @@
"Instructions": {
"Block1": {
"x86InstructionCount": 911,
"ExpectedInstructionCount": 1695,
"ExpectedInstructionCount": 1697,
"x86Insts": [
"sub esp,0x118",
"fld dword [ecx + 0x1084]",
@@ -2364,211 +2364,213 @@
"fcvt s2, d2",
"str s2, [x8, #60]",
"ldr s2, [x8, #60]",
"fcvt d3, s2",
"ldr s4, [x8, #28]",
"fcvt d2, s2",
"ldr s3, [x8, #28]",
"fcvt d3, s3",
"fadd d4, d2, d3",
"strb wzr, [x28, #1049]",
"fcvt s4, d4",
"str s4, [x8, #196]",
"ldr s4, [x8, #196]",
"fcvt d4, s4",
"fadd d5, d3, d4",
"ldr s5, [x8, #44]",
"fcvt d5, s5",
"fadd d6, d4, d5",
"strb wzr, [x28, #1049]",
"fcvt s6, d6",
"str s6, [x8, #188]",
"ldr s6, [x8, #188]",
"fcvt d6, s6",
"ldr s7, [x8, #20]",
"fcvt d7, s7",
"fadd d6, d6, d7",
"ldr s7, [x8, #52]",
"fcvt d7, s7",
"fadd d6, d6, d7",
"strb wzr, [x28, #1049]",
"fcvt s6, d6",
"str s6, [x8, #164]",
"fadd d6, d2, d5",
"ldr s8, [x8, #12]",
"fcvt d8, s8",
"fadd d6, d6, d8",
"fcvt s6, d6",
"str s6, [x8, #180]",
"ldr s6, [x8, #180]",
"fcvt d6, s6",
"fadd d6, d6, d7",
"fcvt s6, d6",
"str s6, [x8, #172]",
"fadd d6, d2, d7",
"ldr s8, [x8, #36]",
"fcvt d8, s8",
"fadd d6, d6, d8",
"fcvt s6, d6",
"str s6, [x8, #64]",
"ldr s6, [x8, #64]",
"fcvt d6, s6",
"ldr s8, [x8, #4]",
"fcvt d8, s8",
"fadd d6, d6, d8",
"fcvt s6, d6",
"str s6, [x8, #148]",
"ldr s6, [x8, #148]",
"fcvt d6, s6",
"fneg v6.2d, v6.2d",
"fcvt s6, d6",
"str s6, [x8, #272]",
"ldr s6, [x8, #272]",
"fcvt d6, s6",
"ldr s8, [x8, #56]",
"fcvt d8, s8",
"fsub d6, d6, d8",
"strb wzr, [x28, #1049]",
"fcvt s6, d6",
"str s6, [x8, #208]",
"ldr s6, [x8, #64]",
"fcvt d6, s6",
"ldr s9, [x8, #20]",
"fcvt d9, s9",
"fadd d6, d6, d9",
"fadd d6, d6, d3",
"fcvt s6, d6",
"str s6, [x8, #156]",
"ldr s6, [x8, #156]",
"fcvt d6, s6",
"fneg v6.2d, v6.2d",
"fcvt s6, d6",
"str s6, [x8, #136]",
"ldr s6, [x8, #136]",
"fcvt d6, s6",
"ldr s9, [x8, #24]",
"fcvt d9, s9",
"fsub d6, d6, d9",
"fsub d6, d6, d8",
"fcvt s6, d6",
"str s6, [x8, #216]",
"ldr s6, [x8, #40]",
"fcvt d6, s6",
"fneg v6.2d, v6.2d",
"fsub d5, d6, d5",
"fsub d5, d5, d8",
"fsub d5, d5, d2",
"strb wzr, [x28, #1049]",
"fcvt s5, d5",
"str s5, [x8, #196]",
"ldr s5, [x8, #196]",
"fcvt d6, s5",
"ldr s7, [x8, #44]",
"str s5, [x8, #64]",
"ldr s5, [x8, #64]",
"fcvt d5, s5",
"fsub d6, d5, d7",
"ldr s7, [x8, #8]",
"fcvt d7, s7",
"fadd d8, d6, d7",
"strb wzr, [x28, #1049]",
"fcvt s8, d8",
"str s8, [x8, #188]",
"ldr s8, [x8, #188]",
"fcvt d8, s8",
"ldr s9, [x8, #20]",
"fsub d7, d6, d7",
"ldr s9, [x8, #12]",
"fcvt d9, s9",
"fadd d8, d8, d9",
"ldr s9, [x8, #52]",
"fcvt d9, s9",
"fadd d8, d8, d9",
"strb wzr, [x28, #1049]",
"fcvt s8, d8",
"str s8, [x8, #164]",
"fadd d8, d3, d7",
"ldr s10, [x8, #12]",
"fcvt d10, s10",
"fadd d8, d8, d10",
"fcvt s8, d8",
"str s8, [x8, #180]",
"ldr s8, [x8, #180]",
"fcvt d8, s8",
"fadd d8, d8, d9",
"fcvt s8, d8",
"str s8, [x8, #172]",
"fadd d8, d3, d9",
"ldr s10, [x8, #36]",
"fcvt d10, s10",
"fadd d8, d8, d10",
"fcvt s8, d8",
"str s8, [x8, #64]",
"ldr s8, [x8, #64]",
"fcvt d8, s8",
"ldr s10, [x8, #4]",
"fcvt d10, s10",
"fadd d8, d8, d10",
"fcvt s8, d8",
"str s8, [x8, #148]",
"ldr s8, [x8, #148]",
"fcvt d8, s8",
"fneg v8.2d, v8.2d",
"fcvt s8, d8",
"str s8, [x8, #272]",
"ldr s8, [x8, #272]",
"fcvt d8, s8",
"ldr s10, [x8, #56]",
"fcvt d10, s10",
"fsub d8, d8, d10",
"strb wzr, [x28, #1049]",
"fcvt s8, d8",
"str s8, [x8, #208]",
"ldr s8, [x8, #64]",
"fcvt d8, s8",
"ldr s11, [x8, #20]",
"fcvt d11, s11",
"fadd d8, d8, d11",
"fadd d8, d8, d4",
"fcvt s8, d8",
"str s8, [x8, #156]",
"ldr s8, [x8, #156]",
"fcvt d8, s8",
"fneg v8.2d, v8.2d",
"fcvt s8, d8",
"str s8, [x8, #136]",
"ldr s8, [x8, #136]",
"fcvt d8, s8",
"ldr s11, [x8, #24]",
"fcvt d11, s11",
"fsub d8, d8, d11",
"fsub d8, d8, d10",
"fcvt s8, d8",
"str s8, [x8, #216]",
"ldr s8, [x8, #40]",
"fcvt d8, s8",
"fneg v8.2d, v8.2d",
"fsub d7, d8, d7",
"fsub d7, d7, d10",
"fsub d7, d7, d3",
"strb wzr, [x28, #1049]",
"fsub d7, d7, d9",
"fcvt s7, d7",
"str s7, [x8, #64]",
"ldr s7, [x8, #64]",
"str s7, [x8, #232]",
"ldr s7, [x8, #20]",
"fcvt d7, s7",
"fsub d8, d7, d9",
"ldr s9, [x8, #8]",
"fcvt d9, s9",
"fsub d9, d8, d9",
"ldr s11, [x8, #12]",
"fcvt d11, s11",
"fsub d9, d9, d11",
"fcvt s9, d9",
"str s9, [x8, #232]",
"ldr s9, [x8, #20]",
"fcvt d9, s9",
"fsub d8, d8, d9",
"ldr s9, [x8, #24]",
"fcvt d9, s9",
"fsub d8, d8, d9",
"fsub d8, d8, d4",
"fsub d6, d6, d7",
"ldr s7, [x8, #24]",
"fcvt d7, s7",
"fsub d6, d6, d7",
"fsub d6, d6, d3",
"mov w20, #0x0",
"strb wzr, [x28, #1049]",
"fcvt s8, d8",
"str s8, [x8, #224]",
"ldr s8, [x8, #48]",
"fcvt d8, s8",
"fsub d7, d7, d8",
"ldr s9, [x8, #8]",
"fcvt s6, d6",
"str s6, [x8, #224]",
"ldr s6, [x8, #48]",
"fcvt d6, s6",
"fsub d5, d5, d6",
"ldr s7, [x8, #8]",
"fcvt d7, s7",
"fsub d7, d5, d7",
"ldr s9, [x8, #12]",
"fcvt d9, s9",
"fsub d9, d7, d9",
"ldr s11, [x8, #12]",
"fcvt d11, s11",
"fsub d9, d9, d11",
"fcvt s9, d9",
"str s9, [x8, #240]",
"ldr s9, [x8, #24]",
"fsub d7, d7, d9",
"fcvt s7, d7",
"str s7, [x8, #240]",
"ldr s7, [x8, #24]",
"fcvt d7, s7",
"ldr s9, [x8, #16]",
"fcvt d9, s9",
"ldr s11, [x8, #16]",
"fcvt d11, s11",
"fadd d9, d9, d11",
"fadd d4, d4, d9",
"fadd d7, d7, d9",
"fadd d3, d3, d7",
"ldr w4, [x7, #4100]",
"ldr w5, [x7, #4096]",
"strb wzr, [x28, #1049]",
"add w4, w5, w4, lsl #2",
"fcvt s4, d4",
"str s4, [x8, #64]",
"ldr s4, [x8, #64]",
"fcvt d4, s4",
"fsub d7, d7, d4",
"fcvt s3, d3",
"str s3, [x8, #64]",
"ldr s3, [x8, #64]",
"fcvt d3, s3",
"fsub d5, d5, d3",
"strb wzr, [x28, #1049]",
"fcvt s7, d7",
"str s7, [x8, #248]",
"ldr s7, [x8, #32]",
"fcvt d7, s7",
"fneg v7.2d, v7.2d",
"fsub d7, d7, d8",
"fcvt s5, d5",
"str s5, [x8, #248]",
"ldr s5, [x8, #32]",
"fcvt d5, s5",
"fneg v5.2d, v5.2d",
"fsub d5, d5, d6",
"strb wzr, [x28, #1049]",
"fsub d7, d7, d10",
"fsub d7, d7, d3",
"fcvt s7, d7",
"str s7, [x8, #64]",
"ldr s7, [x8, #64]",
"fcvt d7, s7",
"ldr s8, [x8]",
"fcvt d8, s8",
"fsub d8, d7, d8",
"fcvt s8, d8",
"str s8, [x8, #264]",
"fsub d4, d7, d4",
"fsub d5, d5, d8",
"fsub d5, d5, d2",
"fcvt s5, d5",
"str s5, [x8, #64]",
"ldr s5, [x8, #64]",
"fcvt d5, s5",
"ldr s6, [x8]",
"fcvt d6, s6",
"fsub d6, d5, d6",
"fcvt s6, d6",
"str s6, [x8, #264]",
"fsub d3, d5, d3",
"strb wzr, [x28, #1049]",
"fcvt s4, d4",
"str s4, [x8, #256]",
"ldr s4, [x8, #144]",
"str s4, [x4]",
"ldr s4, [x8, #148]",
"str s4, [x4, #64]",
"ldr s4, [x8, #152]",
"str s4, [x4, #128]",
"ldr s4, [x8, #156]",
"str s4, [x4, #192]",
"ldr s4, [x8, #160]",
"str s4, [x4, #256]",
"ldr s4, [x8, #164]",
"fcvt d7, s4",
"str s4, [x4, #320]",
"ldr s4, [x8, #168]",
"fcvt d8, s4",
"str s4, [x4, #384]",
"ldr s4, [x8, #172]",
"fcvt d9, s4",
"str s4, [x4, #448]",
"ldr s4, [x8, #176]",
"fcvt d10, s4",
"str s4, [x4, #512]",
"ldr s4, [x8, #180]",
"str s4, [x4, #576]",
"ldr s4, [x8, #184]",
"fcvt d11, s4",
"str s4, [x4, #640]",
"ldr s4, [x8, #188]",
"str s4, [x4, #704]",
"ldr s4, [x8, #192]",
"str s4, [x4, #768]",
"fcvt s3, d3",
"str s3, [x8, #256]",
"ldr s3, [x8, #144]",
"str s3, [x4]",
"ldr s3, [x8, #148]",
"str s3, [x4, #64]",
"ldr s3, [x8, #152]",
"str s3, [x4, #128]",
"ldr s3, [x8, #156]",
"str s3, [x4, #192]",
"ldr s3, [x8, #160]",
"str s3, [x4, #256]",
"ldr s3, [x8, #164]",
"fcvt d5, s3",
"str s3, [x4, #320]",
"ldr s3, [x8, #168]",
"fcvt d6, s3",
"str s3, [x4, #384]",
"ldr s3, [x8, #172]",
"fcvt d7, s3",
"str s3, [x4, #448]",
"ldr s3, [x8, #176]",
"fcvt d8, s3",
"str s3, [x4, #512]",
"ldr s3, [x8, #180]",
"str s3, [x4, #576]",
"ldr s3, [x8, #184]",
"fcvt d9, s3",
"str s3, [x4, #640]",
"ldr s3, [x8, #188]",
"str s3, [x4, #704]",
"ldr s3, [x8, #192]",
"str s3, [x4, #768]",
"strb wzr, [x28, #1049]",
"str s5, [x4, #832]",
"ldr s4, [x8, #200]",
"str s4, [x4, #896]",
"fcvt s3, d4",
"str s3, [x4, #832]",
"ldr s3, [x8, #200]",
"str s3, [x4, #896]",
"strb wzr, [x28, #1049]",
"str s2, [x4, #960]",
"fmov d2, x20",
"fcvt s2, d2",
"str s2, [x4, #1024]",
"fneg v2.2d, v3.2d",
"fcvt s3, d2",
"str s3, [x4, #960]",
"fmov d3, x20",
"fcvt s3, d3",
"str s3, [x4, #1024]",
"fneg v2.2d, v2.2d",
"fcvt s2, d2",
"str s2, [x4, #1088]",
"ldr s2, [x8, #200]",
@@ -2577,7 +2579,7 @@
"fcvt s2, d2",
"str s2, [x4, #1152]",
"strb wzr, [x28, #1049]",
"fneg v2.2d, v6.2d",
"fneg v2.2d, v4.2d",
"fcvt s2, d2",
"str s2, [x4, #1216]",
"ldr s2, [x8, #192]",
@@ -2591,7 +2593,7 @@
"fcvt s2, d2",
"str s2, [x4, #1344]",
"strb wzr, [x28, #1049]",
"fneg v2.2d, v11.2d",
"fneg v2.2d, v9.2d",
"fcvt s2, d2",
"str s2, [x4, #1408]",
"ldr s2, [x8, #180]",
@@ -2600,18 +2602,18 @@
"fcvt s2, d2",
"str s2, [x4, #1472]",
"strb wzr, [x28, #1049]",
"fneg v2.2d, v10.2d",
"fneg v2.2d, v8.2d",
"fcvt s2, d2",
"str s2, [x4, #1536]",
"strb wzr, [x28, #1049]",
"fneg v2.2d, v9.2d",
"fneg v2.2d, v7.2d",
"fcvt s2, d2",
"str s2, [x4, #1600]",
"strb wzr, [x28, #1049]",
"fneg v2.2d, v8.2d",
"fneg v2.2d, v6.2d",
"fcvt s2, d2",
"str s2, [x4, #1664]",
"fneg v2.2d, v7.2d",
"fneg v2.2d, v5.2d",
"fcvt s2, d2",
"str s2, [x4, #1728]",
"ldr s2, [x8, #140]",
@@ -4207,7 +4209,7 @@
},
"Block3": {
"x86InstructionCount": 649,
"ExpectedInstructionCount": 958,
"ExpectedInstructionCount": 982,
"x86Insts": [
"fld dword [esi + 0x64]",
"mov eax,dword [esi + 0x88]",
@@ -4937,59 +4939,62 @@
"fcvt s3, d3",
"str s3, [x8, #40]",
"ldr s3, [x8, #40]",
"fcvt d4, s3",
"str s3, [x8, #16]",
"ldr s4, [x8, #80]",
"fcvt d4, s4",
"fmul d4, d4, d2",
"fcvt s4, d4",
"str s4, [x8, #40]",
"ldr s4, [x8, #40]",
"str s4, [x8, #44]",
"ldr s5, [x8, #88]",
"fcvt d5, s5",
"fmul d5, d5, d2",
"fcvt s5, d5",
"str s5, [x8, #40]",
"ldr s5, [x8, #40]",
"str s5, [x8, #20]",
"ldr s6, [x8, #140]",
"fcvt d6, s6",
"fmul d6, d6, d2",
"fcvt s6, d6",
"str s6, [x8, #76]",
"ldr s6, [x8, #76]",
"str s6, [x8, #68]",
"ldr s6, [x8, #124]",
"fcvt d6, s6",
"fmul d6, d6, d2",
"fcvt s6, d6",
"str s6, [x8, #140]",
"ldr s6, [x8, #140]",
"str s6, [x8, #72]",
"ldr s6, [x8, #132]",
"fcvt d6, s6",
"fmul d6, d6, d2",
"fcvt s6, d6",
"str s6, [x8, #88]",
"ldr s6, [x8, #88]",
"str s6, [x8, #64]",
"ldr s6, [x8, #92]",
"fcvt d6, s6",
"fmul d6, d6, d2",
"fcvt s6, d6",
"str s6, [x8, #80]",
"ldr s6, [x8, #80]",
"str s6, [x8, #132]",
"ldr s6, [x8, #96]",
"fcvt d6, s6",
"fmul d6, d6, d2",
"fcvt s6, d6",
"str s6, [x8, #84]",
"ldr s6, [x8, #84]",
"str s6, [x8, #124]",
"ldr s6, [x8, #100]",
"fcvt d6, s6",
"fmul d2, d2, d6",
"ldr s3, [x8, #80]",
"fcvt d3, s3",
"fmul d3, d3, d2",
"fcvt s3, d3",
"str s3, [x8, #40]",
"ldr s3, [x8, #40]",
"fcvt d5, s3",
"str s3, [x8, #44]",
"ldr s3, [x8, #88]",
"fcvt d3, s3",
"fmul d3, d3, d2",
"fcvt s3, d3",
"str s3, [x8, #40]",
"ldr s3, [x8, #40]",
"fcvt d6, s3",
"str s3, [x8, #20]",
"ldr s3, [x8, #140]",
"fcvt d3, s3",
"fmul d3, d3, d2",
"fcvt s3, d3",
"str s3, [x8, #76]",
"ldr s3, [x8, #76]",
"str s3, [x8, #68]",
"ldr s3, [x8, #124]",
"fcvt d3, s3",
"fmul d3, d3, d2",
"fcvt s3, d3",
"str s3, [x8, #140]",
"ldr s3, [x8, #140]",
"str s3, [x8, #72]",
"ldr s3, [x8, #132]",
"fcvt d3, s3",
"fmul d3, d3, d2",
"fcvt s3, d3",
"str s3, [x8, #88]",
"ldr s3, [x8, #88]",
"str s3, [x8, #64]",
"ldr s3, [x8, #92]",
"fcvt d3, s3",
"fmul d3, d3, d2",
"fcvt s3, d3",
"str s3, [x8, #80]",
"ldr s3, [x8, #80]",
"str s3, [x8, #132]",
"ldr s3, [x8, #96]",
"fcvt d3, s3",
"fmul d3, d3, d2",
"fcvt s3, d3",
"str s3, [x8, #84]",
"ldr s3, [x8, #84]",
"str s3, [x8, #124]",
"ldr s3, [x8, #100]",
"fcvt d3, s3",
"fmul d2, d2, d3",
"strb wzr, [x28, #1049]",
"fcvt s2, d2",
"str s2, [x8, #40]",
@@ -4997,16 +5002,16 @@
"str s2, [x8, #92]",
"ldr s2, [x8, #148]",
"fcvt d2, s2",
"ldr s6, [x8, #132]",
"fcvt d6, s6",
"fadd d6, d6, d2",
"fcvt s6, d6",
"str s6, [x8, #132]",
"ldr s6, [x8, #152]",
"fcvt d6, s6",
"ldr s3, [x8, #132]",
"fcvt d3, s3",
"fadd d3, d3, d2",
"fcvt s3, d3",
"str s3, [x8, #132]",
"ldr s3, [x8, #152]",
"fcvt d3, s3",
"ldr s7, [x8, #124]",
"fcvt d7, s7",
"fadd d7, d7, d6",
"fadd d7, d7, d3",
"fcvt s7, d7",
"str s7, [x8, #124]",
"ldr s7, [x8, #156]",
@@ -5071,11 +5076,14 @@
"ldr w5, [x8, #100]",
"strb wzr, [x28, #1049]",
"str w5, [x8, #252]",
"str s3, [x8, #92]",
"fcvt s8, d4",
"str s8, [x8, #92]",
"strb wzr, [x28, #1049]",
"str s4, [x8, #132]",
"fcvt s8, d5",
"str s8, [x8, #132]",
"strb wzr, [x28, #1049]",
"str s5, [x8, #124]",
"fcvt s8, d6",
"str s8, [x8, #124]",
"ldr s8, [x8, #76]",
"str s8, [x8, #64]",
"ldr s8, [x8, #140]",
@@ -5095,7 +5103,7 @@
"str s8, [x8, #20]",
"ldr s8, [x8, #44]",
"fcvt d8, s8",
"fadd d8, d8, d6",
"fadd d8, d8, d3",
"fcvt s8, d8",
"str s8, [x8, #44]",
"ldr s8, [x8, #16]",
@@ -5158,11 +5166,14 @@
"ldr w5, [x8, #28]",
"strb wzr, [x28, #1049]",
"str w5, [x8, #264]",
"str s3, [x8, #92]",
"fcvt s8, d4",
"str s8, [x8, #92]",
"strb wzr, [x28, #1049]",
"str s4, [x8, #132]",
"fcvt s8, d5",
"str s8, [x8, #132]",
"strb wzr, [x28, #1049]",
"str s5, [x8, #124]",
"fcvt s8, d6",
"str s8, [x8, #124]",
"ldr s8, [x8, #76]",
"str s8, [x8, #64]",
"ldr s8, [x8, #140]",
@@ -5182,7 +5193,7 @@
"str s8, [x8, #20]",
"ldr s8, [x8, #44]",
"fcvt d8, s8",
"fadd d8, d8, d6",
"fadd d8, d8, d3",
"fcvt s8, d8",
"str s8, [x8, #44]",
"ldr s8, [x8, #16]",
@@ -5245,32 +5256,35 @@
"ldr w5, [x8, #28]",
"strb wzr, [x28, #1049]",
"str w5, [x8, #276]",
"str s3, [x8, #92]",
"fcvt s4, d4",
"str s4, [x8, #92]",
"strb wzr, [x28, #1049]",
"fcvt s4, d5",
"str s4, [x8, #132]",
"strb wzr, [x28, #1049]",
"str s5, [x8, #124]",
"ldr s3, [x8, #76]",
"str s3, [x8, #64]",
"ldr s3, [x8, #140]",
"str s3, [x8, #72]",
"ldr s3, [x8, #88]",
"str s3, [x8, #68]",
"ldr s3, [x8, #80]",
"str s3, [x8, #20]",
"ldr s3, [x8, #84]",
"str s3, [x8, #44]",
"ldr s3, [x8, #40]",
"str s3, [x8, #16]",
"ldr s3, [x8, #20]",
"fcvt d3, s3",
"fadd d2, d2, d3",
"fcvt s4, d6",
"str s4, [x8, #124]",
"ldr s4, [x8, #76]",
"str s4, [x8, #64]",
"ldr s4, [x8, #140]",
"str s4, [x8, #72]",
"ldr s4, [x8, #88]",
"str s4, [x8, #68]",
"ldr s4, [x8, #80]",
"str s4, [x8, #20]",
"ldr s4, [x8, #84]",
"str s4, [x8, #44]",
"ldr s4, [x8, #40]",
"str s4, [x8, #16]",
"ldr s4, [x8, #20]",
"fcvt d4, s4",
"fadd d2, d2, d4",
"strb wzr, [x28, #1049]",
"fcvt s2, d2",
"str s2, [x8, #20]",
"ldr s2, [x8, #44]",
"fcvt d2, s2",
"fadd d2, d6, d2",
"fadd d2, d3, d2",
"strb wzr, [x28, #1049]",
"fcvt s2, d2",
"str s2, [x8, #44]",
@@ -5408,59 +5422,62 @@
"fcvt s3, d3",
"str s3, [x8, #20]",
"ldr s3, [x8, #20]",
"fcvt d4, s3",
"str s3, [x8, #92]",
"ldr s4, [x8, #44]",
"fcvt d4, s4",
"fmul d4, d4, d2",
"fcvt s4, d4",
"str s4, [x8, #20]",
"ldr s4, [x8, #20]",
"str s4, [x8, #132]",
"ldr s5, [x8, #16]",
"fcvt d5, s5",
"fmul d5, d5, d2",
"fcvt s5, d5",
"str s5, [x8, #20]",
"ldr s5, [x8, #20]",
"str s5, [x8, #124]",
"ldr s6, [x8, #64]",
"fcvt d6, s6",
"fmul d6, d6, d2",
"fcvt s6, d6",
"str s6, [x8, #40]",
"ldr s6, [x8, #40]",
"str s6, [x8, #64]",
"ldr s6, [x8, #72]",
"fcvt d6, s6",
"fmul d6, d6, d2",
"fcvt s6, d6",
"str s6, [x8, #84]",
"ldr s6, [x8, #84]",
"str s6, [x8, #72]",
"ldr s6, [x8, #68]",
"fcvt d6, s6",
"fmul d6, d6, d2",
"fcvt s6, d6",
"str s6, [x8, #80]",
"ldr s6, [x8, #80]",
"str s6, [x8, #68]",
"ldr s6, [x8, #112]",
"fcvt d6, s6",
"fmul d6, d6, d2",
"fcvt s6, d6",
"str s6, [x8, #88]",
"ldr s6, [x8, #88]",
"str s6, [x8, #20]",
"ldr s6, [x8, #116]",
"fcvt d6, s6",
"fmul d6, d6, d2",
"fcvt s6, d6",
"str s6, [x8, #140]",
"ldr s6, [x8, #140]",
"str s6, [x8, #44]",
"ldr s6, [x8, #120]",
"fcvt d6, s6",
"fmul d2, d2, d6",
"ldr s3, [x8, #44]",
"fcvt d3, s3",
"fmul d3, d3, d2",
"fcvt s3, d3",
"str s3, [x8, #20]",
"ldr s3, [x8, #20]",
"fcvt d5, s3",
"str s3, [x8, #132]",
"ldr s3, [x8, #16]",
"fcvt d3, s3",
"fmul d3, d3, d2",
"fcvt s3, d3",
"str s3, [x8, #20]",
"ldr s3, [x8, #20]",
"fcvt d6, s3",
"str s3, [x8, #124]",
"ldr s3, [x8, #64]",
"fcvt d3, s3",
"fmul d3, d3, d2",
"fcvt s3, d3",
"str s3, [x8, #40]",
"ldr s3, [x8, #40]",
"str s3, [x8, #64]",
"ldr s3, [x8, #72]",
"fcvt d3, s3",
"fmul d3, d3, d2",
"fcvt s3, d3",
"str s3, [x8, #84]",
"ldr s3, [x8, #84]",
"str s3, [x8, #72]",
"ldr s3, [x8, #68]",
"fcvt d3, s3",
"fmul d3, d3, d2",
"fcvt s3, d3",
"str s3, [x8, #80]",
"ldr s3, [x8, #80]",
"str s3, [x8, #68]",
"ldr s3, [x8, #112]",
"fcvt d3, s3",
"fmul d3, d3, d2",
"fcvt s3, d3",
"str s3, [x8, #88]",
"ldr s3, [x8, #88]",
"str s3, [x8, #20]",
"ldr s3, [x8, #116]",
"fcvt d3, s3",
"fmul d3, d3, d2",
"fcvt s3, d3",
"str s3, [x8, #140]",
"ldr s3, [x8, #140]",
"str s3, [x8, #44]",
"ldr s3, [x8, #120]",
"fcvt d3, s3",
"fmul d2, d2, d3",
"mov w20, #0x0",
"strb wzr, [x28, #1049]",
"fcvt s2, d2",
@@ -5469,16 +5486,16 @@
"str s2, [x8, #16]",
"ldr s2, [x8, #148]",
"fcvt d2, s2",
"ldr s6, [x8, #20]",
"fcvt d6, s6",
"fadd d6, d6, d2",
"fcvt s6, d6",
"str s6, [x8, #20]",
"ldr s6, [x8, #152]",
"fcvt d6, s6",
"ldr s3, [x8, #20]",
"fcvt d3, s3",
"fadd d3, d3, d2",
"fcvt s3, d3",
"str s3, [x8, #20]",
"ldr s3, [x8, #152]",
"fcvt d3, s3",
"ldr s7, [x8, #44]",
"fcvt d7, s7",
"fadd d7, d7, d6",
"fadd d7, d7, d3",
"fcvt s7, d7",
"str s7, [x8, #44]",
"ldr s7, [x8, #156]",
@@ -5543,11 +5560,14 @@
"ldr w5, [x8, #120]",
"strb wzr, [x28, #1049]",
"str w5, [x8, #192]",
"str s3, [x8, #92]",
"fcvt s8, d4",
"str s8, [x8, #92]",
"strb wzr, [x28, #1049]",
"str s4, [x8, #132]",
"fcvt s8, d5",
"str s8, [x8, #132]",
"strb wzr, [x28, #1049]",
"str s5, [x8, #124]",
"fcvt s8, d6",
"str s8, [x8, #124]",
"ldr s8, [x8, #40]",
"str s8, [x8, #64]",
"ldr s8, [x8, #84]",
@@ -5567,7 +5587,7 @@
"str s8, [x8, #20]",
"ldr s8, [x8, #44]",
"fcvt d8, s8",
"fadd d8, d8, d6",
"fadd d8, d8, d3",
"fcvt s8, d8",
"str s8, [x8, #44]",
"ldr s8, [x8, #16]",
@@ -5630,11 +5650,14 @@
"ldr w5, [x8, #120]",
"strb wzr, [x28, #1049]",
"str w5, [x8, #204]",
"str s3, [x8, #92]",
"fcvt s8, d4",
"str s8, [x8, #92]",
"strb wzr, [x28, #1049]",
"str s4, [x8, #132]",
"fcvt s8, d5",
"str s8, [x8, #132]",
"strb wzr, [x28, #1049]",
"str s5, [x8, #124]",
"fcvt s8, d6",
"str s8, [x8, #124]",
"ldr s8, [x8, #40]",
"str s8, [x8, #64]",
"ldr s8, [x8, #84]",
@@ -5654,7 +5677,7 @@
"str s8, [x8, #20]",
"ldr s8, [x8, #44]",
"fcvt d8, s8",
"fadd d8, d8, d6",
"fadd d8, d8, d3",
"fcvt s8, d8",
"str s8, [x8, #44]",
"ldr s8, [x8, #16]",
@@ -5718,32 +5741,35 @@
"str w5, [x8, #216]",
"strb wzr, [x28, #1049]",
"str w20, [x8, #-4]!",
"str s3, [x8, #96]",
"fcvt s4, d4",
"str s4, [x8, #96]",
"strb wzr, [x28, #1049]",
"fcvt s4, d5",
"str s4, [x8, #136]",
"strb wzr, [x28, #1049]",
"str s5, [x8, #128]",
"ldr s3, [x8, #44]",
"str s3, [x8, #68]",
"ldr s3, [x8, #88]",
"str s3, [x8, #76]",
"ldr s3, [x8, #84]",
"str s3, [x8, #72]",
"ldr s3, [x8, #92]",
"str s3, [x8, #24]",
"ldr s3, [x8, #144]",
"str s3, [x8, #48]",
"ldr s3, [x8, #80]",
"str s3, [x8, #20]",
"ldr s3, [x8, #24]",
"fcvt d3, s3",
"fadd d2, d2, d3",
"fcvt s4, d6",
"str s4, [x8, #128]",
"ldr s4, [x8, #44]",
"str s4, [x8, #68]",
"ldr s4, [x8, #88]",
"str s4, [x8, #76]",
"ldr s4, [x8, #84]",
"str s4, [x8, #72]",
"ldr s4, [x8, #92]",
"str s4, [x8, #24]",
"ldr s4, [x8, #144]",
"str s4, [x8, #48]",
"ldr s4, [x8, #80]",
"str s4, [x8, #20]",
"ldr s4, [x8, #24]",
"fcvt d4, s4",
"fadd d2, d2, d4",
"strb wzr, [x28, #1049]",
"fcvt s2, d2",
"str s2, [x8, #24]",
"ldr s2, [x8, #48]",
"fcvt d2, s2",
"fadd d2, d6, d2",
"fadd d2, d3, d2",
"strb wzr, [x28, #1049]",
"fcvt s2, d2",
"str s2, [x8, #48]",
@@ -2786,7 +2786,7 @@
"mov x0, x5",
"mov x1, x4",
"mov x2, x6",
"ldr x3, [x28, #3568]",
"ldr x3, [x28, #3584]",
"str x30, [sp, #-16]!",
"blr x3",
"ldr x30, [sp], #16",
@@ -2837,7 +2837,7 @@
"mov x0, x5",
"mov x1, x4",
"mov x2, x6",
"ldr x3, [x28, #3576]",
"ldr x3, [x28, #3592]",
"str x30, [sp, #-16]!",
"blr x3",
"ldr x30, [sp], #16",