Compare commits

...
65 Commits
Author SHA1 Message Date
Ryan Houdek 37f1e55ed5 Docs: Update for release FEX-2204 2022-04-19 01:19:00 -07:00
Ryan Houdek 8ad14728f6 Merge pull request #1644 from Sonicadvance1/ldiv_minor_opt
JITArm64: Get long divide out of the hot path
2022-04-01 18:23:30 -07:00
Ryan Houdek 6b3cd3d31d Merge pull request #1645 from Sonicadvance1/update_aarch64_fit
Scripts: Updates AArch64 fit for Clang 14
2022-04-01 18:23:13 -07:00
Ryan Houdek fba698cb74 Scripts: Updates AArch64 fit for Clang 14
Clang now supports these latest ARMv9 CPUs
2022-04-01 18:08:02 -07:00
Ryan Houdek 0946b123bb JITArm64: Get long divide out of the hot path
For 128-bit divides, we can very quickly check at runtime if we can
avoid the long divide and just do a 64-bit divide.

For unsigned just check if the top bits are all zero.
For signed just check if the top bits match bit 63 of the lower bits.

Additionally, keep the long divide handlers inside of the dispatcher.
This keeps the majority of the code bloat out of the code block itself,
significantly reducing block size for something doing these divides.
Also a fairly large icache improvement from this.

Hard performance number improvements here are hard to get since it
heavily depends on the application, also only occurs on x86-64.

Seems to have helped FTL and Dead Cells performance quite a bit though.
2022-03-31 09:33:19 -07:00
Ryan Houdek b43937a7a1 Merge pull request #1643 from Sonicadvance1/fix_termux
SignalDelegator: Adds missing include
2022-03-29 21:05:28 -07:00
Ryan Houdek 4564eba20d SignalDelegator: Adds missing include
Fixes Termux building.
Fixes #1642
2022-03-29 20:47:22 -07:00
Ryan Houdek 5cc0c0a3da Merge pull request #1641 from philpax/docs-remove-stale-text
docs: Remove stale text
2022-03-29 02:56:20 -07:00
Philpax f8e7c75f86 docs: Remove stale text 2022-03-29 11:27:43 +02:00
Ryan Houdek 042cd354dc Merge pull request #1633 from Sonicadvance1/disable_instructions_on_host_missing
OpcodeDispatcher: Fixes SIGILL on unsupported host instructions
2022-03-23 13:48:04 -07:00
Ryan Houdek 977bda97b2 Merge pull request #1635 from Sonicadvance1/4000_0001h
CPUID: Adds 4000_0001h function
2022-03-23 13:42:45 -07:00
Ryan Houdek 4cf48ca9bb CPUID: Adds 4000_0001h function
Exposes the host architecture through this CPUID function. Only exposes
the architectures we support. Not burning 16-bits on using ELF machine
definitions here.

Uses 4 bits still for future expansion.
2022-03-22 16:53:44 -07:00
Mai M a247df50ea Merge pull request #1624 from Sonicadvance1/cleanup_ir_after_use
FEXCore: Delete IR after it is used
2022-03-22 13:23:05 -04:00
Mai M 1f1c214944 Merge pull request #1634 from Sonicadvance1/cpuid_documentation
Documentation: Adds hypervisor CPUID information
2022-03-22 12:57:15 -04:00
Ryan Houdek d16db4ebde OpcodeDispatcher: Fixes SIGILL on unsupported host instructions
If the host doesn't support the instructions required for implementing
an instruction then don't even add them to the opcodedispatcher.

This means that we will never try emitting instructions that the host
doesn't support (For these instructions anyway) and successfully passes
the guest SIGILL for these particular instructions.

Fixes #1631
2022-03-21 23:03:04 -07:00
Ryan Houdek ae1c563082 Documentation: Adds hypervisor CPUID information
Currently we only implement function 4000_0000h. This will expand in the
future but this is all we have right now.
2022-03-21 22:46:48 -07:00
Ryan Houdek ebd0edbab7 Merge pull request #1632 from FEX-Emu/skmp/flush-test-harness
TestHarnessRunner: Flush log on asserts
2022-03-21 12:50:43 -07:00
Stefanos Kornilios Misis Poiitidis e87e9d269a TestHarnessRunner: Flush log on asserts 2022-03-21 21:32:35 +02:00
Ryan Houdek 187c64182b Merge pull request #1628 from Sonicadvance1/fix_finit
OpcodeDispatcher: Fixes FNINIT
2022-03-17 20:37:57 -07:00
Ryan Houdek 60c7ea6e5f Merge pull request #1620 from Sonicadvance1/fix_1618
FEXCore: Fixes #1618
2022-03-17 20:36:06 -07:00
Ryan Houdek 6f1b4b0eee OpcodeDispatcher: Fixes FNINIT
Was incorrectly setting the FCW to 037h when it was supposed to be
037Fh.

Fixes a bug in a visual novel where its CPUID state wouldn't initialize
if this was set incorrectly.
2022-03-17 20:27:22 -07:00
Ryan Houdek fb69300397 FEXCore: Delete IR after it is used
For the JIT cores we don't need to keep IR around, it's only necessary
for the Interpreter. So once the AOT IR service is done dealing with the
IR, check to see if we can delete it.

This causes teeworld's title screen memory usage to go from 730MB to
566MB. 77.5% the memory usage there.

This is effectively an infinite memory leak if the codespace wasn't ever
overwritten or invalidated. So larger memory usage programs would end up
having a larger impact.
2022-03-13 19:01:40 -07:00
Ryan Houdek 5677924525 Merge pull request #1621 from Sonicadvance1/fix_1584
Softfloat: Fixes FSCALE
2022-03-13 18:57:06 -07:00
Ryan Houdek 8422fc632d Merge pull request #1623 from Sonicadvance1/remove_unused_debug_data
FEXCore: Removes unused debug data
2022-03-13 18:56:50 -07:00
Ryan Houdek d33cd744fb FEXCore: Removes unused debug data
This isn't used anywhere. Just remove these.
If we get the imgui debugger running again then we can add even more
stats to sort block costs by.
2022-03-13 18:40:49 -07:00
Ryan Houdek 3b0fb27ae9 Softfloat: Fixes FSCALE
I misread the implementation details of this instruction when
implementing.

The pseudocode says `ST(0) = ST(0) ∗ 2^rndint(ST(1))` so I understood
the instruction to use the current rounding mode of the host to extract
the integer portion of `ST(1)`.

The actual implementation is in the details of the statement `the
integer portion of the floating- point value in ST(1).`

This behaves like round towards zero/truncate, additional hardware
testing and documentation reading confirms this.

Fixes #1584
2022-03-13 14:11:31 -07:00
Ryan Houdek 4603e09a04 FEXCore: Fixes #1618 2022-03-13 13:42:37 -07:00
Ryan Houdek 7b0265ffe2 Merge pull request #1617 from Sonicadvance1/gdbstub_improvements4
GDBServer improvements: Three's a crowd
2022-03-13 13:24:41 -07:00
Ryan Houdek ec54560a38 GDBServer: Fixes memory reading
memory-map is not something we want to use. Adds a comment about it and
disables it.

Also changes core events to wait for an event from GDBStub for waking up
which fixes a hang.
2022-03-13 13:05:00 -07:00
Ryan Houdek 6a5abd3672 Merge pull request #1616 from Sonicadvance1/gdbstub_improvements3
Gdbstub improvements: The sequel
2022-03-13 13:03:25 -07:00
Ryan Houdek 3e6af39c42 GDBServer: Zero initialize some variables to fix connection stability
Otherwise you always had to attempt connecting twice in a row
2022-03-13 12:50:01 -07:00
Ryan Houdek cf82ffc052 GDBServer: Let gdb know when the library map has updated
We need to fetch the full list of map files from the memory map and hand
it over to gdb.
It will then fetch all the libraries from the remote host and give us
backtraces
2022-03-13 12:50:01 -07:00
Ryan Houdek b190150281 FEXCore: Merges redundant string trimming implementations 2022-03-13 12:50:01 -07:00
Ryan Houdek 53ffe5df43 Merge pull request #1613 from Sonicadvance1/gdbstub_improvements2
GDBServer improvements
2022-03-13 12:49:17 -07:00
Ryan Houdek d39df8d3ed Merge pull request #1614 from Sonicadvance1/add_comment
JIT: Adds comment to EmitDetectionString
2022-03-10 14:53:41 -08:00
Ryan Houdek eeb2b928b9 JIT: Adds comment to EmitDetectionString 2022-03-10 14:28:45 -08:00
Ryan Houdek 6cf24a748f GdbServer: Document what PassSignals is for 2022-03-10 14:23:21 -08:00
Ryan Houdek 3b4fd180de GDBServer: Pass auxv better
Fixes 32-bit auxv as well.
2022-03-10 14:21:46 -08:00
Ryan Houdek 7300c7a853 SignalDelegator: Remove anti-pattern usage 2022-03-10 14:00:51 -08:00
Ryan Houdek 376f6db3ac GDBServer: Reformat code to two space tabs
No functional change
2022-03-10 14:00:51 -08:00
Ryan Houdek ad3a960717 GDBServer: Support sending gdb the correct signal
Instead of just sending SIGSEGV, pass the real signal
2022-03-10 13:57:17 -08:00
Ryan Houdek 0a1ef867ee SignalDelegator: Support multiple backend host handlers
This will be necessary for gdbserver to handle signals indepedentally of
the FEX handling.
2022-03-10 13:57:17 -08:00
Ryan Houdek a7fe69deea CPU: Stop trying to initialize signal handlers per thread
These are static per process and only need to be initialized once.
We are going to support multiple signal handlers from the backend after
this, so can only install once.
2022-03-10 13:57:17 -08:00
Ryan Houdek a8c3b6d46f GDBServer: Capture signal capture numbers
This will allow us to wire this to a future signal handler for gdbserver
2022-03-10 13:57:17 -08:00
Ryan Houdek 60db2655bb GDBServer: Encode the return to pread correctly
This encodes the resulting data as raw binary rather than any special
escaped encoding
2022-03-10 13:57:10 -08:00
Ryan Houdek fd717b6995 GDBServer: Expose program offsets better
We were incorrectly returning programing offsets
Get the program offset from the frontend so we can know what to give gdb
2022-03-09 19:07:45 -08:00
Ryan Houdek ba37388fe3 GDBServer: Expose auxv values
We already expose these in the code loader, pump it through gdb
2022-03-09 19:07:45 -08:00
Ryan Houdek 68f32d85c9 CodeLoader: Expose base ELF loaded offset
Useful for gdb
2022-03-09 19:07:45 -08:00
Ryan Houdek 2aa77e85de LinuxSyscalls: Expose CodeLoader through syscall interface 2022-03-09 19:07:45 -08:00
Ryan Houdek 91665fdf0e Merge pull request #1610 from wannacu/main
FileManager: Fix realpath failed on debian buster
2022-03-09 18:58:52 -08:00
Ryan Houdek 23a1c64bf7 Merge pull request #1612 from Sonicadvance1/tag_memory_allocations
JITs: Emit identification string in the code buffers
2022-03-09 18:52:46 -08:00
Ryan Houdek c2dcf06632 JITs: Emit identification string in the code buffers
At the start of each code buffer, emit a small string for letting memory
inspection know if a code region is for the JITs.
2022-03-09 18:29:37 -08:00
Ryan Houdek fad91bb818 Merge pull request #1609 from Sonicadvance1/fix_map_32bit
Linux: Fixes MAP_32BIT supported range
2022-03-08 17:33:34 -08:00
wannacu 0e769ece26 docs: Update Readme_CN.md 2022-03-08 18:10:08 +08:00
wannacu 9a780b40a2 Docs: Add Chinese README 2022-03-08 18:04:09 +08:00
wannacu 898873e9e3 FileManager: Fix realpath failed on debian buster
This happend on debian buster when run realpath(i386) on arm64 host.
2022-03-08 16:42:05 +08:00
Ryan Houdek 52292e5f7e Linux: Fixes MAP_32BIT supported range
I accidentally committed a 32-bit range that was significantly smaller
than what it should be.
While the minimal range worked for simple cases, it didn't work for
anything complex.
Give it the full range it needs.

Fixes #1600
2022-03-06 17:46:17 -08:00
Ryan Houdek 5de6c866b7 Merge pull request #1608 from Sonicadvance1/termux_build_option
Adds a cmake option for forcing a termux build
2022-03-06 13:54:20 -08:00
Ryan Houdek ec0cd3aec4 Adds a cmake option for forcing a termux build
This is necessary when cross-compiling rather than building on-device
2022-03-06 12:56:25 -08:00
Mai M f5f9512d9a Merge pull request #1606 from Sonicadvance1/fhu_page_size
Change page define usages over to self-defined
2022-03-06 15:53:15 -05:00
Mai M fb27cb4356 Merge pull request #1607 from Sonicadvance1/disable_guis_termux
Disables GUI applications in a Termux build
2022-03-06 15:52:37 -05:00
Mai M 94664580c8 Merge pull request #1605 from Sonicadvance1/update_docs_termux
Update ReleaseProcess docs for Termux
2022-03-06 15:52:10 -05:00
Ryan Houdek 99a93fa9ea Disables GUI applications in a Termux build 2022-03-06 08:09:55 -08:00
Ryan Houdek 4cb6918506 Change page define usages over to self-defined
In the case of an AArch64 builder is using 16kb or 64kb pages like is
common on servers then it would fail to compile, even if the resulting
application would only ever run on 4k page hosts.

Resolve this by removing the build check and hardcoding 4kb pages for
each of our uses. We still require 4kb pages to run, so this mostly just
removes the weirdness where it is 16kb builder + 4k runner. Would have
broken some of our assumptions when running.
2022-03-06 07:33:10 -08:00
Ryan Houdek 9cc743bf84 Update ReleaseProcess docs for Termux
FEX hardly works on Termux as-is, but we should make sure to document
how to update the packages otherwise we will quickly become outdated on
their package management.
2022-03-06 05:59:33 -08:00
57 changed files with 1502 additions and 909 deletions

No files matched your search

+2 -53
View File
@@ -21,6 +21,7 @@ option(ENABLE_COMPILE_TIME_TRACE "Enables time trace compile option" FALSE)
option(ENABLE_LIBCXX "Enables LLVM libc++" FALSE)
option(ENABLE_INTERPRETER "Enables FEX's Interpreter" FALSE)
option(ENABLE_CCACHE "Enables ccache for compile caching" TRUE)
option(ENABLE_TERMUX_BUILD "Forces building for Termux on a non-Termux build machine" FALSE)
set (X86_C_COMPILER "x86_64-linux-gnu-gcc" CACHE STRING "c compiler for compiling x86 guest libs")
set (X86_CXX_COMPILER "x86_64-linux-gnu-g++" CACHE STRING "c++ compiler for compiling x86 guest libs")
@@ -115,59 +116,7 @@ if (NOT ENABLE_OFFLINE_TELEMETRY)
add_definitions(-DFEX_DISABLE_TELEMETRY=1)
endif()
# Check if the build target page size is 4096
include(CheckCSourceRuns)
check_c_source_runs(
"#include <unistd.h>
int main(int argc, char* argv[])
{
return getpagesize() == 4096 ? 0 : 1;
}"
PAGEFILE_RESULT
)
if (NOT ${PAGEFILE_RESULT})
message(FATAL_ERROR "Host PAGE_SIZE is not 4096. Can't build on this target")
endif()
include(CheckCXXSourceCompiles)
check_cxx_source_compiles(
"#include <sys/user.h>
int main() {
return PAGE_SIZE;
}
"
HAS_PAGESIZE)
check_cxx_source_compiles(
"#include <sys/user.h>
int main() {
return PAGE_SHIFT;
}
"
HAS_PAGESHIFT)
check_cxx_source_compiles(
"#include <sys/user.h>
int main() {
return PAGE_MASK;
}
"
HAS_PAGEMASK)
if (NOT HAS_PAGESIZE)
add_definitions(-DPAGE_SIZE=4096)
endif()
if (NOT HAS_PAGESHIFT)
add_definitions(-DPAGE_SHIFT=12)
endif()
if (NOT HAS_PAGEMASK)
add_definitions("-DPAGE_MASK=(~(PAGE_SIZE-1))")
endif()
if(DEFINED ENV{TERMUX_VERSION})
if(DEFINED ENV{TERMUX_VERSION} OR ENABLE_TERMUX_BUILD)
add_definitions(-DTERMUX_BUILD=1)
set(TERMUX_BUILD 1)
# Termux doesn't support Jemalloc due to bad interactions between emutls, jemalloc, and scudo
+2 -12
View File
@@ -18,14 +18,8 @@ This project aims to provide a fast and functional x86-64 emulation library that
* Portable library implementation in order to support easy integration in to applications
### Target Host Architecture
The target host architecture for this library is AArch64. Specifically the ARMv8.1 version or newer.
The CPU IR is designed with AArch64 in mind but there is a desire to run the recompiled code on other architectures as well.
Multiple architecture support is desired for easier bringup and debugging, performance isn't as much of a priority there (ex. x86-64(guest) translated to x86-64(host))
### Not currently goals but will be in the future
* 32bit x86 support
* This will be a desire in the future, but to lower the amount of work required, decided to push this off for now.
* Integration in to WINE
* Later generation of x86-64 instruction sets
* Including AVX, F16C, XOP, FMA, AVX2, etc
The CPU IR is designed with AArch64 in mind but should allow for other architectures as well.
x86-64 host support is available for ease of development, but is not a priority.
### Not desired
* Kernel space emulation
* CPL0-2 emulation
@@ -33,7 +27,3 @@ Multiple architecture support is desired for easier bringup and debugging, perfo
* IRQs
* SVM
* "Cycle Accurate" emulation
### Dependencies
* clang-tidy if you want to ensure the code stays tidy
* cmake
* A C++17 compliant compiler (There are assumptions made about using Clang and LTO)
+5 -1
View File
@@ -188,6 +188,10 @@ struct X80SoftFloat {
return extF80_roundToInt(lhs, softfloat_roundingMode, false);
}
static X80SoftFloat FRNDINT(X80SoftFloat const &lhs, uint_fast8_t RoundMode) {
return extF80_roundToInt(lhs, RoundMode, false);
}
static X80SoftFloat FXTRACT_SIG(X80SoftFloat const &lhs) {
#if defined(DEBUG_X86_FLOAT)
BIGFLOAT Result;
@@ -257,7 +261,7 @@ struct X80SoftFloat {
return Result;
#else
X80SoftFloat Int = FRNDINT(rhs);
X80SoftFloat Int = FRNDINT(rhs, softfloat_round_minMag);
BIGFLOAT Src2_d = Int;
Src2_d = exp2l(Src2_d);
X80SoftFloat Src2_X80 = Src2_d;
+29
View File
@@ -0,0 +1,29 @@
#pragma once
#include <string>
namespace FEXCore::StringUtils {
// Trim the left side of the string of whitespace and new lines
[[maybe_unused]] static std::string LeftTrim(std::string String, std::string TrimTokens = " \t\n\r") {
size_t pos = std::string::npos;
if ((pos = String.find_first_not_of(TrimTokens)) != std::string::npos) {
String.erase(0, pos);
}
return String;
}
// Trim the right side of the string of whitespace and new lines
[[maybe_unused]] static std::string RightTrim(std::string String, std::string TrimTokens = " \t\n\r") {
size_t pos = std::string::npos;
if ((pos = String.find_last_not_of(TrimTokens)) != std::string::npos) {
String.erase(String.begin() + pos + 1, String.end());
}
return String;
}
// Trim both the left and right of the string of whitespace and new lines
[[maybe_unused]] static std::string Trim(std::string String, std::string TrimTokens = " \t\n\r") {
return RightTrim(LeftTrim(String, TrimTokens), TrimTokens);
}
}
+2 -24
View File
@@ -1,4 +1,5 @@
#include "Common/StringConv.h"
#include "Common/StringUtils.h"
#include "Common/Paths.h"
#include "Utils/FileLoading.h"
@@ -370,29 +371,6 @@ namespace JSON {
return {};
}
std::string ltrim(std::string String) {
size_t pos = std::string::npos;
if ((pos = String.find_first_not_of(" \t\n\r")) != std::string::npos) {
String.erase(0, pos);
}
return String;
}
std::string rtrim(std::string String) {
size_t pos = std::string::npos;
if ((pos = String.find_last_not_of(" \t\n\r")) != std::string::npos) {
String.erase(String.begin() + pos + 1, String.end());
}
return String;
}
std::string trim(std::string String) {
return rtrim(ltrim(String));
}
std::string FindContainerPrefix() {
// We only support pressure-vessel at the moment
const static std::string ContainerManager = "/run/host/container-manager";
@@ -401,7 +379,7 @@ namespace JSON {
if (FEXCore::FileLoading::LoadFile(Manager, ContainerManager)) {
// Trim the whitespace, may contain a newline
std::string ManagerStr = Manager.data();
ManagerStr = trim(ManagerStr);
ManagerStr = FEXCore::StringUtils::Trim(ManagerStr);
if (strncmp(ManagerStr.data(), "pressure-vessel", Manager.size()) == 0) {
// We are running inside of pressure vessel
// Our $CMAKE_INSTALL_PREFIX paths are now inside of /run/host/$CMAKE_INSTALL_PREFIX
+21 -1
View File
@@ -826,7 +826,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_4000_0000h(uint32_t Leaf) {
// CPUID documentation information:
// 4000_0000h - 4FFF_FFFFh - No existing or future CPU will return information in this range
// Reserved entirely for VMs to do whatever they want.
Res.eax = 0x40000000;
Res.eax = 0x40000001;
// EBX, EDX, ECX become the hypervisor ID signature
constexpr static char HypervisorID[12] = "FEXIFEXIEMU";
@@ -834,6 +834,25 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_4000_0000h(uint32_t Leaf) {
return Res;
}
// Hypervisor CPUID information leaf
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_4000_0001h(uint32_t Leaf) {
FEXCore::CPUID::FunctionResults Res{};
if (Leaf == 0) {
// EAX[3:0] Is the host architecture that FEX is running under
#ifdef _M_X86_64
// EAX[3:0] = 1 = x86_64 host architecture
Res.eax |= 0b0001;
#elif defined(_M_ARM_64)
// EAX[3:0] = 2 = AArch64 host architecture
Res.eax |= 0b0010;
#else
// EAX[3:0] = 0 = Unknown architecture
#endif
}
return Res;
}
// Highest extended function implemented
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0000h(uint32_t Leaf) {
FEXCore::CPUID::FunctionResults Res{};
@@ -1228,6 +1247,7 @@ void CPUIDEmu::Init(FEXCore::Context::Context *ctx) {
#endif
// Hypervisor CPUID information leaf
RegisterFunction(0x4000'0000, &CPUIDEmu::Function_4000_0000h);
RegisterFunction(0x4000'0001, &CPUIDEmu::Function_4000_0001h);
// Largest extended function number
RegisterFunction(0x8000'0000, &CPUIDEmu::Function_8000_0000h);
+1
View File
@@ -80,6 +80,7 @@ private:
FEXCore::CPUID::FunctionResults Function_15h(uint32_t Leaf);
FEXCore::CPUID::FunctionResults Function_1Ah(uint32_t Leaf);
FEXCore::CPUID::FunctionResults Function_4000_0000h(uint32_t Leaf);
FEXCore::CPUID::FunctionResults Function_4000_0001h(uint32_t Leaf);
FEXCore::CPUID::FunctionResults Function_8000_0000h(uint32_t Leaf);
FEXCore::CPUID::FunctionResults Function_8000_0001h(uint32_t Leaf);
FEXCore::CPUID::FunctionResults Function_8000_0002h(uint32_t Leaf);
+32 -4
View File
@@ -190,6 +190,32 @@ namespace FEXCore::Context {
}
FEXCore::Core::InternalThreadState* Context::InitCore(FEXCore::CodeLoader *Loader) {
// Initialize the CPU core signal handlers
switch (Config.Core) {
#ifdef INTERPRETER_ENABLED
case FEXCore::Config::CONFIG_INTERPRETER:
FEXCore::CPU::InitializeInterpreterSignalHandlers(this);
break;
#endif
case FEXCore::Config::CONFIG_IRJIT:
#if (_M_X86_64 && JIT_X86_64)
FEXCore::CPU::InitializeX86JITSignalHandlers(this);
#elif (_M_ARM_64 && JIT_ARM64)
FEXCore::CPU::InitializeArm64JITSignalHandlers(this);
#else
ERROR_AND_DIE_FMT("FEXCore has been compiled without a viable JIT core");
#endif
break;
case FEXCore::Config::CONFIG_CUSTOM:
// Do nothing
break;
default:
ERROR_AND_DIE_FMT("Unknown core configuration");
break;
}
// Initialize GDBServer after the signal handlers are installed
// It may install its own handlers that need to be executed AFTER the CPU cores
if (Config.GdbServer) {
StartGdbServer();
}
@@ -871,10 +897,6 @@ namespace FEXCore::Context {
StartAddr = _StartAddr;
Length = _Length;
// Initialize metadata
DebugData->GuestCodeSize = TotalInstructionsLength;
DebugData->GuestInstructionCount = TotalInstructions;
// Increment stats
Thread->Stats.BlocksCompiled.fetch_add(1);
@@ -1122,10 +1144,16 @@ namespace FEXCore::Context {
void Context::AddNamedRegion(uintptr_t Base, uintptr_t Size, uintptr_t Offset, const std::string &filename) {
IRCaptureCache.AddNamedRegion(Base, Size, Offset, filename);
if (DebugServer) {
DebugServer->AlertLibrariesChanged();
}
}
void Context::RemoveNamedRegion(uintptr_t Base, uintptr_t Size) {
IRCaptureCache.RemoveNamedRegion(Base, Size);
if (DebugServer) {
DebugServer->AlertLibrariesChanged();
}
}
void ConfigureAOTGen(FEXCore::Core::InternalThreadState *Thread, std::set<uint64_t> *ExternalBranches, uint64_t SectionMaxAddress) {
@@ -436,6 +436,88 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
b(&LoopTop);
}
// Long division helpers
uint64_t LUDIVHandler{};
uint64_t LDIVHandler{};
uint64_t LUREMHandler{};
uint64_t LREMHandler{};
{
LUDIVHandler = GetCursorAddress<uint64_t>();
PushDynamicRegsAndLR();
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.LUDIV)));
SpillStaticRegs();
blr(x3);
FillStaticRegs();
// Result is now in x0
// Fix the stack and any values that were stepped on
PopDynamicRegsAndLR();
// Go back to our code block
ret();
}
{
LDIVHandler = GetCursorAddress<uint64_t>();
PushDynamicRegsAndLR();
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.LDIV)));
SpillStaticRegs();
blr(x3);
FillStaticRegs();
// Result is now in x0
// Fix the stack and any values that were stepped on
PopDynamicRegsAndLR();
// Go back to our code block
ret();
}
{
LUREMHandler = GetCursorAddress<uint64_t>();
PushDynamicRegsAndLR();
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.LUREM)));
SpillStaticRegs();
blr(x3);
FillStaticRegs();
// Result is now in x0
// Fix the stack and any values that were stepped on
PopDynamicRegsAndLR();
// Go back to our code block
ret();
}
{
LREMHandler = GetCursorAddress<uint64_t>();
PushDynamicRegsAndLR();
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.LREM)));
SpillStaticRegs();
blr(x3);
FillStaticRegs();
// Result is now in x0
// Fix the stack and any values that were stepped on
PopDynamicRegsAndLR();
// Go back to our code block
ret();
}
place(&l_PagePtr);
place(&l_CTX);
place(&l_Sleep);
@@ -470,6 +552,10 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
Pointers.OverflowExceptionHandler = OverflowExceptionInstructionAddress;
Pointers.SignalReturnHandler = SignalHandlerReturnAddress;
Pointers.L1Pointer = Thread->LookupCache->GetL1Pointer();
Pointers.LUDIVHandler = LUDIVHandler;
Pointers.LDIVHandler = LDIVHandler;
Pointers.LUREMHandler = LUREMHandler;
Pointers.LREMHandler = LREMHandler;
}
}
File diff suppressed because it is too large. Load diff
+14 -1
View File
@@ -6,8 +6,10 @@ $end_info$
#pragma once
#include <FEXCore/Config/Config.h>
#include <FEXCore/Utils/Event.h>
#include <FEXCore/Utils/Threads.h>
#include <atomic>
#include <istream>
#include <memory>
#include <mutex>
@@ -27,6 +29,10 @@ public:
// Public for threading
void GdbServerLoop();
void AlertLibrariesChanged() {
LibraryMapChanged = true;
}
private:
void Break(int signal);
@@ -38,6 +44,9 @@ private:
void SendACK(std::ostream &stream, bool NACK);
Event ThreadBreakEvent{};
void WaitForThreadWakeup();
struct HandledPacketType {
std::string Response{};
enum ResponseType {
@@ -74,9 +83,13 @@ private:
bool NoAckMode{false};
bool NonStopMode{false};
std::string ThreadString{};
std::string MemoryMapString{};
std::string OSDataString{};
void buildLibraryMap();
std::atomic<bool> LibraryMapChanged = true;
std::string LibraryMapString{};
// Used to keep track of which signals to pass to the guest
std::array<bool, SignalDelegator::MAX_SIGNALS + 1> PassSignals{};
uint32_t CurrentDebuggingThread{};
int ListenSocket{};
FEX_CONFIG_OPT(Filename, APP_FILENAME);
+10
View File
@@ -84,6 +84,16 @@ HostFeatures::HostFeatures() {
SupportsCRC = Features.has(Xbyak::util::Cpu::tSSE42);
SupportsRAND = Features.has(Xbyak::util::Cpu::tRDRAND) && Features.has(Xbyak::util::Cpu::tRDSEED);
// xbyak doesn't know how to check for CLZero
uint32_t eax, ebx, ecx, edx;
// First ensure we support a new enough extended CPUID function range
__cpuid(0x8000'0000, eax, ebx, ecx, edx);
if (eax >= 0x8000'0008U) {
// CLZero defined in 8000_00008_EBX[bit 0]
__cpuid(0x8000'0008, eax, ebx, ecx, edx);
SupportsCLZERO = ebx & 1;
}
SupportsFlushInputsToZero = true;
SupportsFloatExceptions = true;
#else
@@ -37,6 +37,10 @@ public:
void CreateAsmDispatch(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread);
static void InitializeSignalHandlers(FEXCore::Context::Context *CTX);
bool NeedsRetainedIRCopy() const override { return true; }
private:
FEXCore::Context::Context *CTX;
FEXCore::Core::InternalThreadState *State;
@@ -41,25 +41,28 @@ InterpreterCore::InterpreterCore(FEXCore::Context::Context *ctx, FEXCore::Core::
if (!CompileThread &&
CTX->Config.Core == FEXCore::Config::CONFIG_INTERPRETER) {
CreateAsmDispatch(ctx, Thread);
CTX->SignalDelegation->RegisterHostSignalHandler(SignalDelegator::SIGNAL_FOR_PAUSE, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
InterpreterCore *Core = reinterpret_cast<InterpreterCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleSignalPause(Signal, info, ucontext);
}, true);
}
}
void InterpreterCore::InitializeSignalHandlers(FEXCore::Context::Context *CTX) {
CTX->SignalDelegation->RegisterHostSignalHandler(SignalDelegator::SIGNAL_FOR_PAUSE, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
InterpreterCore *Core = reinterpret_cast<InterpreterCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleSignalPause(Signal, info, ucontext);
}, true);
#ifdef _M_ARM_64
CTX->SignalDelegation->RegisterHostSignalHandler(SIGBUS, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
return FEXCore::ArchHelpers::Arm64::HandleSIGBUS(true, Signal, info, ucontext);
}, true);
CTX->SignalDelegation->RegisterHostSignalHandler(SIGBUS, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
return FEXCore::ArchHelpers::Arm64::HandleSIGBUS(true, Signal, info, ucontext);
}, true);
#endif
auto GuestSignalHandler = [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack) -> bool {
InterpreterCore *Core = reinterpret_cast<InterpreterCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleGuestSignal(Signal, info, ucontext, GuestAction, GuestStack);
};
auto GuestSignalHandler = [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack) -> bool {
InterpreterCore *Core = reinterpret_cast<InterpreterCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleGuestSignal(Signal, info, ucontext, GuestAction, GuestStack);
};
for (uint32_t Signal = 0; Signal <= SignalDelegator::MAX_SIGNALS; ++Signal) {
CTX->SignalDelegation->RegisterHostSignalHandlerForGuest(Signal, GuestSignalHandler);
}
for (uint32_t Signal = 0; Signal <= SignalDelegator::MAX_SIGNALS; ++Signal) {
CTX->SignalDelegation->RegisterHostSignalHandlerForGuest(Signal, GuestSignalHandler);
}
}
@@ -71,4 +74,8 @@ std::unique_ptr<CPUBackend> CreateInterpreterCore(FEXCore::Context::Context *ctx
return std::make_unique<InterpreterCore>(ctx, Thread, CompileThread);
}
void InitializeInterpreterSignalHandlers(FEXCore::Context::Context *CTX) {
InterpreterCore::InitializeSignalHandlers(CTX);
}
}
@@ -17,4 +17,6 @@ class CPUBackend;
FEXCore::Core::InternalThreadState *Thread,
bool CompileThread);
void InitializeInterpreterSignalHandlers(FEXCore::Context::Context *CTX);
} // namespace FEXCore::CPU
@@ -324,12 +324,6 @@ void InterpreterOps::Op_NoOp(FEXCore::IR::IROp_Header *IROp, IROpData *Data, IR:
void InterpreterOps::InterpretIR(FEXCore::Core::InternalThreadState *Thread, uint64_t Entry, FEXCore::IR::IRListView *CurrentIR, FEXCore::Core::DebugData *DebugData) {
volatile void *StackEntry = alloca(0);
// Debug data is only passed in debug builds
#ifndef NDEBUG
// TODO: should be moved to an IR Op
Thread->Stats.InstructionsExecuted.fetch_add(DebugData->GuestInstructionCount);
#endif
uintptr_t ListSize = CurrentIR->GetSSACount();
static_assert(sizeof(FEXCore::IR::IROp_Header) == 4);
+119 -52
View File
@@ -636,24 +636,42 @@ DEF_OP(LDiv) {
break;
}
case 8: {
PushDynamicRegsAndLR();
auto Upper64Bit = GetReg<RA_64>(Op->Header.Args[1].ID());
auto Lower64Bit = GetReg<RA_64>(Op->Header.Args[0].ID());
auto Divisor = GetReg<RA_64>(Op->Header.Args[2].ID());
Label Only64Bit{};
Label LongDIVRet{};
mov(x0, GetReg<RA_64>(Op->Header.Args[1].ID()));
mov(x1, GetReg<RA_64>(Op->Header.Args[0].ID()));
mov(x2, GetReg<RA_64>(Op->Header.Args[2].ID()));
// Check if the upper bits match the top bit of the lower 64-bits
// Sign extend the top bit of lower bits
sbfx(TMP1, Lower64Bit, 63, 1);
eor(TMP1, TMP1, Upper64Bit);
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.LDIV)));
// If the sign bit matches then the result is zero
cbz(TMP1, &Only64Bit);
SpillStaticRegs();
blr(x3);
FillStaticRegs();
// Long divide
{
mov(x0, Upper64Bit);
mov(x1, Lower64Bit);
mov(x2, Divisor);
// Result is now in x0
// Fix the stack and any values that were stepped on
PopDynamicRegsAndLR();
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.LDIVHandler)));
blr(x3);
// Move result to its destination register
mov(GetReg<RA_64>(Node), x0);
// Move result to its destination register
mov(GetReg<RA_64>(Node), x0);
// Skip 64-bit path
b(&LongDIVRet);
}
bind(&Only64Bit);
// 64-Bit only
{
sdiv(GetReg<RA_64>(Node), Lower64Bit, Divisor);
}
bind(&LongDIVRet);
break;
}
default: LOGMAN_MSG_A_FMT("Unknown LDIV Size: {}", Size); break;
@@ -680,23 +698,38 @@ DEF_OP(LUDiv) {
break;
}
case 8: {
PushDynamicRegsAndLR();
auto Upper64Bit = GetReg<RA_64>(Op->Header.Args[1].ID());
auto Lower64Bit = GetReg<RA_64>(Op->Header.Args[0].ID());
auto Divisor = GetReg<RA_64>(Op->Header.Args[2].ID());
Label Only64Bit{};
Label LongDIVRet{};
mov(x0, GetReg<RA_64>(Op->Header.Args[1].ID()));
mov(x1, GetReg<RA_64>(Op->Header.Args[0].ID()));
mov(x2, GetReg<RA_64>(Op->Header.Args[2].ID()));
// Check the upper bits for zero
// If the upper bits are zero then we can do a 64-bit divide
cbz(Upper64Bit, &Only64Bit);
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.LUDIV)));
SpillStaticRegs();
blr(x3);
FillStaticRegs();
// Long divide
{
mov(x0, Upper64Bit);
mov(x1, Lower64Bit);
mov(x2, Divisor);
// Result is now in x0
// Fix the stack and any values that were stepped on
PopDynamicRegsAndLR();
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.LUDIVHandler)));
blr(x3);
// Move result to its destination register
mov(GetReg<RA_64>(Node), x0);
// Move result to its destination register
mov(GetReg<RA_64>(Node), x0);
// Skip 64-bit path
b(&LongDIVRet);
}
bind(&Only64Bit);
// 64-Bit only
{
udiv(GetReg<RA_64>(Node), Lower64Bit, Divisor);
}
bind(&LongDIVRet);
break;
}
default: LOGMAN_MSG_A_FMT("Unknown LUDIV Size: {}", Size); break;
@@ -733,23 +766,42 @@ DEF_OP(LRem) {
break;
}
case 8: {
PushDynamicRegsAndLR();
auto Upper64Bit = GetReg<RA_64>(Op->Header.Args[1].ID());
auto Lower64Bit = GetReg<RA_64>(Op->Header.Args[0].ID());
auto Divisor = GetReg<RA_64>(Op->Header.Args[2].ID());
Label Only64Bit{};
Label LongDIVRet{};
mov(x0, GetReg<RA_64>(Op->Header.Args[1].ID()));
mov(x1, GetReg<RA_64>(Op->Header.Args[0].ID()));
mov(x2, GetReg<RA_64>(Op->Header.Args[2].ID()));
// Check if the upper bits match the top bit of the lower 64-bits
// Sign extend the top bit of lower bits
sbfx(TMP1, Lower64Bit, 63, 1);
eor(TMP1, TMP1, Upper64Bit);
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.LREM)));
SpillStaticRegs();
blr(x3);
FillStaticRegs();
// If the sign bit matches then the result is zero
cbz(TMP1, &Only64Bit);
// Result is now in x0
// Fix the stack and any values that were stepped on
PopDynamicRegsAndLR();
// Long divide
{
mov(x0, Upper64Bit);
mov(x1, Lower64Bit);
mov(x2, Divisor);
// Move result to its destination register
mov(GetReg<RA_64>(Node), x0);
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.LREMHandler)));
blr(x3);
// Move result to its destination register
mov(GetReg<RA_64>(Node), x0);
// Skip 64-bit path
b(&LongDIVRet);
}
bind(&Only64Bit);
// 64-Bit only
{
sdiv(TMP1, Lower64Bit, Divisor);
msub(GetReg<RA_64>(Node), TMP1, Divisor, Lower64Bit);
}
bind(&LongDIVRet);
break;
}
default: LOGMAN_MSG_A_FMT("Unknown LREM Size: {}", Size); break;
@@ -782,24 +834,39 @@ DEF_OP(LURem) {
break;
}
case 8: {
auto Upper64Bit = GetReg<RA_64>(Op->Header.Args[1].ID());
auto Lower64Bit = GetReg<RA_64>(Op->Header.Args[0].ID());
auto Divisor = GetReg<RA_64>(Op->Header.Args[2].ID());
Label Only64Bit{};
Label LongDIVRet{};
PushDynamicRegsAndLR();
// Check the upper bits for zero
// If the upper bits are zero then we can do a 64-bit divide
cbz(Upper64Bit, &Only64Bit);
mov(x0, GetReg<RA_64>(Op->Header.Args[1].ID()));
mov(x1, GetReg<RA_64>(Op->Header.Args[0].ID()));
mov(x2, GetReg<RA_64>(Op->Header.Args[2].ID()));
// Long divide
{
mov(x0, Upper64Bit);
mov(x1, Lower64Bit);
mov(x2, Divisor);
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.LUREM)));
SpillStaticRegs();
blr(x3);
FillStaticRegs();
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.LUREMHandler)));
blr(x3);
// Move result to its destination register
mov(GetReg<RA_64>(Node), x0);
// Fix the stack and any values that were stepped on
PopDynamicRegsAndLR();
// Skip 64-bit path
b(&LongDIVRet);
}
// Result is now in x0
// Move result to its destination register
mov(GetReg<RA_64>(Node), x0);
bind(&Only64Bit);
// 64-Bit only
{
udiv(TMP1, Lower64Bit, Divisor);
msub(GetReg<RA_64>(Node), TMP1, Divisor, Lower64Bit);
}
bind(&LongDIVRet);
break;
}
default: LOGMAN_MSG_A_FMT("Unknown LUREM Size: {}", OpSize); break;
+94 -83
View File
@@ -361,53 +361,6 @@ Arm64JITCore::Arm64JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::Intern
: Arm64Emitter(ctx, 0)
, CTX {ctx}
, ThreadState {Thread} {
{
// Set up pointers that the JIT needs to load
auto &Pointers = ThreadState->CurrentFrame->Pointers.AArch64;
// Process specific
Pointers.LUDIV = reinterpret_cast<uint64_t>(LUDIV);
Pointers.LDIV = reinterpret_cast<uint64_t>(LDIV);
Pointers.LUREM = reinterpret_cast<uint64_t>(LUREM);
Pointers.LREM = reinterpret_cast<uint64_t>(LREM);
Pointers.PrintValue = reinterpret_cast<uint64_t>(PrintValue);
Pointers.PrintVectorValue = reinterpret_cast<uint64_t>(PrintVectorValue);
Pointers.RemoveCodeEntryFromJIT = reinterpret_cast<uintptr_t>(&Context::Context::RemoveCodeEntryFromJit);
Pointers.CPUIDObj = reinterpret_cast<uint64_t>(&CTX->CPUID);
{
FEXCore::Utils::MemberFunctionToPointerCast PMF(&FEXCore::CPUIDEmu::RunFunction);
Pointers.CPUIDFunction = PMF.GetConvertedPointer();
}
Pointers.SyscallHandlerObj = reinterpret_cast<uint64_t>(CTX->SyscallHandler);
Pointers.SyscallHandlerFunc = reinterpret_cast<uint64_t>(FEXCore::Context::HandleSyscall);
// Fill in the fallback handlers
InterpreterOps::FillFallbackIndexPointers(Pointers.FallbackHandlerPointers);
// Thread Specific
Pointers.SignalHandlerRefCountPointer = reinterpret_cast<uint64_t>(&Dispatcher->SignalHandlerRefCounter);
}
{
DispatcherConfig config;
config.ExitFunctionLink = reinterpret_cast<uintptr_t>(&ExitFunctionLink);
config.ExitFunctionLinkThis = reinterpret_cast<uintptr_t>(this);
config.StaticRegisterAssignment = ctx->Config.StaticRegisterAllocation;
Dispatcher = std::make_unique<Arm64Dispatcher>(CTX, ThreadState, config);
DispatchPtr = Dispatcher->DispatchPtr;
CallbackPtr = Dispatcher->CallbackPtr;
}
// Can't allocate a code buffer until after dispatcher is created
InitialCodeBuffer = AllocateNewCodeBuffer(Arm64JITCore::INITIAL_CODE_SIZE);
*GetBuffer() = vixl::CodeBuffer(InitialCodeBuffer.Ptr, InitialCodeBuffer.Size);
SetAllowAssembler(true);
CurrentCodeBuffer = &InitialCodeBuffer;
RAPass = Thread->PassManager->GetPass<IR::RegisterAllocationPass>("RA");
#if DEBUG
@@ -447,44 +400,100 @@ Arm64JITCore::Arm64JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::Intern
RegisterVectorHandlers();
RegisterEncryptionHandlers();
{
DispatcherConfig config;
config.ExitFunctionLink = reinterpret_cast<uintptr_t>(&ExitFunctionLink);
config.ExitFunctionLinkThis = reinterpret_cast<uintptr_t>(this);
config.StaticRegisterAssignment = ctx->Config.StaticRegisterAllocation;
Dispatcher = std::make_unique<Arm64Dispatcher>(CTX, ThreadState, config);
DispatchPtr = Dispatcher->DispatchPtr;
CallbackPtr = Dispatcher->CallbackPtr;
}
if (!CompileThread) {
ThreadSharedData.SignalHandlerRefCounterPtr = &Dispatcher->SignalHandlerRefCounter;
ThreadSharedData.SignalReturnInstruction = Dispatcher->SignalHandlerReturnAddress;
ThreadSharedData.UnimplementedInstructionAddress = Dispatcher->UnimplementedInstructionAddress;
ThreadSharedData.OverflowExceptionInstructionAddress = Dispatcher->OverflowExceptionInstructionAddress;
ThreadSharedData.Dispatcher = Dispatcher.get();
// This will register the host signal handler per thread, which is fine
CTX->SignalDelegation->RegisterHostSignalHandler(SIGILL, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
Arm64JITCore *Core = reinterpret_cast<Arm64JITCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleSIGILL(Signal, info, ucontext);
}, true);
CTX->SignalDelegation->RegisterHostSignalHandler(SIGBUS, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
Arm64JITCore *Core = reinterpret_cast<Arm64JITCore*>(Thread->CPUBackend.get());
if (!Core->Dispatcher->IsAddressInJITCode(ArchHelpers::Context::GetPc(ucontext))) {
// Wasn't a sigbus in JIT code
return false;
}
return FEXCore::ArchHelpers::Arm64::HandleSIGBUS(Core->CTX->Config.ParanoidTSO(), Signal, info, ucontext);
}, true);
CTX->SignalDelegation->RegisterHostSignalHandler(SignalDelegator::SIGNAL_FOR_PAUSE, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
Arm64JITCore *Core = reinterpret_cast<Arm64JITCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleSignalPause(Signal, info, ucontext);
}, true);
auto GuestSignalHandler = [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack) -> bool {
Arm64JITCore *Core = reinterpret_cast<Arm64JITCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleGuestSignal(Signal, info, ucontext, GuestAction, GuestStack);
};
for (uint32_t Signal = 0; Signal <= SignalDelegator::MAX_SIGNALS; ++Signal) {
CTX->SignalDelegation->RegisterHostSignalHandlerForGuest(Signal, GuestSignalHandler);
}
}
{
// Set up pointers that the JIT needs to load
auto &Pointers = ThreadState->CurrentFrame->Pointers.AArch64;
// Process specific
Pointers.LUDIV = reinterpret_cast<uint64_t>(LUDIV);
Pointers.LDIV = reinterpret_cast<uint64_t>(LDIV);
Pointers.LUREM = reinterpret_cast<uint64_t>(LUREM);
Pointers.LREM = reinterpret_cast<uint64_t>(LREM);
Pointers.PrintValue = reinterpret_cast<uint64_t>(PrintValue);
Pointers.PrintVectorValue = reinterpret_cast<uint64_t>(PrintVectorValue);
Pointers.RemoveCodeEntryFromJIT = reinterpret_cast<uintptr_t>(&Context::Context::RemoveCodeEntryFromJit);
Pointers.CPUIDObj = reinterpret_cast<uint64_t>(&CTX->CPUID);
{
FEXCore::Utils::MemberFunctionToPointerCast PMF(&FEXCore::CPUIDEmu::RunFunction);
Pointers.CPUIDFunction = PMF.GetConvertedPointer();
}
Pointers.SyscallHandlerObj = reinterpret_cast<uint64_t>(CTX->SyscallHandler);
Pointers.SyscallHandlerFunc = reinterpret_cast<uint64_t>(FEXCore::Context::HandleSyscall);
// Fill in the fallback handlers
InterpreterOps::FillFallbackIndexPointers(Pointers.FallbackHandlerPointers);
// Thread Specific
Pointers.SignalHandlerRefCountPointer = reinterpret_cast<uint64_t>(&Dispatcher->SignalHandlerRefCounter);
}
// Can't allocate a code buffer until after dispatcher is created
InitialCodeBuffer = AllocateNewCodeBuffer(Arm64JITCore::INITIAL_CODE_SIZE);
*GetBuffer() = vixl::CodeBuffer(InitialCodeBuffer.Ptr, InitialCodeBuffer.Size);
SetAllowAssembler(true);
EmitDetectionString();
CurrentCodeBuffer = &InitialCodeBuffer;
}
void Arm64JITCore::InitializeSignalHandlers(FEXCore::Context::Context *CTX) {
CTX->SignalDelegation->RegisterHostSignalHandler(SIGILL, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
Arm64JITCore *Core = reinterpret_cast<Arm64JITCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleSIGILL(Signal, info, ucontext);
}, true);
CTX->SignalDelegation->RegisterHostSignalHandler(SIGBUS, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
Arm64JITCore *Core = reinterpret_cast<Arm64JITCore*>(Thread->CPUBackend.get());
if (!Core->Dispatcher->IsAddressInJITCode(ArchHelpers::Context::GetPc(ucontext))) {
// Wasn't a sigbus in JIT code
return false;
}
return FEXCore::ArchHelpers::Arm64::HandleSIGBUS(Core->CTX->Config.ParanoidTSO(), Signal, info, ucontext);
}, true);
CTX->SignalDelegation->RegisterHostSignalHandler(SignalDelegator::SIGNAL_FOR_PAUSE, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
Arm64JITCore *Core = reinterpret_cast<Arm64JITCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleSignalPause(Signal, info, ucontext);
}, true);
auto GuestSignalHandler = [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack) -> bool {
Arm64JITCore *Core = reinterpret_cast<Arm64JITCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleGuestSignal(Signal, info, ucontext, GuestAction, GuestStack);
};
for (uint32_t Signal = 0; Signal <= SignalDelegator::MAX_SIGNALS; ++Signal) {
CTX->SignalDelegation->RegisterHostSignalHandlerForGuest(Signal, GuestSignalHandler);
}
}
void Arm64JITCore::EmitDetectionString() {
const char JITString[] = "FEXJIT::Arm64JITCore::";
auto Buffer = GetBuffer();
Buffer->EmitString(JITString);
Buffer->Align();
}
void Arm64JITCore::ClearCache() {
@@ -527,6 +536,7 @@ void Arm64JITCore::ClearCache() {
EmplaceNewCodeBuffer(NewCodeBuffer);
*Buffer = vixl::CodeBuffer(NewCodeBuffer.Ptr, NewCodeBuffer.Size);
}
EmitDetectionString();
}
Arm64JITCore::~Arm64JITCore() {
@@ -755,6 +765,7 @@ void *Arm64JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IR
LOGMAN_THROW_A_FMT(BlockIROp->Header.Op == IR::OP_CODEBLOCK, "IR type failed to be a code block");
#endif
uintptr_t BlockStartHostCode = GetCursorAddress<uintptr_t>();
{
const auto Node = IR->GetID(BlockNode);
const auto IsTarget = JumpTargets.try_emplace(Node).first;
@@ -769,10 +780,6 @@ void *Arm64JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IR
bind(&IsTarget->second);
}
if (DebugData) {
DebugData->Subblocks.push_back({GetCursorAddress<uintptr_t>(), 0, IR->GetID(BlockNode)});
}
for (auto [CodeNode, IROp] : IR->GetCode(BlockNode)) {
const auto ID = IR->GetID(CodeNode);
@@ -782,7 +789,7 @@ void *Arm64JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IR
}
if (DebugData) {
DebugData->Subblocks.back().HostCodeSize = GetCursorAddress<uintptr_t>() - DebugData->Subblocks.back().HostCodeStart;
DebugData->Subblocks.push_back({BlockStartHostCode, static_cast<uint32_t>(GetCursorAddress<uintptr_t>() - BlockStartHostCode)});
}
}
@@ -859,4 +866,8 @@ uint64_t Arm64JITCore::ExitFunctionLink(Arm64JITCore *core, FEXCore::Core::CpuSt
std::unique_ptr<CPUBackend> CreateArm64JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, bool CompileThread) {
return std::make_unique<Arm64JITCore>(ctx, Thread, CompileThread);
}
void InitializeArm64JITSignalHandlers(FEXCore::Context::Context *CTX) {
Arm64JITCore::InitializeSignalHandlers(CTX);
}
}
@@ -68,6 +68,8 @@ public:
return Dispatcher->IsAddressInJITCode(Address, IncludeDispatcher, IncludeCompileService);
}
static void InitializeSignalHandlers(FEXCore::Context::Context *CTX);
private:
FEX_CONFIG_OPT(ParanoidTSO, PARANOIDTSO);
@@ -183,6 +185,9 @@ private:
};
CompilerSharedData ThreadSharedData;
// This is purely a debugging aid for developers to see if they are in JIT code space when inspecting raw memory
void EmitDetectionString();
IR::RegisterAllocationPass *RAPass;
IR::RegisterAllocationData *RAData;
+3
View File
@@ -16,8 +16,11 @@ class CPUBackend;
[[nodiscard]] std::unique_ptr<CPUBackend> CreateX86JITCore(FEXCore::Context::Context *ctx,
FEXCore::Core::InternalThreadState *Thread,
bool CompileThread);
void InitializeX86JITSignalHandlers(FEXCore::Context::Context *CTX);
[[nodiscard]] std::unique_ptr<CPUBackend> CreateArm64JITCore(FEXCore::Context::Context *ctx,
FEXCore::Core::InternalThreadState *Thread,
bool CompileThread);
void InitializeArm64JITSignalHandlers(FEXCore::Context::Context *CTX);
} // namespace FEXCore::CPU
+58 -42
View File
@@ -304,32 +304,8 @@ X86JITCore::X86JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalTh
, ThreadState {Thread}
, InitialCodeBuffer {Buffer}
{
{
// Set up pointers that the JIT needs to load
auto &Pointers = ThreadState->CurrentFrame->Pointers.X86;
// Process specific
Pointers.PrintValue = reinterpret_cast<uint64_t>(PrintValue);
Pointers.PrintVectorValue = reinterpret_cast<uint64_t>(PrintVectorValue);
Pointers.RemoveCodeEntryFromJIT = reinterpret_cast<uintptr_t>(&Context::Context::RemoveCodeEntryFromJit);
Pointers.CPUIDObj = reinterpret_cast<uint64_t>(&CTX->CPUID);
{
FEXCore::Utils::MemberFunctionToPointerCast PMF(&FEXCore::CPUIDEmu::RunFunction);
Pointers.CPUIDFunction = PMF.GetConvertedPointer();
}
Pointers.SyscallHandlerObj = reinterpret_cast<uint64_t>(CTX->SyscallHandler);
Pointers.SyscallHandlerFunc = reinterpret_cast<uint64_t>(FEXCore::Context::HandleSyscall);
// Fill in the fallback handlers
InterpreterOps::FillFallbackIndexPointers(Pointers.FallbackHandlerPointers);
// Thread Specific
Pointers.SignalHandlerRefCountPointer = reinterpret_cast<uint64_t>(&Dispatcher->SignalHandlerRefCounter);
}
CurrentCodeBuffer = &InitialCodeBuffer;
EmitDetectionString();
RAPass = Thread->PassManager->GetPass<IR::RegisterAllocationPass>("RA");
@@ -362,6 +338,7 @@ X86JITCore::X86JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalTh
DispatcherConfig config;
config.ExitFunctionLink = reinterpret_cast<uintptr_t>(&ExitFunctionLink);
config.ExitFunctionLinkThis = reinterpret_cast<uintptr_t>(this);
config.StaticRegisterAssignment = ctx->Config.StaticRegisterAllocation;
Dispatcher = std::make_unique<X86Dispatcher>(CTX, ThreadState, config);
DispatchPtr = Dispatcher->DispatchPtr;
@@ -373,26 +350,51 @@ X86JITCore::X86JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalTh
ThreadSharedData.OverflowExceptionInstructionAddress = Dispatcher->OverflowExceptionInstructionAddress;
ThreadSharedData.Dispatcher = Dispatcher.get();
}
// This will register the host signal handler per thread, which is fine
CTX->SignalDelegation->RegisterHostSignalHandler(SIGILL, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
X86JITCore *Core = reinterpret_cast<X86JITCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleSIGILL(Signal, info, ucontext);
}, true);
{
// Set up pointers that the JIT needs to load
auto &Pointers = ThreadState->CurrentFrame->Pointers.X86;
// Process specific
Pointers.PrintValue = reinterpret_cast<uint64_t>(PrintValue);
Pointers.PrintVectorValue = reinterpret_cast<uint64_t>(PrintVectorValue);
Pointers.RemoveCodeEntryFromJIT = reinterpret_cast<uintptr_t>(&Context::Context::RemoveCodeEntryFromJit);
Pointers.CPUIDObj = reinterpret_cast<uint64_t>(&CTX->CPUID);
CTX->SignalDelegation->RegisterHostSignalHandler(SignalDelegator::SIGNAL_FOR_PAUSE, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
X86JITCore *Core = reinterpret_cast<X86JITCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleSignalPause(Signal, info, ucontext);
}, true);
auto GuestSignalHandler = [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack) -> bool {
X86JITCore *Core = reinterpret_cast<X86JITCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleGuestSignal(Signal, info, ucontext, GuestAction, GuestStack);
};
for (uint32_t Signal = 0; Signal <= SignalDelegator::MAX_SIGNALS; ++Signal) {
CTX->SignalDelegation->RegisterHostSignalHandlerForGuest(Signal, GuestSignalHandler);
{
FEXCore::Utils::MemberFunctionToPointerCast PMF(&FEXCore::CPUIDEmu::RunFunction);
Pointers.CPUIDFunction = PMF.GetConvertedPointer();
}
Pointers.SyscallHandlerObj = reinterpret_cast<uint64_t>(CTX->SyscallHandler);
Pointers.SyscallHandlerFunc = reinterpret_cast<uint64_t>(FEXCore::Context::HandleSyscall);
// Fill in the fallback handlers
InterpreterOps::FillFallbackIndexPointers(Pointers.FallbackHandlerPointers);
// Thread Specific
Pointers.SignalHandlerRefCountPointer = reinterpret_cast<uint64_t>(&Dispatcher->SignalHandlerRefCounter);
}
}
void X86JITCore::InitializeSignalHandlers(FEXCore::Context::Context *CTX) {
CTX->SignalDelegation->RegisterHostSignalHandler(SIGILL, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
X86JITCore *Core = reinterpret_cast<X86JITCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleSIGILL(Signal, info, ucontext);
}, true);
CTX->SignalDelegation->RegisterHostSignalHandler(SignalDelegator::SIGNAL_FOR_PAUSE, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
X86JITCore *Core = reinterpret_cast<X86JITCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleSignalPause(Signal, info, ucontext);
}, true);
auto GuestSignalHandler = [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack) -> bool {
X86JITCore *Core = reinterpret_cast<X86JITCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleGuestSignal(Signal, info, ucontext, GuestAction, GuestStack);
};
for (uint32_t Signal = 0; Signal <= SignalDelegator::MAX_SIGNALS; ++Signal) {
CTX->SignalDelegation->RegisterHostSignalHandlerForGuest(Signal, GuestSignalHandler);
}
}
@@ -406,6 +408,13 @@ X86JITCore::~X86JITCore() {
FreeCodeBuffer(InitialCodeBuffer);
}
void X86JITCore::EmitDetectionString() {
const char JITString[] = "FEXJIT::X86JITCore::";
for (char c : JITString) {
db(c);
}
}
void X86JITCore::ClearCache() {
if (*ThreadSharedData.SignalHandlerRefCounterPtr == 0) {
if (!CodeBuffers.empty()) {
@@ -444,6 +453,8 @@ void X86JITCore::ClearCache() {
EmplaceNewCodeBuffer(NewCodeBuffer);
setNewBuffer(NewCodeBuffer.Ptr, NewCodeBuffer.Size);
}
EmitDetectionString();
}
IR::PhysicalRegister X86JITCore::GetPhys(IR::NodeID Node) const {
@@ -805,4 +816,9 @@ uint64_t X86JITCore::ExitFunctionLink(X86JITCore *core, FEXCore::Core::CpuStateF
std::unique_ptr<CPUBackend> CreateX86JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, bool CompileThread) {
return std::make_unique<X86JITCore>(ctx, Thread, AllocateNewCodeBuffer(ctx, CompileThread ? X86JITCore::MAX_CODE_SIZE : X86JITCore::INITIAL_CODE_SIZE), CompileThread);
}
void InitializeX86JITSignalHandlers(FEXCore::Context::Context *CTX) {
X86JITCore::InitializeSignalHandlers(CTX);
}
}
@@ -83,6 +83,8 @@ public:
return Dispatcher->IsAddressInJITCode(Address, IncludeDispatcher, IncludeCompileService);
}
static void InitializeSignalHandlers(FEXCore::Context::Context *CTX);
private:
Label* PendingTargetLabel{};
FEXCore::Context::Context *CTX;
@@ -152,6 +154,9 @@ private:
static uint64_t ExitFunctionLink(X86JITCore* code, FEXCore::Core::CpuStateFrame *Frame, uint64_t *record);
// This is purely a debugging aid for developers to see if they are in JIT code space when inspecting raw memory
void EmitDetectionString();
// This is the initial code buffer that we will fall back to
// In a program without signals and code clearing, we will typically
// only have this code buffer
+87 -23
View File
@@ -4888,6 +4888,7 @@ void OpDispatchBuilder::StoreResult(FEXCore::IR::RegisterClassType Class, FEXCor
OpDispatchBuilder::OpDispatchBuilder(FEXCore::Context::Context *ctx)
: CTX {ctx} {
ResetWorkingList();
InstallHostSpecificOpcodeHandlers();
}
void OpDispatchBuilder::ResetWorkingList() {
@@ -5233,6 +5234,91 @@ void OpDispatchBuilder::InvalidOp(OpcodeArgs) {
#undef OpcodeArgs
void OpDispatchBuilder::InstallHostSpecificOpcodeHandlers() {
static bool Initialized = false;
if (!CTX || Initialized) {
// IRCompaction doesn't set a CTX and doesn't need this anyway
return;
}
#define OPD(prefix, opcode) (((prefix) << 8) | opcode)
constexpr uint16_t PF_38_66 = (1U << 0);
constexpr uint16_t PF_38_F2 = (1U << 1);
constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> H0F38_AES[] = {
{OPD(PF_38_66, 0xDB), 1, &OpDispatchBuilder::AESImcOp},
{OPD(PF_38_66, 0xDC), 1, &OpDispatchBuilder::AESEncOp},
{OPD(PF_38_66, 0xDD), 1, &OpDispatchBuilder::AESEncLastOp},
{OPD(PF_38_66, 0xDE), 1, &OpDispatchBuilder::AESDecOp},
{OPD(PF_38_66, 0xDF), 1, &OpDispatchBuilder::AESDecLastOp},
};
constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> H0F38_CRC[] = {
{OPD(PF_38_F2, 0xF0), 1, &OpDispatchBuilder::CRC32},
{OPD(PF_38_F2, 0xF1), 1, &OpDispatchBuilder::CRC32},
{OPD(PF_38_66 | PF_38_F2, 0xF0), 1, &OpDispatchBuilder::CRC32},
{OPD(PF_38_66 | PF_38_F2, 0xF1), 1, &OpDispatchBuilder::CRC32},
};
#undef OPD
#define OPD(REX, prefix, opcode) ((REX << 9) | (prefix << 8) | opcode)
#define PF_3A_NONE 0
#define PF_3A_66 1
constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> H0F3A_AES[] = {
{OPD(0, PF_3A_66, 0xDF), 1, &OpDispatchBuilder::AESKeyGenAssist},
};
#undef PF_3A_NONE
#undef PF_3A_66
#undef OPD
#define OPD(group, prefix, Reg) (((group - FEXCore::X86Tables::TYPE_GROUP_6) << 5) | (prefix) << 3 | (Reg))
constexpr uint16_t PF_NONE = 0;
constexpr uint16_t PF_66 = 2;
constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> SecondaryExtensionOp_RDRAND[] = {
// GROUP 9
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_NONE, 6), 1, &OpDispatchBuilder::RDRANDOp<false>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_NONE, 7), 1, &OpDispatchBuilder::RDRANDOp<true>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_66, 6), 1, &OpDispatchBuilder::RDRANDOp<false>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_66, 7), 1, &OpDispatchBuilder::RDRANDOp<true>},
};
#undef OPD
constexpr std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> SecondaryModRMExtensionOp_CLZero[] = {
{((3 << 3) | 4), 1, &OpDispatchBuilder::CLZeroOp},
};
auto InstallToTable = [](auto& FinalTable, auto& LocalTable) {
for (auto Op : LocalTable) {
auto OpNum = std::get<0>(Op);
auto Dispatcher = std::get<2>(Op);
for (uint8_t i = 0; i < std::get<1>(Op); ++i) {
LOGMAN_THROW_A_FMT(FinalTable[OpNum + i].OpcodeDispatcher == nullptr, "Duplicate Entry");
FinalTable[OpNum + i].OpcodeDispatcher = Dispatcher;
}
}
};
if (CTX->HostFeatures.SupportsCRC) {
InstallToTable(FEXCore::X86Tables::H0F38TableOps, H0F38_CRC);
}
if (CTX->HostFeatures.SupportsAES) {
InstallToTable(FEXCore::X86Tables::H0F38TableOps, H0F38_AES);
InstallToTable(FEXCore::X86Tables::H0F3ATableOps, H0F3A_AES);
}
if (CTX->HostFeatures.SupportsCLZERO) {
InstallToTable(FEXCore::X86Tables::SecondModRMTableOps, SecondaryModRMExtensionOp_CLZero);
}
if (CTX->HostFeatures.SupportsRAND) {
InstallToTable(FEXCore::X86Tables::SecondInstGroupOps, SecondaryExtensionOp_RDRAND);
}
Initialized = true;
}
void InstallOpcodeHandlers(Context::OperatingMode Mode) {
constexpr std::tuple<uint8_t, uint8_t, X86Tables::OpDispatchPtr> BaseOpTable[] = {
// Instructions
@@ -5794,15 +5880,8 @@ constexpr uint16_t PF_F2 = 3;
// GROUP 9
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_NONE, 1), 1, &OpDispatchBuilder::CMPXCHGPairOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_NONE, 6), 1, &OpDispatchBuilder::RDRANDOp<false>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_NONE, 7), 1, &OpDispatchBuilder::RDRANDOp<true>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_F3, 1), 1, &OpDispatchBuilder::CMPXCHGPairOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_66, 1), 1, &OpDispatchBuilder::CMPXCHGPairOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_66, 6), 1, &OpDispatchBuilder::RDRANDOp<false>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_66, 7), 1, &OpDispatchBuilder::RDRANDOp<true>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_F2, 1), 1, &OpDispatchBuilder::CMPXCHGPairOp},
// GROUP 12
@@ -5868,7 +5947,6 @@ constexpr uint16_t PF_F2 = 3;
// REG /7
{((3 << 3) | 1), 1, &OpDispatchBuilder::RDTSCPOp},
{((3 << 3) | 4), 1, &OpDispatchBuilder::CLZeroOp},
};
// Top bit indicating if it needs to be repeated with {0x40, 0x80} or'd in
@@ -6110,7 +6188,6 @@ constexpr uint16_t PF_F2 = 3;
#define OPD(prefix, opcode) (((prefix) << 8) | opcode)
constexpr uint16_t PF_38_NONE = 0;
constexpr uint16_t PF_38_66 = (1U << 0);
constexpr uint16_t PF_38_F2 = (1U << 1);
constexpr uint16_t PF_38_F3 = (1U << 2);
constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> H0F38Table[] = {
@@ -6176,24 +6253,13 @@ constexpr uint16_t PF_F2 = 3;
{OPD(PF_38_66, 0x40), 1, &OpDispatchBuilder::VectorALUOp<IR::OP_VSMUL, 4>},
{OPD(PF_38_66, 0x41), 1, &OpDispatchBuilder::PHMINPOSUWOp},
{OPD(PF_38_66, 0xDB), 1, &OpDispatchBuilder::AESImcOp},
{OPD(PF_38_66, 0xDC), 1, &OpDispatchBuilder::AESEncOp},
{OPD(PF_38_66, 0xDD), 1, &OpDispatchBuilder::AESEncLastOp},
{OPD(PF_38_66, 0xDE), 1, &OpDispatchBuilder::AESDecOp},
{OPD(PF_38_66, 0xDF), 1, &OpDispatchBuilder::AESDecLastOp},
{OPD(PF_38_NONE, 0xF0), 2, &OpDispatchBuilder::MOVBEOp},
{OPD(PF_38_66, 0xF0), 2, &OpDispatchBuilder::MOVBEOp},
{OPD(PF_38_F2, 0xF0), 1, &OpDispatchBuilder::CRC32},
{OPD(PF_38_F2, 0xF1), 1, &OpDispatchBuilder::CRC32},
{OPD(PF_38_66 | PF_38_F2, 0xF0), 1, &OpDispatchBuilder::CRC32},
{OPD(PF_38_66 | PF_38_F2, 0xF1), 1, &OpDispatchBuilder::CRC32},
{OPD(PF_38_66, 0xF6), 1, &OpDispatchBuilder::ADXOp},
{OPD(PF_38_F3, 0xF6), 1, &OpDispatchBuilder::ADXOp},
};
#undef OPD
#define OPD(REX, prefix, opcode) ((REX << 9) | (prefix << 8) | opcode)
@@ -6225,8 +6291,6 @@ constexpr uint16_t PF_F2 = 3;
{OPD(0, PF_3A_66, 0x40), 1, &OpDispatchBuilder::DPPOp<4>},
{OPD(0, PF_3A_66, 0x41), 1, &OpDispatchBuilder::DPPOp<8>},
{OPD(0, PF_3A_66, 0x42), 1, &OpDispatchBuilder::MPSADBWOp},
{OPD(0, PF_3A_66, 0xDF), 1, &OpDispatchBuilder::AESKeyGenAssist},
};
#undef PF_3A_NONE
#undef PF_3A_66
@@ -1175,6 +1175,8 @@ private:
else
return _LoadMem(Class, Size, ssa0, Invalid(), Align, MEM_OFFSET_SXTX, 1);
}
void InstallHostSpecificOpcodeHandlers();
};
void InstallOpcodeHandlers(Context::OperatingMode Mode);
@@ -621,8 +621,8 @@ void OpDispatchBuilder::FXTRACT(OpcodeArgs) {
}
void OpDispatchBuilder::FNINIT(OpcodeArgs) {
// Init FCW to 0x037
auto NewFCW = _Constant(16, 0x037);
// Init FCW to 0x037F
auto NewFCW = _Constant(16, 0x037F);
_F80LoadFCW(NewFCW);
_StoreContext(2, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
+5 -4
View File
@@ -95,10 +95,11 @@ namespace FEXCore {
LogMan::Msg::EFmt("[{}] Thread has received a signal and hasn't registered itself with the delegate! Programming error!", FHU::Syscalls::gettid());
}
else {
if (Handler.Handler &&
Handler.Handler(Thread, Signal, Info, UContext)) {
// If the host handler handled the fault then we can continue now
return;
for (auto &Handler : Handler.Handlers) {
if (Handler(Thread, Signal, Info, UContext)) {
// If the host handler handled the fault then we can continue now
return;
}
}
if (Handler.FrontendHandler &&
+11 -3
View File
@@ -349,9 +349,17 @@ namespace FEXCore::IR {
// Insert to caches if we generated IR
if (GeneratedIR) {
// Add to thread local ir cache
Core::LocalIREntry Entry = {StartAddr, Length, decltype(Entry.IR)(IRList), decltype(Entry.RAData)(RAData), decltype(Entry.DebugData)(DebugData)};
Thread->LocalIRCache.insert({GuestRIP, std::move(Entry)});
if (Thread->CPUBackend->NeedsRetainedIRCopy()) {
// Add to thread local ir cache
Core::LocalIREntry Entry = {StartAddr, Length, decltype(Entry.IR)(IRList), decltype(Entry.RAData)(RAData), decltype(Entry.DebugData)(DebugData)};
Thread->LocalIRCache.insert({GuestRIP, std::move(Entry)});
}
else {
// If the IR doesn't need to be retained then we can just delete it now
delete DebugData;
delete RAData;
delete IRList;
}
}
}
+13 -33
View File
@@ -5,6 +5,8 @@ tags: ir|parser
$end_info$
*/
#include "Common/StringUtils.h"
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/IR/IREmitter.h>
@@ -42,28 +44,6 @@ enum class DecodeFailure {
};
std::string ltrim(std::string String) {
size_t pos = std::string::npos;
if ((pos = String.find_first_not_of(" \t\n\r")) != std::string::npos) {
String.erase(0, pos);
}
return String;
}
std::string rtrim(std::string String) {
size_t pos = std::string::npos;
if ((pos = String.find_last_not_of(" \t\n\r")) != std::string::npos) {
String.erase(String.begin() + pos + 1, String.end());
}
return String;
}
std::string trim(std::string String) {
return rtrim(ltrim(String));
}
std::string DecodeErrorToString(DecodeFailure Failure) {
switch (Failure) {
case DecodeFailure::DECODE_OKAY: return "Okay";
@@ -295,7 +275,7 @@ class IRParser: public FEXCore::IR::IREmitter {
if (Arg.at(0) != '%') return {DecodeFailure::DECODE_INVALIDCHAR, 0};
// Strip off the type qualifier from the ssa value
std::string SSAName = trim(Arg);
std::string SSAName = FEXCore::StringUtils::Trim(Arg);
const size_t ArgEnd = SSAName.find_first_of(' ');
if (ArgEnd != std::string::npos) {
@@ -382,7 +362,7 @@ class IRParser: public FEXCore::IR::IREmitter {
CurrentDef = &Def;
Def.LineNumber = i;
Line = trim(Line);
Line = FEXCore::StringUtils::Trim(Line);
// Skip empty lines
if (Line.empty()) {
@@ -401,7 +381,7 @@ class IRParser: public FEXCore::IR::IREmitter {
size_t DefinitionEnd = std::string::npos;
if ((DefinitionEnd = Line.find_first_of('=', CurrentPos)) != std::string::npos) {
Def.Definition = Line.substr(0, DefinitionEnd);
Def.Definition = trim(Def.Definition);
Def.Definition = FEXCore::StringUtils::Trim(Def.Definition);
Def.HasDefinition = true;
CurrentPos = DefinitionEnd + 1; // +1 to ensure we go past then assignment
}
@@ -421,7 +401,7 @@ class IRParser: public FEXCore::IR::IREmitter {
size_t SSAEnd = std::string::npos;
if ((SSAEnd = Line.find_last_of(' ', DefinitionEnd)) != std::string::npos) {
std::string Type = Line.substr(SSAEnd + 1, DefinitionEnd - SSAEnd - 1);
Type = trim(Type);
Type = FEXCore::StringUtils::Trim(Type);
auto DefinitionSize = DecodeValue<FEXCore::IR::TypeDefinition>(Type);
if (!CheckPrintError(Def, DefinitionSize.first)) {
@@ -430,7 +410,7 @@ class IRParser: public FEXCore::IR::IREmitter {
Def.Size = DefinitionSize.second;
}
Def.Definition = trim(Line.substr(1, std::min(DefinitionEnd, SSAEnd) - 1));
Def.Definition = FEXCore::StringUtils::Trim(Line.substr(1, std::min(DefinitionEnd, SSAEnd) - 1));
CurrentPos = DefinitionEnd + 1;
}
@@ -447,8 +427,8 @@ class IRParser: public FEXCore::IR::IREmitter {
size_t NameEnd = std::string::npos;
if ((NameEnd = Def.Definition.find_first_of(' ')) != std::string::npos) {
std::string Type = Def.Definition.substr(NameEnd + 1);
Type = trim(Type);
Def.Definition = trim(Def.Definition.substr(0, NameEnd));
Type = FEXCore::StringUtils::Trim(Type);
Def.Definition = FEXCore::StringUtils::Trim(Def.Definition.substr(0, NameEnd));
auto DefinitionSize = DecodeValue<FEXCore::IR::TypeDefinition>(Type);
if (!CheckPrintError(Def, DefinitionSize.first)) return false;
@@ -465,11 +445,11 @@ class IRParser: public FEXCore::IR::IREmitter {
// Let's get the IR op
size_t OpNameEnd = std::string::npos;
std::string RemainingLine = trim(Line.substr(CurrentPos));
std::string RemainingLine = FEXCore::StringUtils::Trim(Line.substr(CurrentPos));
CurrentPos = 0;
if ((OpNameEnd = RemainingLine.find_first_of(" \t\n\r\0", CurrentPos)) != std::string::npos) {
Def.IROp = RemainingLine.substr(CurrentPos, OpNameEnd);
Def.IROp = trim(Def.IROp);
Def.IROp = FEXCore::StringUtils::Trim(Def.IROp);
Def.HasArgs = true;
CurrentPos = OpNameEnd;
}
@@ -486,7 +466,7 @@ class IRParser: public FEXCore::IR::IREmitter {
}
if (Def.HasArgs) {
RemainingLine = trim(RemainingLine.substr(CurrentPos));
RemainingLine = FEXCore::StringUtils::Trim(RemainingLine.substr(CurrentPos));
CurrentPos = 0;
if (RemainingLine.empty()) {
// How did we get here?
@@ -495,7 +475,7 @@ class IRParser: public FEXCore::IR::IREmitter {
else {
while (!RemainingLine.empty()) {
const size_t ArgEnd = RemainingLine.find(',');
std::string Arg = trim(RemainingLine.substr(0, ArgEnd));
std::string Arg = FEXCore::StringUtils::Trim(RemainingLine.substr(0, ArgEnd));
Def.Args.emplace_back(std::move(Arg));
@@ -15,6 +15,7 @@ $end_info$
#include <FEXCore/Utils/BucketList.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/MathUtils.h>
#include <FEXHeaderUtils/TypeDefines.h>
#include <algorithm>
#include <cstddef>
@@ -62,7 +63,7 @@ namespace {
};
static_assert(sizeof(RegisterNode) == 128 * 4);
constexpr size_t REGISTER_NODES_PER_PAGE = PAGE_SIZE / sizeof(RegisterNode);
constexpr size_t REGISTER_NODES_PER_PAGE = FHU::FEX_PAGE_SIZE / sizeof(RegisterNode);
struct RegisterSet {
std::vector<RegisterClass> Classes;
+4 -3
View File
@@ -3,6 +3,7 @@
#include <FEXCore/Utils/CompilerDefs.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXHeaderUtils/Syscalls.h>
#include <FEXHeaderUtils/TypeDefines.h>
#include <array>
#include <sys/mman.h>
@@ -110,10 +111,10 @@ namespace FEXCore::Allocator {
for (int i = 0; i < 64; ++i) {
// Try grabbing a some of the top pages of the range
// x86 allocates some high pages in the top end
void *Ptr = ::mmap(reinterpret_cast<void*>(Size - PAGE_SIZE * i), PAGE_SIZE, PROT_NONE, MAP_FIXED_NOREPLACE | MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
void *Ptr = ::mmap(reinterpret_cast<void*>(Size - FHU::FEX_PAGE_SIZE * i), FHU::FEX_PAGE_SIZE, PROT_NONE, MAP_FIXED_NOREPLACE | MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
if (Ptr != (void*)~0ULL) {
::munmap(Ptr, PAGE_SIZE);
if (Ptr == (void*)(Size - PAGE_SIZE * i)) {
::munmap(Ptr, FHU::FEX_PAGE_SIZE);
if (Ptr == (void*)(Size - FHU::FEX_PAGE_SIZE * i)) {
return true;
}
}
+24 -23
View File
@@ -6,6 +6,7 @@
#include <FEXCore/Utils/MathUtils.h>
#include <FEXHeaderUtils/ScopedSignalMask.h>
#include <FEXHeaderUtils/Syscalls.h>
#include <FEXHeaderUtils/TypeDefines.h>
#include <algorithm>
#include <array>
@@ -43,8 +44,8 @@ namespace Alloc::OSAllocator {
// Lower bound is the starting of the range just past the lower 32bits
constexpr static uintptr_t LOWER_BOUND = 0x1'0000'0000ULL;
uintptr_t UPPER_BOUND_PAGE = UPPER_BOUND / PAGE_SIZE;
constexpr static uintptr_t LOWER_BOUND_PAGE = LOWER_BOUND / PAGE_SIZE;
uintptr_t UPPER_BOUND_PAGE = UPPER_BOUND / FHU::FEX_PAGE_SIZE;
constexpr static uintptr_t LOWER_BOUND_PAGE = LOWER_BOUND / FHU::FEX_PAGE_SIZE;
struct ReservedVMARegion {
uintptr_t Base;
@@ -81,19 +82,19 @@ namespace Alloc::OSAllocator {
// 0x100'0000 Pages
// 1 bit per page for tracking means 0x20'0000 (Pages / 8) bytes of flex space
// Which is 2MB of tracking
uint64_t NumElements = (Size >> PAGE_SHIFT) * sizeof(uint64_t);
uint64_t NumElements = (Size >> FHU::FEX_PAGE_SHIFT) * sizeof(uint64_t);
return sizeof(LiveVMARegion) + FEXCore::FlexBitSet<uint64_t>::Size(NumElements);
}
static void InitializeVMARegionUsed(LiveVMARegion *Region, size_t AdditionalSize) {
size_t SizeOfLiveRegion = FEXCore::AlignUp(LiveVMARegion::GetSizeWithFlexSet(Region->SlabInfo->RegionSize), PAGE_SIZE);
size_t SizeOfLiveRegion = FEXCore::AlignUp(LiveVMARegion::GetSizeWithFlexSet(Region->SlabInfo->RegionSize), FHU::FEX_PAGE_SIZE);
size_t SizePlusManagedData = SizeOfLiveRegion + AdditionalSize;
Region->FreeSpace = Region->SlabInfo->RegionSize - SizePlusManagedData;
size_t NumPages = SizePlusManagedData >> PAGE_SHIFT;
size_t NumPages = SizePlusManagedData >> FHU::FEX_PAGE_SHIFT;
// Memset the full tracking to zero to state nothing used
Region->UsedPages.MemSet(Region->SlabInfo->RegionSize >> PAGE_SHIFT);
Region->UsedPages.MemSet(Region->SlabInfo->RegionSize >> FHU::FEX_PAGE_SHIFT);
// Set our reserved pages
for (size_t i = 0; i < NumPages; ++i) {
// Set our used pages
@@ -119,7 +120,7 @@ namespace Alloc::OSAllocator {
ReservedRegions->erase(ReservedIterator);
// mprotect the new region we've allocated
size_t SizeOfLiveRegion = FEXCore::AlignUp(LiveVMARegion::GetSizeWithFlexSet(ReservedRegion->RegionSize), PAGE_SIZE);
size_t SizeOfLiveRegion = FEXCore::AlignUp(LiveVMARegion::GetSizeWithFlexSet(ReservedRegion->RegionSize), FHU::FEX_PAGE_SIZE);
size_t SizePlusManagedData = UsedSize + SizeOfLiveRegion;
[[maybe_unused]] auto Res = mprotect(reinterpret_cast<void*>(ReservedRegion->Base), SizePlusManagedData, PROT_READ | PROT_WRITE);
@@ -147,7 +148,7 @@ void OSAllocator_64Bit::DetermineVASize() {
size_t Bits = FEXCore::Allocator::DetermineVASize();
uintptr_t Size = 1ULL << Bits;
UPPER_BOUND = Size;
UPPER_BOUND_PAGE = UPPER_BOUND / PAGE_SIZE;
UPPER_BOUND_PAGE = UPPER_BOUND / FHU::FEX_PAGE_SIZE;
}
void *OSAllocator_64Bit::Mmap(void *addr, size_t length, int prot, int flags, int fd, off_t offset) {
@@ -160,13 +161,13 @@ void *OSAllocator_64Bit::Mmap(void *addr, size_t length, int prot, int flags, in
uint64_t Addr = reinterpret_cast<uint64_t>(addr);
// Addr must be page aligned
if (Addr & ~PAGE_MASK) {
if (Addr & ~FHU::FEX_PAGE_MASK) {
return reinterpret_cast<void*>(-EINVAL);
}
// If FD is provided then offset must also be page aligned
if (fd != -1 &&
offset & ~PAGE_MASK) {
offset & ~FHU::FEX_PAGE_MASK) {
return reinterpret_cast<void*>(-EINVAL);
}
@@ -176,10 +177,10 @@ void *OSAllocator_64Bit::Mmap(void *addr, size_t length, int prot, int flags, in
}
bool Fixed = (flags & MAP_FIXED) || (flags & MAP_FIXED_NOREPLACE);
length = FEXCore::AlignUp(length, PAGE_SIZE);
length = FEXCore::AlignUp(length, FHU::FEX_PAGE_SIZE);
uint64_t AddrEnd = Addr + length;
size_t NumberOfPages = length / PAGE_SIZE;
size_t NumberOfPages = length / FHU::FEX_PAGE_SIZE;
// This needs a mutex to be thread safe
FHU::ScopedSignalMaskWithMutex lk(AllocationMutex);
@@ -223,14 +224,14 @@ void *OSAllocator_64Bit::Mmap(void *addr, size_t length, int prot, int flags, in
auto CheckIfRangeFits = [&AllocatedOffset](LiveVMARegion *Region, uint64_t length, int prot, int flags, int fd, off_t offset, uint64_t StartingPosition = 0) -> std::pair<LiveVMARegion*, void*> {
uint64_t AllocatedPage{};
uint64_t NumberOfPages = length >> PAGE_SHIFT;
uint64_t NumberOfPages = length >> FHU::FEX_PAGE_SHIFT;
if (Region->FreeSpace >= length) {
uint64_t LastAllocation =
StartingPosition ?
(StartingPosition - Region->SlabInfo->Base) >> PAGE_SHIFT
(StartingPosition - Region->SlabInfo->Base) >> FHU::FEX_PAGE_SHIFT
: Region->LastPageAllocation;
size_t RegionNumberOfPages = Region->SlabInfo->RegionSize >> PAGE_SHIFT;
size_t RegionNumberOfPages = Region->SlabInfo->RegionSize >> FHU::FEX_PAGE_SHIFT;
// Backward scan
// We need to do a backward scan first to fill any holes
@@ -298,7 +299,7 @@ void *OSAllocator_64Bit::Mmap(void *addr, size_t length, int prot, int flags, in
}
if (AllocatedPage) {
AllocatedOffset = Region->SlabInfo->Base + AllocatedPage * PAGE_SIZE;
AllocatedOffset = Region->SlabInfo->Base + AllocatedPage * FHU::FEX_PAGE_SIZE;
// We need to setup protections for this
void *MMapResult = ::mmap(reinterpret_cast<void*>(AllocatedOffset),
@@ -388,7 +389,7 @@ void *OSAllocator_64Bit::Mmap(void *addr, size_t length, int prot, int flags, in
if (!LiveRegion) {
// Couldn't find a fit in the live regions
// Allocate a new reserved region
size_t lengthOfLiveRegion = FEXCore::AlignUp(LiveVMARegion::GetSizeWithFlexSet(length), PAGE_SIZE);
size_t lengthOfLiveRegion = FEXCore::AlignUp(LiveVMARegion::GetSizeWithFlexSet(length), FHU::FEX_PAGE_SIZE);
size_t lengthPlusManagedData = length + lengthOfLiveRegion;
for (auto it = ReservedRegions->begin(); it != ReservedRegions->end(); ++it) {
if ((*it)->RegionSize >= lengthPlusManagedData) {
@@ -402,7 +403,7 @@ void *OSAllocator_64Bit::Mmap(void *addr, size_t length, int prot, int flags, in
if (LiveRegion) {
// Mark the pages as used
uintptr_t RegionBegin = LiveRegion->SlabInfo->Base;
uintptr_t MappedBegin = (AllocatedOffset - RegionBegin) >> PAGE_SHIFT;
uintptr_t MappedBegin = (AllocatedOffset - RegionBegin) >> FHU::FEX_PAGE_SHIFT;
for (size_t i = 0; i < NumberOfPages; ++i) {
LiveRegion->UsedPages.Set(MappedBegin + i);
@@ -428,11 +429,11 @@ int OSAllocator_64Bit::Munmap(void *addr, size_t length) {
uint64_t Addr = reinterpret_cast<uint64_t>(addr);
if (Addr & ~PAGE_MASK) {
if (Addr & ~FHU::FEX_PAGE_MASK) {
return -EINVAL;
}
if (length & ~PAGE_MASK) {
if (length & ~FHU::FEX_PAGE_MASK) {
return -EINVAL;
}
@@ -443,7 +444,7 @@ int OSAllocator_64Bit::Munmap(void *addr, size_t length) {
// This needs a mutex to be thread safe
FHU::ScopedSignalMaskWithMutex lk(AllocationMutex);
length = FEXCore::AlignUp(length, PAGE_SIZE);
length = FEXCore::AlignUp(length, FHU::FEX_PAGE_SIZE);
uintptr_t PtrBegin = reinterpret_cast<uintptr_t>(addr);
uintptr_t PtrEnd = PtrBegin + length;
@@ -457,8 +458,8 @@ int OSAllocator_64Bit::Munmap(void *addr, size_t length) {
// Live region fully encompasses slab range
uint64_t FreedPages{};
uint32_t SlabPageBegin = (PtrBegin - RegionBegin) >> PAGE_SHIFT;
uint64_t PagesToFree = length >> PAGE_SHIFT;
uint32_t SlabPageBegin = (PtrBegin - RegionBegin) >> FHU::FEX_PAGE_SHIFT;
uint64_t PagesToFree = length >> FHU::FEX_PAGE_SHIFT;
for (size_t i = 0; i < PagesToFree; ++i) {
FreedPages += (*it)->UsedPages.TestAndClear(SlabPageBegin + i) ? 1 : 0;
@@ -4,6 +4,7 @@
#include "HostAllocator.h"
#include <FEXCore/Utils/MathUtils.h>
#include <FEXHeaderUtils/TypeDefines.h>
#include <bitset>
#include <cstddef>
@@ -87,9 +88,9 @@ namespace Alloc {
IntrusiveArenaAllocator(void* Ptr, size_t _Size)
: Begin {reinterpret_cast<uintptr_t>(Ptr)}
, Size {_Size} {
uint64_t NumberOfPages = _Size / PAGE_SIZE;
uint64_t NumberOfPages = _Size / FHU::FEX_PAGE_SIZE;
uint64_t UsedBits = FEXCore::AlignUp(sizeof(IntrusiveArenaAllocator) +
Size / PAGE_SIZE / 8, PAGE_SIZE);
Size / FHU::FEX_PAGE_SIZE / 8, FHU::FEX_PAGE_SIZE);
for (size_t i = 0; i < UsedBits; ++i) {
UsedPages.Set(i);
}
@@ -117,7 +118,7 @@ namespace Alloc {
void *do_allocate(std::size_t bytes, std::size_t alignment) override {
std::scoped_lock<std::mutex> lk{AllocationMutex};
size_t NumberPages = FEXCore::AlignUp(bytes, PAGE_SIZE) / PAGE_SIZE;
size_t NumberPages = FEXCore::AlignUp(bytes, FHU::FEX_PAGE_SIZE) / FHU::FEX_PAGE_SIZE;
uintptr_t AllocatedOffset{};
@@ -161,7 +162,7 @@ namespace Alloc {
LastAllocatedPageOffset = AllocatedOffset + NumberPages;
// Now convert this base page to a pointer and return it
return reinterpret_cast<void*>(Begin + AllocatedOffset * PAGE_SIZE);
return reinterpret_cast<void*>(Begin + AllocatedOffset * FHU::FEX_PAGE_SIZE);
}
return nullptr;
@@ -170,8 +171,8 @@ namespace Alloc {
void do_deallocate(void* p, std::size_t bytes, std::size_t alignment) override {
std::scoped_lock<std::mutex> lk{AllocationMutex};
uintptr_t PageOffset = (reinterpret_cast<uintptr_t>(p) - Begin) / PAGE_SIZE;
size_t NumPages = FEXCore::AlignUp(bytes, PAGE_SIZE) / PAGE_SIZE;
uintptr_t PageOffset = (reinterpret_cast<uintptr_t>(p) - Begin) / FHU::FEX_PAGE_SIZE;
size_t NumPages = FEXCore::AlignUp(bytes, FHU::FEX_PAGE_SIZE) / FHU::FEX_PAGE_SIZE;
// Walk the allocation list and deallocate
uint64_t FreedPages{};
+7
View File
@@ -92,6 +92,13 @@ class LLVMCore;
virtual void CopyNecessaryDataForCompileThread(CPUBackend *Original) {}
virtual bool IsAddressInJITCode(uint64_t Address, bool IncludeDispatcher = true, bool IncludeCompileService = true) const { return false; }
/**
* @brief Does this CPUBackend need its IR to stick around for correct emulation
*
* This should only be used on the interpreter, all other backends can clear their IR
*/
virtual bool NeedsRetainedIRCopy() const { return false; }
using AsmDispatch = FEX_NAKED void(*)(FEXCore::Core::CpuStateFrame *Frame);
using JITCallback = FEX_NAKED void(*)(FEXCore::Core::CpuStateFrame *Frame, uint64_t RIP);
+2
View File
@@ -48,6 +48,8 @@ public:
using IRHandler = std::function<void(uint64_t Addr, FEXCore::IR::IREmitter *IR)>;
virtual void AddIR(IRHandler Handler) {}
virtual uint64_t GetBaseOffset() const { return 0; }
};
+4
View File
@@ -113,6 +113,10 @@ namespace FEXCore::Core {
uint64_t OverflowExceptionHandler{};
uint64_t SignalReturnHandler{};
uint64_t L1Pointer{};
uint64_t LUDIVHandler{};
uint64_t LDIVHandler{};
uint64_t LUREMHandler{};
uint64_t LREMHandler{};
/** @} */
} AArch64;
+3 -2
View File
@@ -8,6 +8,7 @@
#include <utility>
#include <signal.h>
#include <stddef.h>
#include <vector>
namespace FEXCore {
namespace Core {
@@ -96,14 +97,14 @@ namespace Core {
private:
struct HostSignalHandler {
FEXCore::HostSignalDelegatorFunction Handler{};
std::vector<FEXCore::HostSignalDelegatorFunction> Handlers{};
FEXCore::HostSignalDelegatorFunction FrontendHandler{};
};
std::array<HostSignalHandler, MAX_SIGNALS + 1> HostHandlers{};
protected:
void SetHostSignalHandler(int Signal, HostSignalDelegatorFunction Func, bool Required) {
HostHandlers[Signal].Handler = std::move(Func);
HostHandlers[Signal].Handlers.push_back(std::move(Func));
}
void SetFrontendHostSignalHandler(int Signal, HostSignalDelegatorFunction Func, bool Required) {
HostHandlers[Signal].FrontendHandler = std::move(Func);
@@ -38,7 +38,6 @@ namespace FEXCore::Core {
struct DebugDataSubblock {
uintptr_t HostCodeStart;
uint32_t HostCodeSize;
IR::NodeID SSAId;
};
/**
@@ -48,10 +47,6 @@ namespace FEXCore::Core {
*/
struct DebugData {
uint64_t HostCodeSize; ///< The size of the code generated in the host JIT
uint64_t GuestCodeSize; ///< The size of the guest side code
uint64_t GuestInstructionCount; ///< Number of guest instructions
uint64_t TimeSpentInCode; ///< How long this code has spent time running
uint64_t RunCount; ///< Number of times this block of code has been run
std::vector<DebugDataSubblock> Subblocks;
};
+5
View File
@@ -4,6 +4,10 @@
#include <FEXCore/IR/IR.h>
namespace FEXCore {
class CodeLoader;
}
namespace FEXCore::Context {
struct Context;
}
@@ -48,6 +52,7 @@ namespace FEXCore::HLE {
virtual FEXCore::IR::SyscallFlags GetSyscallFlags(uint64_t Syscall) const { return FEXCore::IR::SyscallFlags::DEFAULT; }
SyscallOSABI GetOSABI() const { return OSABI; }
virtual FEXCore::CodeLoader *GetCodeLoader() const { return nullptr; }
protected:
SyscallOSABI OSABI;
@@ -0,0 +1,11 @@
#pragma once
#include <cstddef>
namespace FHU {
// FEX assumes an operating page size of 4096
// To work around build systems that build on a 16k/64k page size, define our page size here
// Don't use the system provided PAGE_SIZE define because of this.
constexpr size_t FEX_PAGE_SIZE = 4096;
constexpr size_t FEX_PAGE_SHIFT = 12;
constexpr size_t FEX_PAGE_MASK = ~(FEX_PAGE_SIZE - 1);
}
+3 -37
View File
@@ -1,3 +1,4 @@
[中文](https://github.com/FEX-Emu/FEX/blob/main/docs/Readme_CN.md)
# FEX - Fast x86 emulation frontend
FEX allows you to run x86 and x86-64 binaries on an AArch64 host, similar to qemu-user and box86.
It has native support for a rootfs overlay, so you don't need to chroot, as well as some thunklibs so it can forward things like GL to the host.
@@ -16,7 +17,7 @@ This command will walk you through installing FEX through a PPA, and downloading
Ubuntu PPA is updated with our monthly releases.
### For everyone else
Follow the guide on the official FEX-Emu Wiki [Here](https://wiki.fex-emu.org/index.php/QuickStartGuide)
Please see [Building FEX](#building-fex).
## Getting Started
FEX has been tested to build and run on ARMv8.0, ARMv8.1+, and x86-64(AVX or newer) hardware.
@@ -28,43 +29,8 @@ On AArch64 hosts the user **MUST** have an x86-64 RootFS [Creating a RootFS](#Ro
### Navigating the Source
See the [Source Outline](docs/SourceOutline.md) for more information.
### Dependencies
* cmake (version 3.14 minimum)
* ninja-build
* clang (version 10 minimum for C++20)
* libglfw3-dev (For GUI)
* libsdl2-dev (For GUI)
* libepoxy-dev (For GUI)
* g++-x86-64-linux-gnu (For building thunks)
* nasm (only if building tests)
### Building FEX
After installing the dependencies you can now build FEX.
```Shell
git clone https://github.com/FEX-Emu/FEX.git
cd FEX
git submodule update --init
mkdir Build
cd Build
CC=clang CXX=clang++ cmake -DCMAKE_INSTALL_PREFIX=/usr -DCMAKE_BUILD_TYPE=Release -DENABLE_LTO=True -DBUILD_TESTS=False -G Ninja ..
ninja
```
### Installation
```Shell
sudo ninja install
```
#### On AArch64 Hosts
You can install a binfmt_misc handler for both 32bit and 64bit x86 execution directly from the environment. If you already have box86's 32bit binfmt_misc handler installed then I don't recommend installing FEX's until it is useful. Make sure to have run install prior to this, otherwise binfmt_misc will install an old handler even if the executable has been updated.
```Shell
sudo ninja binfmt_misc_32
sudo ninja binfmt_misc_64
```
### More information
This wiki page can contain more information about setting up FEX on your device
https://wiki.fex-emu.org/index.php/Development:Setting_up_FEX
Follow the guide on the official FEX-Emu Wiki [here](https://wiki.fex-emu.org/index.php/Development:Setting_up_FEX).
### RootFS generation
AArch64 hosts require a rootfs for running applications.
+3 -3
View File
@@ -18,11 +18,11 @@ BigCoreIDs = {
tuple([0x41, 0xd44]): "cortex-x1",
tuple([0x41, 0xd47]):
[ ["cortex-a78", "0.0"],
["cortex-a710", "999.0"], # Doesn't exist in clang as of version 13
["cortex-a710", "14.0"],
],
tuple([0x41, 0xd48]):
[ ["cortex-x1", "0.0"],
["cortex-x2", "999.0"], # Doesn't exist in clang as of version 13
["cortex-x2", "14.0"],
],
tuple([0x41, 0xd0c]): "neoverse-n1",
tuple([0x41, 0xd49]): "neoverse-n2",
@@ -46,7 +46,7 @@ LittleCoreIDs = {
tuple([0x41, 0xd05]): "cortex-a55",
tuple([0x41, 0xd46]):
[ ["cortex-a55", "0.0"],
["cortex-a510", "999.0"], # Doesn't exist in clang as of version 13
["cortex-a510", "14.0"],
],
# Qualcomm
+8
View File
@@ -366,6 +366,9 @@ class ELFCodeLoader2 final : public FEXCore::CodeLoader {
if (auto elf = LoadElfFile(MainElf, &BrkBase, Mapper, Unmapper)) {
LoadBase = *elf;
if (MainElf.ehdr.e_type == ET_DYN) {
BaseOffset = LoadBase;
}
} else {
LogMan::Msg::EFmt("Failed to load elf file");
return false;
@@ -620,6 +623,10 @@ class ELFCodeLoader2 final : public FEXCore::CodeLoader {
size = AuxTabSize;
}
uint64_t GetBaseOffset() const override {
return BaseOffset;
}
bool Is64BitMode() {
return MainElf.type == ::ELFLoader::ELFContainer::TYPE_X86_64;
}
@@ -643,5 +650,6 @@ class ELFCodeLoader2 final : public FEXCore::CodeLoader {
uint64_t AuxTabBase, AuxTabSize;
uint64_t ArgumentBackingSize{};
uint64_t EnvironmentBackingSize{};
uint64_t BaseOffset{};
};
+5 -4
View File
@@ -22,6 +22,7 @@
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/MathUtils.h>
#include <FEXHeaderUtils/Syscalls.h>
#include <FEXHeaderUtils/TypeDefines.h>
#include <fmt/format.h>
@@ -396,14 +397,14 @@ namespace FEX::HarnessHelper {
};
if (LimitedSize) {
DoMMap(0xe000'0000, PAGE_SIZE * 10);
DoMMap(0xe000'0000, FHU::FEX_PAGE_SIZE * 10);
// SIB8
// We test [-128, -126] (Bottom)
// We test [-8, 8] (Middle)
// We test [120, 127] (Top)
// Can fit in two pages
DoMMap(0xe800'0000 - PAGE_SIZE, PAGE_SIZE * 2);
DoMMap(0xe800'0000 - FHU::FEX_PAGE_SIZE, FHU::FEX_PAGE_SIZE * 2);
}
else {
// This is scratch memory location and SIB8 location
@@ -413,7 +414,7 @@ namespace FEX::HarnessHelper {
}
// Map in the memory region for the test file
size_t Length = FEXCore::AlignUp(RawFile.size(), PAGE_SIZE);
size_t Length = FEXCore::AlignUp(RawFile.size(), FHU::FEX_PAGE_SIZE);
Code_start_page = reinterpret_cast<uint64_t>(DoMMap(Code_start_page, Length));
mprotect(reinterpret_cast<void*>(Code_start_page), Length, PROT_READ | PROT_WRITE | PROT_EXEC);
RIP = Code_start_page;
@@ -446,7 +447,7 @@ namespace FEX::HarnessHelper {
bool Is64BitMode() const { return Config.Is64BitMode(); }
private:
constexpr static uint64_t STACK_SIZE = PAGE_SIZE;
constexpr static uint64_t STACK_SIZE = FHU::FEX_PAGE_SIZE;
constexpr static uint64_t STACK_OFFSET = 0xc000'0000;
// Zero is special case to know when we are done
uint64_t Code_start_page = 0x1'0000;
@@ -387,7 +387,7 @@ uint64_t FileManager::Lstat(const char *pathname, void *buf) {
return Result;
}
return ::lstat(SelfPath, reinterpret_cast<struct stat*>(buf));
return ::lstat(pathname, reinterpret_cast<struct stat*>(buf));
}
uint64_t FileManager::Access(const char *pathname, [[maybe_unused]] int mode) {
+38 -34
View File
@@ -3,6 +3,7 @@
#include <FEXCore/Utils/MathUtils.h>
#include <FEXHeaderUtils/Syscalls.h>
#include <FEXHeaderUtils/TypeDefines.h>
#include <bitset>
#include <map>
@@ -20,8 +21,8 @@ namespace FEX::HLE {
class MemAllocator32Bit final : public FEX::HLE::MemAllocator {
private:
static constexpr uint64_t BASE_KEY = 16;
const uint64_t TOP_KEY = 0xFFFF'F000ULL >> PAGE_SHIFT;
const uint64_t TOP_KEY32BIT = 0x1F'F000ULL >> PAGE_SHIFT;
const uint64_t TOP_KEY = 0xFFFF'F000ULL >> FHU::FEX_PAGE_SHIFT;
const uint64_t TOP_KEY32BIT = 0x7FFF'F000ULL >> FHU::FEX_PAGE_SHIFT;
public:
MemAllocator32Bit() {
@@ -132,10 +133,10 @@ uint64_t MemAllocator32Bit::FindPageRange_TopDown(uint64_t Start, size_t Pages)
void *MemAllocator32Bit::mmap(void *addr, size_t length, int prot, int flags, int fd, off_t offset) {
std::scoped_lock<std::mutex> lk{AllocMutex};
size_t PagesLength = FEXCore::AlignUp(length, PAGE_SIZE) >> PAGE_SHIFT;
size_t PagesLength = FEXCore::AlignUp(length, FHU::FEX_PAGE_SIZE) >> FHU::FEX_PAGE_SHIFT;
uintptr_t Addr = reinterpret_cast<uintptr_t>(addr);
uintptr_t PageAddr = Addr >> PAGE_SHIFT;
uintptr_t PageAddr = Addr >> FHU::FEX_PAGE_SHIFT;
// Define MAP_FIXED_NOREPLACE ourselves to ensure we always parse this flag
constexpr int FEX_MAP_FIXED_NOREPLACE = 0x100000;
@@ -143,13 +144,13 @@ void *MemAllocator32Bit::mmap(void *addr, size_t length, int prot, int flags, in
(flags & FEX_MAP_FIXED_NOREPLACE));
// Both Addr and length must be page aligned
if (Addr & ~PAGE_MASK) {
if (Addr & ~FHU::FEX_PAGE_MASK) {
return reinterpret_cast<void*>(-EINVAL);
}
// If we do have an fd then offset must be page aligned
if (fd != -1 &&
offset & ~PAGE_MASK) {
offset & ~FHU::FEX_PAGE_MASK) {
return reinterpret_cast<void*>(-EINVAL);
}
@@ -170,6 +171,9 @@ void *MemAllocator32Bit::mmap(void *addr, size_t length, int prot, int flags, in
bool Map32Bit = flags & FEX::HLE::X86_64_MAP_32BIT;
// Remove the MAP_32BIT flag if it exists now
flags &= ~FEX::HLE::X86_64_MAP_32BIT;
auto AllocateNoHint = [&]() -> void*{
bool Wrapped = false;
uint64_t BottomPage = Map32Bit && (LastScanLocation >= LastKeyLocation32Bit) ? LastKeyLocation32Bit : LastScanLocation;
@@ -190,7 +194,7 @@ restart:
{
// Try and map the range
void *MappedPtr = ::mmap(
reinterpret_cast<void*>(LowerPage<< PAGE_SHIFT),
reinterpret_cast<void*>(LowerPage << FHU::FEX_PAGE_SHIFT),
length,
prot,
flags | FEX_MAP_FIXED_NOREPLACE,
@@ -202,10 +206,10 @@ restart:
return reinterpret_cast<void*>(-errno);
}
else if (MappedPtr == MAP_FAILED ||
MappedPtr >= reinterpret_cast<void*>(TOP_KEY << PAGE_SHIFT)) {
MappedPtr >= reinterpret_cast<void*>(TOP_KEY << FHU::FEX_PAGE_SHIFT)) {
// Handles the case where MAP_FIXED_NOREPLACE failed with MAP_FAILED
// or if the host system's kernel isn't new enough then it returns the wrong pointer
if (MappedPtr >= reinterpret_cast<void*>(TOP_KEY << PAGE_SHIFT)) {
if (MappedPtr >= reinterpret_cast<void*>(TOP_KEY << FHU::FEX_PAGE_SHIFT)) {
// Make sure to munmap this so we don't leak memory
::munmap(MappedPtr, length);
}
@@ -251,14 +255,14 @@ restart:
}
else {
void *MappedPtr = ::mmap(
reinterpret_cast<void*>(PageAddr << PAGE_SHIFT),
PagesLength << PAGE_SHIFT,
reinterpret_cast<void*>(PageAddr << FHU::FEX_PAGE_SHIFT),
PagesLength << FHU::FEX_PAGE_SHIFT,
prot,
flags,
fd,
offset);
if (MappedPtr >= reinterpret_cast<void*>(TOP_KEY << PAGE_SHIFT) &&
if (MappedPtr >= reinterpret_cast<void*>(TOP_KEY << FHU::FEX_PAGE_SHIFT) &&
(flags & FEX_MAP_FIXED_NOREPLACE)) {
// Handles the case where MAP_FIXED_NOREPLACE isn't handled by the host system's
// kernel and returns the wrong pointer
@@ -279,19 +283,19 @@ restart:
int MemAllocator32Bit::munmap(void *addr, size_t length) {
std::scoped_lock<std::mutex> lk{AllocMutex};
size_t PagesLength = FEXCore::AlignUp(length, PAGE_SIZE) >> PAGE_SHIFT;
size_t PagesLength = FEXCore::AlignUp(length, FHU::FEX_PAGE_SIZE) >> FHU::FEX_PAGE_SHIFT;
uintptr_t Addr = reinterpret_cast<uintptr_t>(addr);
uintptr_t PageAddr = Addr >> PAGE_SHIFT;
uintptr_t PageAddr = Addr >> FHU::FEX_PAGE_SHIFT;
uintptr_t PageEnd = PageAddr + PagesLength;
// Both Addr and length must be page aligned
if (Addr & ~PAGE_MASK) {
if (Addr & ~FHU::FEX_PAGE_MASK) {
return -EINVAL;
}
if (length & ~PAGE_MASK) {
if (length & ~FHU::FEX_PAGE_MASK) {
return -EINVAL;
}
@@ -307,7 +311,7 @@ int MemAllocator32Bit::munmap(void *addr, size_t length) {
while (PageAddr != PageEnd) {
// Always pass to munmap, it may be something allocated we aren't tracking
int Result = ::munmap(reinterpret_cast<void*>(PageAddr << PAGE_SHIFT), PAGE_SIZE);
int Result = ::munmap(reinterpret_cast<void*>(PageAddr << FHU::FEX_PAGE_SHIFT), FHU::FEX_PAGE_SIZE);
if (Result != 0) {
return -errno;
}
@@ -323,8 +327,8 @@ int MemAllocator32Bit::munmap(void *addr, size_t length) {
}
void *MemAllocator32Bit::mremap(void *old_address, size_t old_size, size_t new_size, int flags, void *new_address) {
size_t OldPagesLength = FEXCore::AlignUp(old_size, PAGE_SIZE) >> PAGE_SHIFT;
size_t NewPagesLength = FEXCore::AlignUp(new_size, PAGE_SIZE) >> PAGE_SHIFT;
size_t OldPagesLength = FEXCore::AlignUp(old_size, FHU::FEX_PAGE_SIZE) >> FHU::FEX_PAGE_SHIFT;
size_t NewPagesLength = FEXCore::AlignUp(new_size, FHU::FEX_PAGE_SIZE) >> FHU::FEX_PAGE_SHIFT;
{
std::scoped_lock<std::mutex> lk{AllocMutex};
@@ -335,12 +339,12 @@ void *MemAllocator32Bit::mremap(void *old_address, size_t old_size, size_t new_s
if (!(flags & MREMAP_DONTUNMAP)) {
// Unmap the old location
uintptr_t OldAddr = reinterpret_cast<uintptr_t>(old_address);
SetFreePages(OldAddr >> PAGE_SHIFT, OldPagesLength);
SetFreePages(OldAddr >> FHU::FEX_PAGE_SHIFT, OldPagesLength);
}
// Map the new pages
uintptr_t NewAddr = reinterpret_cast<uintptr_t>(MappedPtr);
SetUsedPages(NewAddr >> PAGE_SHIFT, NewPagesLength);
SetUsedPages(NewAddr >> FHU::FEX_PAGE_SHIFT, NewPagesLength);
}
else {
return reinterpret_cast<void*>(-errno);
@@ -348,15 +352,15 @@ void *MemAllocator32Bit::mremap(void *old_address, size_t old_size, size_t new_s
}
else {
uintptr_t OldAddr = reinterpret_cast<uintptr_t>(old_address);
uintptr_t OldPageAddr = OldAddr >> PAGE_SHIFT;
uintptr_t OldPageAddr = OldAddr >> FHU::FEX_PAGE_SHIFT;
if (NewPagesLength < OldPagesLength) {
void *MappedPtr = ::mremap(old_address, old_size, new_size, flags & ~MREMAP_MAYMOVE);
if (MappedPtr != MAP_FAILED) {
// Clear the pages that we just shrunk
size_t NewPagesLength = FEXCore::AlignUp(new_size, PAGE_SIZE) >> PAGE_SHIFT;
uintptr_t NewPageAddr = reinterpret_cast<uintptr_t>(MappedPtr) >> PAGE_SHIFT;
size_t NewPagesLength = FEXCore::AlignUp(new_size, FHU::FEX_PAGE_SIZE) >> FHU::FEX_PAGE_SHIFT;
uintptr_t NewPageAddr = reinterpret_cast<uintptr_t>(MappedPtr) >> FHU::FEX_PAGE_SHIFT;
SetFreePages(NewPageAddr + NewPagesLength, OldPagesLength - NewPagesLength);
return MappedPtr;
}
@@ -380,9 +384,9 @@ void *MemAllocator32Bit::mremap(void *old_address, size_t old_size, size_t new_s
if (MappedPtr != MAP_FAILED) {
// Map the new pages
size_t NewPagesLength = FEXCore::AlignUp(new_size, PAGE_SIZE) >> PAGE_SHIFT;
size_t NewPagesLength = FEXCore::AlignUp(new_size, FHU::FEX_PAGE_SIZE) >> FHU::FEX_PAGE_SHIFT;
uintptr_t NewAddr = reinterpret_cast<uintptr_t>(MappedPtr);
SetUsedPages(NewAddr >> PAGE_SHIFT, NewPagesLength);
SetUsedPages(NewAddr >> FHU::FEX_PAGE_SHIFT, NewPagesLength);
return MappedPtr;
}
else if (!(flags & MREMAP_MAYMOVE)) {
@@ -416,13 +420,13 @@ void *MemAllocator32Bit::mremap(void *old_address, size_t old_size, size_t new_s
// If we have both MREMAP_DONTUNMAP not set and the new pointer is at a new location
// Make sure to clear the old mapping
uintptr_t OldAddr = reinterpret_cast<uintptr_t>(old_address);
SetFreePages(OldAddr >> PAGE_SHIFT , OldPagesLength);
SetFreePages(OldAddr >> FHU::FEX_PAGE_SHIFT , OldPagesLength);
}
// Map the new pages
size_t NewPagesLength = FEXCore::AlignUp(new_size, PAGE_SIZE) >> PAGE_SHIFT;
size_t NewPagesLength = FEXCore::AlignUp(new_size, FHU::FEX_PAGE_SIZE) >> FHU::FEX_PAGE_SHIFT;
uintptr_t NewAddr = reinterpret_cast<uintptr_t>(MappedPtr);
SetUsedPages(NewAddr >> PAGE_SHIFT, NewPagesLength);
SetUsedPages(NewAddr >> FHU::FEX_PAGE_SHIFT, NewPagesLength);
return MappedPtr;
}
@@ -445,7 +449,7 @@ uint64_t MemAllocator32Bit::shmat(int shmid, const void* shmaddr, int shmflg, ui
}
uintptr_t NewAddr = reinterpret_cast<uintptr_t>(Result);
uintptr_t NewPageAddr = NewAddr >> PAGE_SHIFT;
uintptr_t NewPageAddr = NewAddr >> FHU::FEX_PAGE_SHIFT;
// Add to the map
PageToShm[NewPageAddr] = shmid;
@@ -457,7 +461,7 @@ uint64_t MemAllocator32Bit::shmat(int shmid, const void* shmaddr, int shmflg, ui
if (shmctl(shmid, IPC_STAT, &buf) == 0) {
// Map the new pages
size_t NewPagesLength = buf.shm_segsz >> PAGE_SHIFT;
size_t NewPagesLength = buf.shm_segsz >> FHU::FEX_PAGE_SHIFT;
SetUsedPages(NewPageAddr, NewPagesLength);
}
@@ -475,7 +479,7 @@ uint64_t MemAllocator32Bit::shmat(int shmid, const void* shmaddr, int shmflg, ui
uint64_t PagesLength{};
if (shmctl(shmid, IPC_STAT, &buf) == 0) {
PagesLength = FEXCore::AlignUp(buf.shm_segsz, PAGE_SIZE) >> PAGE_SHIFT;
PagesLength = FEXCore::AlignUp(buf.shm_segsz, FHU::FEX_PAGE_SIZE) >> FHU::FEX_PAGE_SHIFT;
}
else {
return -EINVAL;
@@ -501,7 +505,7 @@ restart:
// Try and map the range
void *MappedPtr = ::shmat(
shmid,
reinterpret_cast<const void*>(LowerPage << PAGE_SHIFT),
reinterpret_cast<const void*>(LowerPage << FHU::FEX_PAGE_SHIFT),
shmflg);
if (MappedPtr == MAP_FAILED) {
@@ -544,7 +548,7 @@ restart:
}
}
uint64_t MemAllocator32Bit::shmdt(const void* shmaddr) {
uint32_t AddrPage = reinterpret_cast<uint64_t>(shmaddr) >> PAGE_SHIFT;
uint32_t AddrPage = reinterpret_cast<uint64_t>(shmaddr) >> FHU::FEX_PAGE_SHIFT;
auto it = PageToShm.find(AddrPage);
if (it == PageToShm.end()) {
+1 -1
View File
@@ -133,7 +133,7 @@ public:
uint64_t HandleBRK(FEXCore::Core::CpuStateFrame *Frame, void *Addr);
FEX::HLE::FileManager FM;
FEXCore::CodeLoader *GetCodeLoader() const { return LocalLoader; }
FEXCore::CodeLoader *GetCodeLoader() const override { return LocalLoader; }
void SetCodeLoader(FEXCore::CodeLoader *Loader) { LocalLoader = Loader; }
FEX::HLE::SignalDelegator *GetSignalDelegator() { return SignalDelegation; }
+3
View File
@@ -62,6 +62,9 @@ void MsgHandler(LogMan::DebugLevels Level, char const *Message) {
void AssertHandler(char const *Message) {
fmt::print("[ASSERT] {}\n", Message);
// make sure buffers are flushed
fflush(nullptr);
}
namespace {
+10 -3
View File
@@ -2,11 +2,18 @@ if (ENABLE_VISUAL_DEBUGGER)
add_subdirectory(Debugger/)
endif()
add_subdirectory(FEXConfig/)
if (NOT TERMUX_BUILD)
# Termux builds can't rely on X11 packages
# SDL2 isn't even compiled with GL support so our GUIs wouldn't even work
add_subdirectory(FEXConfig/)
add_subdirectory(FEXLogServer/)
# Disable FEXRootFSFetcher on Termux, it doesn't even work there
add_subdirectory(FEXRootFSFetcher/)
endif()
add_subdirectory(FEXGetConfig/)
add_subdirectory(FEXMountDaemon/)
add_subdirectory(FEXLogServer/)
add_subdirectory(FEXRootFSFetcher/)
set(NAME Opt)
set(SRCS Opt.cpp)
+40
View File
@@ -0,0 +1,40 @@
# FEXCore custom CPUID functions
## 4000_0000h - Hypervisor information function
* Follows VMWare and Microsoft's hypervisor information proposal
* https://lwn.net/Articles/301888/
* https://docs.microsoft.com/en-us/virtualization/hyper-v-on-windows/tlfs/feature-discovery
* EAX - The maximum input value for the hypervisor CPUID information
* 4000_0001h
* EBX - Hypervisor vendor ID signature
* 'FEXI' - 4958_4546h
* ECX - Hypervisor vendor ID signature
* 'FEXI' - 4958_4546h
* EDX - Hypervisor vendor ID signature
* 'EMU\0' - 0055_4d45h
* memcpy ebx:ecx:edx in to a 12 byte string to get 'FEXIFEXIEMU\0' for determining running under FEX
## 4000_0001h - Hypervisor config function
### Sub-Leaf 0: ECX == 0
* EAX:
* Bits EAX[3:0] - Host architecture
* 0 - Unknown architecture
* 1 - x86_64
* 2 - AArch64
* 3-15: **Reserved**
* EBX - **Reserved** - Read as zero
* ECX - **Reserved** - Read as zero
* EDX - **Reserved** - Read as zero
### Sub-Leaf 0000_0001 - FFFF_FFFF: **Reserved**
## 4000_0002h - 4000_000Fh
* **Reserved range**
* Returns zero until implemented
## 4000_0010h - 4FFF_FFFFh
* **Undefined**
* FEX-Emu will return zero until implemented
+73
View File
@@ -0,0 +1,73 @@
[English](https://github.com/FEX-Emu/FEX/blob/main/Readme.md)
# FEX —— 快速的x86模拟器前端
FEX和qemu-user以及box86类似,允许你在AArch64的host端运行x86和x86-64二进制程序。
FEX原生支持rootfs(作为guest程序的运行环境),所以无需使用chroot。同时支持thunklibs将guest程序所用到的库转发到host,例如:libGL。
FEX为guest程序提供Linux 5.0的接口(系统调用),同时支持AArch64和x86-64做为host。
FEX处于重度开发阶段,所以会有很多改善。
## 快速指引
### Ubuntu 20.04, 21.04, 21.10, 22.04
在终端执行以下命令添加PPA去安装FEX。
`curl --silent https://raw.githubusercontent.com/FEX-Emu/FEX/main/Scripts/InstallFEX.py --output /tmp/InstallFEX.py && python3 /tmp/InstallFEX.py && rm /tmp/InstallFEX.py`
这条命令将会引导你通过PPA安装FEX,然后下载FEX所需的RootFS。
Ubuntu下的PPA 随FEX月度发布更新。
### 其他系统
参考[这里](https://wiki.fex-emu.org/index.php/QuickStartGuide)
## 开始
FEX在ARMv8.0,ARMv8.1+和x86-64(支持AVX或更新处理器)硬件上进行过编译和运行测试。
不支持ARMv7以及老旧的x86处理器。
同时需要确保操作系统为Linux。FEX在Ubuntu 20.04,20.10和21.04以及Arch Linux上测试过。
在AArch64 host端,用户需要准备x86-64 RootFS[创建RootFS](#RootFS-Generation)。
### 源码导览
详见[源码大纲](docs/SourceOutline.md)。
### 编译依赖
* cmake (version 3.14 minimum)
* ninja-build
* clang (version 10 minimum for C++20)
* libglfw3-dev (For GUI)
* libsdl2-dev (For GUI)
* libepoxy-dev (For GUI)
* g++-x86-64-linux-gnu (For building thunks)
* nasm (only if building tests)
### 编译FEX
安装完依赖后,通过以下命令进行编译。
```Shell
git clone https://github.com/FEX-Emu/FEX.git
cd FEX
git submodule update --init
mkdir Build
cd Build
CC=clang CXX=clang++ cmake -DCMAKE_INSTALL_PREFIX=/usr -DCMAKE_BUILD_TYPE=Release -DENABLE_LTO=True -DBUILD_TESTS=False -G Ninja ..
ninja
```
### 安装
```Shell
sudo ninja install
```
### 关于AArch64 Hosts
在AArch64使用binfmt_misc(执行下述命令)可以支持32位和64位x86程序直接运行。如果已经安装了box86 binfmt_misc配置,在FEX达到可用状态前我并不建议安装FEX进行替代。请确保install命令在下述命令前执行,不然binfmt_misc将依旧使用旧版本的FEX,即使FEX已经更新。
```Shell
sudo ninja binfmt_misc_32
sudo ninja binfmt_misc_64
```
### 更多信息
更多关于FEX和平台相关的设置信息请参考以下维基页面:
https://wiki.fex-emu.org/index.php/Development:Setting_up_FEX
### 创建RootFS
AArch64 host端需要一个rootfs去运行guest程序。参考以下维基页面从头开始创建一个rootfs
https://wiki.fex-emu.org/index.php/Development:Setting_up_RootFS
![FEX diagram](Diagram.svg)
+8
View File
@@ -60,6 +60,14 @@ Follow the steps in: https://github.com/FEX-Emu/FEX-ppa/blob/main/README_ppa.md
* Requires PPA GPG key signing access
* Wait the 20-30 minutes for Ubuntu PPA to build and publish the binaries
## Termux package update steps
* Clone https://github.com/termux/termux-packages
* Update the package script with thew new version tag (https://github.com/termux/termux-packages/blob/master/packages/fex/build.sh)
* ***!!! Test the build locally on an Android device !!!***
* Check the wiki for a command to run in the local changed repo on-device
* https://github.com/termux/termux-packages/wiki/Building-packages#how-to-build-package
* Submit a PR to upstream on github to get it picked up
## @FEX_Emu twitter account steps
* Requires @FEX_Emu twitter account access
* Create a tweet with some small blurb/sizzle text about some relevant changes in this tagged version
+1 -1
View File
@@ -1,4 +1,4 @@
# FEX-2203
# FEX-2204
## External/FEXCore
See [FEXCore/Readme.md](../External/FEXCore/Readme.md) for more details
-3
View File
@@ -1,6 +1,3 @@
# Needs Precision checking
Test_32Bit_X87/D9_FD.asm
# Relies on undefined behaviour
Test_32Bit_X87/D9_F9.asm
+18
View File
@@ -0,0 +1,18 @@
%ifdef CONFIG
{
"RegData": {
"RAX": "0x037F"
},
"Mode": "32BIT"
}
%endif
mov eax, 0
mov esp, 0xe0000008
fninit
; Ensures that fnstcw after fninit sets the correct value
fnstcw [esp]
mov ax, word [esp]
hlt
-1
View File
@@ -3,7 +3,6 @@ Test_REP/F3_52.asm
Test_REP/F3_53.asm
Test_TwoByte/0F_52.asm
Test_TwoByte/0F_53.asm
Test_X87/D9_FD.asm
# Not supported in userspace in all cases
Test_Secondary/15_F3_00.asm
+3
View File
@@ -0,0 +1,3 @@
# Intel doesn't support CLZero, which the CI uses
Test_SecondaryModRM/Reg_7_4.asm
Test_SecondaryModRM/Reg_7_4_2.asm
+30
View File
@@ -0,0 +1,30 @@
%ifdef CONFIG
{
"RegData": {
"MM6": ["0x8000000000000000", "0x3FFF"],
"MM7": ["0xD000000000000000", "0xC001"]
},
"MemoryRegions": {
"0x100000000": "4096"
}
}
%endif
lea rdx, [rel data2]
fld tword [rdx + 8 * 0]
lea rdx, [rel data]
fld tword [rdx + 8 * 0]
fscale
hlt
align 8
data:
dt 64.0
dq 0
data2:
dt -6.5
dq 0