Compare commits

...
92 Commits
Author SHA1 Message Date
Ryan Houdek b3bc1e23cc Docs: Update for release FEX-2306 2023-06-08 16:38:52 -07:00
Mai 1a9b6a89f4 Merge pull request #2706 from Sonicadvance1/remove_emulated_cores
FEXConfig: Removes Emulated CPU cores option
2023-06-08 07:39:15 -04:00
Ryan Houdek 21bf35d211 FEXConfig: Removes Emulated CPU cores option
This is just confusing end users these days and no longer matters as a
debug option.

Remove from the GUI initially, maybe afterwards we will even remove
setting this at all and always auto-detect.
2023-06-07 17:55:05 -07:00
Ryan Houdek 9473025b18 Merge pull request #2704 from Sonicadvance1/fexrootfsfetcher_arch
FEXRootFSFetcher: Support rolling release distros
2023-06-07 16:34:09 -07:00
Mai 02f15f4099 Merge pull request #2705 from Sonicadvance1/update_installfex_script
InstallFEX: Updates helper install script for Ubuntu 23.04
2023-06-07 16:34:53 -04:00
Ryan Houdek e007789ced InstallFEX: Updates helper install script for Ubuntu 23.04
Also updates the link in the source to the new json file.
2023-06-07 12:58:11 -07:00
Ryan Houdek a2b165043c FEXRootFSFetcher: Support rolling release distros
This basically just means that we detect ArchLinux and set a flag that
it is a rolling release, skipping doing the version check for an "exact"
match in that instance.
2023-06-07 12:55:38 -07:00
Ryan Houdek 5b5808218b Merge pull request #2703 from Sonicadvance1/minor_of_opt
OpcodeDispatcher: Optimize ADC/ADD OF flag calculation
2023-06-07 12:54:55 -07:00
Ryan Houdek 41ec987f3e OpcodeDispatcher: Optimize ADC/ADD OF flag calculation
`eor <reg>, <reg>, #-1` can't be encoded as an instruction. Instead use
mvn which does the same thing.

Removes a single instruction from each OF calculation for ADC and ADD.

Also no reason to use a switch statement for the source size, just use
_Bfe and calculate the offset based on operation size.

SBB caught in the crossfire to ensure it also isn't using a switch
statement.
2023-06-07 12:40:51 -07:00
Mai 0f4a5edf4f Merge pull request #2702 from Sonicadvance1/fix_ssa_dec
IRDumper: Fixes ssa number in arguments.
2023-06-07 14:18:09 -04:00
Ryan Houdek 03f73531d3 IRDumper: Fixes ssa number in arguments.
This can spuriously end up as a hex number which makes it hard to reason
why DCE wasn't deleting IR operations. Ensure it is always a decimal.
2023-06-07 09:52:04 -07:00
Mai 69181d438c Merge pull request #2701 from Sonicadvance1/optimize_flag_unpacking
OpcodeDispatcher: Optimize EFLAG unpacking
2023-06-06 21:43:31 -04:00
Ryan Houdek a2cbfccb3b OpcodeDispatcher: Optimize EFLAG unpacking
Noticed this was slightly unoptimal. Resulting in a 18% code reduction
in the case of of a simple four instruction test ASM case.
2023-06-06 17:56:25 -07:00
Mai 4e01452a65 Merge pull request #2699 from Sonicadvance1/minor_fcmov_opt
X87: Super minor FCMOV optimization
2023-06-06 20:22:40 -04:00
Mai cc7a56b1a6 Merge pull request #2689 from Sonicadvance1/fix_bmi
CPUID: Only enable BMI1 and BMI2 if AVX is supported
2023-06-06 20:21:57 -04:00
Ryan Houdek 0b0dd3891e X87: Super minor FCMOV optimization
This caught my eye as I was skimming, remove one IR op per FCMOV
instruction.

This was just duplicating the generated GPR mask across the FPR.
2023-06-04 06:39:35 -07:00
Ryan Houdek 8bc33e95c1 Merge pull request #2493 from Sonicadvance1/deferred_signals_partial
Implement support for deferred asynchronous signals
2023-06-02 22:07:18 -07:00
Ryan Houdek 96a0364a86 Review comments 2023-06-02 21:53:52 -07:00
Ryan Houdek c0a783997d Convert remaining memory tracking to deferred signals 2023-06-01 11:35:22 -07:00
Ryan Houdek f78537109d Core: Convert mtrack code invalidation over to deferred signals 2023-06-01 11:35:22 -07:00
Ryan Houdek 0c156ed6f9 Context: Switch over to deferred signals 2023-06-01 11:28:04 -07:00
Ryan Houdek 920913cf80 Syscalls: Always install SIGSEGV handler for deferred handler 2023-06-01 11:28:04 -07:00
Ryan Houdek 8840b2154c Allocator: Allow more optimal deferred signals path 2023-06-01 11:28:04 -07:00
Ryan Houdek e02be8073e FEXCore: Support deferred signal mutex
This is part of FEXCore since it pulls in InternalThreadData, but is
related to the FHU signal mutex class.

Necessary to allow deferring signals in C++ code rather than right in
the JIT.
2023-06-01 11:28:04 -07:00
Ryan Houdek f75d3550b4 Jit64: Used deferred signals in dispatcher 2023-06-01 11:28:04 -07:00
Ryan Houdek 802c588695 Arm64: Use deferred signals in dispatcher 2023-06-01 11:28:04 -07:00
Ryan Houdek fd962f40d7 SignalDelegator: Support deferring signals 2023-06-01 11:28:04 -07:00
Ryan Houdek a9b660af69 CoreState: Add new members to track deferred signal capability 2023-06-01 11:28:04 -07:00
Ryan Houdek fd5c36ba9c Docs: Adds a document explaining how FEX's deferred signals works.
This has design considerations as to why choices were made.
2023-06-01 11:28:04 -07:00
Ryan Houdek 5be798e9e6 Merge pull request #2693 from Sonicadvance1/remove_debug
Context: Remove debug namespace
2023-06-01 11:26:05 -07:00
Ryan Houdek 09997cff9c Merge pull request #2692 from Sonicadvance1/remove_debugger
Tools: Removes visual debugger
2023-06-01 11:25:55 -07:00
Ryan Houdek c9d1f0d75a Merge pull request #2687 from Sonicadvance1/telemetry_save_crash
Telemetry: Save on signal terminate
2023-05-30 10:26:03 -07:00
Ryan Houdek 95b7592241 Merge pull request #2690 from Sonicadvance1/vfork_wait
Linux: Make vfork act more similar to how it should.
2023-05-30 10:25:53 -07:00
Ryan Houdek 1dc4f8c429 Context: Remove debug namespace
Unused and broken
2023-05-30 09:00:57 -07:00
Ryan Houdek 1d7fcdb54a Tools: Removes visual debugger
Unused and broken
2023-05-30 08:53:48 -07:00
Ryan Houdek 45d3b83143 Telemetry: Save on signal terminate
When a signal handler is not installed and is a terminal failure, make
sure to save telemetry before faulting.

We know when an application is going down in this case so we can make
sure to have the telemetry data saved.

Adds a telemetry signal mask data point as well to know which signal
took it down.
2023-05-30 08:49:33 -07:00
Ryan Houdek d97fa9af14 Linux: Make vfork act more similar to how it should.
Noticed this while debugging Proton Experimental hanging and thought
this could be related. Didn't solve that issue but this should be merged
anyway.

vfork doesn't fork the host's process space in to the child process.
Saving Copy-On-Write overhead problems. It also puts the parent process
to sleep until the fork terminates or executes.

This is a major issue under FEX where we can't emulate vfork correctly
because we need to do other work before this process terminates or
executes a new process. We have been treating `vfork` as a `fork` this
entire time.

This can likely cause problems for applications that actually use vfork
to wait for a process to complete. So let's actually emulate that
feature by using a pipe with poll to determine when that FD gets
removed.

FEX can't use waitpid to wait for this process to terminate since we
would affect the guest also wanted to use a waitpid.
2023-05-30 08:44:36 -07:00
Ryan Houdek c9101d3f68 CPUID: Only enable BMI1 and BMI2 if AVX is supported
These two extensions rely on AVX being supported to be used. Primarily
because they are VEX encoded.

GTA5 is using these flags to determine if it should enable its AVX
support.
2023-05-26 20:48:36 -07:00
Mai 52f64a0c7b Merge pull request #2685 from Sonicadvance1/remove_ci_warnings
github: Updates some actions to v3
2023-05-22 22:31:40 -04:00
Mai 737f917838 Merge pull request #2686 from Sonicadvance1/xgetbv
FEXCore: Implements support for xgetbv
2023-05-22 22:31:21 -04:00
Ryan Houdek a6c6248bcb ArmEmitter: Fixes bug in SpillStaticRegs
Some code in FEX's Arm64 emitter was making an assumption that once
SpillStaticRegs was called that it was safe to still use the SRA
register state.
This wasn't actually true since FEX was using one SRA register to
optimize FPR stores. Assuming that the SRA registers were safe to use
since they were just saved and no longer necessary.

Correct this assumption hell by forcing users of the function to provide
the temporary register directly. In all cases the users have a temporary
available that it can use.

Probably fixes some very weird edge case bugs.
2023-05-22 16:48:07 -07:00
Ryan Houdek 5646428640 FEXCore: Implements support for xgetbv
This returns the `XFEATURE_ENABLED_MASK` register which reports what
features are enabled on the CPU.
This behaves similarly to CPUID where it uses an index register in ecx.

This is a prerequisite to enabling XSAVE/XRSTOR and AVX since
applications will expect this to exist.

xsetbv is a privileged instruction and doesn't need to be implemented.
2023-05-22 16:48:07 -07:00
Ryan Houdek 0c8df2beaf github: Updates some actions to v3
Removes some annotation warnings that have been showing up on the
actions results page.
v2 is deprecated so going to v3 is necessary. Apparently this upgrades
from Node.js 12 to 16.
2023-05-22 10:39:13 -07:00
Mai de0f3984e9 Merge pull request #2680 from Sonicadvance1/optimize_getdents
Syscalls: Optimize getdents{64,}
2023-05-22 11:46:23 -04:00
Mai ada226bbb4 Merge pull request #2683 from Sonicadvance1/uprev_kernel
FEXLoader: Allow simulated kernel version up to 6.2
2023-05-22 11:45:12 -04:00
Ryan Houdek 4bc5a09e62 FEXLoader: Allow simulated kernel version up to 6.2
Investigation in #2589 shows we can push it to this point.
6.3 adds a new prctl that FEX can't enable yet.
2023-05-21 09:51:34 -07:00
Ryan Houdek 5704b5f23f Syscalls: Optimize getdents{64,}
I originally wrote this emulation prior to me fully understanding how
the syscall works. So there are two optimizations here.

1) No need to consume the incoming buffer at all.
   - Originally I thought the incoming dirent structures were used to
     calculate offset.
   - This is not the case, the FD's file position is used instead.
   - This means we can remove the incoming buffer consuming overhead
     entirely.
2) No need to allocate a temporary buffer at all.
   - With getdents and getdents64 we are guaranteed to be dealing with
     structures that are the same size or smaller than the host
     structure.
   - This lets us encode the real host dirents in to the provided
     buffer.
   - After the `getdents64` host syscall, we then iterate forward
     through the list, modifying as we go.
   - Need to make sure to shift the elements of the structure in order.
   - Need to make sure to use memmove on the `d_name` member since the
     movement region can overlap.

These two optimizations significantly reduce the amount of time spent in
getdents, which has a noticeable impact on load times.

Side-tangent: I noticed a fun quirk of how NFS operates with getdents.
If the FSCache hasn't populated the metadata for that folder, then it
will early return with "some" data, not fully maxing out the buffer. The
kernel will start prefetching metadata assuming directory iterating is
happening. The next `getdents` happens and it should return a larger
number of elements.

Very neat.
2023-05-18 21:56:24 -07:00
Ryan Houdek 6017a9135a Merge pull request #2679 from Sonicadvance1/mostly_revert_2672
Thunks: Mostly reverts #2672
2023-05-18 16:11:59 -07:00
Ryan Houdek 6ef6d9c391 Thunks: Mostly reverts #2672
I forgot that x11 was part of the custom ABI of thunks. #2672 had broken
thunks on ARM64. I thought I had tested a game with them enabled but
apparently I tested the wrong game.

Not a full revert since we can still ldr with a literal, but we also
still need to adr x11 and nop pad. At least removes the data dependency
on x11 from the ldr.
2023-05-18 15:50:55 -07:00
Ryan Houdek 0ad6f98a8c Merge pull request #2666 from Sonicadvance1/wine_testharnessrunner
Wine TestHarnessRunner support
2023-05-18 12:58:35 -07:00
Ryan Houdek 1354f92cc5 Review comments 2023-05-17 21:09:31 -07:00
Ryan Houdek 8b90caad95 unittests: Adds a Linux HostFeatures flag
Disables two tests that don't work under Wine
2023-05-17 21:09:31 -07:00
Ryan Houdek 3a4a965347 TestHarnessRunner: Support exiting on HLT
Currently WINE's longjump doesn't work, so instead set a flag that if
HLT is attempted, just exit the JIT.

This will get our unittests executing at least.
2023-05-17 21:09:31 -07:00
Ryan Houdek 45cdab2ac3 HostFeatures: Use ID registers under Wine
InferFromOS doesn't work under WINE.
InferFromIDRegisters doesn't work under Windows but it will under Wine.

Since we don't support Windows, just use InferFromIDRegisters.
2023-05-17 21:07:40 -07:00
Ryan Houdek b89dc56ae1 unittests: Update test so it can work on wine.
We don't necessarily care where this memory is, just that it can be
allocated. Move it to a memory location that works on both Linux and
Wine.
2023-05-17 21:07:40 -07:00
Ryan Houdek d675b4af6f External: Update vixl 2023-05-17 21:07:40 -07:00
Ryan Houdek d75fb38344 TestHarnessRunner: Get running on Win32 2023-05-17 21:07:40 -07:00
Ryan Houdek 4cb385a27b unittests: Build ASM tests on win32 2023-05-17 21:07:37 -07:00
Ryan Houdek 9a4fdd8059 ArchHelpers: Adds missing WinContext stub 2023-05-17 21:05:55 -07:00
Ryan Houdek 363411f0c7 ArchHelpers: Adds missing stub function 2023-05-17 21:05:55 -07:00
Ryan Houdek 5bc418407c FEXCore: Disable emitter unit tests on win32 2023-05-17 21:05:55 -07:00
Ryan Houdek 46e2dc7498 Common: Disable some Linux specific files on win32 2023-05-17 21:05:55 -07:00
Ryan Houdek 4a54197868 TestHarnessRunner: Use VirtualAlloc for mapping regions.
Needs to be alligned to allocation size. Which is a page on Linux, or
64k on Windows.

In order to map at `0x1'0000` on Wine, we need to use a special case DOS
area allocation path.
2023-05-17 21:05:55 -07:00
Ryan Houdek 3f214dd244 HarnessHelpers: Use FEXCore helper for file loading. 2023-05-17 21:05:55 -07:00
Ryan Houdek 61ca651fe1 FEXCore: Don't initialize ThunkHandler on Win32
Adds a couple pointer checks to ensure it won't crash.

Doesn't work and will cause assertions.
2023-05-17 21:05:55 -07:00
Ryan Houdek cd0a340d29 unittests/ASM: Ensure wine harness runner works
Needs to execute the correct runner through wine, and needs to reserve
the low DOS region so something doesn't get loaded there.
2023-05-17 21:05:55 -07:00
Mai 77e8be1215 Merge pull request #2671 from Sonicadvance1/wine_syscalls
FEXCore: Support Wine syscalls
2023-05-18 00:04:25 -04:00
Ryan Houdek 182010ca97 Merge pull request #2678 from lioncash/strings
OpcodeDispatcher: Handle PCMPESTRM/VPCMPESTRM
2023-05-16 21:52:53 -07:00
Lioncache f7c663240e OpcodeDispatcher: Handle PCMPESTRM/VPCMPESTRM
...and with that all of the SSE4.2 string instructions are implemented now
2023-05-17 00:21:55 -04:00
Ryan Houdek e9244680aa Merge pull request #2677 from lioncash/masked
OpcodeDispatcher: Handle PCMPISTRM/VPCMPISTRM
2023-05-16 20:25:42 -07:00
Lioncache 82b4aef30d OpcodeDispatcher: Handle PCMPISTRM/VPCMPISTRM 2023-05-16 22:59:54 -04:00
Lioncache 22919a5b65 OpcodeDispatcher: Add mask variant handling to PCMPXSTXOpImpl()
Will be used to handle PCMPESTRM/PCMPISTRM instruction variants.
2023-05-16 22:59:52 -04:00
Mai 00dc373bb9 Merge pull request #2676 from Sonicadvance1/fix_at_execfn
ELFCodeLoader: Fixes missing AT_EXECFN
2023-05-15 17:49:54 -04:00
Ryan Houdek e593237670 ELFCodeLoader: Fixes missing AT_EXECFN
New versions of CEF rely on this existing. It will get this value and
run strdup on it, even if it is nullptr.

Fixes a steamwebhelper process constantly crashing with the Steam Beta
client.

Only missing auxv values now
- AT_PAGESZ
- AT_EXECFD (for execveat?)
- AT_PHDR
- All the random cache information values.
2023-05-14 03:10:11 -07:00
Mai 0a4bf10ba5 Merge pull request #2674 from Sonicadvance1/fix_shm_leaks
unittests: Adds step to remove stale SHM regions.
2023-05-12 23:22:23 -04:00
Ryan Houdek f47caf48c6 Merge pull request #2669 from Sonicadvance1/aotir_mutex
AOTIR: Stop passing a mutex around. It's already guarded
2023-05-12 18:56:55 -07:00
Ryan Houdek 5674d3a871 Merge pull request #2667 from Sonicadvance1/fextl_file
FEXCore: Convert Core and Telemetry over to fextl::file::File
2023-05-12 18:56:45 -07:00
Ryan Houdek fde64aedf7 unittests: Adds step to remove stale SHM regions.
Some of the unit tests we run will leak shm regions. Presumably this is
because they never called `shmctl(IPC_RMID)` so the ID is laked forever.

This can be seen by querying `/proc/sysvipc/shm` to see a list of old
shm regions that eventually hit the maximum capacity of 4096 shm ids.

Once CI is done running, run a utility application that all it does is
check for SHM IDs that have zero attachments (thus unused), it was
created by the UID of the runner, and it is older than ten minutes. At
which point it will erase it.

This will fix spurious failures in our CI caused by running out of SHM
IDs, previously I had a cron job setup to restart the CI runners every
hour or so which caused its own spurious failure problems.

FINALLY this bug was triaged which has been annoying us for...years?
2023-05-12 18:54:01 -07:00
Mai e03b859c20 Merge pull request #2673 from Sonicadvance1/remove_warnings_13
OpcodeDispatcher: Removes a warning that cropped up.
2023-05-12 21:49:43 -04:00
Mai ce4e991a6e Merge pull request #2672 from Sonicadvance1/optimize_trampoline
Thunks: Optimize ARM64 trampoline
2023-05-12 20:50:22 -04:00
Ryan Houdek 7d822ba1c8 OpcodeDispatcher: Removes a warning that cropped up. 2023-05-12 17:34:20 -07:00
Ryan Houdek f90dcd2eb1 FEXCore: Convert Core and Telemetry over to FEXCore::File::File
This way telemetry and IR dumping can work under Wine.
2023-05-12 17:32:48 -07:00
Ryan Houdek adbdd33ece fextl/fmt: Adds write handler for FEXCore::File::File 2023-05-12 17:32:48 -07:00
Ryan Houdek 06250d806d FEXCore/Utils: Adds File type
OS agnostic file class since we can't use std::FILE
2023-05-12 17:32:48 -07:00
Ryan Houdek 613ed559e7 Thunks: Optimize ARM64 trampoline
No need to use adr for getting the PC relative literal, we can use LDR
(literal) to load the PC relative address directly.

Reduces trampline instructions from 3 to 2, also reduces trampoline size
from 24-bytes to 16-bytes.
2023-05-12 17:28:36 -07:00
Ryan Houdek 8ac3841946 FEXCore: Support Wine syscalls
Wine syscalls need to end the code block at the point of the syscall.
This is because syscalls may update RIP which means the JIT loop needs
to immediately restart.

Additionally since they can update CPU state, make wine syscalls not
return a result and instead refill the register state from the CPU
state. This will mean the syscall handler will need to update their
result register (RAX?) before returning.
2023-05-12 16:42:26 -07:00
Ryan Houdek 458259bf47 FEXCore: Move EnumOperators to FEXCore
fextl needs this and can't depend on FHU
2023-05-12 15:23:00 -07:00
Ryan Houdek dc65a5ef8c Merge pull request #2668 from Sonicadvance1/fix_sra_disabled
ARM64: Fixes SRA disabled codepath
2023-05-12 15:21:52 -07:00
Ryan Houdek 2fc529d5b7 AOTIR: Stop passing a mutex around. It's already guarded 2023-05-11 03:56:33 -07:00
Ryan Houdek ea489567da ARM64: Fixes SRA disabled codepath
Disabling SRA has been broken a quite a while. Disabling this was
instrumental in figuring out the VC redistributable crash.

Ensure it works by reintroducing non-SRA load/store register handlers,
and by supporting runtime selectable dispatch pointers for the JIT.

Side-bonus, moves the {LOAD,STORE}MEMTSO ops over to this dispatch as
well to make it consistent and probably slightly quicker.
2023-05-11 03:25:19 -07:00
Ryan Houdek ed69eb9f6f Merge pull request #2665 from Sonicadvance1/prctl_tso
FEXCore: Adds support for hardware x86-TSO prctl
2023-05-09 04:10:34 -07:00
Ryan Houdek 6eae064511 FEXCore: Adds support for hardware x86-TSO prctl
From https://github.com/AsahiLinux/linux/commits/bits/220-tso

This fails gracefully in the case the upstream kernel doesn't support
this feature, so can go in early.

This feature allows FEX to use hardware's TSO emulation capability to
reduce emulation overhead from our atomic/lrcpc implementation.
In the case that the TSO emulation feature is enabled in FEX, we will
check if the hardware supports this feature and then enable it.

If the hardware feature is supported it will then use regular memory
accesses with the expectation that these are x86-TSO in strength.

The only hardware that anyone cares about that supports this is Apple's
M class SoCs. Theoretically NVIDIA Denver/Carmel supports sequentially
consistent, which isn't quite the same thing. I haven't cared to check
if multithreaded SC has as strong of guarantees. But also since
Carmel/Denver hardware is fairly rare, it's hard to care about for our
use case.
2023-05-08 20:12:03 -07:00
138 changed files with 4901 additions and 3135 deletions

No files matched your search

+8 -2
View File
@@ -25,7 +25,7 @@ jobs:
fail-fast: false
steps:
- uses: actions/checkout@v2
- uses: actions/checkout@v3
- name: Set runner label
run: echo "runner_label=${{ matrix.arch[1] }}" >> $GITHUB_ENV
@@ -241,13 +241,19 @@ jobs:
# ASM tests get quite close to 10MB
run: truncate --size=<20M ${{runner.workspace}}/build/Testing/Temporary/LastTest_*.log || true
- name: Remove old SHM regions
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: cmake --build . --config $BUILD_TYPE --target remove_old_shm_regions
- name: Set runner name
if: ${{ always() }}
run: echo "runner_name=$(hostname)" >> $GITHUB_ENV
- name: Upload results
if: ${{ always() }}
uses: 'actions/upload-artifact@v2'
uses: 'actions/upload-artifact@v3'
with:
name: Results-${{ env.runner_name }}
path: ${{runner.workspace}}/build/Testing/Temporary/LastTest_*.log
+8 -2
View File
@@ -33,7 +33,7 @@ jobs:
fail-fast: false
steps:
- uses: actions/checkout@v2
- uses: actions/checkout@v3
- name: Set runner label
run: echo "runner_label=${{ matrix.arch[1] }}" >> $GITHUB_ENV
@@ -176,13 +176,19 @@ jobs:
# ASM tests get quite close to 10MB
run: truncate --size=<20M ${{runner.workspace}}/build/Testing/Temporary/LastTest_*.log || true
- name: Remove old SHM regions
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: cmake --build . --config $BUILD_TYPE --target remove_old_shm_regions
- name: Set runner name
if: ${{ always() }}
run: echo "runner_name=$(hostname)" >> $GITHUB_ENV
- name: Upload results
if: ${{ always() }}
uses: 'actions/upload-artifact@v2'
uses: 'actions/upload-artifact@v3'
with:
name: Results-${{ env.runner_name }}
path: ${{runner.workspace}}/build/Testing/Temporary/LastTest_*.log
+2 -2
View File
@@ -25,7 +25,7 @@ jobs:
fail-fast: false
steps:
- uses: actions/checkout@v2
- uses: actions/checkout@v3
- name: Set runner label
run: echo "runner_label=${{ matrix.arch[1] }}" >> $GITHUB_ENV
@@ -111,7 +111,7 @@ jobs:
- name: Upload results
if: ${{ always() }}
uses: 'actions/upload-artifact@v2'
uses: 'actions/upload-artifact@v3'
with:
name: Results-${{ env.runner_name }}
path: ${{runner.workspace}}/build/Testing/Temporary/LastTest_*.log
-1
View File
@@ -18,7 +18,6 @@ option(ENABLE_ASAN "Enables Clang ASAN" FALSE)
option(ENABLE_TSAN "Enables Clang TSAN" FALSE)
option(ENABLE_ASSERTIONS "Enables assertions in build" FALSE)
option(ENABLE_GDB_SYMBOLS "Enables GDBSymbols integration support" ${HAVE_GDB_JIT_READER_H})
option(ENABLE_VISUAL_DEBUGGER "Enables the visual debugger for compiling" FALSE)
option(ENABLE_STRICT_WERROR "Enables stricter -Werror for CI" FALSE)
option(ENABLE_WERROR "Enables -Werror" FALSE)
option(ENABLE_JEMALLOC "Enables jemalloc allocator" TRUE)
+4 -31
View File
@@ -87,38 +87,11 @@ namespace FEXCore::Context {
return CPUID.RunFunction(Function, Leaf);
}
FEXCore::CPUID::XCRResults FEXCore::Context::ContextImpl::RunXCRFunction(uint32_t Function) {
return CPUID.RunXCRFunction(Function);
}
FEXCore::CPUID::FunctionResults FEXCore::Context::ContextImpl::RunCPUIDFunctionName(uint32_t Function, uint32_t Leaf, uint32_t CPU) {
return CPUID.RunFunctionName(Function, Leaf, CPU);
}
namespace Debug {
//void CompileRIP(FEXCore::Context::Context *CTX, uint64_t RIP) {
// CTX->CompileRIP(CTX->ParentThread, RIP);
//}
//uint64_t GetThreadCount(FEXCore::Context::Context *CTX) {
// return CTX->GetThreadCount();
//}
//FEXCore::Core::RuntimeStats *GetRuntimeStatsForThread(FEXCore::Context::Context *CTX, uint64_t Thread) {
// return CTX->GetRuntimeStatsForThread(Thread);
//}
//bool GetDebugDataForRIP(FEXCore::Context::Context *CTX, uint64_t RIP, FEXCore::Core::DebugData *Data) {
// return CTX->GetDebugDataForRIP(RIP, Data);
//}
//bool FindHostCodeForRIP(FEXCore::Context::Context *CTX, uint64_t RIP, uint8_t **Code) {
// return CTX->FindHostCodeForRIP(RIP, Code);
//}
// XXX:
// bool FindIRForRIP(FEXCore::Context::Context *CTX, uint64_t RIP, FEXCore::IR::IntrusiveIRList **ir) {
// return CTX->FindIRForRIP(RIP, ir);
// }
// void SetIRForRIP(FEXCore::Context::Context *CTX, uint64_t RIP, FEXCore::IR::IntrusiveIRList *const ir) {
// CTX->SetIRForRIP(RIP, ir);
// }
}
}
+32 -12
View File
@@ -1,7 +1,6 @@
#pragma once
#include "Common/JitSymbols.h"
#include "FEXHeaderUtils/ScopedSignalMask.h"
#include "Interface/Core/CPUID.h"
#include "Interface/Core/X86HelperGen.h"
#include "Interface/Core/ObjectCache/ObjectCacheService.h"
@@ -14,6 +13,7 @@
#include <FEXCore/Core/SignalDelegator.h>
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXCore/Utils/CompilerDefs.h>
#include <FEXCore/Utils/DeferredSignalMutex.h>
#include <FEXCore/Utils/Event.h>
#include <FEXCore/fextl/memory.h>
#include <FEXCore/fextl/set.h>
@@ -159,6 +159,7 @@ namespace FEXCore::Context {
void SetSyscallHandler(FEXCore::HLE::SyscallHandler *Handler) override;
FEXCore::CPUID::FunctionResults RunCPUIDFunction(uint32_t Function, uint32_t Leaf) override;
FEXCore::CPUID::XCRResults RunXCRFunction(uint32_t Function) override;
FEXCore::CPUID::FunctionResults RunCPUIDFunctionName(uint32_t Function, uint32_t Leaf, uint32_t CPU) override;
FEXCore::IR::AOTIRCacheEntry *LoadAOTIRCacheEntry(const fextl::string& Name) override;
@@ -180,8 +181,8 @@ namespace FEXCore::Context {
void WriteFilesWithCode(std::function<void(const fextl::string& fileid, const fextl::string& filename)> Writer) override {
IRCaptureCache.WriteFilesWithCode(Writer);
}
void InvalidateGuestCodeRange(uint64_t Start, uint64_t Length) override;
void InvalidateGuestCodeRange(uint64_t Start, uint64_t Length, std::function<void(uint64_t start, uint64_t Length)> callback) override;
void InvalidateGuestCodeRange(FEXCore::Core::InternalThreadState *Thread, uint64_t Start, uint64_t Length) override;
void InvalidateGuestCodeRange(FEXCore::Core::InternalThreadState *Thread, uint64_t Start, uint64_t Length, std::function<void(uint64_t start, uint64_t Length)> callback) override;
void MarkMemoryShared() override;
void ConfigureAOTGen(FEXCore::Core::InternalThreadState *Thread, fextl::set<uint64_t> *ExternalBranches, uint64_t SectionMaxAddress) override;
@@ -289,7 +290,8 @@ namespace FEXCore::Context {
template<auto Fn>
static uint64_t ThreadExitFunctionLink(FEXCore::Core::CpuStateFrame *Frame, uint64_t *record) {
FHU::ScopedSignalMaskWithSharedLock lk(static_cast<ContextImpl*>(Frame->Thread->CTX)->CodeInvalidationMutex);
auto Thread = Frame->Thread;
ScopedDeferredSignalWithSharedLock lk(static_cast<ContextImpl*>(Thread->CTX)->CodeInvalidationMutex, Thread);
return Fn(Frame, record);
}
@@ -301,19 +303,13 @@ namespace FEXCore::Context {
LogMan::Throw::AFmt(Thread->ThreadManager.GetTID() == FHU::Syscalls::gettid(), "Must be called from owning thread {}, not {}", Thread->ThreadManager.GetTID(), FHU::Syscalls::gettid());
FHU::ScopedSignalMaskWithUniqueLock lk(static_cast<ContextImpl*>(Thread->CTX)->CodeInvalidationMutex);
ScopedDeferredSignalWithUniqueLock lk(static_cast<ContextImpl*>(Thread->CTX)->CodeInvalidationMutex, Thread);
ThreadRemoveCodeEntry(Thread, GuestRIP);
}
void RemoveCustomIREntrypoint(uintptr_t Entrypoint);
// Debugger interface
uint64_t GetThreadCount() const;
FEXCore::Core::RuntimeStats *GetRuntimeStatsForThread(uint64_t Thread);
bool GetDebugDataForRIP(uint64_t RIP, FEXCore::Core::DebugData *Data);
bool FindHostCodeForRIP(uint64_t RIP, uint8_t **Code);
struct GenerateIRResult {
FEXCore::IR::IRListView* IRList;
FEXCore::IR::RegisterAllocationData::UniquePtr RAData;
@@ -372,11 +368,32 @@ namespace FEXCore::Context {
FEXCore::Utils::PooledAllocatorVirtual OpDispatcherAllocator;
FEXCore::Utils::PooledAllocatorVirtual FrontendAllocator;
bool IsTSOEnabled() { return (IsMemoryShared || !Config.TSOAutoMigration) && Config.TSOEnabled; }
// If Atomic-based TSO emulation is enabled or not.
bool IsAtomicTSOEnabled() const { return AtomicTSOEmulationEnabled; }
void SetHardwareTSOSupport(bool HardwareTSOSupported) override {
SupportsHardwareTSO = HardwareTSOSupported;
UpdateAtomicTSOEmulationConfig();
}
void EnableExitOnHLT() override { ExitOnHLT = true; }
bool ExitOnHLTEnabled() const { return ExitOnHLT; }
protected:
void ClearCodeCache(FEXCore::Core::InternalThreadState *Thread);
void UpdateAtomicTSOEmulationConfig() {
if (SupportsHardwareTSO) {
// If the hardware supports TSO then we don't need to emulate it through atomics.
AtomicTSOEmulationEnabled = false;
}
else {
// Atomic TSO emulation only enabled if the config option is enabled.
AtomicTSOEmulationEnabled = (IsMemoryShared || !Config.TSOAutoMigration) && Config.TSOEnabled;
}
}
private:
/**
* @brief Does some final thread initialization
@@ -411,6 +428,9 @@ namespace FEXCore::Context {
bool StartPaused = false;
bool IsMemoryShared = false;
bool SupportsHardwareTSO = false;
bool AtomicTSOEmulationEnabled = true;
bool ExitOnHLT = false;
FEX_CONFIG_OPT(AppFilename, APP_FILENAME);
std::shared_mutex CustomIRMutex;
@@ -219,7 +219,7 @@ void Arm64Emitter::PopCalleeSavedRegisters() {
}
}
void Arm64Emitter::SpillStaticRegs(bool FPRs, uint32_t GPRSpillMask, uint32_t FPRSpillMask) {
void Arm64Emitter::SpillStaticRegs(FEXCore::ARMEmitter::Register TmpReg, bool FPRs, uint32_t GPRSpillMask, uint32_t FPRSpillMask) {
if (!StaticRegisterAllocation()) {
return;
}
@@ -252,8 +252,6 @@ void Arm64Emitter::SpillStaticRegs(bool FPRs, uint32_t GPRSpillMask, uint32_t FP
} else {
if (GPRSpillMask && FPRSpillMask == ~0U) {
// Optimize the common case where we can spill four registers per instruction
auto TmpReg = SRA64[FindFirstSetBit(GPRSpillMask)];
// Load the sse offset in to the temporary register
add(ARMEmitter::Size::i64Bit, TmpReg, STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[0][0]));
for (size_t i = 0; i < ConfiguredSRAFPRs; i += 4) {
@@ -183,7 +183,7 @@ protected:
// NOTE: These functions WILL clobber the register TMP4 if AVX support is enabled
// and FPRs are being spilled or filled. If only GPRs are spilled/filled, then
// TMP4 is left alone.
void SpillStaticRegs(bool FPRs = true, uint32_t GPRSpillMask = ~0U, uint32_t FPRSpillMask = ~0U);
void SpillStaticRegs(FEXCore::ARMEmitter::Register TmpReg, bool FPRs = true, uint32_t GPRSpillMask = ~0U, uint32_t FPRSpillMask = ~0U);
void FillStaticRegs(bool FPRs = true, uint32_t GPRFillMask = ~0U, uint32_t FPRFillMask = ~0U);
// Register 0-18 + 29 + 30 are caller saved
+22 -11
View File
@@ -76,10 +76,6 @@ static uint32_t GetCPUID() {
return CPU;
}
// TODO: Replace usages with CTX->HostFeatures.EnableAVX
// when AVX implementations are further along.
constexpr uint32_t SUPPORTS_AVX = 0;
#ifdef CPUID_AMD
constexpr uint32_t FAMILY_IDENTIFIER =
0 | // Stepping
@@ -441,7 +437,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_01h(uint32_t Leaf) {
(CTX->HostFeatures.SupportsAES << 25) | // AES
(0 << 26) | // XSAVE
(0 << 27) | // OSXSAVE
(SUPPORTS_AVX << 28) | // AVX
(SupportsAVX() << 28) | // AVX
(0 << 29) | // F16C
(CTX->HostFeatures.SupportsRAND << 30) | // RDRAND
(Hypervisor << 31);
@@ -630,12 +626,12 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_07h(uint32_t Leaf) {
(1 << 0) | // FS/GS support
(0 << 1) | // TSC adjust MSR
(0 << 2) | // SGX
(1 << 3) | // BMI1
(SupportsAVX() << 3) | // BMI1
(0 << 4) | // Intel Hardware Lock Elison
(0 << 5) | // AVX2 support
(1 << 6) | // FPU data pointer updated only on exception
(1 << 7) | // SMEP support
(1 << 8) | // BMI2
(SupportsAVX() << 8) | // BMI2
(0 << 9) | // Enhanced REP MOVSB/STOSB
(1 << 10) | // INVPCID for system software control of process-context
(0 << 11) | // Restricted transactional memory
@@ -736,13 +732,13 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_0Dh(uint32_t Leaf) {
// Leaf 0
FEXCore::CPUID::FunctionResults Res{};
uint32_t XFeatureSupportedSizeMax = SUPPORTS_AVX ? 0x0000'0340 : 0x0000'0240; // XFeatureEnabledSizeMax: Legacy Header + FPU/SSE + AVX
uint32_t XFeatureSupportedSizeMax = SupportsAVX() ? 0x0000'0340 : 0x0000'0240; // XFeatureEnabledSizeMax: Legacy Header + FPU/SSE + AVX
if (Leaf == 0) {
// XFeatureSupportedMask[31:0]
Res.eax =
(1 << 0) | // X87 support
(1 << 1) | // 128-bit SSE support
(SUPPORTS_AVX << 2) | // 256-bit AVX support
(SupportsAVX() << 2) | // 256-bit AVX support
(0b00 << 3) | // MPX State
(0b000 << 5) | // AVX-512 state
(0 << 8) | // "Used for IA32_XSS" ... Used for what?
@@ -776,8 +772,8 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_0Dh(uint32_t Leaf) {
Res.edx = 0;
}
else if (Leaf == 2) {
Res.eax = SUPPORTS_AVX ? 0x0000'0100 : 0; // YmmSaveStateSize
Res.ebx = SUPPORTS_AVX ? 0x0000'0240 : 0; // YmmSaveStateOffset
Res.eax = SupportsAVX() ? 0x0000'0100 : 0; // YmmSaveStateSize
Res.ebx = SupportsAVX() ? 0x0000'0240 : 0; // YmmSaveStateOffset
// Reserved
Res.ecx = 0;
@@ -1212,11 +1208,26 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_Reserved(uint32_t Leaf) {
return Res;
}
FEXCore::CPUID::XCRResults CPUIDEmu::XCRFunction_0h() {
// This just returns XCR0
FEXCore::CPUID::XCRResults Res{
.eax = static_cast<uint32_t>(XCR0),
.edx = static_cast<uint32_t>(XCR0 >> 32),
};
return Res;
}
void CPUIDEmu::Init(FEXCore::Context::ContextImpl *ctx) {
CTX = ctx;
// Setup some state tracking
SetupHostHybridFlag();
// TODO: Enable once AVX is supported.
if (false && CTX->HostFeatures.SupportsAVX) {
XCR0 |= XCR0_AVX;
}
}
}
+41
View File
@@ -63,13 +63,52 @@ public:
return Function_8000_0004h(Leaf, CPU % PerCPUData.size());
}
FEXCore::CPUID::XCRResults RunXCRFunction(uint32_t Function) {
if (Function >= 1) {
// XCR function 1 is not yet supported.
return {};
}
return XCRFunction_0h();
}
private:
FEXCore::Context::ContextImpl *CTX;
bool Hybrid{};
FEX_CONFIG_OPT(Cores, THREADS);
FEX_CONFIG_OPT(HideHypervisorBit, HIDEHYPERVISORBIT);
// XFEATURE_ENABLED_MASK
// Mask that configures what features are enabled on the CPU.
// Affects XSAVE and XRSTOR when modified.
// Bit layout is as follows.
// [0] - x87 enabled
// [1] - SSE enabled
// [2] - YMM enabled (256-bit SSE)
// [8:3] - Reserved. MBZ.
// [9] - MPK
// [10] - Reserved. MBZ.
// [11] - CET_U
// [12] - CET_S
// [61:13] - Reserved. MBZ.
// [62] - LWP (Lightweight profiling)
// [63] - Reserved for XCR bit vector expansion. MBZ.
// Always enable x87 and SSE by default.
constexpr static uint64_t XCR0_X87 = 1ULL << 0;
constexpr static uint64_t XCR0_SSE = 1ULL << 1;
constexpr static uint64_t XCR0_AVX = 1ULL << 2;
uint64_t XCR0 {
XCR0_X87 |
XCR0_SSE
};
uint32_t SupportsAVX() const {
return (XCR0 & XCR0_AVX) ? 1 : 0;
}
using FunctionHandler = FEXCore::CPUID::FunctionResults (CPUIDEmu::*)(uint32_t Leaf);
struct CPUData {
const char *ProductName{};
#ifdef _M_ARM_64
@@ -109,6 +148,8 @@ private:
FEXCore::CPUID::FunctionResults Function_8000_001Dh(uint32_t Leaf);
FEXCore::CPUID::FunctionResults Function_Reserved(uint32_t Leaf);
FEXCore::CPUID::XCRResults XCRFunction_0h();
void SetupHostHybridFlag();
static constexpr std::array<FunctionHandler, 27> Primary = {
// 0: Highest function parameter and ID
+48 -99
View File
@@ -24,6 +24,7 @@ $end_info$
#include "Interface/IR/Passes/RegisterAllocationPass.h"
#include "Interface/IR/Passes.h"
#include "Interface/IR/PassManager.h"
#include "Utils/Allocator/HostAllocator.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/CodeLoader.h>
@@ -42,6 +43,7 @@ $end_info$
#include <FEXCore/IR/RegisterAllocationData.h>
#include <FEXCore/Utils/Allocator.h>
#include <FEXCore/Utils/Event.h>
#include <FEXCore/Utils/File.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/Threads.h>
#include <FEXCore/Utils/Profiler.h>
@@ -162,6 +164,9 @@ namespace FEXCore::Context {
// Only initialize symbols file if enabled. Ensures we don't pollute /tmp with empty files.
Symbols.InitFile();
}
// Track atomic TSO emulation configuration.
UpdateAtomicTSOEmulationConfig();
}
ContextImpl::~ContextImpl() {
@@ -187,29 +192,6 @@ namespace FEXCore::Context {
}
}
static FEXCore::Core::CPUState CreateDefaultCPUState() {
FEXCore::Core::CPUState NewThreadState{};
// Initialize default CPU state
NewThreadState.rip = ~0ULL;
for (auto& greg : NewThreadState.gregs) {
greg = 0;
}
for (auto& xmm : NewThreadState.xmm.avx.data) {
xmm[0] = 0xDEADBEEFULL;
xmm[1] = 0xBAD0DAD1ULL;
xmm[2] = 0xDEADCAFEULL;
xmm[3] = 0xBAD2CAD3ULL;
}
memset(NewThreadState.flags, 0, Core::CPUState::NUM_EFLAG_BITS);
NewThreadState.flags[1] = 1;
NewThreadState.flags[9] = 1;
NewThreadState.FCW = 0x37F;
NewThreadState.FTW = 0xFFFF;
return NewThreadState;
}
uint64_t ContextImpl::RestoreRIPFromHostPC(FEXCore::Core::InternalThreadState *Thread, uint64_t HostPC) {
const auto Frame = Thread->CurrentFrame;
const uint64_t BlockBegin = Frame->State.InlineJITBlockHeader;
@@ -304,12 +286,13 @@ namespace FEXCore::Context {
StopGdbServer();
}
#ifndef _WIN32
ThunkHandler = FEXCore::ThunkHandler::Create();
#endif
using namespace FEXCore::Core;
FEXCore::Core::CPUState NewThreadState = CreateDefaultCPUState();
FEXCore::Core::InternalThreadState *Thread = CreateThread(&NewThreadState, 0);
FEXCore::Core::InternalThreadState *Thread = CreateThread(nullptr, 0);
// We are the parent thread
ParentThread = Thread;
@@ -552,7 +535,9 @@ namespace FEXCore::Context {
Thread->ThreadManager.TID = FHU::Syscalls::gettid();
Thread->ThreadManager.PID = ::getpid();
SignalDelegation->RegisterTLSState(Thread);
ThunkHandler->RegisterTLSState(Thread);
if (ThunkHandler) {
ThunkHandler->RegisterTLSState(Thread);
}
}
void ContextImpl::RunThread(FEXCore::Core::InternalThreadState *Thread) {
@@ -616,7 +601,9 @@ namespace FEXCore::Context {
FEXCore::Core::InternalThreadState *Thread = new FEXCore::Core::InternalThreadState{};
// Copy over the new thread state to the new object
memcpy(Thread->CurrentFrame, NewThreadState, sizeof(FEXCore::Core::CPUState));
if (NewThreadState) {
memcpy(Thread->CurrentFrame, NewThreadState, sizeof(FEXCore::Core::CPUState));
}
Thread->CurrentFrame->Thread = Thread;
// Set up the thread manager state
@@ -625,6 +612,9 @@ namespace FEXCore::Context {
InitializeCompiler(Thread);
InitializeThreadData(Thread);
Thread->CurrentFrame->State.DeferredSignalRefCount.Store(0);
Thread->CurrentFrame->State.DeferredSignalFaultAddress = reinterpret_cast<Core::NonAtomicRefCounter<uint64_t>*>(FEXCore::Allocator::VirtualAlloc(4096));
// Insert after the Thread object has been fully initialized
{
std::lock_guard lk(ThreadCreationMutex);
@@ -650,6 +640,8 @@ namespace FEXCore::Context {
// To be able to delete a thread from itself, we need to detached the std::thread object
Thread->ExecutionThread->detach();
}
FEXCore::Allocator::VirtualFree(reinterpret_cast<void*>(Thread->CurrentFrame->State.DeferredSignalFaultAddress), 4096);
delete Thread;
}
@@ -712,36 +704,30 @@ namespace FEXCore::Context {
}
static void IRDumper(FEXCore::Core::InternalThreadState *Thread, IR::IREmitter *IREmitter, uint64_t GuestRIP, IR::RegisterAllocationData* RA) {
#ifndef _WIN32
int FD {-1};
bool CloseAfter = false;
FEXCore::File::File FD;
const auto DumpIRStr = static_cast<ContextImpl*>(Thread->CTX)->Config.DumpIR();
// DumpIRStr might be no if not dumping but ShouldDump is set in OpDisp
if (DumpIRStr =="stderr" || DumpIRStr =="no") {
FD = STDERR_FILENO;
FD = FEXCore::File::File::GetStdERR();
}
else if (DumpIRStr =="stdout") {
FD = STDOUT_FILENO;
FD = FEXCore::File::File::GetStdOUT();
}
else {
const auto fileName = fextl::fmt::format("{}/{:x}{}", DumpIRStr, GuestRIP, RA ? "-post.ir" : "-pre.ir");
constexpr int USER_PERMS = S_IRWXU | S_IRWXG | S_IRWXO;
FD = open(fileName.c_str(), O_CREAT | O_WRONLY | O_TRUNC | O_CLOEXEC, USER_PERMS, USER_PERMS);
CloseAfter = true;
FD = FEXCore::File::File(fileName.c_str(),
FEXCore::File::FileModes::WRITE |
FEXCore::File::FileModes::CREATE |
FEXCore::File::FileModes::TRUNCATE);
}
if (FD != -1) {
if (FD.IsValid()) {
fextl::stringstream out;
auto NewIR = IREmitter->ViewIR();
FEXCore::IR::Dump(&out, &NewIR, RA);
fextl::fmt::print(FD, "IR-{} 0x{:x}:\n{}\n@@@@@\n", RA ? "post" : "pre", GuestRIP, out.str());
if (CloseAfter) {
close(FD);
}
}
#endif
};
static void ValidateIR(ContextImpl *ctx, IR::IREmitter *IREmitter) {
@@ -795,7 +781,7 @@ namespace FEXCore::Context {
Thread->FrontendDecoder->DecodeInstructionsAtEntry(GuestCode, GuestRIP, [Thread](uint64_t BlockEntry, uint64_t Start, uint64_t Length) {
if (Thread->LookupCache->AddBlockExecutableRange(BlockEntry, Start, Length)) {
static_cast<ContextImpl*>(Thread->CTX)->SyscallHandler->MarkGuestExecutableRange(Start, Length);
static_cast<ContextImpl*>(Thread->CTX)->SyscallHandler->MarkGuestExecutableRange(Thread, Start, Length);
}
});
@@ -974,7 +960,7 @@ namespace FEXCore::Context {
}
if (SourcecodeResolver && Config.GDBSymbols()) {
auto AOTIRCacheEntry = SyscallHandler->LookupAOTIRCacheEntry(GuestRIP);
auto AOTIRCacheEntry = SyscallHandler->LookupAOTIRCacheEntry(Thread, GuestRIP);
if (AOTIRCacheEntry.Entry && !AOTIRCacheEntry.Entry->ContainsCode) {
AOTIRCacheEntry.Entry->SourcecodeMap =
SourcecodeResolver->GenerateMap(AOTIRCacheEntry.Entry->Filename, AOTIRCacheEntry.Entry->FileId);
@@ -983,7 +969,7 @@ namespace FEXCore::Context {
// AOT IR bookkeeping and cache
{
auto [IRCopy, RACopy, DebugDataCopy, _StartAddr, _Length, _GeneratedIR] = IRCaptureCache.PreGenerateIRFetch(GuestRIP, IRList);
auto [IRCopy, RACopy, DebugDataCopy, _StartAddr, _Length, _GeneratedIR] = IRCaptureCache.PreGenerateIRFetch(Thread, GuestRIP, IRList);
if (_GeneratedIR) {
// Setup pointers to internal structures
IRList = IRCopy;
@@ -1006,9 +992,6 @@ namespace FEXCore::Context {
StartAddr = _StartAddr;
Length = _Length;
// Increment stats
Thread->Stats.BlocksCompiled.fetch_add(1);
// These blocks aren't already in the cache
GeneratedIR = true;
}
@@ -1079,7 +1062,7 @@ namespace FEXCore::Context {
auto FragmentBasePtr = reinterpret_cast<uint8_t *>(CodePtr);
if (DebugData) {
auto GuestRIPLookup = SyscallHandler->LookupAOTIRCacheEntry(GuestRIP);
auto GuestRIPLookup = SyscallHandler->LookupAOTIRCacheEntry(Thread, GuestRIP);
if (DebugData->Subblocks.size()) {
for (auto& Subblock: DebugData->Subblocks) {
@@ -1147,6 +1130,7 @@ namespace FEXCore::Context {
Thread->ExitReason = FEXCore::Context::ExitReason::EXIT_WAITING;
InitializeThreadTLSData(Thread);
Alloc::OSAllocator::RegisterTLSData(Thread);
++IdleWaitRefCount;
@@ -1191,6 +1175,7 @@ namespace FEXCore::Context {
--IdleWaitRefCount;
IdleWaitCV.notify_all();
Alloc::OSAllocator::UninstallTLSData(Thread);
SignalDelegation->UninstallTLSState(Thread);
// If the parent thread is waiting to join, then we can't destroy our thread object
@@ -1221,14 +1206,20 @@ namespace FEXCore::Context {
}
}
void ContextImpl::InvalidateGuestCodeRange(uint64_t Start, uint64_t Length) {
FHU::ScopedSignalMaskWithUniqueLock CodeInvalidationLock(CodeInvalidationMutex);
void ContextImpl::InvalidateGuestCodeRange(FEXCore::Core::InternalThreadState *Thread, uint64_t Start, uint64_t Length) {
// Potential deferred since Thread might not be valid.
// Thread object isn't valid very early in frontend's initialization.
// To be more optimal the frontend should provide this code with a valid Thread object earlier.
ScopedPotentialDeferredSignalWithUniqueLock CodeInvalidationLock(CodeInvalidationMutex, Thread);
InvalidateGuestCodeRangeInternal(this, Start, Length);
}
void ContextImpl::InvalidateGuestCodeRange(uint64_t Start, uint64_t Length, std::function<void(uint64_t start, uint64_t Length)> CallAfter) {
FHU::ScopedSignalMaskWithUniqueLock CodeInvalidationLock(CodeInvalidationMutex);
void ContextImpl::InvalidateGuestCodeRange(FEXCore::Core::InternalThreadState *Thread, uint64_t Start, uint64_t Length, std::function<void(uint64_t start, uint64_t Length)> CallAfter) {
// Potential deferred since Thread might not be valid.
// Thread object isn't valid very early in frontend's initialization.
// To be more optimal the frontend should provide this code with a valid Thread object earlier.
ScopedPotentialDeferredSignalWithUniqueLock CodeInvalidationLock(CodeInvalidationMutex, Thread);
InvalidateGuestCodeRangeInternal(this, Start, Length);
CallAfter(Start, Length);
@@ -1237,6 +1228,7 @@ namespace FEXCore::Context {
void ContextImpl::MarkMemoryShared() {
if (!IsMemoryShared) {
IsMemoryShared = true;
UpdateAtomicTSOEmulationConfig();
if (Config.TSOAutoMigration) {
std::lock_guard<std::mutex> lkThreads(ThreadCreationMutex);
@@ -1290,56 +1282,11 @@ namespace FEXCore::Context {
std::scoped_lock lk(CustomIRMutex);
InvalidateGuestCodeRange(Entrypoint, 1, [this](uint64_t Entrypoint, uint64_t) {
InvalidateGuestCodeRange(nullptr, Entrypoint, 1, [this](uint64_t Entrypoint, uint64_t) {
CustomIRHandlers.erase(Entrypoint);
});
}
// Debug interface
void Context::CompileRIP(FEXCore::Core::InternalThreadState *Thread, uint64_t RIP) {
uint64_t RIPBackup = Thread->CurrentFrame->State.rip;
Thread->CurrentFrame->State.rip = RIP;
auto CTX = static_cast<ContextImpl*>(Thread->CTX);
// Erase the RIP from all the storage backings if it exists
CTX->ThreadRemoveCodeEntry(Thread, RIP);
// We don't care if compilation passes or not
CTX->CompileBlock(Thread->CurrentFrame, RIP);
Thread->CurrentFrame->State.rip = RIPBackup;
}
uint64_t ContextImpl::GetThreadCount() const {
return Threads.size();
}
FEXCore::Core::RuntimeStats *ContextImpl::GetRuntimeStatsForThread(uint64_t Thread) {
return &Threads[Thread]->Stats;
}
bool ContextImpl::GetDebugDataForRIP(uint64_t RIP, FEXCore::Core::DebugData *Data) {
std::lock_guard<std::recursive_mutex> lk(ParentThread->LookupCache->WriteLock);
auto it = ParentThread->DebugStore.find(RIP);
if (it == ParentThread->DebugStore.end()) {
return false;
}
memcpy(Data, it->second.DebugData.get(), sizeof(FEXCore::Core::DebugData));
return true;
}
bool ContextImpl::FindHostCodeForRIP(uint64_t RIP, uint8_t **Code) {
uintptr_t HostCode = ParentThread->LookupCache->FindBlock(RIP);
if (!HostCode) {
return false;
}
*Code = reinterpret_cast<uint8_t*>(HostCode);
return true;
}
uint64_t HandleSyscall(FEXCore::HLE::SyscallHandler *Handler, FEXCore::Core::CpuStateFrame *Frame, FEXCore::HLE::SyscallArguments *Args) {
uint64_t Result{};
Result = Handler->HandleSyscall(Frame, Args);
@@ -1362,7 +1309,9 @@ namespace FEXCore::Context {
}
void ContextImpl::AppendThunkDefinitions(fextl::vector<FEXCore::IR::ThunkDefinition> const& Definitions) {
ThunkHandler->AppendThunkDefinitions(Definitions);
if (ThunkHandler) {
ThunkHandler->AppendThunkDefinitions(Definitions);
}
}
void ContextImpl::ConfigureAOTGen(FEXCore::Core::InternalThreadState *Thread, fextl::set<uint64_t> *ExternalBranches, uint64_t SectionMaxAddress) {
@@ -129,13 +129,6 @@ void Arm64Dispatcher::EmitDispatcher() {
and_(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r3, RipReg.R(), ARMEmitter::Reg::r3);
}
#ifdef VIXL_SIMULATOR
// VIXL simulator can't run syscalls.
constexpr bool SignalSafeCompile = false;
#else
constexpr bool SignalSafeCompile = true;
#endif
ARMEmitter::ForwardLabel NoBlock;
{
@@ -183,7 +176,7 @@ void Arm64Dispatcher::EmitDispatcher() {
{
ThreadStopHandlerAddressSpillSRA = GetCursorAddress<uint64_t>();
if (config.StaticRegisterAllocation)
SpillStaticRegs();
SpillStaticRegs(TMP1);
ThreadStopHandlerAddress = GetCursorAddress<uint64_t>();
@@ -197,28 +190,11 @@ void Arm64Dispatcher::EmitDispatcher() {
{
ExitFunctionLinkerAddress = GetCursorAddress<uint64_t>();
if (config.StaticRegisterAllocation)
SpillStaticRegs();
SpillStaticRegs(TMP1);
#ifndef _WIN32
if (SignalSafeCompile) {
// When compiling code, mask all signals to reduce the chance of reentrant allocations
// Args:
// X0: SETMASK
// X1: Pointer to mask value (uint64_t)
// X2: Pointer to old mask value (uint64_t)
// X3: Size of mask, sizeof(uint64_t)
// X8: Syscall
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, ~0ULL);
stp<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::x0, ARMEmitter::XReg::x0, ARMEmitter::Reg::rsp, -16);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, SIG_SETMASK);
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r1, ARMEmitter::Reg::rsp, 0);
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r2, ARMEmitter::Reg::rsp, 0);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r3, 8);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r8, SYS_rt_sigprocmask);
svc(0);
}
#endif
ldr(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
add(ARMEmitter::Size::i64Bit, ARMEmitter::XReg::x0, ARMEmitter::XReg::x0, 1);
str(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
mov(ARMEmitter::XReg::x0, STATE);
mov(ARMEmitter::XReg::x1, ARMEmitter::XReg::lr);
@@ -230,28 +206,17 @@ void Arm64Dispatcher::EmitDispatcher() {
blr(ARMEmitter::Reg::r2);
#endif
#ifndef _WIN32
if (SignalSafeCompile) {
// Now restore the signal mask
// Living in the same location
mov(ARMEmitter::XReg::x4, ARMEmitter::XReg::x0);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, SIG_SETMASK);
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r1, ARMEmitter::Reg::rsp, 0);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r2, 0);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r3, 8);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r8, SYS_rt_sigprocmask);
svc(0);
// Bring stack back
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, ARMEmitter::Reg::rsp, 16);
mov(ARMEmitter::XReg::x0, ARMEmitter::XReg::x4);
}
#endif
if (config.StaticRegisterAllocation)
FillStaticRegs();
ldr(ARMEmitter::XReg::x1, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
subs(ARMEmitter::Size::i64Bit, ARMEmitter::XReg::x1, ARMEmitter::XReg::x1, 1);
str(ARMEmitter::XReg::x1, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
// Trigger segfault if any deferred signals are pending
ldr(ARMEmitter::XReg::x1, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalFaultAddress));
str(ARMEmitter::XReg::zr, ARMEmitter::XReg::x1, 0);
br(ARMEmitter::Reg::r0);
}
@@ -260,31 +225,11 @@ void Arm64Dispatcher::EmitDispatcher() {
Bind(&NoBlock);
if (config.StaticRegisterAllocation)
SpillStaticRegs();
SpillStaticRegs(TMP1);
#ifndef _WIN32
if (SignalSafeCompile) {
// When compiling code, mask all signals to reduce the chance of reentrant allocations
// Args:
// X0: SETMASK
// X1: Pointer to mask value (uint64_t)
// X2: Pointer to old mask value (uint64_t)
// X3: Size of mask, sizeof(uint64_t)
// X8: Syscall
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, ~0ULL);
stp<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::x0, ARMEmitter::XReg::x2, ARMEmitter::Reg::rsp, -16);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, SIG_SETMASK);
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r1, ARMEmitter::Reg::rsp, 0);
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r2, ARMEmitter::Reg::rsp, 0);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r3, 8);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r8, SYS_rt_sigprocmask);
svc(0);
// Reload x2 to bring back RIP
ldr(ARMEmitter::XReg::x2, ARMEmitter::Reg::rsp, 8);
}
#endif
ldr(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
add(ARMEmitter::Size::i64Bit, ARMEmitter::XReg::x0, ARMEmitter::XReg::x0, 1);
str(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
ldr(ARMEmitter::XReg::x0, &l_CTX);
mov(ARMEmitter::XReg::x1, STATE);
@@ -297,25 +242,17 @@ void Arm64Dispatcher::EmitDispatcher() {
blr(ARMEmitter::Reg::r3); // { CTX, Frame, RIP}
#endif
#ifndef _WIN32
if (SignalSafeCompile) {
// Now restore the signal mask
// Living in the same location
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, SIG_SETMASK);
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r1, ARMEmitter::Reg::rsp, 0);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r2, 0);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r3, 8);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r8, SYS_rt_sigprocmask);
svc(0);
// Bring stack back
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, ARMEmitter::Reg::rsp, 16);
}
#endif
if (config.StaticRegisterAllocation)
FillStaticRegs();
ldr(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
subs(ARMEmitter::Size::i64Bit, ARMEmitter::XReg::x0, ARMEmitter::XReg::x0, 1);
str(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
// Trigger segfault if any deferred signals are pending
ldr(TMP1, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalFaultAddress));
str(ARMEmitter::XReg::zr, TMP1, 0);
b(&LoopTop);
}
@@ -341,7 +278,7 @@ void Arm64Dispatcher::EmitDispatcher() {
GuestSignal_SIGILL = GetCursorAddress<uint64_t>();
if (config.StaticRegisterAllocation)
SpillStaticRegs();
SpillStaticRegs(TMP1);
hlt(0);
}
@@ -352,7 +289,7 @@ void Arm64Dispatcher::EmitDispatcher() {
GuestSignal_SIGTRAP = GetCursorAddress<uint64_t>();
if (config.StaticRegisterAllocation)
SpillStaticRegs();
SpillStaticRegs(TMP1);
brk(0);
}
@@ -363,20 +300,28 @@ void Arm64Dispatcher::EmitDispatcher() {
GuestSignal_SIGSEGV = GetCursorAddress<uint64_t>();
if (config.StaticRegisterAllocation)
SpillStaticRegs();
SpillStaticRegs(TMP1);
// hlt/udf = SIGILL
// brk = SIGTRAP
// ??? = SIGSEGV
// Force a SIGSEGV by loading zero
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r1, 0);
ldr(ARMEmitter::XReg::x1, ARMEmitter::Reg::r1);
if (CTX->ExitOnHLTEnabled()) {
ldr(ARMEmitter::XReg::x0, STATE_PTR(CpuStateFrame, ReturningStackLocation));
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, ARMEmitter::Reg::r0, 0);
PopCalleeSavedRegisters();
ret();
}
else {
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r1, 0);
ldr(ARMEmitter::XReg::x1, ARMEmitter::Reg::r1);
}
}
{
ThreadPauseHandlerAddressSpillSRA = GetCursorAddress<uint64_t>();
if (config.StaticRegisterAllocation)
SpillStaticRegs();
SpillStaticRegs(TMP1);
ThreadPauseHandlerAddress = GetCursorAddress<uint64_t>();
// We are pausing, this means the frontend should be waiting for this thread to idle
@@ -453,7 +398,7 @@ void Arm64Dispatcher::EmitDispatcher() {
LUDIVHandlerAddress = GetCursorAddress<uint64_t>();
PushDynamicRegsAndLR(ARMEmitter::Reg::r3);
SpillStaticRegs();
SpillStaticRegs(ARMEmitter::Reg::r3);
ldr(ARMEmitter::XReg::x3, STATE_PTR(CpuStateFrame, Pointers.AArch64.LUDIV));
#ifdef VIXL_SIMULATOR
@@ -475,7 +420,7 @@ void Arm64Dispatcher::EmitDispatcher() {
LDIVHandlerAddress = GetCursorAddress<uint64_t>();
PushDynamicRegsAndLR(ARMEmitter::Reg::r3);
SpillStaticRegs();
SpillStaticRegs(ARMEmitter::Reg::r3);
ldr(ARMEmitter::XReg::x3, STATE_PTR(CpuStateFrame, Pointers.AArch64.LDIV));
#ifdef VIXL_SIMULATOR
@@ -497,7 +442,7 @@ void Arm64Dispatcher::EmitDispatcher() {
LUREMHandlerAddress = GetCursorAddress<uint64_t>();
PushDynamicRegsAndLR(ARMEmitter::Reg::r3);
SpillStaticRegs();
SpillStaticRegs(ARMEmitter::Reg::r3);
ldr(ARMEmitter::XReg::x3, STATE_PTR(CpuStateFrame, Pointers.AArch64.LUREM));
#ifdef VIXL_SIMULATOR
@@ -519,7 +464,7 @@ void Arm64Dispatcher::EmitDispatcher() {
LREMHandlerAddress = GetCursorAddress<uint64_t>();
PushDynamicRegsAndLR(ARMEmitter::Reg::r3);
SpillStaticRegs();
SpillStaticRegs(ARMEmitter::Reg::r3);
ldr(ARMEmitter::XReg::x3, STATE_PTR(CpuStateFrame, Pointers.AArch64.LREM));
@@ -90,6 +90,8 @@ public:
virtual void GetSRAFPRMapping(uint8_t Mapping[16]) const {
}
const DispatcherConfig& GetConfig() const { return config; }
protected:
Dispatcher(FEXCore::Context::ContextImpl *ctx, const DispatcherConfig &Config)
: CTX {ctx}
@@ -170,38 +170,11 @@ X86Dispatcher::X86Dispatcher(FEXCore::Context::ContextImpl *ctx, const Dispatche
ret();
}
constexpr bool SignalSafeCompile = true;
// Block creation
{
L(NoBlock);
#ifndef _WIN32
if (SignalSafeCompile) {
// When compiling code, mask all signals to reduce the chance of reentrant allocations
// RDI: SETMASK
// RSI: Pointer to mask value (uint64_t)
// RDX: Pointer to old mask value (uint64_t)
// R10: Size of mask, sizeof(uint64_t)
// RAX: Syscall
// Backup rdx
mov(r9, rdx);
mov(rdi, ~0ULL);
sub(rsp, 16);
mov(qword [rsp], rdi);
mov(qword [rsp + 8], rdi);
mov(rdi, SIG_SETMASK);
mov(rsi, rsp);
mov(rdx, rsp);
mov(r10, 8);
mov(rax, SYS_rt_sigprocmask);
syscall();
mov(rdx, r9);
}
#endif
inc(qword [STATE + offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount)]);
// {rdi, rsi, rdx}
mov(rdi, reinterpret_cast<uint64_t>(CTX));
@@ -210,26 +183,15 @@ X86Dispatcher::X86Dispatcher(FEXCore::Context::ContextImpl *ctx, const Dispatche
call(rax);
#ifndef _WIN32
if (SignalSafeCompile) {
// Now restore the signal mask
// Living in the same location
// Backup rdx
mov(r9, rdx);
dec(qword [STATE + offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount)]);
mov(rdi, SIG_SETMASK);
mov(rsi, rsp);
mov(rdx, 0); // Don't care about result
mov(r10, 8);
mov(rax, SYS_rt_sigprocmask);
syscall();
Label AfterStore;
// Skip the deferred fault address if the refcount isn't zero
jne(AfterStore);
mov(rax, qword [STATE + offsetof(FEXCore::Core::CPUState, DeferredSignalFaultAddress)]);
mov(rax, qword [rax]);
// Bring stack back
add(rsp, 16);
mov(rdx, r9);
}
#endif
L(AfterStore);
// rdx already contains RIP here
jmp(LoopTop);
@@ -237,34 +199,7 @@ X86Dispatcher::X86Dispatcher(FEXCore::Context::ContextImpl *ctx, const Dispatche
{
ExitFunctionLinkerAddress = getCurr<uint64_t>();
#ifndef _WIN32
if (SignalSafeCompile) {
// When compiling code, mask all signals to reduce the chance of reentrant allocations
// RDI: SETMASK
// RSI: Pointer to mask value (uint64_t)
// RDX: Pointer to old mask value (uint64_t)
// R10: Size of mask, sizeof(uint64_t)
// RAX: Syscall
// Backup rax
mov(r9, rax);
mov(rdi, ~0ULL);
sub(rsp, 16);
mov(qword [rsp], rdi);
mov(qword [rsp + 8], rdi);
mov(rdi, SIG_SETMASK);
mov(rsi, rsp);
mov(rdx, rsp);
mov(r10, 8);
mov(rax, SYS_rt_sigprocmask);
syscall();
mov(rax, r9);
}
#endif
inc(qword [STATE + offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount)]);
// {rdi, rsi}
mov(rdi, STATE);
@@ -272,30 +207,17 @@ X86Dispatcher::X86Dispatcher(FEXCore::Context::ContextImpl *ctx, const Dispatche
call(qword STATE_PTR(CpuStateFrame, Pointers.Common.ExitFunctionLink));
#ifndef _WIN32
if (SignalSafeCompile) {
// Now restore the signal mask
// Living in the same location
// Backup rax
mov(r9, rax);
dec(qword [STATE + offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount)]);
mov(rdi, SIG_SETMASK);
mov(rsi, rsp);
mov(rdx, 0); // Don't care about result
mov(r10, 8);
mov(rax, SYS_rt_sigprocmask);
syscall();
Label AfterStore;
// Skip the deferred fault address if the refcount isn't zero
jne(AfterStore);
mov(rbx, qword [STATE + offsetof(FEXCore::Core::CPUState, DeferredSignalFaultAddress)]);
mov(qword [rbx], rbx);
// Bring stack back
add(rsp, 16);
L(AfterStore);
jmp(r9);
}
else
#endif
{
jmp(rax);
}
jmp(rax);
}
{
@@ -54,7 +54,12 @@ HostFeatures::HostFeatures() {
#ifdef VIXL_SIMULATOR
auto Features = vixl::CPUFeatures::All();
#else
#ifndef _WIN32
auto Features = vixl::CPUFeatures::InferFromOS();
#else
// Need to use ID registers in WINE.
auto Features = vixl::CPUFeatures::InferFromIDRegisters();
#endif
#endif
SupportsAES = Features.Has(vixl::CPUFeatures::Feature::kAES);
SupportsCRC = Features.Has(vixl::CPUFeatures::Feature::kCRC32);
@@ -143,6 +143,15 @@ DEF_OP(CPUID) {
memcpy(DstPtr, &Results, sizeof(uint32_t) * 4);
}
DEF_OP(XGETBV) {
auto Op = IROp->C<IR::IROp_XGetBV>();
uint32_t *DstPtr = GetDest<uint32_t*>(Data->SSAData, Node);
const uint32_t Function = *GetSrc<uint32_t*>(Data->SSAData, Op->Function);
auto Results = Data->State->CTX->RunXCRFunction(Function);
memcpy(DstPtr, &Results, sizeof(uint32_t) * 2);
}
#undef DEF_OP
} // namespace FEXCore::CPU
@@ -118,6 +118,7 @@ constexpr OpHandlerArray InterpreterOpHandlers = [] {
REGISTER_OP(VALIDATECODE, ValidateCode);
REGISTER_OP(THREADREMOVECODEENTRY, ThreadRemoveCodeEntry);
REGISTER_OP(CPUID, CPUID);
REGISTER_OP(XGETBV, XGETBV);
// Conversion ops
REGISTER_OP(VINSGPR, VInsGPR);
@@ -154,6 +154,7 @@ namespace FEXCore::CPU {
DEF_OP(ValidateCode);
DEF_OP(ThreadRemoveCodeEntry);
DEF_OP(CPUID);
DEF_OP(XGETBV);
///< Conversion ops
DEF_OP(VInsGPR);
@@ -494,7 +494,7 @@ DEF_OP(PDep) {
// We sadly need to spill regs for this for the time being
// TODO: Remove when scratch registers can be allocated
// explicitly.
SpillStaticRegs(false, SpillCode);
SpillStaticRegs(TMP1, false, SpillCode);
mov(EmitSize, InputReg, Input);
@@ -558,7 +558,7 @@ DEF_OP(PExt) {
// We sadly need to spill a reg for this for the time being
// TODO: Remove when scratch registers can be allocated
// explicitly.
SpillStaticRegs(false, 1U << Mask.Idx());
SpillStaticRegs(TMP2, false, 1U << Mask.Idx());
mov(EmitSize, Mask, ZeroReg);
// Main loop
@@ -22,7 +22,7 @@ namespace FEXCore::CPU {
DEF_OP(CallbackReturn) {
// spill back to CTX
SpillStaticRegs();
SpillStaticRegs(TMP1);
// First we must reset the stack
ResetStack();
@@ -175,7 +175,7 @@ DEF_OP(Syscall) {
FPRSpillMask = CALLER_FPR_MASK;
}
SpillStaticRegs(true, GPRSpillMask, FPRSpillMask);
SpillStaticRegs(TMP1, true, GPRSpillMask, FPRSpillMask);
// Now that we are spilled, store in the state that we are in a syscall
// Still without overwriting registers that matter
@@ -216,8 +216,11 @@ DEF_OP(Syscall) {
PopDynamicRegsAndLR();
// Move result to its destination register
mov(ARMEmitter::Size::i64Bit, GetReg(Node), ARMEmitter::Reg::r0);
if ((Flags & FEXCore::IR::SyscallFlags::NORETURNEDRESULT) != FEXCore::IR::SyscallFlags::NORETURNEDRESULT) {
// Move result to its destination register.
// Only if `NORETURNEDRESULT` wasn't set, otherwise we might overwrite the CPUState refilled with `FillStaticRegs`
mov(ARMEmitter::Size::i64Bit, GetReg(Node), ARMEmitter::Reg::r0);
}
}
}
@@ -257,7 +260,7 @@ DEF_OP(InlineSyscall) {
// Ordering is incredibly important here
// We must spill any overlapping registers first THEN claim we are in a syscall without invalidating state at all
// Only spill the registers that intersect with our usage
SpillStaticRegs(false, SpillMask);
SpillStaticRegs(TMP1, false, SpillMask);
// Now that we are spilled, store in the state that we are in a syscall
// Still without overwriting registers that matter
@@ -325,7 +328,7 @@ DEF_OP(Thunk) {
// X0: CTX
// X1: Args (from guest stack)
SpillStaticRegs(); // spill to ctx before ra64 spill
SpillStaticRegs(TMP1); // spill to ctx before ra64 spill
PushDynamicRegsAndLR(TMP1);
@@ -400,12 +403,12 @@ DEF_OP(ThreadRemoveCodeEntry) {
// X1: RIP
PushDynamicRegsAndLR(TMP1);
SpillStaticRegs(TMP1);
mov(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, STATE.R());
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r1, Entry);
ldr(ARMEmitter::XReg::x2, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.ThreadRemoveCodeEntryFromJIT));
SpillStaticRegs();
#ifdef VIXL_SIMULATOR
GenerateIndirectRuntimeCall<void, void*, void*>(ARMEmitter::Reg::r2);
#else
@@ -421,7 +424,7 @@ DEF_OP(CPUID) {
auto Op = IROp->C<IR::IROp_CPUID>();
PushDynamicRegsAndLR(TMP1);
SpillStaticRegs();
SpillStaticRegs(TMP1);
// x0 = CPUID Handler
// x1 = CPUID Function
@@ -447,6 +450,34 @@ DEF_OP(CPUID) {
mov(ARMEmitter::Size::i64Bit, Dst.second, ARMEmitter::Reg::r1);
}
DEF_OP(XGETBV) {
auto Op = IROp->C<IR::IROp_XGetBV>();
PushDynamicRegsAndLR(TMP1);
SpillStaticRegs(TMP1);
// x0 = CPUID Handler
// x1 = XCR Function
ldr(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.CPUIDObj));
ldr(ARMEmitter::XReg::x2, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.XCRFunction));
mov(ARMEmitter::Size::i32Bit, ARMEmitter::Reg::r1, GetReg(Op->Function.ID()));
#ifdef VIXL_SIMULATOR
GenerateIndirectRuntimeCall<uint64_t, void*, uint32_t>(ARMEmitter::Reg::r2);
#else
blr(ARMEmitter::Reg::r2);
#endif
FillStaticRegs();
PopDynamicRegsAndLR();
// Results are in x0
// Results want to be in a i32v2 vector
auto Dst = GetRegPair(Node);
mov(ARMEmitter::Size::i32Bit, Dst.first, ARMEmitter::Reg::r0);
lsr(ARMEmitter::Size::i64Bit, Dst.second, ARMEmitter::Reg::r0, 32);
}
#undef DEF_OP
}
+46 -34
View File
@@ -83,7 +83,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header const *IROp, IR::NodeID Node) {
} else {
switch(Info.ABI) {
case FABI_VOID_U16:{
SpillStaticRegs();
SpillStaticRegs(TMP1);
PushDynamicRegsAndLR(TMP1);
@@ -103,7 +103,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header const *IROp, IR::NodeID Node) {
break;
case FABI_F80_F32:{
SpillStaticRegs();
SpillStaticRegs(TMP1);
PushDynamicRegsAndLR(TMP1);
const auto Src1 = GetVReg(IROp->Args[0].ID());
@@ -127,7 +127,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header const *IROp, IR::NodeID Node) {
break;
case FABI_F80_F64:{
SpillStaticRegs();
SpillStaticRegs(TMP1);
PushDynamicRegsAndLR(TMP1);
@@ -153,7 +153,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header const *IROp, IR::NodeID Node) {
case FABI_F80_I16:
case FABI_F80_I32: {
SpillStaticRegs();
SpillStaticRegs(TMP1);
PushDynamicRegsAndLR(TMP1);
@@ -183,7 +183,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header const *IROp, IR::NodeID Node) {
break;
case FABI_F32_F80:{
SpillStaticRegs();
SpillStaticRegs(TMP1);
PushDynamicRegsAndLR(TMP1);
@@ -209,7 +209,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header const *IROp, IR::NodeID Node) {
break;
case FABI_F64_F80:{
SpillStaticRegs();
SpillStaticRegs(TMP1);
PushDynamicRegsAndLR(TMP1);
@@ -235,7 +235,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header const *IROp, IR::NodeID Node) {
break;
case FABI_F64_F64: {
SpillStaticRegs();
SpillStaticRegs(TMP1);
PushDynamicRegsAndLR(TMP1);
@@ -259,7 +259,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header const *IROp, IR::NodeID Node) {
break;
case FABI_F64_F64_F64: {
SpillStaticRegs();
SpillStaticRegs(TMP1);
PushDynamicRegsAndLR(TMP1);
@@ -285,7 +285,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header const *IROp, IR::NodeID Node) {
break;
case FABI_I16_F80:{
SpillStaticRegs();
SpillStaticRegs(TMP1);
PushDynamicRegsAndLR(TMP1);
@@ -310,7 +310,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header const *IROp, IR::NodeID Node) {
}
break;
case FABI_I32_F80:{
SpillStaticRegs();
SpillStaticRegs(TMP1);
PushDynamicRegsAndLR(TMP1);
@@ -335,7 +335,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header const *IROp, IR::NodeID Node) {
}
break;
case FABI_I64_F80:{
SpillStaticRegs();
SpillStaticRegs(TMP1);
PushDynamicRegsAndLR(TMP1);
@@ -360,7 +360,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header const *IROp, IR::NodeID Node) {
}
break;
case FABI_I64_F80_F80:{
SpillStaticRegs();
SpillStaticRegs(TMP1);
PushDynamicRegsAndLR(TMP1);
@@ -388,7 +388,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header const *IROp, IR::NodeID Node) {
}
break;
case FABI_F80_F80:{
SpillStaticRegs();
SpillStaticRegs(TMP1);
PushDynamicRegsAndLR(TMP1);
@@ -415,7 +415,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header const *IROp, IR::NodeID Node) {
}
break;
case FABI_F80_F80_F80:{
SpillStaticRegs();
SpillStaticRegs(TMP1);
PushDynamicRegsAndLR(TMP1);
@@ -446,7 +446,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header const *IROp, IR::NodeID Node) {
}
break;
case FABI_I32_I64_I64_I128_I128_I16: {
SpillStaticRegs();
SpillStaticRegs(TMP1);
PushDynamicRegsAndLR(TMP1);
const auto Op = IROp->C<IR::IROp_VPCMPESTRX>();
@@ -486,7 +486,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header const *IROp, IR::NodeID Node) {
break;
}
case FABI_I32_I128_I128_I16: {
SpillStaticRegs();
SpillStaticRegs(TMP1);
PushDynamicRegsAndLR(TMP1);
const auto Op = IROp->C<IR::IROp_VPCMPISTRX>();
@@ -615,6 +615,11 @@ Arm64JITCore::Arm64JITCore(FEXCore::Context::ContextImpl *ctx, FEXCore::Core::In
Common.CPUIDFunction = PMF.GetConvertedPointer();
}
{
FEXCore::Utils::MemberFunctionToPointerCast PMF(&FEXCore::CPUIDEmu::RunXCRFunction);
Common.XCRFunction = PMF.GetConvertedPointer();
}
Common.SyscallHandlerObj = reinterpret_cast<uint64_t>(CTX->SyscallHandler);
Common.SyscallHandlerFunc = reinterpret_cast<uint64_t>(FEXCore::Context::HandleSyscall);
Common.ExitFunctionLink = reinterpret_cast<uintptr_t>(&Context::ContextImpl::ThreadExitFunctionLink<Arm64JITCore_ExitFunctionLink>);
@@ -634,6 +639,25 @@ Arm64JITCore::Arm64JITCore(FEXCore::Context::ContextImpl *ctx, FEXCore::Core::In
// Must be done after Dispatcher init
ClearCache();
// Setup dynamic dispatch.
if (CTX->Dispatcher->GetConfig().StaticRegisterAllocation) {
RT_LoadRegister = &Arm64JITCore::Op_LoadRegisterSRA;
RT_StoreRegister = &Arm64JITCore::Op_StoreRegisterSRA;
}
else {
RT_LoadRegister = &Arm64JITCore::Op_LoadRegister;
RT_StoreRegister = &Arm64JITCore::Op_StoreRegister;
}
if (ParanoidTSO()) {
RT_LoadMemTSO = &Arm64JITCore::Op_ParanoidLoadMemTSO;
RT_StoreMemTSO = &Arm64JITCore::Op_ParanoidStoreMemTSO;
}
else {
RT_LoadMemTSO = &Arm64JITCore::Op_LoadMemTSO;
RT_StoreMemTSO = &Arm64JITCore::Op_StoreMemTSO;
}
}
void Arm64JITCore::EmitDetectionString() {
@@ -814,6 +838,7 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry,
for (auto [CodeNode, IROp] : IR->GetCode(BlockNode)) {
const auto ID = IR->GetID(CodeNode);
switch (IROp->Op) {
#define REGISTER_OP_RT(op, x) case FEXCore::IR::IROps::OP_##op: std::invoke(RT_##x, this, IROp, ID); break
#define REGISTER_OP(op, x) case FEXCore::IR::IROps::OP_##op: Op_##x(IROp, ID); break
// ALU ops
REGISTER_OP(TRUNCELEMENTPAIR, TruncElementPair);
@@ -891,6 +916,7 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry,
REGISTER_OP(VALIDATECODE, ValidateCode);
REGISTER_OP(THREADREMOVECODEENTRY, ThreadRemoveCodeEntry);
REGISTER_OP(CPUID, CPUID);
REGISTER_OP(XGETBV, XGETBV);
// Conversion ops
REGISTER_OP(VINSGPR, VInsGPR);
@@ -920,8 +946,8 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry,
// Memory ops
REGISTER_OP(LOADCONTEXT, LoadContext);
REGISTER_OP(STORECONTEXT, StoreContext);
REGISTER_OP(LOADREGISTER, LoadRegister);
REGISTER_OP(STOREREGISTER, StoreRegister);
REGISTER_OP_RT(LOADREGISTER, LoadRegister);
REGISTER_OP_RT(STOREREGISTER, StoreRegister);
REGISTER_OP(LOADCONTEXTINDEXED, LoadContextIndexed);
REGISTER_OP(STORECONTEXTINDEXED, StoreContextIndexed);
REGISTER_OP(SPILLREGISTER, SpillRegister);
@@ -930,22 +956,8 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry,
REGISTER_OP(STOREFLAG, StoreFlag);
REGISTER_OP(LOADMEM, LoadMem);
REGISTER_OP(STOREMEM, StoreMem);
case FEXCore::IR::IROps::OP_LOADMEMTSO:
if (ParanoidTSO()) {
Op_ParanoidLoadMemTSO(IROp, ID);
}
else {
Op_LoadMemTSO(IROp, ID);
}
break;
case FEXCore::IR::IROps::OP_STOREMEMTSO:
if (ParanoidTSO()) {
Op_ParanoidStoreMemTSO(IROp, ID);
}
else {
Op_StoreMemTSO(IROp, ID);
}
break;
REGISTER_OP_RT(LOADMEMTSO, LoadMemTSO);
REGISTER_OP_RT(STOREMEMTSO, StoreMemTSO);
REGISTER_OP(VLOADVECTORMASKED, VLoadVectorMasked);
REGISTER_OP(VSTOREVECTORMASKED, VStoreVectorMasked);
@@ -210,6 +210,16 @@ private:
/** @} */
uint32_t SpillSlots{};
using OpType = void (Arm64JITCore::*)(IR::IROp_Header const *IROp, IR::NodeID Node);
// Runtime selection;
// Load and store register style.
OpType RT_LoadRegister;
OpType RT_StoreRegister;
// Load and store TSO memory style
OpType RT_LoadMemTSO;
OpType RT_StoreMemTSO;
#define DEF_OP(x) void Op_##x(IR::IROp_Header const *IROp, IR::NodeID Node)
///< Unhandled handler
@@ -297,6 +307,7 @@ private:
DEF_OP(ValidateCode);
DEF_OP(ThreadRemoveCodeEntry);
DEF_OP(CPUID);
DEF_OP(XGETBV);
///< Conversion ops
DEF_OP(VInsGPR);
@@ -318,6 +329,8 @@ private:
DEF_OP(StoreContext);
DEF_OP(LoadRegister);
DEF_OP(StoreRegister);
DEF_OP(LoadRegisterSRA);
DEF_OP(StoreRegisterSRA);
DEF_OP(LoadContextIndexed);
DEF_OP(StoreContextIndexed);
DEF_OP(SpillRegister);
@@ -128,6 +128,170 @@ DEF_OP(LoadRegister) {
const auto Op = IROp->C<IR::IROp_LoadRegister>();
const auto OpSize = IROp->Size;
if (Op->Class == IR::GPRClass) {
[[maybe_unused]] const auto regId = (Op->Offset / Core::CPUState::GPR_REG_SIZE) - 1;
const auto regOffs = Op->Offset & 7;
LOGMAN_THROW_A_FMT(regId < SRA64.size(), "out of range regId");
switch (OpSize) {
case 1:
LOGMAN_THROW_AA_FMT(regOffs == 0 || regOffs == 1, "unexpected regOffs");
ldrb(GetReg(Node), STATE, Op->Offset);
break;
case 2:
LOGMAN_THROW_AA_FMT(regOffs == 0, "unexpected regOffs");
ldrh(GetReg(Node), STATE, Op->Offset);
break;
case 4:
LOGMAN_THROW_AA_FMT(regOffs == 0, "unexpected regOffs");
ldr(GetReg(Node).W(), STATE, Op->Offset);
break;
case 8:
LOGMAN_THROW_AA_FMT(regOffs == 0, "unexpected regOffs");
ldr(GetReg(Node).X(), STATE, Op->Offset);
break;
default:
LOGMAN_MSG_A_FMT("Unhandled LoadRegister GPR size: {}", OpSize);
break;
}
}
else if (Op->Class == IR::FPRClass) {
const auto regSize = HostSupportsSVE ? Core::CPUState::XMM_AVX_REG_SIZE
: Core::CPUState::XMM_SSE_REG_SIZE;
[[maybe_unused]] const auto regId = (Op->Offset - offsetof(Core::CpuStateFrame, State.xmm.avx.data[0][0])) / regSize;
LOGMAN_THROW_A_FMT(HostSupportsSVE, "Unsupported code path!");
LOGMAN_THROW_A_FMT(regId < SRAFPR.size(), "out of range regId");
const auto host = GetVReg(Node);
const auto regOffs = Op->Offset & 15;
switch (OpSize) {
case 1: {
LOGMAN_THROW_AA_FMT(regOffs == 0, "unexpected regOffs: {}", regOffs);
ldrb(host, STATE, Op->Offset);
break;
}
case 2: {
LOGMAN_THROW_AA_FMT(regOffs == 0, "unexpected regOffs: {}", regOffs);
ldrh(host, STATE, Op->Offset);
break;
}
case 4: {
LOGMAN_THROW_AA_FMT((regOffs & 3) == 0, "unexpected regOffs: {}", regOffs);
ldr(host.S(), STATE, Op->Offset);
break;
}
case 8: {
LOGMAN_THROW_AA_FMT((regOffs & 7) == 0, "unexpected regOffs: {}", regOffs);
ldr(host.D(), STATE, Op->Offset);
break;
}
case 16: {
LOGMAN_THROW_AA_FMT(regOffs == 0, "unexpected regOffs: {}", regOffs);
ldr(host.Q(), STATE, Op->Offset);
break;
}
}
} else {
LOGMAN_THROW_AA_FMT(false, "Unhandled Op->Class {}", Op->Class);
}
}
DEF_OP(StoreRegister) {
const auto Op = IROp->C<IR::IROp_StoreRegister>();
const auto OpSize = IROp->Size;
if (Op->Class == IR::GPRClass) {
[[maybe_unused]] const auto regId = (Op->Offset / Core::CPUState::GPR_REG_SIZE) - 1;
const auto regOffs = Op->Offset & 7;
LOGMAN_THROW_A_FMT(regId < SRA64.size(), "out of range regId");
const auto Src = GetReg(Op->Value.ID());
switch (OpSize) {
case 1:
LOGMAN_THROW_AA_FMT(regOffs == 0 || regOffs == 1, "unexpected regOffs");
strb(Src, STATE, Op->Offset);
break;
case 2:
LOGMAN_THROW_AA_FMT(regOffs == 0, "unexpected regOffs");
strh(Src, STATE, Op->Offset);
break;
case 4:
LOGMAN_THROW_AA_FMT(regOffs == 0, "unexpected regOffs");
str(Src.W(), STATE, Op->Offset);
break;
case 8:
LOGMAN_THROW_AA_FMT(regOffs == 0, "unexpected regOffs");
str(Src.X(), STATE, Op->Offset);
break;
default:
LOGMAN_MSG_A_FMT("Unhandled StoreRegister GPR size: {}", OpSize);
break;
}
} else if (Op->Class == IR::FPRClass) {
const auto regSize = HostSupportsSVE ? Core::CPUState::XMM_AVX_REG_SIZE
: Core::CPUState::XMM_SSE_REG_SIZE;
[[maybe_unused]] const auto regId = (Op->Offset - offsetof(Core::CpuStateFrame, State.xmm.avx.data[0][0])) / regSize;
LOGMAN_THROW_A_FMT(HostSupportsSVE, "Unsupported code path!");
LOGMAN_THROW_A_FMT(regId < SRAFPR.size(), "regId out of range");
const auto host = GetVReg(Op->Value.ID());
const auto regOffs = Op->Offset & 15;
switch (OpSize) {
case 1:
strb(host, STATE, Op->Offset);
break;
case 2:
LOGMAN_THROW_AA_FMT((regOffs & 1) == 0, "unexpected regOffs: {}", regOffs);
strh(host, STATE, Op->Offset);
break;
case 4:
LOGMAN_THROW_AA_FMT((regOffs & 3) == 0, "unexpected regOffs: {}", regOffs);
str(host.S(), STATE, Op->Offset);
break;
case 8:
LOGMAN_THROW_AA_FMT((regOffs & 7) == 0, "unexpected regOffs: {}", regOffs);
str(host.D(), STATE, Op->Offset);
break;
case 16:
LOGMAN_THROW_AA_FMT(regOffs == 0, "unexpected regOffs: {}", regOffs);
str(host.Q(), STATE, Op->Offset);
break;
default:
LOGMAN_MSG_A_FMT("Unhandled StoreRegister FPR size: {}", OpSize);
break;
}
} else {
LOGMAN_THROW_AA_FMT(false, "Unhandled Op->Class {}", Op->Class);
}
}
DEF_OP(LoadRegisterSRA) {
const auto Op = IROp->C<IR::IROp_LoadRegister>();
const auto OpSize = IROp->Size;
if (Op->Class == IR::GPRClass) {
const auto regId = (Op->Offset - offsetof(Core::CpuStateFrame, State.gregs[0])) / Core::CPUState::GPR_REG_SIZE;
const auto regOffs = Op->Offset & 7;
@@ -312,7 +476,7 @@ DEF_OP(LoadRegister) {
}
}
DEF_OP(StoreRegister) {
DEF_OP(StoreRegisterSRA) {
const auto Op = IROp->C<IR::IROp_StoreRegister>();
const auto OpSize = IROp->Size;
@@ -500,7 +664,6 @@ DEF_OP(StoreRegister) {
}
}
DEF_OP(LoadContextIndexed) {
const auto Op = IROp->C<IR::IROp_LoadContextIndexed>();
const auto OpSize = IROp->Size;
@@ -12,6 +12,8 @@ $end_info$
#include "Interface/Core/JIT/Arm64/JITClass.h"
#include "FEXCore/Debug/InternalThreadState.h"
#include <FEXCore/Core/SignalDelegator.h>
namespace FEXCore::CPU {
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const *IROp, IR::NodeID Node)
@@ -37,7 +39,6 @@ DEF_OP(Fence) {
}
}
#ifndef _WIN32
DEF_OP(Break) {
auto Op = IROp->C<IR::IROp_Break>();
@@ -59,15 +60,15 @@ DEF_OP(Break) {
str(ARMEmitter::XReg::x1, STATE, offsetof(FEXCore::Core::CpuStateFrame, SynchronousFaultData));
switch (Op->Reason.Signal) {
case SIGILL:
case Core::FAULT_SIGILL:
ldr(TMP1, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.GuestSignal_SIGILL));
br(TMP1);
break;
case SIGTRAP:
case Core::FAULT_SIGTRAP:
ldr(TMP1, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.GuestSignal_SIGTRAP));
br(TMP1);
break;
case SIGSEGV:
case Core::FAULT_SIGSEGV:
ldr(TMP1, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.GuestSignal_SIGSEGV));
br(TMP1);
break;
@@ -77,11 +78,6 @@ DEF_OP(Break) {
break;
}
}
#else
DEF_OP(Break) {
ERROR_AND_DIE_FMT("Unsupported");
}
#endif
DEF_OP(GetRoundingMode) {
auto Dst = GetReg(Node);
@@ -148,7 +144,7 @@ DEF_OP(Print) {
auto Op = IROp->C<IR::IROp_Print>();
PushDynamicRegsAndLR(TMP1);
SpillStaticRegs();
SpillStaticRegs(TMP1);
if (IsGPR(Op->Value.ID())) {
mov(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, GetReg(Op->Value.ID()));
@@ -174,7 +170,7 @@ DEF_OP(ProcessorID) {
// Ordering is incredibly important here
// We must spill any overlapping registers first THEN claim we are in a syscall without invalidating state at all
// Only spill the registers that intersect with our usage
SpillStaticRegs(false, SpillMask);
SpillStaticRegs(TMP1, false, SpillMask);
// Now that we are spilled, store in the state that we are in a syscall
// Still without overwriting registers that matter
@@ -146,6 +146,7 @@ DEF_OP(Syscall) {
auto Op = IROp->C<IR::IROp_Syscall>();
// XXX: This is very terrible, but I don't care for right now
FEXCore::IR::SyscallFlags Flags = Op->Flags;
auto NumPush = RA64.size();
for (auto &Reg : RA64)
@@ -186,7 +187,11 @@ DEF_OP(Syscall) {
for (uint32_t i = RA64.size(); i > 0; --i)
pop(RA64[i - 1]);
mov (GetDst<RA_64>(Node), rax);
if ((Flags & FEXCore::IR::SyscallFlags::NORETURNEDRESULT) != FEXCore::IR::SyscallFlags::NORETURNEDRESULT) {
// Move result to its destination register.
// Only if `NORETURNEDRESULT` wasn't set, otherwise we might overwrite the CPUState refilled with `FillStaticRegs`
mov (GetDst<RA_64>(Node), rax);
}
}
DEF_OP(Thunk) {
@@ -302,6 +307,42 @@ DEF_OP(CPUID) {
mov(Dst.second, rdx);
}
DEF_OP(XGETBV) {
auto Op = IROp->C<IR::IROp_XGetBV>();
for (auto &Reg : RA64)
push(Reg);
// CPUID ABI
// this: rdi
// Function: rsi
//
// Result: RAX, RDX. 4xi32
// rsi can be in the source registers, so copy argument to edx first
mov (esi, GetSrc<RA_32>(Op->Function.ID()));
mov (rdi, qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.CPUIDObj)]);
auto NumPush = RA64.size();
if (NumPush & 1)
sub(rsp, 8); // Align
// {rdi, rsi, rdx}
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.XCRFunction)]);
if (NumPush & 1)
add(rsp, 8); // Align
for (uint32_t i = RA64.size(); i > 0; --i)
pop(RA64[i - 1]);
auto Dst = GetSrcPair<RA_64>(Node);
mov(Dst.first.cvt32(), eax);
mov(Dst.second, rax);
shr(Dst.second, 32);
}
#undef DEF_OP
void X86JITCore::RegisterBranchHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &X86JITCore::Op_##x
@@ -314,6 +355,7 @@ void X86JITCore::RegisterBranchHandlers() {
REGISTER_OP(VALIDATECODE, ValidateCode);
REGISTER_OP(THREADREMOVECODEENTRY, ThreadRemoveCodeEntry);
REGISTER_OP(CPUID, CPUID);
REGISTER_OP(XGETBV, XGETBV);
#undef REGISTER_OP
}
}
@@ -444,6 +444,11 @@ X86JITCore::X86JITCore(FEXCore::Context::ContextImpl *ctx, FEXCore::Core::Intern
Common.CPUIDFunction = PMF.GetConvertedPointer();
}
{
FEXCore::Utils::MemberFunctionToPointerCast PMF(&FEXCore::CPUIDEmu::RunXCRFunction);
Common.XCRFunction = PMF.GetConvertedPointer();
}
Common.SyscallHandlerObj = reinterpret_cast<uint64_t>(CTX->SyscallHandler);
Common.SyscallHandlerFunc = reinterpret_cast<uint64_t>(FEXCore::Context::HandleSyscall);
Common.ExitFunctionLink = reinterpret_cast<uintptr_t>(&Context::ContextImpl::ThreadExitFunctionLink<X86JITCore_ExitFunctionLink>);
@@ -313,6 +313,7 @@ private:
DEF_OP(ValidateCode);
DEF_OP(ThreadRemoveCodeEntry);
DEF_OP(CPUID);
DEF_OP(XGETBV);
///< Conversion ops
DEF_OP(VInsGPR);
@@ -22,6 +22,9 @@ namespace FEXCore::CodeSerialize {
// Multiblock enabled
unsigned MultiBlock : 1;
// Hardware TSO enabled
unsigned HardwareTSOEnabled : 1;
// TSO enabled
unsigned TSOEnabled : 1;
@@ -48,13 +51,14 @@ namespace FEXCore::CodeSerialize {
// Padding to remove uninitialized data warning from asan
// Shows remaining amount of bits available for config
unsigned _Pad : 18;
unsigned _Pad : 17;
bool operator==(CodeObjectSerializationConfig const &other) const {
return Cookie == other.Cookie &&
MaxInstPerBlock == other.MaxInstPerBlock &&
Arch == other.Arch &&
MultiBlock == other.MultiBlock &&
HardwareTSOEnabled == other.HardwareTSOEnabled &&
TSOEnabled == other.TSOEnabled &&
ABILocalFlags == other.ABILocalFlags &&
ABINoPF == other.ABINoPF &&
@@ -71,6 +75,7 @@ namespace FEXCore::CodeSerialize {
Hash <<= 32; Hash |= other.MaxInstPerBlock;
Hash <<= 1; Hash |= other.Arch;
Hash <<= 1; Hash |= other.MultiBlock;
Hash <<= 1; Hash |= other.HardwareTSOEnabled;
Hash <<= 1; Hash |= other.TSOEnabled;
Hash <<= 1; Hash |= other.ABILocalFlags;
Hash <<= 1; Hash |= other.ABINoPF;
+71 -25
View File
@@ -35,6 +35,7 @@ void OpDispatchBuilder::SyscallOp(OpcodeArgs) {
constexpr size_t SyscallArgs = 7;
using SyscallArray = std::array<uint64_t, SyscallArgs>;
size_t NumArguments{};
const SyscallArray *GPRIndexes {};
static constexpr SyscallArray GPRIndexes_64 = {
FEXCore::X86State::REG_RAX,
@@ -54,13 +55,26 @@ void OpDispatchBuilder::SyscallOp(OpcodeArgs) {
FEXCore::X86State::REG_RDI,
FEXCore::X86State::REG_RBP,
};
static_assert(GPRIndexes_64.size() == GPRIndexes_32.size());
static std::array<uint64_t, SyscallArgs> GPRIndexes_Hangover = {
static constexpr SyscallArray GPRIndexes_Hangover = {
FEXCore::X86State::REG_RCX,
};
size_t NumArguments{};
static constexpr SyscallArray GPRIndexes_Win64 = {
FEXCore::X86State::REG_RAX,
FEXCore::X86State::REG_R10,
FEXCore::X86State::REG_RDX,
FEXCore::X86State::REG_R8,
FEXCore::X86State::REG_R9,
FEXCore::X86State::REG_RSP,
};
static constexpr SyscallArray GPRIndexes_Win32 = {
FEXCore::X86State::REG_RAX,
FEXCore::X86State::REG_RSP,
};
SyscallFlags DefaultSyscallFlags = FEXCore::IR::SyscallFlags::DEFAULT;
const auto OSABI = CTX->SyscallHandler->GetOSABI();
if (OSABI == FEXCore::HLE::SyscallOSABI::OS_LINUX64) {
@@ -68,9 +82,19 @@ void OpDispatchBuilder::SyscallOp(OpcodeArgs) {
GPRIndexes = &GPRIndexes_64;
}
else if (OSABI == FEXCore::HLE::SyscallOSABI::OS_LINUX32) {
NumArguments = GPRIndexes_64.size();
NumArguments = GPRIndexes_32.size();
GPRIndexes = &GPRIndexes_32;
}
else if (OSABI == FEXCore::HLE::SyscallOSABI::OS_WIN64) {
NumArguments = 6;
GPRIndexes = &GPRIndexes_Win64;
DefaultSyscallFlags = FEXCore::IR::SyscallFlags::NORETURNEDRESULT;
}
else if (OSABI == FEXCore::HLE::SyscallOSABI::OS_WIN32) {
NumArguments = 2;
GPRIndexes = &GPRIndexes_Win32;
DefaultSyscallFlags = FEXCore::IR::SyscallFlags::NORETURNEDRESULT;
}
else if (OSABI == FEXCore::HLE::SyscallOSABI::OS_HANGOVER) {
NumArguments = 1;
GPRIndexes = &GPRIndexes_Hangover;
@@ -109,13 +133,20 @@ void OpDispatchBuilder::SyscallOp(OpcodeArgs) {
Arguments[4],
Arguments[5],
Arguments[6],
FEXCore::IR::SyscallFlags::DEFAULT);
DefaultSyscallFlags);
if (OSABI != FEXCore::HLE::SyscallOSABI::OS_HANGOVER) {
if (OSABI != FEXCore::HLE::SyscallOSABI::OS_HANGOVER &&
(DefaultSyscallFlags & FEXCore::IR::SyscallFlags::NORETURNEDRESULT) != FEXCore::IR::SyscallFlags::NORETURNEDRESULT) {
// Hangover doesn't want us returning a result here
// syscall is being abused as a thunk for now.
StoreGPRRegister(X86State::REG_RAX, SyscallOp);
}
if (Op->TableInfo->Flags & X86Tables::InstFlags::FLAGS_BLOCK_END) {
// RIP could have been updated after coming back from the Syscall.
NewRIP = _LoadContext(GPRSize, GPRClass, offsetof(FEXCore::Core::CPUState, rip));
_ExitFunction(NewRIP);
}
}
void OpDispatchBuilder::ThunkOp(OpcodeArgs) {
@@ -1513,7 +1544,7 @@ void OpDispatchBuilder::SAHFOp(OpcodeArgs) {
OrderedNode *Src = LoadGPRRegister(X86State::REG_RAX, 1, 8);
// Clear bits that aren't supposed to be set
Src = _And(Src, _Constant(~0b101000));
Src = _Andn(Src, _Constant(0b101000));
// Set the bit that is always set here
Src = _Or(Src, _Constant(0b10));
@@ -1748,6 +1779,18 @@ void OpDispatchBuilder::CPUIDOp(OpcodeArgs) {
StoreGPRRegister(X86State::REG_RCX, _Bfe(32, 0, Result_Upper));
}
void OpDispatchBuilder::XGetBVOp(OpcodeArgs) {
OrderedNode *Function = LoadGPRRegister(X86State::REG_RCX);
auto Res = _XGetBV(Function);
OrderedNode *Result_Lower = _ExtractElementPair(Res, 0);
OrderedNode *Result_Upper = _ExtractElementPair(Res, 1);
StoreGPRRegister(X86State::REG_RAX, Result_Lower);
StoreGPRRegister(X86State::REG_RDX, Result_Upper);
}
template<bool SHL1Bit>
void OpDispatchBuilder::SHLOp(OpcodeArgs) {
OrderedNode *Src{};
@@ -3807,7 +3850,7 @@ void OpDispatchBuilder::STOSOp(OpcodeArgs) {
OrderedNode *Counter = LoadGPRRegister(X86State::REG_RCX);
auto DF = GetRFLAG(FEXCore::X86State::RFLAG_DF_LOC);
auto Result = _MemSet(CTX->IsTSOEnabled(), Size, Segment ?: InvalidNode, Dest, Src, Counter, DF);
auto Result = _MemSet(CTX->IsAtomicTSOEnabled(), Size, Segment ?: InvalidNode, Dest, Src, Counter, DF);
StoreGPRRegister(X86State::REG_RCX, _Constant(0));
StoreGPRRegister(X86State::REG_RDI, Result);
}
@@ -3834,7 +3877,7 @@ void OpDispatchBuilder::MOVSOp(OpcodeArgs) {
auto DstSegment = GetSegment(0, FEXCore::X86Tables::DecodeFlags::FLAG_ES_PREFIX, true);
auto SrcSegment = GetSegment(Op->Flags, FEXCore::X86Tables::DecodeFlags::FLAG_DS_PREFIX);
auto Result = _MemCpy(CTX->IsTSOEnabled(), Size,
auto Result = _MemCpy(CTX->IsAtomicTSOEnabled(), Size,
DstSegment ?: InvalidNode,
SrcSegment ?: InvalidNode,
DstAddr, SrcAddr, Counter, DF);
@@ -5065,8 +5108,8 @@ void OpDispatchBuilder::StoreResult_WithOpSize(FEXCore::IR::RegisterClassType Cl
// TODO: Fix the instructions doing partial writes rather than dealing with it here.
auto SrcVector = LoadXMMRegister(gprIndex);
LOGMAN_THROW_AA_FMT(Class != IR::GPRClass, "Partial writes from GPR not allowed. Instruction: {}",
Op->TableInfo->Name);
LOGMAN_THROW_A_FMT(Class != IR::GPRClass, "Partial writes from GPR not allowed. Instruction: {}",
Op->TableInfo->Name);
// OpSize of 16 is special in that it is expected to zero the upper bits of the 256-bit operation.
// TODO: Longer term we should enforce the difference between zero and insert.
@@ -5309,7 +5352,6 @@ void OpDispatchBuilder::ALUOp(OpcodeArgs) {
ALUOpImpl(Op, ALUIROp, AtomicFetchOp, RequiresMask);
}
#ifndef _WIN32
void OpDispatchBuilder::INTOp(OpcodeArgs) {
IR::BreakDefinition Reason;
bool SetRIPToNext = false;
@@ -5318,14 +5360,19 @@ void OpDispatchBuilder::INTOp(OpcodeArgs) {
case 0xCD: { // INT imm8
uint8_t Literal = Op->Src[0].Data.Literal.Value;
if (Literal == 0x80) {
#ifndef _WIN32
constexpr uint8_t SYSCALL_LITERAL = 0x80;
#else
constexpr uint8_t SYSCALL_LITERAL = 0x2E;
#endif
if (Literal == SYSCALL_LITERAL) {
// Syscall on linux
SyscallOp(Op);
return;
}
Reason.ErrorRegister = Literal << 3 | (0b010);
Reason.Signal = SIGSEGV;
Reason.Signal = Core::FAULT_SIGSEGV;
// GP is raised when task-gate isn't setup to be valid
Reason.TrapNumber = X86State::X86_TRAPNO_GP;
Reason.si_code = 0x80;
@@ -5333,33 +5380,33 @@ void OpDispatchBuilder::INTOp(OpcodeArgs) {
}
case 0xCE: // INTO
Reason.ErrorRegister = 0;
Reason.Signal = SIGSEGV;
Reason.Signal = Core::FAULT_SIGSEGV;
Reason.TrapNumber = X86State::X86_TRAPNO_OF;
Reason.si_code = 0x80;
break;
case 0xF1: // INT1
Reason.ErrorRegister = 0;
Reason.Signal = SIGTRAP;
Reason.Signal = Core::FAULT_SIGTRAP;
Reason.TrapNumber = X86State::X86_TRAPNO_DB;
Reason.si_code = 1;
SetRIPToNext = true;
break;
case 0xF4: { // HLT
Reason.ErrorRegister = 0;
Reason.Signal = SIGSEGV;
Reason.Signal = Core::FAULT_SIGSEGV;
Reason.TrapNumber = X86State::X86_TRAPNO_GP;
Reason.si_code = 0x80;
break;
}
case 0x0B: // UD2
Reason.ErrorRegister = 0;
Reason.Signal = SIGILL;
Reason.Signal = Core::FAULT_SIGILL;
Reason.TrapNumber = X86State::X86_TRAPNO_UD;
Reason.si_code = 2;
break;
case 0xCC: // INT3
Reason.ErrorRegister = 0;
Reason.Signal = SIGTRAP;
Reason.Signal = Core::FAULT_SIGTRAP;
Reason.TrapNumber = X86State::X86_TRAPNO_BP;
Reason.si_code = 0x80;
SetRIPToNext = true;
@@ -5406,11 +5453,6 @@ void OpDispatchBuilder::INTOp(OpcodeArgs) {
_Break(Reason);
}
}
#else
void OpDispatchBuilder::INTOp(OpcodeArgs) {
ERROR_AND_DIE_FMT("Unknown INTOp instruction?");
}
#endif
void OpDispatchBuilder::TZCNT(OpcodeArgs) {
OrderedNode *Src = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags, -1);
@@ -5992,7 +6034,9 @@ void OpDispatchBuilder::InstallHostSpecificOpcodeHandlers() {
{OPD(3, 0b01, 0x4B), 1, &OpDispatchBuilder::AVXVectorVariableBlend<8>},
{OPD(3, 0b01, 0x4C), 1, &OpDispatchBuilder::AVXVectorVariableBlend<1>},
{OPD(3, 0b01, 0x60), 1, &OpDispatchBuilder::VPCMPESTRMOp},
{OPD(3, 0b01, 0x61), 1, &OpDispatchBuilder::VPCMPESTRIOp},
{OPD(3, 0b01, 0x62), 1, &OpDispatchBuilder::VPCMPISTRMOp},
{OPD(3, 0b01, 0x63), 1, &OpDispatchBuilder::VPCMPISTRIOp},
{OPD(3, 0b01, 0xDF), 1, &OpDispatchBuilder::VAESKeyGenAssistOp},
@@ -6704,7 +6748,7 @@ constexpr uint16_t PF_F2 = 3;
constexpr std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> SecondaryModRMExtensionOpTable[] = {
// REG /2
{((1 << 3) | 0), 1, &OpDispatchBuilder::UnimplementedOp},
{((1 << 3) | 0), 1, &OpDispatchBuilder::XGetBVOp},
// REG /7
{((3 << 3) | 1), 1, &OpDispatchBuilder::RDTSCPOp},
@@ -7283,7 +7327,9 @@ constexpr uint16_t PF_F2 = 3;
{OPD(0, PF_3A_66, 0x41), 1, &OpDispatchBuilder::DPPOp<8>},
{OPD(0, PF_3A_66, 0x42), 1, &OpDispatchBuilder::MPSADBWOp},
{OPD(0, PF_3A_66, 0x60), 1, &OpDispatchBuilder::VPCMPESTRMOp},
{OPD(0, PF_3A_66, 0x61), 1, &OpDispatchBuilder::VPCMPESTRIOp},
{OPD(0, PF_3A_66, 0x62), 1, &OpDispatchBuilder::VPCMPISTRMOp},
{OPD(0, PF_3A_66, 0x63), 1, &OpDispatchBuilder::VPCMPISTRIOp},
{OPD(0, PF_3A_NONE, 0xCC), 1, &OpDispatchBuilder::SHA1RNDS4Op},
+6 -3
View File
@@ -219,6 +219,7 @@ public:
void MOVOffsetOp(OpcodeArgs);
void CMOVOp(OpcodeArgs);
void CPUIDOp(OpcodeArgs);
void XGetBVOp(OpcodeArgs);
template<bool SHL1Bit>
void SHLOp(OpcodeArgs);
void SHLImmediateOp(OpcodeArgs);
@@ -491,7 +492,9 @@ public:
void VPALIGNROp(OpcodeArgs);
void VPCMPESTRIOp(OpcodeArgs);
void VPCMPESTRMOp(OpcodeArgs);
void VPCMPISTRIOp(OpcodeArgs);
void VPCMPISTRMOp(OpcodeArgs);
void VPERM2Op(OpcodeArgs);
void VPERMDOp(OpcodeArgs);
@@ -862,7 +865,7 @@ private:
const X86Tables::DecodedOperand& Src2,
const X86Tables::DecodedOperand& Imm);
void PCMPXSTRXOpImpl(OpcodeArgs, bool IsExplicit);
void PCMPXSTRXOpImpl(OpcodeArgs, bool IsExplicit, bool IsMask);
OrderedNode* PHADDSOpImpl(OpcodeArgs, const X86Tables::DecodedOperand& Src1,
const X86Tables::DecodedOperand& Src2);
@@ -1555,14 +1558,14 @@ private:
uint64_t Entry;
OrderedNode* _StoreMemAutoTSO(FEXCore::IR::RegisterClassType Class, uint8_t Size, OrderedNode *Addr, OrderedNode *Value, uint8_t Align = 1) {
if (CTX->IsTSOEnabled())
if (CTX->IsAtomicTSOEnabled())
return _StoreMemTSO(Class, Size, Value, Addr, Invalid(), Align, MEM_OFFSET_SXTX, 1);
else
return _StoreMem(Class, Size, Value, Addr, Invalid(), Align, MEM_OFFSET_SXTX, 1);
}
OrderedNode* _LoadMemAutoTSO(FEXCore::IR::RegisterClassType Class, uint8_t Size, OrderedNode *ssa0, uint8_t Align = 1) {
if (CTX->IsTSOEnabled())
if (CTX->IsAtomicTSOEnabled())
return _LoadMemTSO(Class, Size, ssa0, Invalid(), Align, MEM_OFFSET_SXTX, 1);
else
return _LoadMem(Class, Size, ssa0, Invalid(), Align, MEM_OFFSET_SXTX, 1);
@@ -52,10 +52,9 @@ void OpDispatchBuilder::SetPackedRFLAG(bool Lower8, OrderedNode *Src) {
InvalidateDeferredFlags();
}
auto OneConst = _Constant(1);
for (size_t i = 0; i < NumFlags; ++i) {
const auto FlagOffset = FlagOffsets[i];
auto Tmp = _And(_Lshr(Src, _Constant(FlagOffset)), OneConst);
auto Tmp = _Bfe(4, 1, FlagOffset, Src);
SetRFLAG(Tmp, FlagOffset);
}
}
@@ -304,28 +303,11 @@ void OpDispatchBuilder::CalculcateFlags_ADC(uint8_t SrcSize, OrderedNode *Res, O
// OF
// Signed
{
auto NegOne = _Constant(~0ULL);
auto XorOp1 = _Xor(_Xor(Src1, Src2), NegOne);
auto XorOp1 = _Not(_Xor(Src1, Src2));
auto XorOp2 = _Xor(Res, Src1);
OrderedNode *AndOp1 = _And(XorOp1, XorOp2);
switch (Size) {
case 8:
AndOp1 = _Bfe(1, 7, AndOp1);
break;
case 16:
AndOp1 = _Bfe(1, 15, AndOp1);
break;
case 32:
AndOp1 = _Bfe(1, 31, AndOp1);
break;
case 64:
AndOp1 = _Bfe(1, 63, AndOp1);
break;
default:
LOGMAN_MSG_A_FMT("Unknown BFE size: {}", Size);
break;
}
AndOp1 = _Bfe(1, Size - 1, AndOp1);
SetRFLAG<FEXCore::X86State::RFLAG_OF_LOC>(AndOp1);
}
}
@@ -376,24 +358,7 @@ void OpDispatchBuilder::CalculcateFlags_SBB(uint8_t SrcSize, OrderedNode *Res, O
auto XorOp1 = _Xor(Src1, Src2);
auto XorOp2 = _Xor(Res, Src1);
OrderedNode *AndOp1 = _And(XorOp1, XorOp2);
switch (SrcSize) {
case 1:
AndOp1 = _Bfe(1, 7, AndOp1);
break;
case 2:
AndOp1 = _Bfe(1, 15, AndOp1);
break;
case 4:
AndOp1 = _Bfe(1, 31, AndOp1);
break;
case 8:
AndOp1 = _Bfe(1, 63, AndOp1);
break;
default:
LOGMAN_MSG_A_FMT("Unknown BFE size: {}", SrcSize);
break;
}
AndOp1 = _Bfe(1, SrcSize * 8 - 1, AndOp1);
SetRFLAG<FEXCore::X86State::RFLAG_OF_LOC>(AndOp1);
}
}
@@ -492,29 +457,12 @@ void OpDispatchBuilder::CalculcateFlags_ADD(uint8_t SrcSize, OrderedNode *Res, O
// OF
{
auto NegOne = _Constant(~0ULL);
auto XorOp1 = _Xor(_Xor(Src1, Src2), NegOne);
auto XorOp1 = _Not(_Xor(Src1, Src2));
auto XorOp2 = _Xor(Res, Src1);
OrderedNode *AndOp1 = _And(XorOp1, XorOp2);
switch (SrcSize) {
case 1:
AndOp1 = _Bfe(1, 7, AndOp1);
break;
case 2:
AndOp1 = _Bfe(1, 15, AndOp1);
break;
case 4:
AndOp1 = _Bfe(1, 31, AndOp1);
break;
case 8:
AndOp1 = _Bfe(1, 63, AndOp1);
break;
default:
LOGMAN_MSG_A_FMT("Unknown BFE size: {}", SrcSize);
break;
}
AndOp1 = _Bfe(1, SrcSize * 8 - 1, AndOp1);
SetRFLAG<FEXCore::X86State::RFLAG_OF_LOC>(AndOp1);
}
}
@@ -4591,7 +4591,7 @@ void OpDispatchBuilder::VPERMILRegOp<4>(OpcodeArgs);
template
void OpDispatchBuilder::VPERMILRegOp<8>(OpcodeArgs);
void OpDispatchBuilder::PCMPXSTRXOpImpl(OpcodeArgs, bool IsExplicit) {
void OpDispatchBuilder::PCMPXSTRXOpImpl(OpcodeArgs, bool IsExplicit, bool IsMask) {
LOGMAN_THROW_A_FMT(Op->Src[1].IsLiteral(), "Src[1] needs to be a literal");
const auto Control = Op->Src[1].Data.Literal.Value;
@@ -4620,23 +4620,53 @@ void OpDispatchBuilder::PCMPXSTRXOpImpl(OpcodeArgs, bool IsExplicit) {
} else {
IntermediateResult = _VPCMPISTRX(Src1, Src2, Control);
}
OrderedNode *ResultNoFlags = _And(IntermediateResult, _Constant(0xFFFF));
// For the indexed variant of the instructions, if control[6] is set, then we
// store the index of the most significant bit into ECX. If it's not set,
// then we store the least significant bit.
OrderedNode *ZeroConst = _Constant(0);
const auto ECXResult = [&]() -> OrderedNode* {
if (IsMask) {
// For the masked variant of the instructions, if control[6] is set, then we
// need to expand the intermediate result into a byte or word mask (depending
// on data size specified in control[1]) along the entire length of XMM0,
// where set bits in the intermediate result set the corresponding entry
// in XMM0 to all 1s and unset bits set the corresponding entry to all 0s.
//
// If control[6] is not set, then we just store the intermediate result as-is
// into the least significant bits of XMM0 and zero extend it.
const auto IsExpandedMask = (Control & 0b0100'0000) != 0;
if (IsExpandedMask) {
// We need to iterate over the intermediate result and
// expand the mask into XMM0 elements.
const auto ElementSize = 1U << (Control & 1);
const auto NumElements = 16U >> (Control & 1);
OrderedNode *Result = _VectorZero(Core::CPUState::XMM_SSE_REG_SIZE);
for (uint32_t i = 0; i < NumElements; i++) {
OrderedNode *SignBit = _Sbfe(1, i, IntermediateResult);
Result = _VInsGPR(Core::CPUState::XMM_SSE_REG_SIZE, ElementSize, i, Result, SignBit);
}
StoreXMMRegister(0, Result);
} else {
// We insert the intermediate result as-is.
StoreXMMRegister(0, _VCastFromGPR(16, 2, IntermediateResult));
}
} else {
// For the indexed variant of the instructions, if control[6] is set, then we
// store the index of the most significant bit into ECX. If it's not set,
// then we store the least significant bit.
const auto UseMSBIndex = (Control & 0b0100'0000) != 0;
OrderedNode *ResultNoFlags = _Bfe(16, 0, IntermediateResult);
OrderedNode *IfZero = _Constant(16 >> (Control & 1));
OrderedNode *IfNotZero = UseMSBIndex ? _FindMSB(ResultNoFlags)
: _FindLSB(ResultNoFlags);
return _Select(IR::COND_EQ, ResultNoFlags, ZeroConst,
IfZero, IfNotZero);
}();
OrderedNode *Result = _Select(IR::COND_EQ, ResultNoFlags, ZeroConst,
IfZero, IfNotZero);
StoreGPRRegister(X86State::REG_RCX, Result, 4);
}
// Set all of the necessary flags.
// We use the top 16-bits of the result to store the flags
@@ -4656,16 +4686,19 @@ void OpDispatchBuilder::PCMPXSTRXOpImpl(OpcodeArgs, bool IsExplicit) {
SetRFLAG<X86State::RFLAG_AF_LOC>(ZeroConst);
SetRFLAG<X86State::RFLAG_PF_LOC>(ZeroConst);
// ... and we're done!
StoreGPRRegister(X86State::REG_RCX, ECXResult, 4);
}
void OpDispatchBuilder::VPCMPESTRIOp(OpcodeArgs) {
PCMPXSTRXOpImpl(Op, true);
PCMPXSTRXOpImpl(Op, true, false);
}
void OpDispatchBuilder::VPCMPESTRMOp(OpcodeArgs) {
PCMPXSTRXOpImpl(Op, true, true);
}
void OpDispatchBuilder::VPCMPISTRIOp(OpcodeArgs) {
PCMPXSTRXOpImpl(Op, false);
PCMPXSTRXOpImpl(Op, false, false);
}
void OpDispatchBuilder::VPCMPISTRMOp(OpcodeArgs) {
PCMPXSTRXOpImpl(Op, false, true);
}
} // namespace FEXCore::IR
@@ -1374,8 +1374,7 @@ void OpDispatchBuilder::X87FCMOV(OpcodeArgs) {
SrcCond = _Sbfe(1, 0, SrcCond);
OrderedNode *VecCond = _VCastFromGPR(16, 8, SrcCond);
VecCond = _VInsGPR(16, 8, 1, VecCond, SrcCond);
OrderedNode *VecCond = _VDupFromGPR(16, 8, SrcCond);
auto top = GetX87Top();
OrderedNode* arg;
@@ -169,7 +169,7 @@ void InitializeBaseTables(Context::OperatingMode Mode) {
{0xC9, 1, X86InstInfo{"LEAVE", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_DEBUG_MEM_ACCESS , 0, nullptr}},
{0xCA, 2, X86InstInfo{"RETF", TYPE_PRIV, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_BLOCK_END, 0, nullptr}},
{0xCC, 1, X86InstInfo{"INT3", TYPE_INST, FLAGS_DEBUG, 0, nullptr}},
{0xCD, 1, X86InstInfo{"INT", TYPE_INST, FLAGS_DEBUG , 1, nullptr}},
{0xCD, 1, X86InstInfo{"INT", TYPE_INST, DEFAULT_SYSCALL_FLAGS, 1, nullptr}},
{0xCF, 1, X86InstInfo{"IRET", TYPE_INST, FLAGS_SETS_RIP | FLAGS_BLOCK_END, 0, nullptr}},
{0xD7, 1, X86InstInfo{"XLAT", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS, 0, nullptr}},
@@ -43,9 +43,9 @@ void InitializeH0F3ATables(Context::OperatingMode Mode) {
{OPD(0, PF_3A_66, 0x42), 1, X86InstInfo{"MPSADBW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x44), 1, X86InstInfo{"PCLMULQDQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x60), 1, X86InstInfo{"PCMPESTRM", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(0, PF_3A_66, 0x60), 1, X86InstInfo{"PCMPESTRM", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x61), 1, X86InstInfo{"PCMPESTRI", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x62), 1, X86InstInfo{"PCMPISTRM", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(0, PF_3A_66, 0x62), 1, X86InstInfo{"PCMPISTRM", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x63), 1, X86InstInfo{"PCMPISTRI", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_NONE, 0xCC), 1, X86InstInfo{"SHA1RNDS4", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
@@ -22,7 +22,7 @@ void InitializeSecondaryTables(Context::OperatingMode Mode) {
{0x02, 1, X86InstInfo{"LAR", TYPE_UNDEC, FLAGS_NO_OVERLAY, 0, nullptr}},
{0x03, 1, X86InstInfo{"LSL", TYPE_UNDEC, FLAGS_NO_OVERLAY, 0, nullptr}},
{0x04, 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NO_OVERLAY, 0, nullptr}},
{0x05, 1, X86InstInfo{"SYSCALL", TYPE_INST, FLAGS_NO_OVERLAY, 0, nullptr}},
{0x05, 1, X86InstInfo{"SYSCALL", TYPE_INST, DEFAULT_SYSCALL_FLAGS, 0, nullptr}},
{0x06, 1, X86InstInfo{"CLTS", TYPE_PRIV, FLAGS_NO_OVERLAY, 0, nullptr}},
{0x07, 1, X86InstInfo{"SYSRET", TYPE_PRIV, FLAGS_NO_OVERLAY, 0, nullptr}},
{0x08, 1, X86InstInfo{"INVD", TYPE_PRIV, FLAGS_NO_OVERLAY, 0, nullptr}},
@@ -456,9 +456,9 @@ void InitializeVEXTables() {
{OPD(3, 0b01, 0x5E), 1, X86InstInfo{"VMFSUBADDPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(3, 0b01, 0x5F), 1, X86InstInfo{"VFMSUBADDPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(3, 0b01, 0x60), 1, X86InstInfo{"VPCMPESTRM", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(3, 0b01, 0x60), 1, X86InstInfo{"VPCMPESTRM", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x61), 1, X86InstInfo{"VPCMPESTRI", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x62), 1, X86InstInfo{"VPCMPISTRM", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(3, 0b01, 0x62), 1, X86InstInfo{"VPCMPISTRM", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x63), 1, X86InstInfo{"VPCMPISTRI", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(3, 0b01, 0x68), 1, X86InstInfo{"VFMADDPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
@@ -361,6 +361,13 @@ constexpr InstFlagType SIZE_128BIT = 0b101;
constexpr InstFlagType SIZE_256BIT = 0b110;
constexpr InstFlagType SIZE_64BITDEF = 0b111; // Default mode is 64bit instead of typical 32bit
#ifndef _WIN32
constexpr uint32_t DEFAULT_SYSCALL_FLAGS = FLAGS_NO_OVERLAY;
#else
// Syscall ends a block on WIN32 because the instruction can update the CPU's RIP.
constexpr uint32_t DEFAULT_SYSCALL_FLAGS = FLAGS_NO_OVERLAY | FLAGS_BLOCK_END;
#endif
constexpr InstFlagType GetSizeDstFlags(InstFlagType Flags) { return (Flags >> FLAGS_SIZE_DST_OFF) & SIZE_MASK; }
constexpr InstFlagType GetSizeSrcFlags(InstFlagType Flags) { return (Flags >> FLAGS_SIZE_SRC_OFF) & SIZE_MASK; }
+2 -1
View File
@@ -53,8 +53,9 @@ static __attribute__((aligned(16), naked, section("HostToGuestTrampolineTemplate
);
#elif defined(_M_ARM_64)
asm(
// x11 is part of the custom ABI and needs to point to the TrampolineInstanceInfo.
"ldr x16, 0f \n"
"adr x11, 0f \n"
"ldr x16, [x11] \n"
"br x16 \n"
// Manually align to the next 8-byte boundary
// NOTE: GCC over-aligns to a full page when using .align directives on ARM (last tested on GCC 11.2)
+3 -3
View File
@@ -251,8 +251,8 @@ namespace FEXCore::IR {
}
}
AOTIRCaptureCache::PreGenerateIRFetchResult AOTIRCaptureCache::PreGenerateIRFetch(uint64_t GuestRIP, FEXCore::IR::IRListView *IRList) {
auto AOTIRCacheEntry = CTX->SyscallHandler->LookupAOTIRCacheEntry(GuestRIP);
AOTIRCaptureCache::PreGenerateIRFetchResult AOTIRCaptureCache::PreGenerateIRFetch(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP, FEXCore::IR::IRListView *IRList) {
auto AOTIRCacheEntry = CTX->SyscallHandler->LookupAOTIRCacheEntry(Thread, GuestRIP);
PreGenerateIRFetchResult Result{};
@@ -306,7 +306,7 @@ namespace FEXCore::IR {
// Both generated ir and LibraryJITName need a named region lookup
if (GeneratedIR || CTX->Config.LibraryJITNaming() || CTX->Config.GDBSymbols()) {
auto AOTIRCacheEntry = CTX->SyscallHandler->LookupAOTIRCacheEntry(GuestRIP);
auto AOTIRCacheEntry = CTX->SyscallHandler->LookupAOTIRCacheEntry(Thread, GuestRIP);
if (AOTIRCacheEntry.Entry) {
if (DebugData && CTX->Config.LibraryJITNaming()) {
+1 -1
View File
@@ -105,7 +105,7 @@ namespace FEXCore::IR {
uint64_t Length {};
bool GeneratedIR {};
};
[[nodiscard]] PreGenerateIRFetchResult PreGenerateIRFetch(uint64_t GuestRIP, FEXCore::IR::IRListView *IRList);
[[nodiscard]] PreGenerateIRFetchResult PreGenerateIRFetch(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP, FEXCore::IR::IRListView *IRList);
bool PostCompileCode(FEXCore::Core::InternalThreadState *Thread,
void* CodePtr,
+9 -1
View File
@@ -297,6 +297,13 @@
],
"DestSize": "16",
"NumElements": "2"
},
"GPRPair = XGetBV GPR:$Function": {
"Desc": ["Calls in to the XCR handler function to return emulated XCR",
"Returns a 64bit GPR pair that fits emulated EAX, EDX respectively"
],
"DestSize": "8",
"NumElements": "2"
}
},
"Moves": {
@@ -714,7 +721,8 @@
"Desc": ["Integer binary not",
"op:",
"Dest = ~Src"
]
],
"DestSize": "std::max<uint8_t>(4, GetOpSize(_Src))"
},
"GPR = Popcount GPR:$Src": {
"Desc": ["Population count of source register",
+1 -1
View File
@@ -103,7 +103,7 @@ static void PrintArg(fextl::stringstream *out, IRListView const* IR, OrderedNode
if (ArgID.IsInvalid()) {
*out << "%Invalid";
} else {
*out << "%ssa" << ArgID;
*out << "%ssa" << std::dec << ArgID;
if (RAData) {
auto PhyReg = RAData->GetNodeRegister(ArgID);
@@ -314,6 +314,36 @@ namespace {
FEXCore::IR::InvalidClass,
});
// _pad2
ContextClassification->emplace_back(ContextMemberInfo {
ContextMemberClassification {
offsetof(FEXCore::Core::CPUState, _pad2),
sizeof(FEXCore::Core::CPUState::_pad2),
},
ACCESS_NONE,
FEXCore::IR::InvalidClass,
});
// DeferredSignalRefCount
ContextClassification->emplace_back(ContextMemberInfo {
ContextMemberClassification {
offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount),
sizeof(FEXCore::Core::CPUState::DeferredSignalRefCount),
},
ACCESS_NONE,
FEXCore::IR::InvalidClass,
});
// DeferredSignalFaultAddress
ContextClassification->emplace_back(ContextMemberInfo {
ContextMemberClassification {
offsetof(FEXCore::Core::CPUState, DeferredSignalFaultAddress),
sizeof(FEXCore::Core::CPUState::DeferredSignalFaultAddress),
},
ACCESS_NONE,
FEXCore::IR::InvalidClass,
});
[[maybe_unused]] size_t ClassifiedStructSize{};
ContextClassificationInfo->Lookup.reserve(sizeof(FEXCore::Core::CPUState));
@@ -393,6 +423,10 @@ namespace {
SetAccess(Offset++, ACCESS_NONE);
SetAccess(Offset++, ACCESS_NONE);
SetAccess(Offset++, ACCESS_INVALID);
SetAccess(Offset++, ACCESS_INVALID);
SetAccess(Offset++, ACCESS_INVALID);
}
struct BlockInfo {
+14 -4
View File
@@ -5,7 +5,7 @@
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/MathUtils.h>
#include <FEXCore/fextl/sstream.h>
#include <FEXHeaderUtils/ScopedSignalMask.h>
#include <FEXCore/Utils/DeferredSignalMutex.h>
#include <FEXHeaderUtils/Syscalls.h>
#include <FEXHeaderUtils/TypeDefines.h>
#include <FEXCore/fextl/memory.h>
@@ -29,6 +29,16 @@
namespace Alloc::OSAllocator {
thread_local FEXCore::Core::InternalThreadState *TLSThread{};
void RegisterTLSData(FEXCore::Core::InternalThreadState *Thread) {
TLSThread = Thread;
}
void UninstallTLSData(FEXCore::Core::InternalThreadState *Thread) {
TLSThread = nullptr;
}
class OSAllocator_64Bit final : public Alloc::HostAllocator {
public:
OSAllocator_64Bit();
@@ -248,7 +258,7 @@ void *OSAllocator_64Bit::Mmap(void *addr, size_t length, int prot, int flags, in
size_t NumberOfPages = length / FHU::FEX_PAGE_SIZE;
// This needs a mutex to be thread safe
FHU::ScopedSignalMaskWithMutex lk(AllocationMutex);
FEXCore::ScopedPotentialDeferredSignalWithMutex lk(AllocationMutex, TLSThread);
uint64_t AllocatedOffset{};
LiveVMARegion *LiveRegion{};
@@ -436,7 +446,7 @@ int OSAllocator_64Bit::Munmap(void *addr, size_t length) {
}
// This needs a mutex to be thread safe
FHU::ScopedSignalMaskWithMutex lk(AllocationMutex);
FEXCore::ScopedPotentialDeferredSignalWithMutex lk(AllocationMutex, TLSThread);
length = FEXCore::AlignUp(length, FHU::FEX_PAGE_SIZE);
@@ -561,7 +571,7 @@ OSAllocator_64Bit::OSAllocator_64Bit() {
OSAllocator_64Bit::~OSAllocator_64Bit() {
// This needs a mutex to be thread safe
FHU::ScopedSignalMaskWithMutex lk(AllocationMutex);
FEXCore::ScopedPotentialDeferredSignalWithMutex lk(AllocationMutex, TLSThread);
// Walk the pages and deallocate
// First walk the live regions
@@ -6,6 +6,10 @@
#include <cstdint>
#include <sys/types.h>
namespace FEXCore::Core {
struct InternalThreadState;
}
namespace Alloc {
// HostAllocator is just a page pased slab allocator
// Similar to mmap and munmap only mapping at the page level
@@ -36,5 +40,7 @@ namespace Alloc {
}
namespace Alloc::OSAllocator {
void RegisterTLSData(FEXCore::Core::InternalThreadState *Thread);
void UninstallTLSData(FEXCore::Core::InternalThreadState *Thread);
fextl::unique_ptr<Alloc::HostAllocator> Create64BitAllocator();
}
@@ -22,6 +22,11 @@ bool HandleCASAL(void *_ucontext, void *_info, uint32_t Instr) {
bool HandleAtomicMemOp(void *_ucontext, void *_info, uint32_t Instr) {
ERROR_AND_DIE_FMT("HandleAtomicMemOp Not Implemented");
}
std::pair<bool, int32_t> HandleUnalignedAccess(bool ParanoidTSO, uintptr_t ProgramCounter, uint64_t *GPRs) {
ERROR_AND_DIE_FMT("HandleAtomicMemOp Not Implemented");
}
#endif
}
+9 -9
View File
@@ -1,4 +1,5 @@
#include <FEXCore/Config/Config.h>
#include <FEXCore/Utils/File.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/Telemetry.h>
#include <FEXCore/fextl/fmt.h>
@@ -22,6 +23,7 @@ namespace FEXCore::Telemetry {
"32bit CAS Tear",
"64bit CAS Tear",
"128bit CAS Tear",
"Crash mask",
};
void Initialize() {
auto DataDirectory = Config::GetDataDirectory();
@@ -35,7 +37,6 @@ namespace FEXCore::Telemetry {
}
void Shutdown(fextl::string const &ApplicationName) {
#ifndef _WIN32
auto DataDirectory = Config::GetDataDirectory();
DataDirectory += "Telemetry/" + ApplicationName + ".telem";
@@ -45,20 +46,19 @@ namespace FEXCore::Telemetry {
FHU::Filesystem::CopyFile(DataDirectory, Backup, FHU::Filesystem::CopyOptions::OVERWRITE_EXISTING);
}
constexpr int USER_PERMS = S_IRWXU | S_IRWXG | S_IRWXO;
int fd = open(DataDirectory.c_str(), O_CREAT | O_WRONLY | O_TRUNC | O_CLOEXEC, USER_PERMS);
auto File = FEXCore::File::File(DataDirectory.c_str(),
FEXCore::File::FileModes::WRITE |
FEXCore::File::FileModes::CREATE |
FEXCore::File::FileModes::TRUNCATE);
if (fd != -1) {
if (File.IsValid()) {
for (size_t i = 0; i < TelemetryType::TYPE_LAST; ++i) {
auto &Name = TelemetryNames.at(i);
auto &Data = TelemetryValues.at(i);
auto Output = fextl::fmt::format("{}: {}\n", Name, *Data);
write(fd, Output.c_str(), Output.size());
fextl::fmt::print(File, "{}: {}\n", Name, *Data);
}
fsync(fd);
close(fd);
File.Flush();
}
#endif
}
Value &GetObject(TelemetryType Type) {
+4
View File
@@ -5,5 +5,9 @@ namespace FEXCore::CPUID {
struct FunctionResults {
uint32_t eax, ebx, ecx, edx;
};
struct XCRResults {
uint32_t eax, edx;
};
}
+19 -2
View File
@@ -259,6 +259,7 @@ namespace FEXCore::Context {
FEX_DEFAULT_VISIBILITY virtual void SetSyscallHandler(FEXCore::HLE::SyscallHandler *Handler) = 0;
FEX_DEFAULT_VISIBILITY virtual FEXCore::CPUID::FunctionResults RunCPUIDFunction(uint32_t Function, uint32_t Leaf) = 0;
FEX_DEFAULT_VISIBILITY virtual FEXCore::CPUID::XCRResults RunXCRFunction(uint32_t Function) = 0;
FEX_DEFAULT_VISIBILITY virtual FEXCore::CPUID::FunctionResults RunCPUIDFunctionName(uint32_t Function, uint32_t Leaf, uint32_t CPU) = 0;
FEX_DEFAULT_VISIBILITY virtual FEXCore::IR::AOTIRCacheEntry *LoadAOTIRCacheEntry(const fextl::string& Name) = 0;
@@ -270,8 +271,8 @@ namespace FEXCore::Context {
FEX_DEFAULT_VISIBILITY virtual void FinalizeAOTIRCache() = 0;
FEX_DEFAULT_VISIBILITY virtual void WriteFilesWithCode(std::function<void(const fextl::string& fileid, const fextl::string& filename)> Writer) = 0;
FEX_DEFAULT_VISIBILITY virtual void InvalidateGuestCodeRange(uint64_t Start, uint64_t Length) = 0;
FEX_DEFAULT_VISIBILITY virtual void InvalidateGuestCodeRange(uint64_t Start, uint64_t Length, std::function<void(uint64_t start, uint64_t Length)> callback) = 0;
FEX_DEFAULT_VISIBILITY virtual void InvalidateGuestCodeRange(FEXCore::Core::InternalThreadState *Thread, uint64_t Start, uint64_t Length) = 0;
FEX_DEFAULT_VISIBILITY virtual void InvalidateGuestCodeRange(FEXCore::Core::InternalThreadState *Thread, uint64_t Start, uint64_t Length, std::function<void(uint64_t start, uint64_t Length)> callback) = 0;
FEX_DEFAULT_VISIBILITY virtual void MarkMemoryShared() = 0;
FEX_DEFAULT_VISIBILITY virtual void ConfigureAOTGen(FEXCore::Core::InternalThreadState *Thread, fextl::set<uint64_t> *ExternalBranches, uint64_t SectionMaxAddress) = 0;
@@ -288,6 +289,22 @@ namespace FEXCore::Context {
FEX_DEFAULT_VISIBILITY virtual void GetVDSOSigReturn(VDSOSigReturn *VDSOPointers) = 0;
FEX_DEFAULT_VISIBILITY virtual void IncrementIdleRefCount() = 0;
/**
* @brief Informs the context if hardware TSO is supported.
* Once hardware TSO is enabled, then TSO emulation through atomics is disabled and relies on the hardware.
*
* @param HardwareTSOSupported If the hardware supports the TSO memory model or not.
*/
FEX_DEFAULT_VISIBILITY virtual void SetHardwareTSOSupport(bool HardwareTSOSupported) = 0;
/**
* @brief Enable exiting the JIT when HLT is hit.
*
* This is to workaround a bug in Wine's longjump function which breaks our unittests.
*
*/
FEX_DEFAULT_VISIBILITY virtual void EnableExitOnHLT() = 0;
private:
};
+87
View File
@@ -6,10 +6,69 @@
#include <atomic>
#include <cstddef>
#include <cstring>
#include <stdint.h>
#include <string_view>
#include <type_traits>
namespace FEXCore::Core {
// Wrapper around std::atomic using std::memory_order_relaxed.
// This allows compilers to emit more performant code at the expense of visibly tearing.
// In particular, increments/decrements may visibly tear if a signal is received half-way through.
//
// Prefer std::atomic with default memory ordering unless you really know what you're doing.
// Primarily this ensure program ordering when signals are concerned.
template<typename T>
class NonAtomicRefCounter {
public:
void Increment(T Value) {
// Specifically avoiding fetch_add here because that will turn in to ldxr+stxr or lock xadd.
// FEX very specifically wants to use simple loadstore instructions for this
//
// ARM64 ex:
// ldr x0, [x1];
// add x0, x0, #1;
// str x0, [x1];
//
// x86-64 ex:
// inc qword [rax];
auto Current = AtomicVariable.load(std::memory_order_relaxed);
AtomicVariable.store(Current + Value, std::memory_order_relaxed);
}
// Returns original value.
// x86-64 needs to know the result on decrement.
T Decrement(T Value) {
// Specifically avoiding fetch_sub here because that will turn into ldxr+stxr or lock xadd.
// FEX very specifically wants to use simple loadstore instructions for this
//
// ARM64 ex:
// ldr x0, [x1];
// sub x0, x0, #1;
// str x0, [x1];
//
// x86-64 ex:
// dec qword [rax];
auto Current = AtomicVariable.load(std::memory_order_relaxed);
AtomicVariable.store(Current - Value, std::memory_order_relaxed);
return Current;
}
T Load() const {
return AtomicVariable.load(std::memory_order_relaxed);
}
void Store(T Value) {
AtomicVariable.store(Value, std::memory_order_relaxed);
}
private:
std::atomic<T> AtomicVariable;
};
static_assert(std::is_standard_layout_v<NonAtomicRefCounter<uint64_t>>, "Needs to be standard layout");
static_assert(std::is_trivially_copyable_v<NonAtomicRefCounter<uint64_t>>, "needs to be trivially copyable");
static_assert(sizeof(NonAtomicRefCounter<uint64_t>) == sizeof(uint64_t), "Needs to be correct size");
struct FEX_PACKED CPUState {
// Allows more efficient handling of the register
// file in the event AVX is not supported.
@@ -49,6 +108,13 @@ namespace FEXCore::Core {
uint16_t FCW;
uint16_t FTW;
uint32_t _pad2[1];
// Reference counter for FEX's per-thread deferred signals.
// Counts the nesting depth of program sections that cause signals to be deferred.
NonAtomicRefCounter<uint64_t> DeferredSignalRefCount;
// Since this memory region is thread local, we use NonAtomicRefCounter for fast atomic access.
NonAtomicRefCounter<uint64_t> *DeferredSignalFaultAddress;
static constexpr size_t FLAG_SIZE = sizeof(flags[0]);
static constexpr size_t GDT_SIZE = sizeof(gdt[0]);
static constexpr size_t GPR_REG_SIZE = sizeof(gregs[0]);
@@ -63,9 +129,29 @@ namespace FEXCore::Core {
static constexpr size_t NUM_GPRS = sizeof(gregs) / GPR_REG_SIZE;
static constexpr size_t NUM_XMMS = sizeof(xmm) / XMM_AVX_REG_SIZE;
static constexpr size_t NUM_MMS = sizeof(mm) / MM_REG_SIZE;
CPUState() {
// Initialize default CPU state
rip = ~0ULL;
memset(gregs, 0, sizeof(gregs));
for (auto& xmm : xmm.avx.data) {
xmm[0] = 0xDEADBEEFULL;
xmm[1] = 0xBAD0DAD1ULL;
xmm[2] = 0xDEADCAFEULL;
xmm[3] = 0xBAD2CAD3ULL;
}
memset(&flags, 0, Core::CPUState::NUM_EFLAG_BITS);
flags[1] = 1; ///< Reserved - Always 1.
flags[9] = 1; ///< Interrupt flag - Always 1.
FCW = 0x37F;
FTW = 0xFFFF;
}
};
static_assert(std::is_trivially_copyable_v<CPUState>, "Needs to be trivial");
static_assert(std::is_standard_layout_v<CPUState>, "This needs to be standard layout");
static_assert(offsetof(CPUState, xmm) % 32 == 0, "xmm needs to be 256-bit aligned!");
static_assert(offsetof(CPUState, mm) % 16 == 0, "mm needs to be 128-bit aligned!");
static_assert(offsetof(CPUState, DeferredSignalRefCount) % 8 == 0, "Needs to be 8-byte aligned");
struct InternalThreadState;
@@ -144,6 +230,7 @@ namespace FEXCore::Core {
uint64_t ThreadRemoveCodeEntryFromJIT{};
uint64_t CPUIDObj{};
uint64_t CPUIDFunction{};
uint64_t XCRFunction{};
uint64_t SyscallHandlerObj{};
uint64_t SyscallHandlerFunc{};
uint64_t ExitFunctionLink{};
+12
View File
@@ -21,6 +21,18 @@ namespace Core {
Return,
ReturnRT,
};
enum SignalNumber {
#ifndef _WIN32
FAULT_SIGSEGV = SIGSEGV,
FAULT_SIGTRAP = SIGTRAP,
FAULT_SIGILL = SIGILL,
#else
FAULT_SIGSEGV = 11,
FAULT_SIGTRAP = 5,
FAULT_SIGILL = 4,
#endif
};
}
using HostSignalDelegatorFunction = std::function<bool(FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext)>;
-29
View File
@@ -1,29 +0,0 @@
#pragma once
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Debug/InternalThreadState.h>
#include <stdint.h>
namespace FEXCore::Core {
struct RuntimeStats;
}
namespace FEXCore::Context {
class Context;
namespace Debug {
void CompileRIP(FEXCore::Context::Context *CTX, uint64_t RIP);
uint64_t GetThreadCount(FEXCore::Context::Context *CTX);
FEXCore::Core::RuntimeStats *GetRuntimeStatsForThread(FEXCore::Context::Context *CTX, uint64_t Thread);
bool GetDebugDataForRIP(FEXCore::Context::Context *CTX, uint64_t RIP, FEXCore::Core::DebugData *Data);
bool FindHostCodeForRIP(FEXCore::Context::Context *CTX, uint64_t RIP, uint8_t **Code);
// XXX:
// bool FindIRForRIP(FEXCore::Context::Context *CTX, uint64_t RIP, FEXCore::IR::IntrusiveIRList **ir);
// void SetIRForRIP(FEXCore::Context::Context *CTX, uint64_t RIP, FEXCore::IR::IntrusiveIRList *const ir);
}
}
@@ -38,12 +38,6 @@ namespace FEXCore::IR{
}
namespace FEXCore::Core {
struct RuntimeStats {
std::atomic_uint64_t InstructionsExecuted;
std::atomic_uint64_t BlocksCompiled;
};
struct DebugDataSubblock {
uint32_t HostCodeOffset;
uint32_t HostCodeSize;
@@ -103,8 +97,6 @@ namespace FEXCore::Core {
fextl::unique_ptr<FEXCore::IR::PassManager> PassManager;
FEXCore::HLE::ThreadManagement ThreadManager;
RuntimeStats Stats{};
int StatusCode{};
FEXCore::Context::ExitReason ExitReason {FEXCore::Context::ExitReason::EXIT_WAITING};
std::shared_ptr<FEXCore::CompileService> CompileService;
@@ -112,8 +104,19 @@ namespace FEXCore::Core {
std::shared_mutex ObjectCacheRefCounter{};
bool DestroyedByParent{false}; // Should the parent destroy this thread, or it destory itself
alignas(16) FEXCore::Core::CpuStateFrame BaseFrameState{};
struct DeferredSignalState {
#ifndef _WIN32
siginfo_t Info;
#endif
int Signal;
};
// Queue of thread local signal frames that have been deferred.
// Async signals aren't guaranteed to be delivered in any particular order, but FEX treats them as FILO.
fextl::vector<DeferredSignalState> DeferredSignalFrames;
// BaseFrameState should always be at the end.
alignas(16) FEXCore::Core::CpuStateFrame BaseFrameState{};
};
// static_assert(std::is_standard_layout<InternalThreadState>::value, "This needs to be standard layout");
}
+4 -8
View File
@@ -3,7 +3,6 @@
#include <shared_mutex>
#include <FEXCore/IR/IR.h>
#include <FEXHeaderUtils/ScopedSignalMask.h>
namespace FEXCore {
class CodeLoader;
@@ -52,9 +51,8 @@ namespace FEXCore::HLE {
class SourcecodeResolver;
struct AOTIRCacheEntryLookupResult {
AOTIRCacheEntryLookupResult(FEXCore::IR::AOTIRCacheEntry *Entry, uintptr_t VAFileStart, FHU::ScopedSignalMaskWithSharedLock &&lk)
: Entry(Entry), VAFileStart(VAFileStart), lk(std::move(lk))
{
AOTIRCacheEntryLookupResult(FEXCore::IR::AOTIRCacheEntry *Entry, uintptr_t VAFileStart)
: Entry(Entry), VAFileStart(VAFileStart) {
}
@@ -64,8 +62,6 @@ namespace FEXCore::HLE {
uintptr_t VAFileStart;
friend class SyscallHandler;
protected:
FHU::ScopedSignalMaskWithSharedLock lk;
};
class SyscallHandler {
@@ -78,8 +74,8 @@ namespace FEXCore::HLE {
SyscallOSABI GetOSABI() const { return OSABI; }
virtual FEXCore::CodeLoader *GetCodeLoader() const { return nullptr; }
virtual void MarkGuestExecutableRange(uint64_t Start, uint64_t Length) { }
virtual AOTIRCacheEntryLookupResult LookupAOTIRCacheEntry(uint64_t GuestAddr) = 0;
virtual void MarkGuestExecutableRange(FEXCore::Core::InternalThreadState *Thread, uint64_t Start, uint64_t Length) { }
virtual AOTIRCacheEntryLookupResult LookupAOTIRCacheEntry(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestAddr) = 0;
virtual SourcecodeResolver *GetSourcecodeResolver() { return nullptr; }
protected:
+12 -2
View File
@@ -4,7 +4,7 @@
#include <FEXCore/Utils/ThreadPoolAllocator.h>
#include <FEXCore/fextl/memory.h>
#include <FEXCore/fextl/sstream.h>
#include <FEXHeaderUtils/EnumOperators.h>
#include <FEXCore/Utils/EnumOperators.h>
#include <array>
#include <cassert>
@@ -491,10 +491,20 @@ protected:
enum class SyscallFlags : uint8_t {
DEFAULT = 0,
// Syscalldoesn't care about CPUState being serialized up to the syscall instruction.
// Means DeadCodeElimination can optimize through a syscall operation.
OPTIMIZETHROUGH = 1 << 0,
// Syscall only reads the passed in arguments. Doesn't read CPUState.
NOSYNCSTATEONENTRY = 1 << 1,
// Syscall doesn't return. Code generation after syscall return can be removed.
NORETURN = 1 << 2,
NOSIDEEFFECTS = 1 << 3
// Syscall doesn't have any side-effects, so if the result isn't used then it can be removed.
NOSIDEEFFECTS = 1 << 3,
// Syscall doesn't return a result.
// Means the resulting register shouldn't be written (Usually RAX).
// Usually used with !NOSYNCSTATEONENTRY, so the syscall can modify CPU state entirely.
// Then on return FEXCore picks up the new state.
NORETURNEDRESULT = 1 << 4,
};
FEX_DEF_NUM_OPS(SyscallFlags)
@@ -0,0 +1,134 @@
#pragma once
#include <FEXCore/Debug/InternalThreadState.h>
#include <atomic>
#include <cstdint>
#include <mutex>
#include <shared_mutex>
#include <signal.h>
#include <sys/syscall.h>
#include <unistd.h>
namespace FEXCore {
template<typename MutexType, void (MutexType::*lock_fn)(), void (MutexType::*unlock_fn)()>
class ScopedDeferredSignalWithMutexBase final {
public:
ScopedDeferredSignalWithMutexBase(MutexType &_Mutex, FEXCore::Core::InternalThreadState *Thread)
: Mutex {&_Mutex}
, Thread {Thread} {
// Needs to be atomic so that operations can't end up getting reordered around this.
Thread->CurrentFrame->State.DeferredSignalRefCount.Increment(1);
// Lock the mutex
(Mutex->*lock_fn)();
}
// No copy or assignment possible
ScopedDeferredSignalWithMutexBase(const ScopedDeferredSignalWithMutexBase&) = delete;
ScopedDeferredSignalWithMutexBase& operator=(ScopedDeferredSignalWithMutexBase&) = delete;
// Only move
ScopedDeferredSignalWithMutexBase(ScopedDeferredSignalWithMutexBase &&rhs)
: Mutex {rhs.Mutex}
, Thread {rhs.Thread} {
rhs.Mutex = nullptr;
}
~ScopedDeferredSignalWithMutexBase() {
if (Mutex != nullptr) {
// Unlock the mutex
(Mutex->*unlock_fn)();
#ifdef _M_X86_64
// Needs to be atomic so that operations can't end up getting reordered around this.
// Without this, the recount and the signal access could get reordered.
auto Result = Thread->CurrentFrame->State.DeferredSignalRefCount.Decrement(1);
// X86-64 must do an additional check around the store.
if ((Result - 1) == 0) {
// Must happen after the refcount store
Thread->CurrentFrame->State.DeferredSignalFaultAddress->Store(0);
}
#else
Thread->CurrentFrame->State.DeferredSignalRefCount.Decrement(1);
Thread->CurrentFrame->State.DeferredSignalFaultAddress->Store(0);
#endif
}
}
private:
MutexType *Mutex;
FEXCore::Core::InternalThreadState *Thread;
};
using ScopedDeferredSignalWithMutex = ScopedDeferredSignalWithMutexBase<std::mutex, &std::mutex::lock, &std::mutex::unlock>;
using ScopedDeferredSignalWithSharedLock = ScopedDeferredSignalWithMutexBase<std::shared_mutex, &std::shared_mutex::lock_shared, &std::shared_mutex::unlock_shared>;
using ScopedDeferredSignalWithUniqueLock = ScopedDeferredSignalWithMutexBase<std::shared_mutex, &std::shared_mutex::lock, &std::shared_mutex::unlock>;
template<typename MutexType, void (MutexType::*lock_fn)(), void (MutexType::*unlock_fn)()>
class ScopedPotentialDeferredSignalWithMutexBase final {
public:
ScopedPotentialDeferredSignalWithMutexBase(MutexType &_Mutex, FEXCore::Core::InternalThreadState *Thread, uint64_t Mask = ~0ULL)
: Mutex {&_Mutex}
, Thread {Thread} {
if (Thread) {
Thread->CurrentFrame->State.DeferredSignalRefCount.Increment(1);
}
else {
// Mask all signals, storing the original incoming mask
::syscall(SYS_rt_sigprocmask, SIG_SETMASK, &Mask, &OriginalMask, sizeof(OriginalMask));
}
// Lock the mutex
(Mutex->*lock_fn)();
}
// No copy or assignment possible
ScopedPotentialDeferredSignalWithMutexBase(const ScopedPotentialDeferredSignalWithMutexBase&) = delete;
ScopedPotentialDeferredSignalWithMutexBase& operator=(ScopedPotentialDeferredSignalWithMutexBase&) = delete;
// Only move
ScopedPotentialDeferredSignalWithMutexBase(ScopedPotentialDeferredSignalWithMutexBase &&rhs)
: Mutex {rhs.Mutex}
, Thread {rhs.Thread} {
rhs.Mutex = nullptr;
}
~ScopedPotentialDeferredSignalWithMutexBase() {
if (Mutex != nullptr) {
// Unlock the mutex
(Mutex->*unlock_fn)();
if (Thread) {
#ifdef _M_X86_64
// Needs to be atomic so that operations can't end up getting reordered around this.
// Without this, the refcount and the signal access could get reordered.
auto Result = Thread->CurrentFrame->State.DeferredSignalRefCount.Decrement(1);
// X86-64 must do an additional check around the store.
if ((Result - 1) == 0) {
// Must happen after the refcount store
Thread->CurrentFrame->State.DeferredSignalFaultAddress->Store(0);
}
#else
Thread->CurrentFrame->State.DeferredSignalRefCount.Decrement(1);
Thread->CurrentFrame->State.DeferredSignalFaultAddress->Store(0);
#endif
}
else {
// Unmask back to the original signal mask
::syscall(SYS_rt_sigprocmask, SIG_SETMASK, &OriginalMask, nullptr, sizeof(OriginalMask));
}
}
}
private:
MutexType *Mutex;
uint64_t OriginalMask{};
FEXCore::Core::InternalThreadState *Thread;
};
using ScopedPotentialDeferredSignalWithMutex = ScopedPotentialDeferredSignalWithMutexBase<std::mutex, &std::mutex::lock, &std::mutex::unlock>;
using ScopedPotentialDeferredSignalWithSharedLock = ScopedPotentialDeferredSignalWithMutexBase<std::shared_mutex, &std::shared_mutex::lock_shared, &std::shared_mutex::unlock_shared>;
using ScopedPotentialDeferredSignalWithUniqueLock = ScopedPotentialDeferredSignalWithMutexBase<std::shared_mutex, &std::shared_mutex::lock, &std::shared_mutex::unlock>;
}
+256
View File
@@ -0,0 +1,256 @@
#pragma once
#include <FEXCore/fextl/allocator.h>
#include <FEXCore/Utils/EnumOperators.h>
#ifndef _WIN32
#include <fcntl.h>
#include <unistd.h>
#else
#define WIN32_LEAN_AND_MEAN
#include <windows.h>
#undef ERROR
#endif
namespace FEXCore::File {
enum class FileModes : uint32_t {
READ = (1U << 0),
WRITE = (1U << 1),
CREATE = (1U << 2),
TRUNCATE = (1U << 3),
};
enum class SeekOp {
BEGIN,
CURRENT,
END,
};
FEX_DEF_NUM_OPS(FileModes)
class File final {
public:
#ifndef _WIN32
using FileHandleType = int;
#else
using FileHandleType = HANDLE;
#endif
File() = default;
File(const char *Filepath, FileModes Modes) {
#ifndef _WIN32
auto Disp = TranslateModes(Modes);
Handle = open(Filepath, Disp, DEFAULT_USER_PERMS);
IsValidHandle = Handle != -1;
#else
auto Disp = TranslateModes(Modes);
if (Disp.CreationFlag == OPEN_ALWAYS && Disp.TruncateOnExist) {
// If Open + Truncate then try to open with truncate behaviour first.
Handle = CreateFileA(Filepath, Disp.Access, DEFAULT_SHARE_MODE, nullptr, TRUNCATE_EXISTING, FILE_ATTRIBUTE_NORMAL, nullptr);
if (Handle == INVALID_HANDLE_VALUE && GetLastError() == ERROR_FILE_NOT_FOUND) {
// File didn't exist, just open.
Handle = CreateFileA(Filepath, Disp.Access, DEFAULT_SHARE_MODE, nullptr, CREATE_NEW, FILE_ATTRIBUTE_NORMAL, nullptr);
}
}
else {
Handle = CreateFileA(Filepath, Disp.Access, DEFAULT_SHARE_MODE, nullptr, Disp.CreationFlag, FILE_ATTRIBUTE_NORMAL, nullptr);
}
IsValidHandle = Handle != INVALID_HANDLE_VALUE;
#endif
}
/**
* @brief Write Bytes to File
*
* @param Buffer The buffer to write.
* @param Bytes The number of bytes to write.
*
* @return The number of bytes actually written or -1 on error.
*/
ssize_t Write(void const* Buffer, size_t Bytes) {
#ifndef _WIN32
return write(Handle, Buffer, Bytes);
#else
DWORD BytesWritten{};
auto Result = WriteFile(Handle, Buffer, Bytes, &BytesWritten, nullptr);
if (Result) {
return BytesWritten;
}
// Some error, match Linux side.
return -1;
#endif
}
/**
* @brief Read at most Bytes in to the buffer.
*
* @param Buffer The buffer where the data is read in to.
* @param Bytes The size of the buffer.
*
* @return The number of bytes read or -1 on error.
*/
ssize_t Read(void *Buffer, size_t Bytes) {
#ifndef _WIN32
return read(Handle, Buffer, Bytes);
#else
DWORD BytesRead{};
auto Result = ReadFile(Handle, Buffer, Bytes, &BytesRead, nullptr);
if (Result) {
return BytesRead;
}
// Some error, match Linux side.
return -1;
#endif
}
~File() {
if (!IsValidHandle) return;
if (!ShouldClose) return;
#ifndef _WIN32
close(Handle);
#else
CloseHandle(Handle);
#endif
}
/**
* @brief Gets a File object that points to stdout
*/
static File GetStdOUT() {
#ifndef _WIN32
return File(STDOUT_FILENO, false);
#else
return File(GetStdHandle(STD_OUTPUT_HANDLE), false);
#endif
}
/**
* @brief Gets a File object that points to stderr
*/
static File GetStdERR() {
#ifndef _WIN32
return File(STDERR_FILENO, false);
#else
return File(GetStdHandle(STD_ERROR_HANDLE), false);
#endif
}
/**
* @brief Returns if the file handle is valid.
*/
bool IsValid() const { return IsValidHandle; }
/**
* @brief Flush the file contents to the output file backing.
*
* @return True if the flush occured.
*/
bool Flush() {
#ifndef _WIN32
return fsync(Handle) == 0;
#else
return FlushFileBuffers(Handle);
#endif
}
/**
* @brief Seek the file pointer location.
*
* @param Distance The distance to travel.
* @param Op The operation from where to start the travel.
*
* @return The current file pointer location or -1.
*/
ssize_t Seek(ssize_t Distance, SeekOp Op) {
#ifndef _WIN32
return lseek(Handle, Distance, TranslateSeek(Op));
#else
LARGE_INTEGER NewDistance {
.QuadPart = Distance
};
LARGE_INTEGER NewPointer;
auto Result = SetFilePointerEx(Handle, NewDistance, &NewPointer, TranslateSeek(Op));
if (Result) {
return NewPointer.QuadPart;
}
// Some error, match Linux side.
return -1;
#endif
}
protected:
File(FileHandleType Handle, bool ShouldClose)
: ShouldClose {ShouldClose}
, IsValidHandle {true}
, Handle {Handle}
{}
private:
bool ShouldClose{};
bool IsValidHandle{};
FileHandleType Handle;
#ifndef _WIN32
static constexpr int DEFAULT_USER_PERMS = S_IRWXU | S_IRWXG | S_IRWXO;
static uint32_t TranslateModes(FileModes Modes) {
uint32_t Mode{};
if ((Modes & FileModes::READ) == FileModes::READ)
Mode |= O_RDONLY;
if ((Modes & FileModes::WRITE) == FileModes::WRITE)
Mode |= O_WRONLY;
if ((Modes & FileModes::CREATE) == FileModes::CREATE)
Mode |= O_CREAT;
if ((Modes & FileModes::TRUNCATE) == FileModes::TRUNCATE)
Mode |= O_TRUNC;
// Always enable CLOEXEC so that the FD is closed on execve.
// FEXCore never wants to leak FDs across execve using this interface.
Mode |= O_CLOEXEC;
return Mode;
}
static uint32_t TranslateSeek(SeekOp Op) {
switch (Op) {
case SeekOp::BEGIN: return SEEK_SET;
case SeekOp::CURRENT: return SEEK_CUR;
case SeekOp::END: return SEEK_END;
default: FEX_UNREACHABLE;
}
}
#else
static constexpr int DEFAULT_SHARE_MODE = FILE_SHARE_READ | FILE_SHARE_WRITE | FILE_SHARE_DELETE;
struct Disposition {
uint32_t CreationFlag;
uint32_t Access;
bool TruncateOnExist;
};
static Disposition TranslateModes(FileModes Modes) {
Disposition Disp{};
if ((Modes & FileModes::READ) == FileModes::READ)
Disp.Access |= GENERIC_READ;
if ((Modes & FileModes::WRITE) == FileModes::WRITE)
Disp.Access |= GENERIC_WRITE;
if ((Modes & FileModes::CREATE) == FileModes::CREATE)
Disp.CreationFlag = CREATE_ALWAYS;
else
Disp.CreationFlag = OPEN_ALWAYS;
if ((Modes & FileModes::TRUNCATE) == FileModes::TRUNCATE)
Disp.TruncateOnExist = true;
return Disp;
}
static uint32_t TranslateSeek(SeekOp Op) {
switch (Op) {
case SeekOp::BEGIN: return FILE_BEGIN;
case SeekOp::CURRENT: return FILE_CURRENT;
case SeekOp::END: return FILE_END;
default: FEX_UNREACHABLE;
}
}
#endif
};
}
+1
View File
@@ -37,6 +37,7 @@ namespace FEXCore::Telemetry {
TYPE_CAS_32BIT_TEAR,
TYPE_CAS_64BIT_TEAR,
TYPE_CAS_128BIT_TEAR,
TYPE_CRASH_MASK,
TYPE_LAST,
};
+24
View File
@@ -1,6 +1,7 @@
#pragma once
#include <FEXCore/fextl/allocator.h>
#include <FEXCore/fextl/string.h>
#include <FEXCore/Utils/File.h>
#include <fmt/format.h>
#include <unistd.h>
@@ -38,6 +39,7 @@ namespace fextl::fmt {
return fextl::fmt::vformat(fmt, ::fmt::make_format_args(args...));
}
#ifndef _WIN32
template <typename... T>
FMT_INLINE auto print(::fmt::format_string<T...> fmt, T&&... args)
-> void {
@@ -51,6 +53,28 @@ namespace fextl::fmt {
auto String = fextl::fmt::vformat(fmt, ::fmt::make_format_args(args...));
write(FD, String.c_str(), String.size());
}
#else
template <typename... T>
FMT_INLINE auto print(::fmt::format_string<T...> fmt, T&&... args)
-> void {
auto String = fextl::fmt::vformat(fmt, ::fmt::make_format_args(args...));
auto f = fextl::file::File::GetStdOUT();
f.Write(String.c_str(), String.size());
}
template <typename... T>
FMT_INLINE auto print(HANDLE File, ::fmt::format_string<T...> fmt, T&&... args)
-> void {
auto String = fextl::fmt::vformat(fmt, ::fmt::make_format_args(args...));
WriteFile(File, String.c_str(), String.size(), nullptr, nullptr);
}
#endif
template <typename... T>
FMT_INLINE auto print(FEXCore::File::File& f, ::fmt::format_string<T...> fmt, T&&... args)
-> void {
auto String = fextl::fmt::vformat(fmt, ::fmt::make_format_args(args...));
f.Write(String.c_str(), String.size());
}
template <typename... T>
FMT_INLINE auto print(std::FILE* f, ::fmt::format_string<T...> fmt, T&&... args)
+3 -1
View File
@@ -1 +1,3 @@
add_subdirectory(Emitter/)
if (NOT MINGW_BUILD)
add_subdirectory(Emitter/)
endif()
+1 -1
@@ -1,8 +1,8 @@
#pragma once
#include <FEXCore/Utils/CompilerDefs.h>
#include <FEXCore/Utils/LogManager.h>
#include <atomic>
#include <cstdint>
#include <mutex>
#include <shared_mutex>
+3 -1
View File
@@ -305,7 +305,7 @@ def GetRootFSPath():
return _RootFSPath
def CheckRootFSInstallStatus():
# Matches what is available on https://rootfs.fex-emu.com/file/fex-rootfs/RootFS_links.json
# Matches what is available on https://rootfs.fex-emu.gg/RootFS_links.json
UbuntuVersionToRootFS = {
"20.04": "Ubuntu_20_04.sqsh",
"20.04": "Ubuntu_20_04.ero",
@@ -313,6 +313,8 @@ def CheckRootFSInstallStatus():
"22.04": "Ubuntu_22_04.ero",
"22.10": "Ubuntu_22_10.sqsh",
"22.10": "Ubuntu_22_10.ero",
"23.04": "Ubuntu_23_04.sqsh",
"23.04": "Ubuntu_23_04.ero",
}
return os.path.exists(GetRootFSPath() + UbuntuVersionToRootFS[GetDistro()[1]])
+2
View File
@@ -73,6 +73,7 @@ class HostFeatures(Flag) :
FEATURE_BMI1 = (1 << 6)
FEATURE_BMI2 = (1 << 7)
FEATURE_CLWB = (1 << 8)
FEATURE_LINUX = (1 << 9)
RegStringLookup = {
"NONE": Regs.REG_NONE,
@@ -145,6 +146,7 @@ HostFeaturesLookup = {
"BMI1" : HostFeatures.FEATURE_BMI1,
"BMI2" : HostFeatures.FEATURE_BMI2,
"CLWB" : HostFeatures.FEATURE_CLWB,
"LINUX" : HostFeatures.FEATURE_LINUX,
}
def parse_hexstring(s):
+3 -2
View File
@@ -1,5 +1,6 @@
add_subdirectory(Common/)
add_subdirectory(Tools/)
if (NOT MINGW_BUILD)
add_subdirectory(Common/)
add_subdirectory(Linux/)
add_subdirectory(Tools/)
endif()
+7 -3
View File
@@ -3,12 +3,16 @@ add_subdirectory(cpp-optparse/)
set(NAME Common)
set(SRCS
ArgumentLoader.cpp
Config.cpp
EnvironmentLoader.cpp
FEXServerClient.cpp
FileFormatCheck.cpp
StringUtil.cpp)
if (NOT MINGW_BUILD)
list (APPEND SRCS
Config.cpp
FEXServerClient.cpp
FileFormatCheck.cpp)
endif()
add_library(${NAME} STATIC ${SRCS})
target_link_libraries(${NAME} FEXCore_Base cpp-optparse json-maker FEXHeaderUtils)
target_include_directories(${NAME} PRIVATE ${CMAKE_CURRENT_SOURCE_DIR}/External/cpp-optparse/)
+9 -12
View File
@@ -1,7 +1,3 @@
if (ENABLE_VISUAL_DEBUGGER)
add_subdirectory(Debugger/)
endif()
if (NOT MINGW_BUILD)
if (NOT TERMUX_BUILD)
# Termux builds can't rely on X11 packages
@@ -19,13 +15,14 @@ if (NOT MINGW_BUILD)
add_subdirectory(FEXGetConfig/)
add_subdirectory(FEXServer/)
add_subdirectory(FEXBash/)
add_subdirectory(FEXLoader/)
set(NAME Opt)
set(SRCS Opt.cpp)
add_executable(${NAME} ${SRCS})
target_include_directories(${NAME} PRIVATE ${CMAKE_CURRENT_SOURCE_DIR}/Source/)
target_link_libraries(${NAME} FEXCore Common pthread)
endif()
set(NAME Opt)
set(SRCS Opt.cpp)
add_executable(${NAME} ${SRCS})
target_include_directories(${NAME} PRIVATE ${CMAKE_CURRENT_SOURCE_DIR}/Source/)
target_link_libraries(${NAME} FEXCore Common pthread)
add_subdirectory(FEXLoader/)
-32
View File
@@ -1,32 +0,0 @@
set(NAME Debugger)
set(SRCS Main.cpp
DebuggerState.cpp
Context.cpp
FEXImGui.cpp
IMGui_I.cpp
IRLexer.cpp
GLUtils.cpp
MainWindow.cpp
Disassembler.cpp
${CMAKE_SOURCE_DIR}/External/imgui/examples/imgui_impl_glfw.cpp
${CMAKE_SOURCE_DIR}/External/imgui/examples/imgui_impl_opengl3.cpp
)
find_library(EPOXY_LIBRARY epoxy)
find_library(GLFW_LIBRARY glfw3)
find_package(LLVM CONFIG QUIET)
if(LLVM_FOUND AND TARGET LLVM)
message(STATUS "LLVM found!")
include_directories(${LLVM_INCLUDE_DIRS})
endif()
add_definitions(-DIMGUI_IMPL_OPENGL_LOADER_CUSTOM=<epoxy/gl.h>)
add_executable(${NAME} ${SRCS})
target_link_libraries(${NAME} PRIVATE LLVM)
target_include_directories(${NAME} PRIVATE ${LLVM_INCLUDE_DIRS})
target_include_directories(${NAME} PRIVATE ${CMAKE_SOURCE_DIR}/Source/)
target_include_directories(${NAME} PRIVATE ${CMAKE_SOURCE_DIR}/External/imgui/examples/)
target_link_libraries(${NAME} PRIVATE FEXCore Common pthread LLVM epoxy glfw X11 EGL imgui tiny-json json-maker)
-91
View File
@@ -1,91 +0,0 @@
#include "Context.h"
#include <cassert>
#include <GLFW/glfw3.h>
#include <vector>
#include <cstdio>
namespace GLContext {
void glfw_error_callback(int error, const char* description)
{
fprintf(stderr, "Glfw Error %d: %s\n", error, description);
}
class GLFWContext final : public GLContext::Context {
public:
void Create(const char *Title) override {
glfwSetErrorCallback(glfw_error_callback);
if (!glfwInit()) {
assert(0 && "Couldn't init glfw");
}
glfwWindowHint(GLFW_CONTEXT_VERSION_MAJOR, 4);
glfwWindowHint(GLFW_CONTEXT_VERSION_MINOR, 1);
glfwWindowHint(GLFW_OPENGL_PROFILE, GLFW_OPENGL_CORE_PROFILE);
glfwWindowHint(GLFW_OPENGL_FORWARD_COMPAT, GL_TRUE);
glfwWindowHint(GLFW_RED_BITS, 8);
glfwWindowHint(GLFW_GREEN_BITS, 8);
glfwWindowHint(GLFW_BLUE_BITS, 8);
glfwWindowHint(GLFW_ALPHA_BITS, 8);
glfwWindowHint(GLFW_RESIZABLE, 1);
glfwWindowHint(GLFW_DOUBLEBUFFER, 1);
Window = glfwCreateWindow(640, 640, Title, nullptr, nullptr);
if (!Window) {
assert(0 && "Couldn't create window");
}
glfwMakeContextCurrent(Window);
glfwSwapInterval(1);
}
void Shutdown() override {
glfwMakeContextCurrent(nullptr);
glfwDestroyWindow(Window);
glfwTerminate();
}
void Swap() override {
glfwSwapBuffers(Window);
CheckWindowDimensions();
}
void RegisterResizeEvent(ResizeEvent Event) override {
ResizeEvents.emplace_back(Event);
}
void GetDim(uint32_t *Dim) override {
Dim[0] = Width;
Dim[1] = Height;
}
private:
void CheckWindowDimensions() {
int LocalWidth;
int LocalHeight;
glfwGetWindowSize(Window, &LocalWidth, &LocalHeight);
if (LocalHeight != Height || LocalWidth != Width) {
Width = LocalWidth;
Height = LocalHeight;
for (auto const &Event : ResizeEvents) {
Event(Width, Height);
}
}
}
void* GetWindow() override {
return Window;
}
std::vector<ResizeEvent> ResizeEvents;
GLFWwindow *Window;
uint32_t Width{};
uint32_t Height{};
};
std::unique_ptr<Context> CreateContext() {
return std::make_unique<GLFWContext>();
}
}
-20
View File
@@ -1,20 +0,0 @@
#pragma once
#include <functional>
#include <memory>
namespace GLContext {
class Context {
public:
virtual ~Context() {}
virtual void Create(const char *Title) = 0;
virtual void Shutdown() = 0;
virtual void Swap() = 0;
virtual void GetDim(uint32_t *Dim) = 0;
using ResizeEvent = std::function<void(uint32_t, uint32_t)>;
virtual void RegisterResizeEvent(ResizeEvent Event) = 0;
virtual void* GetWindow() = 0;
};
std::unique_ptr<Context> CreateContext();
}
-150
View File
@@ -1,150 +0,0 @@
#include "DebuggerState.h"
#include <FEXCore/Core/Context.h>
#include <FEXCore/Debug/ContextDebug.h>
namespace FEX::DebuggerState {
FEXCore::Context::Context *s_CTX{};
FEXCore::Config::ConfigCore s_CoreType = FEXCore::Config::ConfigCore::CONFIG_INTERPRETER;
int s_IsStepping = 0;
bool NewState = false;
std::function<void()> StepCallback;
std::function<void()> PauseCallback;
std::function<void()> ContinueCallback;
std::function<void()> NewStateCallback;
std::function<void()> CloseCallback;
std::function<void(char const *, bool)> CreateCallback;
std::function<void(uint64_t)> CompileRIPCallback;
std::function<void(std::stringstream *out, uint64_t PC)> GetIRCallback;
bool ActiveCore() {
return s_CTX != nullptr;
}
void SetContext(FEXCore::Context::Context *ctx) {
s_CTX = ctx;
}
FEXCore::Context::Context *GetContext() {
return s_CTX;
}
FEXCore::Config::ConfigCore GetCoreType() {
return s_CoreType;
}
void SetCoreType(FEXCore::Config::ConfigCore CoreType) {
s_CoreType = CoreType;
}
int GetRunningMode() {
return s_IsStepping;
}
int GetCoreCurrentRunningMode() {
if (ActiveCore()) {
return FEXCore::Config::GetConfig(s_CTX, FEXCore::Config::CONFIG_SINGLESTEP);
}
return s_IsStepping;
}
void SetRunningMode(int RunningMode) {
s_IsStepping = RunningMode;
if (ActiveCore()) {
FEXCore::Config::SetConfig(s_CTX, FEXCore::Config::CONFIG_SINGLESTEP, RunningMode);
}
}
FEXCore::Core::CPUState GetCPUState() {
auto ret = FEXCore::Core::CPUState{};
if (ActiveCore()) {
FEXCore::Context::GetCPUState(s_CTX, &ret);
}
return ret;
}
bool IsCoreRunning() {
if (!ActiveCore()) {
return false;
}
return !FEXCore::Context::IsDone(s_CTX);
}
// Client interface
void RegisterStepCallback(std::function<void()> Callback) {
StepCallback = std::move(Callback);
}
void RegisterPauseCallback(std::function<void()> Callback) {
PauseCallback = std::move(Callback);
}
void RegisterContinueCallback(std::function<void()> Callback) {
ContinueCallback = std::move(Callback);
}
void Step() {
StepCallback();
}
void Pause() {
PauseCallback();
}
void Continue() {
ContinueCallback();
}
void RegisterCreateCallback(std::function<void(char const *Filename, bool)> Callback)
{
CreateCallback = std::move(Callback);
}
void Create(char const *Filename, bool ELF) {
CreateCallback(Filename, ELF);
}
void RegisterCompileRIPCallback(std::function<void(uint64_t)> Callback) {
CompileRIPCallback = std::move(Callback);
}
void CompileRIP(uint64_t RIP) {
CompileRIPCallback(RIP);
}
void RegisterCloseCallback(std::function<void()> Callback) {
CloseCallback = std::move(Callback);
}
void Close() {
CloseCallback();
}
void RegisterNewStateCallback(std::function<void()> Callback) {
NewStateCallback = std::move(Callback);
}
void RegisterGetIRCallback(std::function<void(std::stringstream *out, uint64_t PC)> Callback) {
GetIRCallback = std::move(Callback);
}
void GetIR(std::stringstream *out, uint64_t PC) {
GetIRCallback(out, PC);
}
void CallNewState() {
NewStateCallback();
}
bool HasNewState() {
return NewState;
}
void SetHasNewState(bool State) {
NewState = State;
}
}
-49
View File
@@ -1,49 +0,0 @@
#pragma once
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/Context.h>
#include <FEXCore/Core/CoreState.h>
#include <sstream>
namespace FEX::DebuggerState {
bool ActiveCore();
void SetContext(FEXCore::Context::Context *ctx);
FEXCore::Context::Context *GetContext();
FEXCore::Config::ConfigCore GetCoreType();
void SetCoreType(FEXCore::Config::ConfigCore CoreType);
int GetRunningMode(); ///< This is the running mode we've set
int GetCoreCurrentRunningMode(); ///< This typically matches `GetRunningMode()` but can differ when we are in the middle of stepping
void SetRunningMode(int RunningMode);
FEXCore::Core::CPUState GetCPUState();
bool IsCoreRunning();
// Client interface
void RegisterStepCallback(std::function<void()> Callback);
void RegisterPauseCallback(std::function<void()> Callback);
void RegisterContinueCallback(std::function<void()> Callback);
void Step();
void Pause();
void Continue();
void RegisterCreateCallback(std::function<void(char const *Filename, bool ELF)> Callback);
void Create(char const *Filename, bool ELF);
void RegisterCompileRIPCallback(std::function<void(uint64_t)> Callback);
void CompileRIP(uint64_t RIP);
void RegisterCloseCallback(std::function<void()> Callback);
void Close();
void RegisterNewStateCallback(std::function<void()> Callback);
void CallNewState();
void RegisterGetIRCallback(std::function<void(std::stringstream *out, uint64_t PC)>);
void GetIR(std::stringstream *out, uint64_t PC);
bool HasNewState();
void SetHasNewState(bool State = true);
}
-94
View File
@@ -1,94 +0,0 @@
#include "Disassembler.h"
#include <llvm-c/Disassembler.h>
#include <llvm-c/Target.h>
#include <memory>
#include <sstream>
#include <iomanip>
namespace FEX::Debugger {
class LLVMDisassembler final : public Disassembler {
public:
~LLVMDisassembler() override {
LLVMDisasmDispose(LLVMContext);
}
explicit LLVMDisassembler(const char *Arch);
std::string Disassemble(uint8_t *Code, uint32_t CodeSize, uint32_t MaxInst, uint64_t StartingPC, uint32_t *InstructionCount) override;
private:
LLVMDisasmContextRef LLVMContext;
};
LLVMDisassembler::LLVMDisassembler(const char *Arch) {
LLVMInitializeAllTargetInfos();
LLVMInitializeAllTargetMCs();
LLVMInitializeAllDisassemblers();
LLVMContext = LLVMCreateDisasmCPU(Arch, "", nullptr, 0, nullptr, nullptr);
if (!LLVMContext)
return;
LLVMSetDisasmOptions(LLVMContext, LLVMDisassembler_Option_AsmPrinterVariant |
LLVMDisassembler_Option_PrintLatency);
}
std::string LLVMDisassembler::Disassemble(uint8_t *Code, uint32_t CodeSize, uint32_t MaxInst, uint64_t StartingPC, uint32_t *InstructionCount) {
std::ostringstream Output;
uint8_t *CurrentRIPAddr = Code;
uint8_t *EndRIPAddr = Code + CodeSize;
uint64_t CurrentRIP = StartingPC;
uint32_t NumberOfInstructions = 0;
while (CurrentRIPAddr <= EndRIPAddr) {
char OutputText[128];
size_t InstSize = LLVMDisasmInstruction(LLVMContext,
CurrentRIPAddr,
static_cast<uint64_t>(EndRIPAddr - CurrentRIPAddr),
CurrentRIP,
OutputText,
128);
Output << "0x" << std::hex << CurrentRIP << ": ";
if (!InstSize) {
Output << "<Invalid Inst>" << std::endl;
break;
}
else {
// Print instruction hex encodings if we want
if (!true) {
for (size_t i = 0; i < InstSize; ++i) {
Output << std::setw(2) << std::setfill('0') << static_cast<uint32_t>(CurrentRIPAddr[i]);
if ((i + 1) == InstSize) {
Output << ": ";
}
else {
Output << " ";
}
}
}
Output << OutputText << std::endl;
}
CurrentRIP += InstSize;
CurrentRIPAddr += InstSize;
NumberOfInstructions++;
if (NumberOfInstructions >= MaxInst) {
break;
}
}
*InstructionCount = NumberOfInstructions;
return Output.str();
}
std::unique_ptr<Disassembler> CreateHostDisassembler() {
return std::make_unique<LLVMDisassembler>("x86_64-none-unknown");
}
std::unique_ptr<Disassembler> CreateGuestDisassembler() {
return std::make_unique<LLVMDisassembler>("x86_64-none-unknown");
}
}
-15
View File
@@ -1,15 +0,0 @@
#pragma once
#include <memory>
namespace FEX::Debugger {
class Disassembler {
public:
virtual ~Disassembler() {}
virtual std::string Disassemble(uint8_t *Code, uint32_t CodeSize, uint32_t MaxInst, uint64_t StartingPC, uint32_t *InstructionCount) = 0;
};
std::unique_ptr<Disassembler> CreateHostDisassembler();
std::unique_ptr<Disassembler> CreateGuestDisassembler();
}
-153
View File
@@ -1,153 +0,0 @@
#include "FEXImGui.h"
#include <imgui.h>
#ifndef IMGUI_DEFINE_MATH_OPERATORS
#define IMGUI_DEFINE_MATH_OPERATORS
#endif
#include <imgui_internal.h>
#include <vector>
namespace FEXImGui {
using namespace ImGui;
// FIXME: In principle this function should be called EndListBox(). We should rename it after re-evaluating if we want to keep the same signature.
bool ListBoxHeader(const char* label, int items_count)
{
// Size default to hold ~7.25 items.
// We add +25% worth of item height to allow the user to see at a glance if there are more items up/down, without looking at the scrollbar.
// We don't add this extra bit if items_count <= height_in_items. It is slightly dodgy, because it means a dynamic list of items will make the widget resize occasionally when it crosses that size.
// I am expecting that someone will come and complain about this behavior in a remote future, then we can advise on a better solution.
const ImGuiStyle& style = GetStyle();
// We include ItemSpacing.y so that a list sized for the exact number of items doesn't make a scrollbar appears. We could also enforce that by passing a flag to BeginChild().
ImVec2 size;
size.x = 0.0f;
auto Window = GetCurrentWindowRead();
ImVec2 Base = Window->Rect().Min;
ImVec2 Size = Window->Rect().Max;
ImVec2 Height = Size - Base;
size.y = Height.y - style.FramePadding.y * 2.0f;
return ImGui::ListBoxHeader(label, size);
}
bool ListBox(const char* label, int* current_item, bool (*items_getter)(void*, int, const char**), void* data, int items_count)
{
if (!ListBoxHeader(label, items_count))
return false;
// Assume all items have even height (= 1 line of text). If you need items of different or variable sizes you can create a custom version of ListBox() in your code without using the clipper.
ImGuiContext& g = *GImGui;
bool value_changed = false;
ImGuiListClipper clipper(items_count, GetTextLineHeightWithSpacing()); // We know exactly our line height here so we pass it as a minor optimization, but generally you don't need to.
while (clipper.Step())
for (int i = clipper.DisplayStart; i < clipper.DisplayEnd; i++)
{
const bool item_selected = (i == *current_item);
const char* item_text;
if (!items_getter(data, i, &item_text))
item_text = "*Unknown item*";
PushID(i);
if (Selectable(item_text, item_selected))
{
*current_item = i;
value_changed = true;
}
if (item_selected)
SetItemDefaultFocus();
PopID();
}
ListBoxFooter();
if (value_changed)
MarkItemEdited(g.CurrentWindow->DC.LastItemId);
return value_changed;
}
bool CustomIRViewer(char const *buf, size_t buf_size, std::vector<IRLines> *lines) {
ImGuiContext& g = *GImGui;
const float inner_spacing = g.Style.ItemInnerSpacing.x;
ImGui::Columns(2);
ImGui::SetColumnWidth(0, 100.0f + inner_spacing * 2.0f);
static float TextOffset = 0.0f;
if (ImGui::BeginChildFrame(GetID(""), ImVec2(100, -1))) {
const std::vector<ImVec4> LineColors = {
ImVec4(1, 0, 0, 1),
ImVec4(1, 1, 0, 1),
ImVec4(0, 0, 0, 1),
ImVec4(0, 0, 1, 1),
ImVec4(1, 0, 1, 1),
ImVec4(0, 1, 1, 1),
};
auto Window = GetCurrentWindowRead();
auto DrawList = GetWindowDrawList();
auto DrawIRLine = [&](size_t Index, float From, float To, ImVec4 Color) {
ImVec2 Size = Window->Rect().Max;
float LineWidth = 3.0;
float LineSpacing = LineWidth * 5;
// Zero indexing, add one
Index = Index % LineColors.size();
Index++;
From -= TextOffset;
To -= TextOffset;
ImVec2 LeftFrom = ImVec2(Size.x - LineSpacing * static_cast<float>(Index), From);
ImVec2 RightFrom = ImVec2(Size.x, From);
ImVec2 LeftTo = ImVec2(Size.x - LineSpacing * static_cast<float>(Index), To);
ImVec2 RightTo = ImVec2(Size.x, To);
// We want to draw a three lines
// One horizontal one from the "FROM" offset
// One vertical one down the middle. Offset by Index
// Another horizontol to the "TO" offset
//
// Additionally we want a triangle pointing to the TO location
// Draw the from line
DrawList->AddLine(LeftFrom, RightFrom, GetColorU32(Color), LineWidth);
// Draw the to line
DrawList->AddLine(LeftTo, RightTo, GetColorU32(Color), LineWidth);
// Draw a small filled rectangle at the target
DrawList->AddTriangleFilled(ImVec2(RightTo.x - LineWidth * 2, RightTo.y - LineWidth), ImVec2(RightTo.x - LineWidth * 2, RightTo.y + LineWidth), RightTo, GetColorU32(Color));
// Draw the vertical line between the two
DrawList->AddLine(LeftFrom, LeftTo, GetColorU32(Color), LineWidth);
};
ImVec2 Base = Window->Rect().Min;
ImVec2 Size = Window->Rect().Max;
DrawList->AddRectFilled(Base, Size, GetColorU32(ImVec4(0.5, 0.5, 0.5, 1.0)));
float FontSize = Window->CalcFontSize();
for (size_t i = 0; i < lines->size(); ++i) {
float BaseOffset = Base.y + inner_spacing + FontSize / 2.0f;
DrawIRLine(i, BaseOffset + static_cast<float>(lines->at(i).From) * FontSize, BaseOffset + static_cast<float>(lines->at(i).To) * FontSize, LineColors[i % LineColors.size()]);
}
}
ImGui::EndChildFrame();
ImGui::NextColumn();
ImGui::InputTextMultiline("##IR", const_cast<char*>(buf), buf_size, ImVec2(-1, -1), ImGuiInputTextFlags_ReadOnly);
// This is a bit filthy
// Pulls the window and if the child is available then we can pull its child and read the scroll amount
// This lets the IR CFG lines match up with textbox, albiet a frame behind so it gets a bit of a wiggle
auto win = GetCurrentWindow();
if (win->DC.ChildWindows.size() == 2) {
win = win->DC.ChildWindows[win->DC.ChildWindows.size() - 1];
TextOffset = win->Scroll.y;
}
return false;
}
}
-18
View File
@@ -1,18 +0,0 @@
#pragma once
#include <imgui.h>
#include <cstdint>
#include <vector>
namespace FEXImGui {
// Listbox that fills the child window
bool ListBox(const char* label, int* current_item, bool (*items_getter)(void*, int, const char**), void* data, int items_count);
struct IRLines {
size_t From;
size_t To;
};
bool CustomIRViewer(char const *buf, size_t buf_size, std::vector<IRLines> *lines);
}
-5
View File
@@ -1,5 +0,0 @@
#include "GLUtils.h"
#include <cstdio>
namespace GLUtils {
}
-7
View File
@@ -1,7 +0,0 @@
#pragma once
#include <epoxy/gl.h>
#include <imgui.h>
namespace GLUtils {
}
File diff suppressed because it is too large. Load diff
-8
View File
@@ -1,8 +0,0 @@
#include "Context.h"
namespace FEX::Debugger {
void Init();
void Shutdown();
void DrawDebugUI(GLContext::Context *Context);
}
-43
View File
@@ -1,43 +0,0 @@
#include <cstring>
#include <sstream>
#include "IRLexer.h"
#include "Common/StringUtil.h"
#include <FEXCore/IR/IR.h>
#include <FEXCore/Utils/LogManager.h>
namespace FEX::Debugger::IR {
bool Lexer::Lex(char const *IR) {
// std::istringstream iss {std::string(IR)};
// HadError = false;
//
// [[maybe_unused]] int CurrentLine {};
// [[maybe_unused]] int CurrentColumn {};
// while (true) {
// char Line[256];
// char *Str = Line;
// std::string LineStr;
// std::getline(iss, LineStr);
// FEX::StringUtil::trim(LineStr);
// strncpy(Line, LineStr.c_str(), 256);
//
// struct {
// bool HadDest;
// FEXCore::IR::AlignmentType DestLoc;
// } IRData;
//
// if (strstr(Str, "%ssa") == nullptr) {
// IRData.HadDest = true;
// sscanf(Str, "%%ssa%d =", reinterpret_cast<int*>(&IRData.DestLoc));
// strtok(Str, " = ");
// strtok(nullptr, " = ");
// }
//
// CurrentColumn = false;
// ++CurrentLine;
// }
// // Our IR is fairly simple to lex, just spin through it
// return HadError;
return false;
}
}
-11
View File
@@ -1,11 +0,0 @@
#pragma once
namespace FEX::Debugger::IR {
class Lexer {
public:
bool Lex(char const *IR);
private:
};
}
-188
View File
@@ -1,188 +0,0 @@
#include "Common/ArgumentLoader.h"
#include "Common/EnvironmentLoader.h"
#include "Common/Config.h"
#include "Tests/HarnessHelpers.h"
#include "MainWindow.h"
#include "Context.h"
#include "GLUtils.h"
#include "IMGui_I.h"
#include "DebuggerState.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/Context.h>
#include <FEXCore/Core/CodeLoader.h>
#include <FEXCore/Debug/ContextDebug.h>
#include <FEXCore/Utils/Event.h>
#include <FEXCore/Utils/LogManager.h>
#include <imgui.h>
#include <epoxy/gl.h>
#include <memory>
#include <GLFW/glfw3.h>
#include <thread>
#include "imgui_impl_glfw.h"
#include "imgui_impl_opengl3.h"
Event SteppingEvent;
FEXCore::Context::Context *CTX{};
std::atomic_bool ShouldClose = false;
std::thread CoreThread;
void ExitHandler(uint64_t thread, FEXCore::Context::ExitReason reason) {
if (ShouldClose)
FEXCore::Context::Stop(CTX);
FEX::DebuggerState::SetHasNewState();
FEX::DebuggerState::CallNewState();
}
void StepCallback() {
FEXCore::Context::Step(CTX);
}
void CreateCoreCallback(char const *Filename, bool ELF) {
CTX = FEXCore::Context::CreateNewContext();
FEXCore::Config::SetConfig(CTX, FEXCore::Config::CONFIG_DEFAULTCORE, FEX::DebuggerState::GetCoreType());
FEXCore::Context::InitializeContext(CTX);
bool Result{};
if (ELF) {
FEX::HarnessHelper::ELFCodeLoader Loader{Filename, {}, {}, {}};
Result = FEXCore::Context::InitCore(CTX, Loader.DefaultRIP(), Loader.GetStackPointer());
}
else {
std::string ConfigName = Filename;
ConfigName.erase(ConfigName.end() - 3, ConfigName.end());
ConfigName += "config.bin";
LogMan::Msg::IFmt("Opening '{}'", Filename);
LogMan::Msg::IFmt("Opening '{}'", ConfigName);
FEX::HarnessHelper::HarnessCodeLoader Loader{Filename, ConfigName.c_str()};
Result = FEXCore::Context::InitCore(CTX, Loader.DefaultRIP(), Loader.GetStackPointer());
}
LOGMAN_THROW_A_FMT(Result, "Couldn't initialize CTX");
FEX::DebuggerState::SetContext(CTX);
FEX::DebuggerState::SetHasNewState();
FEX::DebuggerState::CallNewState();
FEXCore::Context::SetExitHandler(CTX, ExitHandler);
}
void CompileRIPCallback(uint64_t RIP) {
FEXCore::Context::Debug::CompileRIP(CTX, RIP);
FEX::DebuggerState::SetHasNewState();
FEX::DebuggerState::CallNewState();
}
void CloseCallback() {
FEXCore::Context::Stop(CTX);
if (CTX) {
FEXCore::Context::DestroyContext(CTX);
}
CTX = nullptr;
ShouldClose = false;
FEX::DebuggerState::SetContext(CTX);
FEX::DebuggerState::SetHasNewState();
FEX::DebuggerState::CallNewState();
}
void PauseCallback() {
FEXCore::Context::Pause(CTX);
FEX::DebuggerState::SetHasNewState();
FEX::DebuggerState::CallNewState();
}
void ContinueCallback() {
FEXCore::Config::SetConfig(CTX, FEXCore::Config::CONFIG_SINGLESTEP, FEX::DebuggerState::GetRunningMode());
SteppingEvent.NotifyAll();
}
void GetIRCallback(std::stringstream *out, uint64_t PC) {
if (!CTX) {
*out << "<No Core>";
return;
}
// FEXCore::IR::IntrusiveIRList *ir;
// bool HadIR = FEXCore::Context::Debug::FindIRForRIP(FEX::DebuggerState::GetContext(), PC, &ir);
// if (HadIR) {
// FEXCore::IR::Dump(out, ir);
// }
// else {
// *out << "<No IR Found>";
// }
}
int main(int argc, char **argv, char **const envp) {
FEX::EnvLoader::Load(envp);
FEXCore::Config::Initialize();
FEXCore::Config::AddLayer(std::make_unique<FEX::Config::MainLoader>());
FEXCore::Config::AddLayer(std::make_unique<FEX::ArgLoader::ArgLoader>(argc, argv));
FEXCore::Config::AddLayer(std::make_unique<FEX::Config::EnvLoader>(envp));
FEXCore::Config::Load();
FEXCore::Context::InitializeStaticTables();
auto Context = GLContext::CreateContext();
Context->Create("FEX Debugger");
FEX::Debugger::Init();
FEX::DebuggerState::RegisterStepCallback(StepCallback);
FEX::DebuggerState::RegisterPauseCallback(PauseCallback);
FEX::DebuggerState::RegisterContinueCallback(ContinueCallback);
FEX::DebuggerState::RegisterCreateCallback(CreateCoreCallback);
FEX::DebuggerState::RegisterCompileRIPCallback(CompileRIPCallback);
FEX::DebuggerState::RegisterCloseCallback(CloseCallback);
FEX::DebuggerState::RegisterGetIRCallback(GetIRCallback);
// Setup Dear ImGui context
IMGUI_CHECKVERSION();
ImGui::CreateContext();
ImGuiIO& io = ImGui::GetIO();
io.ConfigFlags |= ImGuiConfigFlags_NavEnableKeyboard; // Enable Keyboard Controls
io.ConfigFlags |= ImGuiConfigFlags_DockingEnable; // Enable Docking
io.ConfigFlags |= ImGuiConfigFlags_ViewportsEnable; // Enable Multi-Viewport / Platform Windows
GLFWwindow *Window = static_cast<GLFWwindow*>(Context->GetWindow());
// Setup Platform/Renderer bindings
ImGui_ImplGlfw_InitForOpenGL(Window, true);
const char* glsl_version = "#version 410";
ImGui_ImplOpenGL3_Init(glsl_version);
while (!glfwWindowShouldClose(Window)) {
glfwPollEvents();
// Start the Dear ImGui frame
ImGui_ImplOpenGL3_NewFrame();
ImGui_ImplGlfw_NewFrame();
FEX::Debugger::DrawDebugUI(Context.get());
glClear(GL_COLOR_BUFFER_BIT);
ImGui_ImplOpenGL3_RenderDrawData(ImGui::GetDrawData());
if (io.ConfigFlags & ImGuiConfigFlags_ViewportsEnable)
{
GLFWwindow* backup_current_context = glfwGetCurrentContext();
ImGui::UpdatePlatformWindows();
ImGui::RenderPlatformWindowsDefault();
glfwMakeContextCurrent(backup_current_context);
}
Context->Swap();
}
FEX::Debugger::Shutdown();
// Cleanup
ImGui_ImplOpenGL3_Shutdown();
ImGui_ImplGlfw_Shutdown();
ImGui::DestroyContext();
Context->Shutdown();
return 0;
}
-1
View File
@@ -1 +0,0 @@
#include "MainWindow.h"
-3
View File
@@ -1,3 +0,0 @@
#pragma once
#include <memory>
@@ -1,55 +0,0 @@
#pragma once
#include <cstddef>
#include <sys/types.h>
#include <vector>
namespace FEX::Debugger::Util {
template<typename T>
class DataRingBuffer {
public:
DataRingBuffer(size_t Size)
: RingSize {Size}
, BufferSize {RingSize * SIZE_MULTIPLE} {
Data.resize(BufferSize);
}
T const* operator()() const {
return &Data.at(ReadOffset);
}
T back() const {
return Data.at(WriteOffset - 1);
}
bool empty() const {
return WriteOffset == 0;
}
size_t size() const {
return WriteOffset - ReadOffset;
}
void push_back(T Val) {
if (WriteOffset + 1 >= BufferSize) {
// We reached the end of the buffer size. Time to wrap around
// Copy the ring buffer expected size to the start
memcpy(&Data.at(0), &Data.at(ReadOffset), sizeof(T) * RingSize);
WriteOffset = RingSize;
ReadOffset = 0;
}
Data.at(WriteOffset) = Val;
++WriteOffset;
if (WriteOffset - ReadOffset > RingSize) {
++ReadOffset;
}
}
private:
constexpr static ssize_t SIZE_MULTIPLE = 3;
size_t RingSize;
size_t WriteOffset{};
size_t ReadOffset{};
size_t BufferSize;
std::vector<T> Data;
};
}
-10
View File
@@ -237,7 +237,6 @@ namespace {
void FillCPUConfig() {
char BlockSize[32]{};
char EmulatedCPUCores[32]{};
if (ImGui::BeginTabItem("CPU")) {
std::optional<fextl::string*> Value{};
@@ -272,15 +271,6 @@ namespace {
ConfigChanged = true;
}
Value = LoadedConfig->Get(FEXCore::Config::ConfigOption::CONFIG_THREADS);
if (Value.has_value() && !(*Value)->empty()) {
strncpy(EmulatedCPUCores, &(*Value)->at(0), 32);
}
if (ImGui::InputText("Emulated CPU cores:", EmulatedCPUCores, 32, ImGuiInputTextFlags_EnterReturnsTrue)) {
LoadedConfig->EraseSet(FEXCore::Config::ConfigOption::CONFIG_THREADS, EmulatedCPUCores);
ConfigChanged = true;
}
ImGui::EndTabItem();
}
}
@@ -1,5 +1,6 @@
#pragma once
#ifdef _WIN32
#include <FEXCore/Utils/LogManager.h>
#include <winnt.h>
namespace FEX::ArchHelpers::Context {
@@ -58,6 +59,10 @@ static inline void SetState(PCONTEXT Context, uint64_t val) {
Context->R14 = val;
}
static inline uint64_t *GetArmGPRs(PCONTEXT Context) {
ERROR_AND_DIE_FMT("Not implemented for x86 host");
}
#endif
}
+137 -130
View File
@@ -1,148 +1,154 @@
add_subdirectory(LinuxSyscalls)
list(APPEND LIBS FEXCore Common)
if (TERMUX_BUILD)
# Termux needs android-shmem to get the shm emulation library.
list(APPEND LIBS android-shmem)
endif()
if (NOT MINGW_BUILD)
add_subdirectory(LinuxSyscalls)
function(GenerateInterpreter NAME AsInterpreter)
add_executable(${NAME}
FEXLoader.cpp
VDSO_Emulation.cpp
AOT/AOTGenerator.cpp)
# Enable FEX APIs to be used by targets that use target_link_libraries on FEXLoader
set_target_properties(${NAME} PROPERTIES ENABLE_EXPORTS 1)
target_include_directories(${NAME}
PRIVATE
${CMAKE_CURRENT_SOURCE_DIR}/
${CMAKE_BINARY_DIR}/generated
)
target_link_libraries(${NAME}
PRIVATE
${LIBS}
LinuxEmulation
${PTHREAD_LIB}
fmt::fmt
)
target_compile_definitions(${NAME} PRIVATE -DFEXLOADER_AS_INTERPRETER=${AsInterpreter})
if (CMAKE_BUILD_TYPE MATCHES "RELEASE")
target_link_options(${NAME}
PRIVATE
"LINKER:--gc-sections"
"LINKER:--strip-all"
"LINKER:--as-needed"
)
if (TERMUX_BUILD)
# Termux needs android-shmem to get the shm emulation library.
list(APPEND LIBS android-shmem)
endif()
if (NOT MINGW_BUILD)
function(GenerateInterpreter NAME AsInterpreter)
add_executable(${NAME}
FEXLoader.cpp
VDSO_Emulation.cpp
AOT/AOTGenerator.cpp)
# Enable FEX APIs to be used by targets that use target_link_libraries on FEXLoader
set_target_properties(${NAME} PROPERTIES ENABLE_EXPORTS 1)
target_include_directories(${NAME}
PRIVATE
${CMAKE_CURRENT_SOURCE_DIR}/
${CMAKE_BINARY_DIR}/generated
)
target_link_libraries(${NAME}
PRIVATE
${LIBS}
LinuxEmulation
${PTHREAD_LIB}
fmt::fmt
)
target_compile_definitions(${NAME} PRIVATE -DFEXLOADER_AS_INTERPRETER=${AsInterpreter})
if (CMAKE_BUILD_TYPE MATCHES "RELEASE")
target_link_options(${NAME}
PRIVATE
"LINKER:--gc-sections"
"LINKER:--strip-all"
"LINKER:--as-needed"
)
endif()
install(TARGETS ${NAME}
RUNTIME
DESTINATION bin
COMPONENT runtime
)
endif()
endfunction()
endfunction()
GenerateInterpreter(FEXLoader 0)
GenerateInterpreter(FEXInterpreter 1)
GenerateInterpreter(FEXLoader 0)
GenerateInterpreter(FEXInterpreter 1)
install(PROGRAMS "${PROJECT_SOURCE_DIR}/Scripts/FEXUpdateAOTIRCache.sh" DESTINATION bin RENAME FEXUpdateAOTIRCache)
install(PROGRAMS "${PROJECT_SOURCE_DIR}/Scripts/FEXUpdateAOTIRCache.sh" DESTINATION bin RENAME FEXUpdateAOTIRCache)
if (CMAKE_SYSTEM_PROCESSOR MATCHES "aarch64")
# Check for conflicting binfmt before installing
set (CONFLICTING_BINFMTS_32
${CMAKE_INSTALL_PREFIX}/share/binfmts/qemu-i386
${CMAKE_INSTALL_PREFIX}/share/binfmts/box86)
set (CONFLICTING_BINFMTS_64
${CMAKE_INSTALL_PREFIX}/share/binfmts/qemu-x86_64
${CMAKE_INSTALL_PREFIX}/share/binfmts/box64)
if (CMAKE_SYSTEM_PROCESSOR MATCHES "aarch64")
# Check for conflicting binfmt before installing
set (CONFLICTING_BINFMTS_32
${CMAKE_INSTALL_PREFIX}/share/binfmts/qemu-i386
${CMAKE_INSTALL_PREFIX}/share/binfmts/box86)
set (CONFLICTING_BINFMTS_64
${CMAKE_INSTALL_PREFIX}/share/binfmts/qemu-x86_64
${CMAKE_INSTALL_PREFIX}/share/binfmts/box64)
find_program(UPDATE_BINFMTS_PROGRAM update-binfmts)
if (UPDATE_BINFMTS_PROGRAM)
add_custom_target(binfmt_misc_32
echo "Attempting to install FEX-x86 misc now."
COMMAND "${CMAKE_SOURCE_DIR}/Scripts/CheckBinfmtNotInstall.sh" ${CONFLICTING_BINFMTS_32}
COMMAND "update-binfmts" "--importdir=${CMAKE_INSTALL_PREFIX}/share/binfmts/" "--import" "FEX-x86"
COMMAND ${CMAKE_COMMAND} -E
echo "binfmt_misc FEX-x86 installed"
)
add_custom_target(binfmt_misc_64
COMMAND ${CMAKE_COMMAND} -E
echo "Attempting to install FEX-x86_64 misc now."
COMMAND "${CMAKE_SOURCE_DIR}/Scripts/CheckBinfmtNotInstall.sh" ${CONFLICTING_BINFMTS_64}
COMMAND "update-binfmts" "--importdir=${CMAKE_INSTALL_PREFIX}/share/binfmts/" "--import" "FEX-x86_64"
COMMAND ${CMAKE_COMMAND} -E
echo "binfmt_misc FEX-x86_64 installed"
)
if(TARGET uninstall)
add_custom_target(uninstall_binfmt_misc_32
COMMAND update-binfmts --unimport FEX-x86 || (exit 0)
)
add_custom_target(uninstall_binfmt_misc_64
COMMAND update-binfmts --unimport FEX-x86_64 || (exit 0)
)
add_dependencies(uninstall uninstall_binfmt_misc_32)
add_dependencies(uninstall uninstall_binfmt_misc_64)
endif()
else()
# In the case of update-binfmts not being available (Arch for example) then we need to install manually
add_custom_target(binfmt_misc_32
COMMAND ${CMAKE_COMMAND} -E
echo "Attempting to remove FEX-x86 misc prior to install. Ignore permission denied"
COMMAND ${CMAKE_COMMAND} -E
echo -1 > /proc/sys/fs/binfmt_misc/FEX-x86 || (exit 0)
COMMAND ${CMAKE_COMMAND} -E
find_program(UPDATE_BINFMTS_PROGRAM update-binfmts)
if (UPDATE_BINFMTS_PROGRAM)
add_custom_target(binfmt_misc_32
echo "Attempting to install FEX-x86 misc now."
COMMAND ${CMAKE_COMMAND} -E
echo
':FEX-x86:M:0:\\x7fELF\\x01\\x01\\x01\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x02\\x00\\x03\\x00:\\xff\\xff\\xff\\xff\\xff\\xfe\\xfe\\x00\\x00\\x00\\x00\\xff\\xff\\xff\\xff\\xff\\xfe\\xff\\xff\\xff:${CMAKE_INSTALL_PREFIX}/bin/FEXInterpreter:POCF' > /proc/sys/fs/binfmt_misc/register
COMMAND ${CMAKE_COMMAND} -E
COMMAND "${CMAKE_SOURCE_DIR}/Scripts/CheckBinfmtNotInstall.sh" ${CONFLICTING_BINFMTS_32}
COMMAND "update-binfmts" "--importdir=${CMAKE_INSTALL_PREFIX}/share/binfmts/" "--import" "FEX-x86"
COMMAND ${CMAKE_COMMAND} -E
echo "binfmt_misc FEX-x86 installed"
)
add_custom_target(binfmt_misc_64
COMMAND ${CMAKE_COMMAND} -E
echo "Attempting to remove FEX-x86_64 misc prior to install. Ignore permission denied"
COMMAND ${CMAKE_COMMAND} -E
echo -1 > /proc/sys/fs/binfmt_misc/FEX-x86_64 || (exit 0)
COMMAND ${CMAKE_COMMAND} -E
add_custom_target(binfmt_misc_64
COMMAND ${CMAKE_COMMAND} -E
echo "Attempting to install FEX-x86_64 misc now."
COMMAND ${CMAKE_COMMAND} -E
echo
':FEX-x86_64:M:0:\\x7fELF\\x02\\x01\\x01\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x02\\x00\\x3e\\x00:\\xff\\xff\\xff\\xff\\xff\\xfe\\xfe\\x00\\x00\\x00\\x00\\xff\\xff\\xff\\xff\\xff\\xfe\\xff\\xff\\xff:${CMAKE_INSTALL_PREFIX}/bin/FEXInterpreter:POCF' > /proc/sys/fs/binfmt_misc/register
COMMAND ${CMAKE_COMMAND} -E
COMMAND "${CMAKE_SOURCE_DIR}/Scripts/CheckBinfmtNotInstall.sh" ${CONFLICTING_BINFMTS_64}
COMMAND "update-binfmts" "--importdir=${CMAKE_INSTALL_PREFIX}/share/binfmts/" "--import" "FEX-x86_64"
COMMAND ${CMAKE_COMMAND} -E
echo "binfmt_misc FEX-x86_64 installed"
)
if(TARGET uninstall)
add_custom_target(uninstall_binfmt_misc_32
COMMAND update-binfmts --unimport FEX-x86 || (exit 0)
)
add_custom_target(uninstall_binfmt_misc_64
COMMAND update-binfmts --unimport FEX-x86_64 || (exit 0)
)
if(TARGET uninstall)
add_custom_target(uninstall_binfmt_misc_32
add_dependencies(uninstall uninstall_binfmt_misc_32)
add_dependencies(uninstall uninstall_binfmt_misc_64)
endif()
else()
# In the case of update-binfmts not being available (Arch for example) then we need to install manually
add_custom_target(binfmt_misc_32
COMMAND ${CMAKE_COMMAND} -E
echo "Attempting to remove FEX-x86 misc prior to install. Ignore permission denied"
COMMAND ${CMAKE_COMMAND} -E
echo -1 > /proc/sys/fs/binfmt_misc/FEX-x86 || (exit 0)
)
add_custom_target(uninstall_binfmt_misc_64
COMMAND ${CMAKE_COMMAND} -E
echo "Attempting to install FEX-x86 misc now."
COMMAND ${CMAKE_COMMAND} -E
echo
':FEX-x86:M:0:\\x7fELF\\x01\\x01\\x01\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x02\\x00\\x03\\x00:\\xff\\xff\\xff\\xff\\xff\\xfe\\xfe\\x00\\x00\\x00\\x00\\xff\\xff\\xff\\xff\\xff\\xfe\\xff\\xff\\xff:${CMAKE_INSTALL_PREFIX}/bin/FEXInterpreter:POCF' > /proc/sys/fs/binfmt_misc/register
COMMAND ${CMAKE_COMMAND} -E
echo "binfmt_misc FEX-x86 installed"
)
add_custom_target(binfmt_misc_64
COMMAND ${CMAKE_COMMAND} -E
echo "Attempting to remove FEX-x86_64 misc prior to install. Ignore permission denied"
COMMAND ${CMAKE_COMMAND} -E
echo -1 > /proc/sys/fs/binfmt_misc/FEX-x86_64 || (exit 0)
)
COMMAND ${CMAKE_COMMAND} -E
echo "Attempting to install FEX-x86_64 misc now."
COMMAND ${CMAKE_COMMAND} -E
echo
':FEX-x86_64:M:0:\\x7fELF\\x02\\x01\\x01\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x02\\x00\\x3e\\x00:\\xff\\xff\\xff\\xff\\xff\\xfe\\xfe\\x00\\x00\\x00\\x00\\xff\\xff\\xff\\xff\\xff\\xfe\\xff\\xff\\xff:${CMAKE_INSTALL_PREFIX}/bin/FEXInterpreter:POCF' > /proc/sys/fs/binfmt_misc/register
COMMAND ${CMAKE_COMMAND} -E
echo "binfmt_misc FEX-x86_64 installed"
)
add_dependencies(uninstall uninstall_binfmt_misc_32)
add_dependencies(uninstall uninstall_binfmt_misc_64)
if(TARGET uninstall)
add_custom_target(uninstall_binfmt_misc_32
COMMAND ${CMAKE_COMMAND} -E
echo -1 > /proc/sys/fs/binfmt_misc/FEX-x86 || (exit 0)
)
add_custom_target(uninstall_binfmt_misc_64
COMMAND ${CMAKE_COMMAND} -E
echo -1 > /proc/sys/fs/binfmt_misc/FEX-x86_64 || (exit 0)
)
add_dependencies(uninstall uninstall_binfmt_misc_32)
add_dependencies(uninstall uninstall_binfmt_misc_64)
endif()
endif()
endif()
add_custom_target(binfmt_misc
DEPENDS binfmt_misc_32
DEPENDS binfmt_misc_64
)
add_custom_target(binfmt_misc
DEPENDS binfmt_misc_32
DEPENDS binfmt_misc_64
)
endif()
endif()
if (BUILD_TESTS)
add_executable(TestHarnessRunner TestHarnessRunner/HostRunner.cpp TestHarnessRunner.cpp)
set (SRCS TestHarnessRunner.cpp)
if (NOT MINGW_BUILD)
list(APPEND SRCS TestHarnessRunner/HostRunner.cpp)
list(APPEND LIBS LinuxEmulation)
endif()
add_executable(TestHarnessRunner ${SRCS})
target_include_directories(TestHarnessRunner
PRIVATE
${CMAKE_CURRENT_SOURCE_DIR}/Source/
@@ -151,23 +157,24 @@ if (BUILD_TESTS)
target_link_libraries(TestHarnessRunner
PRIVATE
${LIBS}
LinuxEmulation
${PTHREAD_LIB}
)
add_executable(IRLoader
IRLoader.cpp
)
target_include_directories(IRLoader
PRIVATE
${CMAKE_CURRENT_SOURCE_DIR}/Source/
${CMAKE_BINARY_DIR}/generated
)
target_link_libraries(IRLoader
PRIVATE
${LIBS}
LinuxEmulation
${PTHREAD_LIB}
fmt::fmt
)
if (NOT MINGW_BUILD)
add_executable(IRLoader
IRLoader.cpp
)
target_include_directories(IRLoader
PRIVATE
${CMAKE_CURRENT_SOURCE_DIR}/Source/
${CMAKE_BINARY_DIR}/generated
)
target_link_libraries(IRLoader
PRIVATE
${LIBS}
LinuxEmulation
${PTHREAD_LIB}
fmt::fmt
)
endif()
endif()
+19 -16
View File
@@ -77,7 +77,7 @@ class ELFCodeLoader final : public FEXCore::CodeLoader {
return true;
}
void *rv = Handler->GuestMmap((void*)addr, size, prot, flags, file.fd, off);
void *rv = Handler->GuestMmap(nullptr, (void*)addr, size, prot, flags, file.fd, off);
if (rv == MAP_FAILED) {
// uhoh, something went wrong
@@ -120,7 +120,7 @@ class ELFCodeLoader final : public FEXCore::CodeLoader {
if (Elf.ehdr.e_type == ET_DYN) {
// needs base address
auto TotalSize = CalculateTotalElfSize(Elf.phdrs) + (BrkBase ? BRK_SIZE : 0);
LoadBase = (uintptr_t)Handler->GuestMmap(reinterpret_cast<void*>(LoadHint), TotalSize, PROT_NONE, MAP_ANONYMOUS | MAP_PRIVATE, -1, 0);
LoadBase = (uintptr_t)Handler->GuestMmap(nullptr, reinterpret_cast<void*>(LoadHint), TotalSize, PROT_NONE, MAP_ANONYMOUS | MAP_PRIVATE, -1, 0);
if ((void*)LoadBase == MAP_FAILED) {
return {};
}
@@ -154,7 +154,7 @@ class ELFCodeLoader final : public FEXCore::CodeLoader {
}
if (BSSPageStart != BSSPageEnd) {
auto bss = Handler->GuestMmap((void*)BSSPageStart, BSSPageEnd - BSSPageStart, MapProt, MapType | MAP_ANONYMOUS, -1, 0);
auto bss = Handler->GuestMmap(nullptr, (void*)BSSPageStart, BSSPageEnd - BSSPageStart, MapProt, MapType | MAP_ANONYMOUS, -1, 0);
if ((void*)bss == MAP_FAILED) {
LogMan::Msg::EFmt("Failed to allocate BSS @ {}, {}\n", fmt::ptr(bss), errno);
return {};
@@ -404,7 +404,7 @@ class ELFCodeLoader final : public FEXCore::CodeLoader {
uint64_t StackHint = Is64BitMode() ? STACK_HINT_64 : STACK_HINT_32;
// Allocate the base of the full 128MB stack range.
StackPointerBase = Handler->GuestMmap(reinterpret_cast<void*>(StackHint), FULL_STACK_SIZE, PROT_NONE, MAP_PRIVATE | MAP_ANONYMOUS | MAP_STACK | MAP_GROWSDOWN | MAP_NORESERVE, -1, 0);
StackPointerBase = Handler->GuestMmap(nullptr, reinterpret_cast<void*>(StackHint), FULL_STACK_SIZE, PROT_NONE, MAP_PRIVATE | MAP_ANONYMOUS | MAP_STACK | MAP_GROWSDOWN | MAP_NORESERVE, -1, 0);
if (StackPointerBase == reinterpret_cast<void*>(~0ULL)) {
LogMan::Msg::EFmt("Allocating stack failed");
@@ -412,7 +412,7 @@ class ELFCodeLoader final : public FEXCore::CodeLoader {
}
// Allocate with permissions the 8MB of regular stack size.
StackPointer = reinterpret_cast<uintptr_t>(Handler->GuestMmap(
StackPointer = reinterpret_cast<uintptr_t>(Handler->GuestMmap(nullptr,
reinterpret_cast<void*>(reinterpret_cast<uint64_t>(StackPointerBase) + FULL_STACK_SIZE - StackSize()),
StackSize(), PROT_READ | PROT_WRITE, MAP_FIXED | MAP_PRIVATE | MAP_ANONYMOUS | MAP_STACK | MAP_GROWSDOWN, -1, 0));
@@ -539,7 +539,7 @@ class ELFCodeLoader final : public FEXCore::CodeLoader {
// XXX Randomise brk?
BrkStart = (uint64_t)Handler->GuestMmap((void*)BrkBase, BRK_SIZE, PROT_READ | PROT_WRITE, MAP_ANONYMOUS | MAP_PRIVATE | MAP_FIXED, -1, 0);
BrkStart = (uint64_t)Handler->GuestMmap(nullptr, (void*)BrkBase, BRK_SIZE, PROT_READ | PROT_WRITE, MAP_ANONYMOUS | MAP_PRIVATE | MAP_FIXED, -1, 0);
if ((void*)BrkStart == MAP_FAILED) {
LogMan::Msg::EFmt("Failed to allocate BRK @ {:x}, {}\n", BrkBase, errno);
@@ -571,6 +571,7 @@ class ELFCodeLoader final : public FEXCore::CodeLoader {
AuxVariables.emplace_back(auxv_t{26, HWCap2}); // AT_HWCAP2
AuxVariables.emplace_back(auxv_t{51, CalculateSignalStackSize()}); // AT_MINSIGSTKSZ
AuxPlatform = &AuxVariables.emplace_back(auxv_t{24, ~0ULL}); // AT_PLATFORM
AuxExecFN = &AuxVariables.emplace_back(auxv_t{AT_EXECFN, ~0ULL}); // AT_EXECFN
if (Is64BitMode()) {
AuxVariables.emplace_back(auxv_t{4, 0x38}); // AT_PHENT
@@ -582,7 +583,7 @@ class ELFCodeLoader final : public FEXCore::CodeLoader {
if (!VSyscallEntry) [[unlikely]] {
// If the VDSO thunk doesn't exist then we might not have a vsyscall entry.
// Newer glibc requires vsyscall to exist now. So let's allocate a buffer and stick a vsyscall in to it.
auto VSyscallPage = Handler->GuestMmap(nullptr, FHU::FEX_PAGE_SIZE, PROT_READ | PROT_WRITE, MAP_ANONYMOUS | MAP_PRIVATE, -1, 0);
auto VSyscallPage = Handler->GuestMmap(nullptr, nullptr, FHU::FEX_PAGE_SIZE, PROT_READ | PROT_WRITE, MAP_ANONYMOUS | MAP_PRIVATE, -1, 0);
constexpr static uint8_t VSyscallCode[] = {
0xcd, 0x80, // int 0x80
0xc3, // ret
@@ -624,9 +625,7 @@ class ELFCodeLoader final : public FEXCore::CodeLoader {
const fextl::vector<fextl::string> &EnvironmentVariables,
const fextl::list<auxv_t> &AuxVariables,
uint64_t *AuxTabBase,
uint64_t *AuxTabSize,
PointerType RandomNumberOffset,
PointerType PlatformNameOffset
uint64_t *AuxTabSize
) {
// Pointer list offsets
PointerType *ArgumentPointers = reinterpret_cast<PointerType*>(StackPointer + PointerSize);
@@ -725,6 +724,9 @@ class ELFCodeLoader final : public FEXCore::CodeLoader {
uint64_t PlatformNameLocation = TotalArgumentMemSize;
TotalArgumentMemSize += platform_string_max_size;
uint64_t ExecFNLocation = TotalArgumentMemSize;
TotalArgumentMemSize += Args[0].size() + 1;
// Offset the stack by how much memory we need
StackPointer -= TotalArgumentMemSize;
@@ -758,6 +760,10 @@ class ELFCodeLoader final : public FEXCore::CodeLoader {
}
}
// Setup ExecFN aux
AuxExecFN->val = StackPointer + ExecFNLocation;
strncpy(reinterpret_cast<char*>(AuxExecFN->val), Args[0].c_str(), Args[0].size() + 1);
// Stack setup
// [0, 8): Argument Count
// [8, 16): Argument Pointer 0
@@ -782,9 +788,7 @@ class ELFCodeLoader final : public FEXCore::CodeLoader {
EnvironmentVariables,
AuxVariables,
&AuxTabBase,
&AuxTabSize,
RandomNumberLocation,
PlatformNameLocation
&AuxTabSize
);
}
else {
@@ -797,9 +801,7 @@ class ELFCodeLoader final : public FEXCore::CodeLoader {
EnvironmentVariables,
AuxVariables,
&AuxTabBase,
&AuxTabSize,
RandomNumberLocation,
PlatformNameLocation
&AuxTabSize
);
}
}
@@ -917,6 +919,7 @@ class ELFCodeLoader final : public FEXCore::CodeLoader {
auxv_t *AuxRandom{};
auxv_t *AuxPlatform{};
auxv_t *AuxExecFN{};
static constexpr std::string_view platform_name_x86_64 = "x86_64";
static constexpr std::string_view platform_name_i686 = "i686";
+46 -1
View File
@@ -46,6 +46,7 @@ $end_info$
#include <queue>
#include <set>
#include <sys/auxv.h>
#include <sys/prctl.h>
#include <sys/resource.h>
#include <sys/select.h>
#include <system_error>
@@ -234,6 +235,47 @@ bool IsInterpreterInstalled() {
access("/proc/sys/fs/binfmt_misc/FEX-x86_64", F_OK) == 0);
}
namespace FEX::TSO {
void SetupTSOEmulation(FEXCore::Context::Context *CTX) {
// We need to check if these are defined or not. This is a very fresh feature.
#ifndef PR_GET_MEM_MODEL
#define PR_GET_MEM_MODEL 0x6d4d444c
#endif
#ifndef PR_SET_MEM_MODEL
#define PR_SET_MEM_MODEL 0x4d4d444c
#endif
#ifndef PR_SET_MEM_MODEL_DEFAULT
#define PR_SET_MEM_MODEL_DEFAULT 0
#endif
#ifndef PR_SET_MEM_MODEL_TSO
#define PR_SET_MEM_MODEL_TSO 1
#endif
// Check to see if this is supported.
auto Result = prctl(PR_GET_MEM_MODEL, 0, 0, 0, 0);
if (Result == -1) {
// Unsupported, early exit.
return;
}
FEX_CONFIG_OPT(TSOEnabled, TSOENABLED);
if (!TSOEnabled()) {
// TSO emulation isn't even enabled, early exit.
return;
}
if (Result == PR_SET_MEM_MODEL_DEFAULT) {
// Try to set the TSO mode if we are currently default.
Result = prctl(PR_SET_MEM_MODEL, PR_SET_MEM_MODEL_TSO, 0, 0, 0);
if (Result == 0) {
// TSO mode successfully enabled. Tell the context to disable TSO emulation through atomics.
// This flag gets inherited on thread creation, so FEX only needs to set it at the start.
CTX->SetHardwareTSOSupport(true);
}
}
}
}
int main(int argc, char **argv, char **const envp) {
auto SBRKPointer = FEXCore::Allocator::DisableSBRKAllocations();
FEXCore::Allocator::GLIBCScopedFault GLIBFaultScope;
@@ -427,7 +469,10 @@ int main(int argc, char **argv, char **const envp) {
auto CTX = FEXCore::Context::Context::CreateNewContext();
CTX->InitializeContext();
auto SignalDelegation = FEX::HLE::CreateSignalDelegator(CTX.get());
// Setup TSO hardware emulation immediately after initializing the context.
FEX::TSO::SetupTSOEmulation(CTX.get());
auto SignalDelegation = FEX::HLE::CreateSignalDelegator(CTX.get(), Program.ProgramName);
auto SyscallHandler = Loader.Is64BitMode() ? FEX::HLE::x64::CreateHandler(CTX.get(), SignalDelegation.get())
: FEX::HLE::x32::CreateHandler(CTX.get(), SignalDelegation.get(), std::move(Allocator));
+32 -74
View File
@@ -1,17 +1,12 @@
#pragma once
#include "Common/Config.h"
#include "LinuxSyscalls/Syscalls.h"
#include "Linux/Utils/ELFContainer.h"
#include "Linux/Utils/ELFSymbolDatabase.h"
#include <array>
#include <bitset>
#include <cassert>
#include <cstring>
#include <fcntl.h>
#include <sys/mman.h>
#include <sys/user.h>
#include <FEXCore/Core/CodeLoader.h>
#include <FEXCore/Core/CoreState.h>
@@ -19,6 +14,7 @@
#include <FEXCore/Utils/Allocator.h>
#include <FEXCore/Utils/BitUtils.h>
#include <FEXCore/Utils/CompilerDefs.h>
#include <FEXCore/Utils/FileLoading.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/MathUtils.h>
#include <FEXCore/fextl/fmt.h>
@@ -157,36 +153,10 @@ namespace FEX::HarnessHelper {
return Matches;
}
inline void ReadFile(fextl::string const &Filename, fextl::vector<char> *Data) {
int fd = open(Filename.c_str(), O_RDONLY | O_CLOEXEC);
if (fd == -1) {
LogMan::Msg::AFmt("Failed to open file");
}
struct stat buf;
if (fstat(fd, &buf) != 0) {
close(fd);
LogMan::Msg::AFmt("Failed to open file");
}
auto FileSize = buf.st_size;
if (FileSize <= 0) {
close(fd);
LogMan::Msg::AFmt("Failed to open file");
}
Data->resize(FileSize);
const auto ReadSize = pread(fd, Data->data(), FileSize, 0);
close(fd);
LOGMAN_THROW_AA_FMT(ReadSize == FileSize, "Failed to open file");
}
class ConfigLoader final {
public:
void Init(fextl::string const &ConfigFilename) {
ReadFile(ConfigFilename, &RawConfigFile);
FEXCore::FileLoading::LoadFile(RawConfigFile, ConfigFilename);
memcpy(&BaseConfig, RawConfigFile.data(), sizeof(ConfigStructBase));
GetEnvironmentOptions();
}
@@ -400,6 +370,8 @@ namespace FEX::HarnessHelper {
FEATURE_BMI1 = (1 << 6),
FEATURE_BMI2 = (1 << 7),
FEATURE_CLWB = (1 << 8),
FEATURE_LINUX = (1 << 9),
};
bool Requires3DNow() const { return BaseConfig.OptionHostFeatures & HostFeatures::FEATURE_3DNOW; }
@@ -411,6 +383,7 @@ namespace FEX::HarnessHelper {
bool RequiresBMI1() const { return BaseConfig.OptionHostFeatures & HostFeatures::FEATURE_BMI1; }
bool RequiresBMI2() const { return BaseConfig.OptionHostFeatures & HostFeatures::FEATURE_BMI2; }
bool RequiresCLWB() const { return BaseConfig.OptionHostFeatures & HostFeatures::FEATURE_CLWB; }
bool RequiresLinux() const { return BaseConfig.OptionHostFeatures & HostFeatures::FEATURE_LINUX; }
private:
FEX_CONFIG_OPT(ConfigDumpGPRs, DUMPGPRS);
@@ -459,9 +432,7 @@ namespace FEX::HarnessHelper {
public:
HarnessCodeLoader(fextl::string const &Filename, fextl::string const &ConfigFilename) {
TestFD = open(Filename.c_str(), O_CLOEXEC | O_RDONLY);
TestFileSize = lseek(TestFD, 0, SEEK_END);
lseek(TestFD, 0, SEEK_SET);
FEXCore::FileLoading::LoadFile(RawASMFile, Filename);
Config.Init(ConfigFilename);
}
@@ -472,10 +443,10 @@ namespace FEX::HarnessHelper {
uint64_t GetStackPointer() override {
if (Config.Is64BitMode()) {
return reinterpret_cast<uint64_t>(FEXCore::Allocator::mmap(nullptr, STACK_SIZE, PROT_READ | PROT_WRITE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0)) + STACK_SIZE;
return reinterpret_cast<uint64_t>(FEXCore::Allocator::VirtualAlloc(STACK_SIZE)) + STACK_SIZE;
}
else {
uint64_t Result = reinterpret_cast<uint64_t>(FEXCore::Allocator::mmap(reinterpret_cast<void*>(STACK_OFFSET), STACK_SIZE, PROT_READ | PROT_WRITE, MAP_FIXED_NOREPLACE | MAP_PRIVATE | MAP_ANONYMOUS, -1, 0));
uint64_t Result = reinterpret_cast<uint64_t>(FEXCore::Allocator::VirtualAlloc(reinterpret_cast<void*>(STACK_OFFSET), STACK_SIZE));
LOGMAN_THROW_AA_FMT(Result != ~0ULL, "Stack Pointer mmap failed");
return Result + STACK_SIZE;
}
@@ -485,39 +456,51 @@ namespace FEX::HarnessHelper {
return RIP;
}
bool MapMemoryInternal(auto Mapper, auto Munmap) {
bool MapMemory() {
bool LimitedSize = true;
auto DoMMap = [](auto This, auto Mapper, uint64_t Address, size_t Size) -> void* {
void *Result = (This->*Mapper)(reinterpret_cast<void*>(Address), Size, PROT_READ | PROT_WRITE | PROT_EXEC, MAP_FIXED | MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
auto DoMMap = [](uint64_t Address, size_t Size) -> void* {
void *Result = FEXCore::Allocator::VirtualAlloc(reinterpret_cast<void*>(Address), Size, true);
LOGMAN_THROW_AA_FMT(Result == reinterpret_cast<void*>(Address), "Map Memory mmap failed");
return Result;
};
#ifndef _WIN32
const auto AllocPageSize = FHU::FEX_PAGE_SIZE;
#else
const auto AllocPageSize = 64 * 1024;
#endif
if (LimitedSize) {
DoMMap(this, Mapper, 0xe000'0000, FHU::FEX_PAGE_SIZE * 10);
DoMMap(0xe000'0000, AllocPageSize * 10);
// SIB8
// We test [-128, -126] (Bottom)
// We test [-8, 8] (Middle)
// We test [120, 127] (Top)
// Can fit in two pages
DoMMap(this, Mapper, 0xe800'0000 - FHU::FEX_PAGE_SIZE, FHU::FEX_PAGE_SIZE * 2);
DoMMap(0xe800'0000 - AllocPageSize, AllocPageSize * 2);
}
else {
// This is scratch memory location and SIB8 location
DoMMap(this, Mapper, 0xe000'0000, 0x1000'0000);
DoMMap(0xe000'0000, 0x1000'0000);
// This is for large SIB 32bit displacement testing
DoMMap(this, Mapper, 0x2'0000'0000, 0x1'0000'1000);
DoMMap(0x2'0000'0000, 0x1'0000'1000);
}
// Map in the memory region for the test file
size_t Length = FEXCore::AlignUp(TestFileSize, FHU::FEX_PAGE_SIZE);
(this->*Mapper)(reinterpret_cast<void*>(Code_start_page), Length, PROT_READ | PROT_WRITE | PROT_EXEC, MAP_FIXED | MAP_PRIVATE, TestFD, 0);
#ifndef _WIN32
size_t Length = FEXCore::AlignUp(RawASMFile.size(), FHU::FEX_PAGE_SIZE);
auto ASMPtr = FEXCore::Allocator::VirtualAlloc(reinterpret_cast<void*>(Code_start_page), Length, true);
#else
// Special magic DOS area that starts at 0x1'0000
auto ASMPtr = FEXCore::Allocator::VirtualAlloc(reinterpret_cast<void*>(1), 0x110000 - 1, true);
#endif
LOGMAN_THROW_A_FMT((uint64_t)ASMPtr == Code_start_page, "Couldn't allocate code at expected page: 0x{:x} != 0x{:x}", (uint64_t)ASMPtr, Code_start_page);
memcpy(ASMPtr, RawASMFile.data(), RawASMFile.size());
RIP = Code_start_page;
// Map the memory regions the test file asks for
for (auto& [region, size] : Config.GetMemoryRegions()) {
DoMMap(this, Mapper, region, size);
DoMMap(region, size);
}
LoadMemory();
@@ -525,15 +508,6 @@ namespace FEX::HarnessHelper {
return true;
}
bool MapMemory(FEX::HLE::SyscallHandler *const Handler) {
SyscallHandler = Handler;
return MapMemoryInternal(&FEX::HarnessHelper::HarnessCodeLoader::SyscallMMap, &FEX::HarnessHelper::HarnessCodeLoader::SyscallMunmap);
}
bool MapMemory() {
return MapMemoryInternal(&FEX::HarnessHelper::HarnessCodeLoader::MMap, &FEX::HarnessHelper::HarnessCodeLoader::Munmap);
}
void LoadMemory() {
// Memory base here starts at the start location we passed back with GetLayout()
// This will write at [CODE_START_RANGE + 0, RawFile.size() )
@@ -558,32 +532,16 @@ namespace FEX::HarnessHelper {
bool RequiresBMI1() const { return Config.RequiresBMI1(); }
bool RequiresBMI2() const { return Config.RequiresBMI2(); }
bool RequiresCLWB() const { return Config.RequiresCLWB(); }
bool RequiresLinux() const { return Config.RequiresLinux(); }
protected:
void *SyscallMMap(void *addr, size_t length, int prot, int flags, int fd, off_t offset) {
return SyscallHandler->GuestMmap(addr, length, prot, flags, fd, offset);
}
int SyscallMunmap(void *addr, size_t length) {
return SyscallHandler->GuestMunmap(addr, length);
}
void *MMap(void *addr, size_t length, int prot, int flags, int fd, off_t offset) {
return mmap(addr, length, prot, flags, fd, offset);
}
int Munmap(void *addr, size_t length) {
return munmap(addr, length);
}
private:
FEX::HLE::SyscallHandler *SyscallHandler{};
constexpr static uint64_t STACK_SIZE = FHU::FEX_PAGE_SIZE;
constexpr static uint64_t STACK_OFFSET = 0xc000'0000;
// Zero is special case to know when we are done
uint64_t Code_start_page = 0x1'0000;
uint64_t RIP {};
int TestFD{};
size_t TestFileSize{};
fextl::vector<char> RawASMFile;
ConfigLoader Config;
};
+3 -4
View File
@@ -158,9 +158,8 @@ class DummySyscallHandler: public FEXCore::HLE::SyscallHandler, public FEXCore::
}
// These are no-ops implementations of the SyscallHandler API
std::shared_mutex StubMutex;
FEXCore::HLE::AOTIRCacheEntryLookupResult LookupAOTIRCacheEntry(uint64_t GuestAddr) override {
return {0, 0, FHU::ScopedSignalMaskWithSharedLock {StubMutex}};
FEXCore::HLE::AOTIRCacheEntryLookupResult LookupAOTIRCacheEntry(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestAddr) override {
return {0, 0};
}
};
@@ -188,7 +187,7 @@ int main(int argc, char **argv, char **const envp)
auto CTX = FEXCore::Context::Context::CreateNewContext();
CTX->InitializeContext();
auto SignalDelegation = FEX::HLE::CreateSignalDelegator(CTX.get());
auto SignalDelegation = FEX::HLE::CreateSignalDelegator(CTX.get(), {});
CTX->SetSignalDelegator(SignalDelegation.get());
CTX->SetSyscallHandler(new DummySyscallHandler());
@@ -20,7 +20,6 @@ $end_info$
#include <FEXCore/fextl/string.h>
#include <FEXCore/fextl/vector.h>
#include <FEXHeaderUtils/Filesystem.h>
#include <FEXHeaderUtils/ScopedSignalMask.h>
#include <FEXHeaderUtils/Syscalls.h>
#include <algorithm>
@@ -1269,6 +1269,13 @@ namespace FEX::HLE {
// Ref count our faults
// We use this to track if it is safe to clear cache
--Thread->CurrentFrame->SignalHandlerRefCounter;
if (Thread->DeferredSignalFrames.size() != 0) {
// If we have more deferred frames to process then mprotect back to PROT_NONE.
// It will have been RW coming in to this sigreturn and now we need to remove permissions
// to ensure FEX trampolines back to the SIGSEGV deferred handler.
mprotect(reinterpret_cast<void*>(Thread->CurrentFrame->State.DeferredSignalFaultAddress), 4096, PROT_NONE);
}
return true;
}
@@ -1376,13 +1383,104 @@ namespace FEX::HLE {
/** @} */
static bool IsAsyncSignal(const siginfo_t* Info, int Signal) {
if (Info->si_code <= SI_USER) {
// If the signal is not from the kernel then it is always async.
// This is because synchronous signals can be sent through tgkill,sigqueue and other methods.
// SI_USER == 0 and all negative si_code values come from the user.
return true;
}
else {
// If the signal is from the kernel then it is async only if it isn't an explicit synchronous signal.
switch (Signal) {
// These are all synchronous signals.
case SIGBUS:
case SIGFPE:
case SIGILL:
case SIGSEGV:
case SIGTRAP:
return false;
default: break;
}
}
// Everything else is async and can be deferred.
return true;
}
void SignalDelegator::HandleGuestSignal(FEXCore::Core::InternalThreadState *Thread, int Signal, void *Info, void *UContext) {
ucontext_t* _context = (ucontext_t*)UContext;
auto SigInfo = *static_cast<siginfo_t*>(Info);
constexpr bool SupportDeferredSignals = true;
if (SupportDeferredSignals) {
auto MustDeferSignal = (Thread->CurrentFrame->State.DeferredSignalRefCount.Load() != 0);
if (Signal == SIGSEGV &&
SigInfo.si_code == SEGV_ACCERR &&
SigInfo.si_addr == reinterpret_cast<void*>(Thread->CurrentFrame->State.DeferredSignalFaultAddress)) {
if (!MustDeferSignal) {
// We just reached the end of the outermost signal-deferring section and faulted to check for pending signals.
// Pull a signal frame off the stack.
mprotect(reinterpret_cast<void*>(Thread->CurrentFrame->State.DeferredSignalFaultAddress), 4096, PROT_READ | PROT_WRITE);
if (Thread->DeferredSignalFrames.empty()) {
// No signals to defer. Just set the fault page back to RW and continue execution.
// This occurs as a minor race condition between the refcount decrement and the access to the fault page.
return;
}
auto Top = Thread->DeferredSignalFrames.back();
Signal = Top.Signal;
SigInfo = Top.Info;
Thread->DeferredSignalFrames.pop_back();
// Until we re-protect the page to PROT_NONE, FEX will now *permanently* defer signals and /not/ check them.
//
// In order to return /back/ to a sane state, we wait for the rt_sigreturn to happen.
// rt_sigreturn will check if there are any more deferred signals to handle
// - If there are deferred signals
// - mprotect back to PROT_NONE
// - sigreturn will trampoline out to the previous fault address check, SIGSEGV and restart
// - If there are *no* deferred signals
// - No need to mprotect, it is already RW
}
else {
#ifdef _M_ARM_64
// If RefCount != 0 then that means we hit an access with nested signal-deferring sections.
// Increment the PC past the `str zr, [x1]` to continue code execution until we reach the outermost section.
ArchHelpers::Context::SetPc(UContext, ArchHelpers::Context::GetPc(UContext) + 4);
return;
#else
// X86 should always be doing a refcount compare and branch since we can't guarantee instruction size.
// ARM64 just always does the access to reduce branching overhead.
ERROR_AND_DIE_FMT("X86 shouldn't hit this DeferredSignalFaultAddress");
#endif
}
}
else {
if (IsAsyncSignal(&SigInfo, Signal) && MustDeferSignal) {
// If the signal is asynchronous (as determined by si_code) and FEX is in a state of needing
// to defer the signal, then add the signal to the thread's signal queue.
LOGMAN_THROW_A_FMT(Thread->DeferredSignalFrames.size() != Thread->DeferredSignalFrames.capacity(),
"Deferred signals vector hit capacity size. This will likely crash! Asserting now!");
Thread->DeferredSignalFrames.emplace_back(FEXCore::Core::InternalThreadState::DeferredSignalState {
.Info = SigInfo,
.Signal = Signal,
});
// Now update the faulting page permissions so it will fault on write.
mprotect(reinterpret_cast<void*>(Thread->CurrentFrame->State.DeferredSignalFaultAddress), 4096, PROT_NONE);
// Postpone the remainder of signal handling logic until we process the SIGSEGV triggered by writing to DeferredSignalFaultAddress.
return;
}
}
}
// Let the host take first stab at handling the signal
SignalHandler &Handler = HostHandlers[Signal];
ucontext_t* _context = (ucontext_t*)UContext;
const auto SigInfo = static_cast<siginfo_t*>(Info);
// Remove the pending signal
ThreadData.PendingSignals &= ~(1ULL << (Signal - 1));
@@ -1399,8 +1497,7 @@ namespace FEX::HLE {
}
else {
if (Handler.GuestHandler &&
Handler.GuestHandler(Thread, Signal, Info, UContext, &Handler.GuestAction, &ThreadData.GuestAltStack)) {
Handler.GuestHandler(Thread, Signal, &SigInfo, UContext, &Handler.GuestAction, &ThreadData.GuestAltStack)) {
// Set up a new mask based on this signals signal mask
uint64_t NewMask = Handler.GuestAction.sa_mask.Val;
@@ -1430,7 +1527,7 @@ namespace FEX::HLE {
// Unhandled crash
// Call back in to the previous handler
if (Handler.OldAction.sa_flags & SA_SIGINFO) {
Handler.OldAction.sigaction(Signal, SigInfo, UContext);
Handler.OldAction.sigaction(Signal, &SigInfo, UContext);
}
else if (Handler.OldAction.handler == SIG_IGN ||
(Handler.OldAction.handler == SIG_DFL &&
@@ -1440,9 +1537,18 @@ namespace FEX::HLE {
else if (Handler.OldAction.handler == SIG_DFL &&
(Handler.DefaultBehaviour == DEFAULT_COREDUMP ||
Handler.DefaultBehaviour == DEFAULT_TERM)) {
// In the case of signals that cause coredump or terminate, save telemetry early.
// FEX is hard crashing at this point and won't hit regular shutdown routines.
// Add the signal to the crash mask.
CrashMask |= (1ULL << Signal);
if (!ApplicationName.empty()) {
FEXCore::Telemetry::Shutdown(ApplicationName);
}
// Reassign back to DFL and crash
signal(Signal, SIG_DFL);
if (SigInfo->si_code != SI_KERNEL) {
if (SigInfo.si_code != SI_KERNEL) {
// If the signal wasn't sent by the kernel then we need to reraise it.
// This is necessary since returning from this signal handler now might just continue executing.
// eg: If sent from tgkill then the signal gets dropped and returns.
@@ -1550,8 +1656,9 @@ namespace FEX::HLE {
::syscall(SYS_rt_sigaction, Signal, &SignalHandler.OldAction, nullptr, 8);
}
SignalDelegator::SignalDelegator(FEXCore::Context::Context *_CTX)
: CTX {_CTX} {
SignalDelegator::SignalDelegator(FEXCore::Context::Context *_CTX, const std::string_view ApplicationName)
: CTX {_CTX}
, ApplicationName {ApplicationName} {
// Register this delegate
LOGMAN_THROW_AA_FMT(!GlobalDelegator, "Can't register global delegator multiple times!");
GlobalDelegator = this;
@@ -1702,6 +1809,12 @@ namespace FEX::HLE {
// Get the current host signal mask
::syscall(SYS_rt_sigprocmask, 0, nullptr, &ThreadData.CurrentSignalMask.Val, 8);
if (Thread != (FEXCore::Core::InternalThreadState*)UINTPTR_MAX) {
// Reserve a small amount of deferred signal frames. Usually the stack won't be utilized beyond
// 1 or 2 signals but add a few more just in case.
Thread->DeferredSignalFrames.reserve(8);
}
}
void SignalDelegator::UninstallFrontendTLSState(FEXCore::Core::InternalThreadState *Thread) {
@@ -2020,7 +2133,7 @@ namespace FEX::HLE {
return Result == -1 ? -errno : Result;
}
fextl::unique_ptr<FEX::HLE::SignalDelegator> CreateSignalDelegator(FEXCore::Context::Context *CTX) {
return fextl::make_unique<FEX::HLE::SignalDelegator>(CTX);
fextl::unique_ptr<FEX::HLE::SignalDelegator> CreateSignalDelegator(FEXCore::Context::Context *CTX, const std::string_view ApplicationName) {
return fextl::make_unique<FEX::HLE::SignalDelegator>(CTX, ApplicationName);
}
}
@@ -22,6 +22,7 @@ $end_info$
#include <mutex>
#include <FEXCore/Core/SignalDelegator.h>
#include <FEXCore/Utils/Telemetry.h>
namespace FEXCore {
namespace Context {
@@ -40,7 +41,7 @@ namespace FEX::HLE {
public:
// Returns true if the host handled the signal
// Arguments are the same as sigaction handler
SignalDelegator(FEXCore::Context::Context *_CTX);
SignalDelegator(FEXCore::Context::Context *_CTX, const std::string_view ApplicationName);
~SignalDelegator() override;
/**
@@ -118,6 +119,8 @@ namespace FEX::HLE {
private:
FEX_CONFIG_OPT(Is64BitMode, IS64BIT_MODE);
FEX_CONFIG_OPT(Core, CORE);
fextl::string const ApplicationName;
FEXCORE_TELEMETRY_INIT(CrashMask, TYPE_CRASH_MASK);
enum DefaultBehaviour {
DEFAULT_TERM,
@@ -231,5 +234,5 @@ namespace FEX::HLE {
std::mutex GuestDelegatorMutex;
};
fextl::unique_ptr<FEX::HLE::SignalDelegator> CreateSignalDelegator(FEXCore::Context::Context *CTX);
fextl::unique_ptr<FEX::HLE::SignalDelegator> CreateSignalDelegator(FEXCore::Context::Context *CTX, const std::string_view ApplicationName);
}
@@ -37,7 +37,6 @@ $end_info$
#include <FEXCore/fextl/vector.h>
#include <FEXHeaderUtils/Filesystem.h>
#include <FEXHeaderUtils/Syscalls.h>
#include <FEXHeaderUtils/ScopedSignalMask.h>
#include <FEXHeaderUtils/TypeDefines.h>
#include <algorithm>
@@ -61,56 +60,21 @@ namespace FEX::HLE {
class SignalDelegator;
SyscallHandler *_SyscallHandler{};
template<bool IncrementOffset, typename T>
uint64_t GetDentsEmulation(int fd, T *dirp, uint32_t count) {
fextl::vector<uint8_t> TmpVector(count);
void *TmpPtr = reinterpret_cast<void*>(&TmpVector.at(0));
uint64_t Offset = 0;
uint64_t TmpOffset = 0;
// Copy the incoming structures to our temporary array
while (Offset < count) {
T *Incoming = (T*)(reinterpret_cast<uint64_t>(dirp) + Offset);
FEX::HLE::x64::linux_dirent_64 *Tmp = (FEX::HLE::x64::linux_dirent_64*)(reinterpret_cast<uint64_t>(TmpPtr) + TmpOffset);
if (!Incoming->d_reclen ||
(Offset + Incoming->d_reclen) > count) {
break;
}
size_t NewRecLen = FEXCore::AlignUp(Incoming->d_reclen + (sizeof(std::remove_reference<decltype(*Tmp)>::type) - sizeof(*Incoming)),
alignof(decltype(Tmp->d_ino)));
Tmp->d_ino = Incoming->d_ino;
Tmp->d_off = Incoming->d_off;
Tmp->d_reclen = NewRecLen;
// d_type is hidden at the very end of reclen
Tmp->d_type = Incoming->d_name[Incoming->d_reclen - offsetof(T, d_name) - 1];
// This actually copies one more byte than the string of d_name
// Copies a null byte for the string
size_t CopySize = std::clamp<uint32_t>(Incoming->d_reclen - offsetof(T, d_name) - 1, 0U, count - Offset);
memcpy(Tmp->d_name, Incoming->d_name, CopySize);
// We take up 8 more bytes of space
TmpOffset += NewRecLen;
Offset += Incoming->d_reclen;
}
uint64_t Result = syscall(SYSCALL_DEF(getdents64),
static_cast<uint64_t>(fd),
TmpPtr,
dirp,
static_cast<uint64_t>(count));
// Now copy back in to the array we were given
if (Result != -1) {
// If the outgoing d_ino is smaller than the incoming d_ino from the kernel
// Then we need to check for overflow before writing any of the data back
if (sizeof(decltype(FEX::HLE::x64::linux_dirent_64::d_ino)) > sizeof(decltype(T::d_ino))) {
if constexpr (sizeof(decltype(FEX::HLE::x64::linux_dirent_64::d_ino)) > sizeof(decltype(T::d_ino))) {
uint64_t TmpOffset = 0;
while (TmpOffset < Result) {
FEX::HLE::x64::linux_dirent_64 *Tmp = (FEX::HLE::x64::linux_dirent_64*)(reinterpret_cast<uint64_t>(TmpPtr) + TmpOffset);
FEX::HLE::x64::linux_dirent_64 *Tmp = (FEX::HLE::x64::linux_dirent_64*)(reinterpret_cast<uint64_t>(dirp) + TmpOffset);
decltype(T::d_ino) Result_d_ino = Tmp->d_ino;
if (Result_d_ino != Tmp->d_ino) {
@@ -124,10 +88,13 @@ uint64_t GetDentsEmulation(int fd, T *dirp, uint32_t count) {
uint64_t Offset = 0;
uint64_t TmpOffset = 0;
size_t OffsetIndex = 1;
// With how the emulation occurs we will always return a smaller buffer than what was given to us
// With how the emulation occurs we will always return a smaller buffer than what was given to us.
// We need to be careful with the in-place translation that occurs here, the data returning to the guest is guaranteed to be smaller
// than the data returned by getdents64.
// This means FEX is guaranteed to /never/ fill the full getdents buffer to the guest, but we may temporarily use it all.
while (TmpOffset < Result) {
T *Outgoing = (T*)(reinterpret_cast<uint64_t>(dirp) + Offset);
FEX::HLE::x64::linux_dirent_64 *Tmp = (FEX::HLE::x64::linux_dirent_64*)(reinterpret_cast<uint64_t>(TmpPtr) + TmpOffset);
FEX::HLE::x64::linux_dirent_64 *Tmp = (FEX::HLE::x64::linux_dirent_64*)(reinterpret_cast<uint64_t>(dirp) + TmpOffset);
if (!Tmp->d_reclen) {
break;
@@ -145,7 +112,7 @@ uint64_t GetDentsEmulation(int fd, T *dirp, uint32_t count) {
// Copies null character as well
size_t NameLength = Tmp->d_reclen - OffsetOfName - 1;
memcpy(Outgoing->d_name, Tmp->d_name, NameLength);
memmove(Outgoing->d_name, Tmp->d_name, NameLength);
// Copy the hidden d_type flag
Outgoing->d_name[Outgoing->d_reclen - offsetof(T, d_name) - 1] = Tmp->d_type;
@@ -703,7 +670,7 @@ uint64_t SyscallHandler::HandleBRK(FEXCore::Core::CpuStateFrame *Frame, void *Ad
uint64_t RemainingSize = DataSpaceMaxSize - NewSizeAligned;
// We have pages we can unmap
auto ok = GuestMunmap(reinterpret_cast<void*>(DataSpace + NewSizeAligned), RemainingSize);
auto ok = GuestMunmap(Frame->Thread, reinterpret_cast<void*>(DataSpace + NewSizeAligned), RemainingSize);
LOGMAN_THROW_A_FMT(ok != -1, "Munmap failed");
DataSpaceMaxSize = NewSizeAligned;
@@ -718,13 +685,13 @@ uint64_t SyscallHandler::HandleBRK(FEXCore::Core::CpuStateFrame *Frame, void *Ad
}
uint64_t NewBRK{};
NewBRK = (uint64_t)GuestMmap((void*)(DataSpace + DataSpaceMaxSize), AllocateNewSize, PROT_READ | PROT_WRITE, MAP_FIXED_NOREPLACE | MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
NewBRK = (uint64_t)GuestMmap(Frame->Thread, (void*)(DataSpace + DataSpaceMaxSize), AllocateNewSize, PROT_READ | PROT_WRITE, MAP_FIXED_NOREPLACE | MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
if (NewBRK != ~0ULL && NewBRK != (DataSpace + DataSpaceMaxSize)) {
// Couldn't allocate that the region we wanted
// Can happen if MAP_FIXED_NOREPLACE isn't understood by the kernel
int ok = GuestMunmap(reinterpret_cast<void*>(NewBRK), AllocateNewSize);
int ok = GuestMunmap(Frame->Thread, reinterpret_cast<void*>(NewBRK), AllocateNewSize);
LOGMAN_THROW_A_FMT(ok != -1, "Munmap failed");
NewBRK = ~0ULL;
}
@@ -761,9 +728,7 @@ SyscallHandler::SyscallHandler(FEXCore::Context::Context *_CTX, FEX::HLE::Signal
GuestKernelVersion = CalculateGuestKernelVersion();
Alloc32Handler = FEX::HLE::Create32BitAllocator();
if (SMCChecks == FEXCore::Config::CONFIG_SMC_MTRACK) {
SignalDelegation->RegisterHostSignalHandler(SIGSEGV, HandleSegfault, true);
}
SignalDelegation->RegisterHostSignalHandler(SIGSEGV, HandleSegfault, true);
}
SyscallHandler::~SyscallHandler() {
@@ -790,8 +755,8 @@ uint32_t SyscallHandler::CalculateHostKernelVersion() {
}
uint32_t SyscallHandler::CalculateGuestKernelVersion() {
// We currently only emulate a kernel between the ranges of Kernel 5.0.0 and 5.18.0
return std::max(KernelVersion(5, 0), std::min(KernelVersion(5, 18), GetHostKernelVersion()));
// We currently only emulate a kernel between the ranges of Kernel 5.0.0 and 6.2.0
return std::max(KernelVersion(5, 0), std::min(KernelVersion(6, 2), GetHostKernelVersion()));
}
uint64_t SyscallHandler::HandleSyscall(FEXCore::Core::CpuStateFrame *Frame, FEXCore::HLE::SyscallArguments *Args) {
+11 -11
View File
@@ -198,25 +198,25 @@ public:
FEX::HLE::MemAllocator *Get32BitAllocator() { return Alloc32Handler.get(); }
// does a mmap as if done via a guest syscall
virtual void *GuestMmap(void *addr, size_t length, int prot, int flags, int fd, off_t offset) = 0;
virtual void *GuestMmap(FEXCore::Core::InternalThreadState *Thread, void *addr, size_t length, int prot, int flags, int fd, off_t offset) = 0;
// does a guest munmap as if done via a guest syscall
virtual int GuestMunmap(void *addr, uint64_t length) = 0;
virtual int GuestMunmap(FEXCore::Core::InternalThreadState *Thread, void *addr, uint64_t length) = 0;
///// Memory Manager tracking /////
void TrackMmap(uintptr_t Base, uintptr_t Size, int Prot, int Flags, int fd, off_t Offset);
void TrackMunmap(uintptr_t Base, uintptr_t Size);
void TrackMprotect(uintptr_t Base, uintptr_t Size, int Prot);
void TrackMremap(uintptr_t OldAddress, size_t OldSize, size_t NewSize, int flags, uintptr_t NewAddress);
void TrackShmat(int shmid, uintptr_t Base, int shmflg);
void TrackShmdt(uintptr_t Base);
void TrackMadvise(uintptr_t Base, uintptr_t Size, int advice);
void TrackMmap(FEXCore::Core::InternalThreadState *Thread, uintptr_t Base, uintptr_t Size, int Prot, int Flags, int fd, off_t Offset);
void TrackMunmap(FEXCore::Core::InternalThreadState *Thread, uintptr_t Base, uintptr_t Size);
void TrackMprotect(FEXCore::Core::InternalThreadState *Thread, uintptr_t Base, uintptr_t Size, int Prot);
void TrackMremap(FEXCore::Core::InternalThreadState *Thread, uintptr_t OldAddress, size_t OldSize, size_t NewSize, int flags, uintptr_t NewAddress);
void TrackShmat(FEXCore::Core::InternalThreadState *Thread, int shmid, uintptr_t Base, int shmflg);
void TrackShmdt(FEXCore::Core::InternalThreadState *Thread, uintptr_t Base);
void TrackMadvise(FEXCore::Core::InternalThreadState *Thread, uintptr_t Base, uintptr_t Size, int advice);
///// VMA (Virtual Memory Area) tracking /////
static bool HandleSegfault(FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext);
void MarkGuestExecutableRange(uint64_t Start, uint64_t Length) override;
void MarkGuestExecutableRange(FEXCore::Core::InternalThreadState *Thread, uint64_t Start, uint64_t Length) override;
// AOTIRCacheEntryLookupResult also includes a shared lock guard, so the pointed AOTIRCacheEntry return can be safely used
FEXCore::HLE::AOTIRCacheEntryLookupResult LookupAOTIRCacheEntry(uint64_t GuestAddr) final override;
FEXCore::HLE::AOTIRCacheEntryLookupResult LookupAOTIRCacheEntry(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestAddr) final override;
///// FORK tracking /////
void LockBeforeFork();
Loaded 100 of 138 files, more files were not shown because too many files have changed in this diff. Show more