Compare commits

..
901 Commits
Author SHA1 Message Date
Ryan Houdek b3bc1e23cc Docs: Update for release FEX-2306 2023-06-08 16:38:52 -07:00
Mai 1a9b6a89f4 Merge pull request #2706 from Sonicadvance1/remove_emulated_cores
FEXConfig: Removes Emulated CPU cores option
2023-06-08 07:39:15 -04:00
Ryan Houdek 21bf35d211 FEXConfig: Removes Emulated CPU cores option
This is just confusing end users these days and no longer matters as a
debug option.

Remove from the GUI initially, maybe afterwards we will even remove
setting this at all and always auto-detect.
2023-06-07 17:55:05 -07:00
Ryan Houdek 9473025b18 Merge pull request #2704 from Sonicadvance1/fexrootfsfetcher_arch
FEXRootFSFetcher: Support rolling release distros
2023-06-07 16:34:09 -07:00
Mai 02f15f4099 Merge pull request #2705 from Sonicadvance1/update_installfex_script
InstallFEX: Updates helper install script for Ubuntu 23.04
2023-06-07 16:34:53 -04:00
Ryan Houdek e007789ced InstallFEX: Updates helper install script for Ubuntu 23.04
Also updates the link in the source to the new json file.
2023-06-07 12:58:11 -07:00
Ryan Houdek a2b165043c FEXRootFSFetcher: Support rolling release distros
This basically just means that we detect ArchLinux and set a flag that
it is a rolling release, skipping doing the version check for an "exact"
match in that instance.
2023-06-07 12:55:38 -07:00
Ryan Houdek 5b5808218b Merge pull request #2703 from Sonicadvance1/minor_of_opt
OpcodeDispatcher: Optimize ADC/ADD OF flag calculation
2023-06-07 12:54:55 -07:00
Ryan Houdek 41ec987f3e OpcodeDispatcher: Optimize ADC/ADD OF flag calculation
`eor <reg>, <reg>, #-1` can't be encoded as an instruction. Instead use
mvn which does the same thing.

Removes a single instruction from each OF calculation for ADC and ADD.

Also no reason to use a switch statement for the source size, just use
_Bfe and calculate the offset based on operation size.

SBB caught in the crossfire to ensure it also isn't using a switch
statement.
2023-06-07 12:40:51 -07:00
Mai 0f4a5edf4f Merge pull request #2702 from Sonicadvance1/fix_ssa_dec
IRDumper: Fixes ssa number in arguments.
2023-06-07 14:18:09 -04:00
Ryan Houdek 03f73531d3 IRDumper: Fixes ssa number in arguments.
This can spuriously end up as a hex number which makes it hard to reason
why DCE wasn't deleting IR operations. Ensure it is always a decimal.
2023-06-07 09:52:04 -07:00
Mai 69181d438c Merge pull request #2701 from Sonicadvance1/optimize_flag_unpacking
OpcodeDispatcher: Optimize EFLAG unpacking
2023-06-06 21:43:31 -04:00
Ryan Houdek a2cbfccb3b OpcodeDispatcher: Optimize EFLAG unpacking
Noticed this was slightly unoptimal. Resulting in a 18% code reduction
in the case of of a simple four instruction test ASM case.
2023-06-06 17:56:25 -07:00
Mai 4e01452a65 Merge pull request #2699 from Sonicadvance1/minor_fcmov_opt
X87: Super minor FCMOV optimization
2023-06-06 20:22:40 -04:00
Mai cc7a56b1a6 Merge pull request #2689 from Sonicadvance1/fix_bmi
CPUID: Only enable BMI1 and BMI2 if AVX is supported
2023-06-06 20:21:57 -04:00
Ryan Houdek 0b0dd3891e X87: Super minor FCMOV optimization
This caught my eye as I was skimming, remove one IR op per FCMOV
instruction.

This was just duplicating the generated GPR mask across the FPR.
2023-06-04 06:39:35 -07:00
Ryan Houdek 8bc33e95c1 Merge pull request #2493 from Sonicadvance1/deferred_signals_partial
Implement support for deferred asynchronous signals
2023-06-02 22:07:18 -07:00
Ryan Houdek 96a0364a86 Review comments 2023-06-02 21:53:52 -07:00
Ryan Houdek c0a783997d Convert remaining memory tracking to deferred signals 2023-06-01 11:35:22 -07:00
Ryan Houdek f78537109d Core: Convert mtrack code invalidation over to deferred signals 2023-06-01 11:35:22 -07:00
Ryan Houdek 0c156ed6f9 Context: Switch over to deferred signals 2023-06-01 11:28:04 -07:00
Ryan Houdek 920913cf80 Syscalls: Always install SIGSEGV handler for deferred handler 2023-06-01 11:28:04 -07:00
Ryan Houdek 8840b2154c Allocator: Allow more optimal deferred signals path 2023-06-01 11:28:04 -07:00
Ryan Houdek e02be8073e FEXCore: Support deferred signal mutex
This is part of FEXCore since it pulls in InternalThreadData, but is
related to the FHU signal mutex class.

Necessary to allow deferring signals in C++ code rather than right in
the JIT.
2023-06-01 11:28:04 -07:00
Ryan Houdek f75d3550b4 Jit64: Used deferred signals in dispatcher 2023-06-01 11:28:04 -07:00
Ryan Houdek 802c588695 Arm64: Use deferred signals in dispatcher 2023-06-01 11:28:04 -07:00
Ryan Houdek fd962f40d7 SignalDelegator: Support deferring signals 2023-06-01 11:28:04 -07:00
Ryan Houdek a9b660af69 CoreState: Add new members to track deferred signal capability 2023-06-01 11:28:04 -07:00
Ryan Houdek fd5c36ba9c Docs: Adds a document explaining how FEX's deferred signals works.
This has design considerations as to why choices were made.
2023-06-01 11:28:04 -07:00
Ryan Houdek 5be798e9e6 Merge pull request #2693 from Sonicadvance1/remove_debug
Context: Remove debug namespace
2023-06-01 11:26:05 -07:00
Ryan Houdek 09997cff9c Merge pull request #2692 from Sonicadvance1/remove_debugger
Tools: Removes visual debugger
2023-06-01 11:25:55 -07:00
Ryan Houdek c9d1f0d75a Merge pull request #2687 from Sonicadvance1/telemetry_save_crash
Telemetry: Save on signal terminate
2023-05-30 10:26:03 -07:00
Ryan Houdek 95b7592241 Merge pull request #2690 from Sonicadvance1/vfork_wait
Linux: Make vfork act more similar to how it should.
2023-05-30 10:25:53 -07:00
Ryan Houdek 1dc4f8c429 Context: Remove debug namespace
Unused and broken
2023-05-30 09:00:57 -07:00
Ryan Houdek 1d7fcdb54a Tools: Removes visual debugger
Unused and broken
2023-05-30 08:53:48 -07:00
Ryan Houdek 45d3b83143 Telemetry: Save on signal terminate
When a signal handler is not installed and is a terminal failure, make
sure to save telemetry before faulting.

We know when an application is going down in this case so we can make
sure to have the telemetry data saved.

Adds a telemetry signal mask data point as well to know which signal
took it down.
2023-05-30 08:49:33 -07:00
Ryan Houdek d97fa9af14 Linux: Make vfork act more similar to how it should.
Noticed this while debugging Proton Experimental hanging and thought
this could be related. Didn't solve that issue but this should be merged
anyway.

vfork doesn't fork the host's process space in to the child process.
Saving Copy-On-Write overhead problems. It also puts the parent process
to sleep until the fork terminates or executes.

This is a major issue under FEX where we can't emulate vfork correctly
because we need to do other work before this process terminates or
executes a new process. We have been treating `vfork` as a `fork` this
entire time.

This can likely cause problems for applications that actually use vfork
to wait for a process to complete. So let's actually emulate that
feature by using a pipe with poll to determine when that FD gets
removed.

FEX can't use waitpid to wait for this process to terminate since we
would affect the guest also wanted to use a waitpid.
2023-05-30 08:44:36 -07:00
Ryan Houdek c9101d3f68 CPUID: Only enable BMI1 and BMI2 if AVX is supported
These two extensions rely on AVX being supported to be used. Primarily
because they are VEX encoded.

GTA5 is using these flags to determine if it should enable its AVX
support.
2023-05-26 20:48:36 -07:00
Mai 52f64a0c7b Merge pull request #2685 from Sonicadvance1/remove_ci_warnings
github: Updates some actions to v3
2023-05-22 22:31:40 -04:00
Mai 737f917838 Merge pull request #2686 from Sonicadvance1/xgetbv
FEXCore: Implements support for xgetbv
2023-05-22 22:31:21 -04:00
Ryan Houdek a6c6248bcb ArmEmitter: Fixes bug in SpillStaticRegs
Some code in FEX's Arm64 emitter was making an assumption that once
SpillStaticRegs was called that it was safe to still use the SRA
register state.
This wasn't actually true since FEX was using one SRA register to
optimize FPR stores. Assuming that the SRA registers were safe to use
since they were just saved and no longer necessary.

Correct this assumption hell by forcing users of the function to provide
the temporary register directly. In all cases the users have a temporary
available that it can use.

Probably fixes some very weird edge case bugs.
2023-05-22 16:48:07 -07:00
Ryan Houdek 5646428640 FEXCore: Implements support for xgetbv
This returns the `XFEATURE_ENABLED_MASK` register which reports what
features are enabled on the CPU.
This behaves similarly to CPUID where it uses an index register in ecx.

This is a prerequisite to enabling XSAVE/XRSTOR and AVX since
applications will expect this to exist.

xsetbv is a privileged instruction and doesn't need to be implemented.
2023-05-22 16:48:07 -07:00
Ryan Houdek 0c8df2beaf github: Updates some actions to v3
Removes some annotation warnings that have been showing up on the
actions results page.
v2 is deprecated so going to v3 is necessary. Apparently this upgrades
from Node.js 12 to 16.
2023-05-22 10:39:13 -07:00
Mai de0f3984e9 Merge pull request #2680 from Sonicadvance1/optimize_getdents
Syscalls: Optimize getdents{64,}
2023-05-22 11:46:23 -04:00
Mai ada226bbb4 Merge pull request #2683 from Sonicadvance1/uprev_kernel
FEXLoader: Allow simulated kernel version up to 6.2
2023-05-22 11:45:12 -04:00
Ryan Houdek 4bc5a09e62 FEXLoader: Allow simulated kernel version up to 6.2
Investigation in #2589 shows we can push it to this point.
6.3 adds a new prctl that FEX can't enable yet.
2023-05-21 09:51:34 -07:00
Ryan Houdek 5704b5f23f Syscalls: Optimize getdents{64,}
I originally wrote this emulation prior to me fully understanding how
the syscall works. So there are two optimizations here.

1) No need to consume the incoming buffer at all.
   - Originally I thought the incoming dirent structures were used to
     calculate offset.
   - This is not the case, the FD's file position is used instead.
   - This means we can remove the incoming buffer consuming overhead
     entirely.
2) No need to allocate a temporary buffer at all.
   - With getdents and getdents64 we are guaranteed to be dealing with
     structures that are the same size or smaller than the host
     structure.
   - This lets us encode the real host dirents in to the provided
     buffer.
   - After the `getdents64` host syscall, we then iterate forward
     through the list, modifying as we go.
   - Need to make sure to shift the elements of the structure in order.
   - Need to make sure to use memmove on the `d_name` member since the
     movement region can overlap.

These two optimizations significantly reduce the amount of time spent in
getdents, which has a noticeable impact on load times.

Side-tangent: I noticed a fun quirk of how NFS operates with getdents.
If the FSCache hasn't populated the metadata for that folder, then it
will early return with "some" data, not fully maxing out the buffer. The
kernel will start prefetching metadata assuming directory iterating is
happening. The next `getdents` happens and it should return a larger
number of elements.

Very neat.
2023-05-18 21:56:24 -07:00
Ryan Houdek 6017a9135a Merge pull request #2679 from Sonicadvance1/mostly_revert_2672
Thunks: Mostly reverts #2672
2023-05-18 16:11:59 -07:00
Ryan Houdek 6ef6d9c391 Thunks: Mostly reverts #2672
I forgot that x11 was part of the custom ABI of thunks. #2672 had broken
thunks on ARM64. I thought I had tested a game with them enabled but
apparently I tested the wrong game.

Not a full revert since we can still ldr with a literal, but we also
still need to adr x11 and nop pad. At least removes the data dependency
on x11 from the ldr.
2023-05-18 15:50:55 -07:00
Ryan Houdek 0ad6f98a8c Merge pull request #2666 from Sonicadvance1/wine_testharnessrunner
Wine TestHarnessRunner support
2023-05-18 12:58:35 -07:00
Ryan Houdek 1354f92cc5 Review comments 2023-05-17 21:09:31 -07:00
Ryan Houdek 8b90caad95 unittests: Adds a Linux HostFeatures flag
Disables two tests that don't work under Wine
2023-05-17 21:09:31 -07:00
Ryan Houdek 3a4a965347 TestHarnessRunner: Support exiting on HLT
Currently WINE's longjump doesn't work, so instead set a flag that if
HLT is attempted, just exit the JIT.

This will get our unittests executing at least.
2023-05-17 21:09:31 -07:00
Ryan Houdek 45cdab2ac3 HostFeatures: Use ID registers under Wine
InferFromOS doesn't work under WINE.
InferFromIDRegisters doesn't work under Windows but it will under Wine.

Since we don't support Windows, just use InferFromIDRegisters.
2023-05-17 21:07:40 -07:00
Ryan Houdek b89dc56ae1 unittests: Update test so it can work on wine.
We don't necessarily care where this memory is, just that it can be
allocated. Move it to a memory location that works on both Linux and
Wine.
2023-05-17 21:07:40 -07:00
Ryan Houdek d675b4af6f External: Update vixl 2023-05-17 21:07:40 -07:00
Ryan Houdek d75fb38344 TestHarnessRunner: Get running on Win32 2023-05-17 21:07:40 -07:00
Ryan Houdek 4cb385a27b unittests: Build ASM tests on win32 2023-05-17 21:07:37 -07:00
Ryan Houdek 9a4fdd8059 ArchHelpers: Adds missing WinContext stub 2023-05-17 21:05:55 -07:00
Ryan Houdek 363411f0c7 ArchHelpers: Adds missing stub function 2023-05-17 21:05:55 -07:00
Ryan Houdek 5bc418407c FEXCore: Disable emitter unit tests on win32 2023-05-17 21:05:55 -07:00
Ryan Houdek 46e2dc7498 Common: Disable some Linux specific files on win32 2023-05-17 21:05:55 -07:00
Ryan Houdek 4a54197868 TestHarnessRunner: Use VirtualAlloc for mapping regions.
Needs to be alligned to allocation size. Which is a page on Linux, or
64k on Windows.

In order to map at `0x1'0000` on Wine, we need to use a special case DOS
area allocation path.
2023-05-17 21:05:55 -07:00
Ryan Houdek 3f214dd244 HarnessHelpers: Use FEXCore helper for file loading. 2023-05-17 21:05:55 -07:00
Ryan Houdek 61ca651fe1 FEXCore: Don't initialize ThunkHandler on Win32
Adds a couple pointer checks to ensure it won't crash.

Doesn't work and will cause assertions.
2023-05-17 21:05:55 -07:00
Ryan Houdek cd0a340d29 unittests/ASM: Ensure wine harness runner works
Needs to execute the correct runner through wine, and needs to reserve
the low DOS region so something doesn't get loaded there.
2023-05-17 21:05:55 -07:00
Mai 77e8be1215 Merge pull request #2671 from Sonicadvance1/wine_syscalls
FEXCore: Support Wine syscalls
2023-05-18 00:04:25 -04:00
Ryan Houdek 182010ca97 Merge pull request #2678 from lioncash/strings
OpcodeDispatcher: Handle PCMPESTRM/VPCMPESTRM
2023-05-16 21:52:53 -07:00
Lioncache f7c663240e OpcodeDispatcher: Handle PCMPESTRM/VPCMPESTRM
...and with that all of the SSE4.2 string instructions are implemented now
2023-05-17 00:21:55 -04:00
Ryan Houdek e9244680aa Merge pull request #2677 from lioncash/masked
OpcodeDispatcher: Handle PCMPISTRM/VPCMPISTRM
2023-05-16 20:25:42 -07:00
Lioncache 82b4aef30d OpcodeDispatcher: Handle PCMPISTRM/VPCMPISTRM 2023-05-16 22:59:54 -04:00
Lioncache 22919a5b65 OpcodeDispatcher: Add mask variant handling to PCMPXSTXOpImpl()
Will be used to handle PCMPESTRM/PCMPISTRM instruction variants.
2023-05-16 22:59:52 -04:00
Mai 00dc373bb9 Merge pull request #2676 from Sonicadvance1/fix_at_execfn
ELFCodeLoader: Fixes missing AT_EXECFN
2023-05-15 17:49:54 -04:00
Ryan Houdek e593237670 ELFCodeLoader: Fixes missing AT_EXECFN
New versions of CEF rely on this existing. It will get this value and
run strdup on it, even if it is nullptr.

Fixes a steamwebhelper process constantly crashing with the Steam Beta
client.

Only missing auxv values now
- AT_PAGESZ
- AT_EXECFD (for execveat?)
- AT_PHDR
- All the random cache information values.
2023-05-14 03:10:11 -07:00
Mai 0a4bf10ba5 Merge pull request #2674 from Sonicadvance1/fix_shm_leaks
unittests: Adds step to remove stale SHM regions.
2023-05-12 23:22:23 -04:00
Ryan Houdek f47caf48c6 Merge pull request #2669 from Sonicadvance1/aotir_mutex
AOTIR: Stop passing a mutex around. It's already guarded
2023-05-12 18:56:55 -07:00
Ryan Houdek 5674d3a871 Merge pull request #2667 from Sonicadvance1/fextl_file
FEXCore: Convert Core and Telemetry over to fextl::file::File
2023-05-12 18:56:45 -07:00
Ryan Houdek fde64aedf7 unittests: Adds step to remove stale SHM regions.
Some of the unit tests we run will leak shm regions. Presumably this is
because they never called `shmctl(IPC_RMID)` so the ID is laked forever.

This can be seen by querying `/proc/sysvipc/shm` to see a list of old
shm regions that eventually hit the maximum capacity of 4096 shm ids.

Once CI is done running, run a utility application that all it does is
check for SHM IDs that have zero attachments (thus unused), it was
created by the UID of the runner, and it is older than ten minutes. At
which point it will erase it.

This will fix spurious failures in our CI caused by running out of SHM
IDs, previously I had a cron job setup to restart the CI runners every
hour or so which caused its own spurious failure problems.

FINALLY this bug was triaged which has been annoying us for...years?
2023-05-12 18:54:01 -07:00
Mai e03b859c20 Merge pull request #2673 from Sonicadvance1/remove_warnings_13
OpcodeDispatcher: Removes a warning that cropped up.
2023-05-12 21:49:43 -04:00
Mai ce4e991a6e Merge pull request #2672 from Sonicadvance1/optimize_trampoline
Thunks: Optimize ARM64 trampoline
2023-05-12 20:50:22 -04:00
Ryan Houdek 7d822ba1c8 OpcodeDispatcher: Removes a warning that cropped up. 2023-05-12 17:34:20 -07:00
Ryan Houdek f90dcd2eb1 FEXCore: Convert Core and Telemetry over to FEXCore::File::File
This way telemetry and IR dumping can work under Wine.
2023-05-12 17:32:48 -07:00
Ryan Houdek adbdd33ece fextl/fmt: Adds write handler for FEXCore::File::File 2023-05-12 17:32:48 -07:00
Ryan Houdek 06250d806d FEXCore/Utils: Adds File type
OS agnostic file class since we can't use std::FILE
2023-05-12 17:32:48 -07:00
Ryan Houdek 613ed559e7 Thunks: Optimize ARM64 trampoline
No need to use adr for getting the PC relative literal, we can use LDR
(literal) to load the PC relative address directly.

Reduces trampline instructions from 3 to 2, also reduces trampoline size
from 24-bytes to 16-bytes.
2023-05-12 17:28:36 -07:00
Ryan Houdek 8ac3841946 FEXCore: Support Wine syscalls
Wine syscalls need to end the code block at the point of the syscall.
This is because syscalls may update RIP which means the JIT loop needs
to immediately restart.

Additionally since they can update CPU state, make wine syscalls not
return a result and instead refill the register state from the CPU
state. This will mean the syscall handler will need to update their
result register (RAX?) before returning.
2023-05-12 16:42:26 -07:00
Ryan Houdek 458259bf47 FEXCore: Move EnumOperators to FEXCore
fextl needs this and can't depend on FHU
2023-05-12 15:23:00 -07:00
Ryan Houdek dc65a5ef8c Merge pull request #2668 from Sonicadvance1/fix_sra_disabled
ARM64: Fixes SRA disabled codepath
2023-05-12 15:21:52 -07:00
Ryan Houdek 2fc529d5b7 AOTIR: Stop passing a mutex around. It's already guarded 2023-05-11 03:56:33 -07:00
Ryan Houdek ea489567da ARM64: Fixes SRA disabled codepath
Disabling SRA has been broken a quite a while. Disabling this was
instrumental in figuring out the VC redistributable crash.

Ensure it works by reintroducing non-SRA load/store register handlers,
and by supporting runtime selectable dispatch pointers for the JIT.

Side-bonus, moves the {LOAD,STORE}MEMTSO ops over to this dispatch as
well to make it consistent and probably slightly quicker.
2023-05-11 03:25:19 -07:00
Ryan Houdek ed69eb9f6f Merge pull request #2665 from Sonicadvance1/prctl_tso
FEXCore: Adds support for hardware x86-TSO prctl
2023-05-09 04:10:34 -07:00
Ryan Houdek 6eae064511 FEXCore: Adds support for hardware x86-TSO prctl
From https://github.com/AsahiLinux/linux/commits/bits/220-tso

This fails gracefully in the case the upstream kernel doesn't support
this feature, so can go in early.

This feature allows FEX to use hardware's TSO emulation capability to
reduce emulation overhead from our atomic/lrcpc implementation.
In the case that the TSO emulation feature is enabled in FEX, we will
check if the hardware supports this feature and then enable it.

If the hardware feature is supported it will then use regular memory
accesses with the expectation that these are x86-TSO in strength.

The only hardware that anyone cares about that supports this is Apple's
M class SoCs. Theoretically NVIDIA Denver/Carmel supports sequentially
consistent, which isn't quite the same thing. I haven't cared to check
if multithreaded SC has as strong of guarantees. But also since
Carmel/Denver hardware is fairly rare, it's hard to care about for our
use case.
2023-05-08 20:12:03 -07:00
Ryan Houdek f1eb98548a Docs: Update for release FEX-2305 2023-05-07 05:00:11 -07:00
Mai 86af1f6a68 Merge pull request #2654 from Sonicadvance1/move_unaligned_handler
FEXCore: Moves SIGBUS handler to FEXCore/Utils
2023-05-05 21:29:37 -04:00
Ryan Houdek 2d4bf97cac FEXCore: Moves SIGBUS handler to FEXCore/Utils
This can be done in an OS agnostic fashion. FEXCore knows the details of
its JIT and should be done in FEXCore itself.

The frontend is only necessary to inform FEXCore where the fault occured
and provide the array of GPRs for accessing and modifying the signal
state.

This is necessary for supporting both Linux and Wine signal contexts
with their unaligned access handlers.
2023-05-05 17:04:26 -07:00
Ryan Houdek 542aeed7b9 Merge pull request #2663 from Sonicadvance1/fix_xcb_thunk
Thunks: Fixes xcb helper thread creation
2023-05-05 16:00:02 -07:00
Mai f7d827a26a Merge pull request #2662 from Sonicadvance1/disable_rdtscp
CPUID: Disable RDTSCP under wine
2023-05-05 17:33:20 -04:00
Ryan Houdek 37b5bc49c6 Merge pull request #2656 from Sonicadvance1/fexcore_no_exceptions
FEXCore: Compile without exceptions
2023-05-05 14:32:37 -07:00
Ryan Houdek dcb3f182d6 CPUID: Disable RDTSCP under wine
We don't have a sane way to query cpu index under wine. We could
technically still use the syscall since we know that we are still
executing under Linux, but that seems a bit terrible.

Disable for now until something can be worked out. Not like it is used
heavily anyway.
2023-05-05 13:52:39 -07:00
Mai ba45bf4ae7 Merge pull request #2661 from Sonicadvance1/virtual_alloc_base
Allocator: Adds VirtualAlloc with memory Base hint function
2023-05-05 14:35:24 -04:00
Mai 121d9fda2d Merge pull request #2660 from Sonicadvance1/arm64_win32_ra
Arm64Emitter: Replace x18 usage with x30
2023-05-05 14:35:02 -04:00
Mai 73ede9d000 Merge pull request #2659 from Sonicadvance1/save_platform_register
ARM64Emitter: Ensure platform register is saved on win32
2023-05-05 14:34:25 -04:00
Mai 6dfea8a80f Merge pull request #2657 from Sonicadvance1/remove_unnecessary_guard
LookupCache: Removes unnecessary recursive lock_guard
2023-05-05 14:33:50 -04:00
Ryan Houdek f06972c9e2 Thunks: Fixes xcb helper thread creation
Forgot to initialize CBDone to false before the helper thread is
started.
This fixes an issue where an XCB context is created, then stopped, then
another is created but immediately exits because CBDone was still true
from the previous run.

Also adds the handler for `xcb_connect_to_fd` so we don't miss that
usage.
2023-05-05 00:39:07 -07:00
Ryan Houdek ef6c220a75 Allocator: Adds VirtualAlloc with memory Base hint function
This will be used with the TestHarnessRunner in the future to map
specific memory regions.

This is only used as a hint rather than exact placement with failure on
inability to map. This also hits the fun quirk of 64k allocation
granularity which developers need to be careful about.
2023-05-04 15:39:32 -07:00
Ryan Houdek 1e4a6d432c Merge pull request #2658 from Sonicadvance1/remove_unused_log
LogManager: Remove unused handler
2023-05-04 15:32:04 -07:00
Ryan Houdek 4ebd180147 Arm64Emitter: Replace x18 usage with x30
Related to #2659 but not necessary directly.

Currently x30(LR) is unused in our RA. In all locations that call out to
code, we are already preserving LR and bringing it back after the fact.
This was just a missed opportunity since we aren't doing any call-ret
stack manipulations that would facilitate LR needing to stick around.

Since x18 is a reserved platform register on win32, we can replace its
usage with r19, and then replace r19 usage with x30 and everything just
works happily. Now x18 is the unused register instead of x30 and we can
come back in the future to gain one more register for RA on Linux
platforms.
2023-05-04 15:25:47 -07:00
Ryan Houdek ac4ef63ae6 ARM64Emitter: Ensure platform register is saved on win32
Platform register stores the TEB region on win32 and needs to be
preserved if we're going to overwrite it.

Ensure we do so.
2023-05-04 15:12:52 -07:00
Mai 65fa495890 Merge pull request #2655 from Sonicadvance1/unique_no_array
fextl/memory: Don't allow arrays in fextl::make_unique
2023-05-04 17:53:17 -04:00
Ryan Houdek b2392ef1c6 LogManager: Remove unused handler
This non-fmt handler is now entirely unused and can be removed.
2023-05-04 14:52:45 -07:00
Ryan Houdek 8e4d52396b LookupCache: Removes unnecessary recursive lock_guard
All code paths to this are already guaranteed to own the lock.

The rest of the codepaths haven't been vetted to actually need
recursive_mutex yet, but seems likely that it will be able to get
converted to a regular mutex with some more work.
2023-05-04 14:45:19 -07:00
Ryan Houdek 6eeb45b2dc FEXCore: Compile without exceptions
This disables some unwinding overhead when FEXCore is already guaranteed
to not throw.
2023-05-04 14:42:02 -07:00
Ryan Houdek 22cf2696da fextl/memory: Don't allow arrays in fextl::make_unique
This ensures we don't hit a programming error since we don't support the
array version of this.
2023-05-04 14:38:12 -07:00
Ryan Houdek 468f7471e1 Merge pull request #2653 from julliard/win32-fixes
Win32 memory allocation fixes
2023-05-03 11:18:57 -07:00
Alexandre Julliard 8081ac61e5 AllocatorHooks: Fix parameter order for Win32 _aligned_malloc.
The prototype is the opposite of memalign().
2023-05-03 16:15:07 +02:00
Alexandre Julliard 435b4daae1 AllocatorHooks: Pass valid parameters to the Win32 VirtualAlloc. 2023-05-03 16:13:37 +02:00
Ryan Houdek 30974fb2c9 Merge pull request #2652 from lioncash/str
OpcodeDispatcher: Simplify PCMPXSTRIOpImpl
2023-05-02 18:09:44 -07:00
Lioncache 5ee913bc75 OpcodeDispatcher: Simplify PCMPXSTRIOpImpl
All variants of the PCMPXSTRX instructions will take their arguments in
the same manner, so we don't need to specify them for each handler.

We can also rename the function to PCMPXSTRXOpImpl, since this will
be extended to handle the masking variants of the string instructions.
2023-05-02 18:48:35 -04:00
Ryan Houdek db706bb28f Merge pull request #2651 from Sonicadvance1/default_drm_handler
Ioctl: Add default handler for drm
2023-05-02 15:40:16 -07:00
Ryan Houdek 0d1810c159 Ioctl: Add default handler for drm
In the case that an unknown drm device shows up, send it down the
default handler. This handler is just a passthrough and assumes that the
kernel doesn't have any compat handlers for that device.

This is nicer than crashing or returning EPERM, since then downstream
drm drivers like Xe, Asahi, and PVR and still try to run.

We of course still want to run their struct definitions through CI once
they go upstream.
2023-05-02 15:08:20 -07:00
Mai c1dab3e6cc Merge pull request #2650 from Sonicadvance1/ioctl_fix_null
Ioctl: Ensure DRM name check uses strncmp
2023-05-02 17:57:11 -04:00
Ryan Houdek e8e70c4faf Ioctl: Ensure DRM name check uses strncmp
These strings don't actually null-terminate and previous checks were
working just because it would usually be null terminated due to
initialization.

Since this isn't guaranteed, I noticed a failure to determine drm device
due to some trailing garbage in the string.
2023-05-02 14:36:32 -07:00
Ryan Houdek 88247141d7 Merge pull request #2649 from lioncash/istri
OpcodeDispatcher: Handle PCMPISTRI/VPCMPISTRI
2023-05-02 14:35:47 -07:00
Lioncache f502154f96 OpcodeDispatcher: Handle VPCMPISTRI 2023-05-02 14:00:05 -04:00
Lioncache 7a59fb3e25 IR: Add IR fallback for VPCMPISTRX
Will be the fallback that handles the implicit length string instruction emulation.
2023-05-02 13:52:30 -04:00
Ryan Houdek e71f3e898f Merge pull request #2648 from lioncash/store
unittests: Add missing VPMASKMOVQ store test
2023-05-02 10:49:17 -07:00
Lioncache 8369f9c25b unittests: Add missing VPMASKMOVQ store test
Realized I forgot to add this in the commit that added
VPMASKMOVD/VPMASKMOVQ support.
2023-05-02 11:13:44 -04:00
Mai 456e9dbdea Merge pull request #2647 from Sonicadvance1/fix_vc_redist
SignalDelegator: Make sure to save and restore `InSyscallInfo`
2023-05-01 20:04:03 -04:00
Ryan Houdek 41d00c8dc6 SignalDelegator: Make sure to save and restore InSyscallInfo
Fixes #2560

This was a forgotten member of the context that needs to be saved and
restored when jumping around the signal state.
When FEX was receiving signals back to back, there was a chance that the
signals would ride the edge of having set `InSyscallInfo` in the JIT,
which meant the SRA state would get saved once, then another signal
would occur with the previous SRA data, saving SRA again. Then when
unwinding the frames it would corrupt the SRA registers. This would
result in trying to load SRA state that was no longer valid, looking
like a crash in the JIT that was hard to see what happens.

Should also make Mono games a little less crash happy.

Cleans up CPUState memcpy as well, since it was a little weird looking.
2023-05-01 12:59:19 -07:00
Mai 886c562882 Merge pull request #2646 from Sonicadvance1/siginfo_layout
SignalDelegator: Calculate siginfo layout
2023-04-26 12:29:51 -04:00
Ryan Houdek 8be2ad8c69 SignalDelegator: Calculate siginfo layout
There are a handful of siginfo layout types. To determine which layout
to use we need to check a combination of si_code and signal number
because each one in isolation doesn't explain the layout type.

Once the layout is calculated then calculate the siginfo using that
layout type. Adds a couple of different layout types to our guest
siginfo_t to handle the previously missing types.
2023-04-26 09:04:56 -07:00
Mai b2a3c6a043 Merge pull request #2584 from Sonicadvance1/remove_double_syscall_overhead
FM: Removes double syscall issue with `GetEmulatedFDPath`
2023-04-26 06:54:31 -04:00
Ryan Houdek e2db607769 FM: Removes double syscall issue with GetEmulatedFDPath
In the case that the symlink following would immediately fail (Due to
the first query not existing or not being a symlink), then FEX would
immediately return the first query. Then the resulting syscall using
this result would try to run the syscall on that file, immediately
error, and try again on the file outside of the rootfs. This gives us
three syscalls for the price of one.

This adds an additional syscall cost to every syscall that that takes
the `FollowSymlink` code path, which is quite a few.

Instead now, check to see if the filepath exists at all, and if the
filepath doesn't even exist for the first fstatat, return a `NoEntry`
immediately. This will cause the resulting syscall waiting for the
result to skip the EmulatedFDPath check and run regular syscall.

If the filepath is symlink it will still loop to track through the
entire symlink tree to find the one that is in the rootfs.

If the filepath is just a file, then the `IsLink` step will fail,
returning the filepath to the file inside of the rootfs. (Getting this
step wrong will make wine/proton immediately break).

With this resolved, this converts a lot of three syscall operations down
to two.
2023-04-26 02:13:29 -07:00
Mai 590422b295 Merge pull request #2641 from Sonicadvance1/remove_unittest_gen
FEXCore: Stop exposing the x86 table data symbols
2023-04-26 05:09:23 -04:00
Mai 7552ad29fa Merge pull request #2640 from Sonicadvance1/mingw_more
Get mingw compiling libFEXCore
2023-04-26 05:08:20 -04:00
Ryan Houdek 9d268df91f Softfloat: Disable some duplicate BIGFLOAT handlers
Since mingw has its reduced precision has double, these handlers were
duplicated and causing compile failure.
2023-04-26 01:48:37 -07:00
Ryan Houdek 699541485d Arm64: Disable ProcessorID and Break on mingw
Currently unsupported on mingw
2023-04-26 01:48:37 -07:00
Ryan Houdek 46a63186a2 FEXCore: Name libFEXCore correctly and use sync library 2023-04-26 01:48:37 -07:00
Ryan Houdek 520441c262 FHU: Implement ScopedSignalMaskWithMutexBase without signal mask for mingw 2023-04-26 01:48:37 -07:00
Ryan Houdek 90f347839d InterruptableConditionVariable: Implement for mingw 2023-04-26 01:48:37 -07:00
Ryan Houdek c9e7d9f331 FEXCore: Disable IRDumper on mingw 2023-04-26 01:48:37 -07:00
Ryan Houdek 8c3a3bfb7c FEXCore: Resolve some header includes
Some aren't necessary anymore. Some need to not exist on mingw.
2023-04-26 01:48:37 -07:00
Ryan Houdek 9034946b43 Move UContext from FEXCore to frontend.
FEXCore no longer needs this since all the signal handling is done in
the frontend.
2023-04-26 01:48:37 -07:00
Mai 9432a84cb4 Merge pull request #2638 from Sonicadvance1/mingw_signals
SignalDelegator: Moves all signal handling to the frontend
2023-04-26 04:46:06 -04:00
Ryan Houdek 056f44be0b SignalDelegator: Moves all signal handling to the frontend
This is a very OS specific operation and it living in FEXCore doesn't
make much sense. This still requires some strong collaboration between
FEXCore and the frontend but it is now split between the locations.

There's still a bit more cleanup work that can be done after this is
merged, but we need to get this burning fire out of the way.

This is necessary for llvm-mingw, this requires all previous PRs to be
merged first.

After this is merged, most of the llvm-mingw work is complete, just some
minor cleanups.

To be merged first:
- #2602
- #2604
- #2605
- #2607
- #2610
- #2615
- #2619
- #2621
- #2622
- #2624
- #2625
- #2626
- #2627
- #2628
- #2629
2023-04-26 01:24:11 -07:00
Mai b5420f5db3 Merge pull request #2629 from Sonicadvance1/fexcore_cmake_mingw
FEXCore: Fixup cmake file for mingw
2023-04-25 10:12:17 -04:00
Mai c94268789b Merge pull request #2619 from Sonicadvance1/fileloading_mingw
FileLoading: Add WIN32 specific loading path
2023-04-25 10:11:14 -04:00
Mai af15277fc4 Merge pull request #2615 from Sonicadvance1/fhu_mingw
FHU/FS: Create WIN32 helpers for some functions.
2023-04-25 10:09:35 -04:00
Mai 86e09a00f0 Merge pull request #2610 from Sonicadvance1/mingw_virtual_alloc
AllocatorHooks: Adds some mingw allocator helpers
2023-04-25 10:08:44 -04:00
Ryan Houdek 238ffdf893 Merge pull request #2643 from lioncash/vp
OpcodeDispatcher: Handle VPMASKMOVD/VPMASKMOVQ
2023-04-24 10:10:24 -07:00
Lioncache c94721a04b OpcodeDispatcher: Handle VPMASKMOVD/VPMASKMOVQ
We can reuse the same helper we have for handling VMASKMOVPD and VMASKMOVPS,
though we need to move some handling around to account for the fact that
VPMASKMOVD and VPMASKMOVQ 'hijack' the REX.W bit to signify the element
size of the operation.
2023-04-24 10:50:11 -04:00
Ryan Houdek c87f361bb5 FEXCore: Stop exposing the x86 table data symbols
This was only used for the unit test fuzzing framework. Which has been
removed and unused for pretty much its entire lifespan.

These can now be internal only.
2023-04-23 09:38:03 -07:00
Ryan Houdek ea52ae3bbc UnitTestGenerator: Remove this unused test
This is supposed to be a fuzzing based test but has been unused and not
fully supported for a long time.

Just remove it since we won't be coming back to this.
2023-04-23 09:30:40 -07:00
Mai 0fa4390e47 Merge pull request #2622 from Sonicadvance1/dispatcher_signals
Dispatcher: Disable signal handling under mingw
2023-04-21 21:43:30 -04:00
Mai 7a774a8d80 Merge pull request #2624 from Sonicadvance1/fexcore_cpuid
FEXCore: Switch to xbyak for CPUID fetch helpers.
2023-04-21 21:42:54 -04:00
Mai 361e684c64 Merge pull request #2628 from Sonicadvance1/objectcache_mingw
Disable AOT and object cache under mingw
2023-04-21 21:42:25 -04:00
Mai 4c74913edf Merge pull request #2627 from Sonicadvance1/disable_break_mingw
Disable Break/INT operations on mingw
2023-04-21 21:42:08 -04:00
Mai 059472fcef Merge pull request #2621 from Sonicadvance1/object_cache_packed
ObjectCache: Ensure correctly packed config option
2023-04-21 21:41:34 -04:00
Mai 4a11111abd Merge pull request #2626 from Sonicadvance1/thunks_mingw
Thunks: Disable under mingw
2023-04-21 21:40:40 -04:00
Mai f673afc38f Merge pull request #2625 from Sonicadvance1/gdbserver_mingw
GdbServer: Disable under mingw
2023-04-21 21:40:24 -04:00
Mai 1fad26d72f Merge pull request #2613 from Sonicadvance1/cpuinfo_mingw
CPUInfo: Add mingw helper for CalculateNumberOfCPUs
2023-04-21 21:39:52 -04:00
Mai 2b5ddb6b93 Merge pull request #2607 from Sonicadvance1/mingw_softflow
llvm-mingw: Fix SoftFloat compiling
2023-04-21 21:38:27 -04:00
Mai c140dd7da8 Merge pull request #2605 from Sonicadvance1/aligned_alloc
Allocator: Ensure uses of aligned allocations use aligned_free
2023-04-21 21:38:08 -04:00
Mai da126141d3 Merge pull request #2604 from Sonicadvance1/move_config_paths
Config: Move path generation to the frontend
2023-04-21 21:37:23 -04:00
Mai 874ae5b0fc Merge pull request #2602 from Sonicadvance1/move_thread_creation
Threads: Moves pthread logic to FEXLoader
2023-04-21 21:36:28 -04:00
Ryan Houdek 35f192b6fd Merge pull request #2636 from lioncash/pd
OpcodeDispatcher: Handle VCVTPD2PS/VCVTPS2PD
2023-04-18 07:44:28 -07:00
Lioncache 651c6f8ddf OpcodeDispatcher: Handle VCVTPS2PD/VCVTPD2PS 2023-04-18 10:29:57 -04:00
Lioncache 73ca4e5687 OpcodeDispatcher: Move vector float conversion to helper
Will be used for implementing the equivalent AVX instructions.
2023-04-18 10:07:30 -04:00
Lioncache cb9cc74fcc OpcodeDispatcher: Handle AVX variants of float-to-float conversions
Adds in the handling of destination type size differences with AVX.

Also fixes cases where the SSE operations would load 128-bit vectors
from meory, rather than only loading 64-bit vectors with VCVTPS2PD.
2023-04-18 09:52:28 -04:00
Ryan Houdek 512d6d0069 Merge pull request #2635 from lioncash/sd
OpcodeDispatcher: Handle VCVTSD2SS/VCVTSS2SD
2023-04-18 05:26:59 -07:00
Lioncache d1116456fc OpcodeDispatcher: Handle VCVTSD2SS/VCVTSS2SD 2023-04-18 08:13:23 -04:00
Lioncache 84985952c9 OpcodeDispatcher: Factor out scalar floating-point conversion to helper
Will be used to implement the AVX variants of VCVTSD2SS and VCVTSS2SD
2023-04-18 07:16:37 -04:00
Mai a351620c60 Merge pull request #2634 from Sonicadvance1/missing_avx
VEXTables: Adds a missing class of AVX instructions
2023-04-18 06:55:35 -04:00
Ryan Houdek 9117f7e724 VEXTables: Adds a missing class of AVX instructions
These are all AVX1, not sure how I missed this.
Sorry @lioncash, four more instructions.
2023-04-17 20:39:59 -07:00
Ryan Houdek 6e52a16ef3 Merge pull request #2633 from lioncash/reorg
Interpreter: Separate fallback OpHandler from F80 fallbacks
2023-04-17 20:26:30 -07:00
Lioncache 8e391e7a61 Interpreter: Move PCMPESTRX fallback to VectorFallbacks
Now that OpHandlers isn't coupled to the F80 ops anymore, we can
move this over to its own file dedicated to vector fallbacks.
2023-04-17 22:57:09 -04:00
Lioncache 98fbc4a46d Interpreter: Move OpHandler struct into its own header
We can also provide a general rundown for hooking up interpreter fallbacks here for the uninitiated.
2023-04-17 22:55:02 -04:00
Lioncache 8481aeccb5 Interpreter: Move F80Ops.h into Fallback directory
We can also rename it to F80Fallbacks.h to make the file purpose a little more explicit.
2023-04-17 22:54:58 -04:00
Lioncache b1df63f425 Interpreter: Move fallbacks into new directory
Will be used to store fallbacks and separate the definition struct from the F80 fallbacks
2023-04-17 22:05:00 -04:00
Ryan Houdek a12802e74c Merge pull request #2632 from lioncash/string
OpcodeDispatcher: Handle PCMPESTRI/VPCMPESTRI
2023-04-17 18:55:24 -07:00
Lioncache 39c73d975b OpcodeDispatcher: Handle PCMPESTRI/VPCMPESTRI 2023-04-17 21:42:58 -04:00
Lioncache 30cb1aaaed IR: Add VPCMPESTRX fallback
In order to implement the SSE4.2 string instructions in a reasonable
manner, we can make use of a fallback implementation for the time
being.

This implementation just returns the intermediate result and leaves it
up to the function making use of it to derive the final result from said
intermediate result. This is fine, considering we have the immediate
control byte that tells us exactly what is desired as far as output
formats go.

Given that the result of this IR op will never take up more than
16-bits, we store the flags we need to set in the upper 16 bits of the
result to avoid needing to implement multiple return values in the JIT.

Also, since the IR op just returns the intermediate result, this can be
used to implement all of the explicit string instructions with a single IR op.

The implementation is pretty heavily documented to help make heads or
tails of these monster instructions.
2023-04-17 21:39:32 -04:00
Ryan Houdek cbf41448fc Thunks: Disable under mingw 2023-04-17 03:10:04 -07:00
Ryan Houdek 47bdc9af12 Config: Move realpath usage to FHU 2023-04-17 03:05:25 -07:00
Ryan Houdek 0fad5b88c1 FEXCore: Fixup cmake file for mingw
- 64-bit allocator doesn't work under mingw atm.
- Can't link against libdl
- Can't have a SONAME because it is a PE, not a shared library.
2023-04-17 02:57:27 -07:00
Ryan Houdek fb93fa573c FHU: Implements a couple of helpers for mingw
tgkill as a stub because that cna't be handled.
2023-04-17 02:56:07 -07:00
Ryan Houdek 8c9fe0dd31 AOTIR: Disable loading and saving on mingw 2023-04-17 02:55:15 -07:00
Ryan Houdek dda3afcfaf ObjectCache: Disable job handling on mingw
This isn't wired up anyway, but this needs to be disabled for now.
2023-04-17 02:55:11 -07:00
Ryan Houdek 34ceefb2c3 JIT64: Disable Break op on mingw
No way to handle this currently.
2023-04-17 02:54:30 -07:00
Ryan Houdek 25ef63a069 OpcodeDispatcher: Disable INT instruction entirely under mingw
Not yet able to handle this there.
2023-04-17 02:54:25 -07:00
Ryan Houdek 99a9c88f3f GdbServer: Disable under mingw
This needs to move to the frontend at some point.
2023-04-17 02:53:40 -07:00
Ryan Houdek 77f56199e8 Dispatcher: Disable signal handling under mingw
This needs some hefty reconstructing
2023-04-17 02:53:03 -07:00
Ryan Houdek 6c13b629af FEXCore: Switch to xbyak for CPUID fetch helpers.
This will use the correct `__cpuid` define, either in cpuid.h or
self-defined depending on environment.

Otherwise we would need to define our own cpuid helpers to match the
difference between mingw and linux.
2023-04-17 02:52:17 -07:00
Ryan Houdek 3ebe9f7b04 CPUInfo: Add mingw helper for CalculateNumberOfCPUs 2023-04-16 17:30:30 -07:00
Ryan Houdek 1979273ce5 FHU/FS: Create WIN32 helpers for some functions.
These use std::filesystem and should be moved over to WIN32 specific
code.

Added comments to explain that these are currently placeholders and
since we don't do mingw+glibc fault testing these won't get picked up
anyway.
2023-04-16 00:36:20 -07:00
Ryan Houdek 005389f8c1 llvm-mingw: Fix SoftFloat compiling 2023-04-16 00:30:28 -07:00
Mai a33443db62 Merge pull request #2611 from Sonicadvance1/arm64_mingw
ARM64Dispatcher: Fix compiling with mingw
2023-04-16 03:30:09 -04:00
Mai 68599bf124 Merge pull request #2618 from Sonicadvance1/corestate_mingw
CoreState: Fix SynchronousFaultData padding type
2023-04-16 03:26:32 -04:00
Mai d7f9c7ece2 Merge pull request #2614 from Sonicadvance1/disable_frontend_mingw
mingw: Disable compiling Common/Linux/Tools
2023-04-16 03:25:47 -04:00
Mai 4bffdc6345 Merge pull request #2612 from Sonicadvance1/frontend_mingw
Frontend: Remove errant header
2023-04-16 03:25:03 -04:00
Mai 892c07a5ed Merge pull request #2603 from Sonicadvance1/mingw_no_fexloader
FEXLoader: Don't build with mingw
2023-04-16 03:24:36 -04:00
Mai 780491d61b Merge pull request #2606 from Sonicadvance1/disable_timerfd_test
gvisor: Disable timerfd test
2023-04-16 03:23:49 -04:00
Mai cbe55b0765 Merge pull request #2616 from Sonicadvance1/telemetry_mingw
Telemetry: Disable on WIN32
2023-04-16 03:23:20 -04:00
Mai 2ff5096103 Merge pull request #2617 from Sonicadvance1/netstream_mingw
Netstream: Disable on WIN32
2023-04-16 03:23:06 -04:00
Mai 797737a84d Merge pull request #2609 from Sonicadvance1/mingw_threadname
Threads: Adds SetThreadName helper
2023-04-16 03:22:46 -04:00
Mai cfc1aa593b Merge pull request #2620 from Sonicadvance1/ra_helper
RA: Use FindFirstSetBit helper
2023-04-16 03:20:09 -04:00
Mai fbc5d583a2 Merge pull request #2608 from Sonicadvance1/cpuid_mingw
CPUID: Fix std::min type cast
2023-04-16 03:19:48 -04:00
Ryan Houdek 6b964f70e0 CPUID: Fix std::min type cast 2023-04-16 00:09:36 -07:00
Ryan Houdek 78844ee975 ObjectCache: Ensure correctly packed config option 2023-04-15 18:41:57 -07:00
Ryan Houdek 51afcb7143 RA: Use FindFirstSetBit helper 2023-04-15 18:41:35 -07:00
Ryan Houdek 5258b1972b FileLoading: Add WIN32 specific loading path 2023-04-15 18:41:11 -07:00
Ryan Houdek c6616d64d8 CoreState: Fix SynchronousFaultData padding type 2023-04-15 18:40:34 -07:00
Ryan Houdek d9b9ce804b Netstream: Disable on WIN32 2023-04-15 18:40:11 -07:00
Ryan Houdek 132aa7e4d3 Telemetry: Disable on WIN32 2023-04-15 18:39:44 -07:00
Ryan Houdek 0419d065b5 mingw: Disable compiling Common/Linux/Tools 2023-04-15 18:38:40 -07:00
Ryan Houdek fc00a31aee Frontend: Remove errant header 2023-04-15 18:37:53 -07:00
Ryan Houdek 1de84110e8 ARM64Dispatcher: Fix compiling with mingw 2023-04-15 18:37:31 -07:00
Ryan Houdek 1962f036e1 ObjectCacheService: Use ThreadName helper 2023-04-15 18:21:42 -07:00
Ryan Houdek 879a081556 Threads: Adds SetThreadName helper 2023-04-15 18:21:42 -07:00
Ryan Houdek 105060363f FEXCore: Move mmap allocators over to VirtualAlloc 2023-04-15 18:07:54 -07:00
Ryan Houdek bbb3a6439f AllocatorHooks: Adds some mingw allocator helpers 2023-04-15 18:07:49 -07:00
Ryan Houdek 40b67462b7 Allocator: Ensure uses of aligned allocations use aligned_free
This will be used by mingw.
2023-04-15 15:25:17 -07:00
Ryan Houdek d853de39ff Config: Move path generation to the frontend
This lets all the path generation for the config to be in the frontend.
This then informs FEXCore where things should live.

This is for llvm-mingw. While paths aren't quite generated correctly,
this gets the code closer to compiling.
2023-04-15 15:25:01 -07:00
Ryan Houdek 401d89ee40 FEXLoader: Don't build with mingw
FEXLoader won't ever work in this configuration. This was a mistake to
enable.
2023-04-15 15:24:46 -07:00
Ryan Houdek 75a62f856b Threads: Moves pthread logic to FEXLoader
This is not an attempt to clean up the various issues with the pthread
logic, instead just moving the pthread specific logic out of FEXCore in
to FEXLoader.

FEXCore needs to know how to create threads in an agnostic way. Which is
why we obfuscate the details with this inteface.

Initially this was implemented with the pthread handlers in FEXCore and
expected eventually for those to get moved to the frontend. This is the
time when it has been moved.

This is the first step towards compiling with llvm-mingw.

Still a long way to go.
2023-04-15 15:24:30 -07:00
Ryan Houdek e8aaadb2d0 gvisor: Disable timerfd test
Somewhere between kernel 5.15 and 6.0 this syscall's behaviour has
changed.
The unit test expects -ESPIPE result but new kernels will return 0
(noop_llseek in the source).

The new solidrun board is our first CI machine to be running kernel 6.1
so this needs to be merged before that CI machine can be enabled.
2023-04-15 15:23:58 -07:00
Ryan Houdek 27c03f98d8 Merge pull request #2580 from Sonicadvance1/glibc_fault_ci
CI: Adds glibc faulting testing
2023-04-15 15:07:09 -07:00
Ryan Houdek bfd606ec3d Merge pull request #2552 from Sonicadvance1/glibc_docs_and_ci
Docs: Adds programming concerns documentation
2023-04-15 03:36:21 -07:00
Ryan Houdek 9823a64164 CI: Adds glibc faulting testing
This is based on our regular CI runner file with a couple of things
stripped out.

- gvisor tests removed because they cause our CI machines pain with
  tmp/shm mounts
- Thunks disabled since glibc fault testing is incompatible with it
- ARMEmitter tests removed since we don't want to test vixl.

All glibc conversion PRs will need to be merged before this and then
this needs to be rebased.
2023-04-15 03:18:09 -07:00
Ryan Houdek 72483ea21d Docs: Adds programming concerns documentation
Explaining FEX's memory concerns.
2023-04-15 02:29:31 -07:00
Ryan Houdek 1a91d849f0 Merge pull request #2591 from Sonicadvance1/new_jemalloc
Add in jemalloc glibc hooking again
2023-04-14 13:32:11 -07:00
Ryan Houdek 6d4cef723a Merge pull request #2595 from Sonicadvance1/fix_thread_cancel_fault_test
Threads: Disable glibc allocator fault testing with `exit`
2023-04-14 13:24:04 -07:00
Ryan Houdek 1c2fd72c84 Merge pull request #2590 from Sonicadvance1/drm_v6.2
Update drm headers to v6.2
2023-04-14 13:20:07 -07:00
Ryan Houdek 77f2378080 docs: Adds jemalloc document 2023-04-14 13:16:22 -07:00
Ryan Houdek 1cc9f2107d Add in jemalloc glibc hooking again
We still need to hook glibc for thunks to work with
`IsHostHeapAllocation`.
So now we link in two jemalloc allocators in different namespaces.

As usual we have multiple heap allocators that we need to be careful about.

1. jemalloc with `je_` namespace.
  - This is FEX's regular heap allocator and what gets used for all the
    fextl objects.
  - This allocator is the one that the FEX mmap/munmap hooks hook in to
     - This mmap hooking gives this allocator the full 48-bit VA even in
       32-bit space.

2. jemalloc with `glibc_je_` namespace.
  - This is the allocator that overrides the glibc allocator routines
  - This is the allocator that thunks will use.
  - This is what `IsHostHeapAllocation` will check for.

3. Guest glibc allocator
  - We don't touch this one. But it is distinct from the host side
    allocators.
  - The guest side of thunks will use this heap allocator.

4. Host glibc allocator
  - #2 replaces this one unless explicitly disabled.
  - Always expected to override the allocator, so this configuration
    isn't expected.

Already tested this with Dota Underlords to ensure this works with
thunks.
2023-04-14 13:16:22 -07:00
Ryan Houdek b979b339fc Threads: Disable glibc allocator fault testing with exit
FEX might not ever make it to its cleanup routines when exit is used. So
make sure to disable the fault checking if exit is used.

Opened #2594 to track this investigation.
2023-04-14 13:10:14 -07:00
Ryan Houdek 46e5343a0e External/drm: Update to v6.2 2023-04-14 13:07:52 -07:00
Ryan Houdek d519883dbe StructPackVerifier: Pass on Aligned attr 2023-04-14 13:07:52 -07:00
Ryan Houdek beb2c36fc2 drm: Define __user define which drm headers use. 2023-04-14 13:07:52 -07:00
Ryan Houdek f45ea1e0f6 Merge pull request #2600 from Sonicadvance1/fix_stringconv_allocate
StringUtils: Stop allocating TrimTokens
2023-04-12 03:50:20 -07:00
Ryan Houdek 92593162b0 StringUtils: Stop allocating TrimTokens
Use a string_view instead of a fextl::string so this stops allocating
memory (stack in the cases I have seen).

Fixes #2562
2023-04-12 03:34:58 -07:00
Ryan Houdek 307158d425 Merge pull request #2601 from Sonicadvance1/move_fexloader
FEXLoader: Move to Tools folder
2023-04-12 03:33:02 -07:00
Ryan Houdek 5c62ea21f4 Merge pull request #2598 from Sonicadvance1/stop_leaking_avx
FEXCore: Stop leaking AVX configuration state
2023-04-11 19:57:58 -07:00
Ryan Houdek 73caa6725f Merge pull request #2597 from Sonicadvance1/stop_sending_uninitialized_data
FEXServerClient: Insert missing padding in message packet
2023-04-11 19:57:50 -07:00
Ryan Houdek b0b23abedd StructVerifier: Update include path 2023-04-11 16:56:53 -07:00
Mai f55a653e99 Merge pull request #2599 from Sonicadvance1/andn_usage
OpcodeDispatcher: Move usages of `And(Not(` to Andn
2023-04-11 19:53:44 -04:00
Ryan Houdek 4cb35d2f5e Fix header includes from the move 2023-04-11 16:44:06 -07:00
Ryan Houdek 7e4334c98a FEXLoader: Move to Tools folder
This has confused people for quite a long time thinking these are unit
tests.

First step, move it without any changes. Only Renamed.
2023-04-11 16:37:21 -07:00
Ryan Houdek 0d0b99f344 OpcodeDispatcher: Move usages of And(Not( to Andn
Fixes #2199

Very few uses actually, we were pretty good at this already.
2023-04-11 15:35:12 -07:00
Ryan Houdek 44e06185b7 FEXCore: Stop leaking AVX configuration state
The dispatcher was saving AVX state even though FEX doesn't support it
currently. This is due to it checking for the config option rather than
the HostFeatures option.

The `EnableAVX` config option is supposed to be used to inform FEXCore
if we want AVX disabled or not when the host supports the feature. In
this case it is universally enabled because we haven't encountered any
games that have issues with AVX state being saved with signals. (We know
they exist, we just don't have configurations for them).

The HostFeatures option `SupportsAVX` is the option that is supposed to
be getting used for determining if the runtime AVX feature is enabled.
This also had an issue though that this was **also** always enabled if
running on an x86 host with AVX, or an ARM host with SVE2-256bit.
It was then disabled if the config option was disabled; But, since
FEX-Emu doesn't support AVX fully yet, we need to ensure this isn't yet
enabled.

But this only solves half the problem. In order for our CI to test AVX
features before fully supporting AVX, it needs to be able to enable AVX
so that the CPU state is correctly saved.

So we need to change the default configuration option to be false, and
have CI enable it for the tests that matter before AVX is fully
implemented.
2023-04-11 15:21:32 -07:00
Ryan Houdek 3f89cf1512 FEXServerClient: Insert missing padding in message packet
This was sending uninitialized data across the wire.
2023-04-11 14:27:40 -07:00
Ryan Houdek 18154183ad Merge pull request #2585 from Sonicadvance1/jemalloc_ptr_cost
Allocator: Remove pointer indirection overhead
2023-04-11 10:47:24 -07:00
Ryan Houdek 49abe8afb5 Allocator: Remove pointer indirection overhead
Every time we are calling a function in `FEXCore::Allocator::` this is a
pointer indirection. Which means on x86 it is always a `call [rdi]` and
on AArch64 it is a `ldr x17, [x0]; blr x17;`.

Instead of doing this, use inline functions in the header that call the
correct allocation function directly. This function gets inlined and is
no longer an indirect call.

When compiling with jemalloc, we forward declare the jemalloc function
definitions so we don't have to pull in the entire jemalloc interface in
to the public header definitions.
2023-04-11 10:29:35 -07:00
Ryan Houdek 7916281ee7 Merge pull request #2593 from Sonicadvance1/cpp_optparse_update
cpp-optparse: Update to latest optparse
2023-04-11 09:25:28 -07:00
Ryan Houdek d7bc0370ee Merge pull request #2592 from Sonicadvance1/remove_fwrite
fextl::fmt: Remove fwrite usage
2023-04-11 03:32:06 -07:00
Ryan Houdek c5da0e7ac1 Merge pull request #2596 from Sonicadvance1/fix_stringconv_handlers
StringConv: Convert to conversion functions that don't use std::string
2023-04-11 03:31:58 -07:00
Ryan Houdek 0e007d2724 StringConv: Convert to conversion functions that don't use std::string
`std::stoul` and `std::stroull` take a std::string which was converting
the string_view to a std::string first. Causing glibc fault testing to
catch this since not much uses this.

These will be added to the documentation.
2023-04-10 18:56:05 -07:00
Ryan Houdek 63b31d54c4 cpp-optparse: Update to latest optparse
Changes std::set and std::map usage over to fextl.

I missed pushing this change before.
2023-04-10 18:12:18 -07:00
Ryan Houdek 466edf7744 fextl::fmt: Remove fwrite usage
fwrite allocates some backing memory for buffering outputs which FEX
can't track.

Switch to using `fileno` to get the fd from the FILE and write directly.
This will need to be changed for llvm-mingw support but that will come after
this.

This will be added to the documentation that we can't use fwrite.
2023-04-10 17:43:47 -07:00
Ryan Houdek 0bb59ade53 Updates jemalloc
Needed to change some symbol names due to proper jemalloc namespacing
now.
2023-04-10 16:21:33 -07:00
Ryan Houdek c03ed529e6 Merge pull request #2582 from Sonicadvance1/ccache_thunks
Thunks: Enable ccache if available
2023-04-10 15:57:23 -07:00
Ryan Houdek f806ca688c Merge pull request #2564 from Sonicadvance1/more_glibc_allocations
More glibc allocation removals.
2023-04-10 15:57:05 -07:00
Ryan Houdek 48531e1dd2 FEXLoader: Use the std::FILE print handlers again
Just to match prior behaviour.
2023-04-10 15:40:51 -07:00
Ryan Houdek 1864a1d3b5 fextl::fmt: Add print with std::FILE* handler
This is to match prior behaviour, but untested if fwrite itself
allocates any memory so far.
2023-04-10 15:38:55 -07:00
Ryan Houdek e98a46aa5f Review comments 2023-04-07 17:01:53 -07:00
Ryan Houdek 96e9b5cf51 FEXServer: Fixes some missing fextl usage
Not sure how this was working when the type declaration didn't match.
2023-04-07 17:01:53 -07:00
Ryan Houdek 265c918d90 Move fextl::String_from_path to the only usage in FEXConfig
Ensures that people won't be tempted to use this elsewhere.
2023-04-07 17:01:53 -07:00
Ryan Houdek 76a5e4e66e ELFContainer: Removes some c_str 2023-04-07 17:01:53 -07:00
Ryan Houdek dce6389d87 FHU: Convert relative/absolute check to string_view 2023-04-07 17:01:52 -07:00
Ryan Houdek 46b306e861 Config: Remove to_string usage 2023-04-07 17:01:52 -07:00
Ryan Houdek 32d7fae373 GdbServer: Convert to_string usage 2023-04-07 17:01:52 -07:00
Ryan Houdek 47e5096676 Syscalls: Remove usage of std::string 2023-04-07 17:01:52 -07:00
Ryan Houdek f25b0461d0 HarnessHelpers: Remove std::to_string usage 2023-04-07 17:01:52 -07:00
Ryan Houdek 11402b637a fextl: Remove now unused string_from_string 2023-04-07 17:01:52 -07:00
Ryan Houdek a9c27646a0 FEXRootFSFetcher: Remove usage of fextl::string_from_string 2023-04-07 17:01:52 -07:00
Ryan Houdek 278f411cfa FEXGetConfig: Convert to fextl 2023-04-07 17:01:52 -07:00
Ryan Houdek cafbcfec69 Config: Remove string_from_path 2023-04-07 17:01:52 -07:00
Ryan Houdek f5ed9c4ff3 CodeCache: Convert std::fs to FHU 2023-04-07 17:01:52 -07:00
Ryan Houdek e027521941 ELFContainer: Convert std::fs to FHU 2023-04-07 17:01:52 -07:00
Ryan Houdek cb6680ef57 ELFSymbolDB: Convert to FHU 2023-04-07 17:01:52 -07:00
Ryan Houdek 613a368f5f FEXLoader: Convert unique_ptr to fextl 2023-04-07 17:01:52 -07:00
Ryan Houdek 13500e59a3 TestHarnessRunner: Convert unique_ptr to fextl 2023-04-07 17:01:52 -07:00
Ryan Houdek 1306e597dd CodeSerialize: Convert unique_ptr to fextl 2023-04-07 17:01:52 -07:00
Ryan Houdek 87f8a6655a FEXConfig: Fixes crash that will occur once glibc hooking is removed.
This unique_ptr would end up getting allocated through jemalloc then
fread through glibc.

Convert over to fextl::unique_ptr to ensure it gets allocated and
deallocated through jemalloc correctly.
2023-04-07 17:01:52 -07:00
Ryan Houdek 4d70f4fc4e Remove some unused headers now. 2023-04-07 17:01:52 -07:00
Ryan Houdek 9dd715573c Syscalls: Convert Sourcecodemap fstram usage to raw files 2023-04-07 17:01:52 -07:00
Ryan Houdek 87772efb31 IRLoader: Convert to fextl 2023-04-07 17:01:52 -07:00
Ryan Houdek bb922f9a9e External: Update robin-map 2023-04-07 17:01:52 -07:00
Ryan Houdek 7fe14b0748 FileManagement: Remove extraneous string_from_string 2023-04-07 17:01:52 -07:00
Ryan Houdek 4ab822aebb IRParser: Convert to fextl 2023-04-07 17:01:52 -07:00
Ryan Houdek ecb7956d78 Syscalls: Convert std::filesystem to FHU 2023-04-07 17:01:52 -07:00
Ryan Houdek e232a10442 FileFormatCheck: Convert fstream to raw files. 2023-04-07 17:01:52 -07:00
Ryan Houdek efada6a0ea Config: Convert some std::filesystem to FHU 2023-04-07 17:01:52 -07:00
Ryan Houdek cfbd74e17b FEXLoader: Convert some std::filesystem over to FHU 2023-04-07 17:01:52 -07:00
Ryan Houdek 4eb91ef7f1 ELFContainer: Remove fstream usage 2023-04-07 17:01:52 -07:00
Ryan Houdek 2d18156e15 AOT: Convert fstream to fextl and raw files 2023-04-07 17:01:52 -07:00
Ryan Houdek dff0b45f29 TestHarnessRunner: Remove extraneous fextl::string creation 2023-04-07 17:01:51 -07:00
Ryan Houdek 8fb7e8d80e FHU: Add RenameFile helper 2023-04-07 17:01:51 -07:00
Ryan Houdek 60fe987e09 NetStream: Add operator new/delete because of raw pointer usage. 2023-04-07 17:01:51 -07:00
Ryan Houdek 7180bb1496 GdbServer: Convert fstream to fextl 2023-04-07 17:01:51 -07:00
Ryan Houdek 001a086d85 Convert remaining fmt::format to fextl 2023-04-07 17:01:51 -07:00
Ryan Houdek 8a711383bb fextl/fmt: Adds print helper that takes FD 2023-04-07 17:01:51 -07:00
Ryan Houdek 257a3a54dc Context: Convert over to a unique_ptr 2023-04-07 17:01:51 -07:00
Ryan Houdek e4fadd6992 Merge pull request #2587 from Sonicadvance1/disable_sbrk
Allocator: Disable glibc sbrk allocations
2023-04-07 17:01:19 -07:00
Ryan Houdek 9e5971b89c Allocator: Disable glibc sbrk allocations
This is done by consuming a single page at the end of the current sbrk
memory region. Then consuming any remaining bytes that could have
potentially ended up in it.

This ensures that glibc won't be able to return 64-bit pointers to
32-bit thunks once the remaining work is in place.
2023-04-06 12:27:44 -07:00
Ryan Houdek 28c168ea0c Merge pull request #2586 from AndreRH/main
Dispatcher: Fixes restoring of AVX state
2023-04-05 12:42:36 -07:00
André Zwing f944709139 Dispatcher: Fixes restoring of AVX state 2023-04-05 21:15:05 +02:00
Ryan Houdek e8bf7a1a46 Merge pull request #2583 from Sonicadvance1/xcb_thunk
Thunks: Make xcb's callback more robust.
2023-04-03 13:30:24 -07:00
Ryan Houdek f87b00a7e4 Thunks: Make xcb's callback more robust.
Fixes a crash in xcb thunks since more stuff on FEX side has moved over
to jemalloc.

Destructors don't actually get called when a shared library is removed.
It's some weirdo quirk that we can't work around. Instead refcount
Display connections being created and disconnected. Creating the thread
on the first display creation, and tearing down on final display
teardown.

Must be merged before #2564 otherwise that PR will break thunks.
2023-04-03 13:16:50 -07:00
Ryan Houdek a9f0cb15bf Thunks: Enable ccache if available
More vrooms.
2023-04-03 09:39:42 -07:00
Mai a78860c194 Merge pull request #2581 from Sonicadvance1/disable_mcount_pic
unittests/gcc: Disable mcount_pic test
2023-04-01 18:09:47 -04:00
Ryan Houdek f056cc790e unittests/gcc: Disable mcount_pic test
This has a race condition with SIGPROF which the Tegra Orin hits
consistently. Probably due to a faster timer or something?
2023-04-01 14:48:44 -07:00
Ryan Houdek aac4e25ca4 Merge pull request #2549 from Sonicadvance1/glibc_remaining_allocations
Move FEX away from the remaining glibc allocations that we can
2023-04-01 09:46:29 -07:00
Ryan Houdek 546a1edb55 CodeReview 2023-04-01 09:27:01 -07:00
Tony Wasserka 21838fe03f Merge pull request #2574 from neobrain/feature_thunk_wayland
Add support for thunking Wayland
2023-04-01 17:15:49 +02:00
Tony Wasserka 42e2a5421a Thunks: Add wayland-client 2023-04-01 16:58:30 +02:00
Tony Wasserka 557df4f0a7 Thunks/gen: Add support for thunking APIs with pointer-to-function-pointer arguments 2023-04-01 16:52:49 +02:00
Tony Wasserka 99ba648a71 Thunks: Fix thunking libraries with "-" in their name
The LOAD_LIB and EXPORTS macros behave slightly differently in this regard:
* Use LOAD_LIB(libwayland-client) in Guest.cpp (library name with dash)
* Use EXPORTS(libwayland_client) in Host.cpp (library name with underscore)
2023-04-01 16:52:49 +02:00
Ryan Houdek 97daec3dba Review comments 2023-03-31 06:03:06 -07:00
Mai 4f20dba505 Merge pull request #2566 from Sonicadvance1/fexconfig_fixes
FEXConfig: Fixes misalignment when advanced option is changed.
2023-03-31 03:22:18 -04:00
Ryan Houdek 2990a9d820 FaultingAllocator: Review comments 2023-03-30 16:28:34 -07:00
Ryan Houdek 53bbbd5a4f Review code 2023-03-30 16:28:34 -07:00
Ryan Houdek 047dddb023 Rebase patching 2023-03-30 16:28:34 -07:00
Ryan Houdek 4f66ff6ec4 Paths: Remove unique_ptr usage 2023-03-30 16:28:34 -07:00
Ryan Houdek e46dbf7b7e TestHarnessRunner: Convert to fextl 2023-03-30 16:28:34 -07:00
Ryan Houdek 067f807405 InternalThreadState: Convert tsl to fextl 2023-03-30 16:28:34 -07:00
Ryan Houdek 463b4b748c fextl: Add fextl::fmt::print 2023-03-30 16:28:34 -07:00
Ryan Houdek 3eae668cec X86Jit: Fix xbyak allocating through glibc 2023-03-30 16:28:34 -07:00
Ryan Houdek fc8bf9f0f6 Config: Remove static which is allocated 2023-03-30 16:28:34 -07:00
Ryan Houdek 3cfc1de410 Common: Convert cpp-optparse over to fextl and use. 2023-03-30 16:28:34 -07:00
Ryan Houdek a14353bc3c FEXLoader: Convert remaining usages away from glibc 2023-03-30 16:28:34 -07:00
Ryan Houdek 629e547e5f AllocatorOverride: Remove AFmt, it can try allocating memory and infinite loop. 2023-03-30 16:28:34 -07:00
Ryan Houdek 12b710ed16 APITests: Test FHU::LexicallyNormal 2023-03-30 16:28:34 -07:00
Ryan Houdek ea275bbdcc FDUtils: Remove std::fs get_fdpath, no longer used and avoids glibc 2023-03-30 16:28:34 -07:00
Ryan Houdek e6a48ad7f3 Syscalls: Convert to get_fdpath with temp to avoid glibc 2023-03-30 16:28:34 -07:00
Ryan Houdek 114b626cff FileManager: Convert to FHU to avoid glibc 2023-03-30 16:28:34 -07:00
Ryan Houdek 4d35d550c1 EmulatedFiles: Convert to FHU to avoid glibc 2023-03-30 16:28:34 -07:00
Ryan Houdek e8c1dfa03a FEXLoader: Convert to FHU to avoid glibc 2023-03-30 16:28:34 -07:00
Ryan Houdek 9d04c4daee ELFCodeLoader: Convert to get_fdpath with temp to avoid glibc 2023-03-30 16:28:34 -07:00
Ryan Houdek 1ec31c610c FEXServerClient: Convert to FHU to remove glibc 2023-03-30 16:28:34 -07:00
Ryan Houdek 86f8ebf0ee FEX/Config: Convert to FHU to remove glibc 2023-03-30 16:28:34 -07:00
Ryan Houdek bbd0d26c16 Telemetry: Convert to FHU to remove glibc 2023-03-30 16:28:34 -07:00
Ryan Houdek 1eac7e7105 AOTIR: Convert to FHU to remove glibc 2023-03-30 16:28:34 -07:00
Ryan Houdek 1f9458a3c3 Config: Convert to FHU to remove glibc 2023-03-30 16:28:34 -07:00
Ryan Houdek a3be4b77fa Paths: Convert to FHU to remove glibc 2023-03-30 16:28:33 -07:00
Ryan Houdek 4cc14cf0e9 FHU: Add more utilities 2023-03-30 16:28:33 -07:00
Ryan Houdek b2ec28503d LookupCache: Move over to fextl::pmr 2023-03-30 16:28:33 -07:00
Ryan Houdek 170c9ee9e4 LookupCache: Switch to fextl 2023-03-30 16:28:33 -07:00
Ryan Houdek c77c2faaea FEXLoader: Add glibc hook faulting setup 2023-03-30 16:28:33 -07:00
Ryan Houdek 1eb36b8b31 Convert a ton of things over to fextl 2023-03-30 16:28:33 -07:00
Ryan Houdek 465ecd9b19 Mark code regions that require glibc memory allocations.
This ensures that when we enable glibc fault testing these sections
won't break CI.
2023-03-30 16:28:33 -07:00
Mai df354e37dd Merge pull request #2578 from Sonicadvance1/support_salc
OpcodeDispatcher: Implement support for 32-bit SALC instruction
2023-03-30 18:12:15 -04:00
Ryan Houdek 43e6d398b6 X86Dispatcher: Move xbyak to custom types 2023-03-30 08:49:26 -07:00
Ryan Houdek 0d7c856775 Update xbyak 2023-03-30 08:49:26 -07:00
Ryan Houdek 141dddc83e CMake: Adds glibc allocator fault option
This will be used for CI to ensure FEX doesn't use the glibc allocator
2023-03-30 08:49:26 -07:00
Ryan Houdek 64aa3bfabe Switch FEX to fextl::fmt 2023-03-30 08:49:26 -07:00
Ryan Houdek f02a111d33 fextl: add memory for unique_ptr and make_unique 2023-03-30 08:49:26 -07:00
Ryan Houdek 79f7baffe3 fextl: Add pmr default resource 2023-03-30 08:49:26 -07:00
Ryan Houdek 7150c532f3 fextl: add robin_map 2023-03-30 08:49:26 -07:00
Ryan Houdek e48fb1850e fextl: add unordered_multimap 2023-03-30 08:49:26 -07:00
Mai 88dba60bee Merge pull request #2579 from Sonicadvance1/invalid_instruction_log
Core: Add a new log message for unsupported instruction
2023-03-29 22:52:08 -04:00
Ryan Houdek c9fb9c4cae Merge pull request #2577 from lioncash/wide
ARMEmitter: Handle SVE2 integer add/subtract wide category
2023-03-29 15:17:17 -07:00
Ryan Houdek d615ae9c6a Core: Add a new log message for unsupported instruction
The previous log in the frontend is super useful when an instruction
decoding wasn't supported.
Now that most of AVX is covered, a game will crash on SIGILL (and
usually catch it) and close without any indication.

Now if the instruction is decoded but it is invalid for the
configuration, still output a message as a good indicator that the game
is using instructions that the host doesn't support.

Will let us still pick up on games crashing due to lack of SVE very
easily.
2023-03-29 14:38:16 -07:00
Ryan Houdek 7629edcf61 OpcodeDispatcher: Implement support for 32-bit SALC instruction
This is an undocumented but supported instruction. It behaves just like
an `sbb al, al` but doesn't set flags and is one byte shorter.

The end result is that al is set to 0xFF or 0 depending on if CF is set
or not.
2023-03-29 14:34:51 -07:00
Lioncache 73d250c555 ARMEmitter: Handle SVE2 integer add/subtract wide category 2023-03-29 17:02:18 -04:00
Lioncache 337f8b06a3 ARMEmitter: Convert SVE2 integer multiply long to wide helper
Unifies the emitted ops under the same underlying emitter function.
2023-03-29 16:52:37 -04:00
Lioncache 87fa545bd0 ARMEmitter: Convert SVE2 integer add/subtract long to wide helper
The generic helper will be used to implement the remaining unimplemented
category from this group
2023-03-29 16:52:34 -04:00
Ryan Houdek 7747ac8de8 Merge pull request #2576 from lioncash/mul
ARMEmitter: Handle SVE Integer Multiply-Add - Unpredicated group
2023-03-29 13:14:29 -07:00
Lioncache deb1c9e933 ARMEmitter: Handle SVE mixed sign dot product category 2023-03-29 15:43:07 -04:00
Lioncache 672a88395d ARMEmitter: Handle SVE2 saturating multiply-add high category 2023-03-29 15:38:38 -04:00
Lioncache 88524ce718 ARMEmitter: Handle SVE2 saturating multiply-add long category 2023-03-29 15:33:49 -04:00
Lioncache bb153054f9 ARMEmitter: Handle SVE2 integer multiply-add long category 2023-03-29 15:28:32 -04:00
Lioncache 22a7a49042 ARMEmitter: Handle SVE2 complex integer multiply-add 2023-03-29 15:19:09 -04:00
Lioncache 9876f3eb5c ARMEmitter: Handle SVE2 saturating multiply-add interleaved long category 2023-03-29 15:08:30 -04:00
Lioncache b9e4ce4029 ARMEmitter: Handle SVE integer dot product (unpredicated) category 2023-03-29 14:55:06 -04:00
Lioncache 0c048772e0 ARMEmitter: Handle CDOT (vectors) 2023-03-29 14:54:35 -04:00
Ryan Houdek cf66643c60 Merge pull request #2575 from lioncash/store
OpcodeDispatcher: Handle store variants of VMASKMOVPD/VMASKMOVPS
2023-03-29 11:26:22 -07:00
Lioncache 830c1884d1 OpcodeDispatcher: Handle store variants of VMASKMOVPD/VMASKMOVPS
And with that, we support all of the AVX1-only instructions.

The remaining instructions for full AVX1 support is now just the SSE4.2
string instructions.
2023-03-29 14:03:23 -04:00
Lioncache 5abf9de8a5 IR: Add VStoreVectorMasked IR op
Will be used to implement the store variants of VPMASKMOV and
VMASKMOVP{D, S}
2023-03-29 14:03:20 -04:00
Tony Wasserka fa21944428 Thunks/gen: Print stacktrace on crash
This is useful information e.g. when map lookup from clang::Type pointers
fails during code generation.
2023-03-29 18:39:50 +02:00
Ryan Houdek 5cdde0bef1 Merge pull request #2572 from lioncash/mask
OpcodeDispatcher: Handle load variants of VMASKMOVPD/VMASKMOVPS
2023-03-28 08:01:44 -07:00
Lioncache 25960fe6b1 OpcodeDispatcher: Handle load variants of VMASKMOVP{D, S} 2023-03-28 10:35:23 -04:00
Lioncache eb8626c1f7 IR: Add VLoadVectorMasked IR op
Will be used to implement the load variants of VMASKMOVP{D, S} and
VPMASKMOV{D, Q}

Particularly useful, since with SVE this behavior can be collapsed into
two instructions (CMPGT followed by the relevant LD1 load instruction)
2023-03-28 01:57:25 -04:00
Ryan Houdek 100b4d4a5b Merge pull request #2571 from lioncash/asm
ARMEmitter: Fix treating 32-bit elements as 64-bit with ld1w
2023-03-27 22:11:13 -07:00
Lioncache ef7853ca4a ARMEmitter: Fix treating 32-bit elements as 64-bit with ld1w
These conditionals were accidentally inverted and were treating 32-bit
elements as 64-bit ones, when this is unintended.

Also add missing tests to ensure this doesn't slip through in the
future.
2023-03-28 00:55:04 -04:00
Ryan Houdek 55d65f3aea Merge pull request #2570 from lioncash/mov
Arm64/VectorOps: Remove a few unnecessary EORs from comparisons in SVE path
2023-03-27 13:35:48 -07:00
Ryan Houdek 2c2abc550b Merge pull request #2569 from lioncash/psadbw
OpcodeDispatcher: Handle VMPSADBW
2023-03-27 13:31:52 -07:00
Lioncache a1dc132f03 Arm64/VectorOps: Eliminate unnecessary EOR and MOV in FP compares
We can use the zeroing variant of MOVPRFX to perform the same behavior.
2023-03-27 16:16:28 -04:00
Ryan Houdek 719803bc5a Merge pull request #2568 from lioncash/int
ARMEmitter: Finish off SVE2 Integer - Predicated group
2023-03-27 13:01:07 -07:00
Lioncache ef31e0c7c7 OpcodeDispatcher: Handle VMPSADBW 2023-03-27 16:00:24 -04:00
Lioncache eecd016ba8 Arm64/VectorOps: Eliminate unnecessary EOR and MOV in VCMP{EQ,GT}
We can use the zeroing version of MOVPRFX to perform the same behavior.
2023-03-27 15:38:03 -04:00
Lioncache b2ec6d5208 OpcodeDispatcher: Move MPSADBW implementation into helper
This will be used for implementing the AVX variant of this instruction.
2023-03-27 12:40:28 -04:00
Lioncache 416a7b825d ARMEmitter: Move SVE2IntegerSaturatingAddSub over to using SVE2IntegerPredicated helper
Deduplicates some code.
2023-03-27 12:21:42 -04:00
Lioncache c3511ffa48 ARMEmitter: Move SVEIntegerPairwiseArithmetic over to using SVE2IntegerPredicated helper
Deduplicates some code.
2023-03-27 12:21:42 -04:00
Lioncache b9c3277b09 ARMEmitter: Move SVE2IntegerHalvingPredicated over to using SVE2IntegerPredicated helper
Deduplicates some code.
2023-03-27 12:21:38 -04:00
Lioncache 09f259c458 ARMEmitter: Handle SVE2 saturating/rounding bitwise shift left (predicated) category 2023-03-27 11:58:02 -04:00
Lioncache 7a2d23e189 ARMEmitter: Handle SVE2 integer unary operations (predicated) category 2023-03-27 11:36:06 -04:00
Lioncache 1e10561892 ARMEmitter: Handle SVE2 integer pairwise add accumulate long category 2023-03-27 11:22:39 -04:00
Ryan Houdek dcc8d2a1c7 FEXConfig: Fixes misalignment when advanced option is changed.
When the textual field is changed, that order changes and there would be
a partial resort where the labels would misalign compared to the text
entries.
Fixes that.
2023-03-26 03:37:40 -07:00
Ryan Houdek 9183cf144f Merge pull request #2547 from Sonicadvance1/standard_string_theory_versus_fex
Convert most std::string over to fextl
2023-03-24 04:37:40 -07:00
Ryan Houdek 043a03547f Resolve review comments 2023-03-24 04:09:03 -07:00
Ryan Houdek 7022b3b825 Review c_str() changes 2023-03-23 12:45:14 -07:00
Ryan Houdek 6016c48f98 SourcecodeResolver: Resolve review 2023-03-23 08:19:12 -07:00
Ryan Houdek 60408241c5 EmulatedFiles: Resolve review 2023-03-23 08:18:56 -07:00
Ryan Houdek bed10000e4 fextl: Add fmt 2023-03-23 08:15:47 -07:00
Ryan Houdek a899444ddd StrConv: Address review comments 2023-03-23 08:11:01 -07:00
Ryan Houdek 72df6a26d5 Merge pull request #2559 from lioncash/shift
Arm64/VectorOps: Handle 64-bit elements in VSShr
2023-03-21 19:40:24 -07:00
Lioncache dfeee6a1ed Arm64/VectorOps: Fix comment typo in VUShr
Meant to write USHL here, rather than SSHL
2023-03-21 22:25:34 -04:00
Lioncache 91d6dc0528 Arm64/VectorOps: Handle 64-bit elements in VSShr
Makes it consistent with all of the other variable shift IR ops.

Now this won't explode if we ever implement VPSRAVQ from AVX-512
2023-03-21 22:23:12 -04:00
Ryan Houdek 54e784742c Merge pull request #2558 from lioncash/shift
OpcodeDispatcher: Handle VPSLLVD/VPSLLVQ
2023-03-21 15:15:40 -07:00
Ryan Houdek 22936144ce Merge pull request #2557 from lioncash/pred
ARMEmitter: Handle SVE Inc/Dec by Predicate Count category
2023-03-21 15:12:39 -07:00
Lioncache fea0162096 OpcodeDispatcher: Handle VPSLLVD/VPSLLVQ 2023-03-21 17:08:03 -04:00
Lioncache 8287375117 x86_64/VectorOps: Implement VUShl IR op
This will be used to implement VPSLLVD/VPSLLVQ
2023-03-21 16:57:48 -04:00
Lioncache 0106f03b44 Arm64/VectorOps: Implement VUShl IR op
This will be used to implement VPSLLVD/VPSLLVQ
2023-03-21 16:55:33 -04:00
Lioncache 3e2d1cf8c0 OpcodeDispatcher: Factor out variable shift code into helper
We can slightly modify the table flags for VPSRAVD in order to
consolidate the variable shift code into one place.

Useful, since we still need to implement VPSLLVD and VPSLLVQ.

Also in the event we start supporting AVX-512, then the code can just be
extended from a single point.
2023-03-21 16:25:20 -04:00
Lioncache 392438d70a ARMEmitter: Handle SVE inc/dec register by predicate count group 2023-03-21 16:16:13 -04:00
Lioncache 0930a10026 ARMEmitter: Handle SVE inc/dec vector by predicate count group 2023-03-21 16:16:13 -04:00
Lioncache cb8ab3d328 ARMEmitter: Handle SVE saturating inc/dec register by predicate count group 2023-03-21 16:16:13 -04:00
Lioncache a670c1d762 ARMEmitter: Handle SVE saturating inc/dec vector by predicate count group 2023-03-21 16:16:11 -04:00
Ryan Houdek 91f056bd2d Merge pull request #2556 from lioncash/shift
OpcodeDispatcher: Handle VPSRLVD/VPSRLVQ
2023-03-21 12:29:19 -07:00
Lioncache efba2aed7d Arm64/VectorOps: Remove unnecessary mov in VSShr
We can just use the movprfx to move the data to shift into the
destination and then perform the shift directly with Dst as an operand.

This is safe, since no original vectors are used in the ASR operation,
only temporaries are used.

Lets us eliminate a temporary and keep things simpler.
2023-03-21 15:03:59 -04:00
Lioncache 233aef5289 OpcodeDispatcher: Handle VPSRLVD/VPSRLVQ 2023-03-21 15:03:59 -04:00
Lioncache 9d38bc0c81 x86_64/VectorOps: Implement VUShr IR op
Will be used for implementing VPSRLV AVX ops
2023-03-21 15:03:59 -04:00
Lioncache a76967650a Arm64/VectorOps: Implement VUShr IR op
Will be used for implementing VPSRLV AVX ops
2023-03-21 15:03:55 -04:00
Ryan Houdek 7f0cccfbe0 Merge pull request #2555 from lioncash/pred
ARMEmitter: Handle SVE Element Count category
2023-03-21 10:49:21 -07:00
Lioncache 437ea926b4 ARMEmitter: Handle SVE saturating inc/dec register by element count 2023-03-21 12:55:48 -04:00
Lioncache 4359f7439a ARMEmitter: Handle SVE inc/dec register by element count 2023-03-21 12:32:28 -04:00
Lioncache a1712ab455 ARMEmitter: Handle SVE inc/dec vector by element count group 2023-03-21 12:24:39 -04:00
Lioncache 0553f69eed ARMEmitter: Handle SVE element count category 2023-03-21 12:13:20 -04:00
Lioncache d03dfe2b33 ARMEmitter: Handle SVE saturating inc/dec vector by element count group 2023-03-21 12:06:50 -04:00
Ryan Houdek 5e105e86af Merge pull request #2554 from lioncash/warn
HostFeatures: Mark DCZID related utilities as [[maybe_unused]]
2023-03-21 08:41:56 -07:00
Lioncache 97bc3afb84 HostFeatures: Mark DCZID related utilities as [[maybe_unused]]
When building with the simulator enabled, these aren't used and cause
some unused funtion/variable warnings.

We can silence these with [[maybe_unused]], since these are used in
non-simulator builds.
2023-03-21 11:26:55 -04:00
Ryan Houdek bccc9f1656 HarnessHelpers: Convert string to fextl 2023-03-16 03:21:46 -07:00
Ryan Houdek 834e862dfb FEX: Convert string to fextl 2023-03-16 03:21:46 -07:00
Ryan Houdek b433d7b4cb ELFSymbolDB: Convert string to fextl 2023-03-16 03:21:46 -07:00
Ryan Houdek b6d36f123a Common: Convert string to fextl 2023-03-16 03:21:46 -07:00
Ryan Houdek 0ab2a550b1 FEXCore: Convert string to fextl 2023-03-16 03:21:46 -07:00
Ryan Houdek 1877457c4d FEX: Convert sstream to fextl 2023-03-16 03:21:46 -07:00
Ryan Houdek e5867b89ed FEXCore: Convert sstream to fextl 2023-03-16 03:21:46 -07:00
Ryan Houdek 487785de41 StringConv: Remove now redundant duplicated functions 2023-03-16 03:21:46 -07:00
Ryan Houdek e741ac37b1 StringUtils: Remove now redundant duplicated StringUtils functions 2023-03-16 03:21:46 -07:00
Ryan Houdek 5d928fbdee FileLoading: Remove now unnecessary duplicated functions 2023-03-16 03:21:46 -07:00
Ryan Houdek d1a14bf321 CPUID: Convert string to fextl 2023-03-16 03:21:46 -07:00
Ryan Houdek 3f98eff6e5 AOT: Convert string to fextl 2023-03-16 03:21:46 -07:00
Ryan Houdek 46d92b7d78 Dispatcher: Convert string to fextl 2023-03-16 03:21:46 -07:00
Ryan Houdek 4f6e12c460 GdbServer: Convert stringstream to fextl 2023-03-16 03:21:46 -07:00
Ryan Houdek 45493ad3ef fextl: Add sstream 2023-03-16 03:21:46 -07:00
Ryan Houdek fe886716a4 Config: Convert string to fextl
This sprawled out to lots of places as expected
2023-03-16 03:21:45 -07:00
Ryan Houdek e9a0d95c65 Telemetry: Convert string to fextl
Kind of sprawled a bit.
2023-03-16 03:21:45 -07:00
Ryan Houdek bb0757c2c9 Profiler: Convert string to fextl 2023-03-16 03:21:45 -07:00
Ryan Houdek 24101d99ac CPUBackend: Convert string to fextl 2023-03-16 03:21:45 -07:00
Ryan Houdek 7d226a6b18 SourcecodeResolver: Convert string to fextl 2023-03-16 03:21:45 -07:00
Ryan Houdek 5143579c20 SoftFloat: Convert string to fextl 2023-03-16 03:21:45 -07:00
Ryan Houdek a649e6aedf IRParser: Convert string to fextl 2023-03-16 03:21:45 -07:00
Ryan Houdek 3c021d64f8 ObjectCache: Convert string to fextl 2023-03-16 03:21:45 -07:00
Ryan Houdek 5ed0028d1a StringConv: Convert string to fextl 2023-03-16 03:21:45 -07:00
Ryan Houdek 61854c3655 StringUtils: Add fextl helpers 2023-03-16 03:21:45 -07:00
Ryan Houdek f77f243ae6 PassManager: Convert string to fextl 2023-03-16 03:21:45 -07:00
Ryan Houdek 5092675179 fextl/string: Add hash so it can exist in hashmaps 2023-03-16 03:21:45 -07:00
Ryan Houdek 9883f8fced unittests/Emitter: Convert string to fextl 2023-03-16 03:21:45 -07:00
Ryan Houdek e3f6ef6b18 Merge pull request #2545 from Sonicadvance1/more_fextl
Convert most things to fextl
2023-03-16 03:21:25 -07:00
Ryan Houdek 606242472a Convert the rest of map to fextl 2023-03-16 03:03:08 -07:00
Ryan Houdek 8941b8a312 Convert the rest of vector to fextl 2023-03-16 03:03:08 -07:00
Ryan Houdek 37832af818 Allocator: Convert vector to fextl 2023-03-16 03:03:08 -07:00
Ryan Houdek a7334608fa fextl: Move allocator hooks to its own header 2023-03-16 03:03:08 -07:00
Ryan Houdek bc42e849e9 TestHarnessRunner: Convert vector to fextl 2023-03-16 03:03:08 -07:00
Ryan Houdek 54c33b07a2 ELFCodeLoader: Convert to fextl
Kind of sprawls all over the place
2023-03-16 03:03:08 -07:00
Ryan Houdek a5b034129e AOTGenerator: Convert vector to fextl 2023-03-16 02:29:05 -07:00
Ryan Houdek 8d57446e88 FEXServerClient: Convert vector to fextl 2023-03-16 02:29:05 -07:00
Ryan Houdek bba8716c7e Merge pull request #2544 from Sonicadvance1/bulk_fextl_discount
fextl: Bulk merge
2023-03-16 02:24:34 -07:00
Mai 1771d086ed Merge pull request #2546 from Sonicadvance1/irdef_depends
FEXCore: IR_INC dependency on FEXCore_Base
2023-03-15 20:45:45 -04:00
Ryan Houdek d416331650 FEXCore: IR_INC dependency on FEXCore_Base
This dependency was on FEXCore only. Add it to the same dependency
definitions that CONFIG_INC is in so that FEXCore_Base picks it up.

Might solve a compile error from FEX Common missing the dependency being
generated.
2023-03-15 15:56:49 -07:00
Ryan Houdek aea90287eb Merge pull request #2520 from lioncash/cnt
ARMEmitter: Handle SVE predicate count group
2023-03-15 12:23:30 -07:00
Ryan Houdek 95a4994200 Cache: Convert over to fextl
Requires #2542 merged first. First commit is cherry-picked from that PR.
2023-03-15 12:22:04 -07:00
Ryan Houdek e30f3bd0c8 fextl: Implement queue 2023-03-15 12:22:04 -07:00
Ryan Houdek bd3facde63 fextl: Implement stack 2023-03-15 12:22:04 -07:00
Ryan Houdek 75e8b6cab8 ThreadHelper: Convert deque to fextl 2023-03-15 12:22:04 -07:00
Ryan Houdek 91330d6b86 IR: Convert deque to fextl 2023-03-15 12:22:04 -07:00
Ryan Houdek 92f2dc0737 ThreadPoolAllocator: Convert list to fextl 2023-03-15 12:22:04 -07:00
Ryan Houdek 5b97e6d354 Config: Convert list to fextl
Needs #2506 merged first. First commit is cherry-picked from it.
2023-03-15 12:22:04 -07:00
Ryan Houdek 728d19dbbe ELFCodeLoader2: Convert list to fextl 2023-03-15 12:22:04 -07:00
Ryan Houdek 4b9c7bd116 FileManager: Convert list to fextl 2023-03-15 12:22:04 -07:00
Ryan Houdek 017f101448 Thunks: Convert set to fextl 2023-03-15 12:22:04 -07:00
Ryan Houdek 4bde9f796a IR: Convert set to fextl 2023-03-15 12:22:03 -07:00
Ryan Houdek 7b790903c9 AOT: Convert set to fextl
This worms its way all the way to the frontend
2023-03-15 12:16:03 -07:00
Ryan Houdek 9c62442f43 Context: Convert unordered_map to fextl 2023-03-15 12:15:18 -07:00
Ryan Houdek a9a2c95202 Jit64: Convert unordered_map to fextl
Interpreter just remove the destination as a map. This isn't useful
anymore.
2023-03-15 12:14:59 -07:00
Ryan Houdek 194f97aefa Thunks: Convert unordered_map to fextl 2023-03-15 12:14:19 -07:00
Ryan Houdek 1c0a0c1fe5 IR: Convert unordered_map to fextl 2023-03-15 12:14:10 -07:00
Ryan Houdek 149fd2a2b6 ELFContainer: Convert unordered_map to fextl 2023-03-15 12:13:06 -07:00
Ryan Houdek 06211d330d LinuxSyscalls: Convert unordered_map to fextl 2023-03-15 12:12:16 -07:00
Ryan Houdek 18c9e84543 Config: Convert unordered_map to fextl
Needs #2506 merged first. First commit is cherry-picked from it.
2023-03-15 12:11:41 -07:00
Ryan Houdek a4831519cf FileManager: Convert unordered_set to fextl 2023-03-15 12:11:16 -07:00
Ryan Houdek bac5eff296 RAPass: Convert unordered_set to fextl 2023-03-15 12:10:32 -07:00
Ryan Houdek 97565b68db LinuxSyscalls: Convert vector to fextl 2023-03-15 12:10:21 -07:00
Ryan Houdek 6b053d298e VDSOEmulation: Convert vector to fextl
Pipes back to thunk definitions in FEXCore as well.
2023-03-15 12:08:58 -07:00
Ryan Houdek 823222702a ELF parsing: Convert vector to fextl 2023-03-15 12:08:14 -07:00
Ryan Houdek e8132b82d0 LookupCache: Convert vector to fextl 2023-03-15 12:05:42 -07:00
Ryan Houdek cdac296ce3 GDBServer: Convert vector to fextl 2023-03-15 12:05:37 -07:00
Ryan Houdek ffbcc52af6 Frontend/OpDispatcher: Convert vector to fextl
These two are intrinsicly linked. Need to cover both at the same time.
2023-03-15 12:05:30 -07:00
Ryan Houdek ddafc172fc CPUID: Convert vector to fextl 2023-03-15 12:05:26 -07:00
Ryan Houdek 670a029228 Logmanager: Convert vector to fextl 2023-03-15 12:05:21 -07:00
Ryan Houdek 30dcf2cf39 Context: Convert Threads vector to fextl
Avoiding `AppendThunkDefinitions` as it interacts with the frontend
right now.
2023-03-15 12:05:17 -07:00
Ryan Houdek df17f7adbb JITs: Convert relocations vector to fextl
Kind of goes places.
2023-03-15 12:05:12 -07:00
Ryan Houdek 9ffbffe463 Interpreter: Convert vector to fextl 2023-03-15 12:05:07 -07:00
Ryan Houdek 1aaa03a10d Dispatcher: Removes signal frame stack
This wasn't used anymore. Removed.
2023-03-15 12:05:02 -07:00
Ryan Houdek e4b2ed72a8 CodeEmitter: Convert vector to fextl
Needs #2506 merged first. First commit is cherry-picked from it.
2023-03-15 12:04:57 -07:00
Ryan Houdek c6d1b645bd Config: Convert vector to fextl 2023-03-15 12:04:50 -07:00
Ryan Houdek 35d126e12d Utils: Temporarily duplicate LoadFile
The std::vector implementation is going to be short-lived while code
churn happens.

Just duplicate it while we smash through the world.
2023-03-15 12:04:46 -07:00
Ryan Houdek d7a43c5553 FEXCore: Move Allocator.cpp to FEXCore_Base
We are going to use this from FileLoading.cpp which is in the base.
2023-03-15 12:04:41 -07:00
Ryan Houdek 8ccd95ff02 GDBJIT: Changes vector to fextl 2023-03-15 12:04:31 -07:00
Ryan Houdek af91430007 IR: Changes vector to fextl 2023-03-15 12:04:26 -07:00
Mai eadce2854f Merge pull request #2522 from Sonicadvance1/remove_originalelfcodeloader
FEXLoader: Move ELFCodeLoader2 to remove 2
2023-03-15 09:47:49 -04:00
Ryan Houdek 29963ad5e2 FEXLoader: Move ELFCodeLoader2 to remove 2
We aren't going back to the old code at this point.
2023-03-14 13:35:28 -07:00
Ryan Houdek 335aedc781 Tests: Remove old ELFCodeLoader that isn't used. 2023-03-14 13:32:52 -07:00
Ryan Houdek 41477db7aa Merge pull request #2511 from lioncash/scalars
OpcodeDispatcher: Handle VCVTSI2SD/VCVTSI2SS
2023-03-14 13:20:16 -07:00
Lioncache 02e245d61f OpcodeDispatcher: Handle VCVTSI2SD 2023-03-14 15:36:54 -04:00
Lioncache 03724c8486 OpcodeDispatcher: Handle VCVTSI2SS 2023-03-14 15:28:55 -04:00
Lioncache f8a575a982 OpcodeDispatcher: Factor out CVTGPR_To_FPR impl to helper
This can be used for the AVX implementations as well.
2023-03-14 15:12:21 -04:00
Ryan Houdek 2d05ffe3aa Merge pull request #2503 from Sonicadvance1/passes_fextl_vector
IR/Passes: Changes vector to fextl
2023-03-14 12:09:02 -07:00
Ryan Houdek 042b511126 IR/Passes: Changes vector to fextl 2023-03-14 11:53:35 -07:00
Ryan Houdek 8b0f66d599 FEXCore/Allocator: Adds missing allocation hook pointers
Just ones that I missed that are pervasive. At the very least with need
`aligned_alloc` for the fextl allocator.
2023-03-14 11:51:47 -07:00
Ryan Houdek 62b5441d56 fextl: Fix incorrect include path for allocator.h 2023-03-14 11:51:24 -07:00
Lioncache 8b20921341 ARMEmitter: Handle SVE predicate count group 2023-03-14 14:49:38 -04:00
Ryan Houdek 8633528cee Merge pull request #2502 from lioncash/flag
OpcodeDispatcher: Handle VPINSRB/VPINSRD/VPINSRQ/VPINSRW
2023-03-14 11:44:06 -07:00
Lioncache ce961a3ef4 OpcodeDispatcher: Handle VPINSRD/VPINSRQ
These are handled from the same encoding via the VEX.W bit.
2023-03-14 14:17:59 -04:00
Lioncache 8ce87d3b27 OpcodeDispatcher: Handle VPINSRW 2023-03-14 14:17:55 -04:00
Lioncache 24f03dd740 OpcodeDispatcher: Handle VPINSRB 2023-03-14 14:16:44 -04:00
Lioncache a1b0d1853b OpcodeDispatcher: Move PINSROp implementation to helper 2023-03-14 14:16:44 -04:00
Lioncache 6531d369cd Frontend: Handle 'op DST_XMM, SRC_XMM, GPR' AVX patterns
Things like VPINSRB and co. use this pattern, as well as VCVTSI2S{D,S}.

No other VEX-encoded instructions expect the first source to be a GPR
aside from some of the BMI1/BMI2 instructions, so we'll need to check
for this case.
2023-03-14 14:16:41 -04:00
Ryan Houdek 7344680672 Merge pull request #2501 from lioncash/ind
ARMEmitter: Handle floating-point multiply add (indexed) groups
2023-03-14 08:58:35 -07:00
Lioncache 16ae9ad1c7 ARMEmitter: Handle SVE floating point matrix multiply accumulate category 2023-03-14 11:37:49 -04:00
Lioncache 737400fd84 ARMEmitter: Handle SVE floating-point multiply (indexed) category 2023-03-14 11:22:51 -04:00
Lioncache 331cf66fda ARMEmitter: Make FCMLA (indexed) use the SVEFPMultiplyAddIndexed helper
Same behavior, but without all of the code duplication.
2023-03-14 11:19:19 -04:00
Lioncache 96c35a7bc7 ARMEmitter: Move fcmla (indexed) down to it's label
Just code movement to get rid of a TODO.
2023-03-14 11:05:14 -04:00
Lioncache 15903f5300 ARMEmitter: Handle SVE floating-point multiply-add (indexed) group 2023-03-14 11:01:49 -04:00
Ryan Houdek a8ed2af658 Merge pull request #2498 from Sonicadvance1/begin_the_stl_torment
FEXCore: Adds fexctl container alias objects
2023-03-14 07:30:12 -07:00
Ryan Houdek 3a81efdb28 Merge pull request #2494 from Sonicadvance1/sra_on_32bit
Arm64: Reclaim SRA registers on 32-bit
2023-03-14 07:30:00 -07:00
Ryan Houdek 3275dabd85 Merge pull request #2499 from lioncash/mla
ARMEmitter: Handle SVE Floating Point Multiply-Add group
2023-03-14 07:16:36 -07:00
Ryan Houdek 64c45ed70a Arm64: Reclaim SRA registers on 32-bit
Causes Portal to go from 85FPS to 120FPS on Lenovo X13s.

When running in 32-bit mode we were wasting 8 GPRs and 8 FPRs by still
allocating the top 8 registers of each even though 32-bit can't use
them.

Reallocate them to be register allocated registers when running 32-bit
applications which reduces spills and lowers the cost of
spilling/filling SRA registers.

Quite a significant speed boost for a little bit of work.
2023-03-14 07:15:05 -07:00
Lioncache 02d7f68094 ARMEmitter: Handle SVE Floating Point Multiply-Add group 2023-03-14 09:57:23 -04:00
Ryan Houdek afdc037110 FEXCore: Adds fexctl container alias objects
Completely unused right now, this will very quickly spread throughout
the entire codebase.

Currently only containers with FEX's global allocator usage.
Any STL class that is allocation free and has been vetted to fit within
the guidelines of not allocating memory will get an alias inside of
`fextl`.

This means that any std:: namespace usage going forward will need to be
scrutinized heavily. As aliasing it will mean it is vetted, otherwise it
needs to be checked with FEX CI (once in place) to ensure it doesn't
allocate memory behind our back.

Aliases everything we allow from std:: reduces mental burden from
needing to check if it is okay to use or not. Additionally it will allow
us to quickly pivot to the EASTL if we find as some point during this
transition that this will in-fact not work.

Once this is committed, we are going to be moving very quickly to
piecemeal convert /everything/ in the codebase over to fextl:: namespace
from std::.

I am sorry for the hellscape that will follow.
2023-03-14 06:56:44 -07:00
Mai 6e0a1a5f14 Merge pull request #2490 from Sonicadvance1/enable_fast_movs_cpuid
CPUID: Enable FAST REP MOVS
2023-03-13 16:20:17 -04:00
Mai ad332e3c30 Merge pull request #2492 from Sonicadvance1/signaldelegator_remove_unused
SignalDelegator: Cleanup unused functions
2023-03-13 16:19:48 -04:00
Mai 1aeab04c9f Merge pull request #2491 from Sonicadvance1/remove_unused_fd_mapping
FM: Remove unused FD to Name mapping
2023-03-13 16:17:45 -04:00
Mai aba42570a1 Merge pull request #2495 from Sonicadvance1/mingw_cmake
CMake: Get past configuration when mingw is used
2023-03-13 16:17:12 -04:00
Mai 7f243ee08a Merge pull request #2496 from Sonicadvance1/remove_get_nprocs
CPUInfo: Switch away from using get_nprocs_conf
2023-03-13 16:16:31 -04:00
Mai 82f7bb282f Merge pull request #2497 from Sonicadvance1/intrusive_ir_alloc
IntrusiveIRList: Ensure this is using the FEX allocator
2023-03-13 16:16:08 -04:00
Ryan Houdek 1248f3573c CPUInfo: Switch away from using get_nprocs_conf
Allocates memory behind our backs. Needed to make CI glibc allocator
clean.
2023-03-13 10:50:25 -07:00
Ryan Houdek d89f53cfd3 IntrusiveIRList: Ensure this is using the FEX allocator
Necessary to be glibc allocator clean.
2023-03-13 10:48:23 -07:00
Ryan Houdek 89f1e61779 CMake: Get past configuration when mingw is used
Right now cmake doesn't even get past the configuration stage when mingw
is used.
If mingw is detected, setup cmake so it can at least configure itself.
Will be necessary for cleaning up the rest of the codebase.
2023-03-12 16:37:35 -07:00
Ryan Houdek 11d396ca06 Merge pull request #2486 from lioncash/zero
ARMEmitter: Handle SVE floating-point compare with zero group
2023-03-11 15:35:06 -08:00
Ryan Houdek 3ff7cf19c0 FM: Remove unused FD to Name mapping
This was completely unused so this just reduces some syscall signal
mutex locking around FD operations.
2023-03-10 16:30:49 -08:00
Ryan Houdek 11bbfe5bed SignalDelegator: Cleanup unused functions
Also changes the error log to an assert log to ensure the crashing
process doesn't spin forever.
2023-03-10 16:18:14 -08:00
Ryan Houdek 07aa5327f4 CPUID: Enable FAST REP MOVS
Even without this being an optimal implementation, enabling this option
reduces the amount of time spent in memmove in Hollow Knight.

Before:
```
10.07%  [JIT] tid 284422  [.] /usr/lib/x86_64-linux-gnu/libc.so.6+0x17d780 (0x7fbcef8d6dfd)
```

After:
```
0.53%  [JIT] tid 288916  [.] /usr/lib/x86_64-linux-gnu/libc.so.6+0x17d780 (0x7f0cbe0dc608)
```
2023-03-10 15:44:51 -08:00
Lioncache c6ba51ee35 ARMEmitter: Handle SVE floating-point compare with zero group 2023-03-09 01:13:33 -05:00
Ryan Houdek cd40a85567 Merge pull request #2485 from lioncash/serial
ARMEmitter: Handle SVE Floating-point Serial Reduction (Predicated) group
2023-03-08 21:52:31 -08:00
Lioncache 19125720c2 ARMEmitter: Handle SVE Floating-point Serial Reduction (Predicated) group 2023-03-09 00:16:38 -05:00
Ryan Houdek 9b70c1d4ac Merge pull request #2484 from lioncash/unary
ARMEmitter: Handle SVE Floating Point Unary Operations - Unpredicated group
2023-03-08 20:58:32 -08:00
Lioncache 02d06ee833 ARMEmitter: Handle SVE Floating Point Unary Operations - Unpredicated group 2023-03-08 23:11:50 -05:00
Ryan Houdek 5e8d9601ed Merge pull request #2483 from lioncash/ldrr
ARMEmitter: Handle LDR (vector)
2023-03-08 19:59:34 -08:00
Lioncache 640cfcd0bd ARMEmitter: Make LDR (predicate) take an XRegister
Matches assembly use more closely.
2023-03-08 22:45:15 -05:00
Lioncache 5d71dfed86 ARMEmitter: Handle LDR (vector)
Now we have the vector variant as well as the predicate variant.
2023-03-08 22:43:13 -05:00
Ryan Houdek 1c3af30e96 Merge pull request #2482 from lioncash/pred
ARMEmitter: Move op and assertions into SVEFloatArithmeticPredicated helper
2023-03-08 19:32:09 -08:00
Lioncache a06fc3641f ARMEmitter: Handle FTMAD
Finishes off the only top-level instruction in the SVE Floating-Point
Arithemtic - Predicated group (also FTMAD is unpredicated, so cool).
2023-03-08 22:13:49 -05:00
Lioncache 5cbf984743 ARMEmitter: Move op and assertions into SVEFloatArithmeticPredicated helper
Deduplicates some code.
2023-03-08 22:13:44 -05:00
Ryan Houdek d1ece88bed Merge pull request #2481 from lioncash/align
unittests: Change alignment directive in 256-bit VPSADBW test to 32
2023-03-08 19:01:38 -08:00
Lioncache b1a00b05c4 unittests: Change alignment directive in 256-bit VPSADBW test to 32
Meant to change this over when writing the test, but forgot.
2023-03-08 21:26:20 -05:00
Ryan Houdek 1087c45b6c Merge pull request #2480 from lioncash/psadbw
OpcodeDispatcher: Handle VPSADBW
2023-03-08 18:25:23 -08:00
Lioncache d9da63d492 OpcodeDispatcher: Handle VPSADBW 2023-03-08 21:02:44 -05:00
Lioncache 597ffe5ba6 OpcodeDispatcher: Factor out PSADBW implementation to helper
This can be reused when implementing the AVX equivalent of the
operation.
2023-03-08 18:40:59 -05:00
Ryan Houdek 6517f7eb30 Merge pull request #2479 from lioncash/test
OpcodeDispatcher: Handle VTESTPD/VTESTPS
2023-03-08 15:25:47 -08:00
Lioncache 73ff932bb6 OpcodeDispatcher: Handle VTESTPD 2023-03-08 17:27:41 -05:00
Lioncache 9b742b1fac OpcodeDispatcher: Handle VTESTPS 2023-03-08 17:27:41 -05:00
Lioncache 8c63afb345 OpcodeDispatcher: Add helper for VTESTP{D,S} 2023-03-08 17:27:38 -05:00
Ryan Houdek c14f4355b7 Merge pull request #2478 from lioncash/maddubsw
OpcodeDispatcher: Handle VPMADDUBSW
2023-03-08 13:27:43 -08:00
Ryan Houdek 6645e68c95 Merge pull request #2477 from lioncash/mskb
OpcodeDispatcher: Handle VPMOVMSKB
2023-03-08 13:19:22 -08:00
Lioncache ff9de85851 OpcodeDispatcher: Handle VPMADDUBSW 2023-03-08 16:01:34 -05:00
Lioncache 4dca609614 OpcodeDispatcher: Factor out PMADDUBSW implementation into helper
This will be used to centralize the SSE and AVX implementations.
2023-03-08 15:40:17 -05:00
Lioncache 4fc365c149 OpcodeDispatcher: Handle VPMOVMSKB 2023-03-08 15:29:58 -05:00
Lioncache 8804e0a62a OpcodeDispatcher: Make MOVMSKOpOne suitable for AVX
Instead of hardcoding most sizes, derive them off the source register
size.

This will allow us to share the implementation between the SSE and AVX.
2023-03-08 15:29:55 -05:00
Ryan Houdek 1f512b098e Merge pull request #2476 from lioncash/redundant
Arm64/VectorOps: Remove unnecessary EOR in VAddP/VFAddP SVE path
2023-03-08 11:21:49 -08:00
Lioncache a31c3ad87a Arm64/VectorOps: Remove unnecessary EOR in VAddP/VFAddP SVE path
These aren't necessary for the operations anymore.
2023-03-08 14:05:52 -05:00
Ryan Houdek e92cfc4d7a Merge pull request #2475 from lioncash/extract
OpcodeDispatcher: Pass full register size through VExtractToGPR
2023-03-08 10:56:10 -08:00
Lioncache 820d743321 OpcodeDispatcher: Pass full register size through VExtractToGPR
While the register size field is currently unused, it's correct to
differentiate between 128-bit and 256-bit registers here.
2023-03-08 13:29:51 -05:00
Ryan Houdek 3b74e34b47 Merge pull request #2474 from lioncash/ffr
ARMEmitter: Handle SVE Write FFR group
2023-03-08 10:04:55 -08:00
Ryan Houdek 054430f004 Merge pull request #2473 from lioncash/vec
ARMEmitter: Move SVEBitwiseShiftbyVector into private section
2023-03-08 09:40:31 -08:00
Lioncache aeff4346e8 ARMEmitter: Handle SVE Write FFR group 2023-03-08 12:37:28 -05:00
Lioncache 5b51440c87 ARMEmitter: Put precondition checks inside SVEBitwiseShiftbyVector
Lets us deduplicate some code.
2023-03-08 12:16:47 -05:00
Lioncache baace30e30 ARMEmitter: Move SVEBitwiseShiftbyVector into private section
This is an internal helper, so it shouldn't be exposed.
2023-03-08 12:08:58 -05:00
Mai 11c6f97643 Merge pull request #2470 from Sonicadvance1/optimize_movs
OpDispatcher: Implements REP MOVS as Memcpy IR op
2023-03-07 15:17:12 -05:00
Mai 1f44037d9a Merge pull request #2471 from Sonicadvance1/fix_32bit_clock_nanosleep
x32: Fixes clock_nanosleep syscall
2023-03-07 15:15:59 -05:00
Ryan Houdek 3bc484b586 x32: Fixes clock_nanosleep syscall
Alwa's Awakening had an issue where it was spamming clock_nanosleep and
getting EINVAL, consuming a CPU thread to 100%.

This is due to two things in the previous implementation
1) glibc helper does some data munging and validation of inputs
2) We were updating remain even if TIMER_ABSTIME was used

This only cropped up because this error would only occur when EINTR
occured. Since the game is a mono game, the thread that would be
sleeping would get interrupted and then spin forever with corrupted
data.
2023-03-07 11:11:40 -08:00
Ryan Houdek 79170a23e5 OpDispatcher: Implements REP MOVS as Memcpy IR op
Just like the `REP STOS` instruction, this removes a fairly nasty bit of
code generation from our dispatcher. This instruction behaves like a
memcpy/memmove when not dealing with the faulting behaviour around the
instruction.

With a microbench things improves this instruction's throughput by 3.5%
to 25.5% depending on operating size.

There is a TODO for implementing this instruction with ARM's future MOPS
instructions which will accelerate this implementation further.
2023-03-06 16:39:10 -08:00
Ryan Houdek 81e52ca19f IR: Implements a memcpy operation
This matches the x86 REP MOVS instruction behaviour.
2023-03-06 16:38:43 -08:00
Ryan Houdek 5301036968 Int/JIT64: Fixes Memset Prefix
Accidentally missed appending the prefix to the address in x86 JIT and
interpreter
2023-03-06 16:37:11 -08:00
Ryan Houdek 2321d2ad97 Merge pull request #2469 from lioncash/change
ARMEmitter/ASIMDOps: Remove unnecessary template constraints
2023-03-06 15:54:47 -08:00
Lioncache 9c2d29ee4e ARMEmitter/ASIMDOps: Remove unnecessary template constraints
Now that our register types are strongly typed and don't implicitly convert
across one another, we can get rid of the templating on some member
functions, turning them into normal member functions.
2023-03-06 18:37:53 -05:00
Ryan Houdek 4b8ac0f8c6 Merge pull request #2468 from lioncash/asimd
ARMEmitter/ScalarOps: Move base opcodes into helper functions
2023-03-06 15:16:29 -08:00
Lioncache 70c93efdfa ScalarOps: Simplify Floating-point data-processing (3 source) category
We can move the opcode into the helper function
2023-03-06 17:26:24 -05:00
Lioncache f5e207c28f ScalarOps: Simplify Floating-point data-processing (2 source)
We can move the base opcode into the helper function
2023-03-06 17:21:36 -05:00
Lioncache 9c64a22280 ScalarOps: Simplify Floating-point conditional select category
We can move the base opcode into the helper
2023-03-06 17:15:47 -05:00
Lioncache 5f57fa4b77 ScalarOps: Simplify Floating-point data-processing (1 source) category
We can move the opcode into the helper function
2023-03-06 17:06:58 -05:00
Lioncache fe81523179 ScalarOps: Simplify Advanced SIMD scalar shift by immediate category
We can move the opcode into the helper
2023-03-06 16:53:04 -05:00
Lioncache 4be61aaaaf ScalarOps: Simplify Advanced SIMD scalar three same category
We can move the opcode into the helper
2023-03-06 16:46:36 -05:00
Lioncache da0681ca81 ScalarOps: Simplify Advanced SIMD scalar three different category
We can move the opcode into the helper function.
2023-03-06 16:37:52 -05:00
Lioncache 02258073e3 ScalarOps: Simplify Advanced SIMD scalar two-register miscellaneous category
We can move the opcode into the helper function.
2023-03-06 16:34:38 -05:00
Lioncache cde55d44ef ScalarOps: Simplify Advanced SIMD scalar two-register miscellaneous FP16 category
The opcode can be moved into the helper function, deduplicating it over
several functions
2023-03-06 16:12:07 -05:00
Lioncache 97c12ef351 ScalarOps: Simplify Advanced SIMD scalar three same FP16 category
We can move the opcode into the helper itself.
2023-03-06 16:07:12 -05:00
Ryan Houdek 94e0591601 Merge pull request #2467 from lioncash/unused
ARMEmitter: Remove unused helper functions
2023-03-06 12:59:06 -08:00
Lioncache b802f64ffe ScalarOps: Simplify Floating-point compare category
This category as well can have its opcode moved into the helper function
2023-03-06 15:58:40 -05:00
Lioncache 0e32d8a8b6 ScalarOps: Simplify Floating-point conditional compare category
We can move the opcode itself into the helper function, so it can be
deduplicated.
2023-03-06 15:57:57 -05:00
Lioncache fe8bd17c88 ARMEmitter: Simplify bit size retrieval helpers
We can collapse the array into a shift, and the others can just use 8 as
the starting value to be shifted instead of multiplying by 8.

Aside from the array change, everything else is likely already constant
folded, but it does make things a little more straightforward
2023-03-06 15:24:58 -05:00
Lioncache f333068d9a ARMEmitter: Remove unused helper functions
The templated variants were never used. Also we can just make these
helpers constexpr to begin with.
2023-03-06 15:21:48 -05:00
Ryan Houdek 3bd4b23af7 Merge pull request #2466 from lioncash/misc
ARMEmitter: Fully handle SVE Integer Misc - Unpredicated
2023-03-06 12:19:03 -08:00
Ryan Houdek 060745f98b Merge pull request #2465 from lioncash/constraint
ARMEmitter: Simplify uses of IsXOrWRegister
2023-03-06 12:08:26 -08:00
Lioncache 77ae7a8a3f ARMEmitter: Fully handle SVE Integer Misc - Unpredicated
Finishes off an integer category that, at present, has more floating point
operations in it than anything else.
2023-03-06 15:05:30 -05:00
Lioncache b989bb9569 ARMEmitter: Simplify uses of IsXOrWRegister
Turns out these can be nicely simplified to being directly in the
template declaration.

Neat little tip pointed out by @neobrain
2023-03-06 12:05:46 -05:00
Ryan Houdek 63ce78c41d Docs: Update for release FEX-2303 2023-03-06 08:50:41 -08:00
Mai fc38df2ff0 Merge pull request #2460 from Sonicadvance1/implement_memset
OpcodeDispatcher: Optimize REP STOS to MemSet operation
2023-03-04 11:31:58 -05:00
Mai 2fca207e14 Merge pull request #2462 from Sonicadvance1/fix_proton_2
FileManagement: Fixes Proton
2023-03-04 11:29:07 -05:00
Mai 308fa76aa3 Merge pull request #2463 from Sonicadvance1/update_rootfslinks
FEXRootFSFetcher: Update link to rootfs links file
2023-03-04 11:27:55 -05:00
Ryan Houdek b6ac26e0e9 FEXRootFSFetcher: Update link to rootfs links file
Switches to the new CDN which is significantly faster and has other
benefits.

In order to make sure we don't break old clients, switch to the new link
for a few months while leaving the old one operational.

The links file in the old CDN still points to the new rootfs links so
they get the performance improvement on old clients still.
2023-03-04 01:59:44 -08:00
Ryan Houdek ecd144de6a FileManagement: Fixes Proton
Need to ensure that dirfd is AT_FDCWD and also need to check flags
correctly.

Flags were incorrectly checking mode for O_WRONLY and also we should
check for O_APPEND. Split it out to a helper function just so it is
easier to see what is going on.

Fixes the issue of proton not finding `/lib64/ld-linux-x86-64.so.2`
2023-03-03 13:44:17 -08:00
Ryan Houdek 7e66508016 OpcodeDispatcher: Optimize REP STOS to MemSet operation
x86's REP STOS instruction is a memset (with element size!) with the
ability to choose a direction of execution.
Additionally it has a feature where if it faults part-way through the
copy, an application can catch the fault and continue afterwards to know
how many bytes got copied.

RCX is the counter which decrements for each element, and RDI is the
memory pointer. On fault these will reflect the last location that was
attempted to be written. FEX doesn't support this behaviour which makes
our lives easier.

Without supporting that feature, this turns in to a directional memset
by element size. Let's remove all the multiple blocks and just emit a
single IR operation to improve performance of the JIT.
Our generated code here was terrible, the IR was terrible, multiblock is
always slow with RA. Just a general overall improvement.

With profiling pressure-vessel this change deletes the hottest block that
appeared in the trace. This instruction is very commonly used for
memsetting a region to zero so it should be quite fast.

We can also optimize REP MOVS in the future with a Memcpy IR operation
in a similar fashion.

Additionally in the future these can be optimized to use ARM's new MOPS
instructions since the most common case is memset by byte. Which is when
we should expose the "Fast REP STOS" CPUID bit. Both setp/setm/sete and
cpyfp/cpyfm/cpyfe match `REP STOS` and `REP MOVS` respectively.
2023-03-03 09:16:28 -08:00
Ryan Houdek d6f50bf7b0 IR: Implement support for MemSet operation
This operation directly matches what the x86 STOS instruction does
without supporting its faulting behaviour.

STOS faulting behaviour is that RCX and RDI get updated to the last word
written. Which is something that FEX hasn't ever supported.
2023-03-03 09:16:28 -08:00
Ryan Houdek e7069f9f95 Merge pull request #2461 from lioncash/pair
ARMEmitter: Tidy up some assertion handling
2023-03-02 08:33:53 -08:00
Lioncache ea96ccb63d ARMEmitter: Add missing SVE floating-point compare vectors instructions
We're missing FACGE/FACGT and the aliases FACLE FACLT
2023-03-02 10:54:54 -05:00
Lioncache 4d4eac0987 ARMEmitter: Simplify SVE floating-point compare vectors
We can move the asserts into the helper function
2023-03-02 10:42:58 -05:00
Lioncache f0ee8a49b2 ARMEmitter: Simplify Emitter: SVE: SVE2 floating-point pairwise operations ops
We can centralize all of the assertion handling in the implementation
function.
2023-03-02 10:35:58 -05:00
Lioncache e0d8fc7c2b ARMEmitter: Simplify SVE2 integer halving add/subtract (predicated) ops
We can centralize all the assertion handling in the implementation
function.
2023-03-02 10:28:07 -05:00
Lioncache adf1de5562 ARMEmitter: Simplify SVE integer pairwise ops
We can centralize everything in the helper function itself, getting rid
of a few duplicated assertions.
2023-03-02 10:17:19 -05:00
Mai e310e29898 Merge pull request #2459 from Sonicadvance1/fix_pressure_vessel
FileManagement: Skip opening emulated writable files
2023-03-02 09:41:39 -05:00
Ryan Houdek 37ec68421c FileManagement: Skip opening emulated writable files
In the case that a file is getting opened to be created or writable then
skip EmuFD and rootfs searching for this file.
This fixes an edge case where if FEX was run with an unpacked rootfs
that was writable then pressure-vessel would break.

Fixes pressure-vessel with unpacked rootfs.
2023-03-02 00:26:34 -08:00
Ryan Houdek 41731e2680 Merge pull request #2458 from lioncash/pred
ARMEmitter: Remove predicate implicit conversion operators
2023-03-01 19:58:13 -08:00
Lioncache 5e6a3c6280 ARMEmitter: Remove predicate implicit conversion operators
Like with the vector registers, we can remove all implicit conversion
operators except the ones that convert down to the base PRegister class.

With this, all of the registers are now adequately constrained, so we
shouldn't have any wonky implicit conversions happening anymore.
2023-03-01 22:44:40 -05:00
Ryan Houdek e71e3ec930 Merge pull request #2457 from lioncash/sxtw
ARMEmitter: Make second sxtw parameter a WRegister
2023-03-01 19:35:31 -08:00
Lioncache 4cac100660 ARMEmitter: Make second sxtw parameter a WRegister
Matches the assembly use of it more closely.
2023-03-01 22:20:42 -05:00
Ryan Houdek 378e0692b9 Merge pull request #2456 from lioncash/reg
ARMEmitter: Remove implicit conversions from Register/XRegister/WRegister
2023-03-01 19:16:26 -08:00
Lioncache 678415c4c9 ARMEmitter: Remove implicit conversions from Register/XRegister/WRegister
Ensures that we're always explicit about the size of a register when
using APIs that enforce it.

The only implicit conversions we keep are conversions that convert down
to Register, but not anything that converts up the hierarchy or across
it.
2023-03-01 21:59:53 -05:00
Ryan Houdek e869b2fe67 Merge pull request #2455 from lioncash/comp
ARMEmitter: Remove predicate uint32_t conversion operators
2023-03-01 18:37:34 -08:00
Lioncache 2194a1027c ARMEmitter: Remove predicate uint32_t conversion operators
Now that we have dedicated comparison operators, we no longer need to
keep these implicit conversion operators around.
2023-03-01 21:16:55 -05:00
Lioncache 52b4378e49 ARMEmitter: Add comparison functions to register types
Gets rid of the need to compare indices directly in order to compare
register equality
2023-03-01 21:15:46 -05:00
Ryan Houdek 0f45318040 Merge pull request #2454 from lioncash/convert
ARMEmitter: Remove most implicit conversion operators for vector register types
2023-03-01 17:57:21 -08:00
Ryan Houdek 21fbcef0bd Merge pull request #2453 from lioncash/explicit
ARMEmitter: Make VRegister constructor explicit
2023-03-01 17:54:49 -08:00
Ryan Houdek ef02083767 Merge pull request #2452 from lioncash/consecutive
ARMEmitter: Handle sequential registers in lists nicer
2023-03-01 17:53:47 -08:00
Ryan Houdek 24904f48c4 Merge pull request #2451 from lioncash/saddl
ARMEmitter: Simplify size handling Advanced SIMD 3 different group
2023-03-01 17:45:42 -08:00
Lioncache 9461ab5094 ARMEmitter: Remove conversion operators for VRegister 2023-03-01 18:40:12 -05:00
Lioncache 66c8b14470 ARMEmitter: Remove conversion operators for QRegister 2023-03-01 18:29:11 -05:00
Lioncache 6ab78ca93b ARMEmitter: Remove conversion operators for DRegister 2023-03-01 18:20:32 -05:00
Lioncache 2fe808f5cd ARMEmitter: Remove conversion operators for SRegister 2023-03-01 18:02:31 -05:00
Lioncache 81a94b9ffe ARMEmitter: Remove conversion operators for BRegister 2023-03-01 18:00:24 -05:00
Lioncache 0d87ed46da ARMEmitter: Remove conversion operators for HRegister 2023-03-01 17:58:06 -05:00
Mai 545a216da6 Merge pull request #2448 from Sonicadvance1/optimize_openat
EmulatedFiles: Optimize openat handler
2023-03-01 17:20:14 -05:00
Lioncache 11f65df554 ARMEmitter: Make VRegister constructor explicit
All other parameter taking constructors for the other register types are
explicit, so this just makes behavior more consistent.
2023-03-01 16:48:10 -05:00
Lioncache 29ff642499 ARMEmitter: Make use of sequential register helper
Fixes assertion behavior on quite a bit of ASIMD load-store operations
as well as a few SVE ops as well
2023-03-01 15:25:46 -05:00
Lioncache d36517a9d3 ARMEmitter: Add helper for determining if vectors are sequential
A few vector instructions that take register lists often require
vector registers within the list to be sequential in the form of an
increasing list modulo the register file size.

For example:

v1,  v2, v3, v4
v31, v0, v1, v2

both fit these requirements.

This will be used to enforce this restriction within the asserts from a
single place.
2023-03-01 15:24:07 -05:00
Lioncache 83419b410d ARMEmitter: Simplify size handling Advanced SIMD 3 different group
A large amount of size handling in this category is just decrementing
the size by 1, so we can tidy up a bunch of conditionals by just doing
that instead.
2023-03-01 11:04:40 -05:00
Mai 77fad28b69 Merge pull request #2447 from Sonicadvance1/add_hypervisorbit_hide_option
CPUID: Adds an config option to hide hypervisor bit
2023-02-28 10:59:21 -05:00
Mai 70aefc9db2 Merge pull request #2450 from Sonicadvance1/fix_fexserver_zombie
FEXServerClient: Fixes instance where FEXServer can create a zombie
2023-02-28 10:58:11 -05:00
Mai d2e0adf540 Merge pull request #2449 from Sonicadvance1/fix_fexserver_daemon_systemd
FEXServer: Change systemd service environment variable key
2023-02-28 10:57:09 -05:00
Ryan Houdek 84060cd947 FEXServerClient: Fixes instance where FEXServer can create a zombie
When FEXServer is daemonizing through an instance of FEXLoader or
FEXInterpreter, it would leave a zombie process which was waiting for us
to read the process status.
Since we don't care about the child status and don't want to get blocked
by waitpid, just ignore the signal.

This tells the kernel that we don't care about the signal and will kill
the zombie process immediately.
Didn't notice this before since FEXServer started failing to daemonize.
2023-02-28 05:09:16 -08:00
Ryan Houdek aaf17b6d41 FEXServer: Change systemd service environment variable key
It looks like `SYSTEMD_EXEC_PID` can leak through to the executable
environment in regular situations. Instead let's key off of
`INVOCATATION_ID` which doesn't leak through.

Fixes an edge case behaviour where FEXServer wouldn't daemonize in some
systemd environments.
2023-02-28 05:07:02 -08:00
Ryan Houdek 8ded25ada7 EmulatedFiles: Optimize openat handler
Fixes #2443
I found out with some profiling that this we were spending a decent
amount of time with the `openat` syscall in heavily utilized situations.
While not super common in active gameplay situations, it matters
significantly in loading screens that this is fairly optimal.

The bulk of the time is spent in the emulated files handler to ensure
that whatever path we are given, we can capture file paths that we need
to emulate. The largest contributor being the std::filesystem::canonical
function call.

A couple of optimizations in place here.
1) Do a quick hashmap check right at the start to see if we exactly fit
2) Change from `std::fs::canonical` to `realpath`
3) Switch `GetEmulatedFDPath` to not use optional so it stops building
   on the stack

I'm still not super happy with the performance of `realpath` and also
not happy that we still need to use `lexically_normal` in one code path.
But short of writing a super hand-optimized `realpath` that fits our
constraints, I don't think we can do better.

Micro benchmark needs to test four different situations due to this
optimization.
1) Non-EmuFD path
2) Non-EmuFD path with dirfs
3) EmuFD path
4) EmuFD path with dirfs

And the performance improvement for each situation respectively
1) 12% performance improvement
  - 213413 openat syscalls/s -> 238999 syscalls/s
2) 17% performance improvement
  - 202085 openat syscalls/s -> 237309 syscalls/s
3) 17% performance improvement (/proc/cpuinfo)
  - 56616 openat syscalls/s -> 66231 syscalls/s
  - Includes overhead of generating temp FD and close syscall
4) 5% performance improvement (/proc/cpuinfo)
  - 51080 openat syscalls/s -> 53956 syscalls/s
  - Includes overhead of generating temp FD and close syscall

And for sake of comparison to the non-emulated system; My test system
can hit around 1-1.1 million openat syscalls per second in the same
microbench.

Nice little performance uplift.
2023-02-28 04:00:28 -08:00
Ryan Houdek 5b9fe8f26b CPUID: Adds an config option to hide hypervisor bit
This is known to cause issues in some cases. We hit the first game that
checks for this bit and early exits if it is found.

Lets the MMORPG Tibia run in non-VM situations.
Looks like they have more checks for VMs other than hypervisor bit, so
running under Parallels still won't work. Running on bare Linux is fine.
2023-02-27 23:11:23 -08:00
Ryan Houdek e65b429c83 Merge pull request #2446 from lioncash/cpy
ARMEmitter: Simplify advanced SIMD copy
2023-02-27 19:52:05 -08:00
Lioncache dd290f129f ARMEmitter: Simplify advanced SIMD copy
Same behavior, but collapses some if statements.
2023-02-27 22:32:39 -05:00
Ryan Houdek 1832cc80d6 Merge pull request #2445 from lioncash/unsigned
ARMEmitter: Centralize handling for unsigned offset load-stores
2023-02-27 18:18:30 -08:00
Ryan Houdek fe1faf9ebe Merge pull request #2444 from lioncash/scalar
ARMEmitter: Handle SVE Integer Compare - Scalars group
2023-02-27 18:16:31 -08:00
Lioncache 12d0a7fa98 ARMEmitter: Use constants for unsigned offset encoding limits
Allows us to give some names to these constants that are used in the
JIT instead of writing them by hand.
2023-02-27 17:44:17 -05:00
Lioncache c1b08079f3 ARMEmitter: Strengthen unsigned immediate load/store helper
Centralizes all the shifting behavior and whatnot into a single
function, making everything much more localized.

Also gets rid of a lot of magic constants related to the encoding limits
of immediates.
2023-02-27 17:44:13 -05:00
Mai d688026fe4 Merge pull request #2442 from Sonicadvance1/fix_misaligned_stack_signals
Dispatcher: Fixes crash with misalign stack returning from signal
2023-02-27 14:48:39 -05:00
Lioncache 1c388b455a ARMEmitter: Move missed SVE public helpers into private section 2023-02-27 14:45:59 -05:00
Lioncache 9426abc98d ARMEmitter: Handle SVE pointer conflict compare group 2023-02-27 14:27:29 -05:00
Lioncache 5ee2db34a7 ARMEmitter: Handle SVE conditionally terminate scalars group 2023-02-27 14:22:57 -05:00
Lioncache ad37c19043 ARMEmitter: Handle SVE integer compare scalar count and limit group 2023-02-27 14:09:40 -05:00
Ryan Houdek 0e6c5911b8 Dispatcher: Fixes crash with misalign stack returning from signal
When we were taking a signal that had a misaligned stack, we would store
the host stack at a weird offset.

After that point when we were trying to sigreturn we wouldn't know the
alignment of the stack coming back and we would try loading the host
stack from the wrong offset. Easy fix is to just align the host stack
location.

Fixes Ender Lilies, which was consistently crashing from a SIGCHLD due
to having a misaligned stack.

Side-change: Move the cookie check to the start of the restore. Doesn't
make sense to check the cookie after restoring state since it could be
quite wrong.
2023-02-26 19:58:22 -08:00
Ryan Houdek f2aa0026b5 Merge pull request #2439 from lioncash/log
Emitter/ALUOps: Fix typos in log messages
2023-02-23 16:44:19 -08:00
Lioncache 553efbeb29 Emitter/ALUOps: Fix typos in log messages
Fixes a few incorrect instruction names in the logs.
2023-02-23 19:14:45 -05:00
Ryan Houdek b39a882a2d Merge pull request #2438 from lioncash/restrict
OpcodeDispatcher: Restrict partial XMM stores to FPRs in StoreResult_WithOpSize
2023-02-23 16:06:53 -08:00
Lioncache 85f7f8e6c0 OpcodeDispatcher: Restrict partial XMM stores to FPRs in StoreResult_WithOpSize
As far as I know, nothing actually uses this path. Partially resolves
the TODO of dealing with partial writes.
2023-02-23 18:48:24 -05:00
Ryan Houdek 4d25de31de Merge pull request #2437 from lioncash/dup
OpcodeDispatcher: Remove now unused _VDupElement path in LoadSource_WithOpSize
2023-02-23 13:58:09 -08:00
Ryan Houdek 9e01730c6c Merge pull request #2436 from lioncash/builtin
Arm64Emitter: Use bit utils wrapper over __builtin_ffs
2023-02-23 13:51:24 -08:00
Lioncache fea3ee1298 OpcodeDispatcher: Remove now unused _VDupElement path in LoadSource_WithOpSize
XMM instances can't use high indices anymore, since we've gotten rid of
the only flag that allows this scenario to occur.
2023-02-23 16:09:03 -05:00
Lioncache e1c42315ed Arm64Emitter: Use bit utils wrapper over __builtin_ffs
Just keeps the use of builtins contained to one place.
2023-02-23 15:23:07 -05:00
Ryan Houdek 165db37c8d Merge pull request #2434 from lioncash/predmisc
ARMEmitter: Finish off SVE Predicate Misc group
2023-02-23 12:20:03 -08:00
Ryan Houdek 9b23ae9133 Merge pull request #2433 from lioncash/subsw
OpcodeDispatcher: Handle VPHSUBSW
2023-02-23 12:17:58 -08:00
Ryan Houdek f951a406e6 Merge pull request #2435 from lioncash/mov
OpcodeDispatcher: Share MOVHPD implementation with MOVHPS
2023-02-23 12:16:27 -08:00
Lioncache 497b5c0561 X86Tables: Reclaim FLAGS_SF_HIGH_XMM_REG as an unused flag
Now that we've moved MOVHPS over to sharing the implementation of
MOVHPD, the FLAGS_SF_HIGH_XMM_REG is now unused.

Since we're supporting AVX, this flag is kind of weird in terms of
behavior, since what determines the high part of a register is now
situationally different.

Also it's much more explicit to perform the insert directly in the
implementation of instructions, than relying on a flag to do it for us.

So, instead of keeping it around, we can reclaim it as unused for use
with any necessary behavior that we would require in the future.
2023-02-23 14:22:25 -05:00
Lioncache 95393b07fb OpcodeDispatcher: Share MOVHPD implementation with MOVHPS
These instructions essentially have the same behavior. This also allows
us to remove the only used instance of FLAGS_SF_HIGH_XMM_REG, which,
given that we now support AVX, has ambiguous use.

While we're at it, we can expand the tests to make use of the store to
memory variant.

Also removes an erroneous copy-pasted comment about ZEXTing. This is
from the MOVQ implementation function. MOVHPS/MOVHPD don't do any
ZEXTing, they either store to memory or insert into a register.
2023-02-23 14:07:22 -05:00
Lioncache 372da1b820 ARMEmitter: Handle PNEXT
Now, with the helper in place, we can implement PNEXT and finish off the
SVE Predicate Misc group.
2023-02-23 11:59:39 -05:00
Lioncache 61a59d0314 ARMEmitter: Unify SVE Predicate Misc group under single helper
Centralizes the implementations and also gets rid of some code in the
process.
2023-02-23 11:51:03 -05:00
Lioncache 1045e05870 OpcodeDispatcher: Handle VPHSUBSW 2023-02-23 10:55:57 -05:00
Lioncache 052872725c OpcodeDispatcher: Factor out PHSUBS implementation into helper
This will allow it to be shared in the AVX implementation.
2023-02-23 10:32:48 -05:00
Ryan Houdek 4d655218ab Merge pull request #2431 from lioncash/brk
ARMEmitter: Handle SVE partition break categories
2023-02-22 21:13:33 -08:00
Lioncache 78ba195b66 ARMEmitter: Handle SVE partition break condition category 2023-02-22 22:49:37 -05:00
Lioncache ae2b28716d ARMEmitter: Handle SVE propagate break to next partition category 2023-02-22 22:41:55 -05:00
Lioncache 326e5e8d57 ARMEmitter: Handle propagate break from previous partition category 2023-02-22 22:35:31 -05:00
Ryan Houdek 0a8fc2cbef Merge pull request #2430 from lioncash/assert
ARMEmitter: Handle SVE integer compare with wide elements category
2023-02-22 18:57:33 -08:00
Lioncache 552293b226 ARMEmitter: Handle SVE integer compare with wide elements category
We can piggy-back on top of the existing SVEIntegerCompareVector to make
these trivial to implement.
2023-02-22 21:03:35 -05:00
Lioncache 751a4c8019 ARMEmitter: Move assertion into SVEIntegerCompareVector
Same behavior, but centralizes the assertion. While we're at it, we can
also add another assert to ensure that only predicates p0-p7 are used.
2023-02-22 20:14:49 -05:00
Ryan Houdek 68b2072eab Merge pull request #2429 from lioncash/align
OpcodeDispatcher: Handle alignment for MOVAPS a little better
2023-02-22 14:40:28 -08:00
Lioncache e3cac40b1b OpcodeDispatcher: Fix SSE MOVAPS variants being treated as MOVUPS
0x10/0x11 in the two byte op table corresponds to MOVUPS
0x28/0x29 in the two byte op table corresponds to MOVAPS
2023-02-22 15:56:15 -05:00
Ryan Houdek 9b123353b3 Merge pull request #2428 from lioncash/hsub
OpcodeDispatcher: Handle VHSUBPD/VHSUBPS
2023-02-22 11:43:10 -08:00
Lioncache f2c0c55b9c OpcodeDispatcher: Handle VHSUBPS 2023-02-22 14:27:51 -05:00
Ryan Houdek a4c694ffc7 Merge pull request #2427 from lioncash/pred
ARMEmitter: Finish off SVE Permute Vector - Predicated group
2023-02-22 11:26:49 -08:00
Lioncache 1eb722dea7 OpcodeDispatcher: Handle VHSUBPD 2023-02-22 14:12:11 -05:00
Lioncache a6746988d7 x86_64/VectorOps: Fix behavior of UnZip2 with 64-bit element 256-bit vectors
The 256-bit variant of vshufpd uses extra immediate bits rather than the
same bits for the lower lane.
2023-02-22 14:12:11 -05:00
Lioncache 0a1707f1bd OpcodeDispatcher: Factor HSUBP implementation into helper
Will be used for implementing the AVX variants of the same instructions.
2023-02-22 12:12:55 -05:00
Lioncache 23b9d8e108 ARMEmitter: Add check for registers being consecutive in constructive SPLICE
Will catch cases where registers aren't consecutive in the constructive
variant. While we're at it, we can also amend EXT's similar but slightly wrong
consecutive check.

Also adds tests to ensure these corner-cases hold.
2023-02-22 11:59:58 -05:00
Lioncache 3fa44604ba ARMEmitter: Make SPLICE use SVEPermuteVectorPredicated
These are in the same instruction category, so we can use the helper to
simplify the implementation.
2023-02-22 11:48:11 -05:00
Lioncache d3bc0c084d ARMEmitter: Make CPY (SIMD&FP) and CPY (scalar) use SVEPermuteVectorPredicated
These fall under the same instruction category, so we can use the helper
to simplify the implementation.
2023-02-22 11:29:07 -05:00
Lioncache e206414919 ARMEmitter: Make COMPACT use SVEPermuteVectorPredicated
This falls under the same category of instructions, so we can use it to
simplify the implementation.
2023-02-22 11:23:37 -05:00
Lioncache e635cc5404 ARMEmitter: Use predicated helper with revb/revh/revw/rbit
Since these are under the same category, we can merge these and get rid
of a now unnecessary helper.
2023-02-22 11:17:12 -05:00
Lioncache 822d67467b ARMEmitter: Handle SVE conditionally extract element to GPR/scalar categories 2023-02-22 11:10:13 -05:00
Lioncache e043d2c0f5 ARMEmitter: Handle SVE conditionally broadcast element to vector category 2023-02-22 10:55:34 -05:00
Lioncache d58c4405f7 ARMEmitter: Handle extract element to general register/scalar categories 2023-02-22 10:47:49 -05:00
Mai 66d879f387 Merge pull request #2400 from Sonicadvance1/rip_reconstruct
Dispatcher: Support reconstructing RIP from block entry
2023-02-22 09:48:02 -05:00
Mai 55d3edb8e6 Merge pull request #2426 from Sonicadvance1/optimize_getemulatedpath
FileManagement: Optimize GetEmulatedFDPath with an FD!
2023-02-22 09:46:44 -05:00
Ryan Houdek 98f0f22f41 FileManagement: Optimize GetEmulatedFDPath with an FD!
Performance stats up front:
This improves pressure-vessel startup time on my test device by 10.1%
Improving the startup time from 9.71425 seconds to 8.7421 seconds.

Most filesystem based syscalls support a file descriptor version with an
*at suffix. This allows us to do these syscalls with pathnames that are
relative to the directory FD that is passed to the syscall.

This is pretty much exactly what we want when we are searching for files
inside of our rootfs. The only quirk ends up being that we are getting
passed absolute paths. This ends up being very simple to workaround by
stripping off the front '/' character. Doing this is just offsetting the
pointer passed to the syscall by one byte.

This does require having two temporary buffers of size PATH_MAX passed
to the handler since just like in the other implementation, we need to
keep the previous result around. The difference being now that we aren't
doing a bunch of std::string temporary manipulation and now we are
returning one of the passed in buffers back depending on the result.
2023-02-22 01:25:40 -08:00
Ryan Houdek 5f574fb935 Merge pull request #2425 from lioncash/xop
VEXTables: Remove VPERMIL2PD and VPERMIL2PS entries
2023-02-20 18:27:32 -08:00
Ryan Houdek 618f5bb869 Merge pull request #2424 from lioncash/permil
OpcodeDispatcher: Handle register variants of VPERMILPD/VPERMILPS
2023-02-20 18:02:14 -08:00
Lioncache b5ca5f173e VEXTables: Remove VPERMIL2PD and VPERMIL2PS entries
These are actually XOP instructions. That, despite being so, are encoded
using a VEX prefix.
2023-02-20 20:58:55 -05:00
Lioncache 5cf6a680bb OpcodeDispatcher: Handle register variants of VPERMILPD/VPERMILPS 2023-02-20 20:32:31 -05:00
Ryan Houdek 645f40bb96 Merge pull request #2423 from lioncash/permd
OpcodeDispatcher: Handle VPERMD/VPERMPS
2023-02-20 15:20:59 -08:00
Ryan Houdek 268deddd09 Merge pull request #2422 from lioncash/phadds
OpcodeDispatcher: Handle VPHADDSW
2023-02-20 15:20:20 -08:00
Ryan Houdek e4488b0cfc Merge pull request #2421 from lioncash/index
ARMEmitter: Handle SVE index generation category
2023-02-20 15:16:33 -08:00
Lioncache 65b9dcd20b OpcodeDispatcher: Handle VPERMPS
With the VPERMD work in place, this is trivial to support.
2023-02-20 17:00:39 -05:00
Lioncache b2c333c383 OpcodeDispatcher: Handle VPERMD 2023-02-20 17:00:35 -05:00
Lioncache 1ea53c65ab x86_64/VectorOps: Handle 8-bit VShlI IR op
Useful for handling VPERMD.
2023-02-20 16:45:13 -05:00
Lioncache 8beae0fce4 OpcodeDispatcher: Add VTrn/VTrn2 IR opcodes
Provides a convenient way to propogate indices at given intervals in
vectors. This makes permutation instructions a little less annoying to
implement.
2023-02-20 16:44:14 -05:00
Lioncache add775c5cd OpcodeDispatcher: Handle VPHADDSW 2023-02-20 12:25:28 -05:00
Lioncache f3e6f62356 OpcodeDispatcher: Factor PHADDS implementation into helper
This will be used to also handle the VEX variant of PHADDSW
2023-02-20 12:00:38 -05:00
Lioncache 59ab10f155 ARMEmitter: Move SVE instruction helpers into privare section
Moves some instruction helpers that existed outside of the private
section of the class back into them, so that we're not exposing
unnecessary things in the interface.
2023-02-20 11:39:21 -05:00
Lioncache 2e1bd4b32b ARMEmitter: Handle SVE index generation category 2023-02-20 11:30:09 -05:00
Mai f71f2445db Merge pull request #2389 from Sonicadvance1/remove_context_c_interface
FEXCore: Removes C wrapper interface
2023-02-20 10:08:07 -05:00
Mai f6e2fe1515 Merge pull request #2420 from Sonicadvance1/fix_syscall_race
Arm64: Fixes a race condition on syscall spilling SRA
2023-02-20 10:06:53 -05:00
Mai 11c8db5a14 Merge pull request #2419 from Sonicadvance1/cortex_c_classify
Scripts: Update fit_native script for X1C/A78C
2023-02-20 10:06:07 -05:00
Mai 65b2da20d6 Merge pull request #2418 from Sonicadvance1/optimize_aluop_dispatcher
OpcodeDispatcher: Optimize ALUOp handler
2023-02-20 10:05:47 -05:00
Ryan Houdek 273f5e1f26 Arm64: Fixes a race condition on syscall spilling SRA
When executing a non-inlined syscall, we spill all static registers.
We weren't storing in to the thread context that we have done this.
If a signal occured between FEX returning from the syscall (after the
blr) and before the `FillStaticRegs` then the signal handler would get
the incorrect register state.

This typically manifested as Steam getting a SIGCHLD, trying to recover
the guest stack pointer, and it that pointer would be zero or some other
corrupt value. Thus crashing inside of the signal handler.

Surprising that we hadn't hit this way more before this point, must have
needed hardware that tickled the race condition *just* right.
2023-02-19 16:12:06 -08:00
Ryan Houdek 35af4bd42a FEXCore: Removes C wrapper interface
This has been a long time coming. The C interface has been a thorn in
our side for no reason for a long time.

The purpose of this step is to remove the C interface without changing
behaviour as much as possible. This means that with this commit there
are still some bad practices but the remaining issues will be solved
with followup PRs.

Primarily, we still have a `DestroyContext(CTX)` static function which calls
the Context implementation's `DestroyContext` and does a raw C++ delete.

Follow up PR will remove that, but I didn't want to touch it yet since
it'll require checking to ensure the unique_ptr changes play nice with
our allocator hooking. Which this is already a huge PR without trying to
change behaviour.
2023-02-19 11:59:11 -08:00
Ryan Houdek 7f1464b135 Scripts: Update fit_native script for X1C/A78C
Cortex-X1C and A78C are relatively minor changes to their non-C
counterparts. Support classifying them in case clang understands them.

Fixes a minor perf regression noticed on the Lenovo X13s while testing.
2023-02-18 23:18:48 -08:00
Ryan Houdek e594b2c4c7 OpcodeDispatcher: Optimize ALUOp handler
Take a leaf from the Vector ops and have the jump entry choose the IR
op.
Also generate one atomic op and modify the IR type in the locked memory
type just like the non locked memory path.

This class of instructions in the number one instruction type percentage
wise, so making this more optimal will be a win.

It's a fairly minor optimization so it should be a small impact.
2023-02-18 03:32:43 -08:00
Ryan Houdek 2aead5aec2 Config: Removes the x86dec_SynchronizeRIPOnAllBlocks option
This is no longer necessary since we reconstruct up to block entry from
the previous commit.
2023-02-18 02:48:55 -08:00
Ryan Houdek c3f1f602fe Dispatcher: Support reconstructing RIP from block entry
This allows us to not update RIP on block entry, but still allow
reconstructing the RIP up until that point.

While still not full RIP reconstruction, this lets us update the signal
context's RIP just like the `x86dec_SynchronizeRIPOnAllBlocks` without
eating the cost of writing to RIP on block entry.
2023-02-18 02:48:55 -08:00
Ryan Houdek 9c256bfe96 Merge pull request #2413 from lioncash/unpred
ARMEmitter: Handle a few more vector permutation categories
2023-02-15 14:51:59 -08:00
Ryan Houdek b5bc8cd294 Merge pull request #2416 from lioncash/mov
VectorOps: Remove unnecessary mov in VUShrNI2/VSQXTN2/VSQXTUN2
2023-02-15 14:46:02 -08:00
Lioncache e78b573610 VectorOps: Remove unnecessary mov in VUShrNI2/VSQXTN2/VSQXTUN2
We can move the initial move down by SPLICE, which not only lets us turn
it into a MOVPRFX, but also we can safely move into the final
destination register directly, since we can be sure there's no
potential dependencies at this point
2023-02-15 17:14:26 -05:00
Mai 81a89ab747 Merge pull request #2415 from Sonicadvance1/spillsra_fix
Dispatcher: Fixes guest stack register usage
2023-02-15 15:51:49 -05:00
Ryan Houdek fd17a3de50 Dispatcher: Fixes guest stack register usage
Fixes #2410

We were pulling the guest RSP before spilling static registers back to
the state.
Move this to after we spill SRA state to fix this bug.

Thanks to @ifquant for diving in, identifying, and finding the exact bug.
2023-02-15 12:21:11 -08:00
Ryan Houdek 2f260ae6ad Merge pull request #2414 from lioncash/sve-ex
Arm64/VectorOps: Use SVE only with 256-bit op sizes
2023-02-15 12:11:05 -08:00
Lioncache bf7118fc85 Arm64/VectorOps: Use SVE only with 256-bit op sizes
Keeps all of the IR ops consistent with each other. Also removes some
redundant scalar checks that weren't really necessary.
2023-02-15 14:46:05 -05:00
Ryan Houdek a90f5363dd Merge pull request #2412 from lioncash/ptest
OpcodeDispatcher: Handle VPTEST
2023-02-15 10:31:12 -08:00
Ryan Houdek ab03e59500 Merge pull request #2411 from lioncash/zero
OpcodeDispatcher: Use VectorZero over VectorImm in InsertPSOpImpl
2023-02-15 10:30:23 -08:00
Lioncache a59d700bbe ARMEmitter: Handle SVE Permute Predicate category 2023-02-15 12:56:39 -05:00
Lioncache 7cf27a7c26 ARMEmitter: Handle SVE Permute Vector - Unpredicated category 2023-02-15 12:18:42 -05:00
Lioncache 14e1d16710 OpcodeDispatcher: Handle VPTEST 2023-02-15 11:24:37 -05:00
Lioncache 203f29a91f OpcodeDispatcher: Use VectorZero over VectorImm in InsertPSOpImpl
A little more straightforward than using VectorImm for the same purpose.
2023-02-15 09:37:54 -05:00
Ryan Houdek 25f0a03ceb Merge pull request #2407 from lioncash/mov
OpcodeDispatcher: Handle VMOVSD/VMOVSS
2023-02-14 22:36:13 -08:00
Lioncache 3ced41414e OpcodeDispatcher: Handle VMOVSD 2023-02-15 01:18:54 -05:00
Lioncache 1a64b26d03 OpcodeDispatcher: Handle VMOVSS 2023-02-15 01:18:15 -05:00
Ryan Houdek efafe0e6e9 Merge pull request #2408 from lioncash/pmaddwd
OpcodeDispatcher: Handle VPMADDWD
2023-02-14 17:52:57 -08:00
Ryan Houdek 35746c7669 Merge pull request #2406 from lioncash/shuffle
OpcodeDispatcher: Handle VSHUFPD/VSHUFPS
2023-02-14 17:47:39 -08:00
Lioncache 4a69b87cb9 OpcodeDispatcher: Handle VPMADDWD 2023-02-14 18:50:37 -05:00
Lioncache fb2de47e73 OpcodeDispatcher: Factor out PMADDWD implementation to helper
This will be used to centralize code to also implement the AVX variant.
2023-02-14 18:38:12 -05:00
Lioncache bcee3e9374 OpcodeDispatcher: Handle VSHUFPS 2023-02-14 16:46:03 -05:00
Lioncache 6d87154ac8 OpcodeDispatcher: Handle VSHUFPD 2023-02-14 16:46:03 -05:00
Lioncache c5d799df8c OpcodeDispatcher: Make SHUFOpImpl suitable for AVX
Drops in the AVX-specific bits into the helper in preparation for
implementing VSHUFPD and VSHUFPS
2023-02-14 16:45:32 -05:00
Lioncache 449645669a OpcodeDispatcher: Move SHUFOp implementation to helper function
Will be useful for handling both the SSE and AVX variants in the same
place.
2023-02-14 16:43:33 -05:00
Ryan Houdek 3ac7b2cddf Merge pull request #2405 from lioncash/shufw
OpcodeDispatcher: Handle VPSHUFD/VPSHUFHW/VPSHUFLW
2023-02-14 10:36:48 -08:00
Lioncache 504d409cf6 OpcodeDispatcher: Handle VPSHUFD 2023-02-14 13:13:09 -05:00
Lioncache 29a6d584a9 OpcodeDispatcher: Handle VPSHUFHW 2023-02-14 12:48:42 -05:00
Lioncache 310fcf969c OpcodeDispatcher: Handle VPSHUFLW 2023-02-14 12:32:47 -05:00
Ryan Houdek b329442c09 Merge pull request #2404 from lioncash/dup
IR: Add VDupFromGPR
2023-02-13 14:37:59 -08:00
Lioncache f4d799abdd OpcodeDispatcher: Make use of VDupFromGPR where applicable
Simplifies some of the IR usage.
2023-02-13 16:52:51 -05:00
Lioncache 4bb7f49c2a IR: Add VDupFromGPR
Allows broadcasting constants into vectors from GPRs. Resolves the only
remaining TODOs within our vector ops.
2023-02-13 16:52:47 -05:00
Ryan Houdek f7f2dc2210 Merge pull request #2403 from lioncash/err
ARMEmitter/ASIMDOps: Amend a few error logs
2023-02-13 12:31:52 -08:00
Ryan Houdek d40812929f Merge pull request #2402 from lioncash/shufb
OpcodeDispatcher: Handle VPSHUFB
2023-02-13 12:18:16 -08:00
Lioncache c381185a7d ARMEmitter/ASIMDOps: Amend a few error logs
A few were logging out the wrong instruction name on a precondition
failure.
2023-02-13 15:17:09 -05:00
Lioncache ac5d09885e OpcodeDispatcher: Handle VPSHUFB 2023-02-13 14:47:37 -05:00
Lioncache d9a505e22e OpcodeDispatcher: Factor PSHUFB implementation into helper
Will let us centralize the implementation for PSHUFB and VPSHUFB
2023-02-13 12:39:27 -05:00
Ryan Houdek a96ad0fc9d Merge pull request #2401 from lioncash/palign
OpcodeDispatcher: Handle VPALIGNR
2023-02-13 09:35:19 -08:00
Lioncache 9268a356f6 OpcodeDispatcher: Handle VPALIGNR 2023-02-13 10:53:02 -05:00
Lioncache 92141d3edc OpcodeDispatcher: Factor PALIGNR code into helper
Will allow us to centralize the implementation of PALIGNR and VPALIGNR.
2023-02-13 09:49:26 -05:00
Ryan Houdek 8c8b680640 Merge pull request #2398 from lioncash/sve2acc
ARMEmitter: Handle SVE2 Accumulate category
2023-02-10 23:42:43 -08:00
Lioncache 504be62a92 ARMEmitter: Handle SVE2 integer absolute difference and accumulate 2023-02-11 00:27:07 -05:00
Lioncache 62e2f1b45d ARMEmitter: Handle SVE2 bitwise shift and insert category 2023-02-11 00:27:04 -05:00
Lioncache 880cc72842 ARMEmitter: Handle SVE2 bitwise shift right and accumulate 2023-02-11 00:24:18 -05:00
Lioncache feacd897fc ARMEmitter: Handle SVE2 integer add/sub long with carry category 2023-02-11 00:24:18 -05:00
Lioncache fdd950e1d1 ARMEmitter: Handle SVE2 integer absolute difference and accumulate long category 2023-02-11 00:24:17 -05:00
Lioncache b02af95629 ARMEmitter: Handle SVE2 complex add category 2023-02-10 22:31:45 -05:00
Ryan Houdek 2bd64ad24e Merge pull request #2396 from lioncash/narrow
ARMEmitter: Finish off SVE Misc category
2023-02-09 10:28:52 -08:00
Lioncache fa95a823c9 ARMEmitter: Handle SVE2 bitwise shift left long category 2023-02-09 06:30:43 -05:00
Lioncache 399ed61380 ARMEmitter: Handle SVE2 integer add/sub interleaved long 2023-02-09 05:44:44 -05:00
Lioncache 71550e29eb ARMEmitter: Handle SVE integer matrix multiply accumulate 2023-02-09 05:34:11 -05:00
Lioncache 1b8d8f8280 ARMEmitter: Handle SVE2 interleaved XOR category 2023-02-09 05:16:59 -05:00
Lioncache ff6c70f5e1 ARMEmitter: Handle SVE2 bitwise permute category 2023-02-09 05:11:48 -05:00
Lioncache 13ee2b5ec4 ARMEmitter: Handle SVE2 add/sub narrow high part 2023-02-09 05:01:14 -05:00
Ryan Houdek 3c1ba846f7 Merge pull request #2394 from lioncash/cpy
ARMEmitter: Handle CPY (scalar) and CPY (SIMD&FP, scalar)
2023-02-09 01:18:17 -08:00
Lioncache 1b4488e7a3 ARMEmitter: Remove outdated histogram TODO
This was implemented along with histcnt
2023-02-09 04:04:39 -05:00
Lioncache 6b4df4c998 ARMEmitter: Handle CPY (SIMD&FP, scalar) 2023-02-09 03:53:31 -05:00
Lioncache 2a1ef0ba56 ARMEmitter: Handle CPY (scalar) 2023-02-09 03:46:07 -05:00
Ryan Houdek dd2e70e4aa Merge pull request #2393 from lioncash/wide2
ARMEmitter: Handle predicated wide shifts
2023-02-08 23:23:38 -08:00
Lioncache 915a8b23ae ARMEmitter: suffix unpredicated wide shifts
Keeps the naming convention consistent while avoiding clashing
overloads.
2023-02-09 02:00:28 -05:00
Lioncache 6eeafd0724 ARMEmitter: Handle predicated wide shifts 2023-02-09 01:58:23 -05:00
Ryan Houdek e6fc159d88 Merge pull request #2390 from lioncash/ext
OpcodeDispatcher: Handle VEXTRACTF128/VEXTRACTI128
2023-02-08 22:04:22 -08:00
Ryan Houdek 5fd68b6f07 Merge pull request #2392 from lioncash/wide
ARMEmitter: Handle unpredicated wide shifts and unpredicated shifts by immediates
2023-02-08 21:52:23 -08:00
Lioncache 5f80702cf1 ARMEmitter: Handle unpredicated bitwise shift by immediate 2023-02-09 00:15:35 -05:00
Lioncache b45b980b3a ARMEmitter: Handle unpredicated shifts by wide elements 2023-02-09 00:00:17 -05:00
Lioncache f341755e3b Externals: Update fex-gcc-target-test-bins
Allows filtering out the AVX-enabled tests on non-AVX capable systems.
2023-02-08 21:54:35 -05:00
Lioncache c53e7d759b guest_test_runner: Handle AVX-only binary tests
Because the binaries have no metadata, we allow a .json file to be
placed alongside a test indicating required features in a requirements
directory

We also check if the system itself supports those features and run tests
based off of that.
2023-02-08 21:42:04 -05:00
Mai 143ef57141 Merge pull request #2345 from Sonicadvance1/user_sigreturn
Support user supplied signal restorer.
2023-02-08 20:11:47 -05:00
Ryan Houdek 8689038533 Merge pull request #2391 from lioncash/aes
IR: Allow specifying register size for AES enc/dec ops and PCLMUL
2023-02-08 17:11:03 -08:00
Lioncache ade34eeda6 gcc tests: Handle pr57275 test
We now handle all instructions that this uses.
2023-02-08 17:49:29 -05:00
Lioncache 0218c966bd IR: Allow specifying register size for PCLMUL
This will allow us to support 256-bit vector operation in the future.
2023-02-08 16:35:20 -05:00
Lioncache ec5bc9cf3e IR: Allow specifying register sizes for AES enc/dec ops
This will allow us to support operating on 256-bit vectors.

Currently only sets up the bits and pieces on the x86-64 side, since
facilities for testing the 256-bit operations on ARM isn't set up yet.
2023-02-08 16:25:19 -05:00
Lioncache 63bf0d5826 OpcodeDispatcher: Handle VEXTRACTI128 2023-02-08 15:46:27 -05:00
Lioncache 2526fa8b6f OpcodeDispatcher: Handle VEXTRACTF128 2023-02-08 15:40:36 -05:00
Mai ef6f5d2003 Merge pull request #2388 from Sonicadvance1/move_fexbash
FEXBash: Move to Tools folder
2023-02-07 16:18:55 -05:00
Ryan Houdek be02cafb05 FEXBash: Move to Tools folder
Just a cleanup, no functional change.
2023-02-07 07:40:46 -08:00
Ryan Houdek e8fd8ef3b7 Merge pull request #2387 from lioncash/prfx
Arm64/VectorOps: Use movprfx with VBSL
2023-02-06 22:53:41 -08:00
Lioncache d7c6ed842d Arm64/VectorOps: Use movprfx with VBSL
We can use movprfx here to allow compressing the move and bsl operation
together on cpus that can handle it.
2023-02-07 00:44:08 -05:00
Ryan Houdek 86a6118b62 Merge pull request #2386 from lioncash/bsl
VectorOps: Only use VBSL 256-bit path if SVE is present
2023-02-06 21:40:15 -08:00
Lioncache 7a75e43125 VectorOps: Only use VBSL 256-bit path if SVE is present
With this in place, a _VMov isn't necessary for variable blends anymore,
since the vector upper lanes are guaranteed to be zeroed out in the 128-bit case.
2023-02-07 00:13:04 -05:00
Ryan Houdek 4ef3066b69 Merge pull request #2385 from lioncash/vblend
OpcodeDispatcher: Handle VPBLENDVB/VBLENDVPD/VBLENDVPS
2023-02-06 20:25:06 -08:00
Lioncache 88fee019a1 OpcodeDispatcher: Handle VPBLENDVB 2023-02-06 23:04:26 -05:00
Lioncache 5d3141dffc OpcodeDispatcher: Handle VBLENDVPD 2023-02-06 23:04:26 -05:00
Lioncache 94e91565b1 OpcodeDispatcher: Handle VBLENDVPS 2023-02-06 23:04:26 -05:00
Lioncache acbfee55b4 IR: Allow provising register size for VBSL
Necessary, since this will now be used with both 256-bit and 128-bit
registers, rather than just 128-bit.
2023-02-06 23:04:26 -05:00
Lioncache a2481d6892 OpcodeDispatcher: Add helper for AVX variable blends
These will be used by following instruction implementations.
2023-02-06 23:04:26 -05:00
Ryan Houdek e255f1cdef Merge pull request #2383 from lioncash/blend
OpcodeDispatcher: Handle VBLENDPD/VPBLENDW
2023-02-06 18:55:31 -08:00
Ryan Houdek cb3cfed9c2 Merge pull request #2384 from lioncash/sqadd
ARMEmitter: Handle SVE2 saturating add/subtract category
2023-02-06 18:55:23 -08:00
Lioncache 2b12a46d2d ARMEmitter: Handle UQSUBR 2023-02-06 21:29:10 -05:00
Lioncache 8f50109501 ARMEmitter: Handle SQSUBR 2023-02-06 21:29:10 -05:00
Lioncache 1d451b8df1 ARMEmitter: Handle USQADD 2023-02-06 21:29:10 -05:00
Lioncache 49772c6826 ARMEmitter: Handle SUQADD 2023-02-06 21:29:10 -05:00
Lioncache 535a2ab2ba ARMEmitter: Handle UQSUB (vectors, predicated) 2023-02-06 21:29:10 -05:00
Lioncache b8b212719f ARMEmitter: Handle SQSUB (vectors, predicated) 2023-02-06 21:29:10 -05:00
Lioncache be593d43ce ARMEmitter: Handle UQADD (vectors, predicated) 2023-02-06 21:29:10 -05:00
Lioncache a89b7c5dbb ARMEmitter: Handle SQADD (vectors, predicated) 2023-02-06 21:29:07 -05:00
Lioncache c682f51811 OpcodeDispatcher: Handle VPBLENDW 2023-02-06 21:23:30 -05:00
Lioncache 2c7562c54c OpcodeDispatcher: Handle VBLENDPD 2023-02-06 21:03:54 -05:00
Lioncache 8f5ec20cb7 OpcodeDispatcher: Add helper for AVX vector blends 2023-02-06 20:38:35 -05:00
Ryan Houdek 582108a68a Merge pull request #2382 from lioncash/dedup
ARMEmitter: Centralize instruction handling for a few categories
2023-02-06 17:30:09 -08:00
Lioncache 79abe2aa64 ARMEmitter: Simplify bitwise shift by immediate (predicated) category
Centralizes the immediate handling in the encoding helper function.

Lets us move all the asserts there as well.
2023-02-06 20:03:21 -05:00
Lioncache 9f3857b3b0 ARMEmitter: Simplify saturating extract narrow category
Centralizes the immediate handling in the encoding function.
2023-02-06 19:21:42 -05:00
Lioncache 9521638910 ARMEmitter: Simplify bitwise shift right narrow category
Centralizes the immediate handling in one place, making everything much
shorter.
2023-02-06 19:21:39 -05:00
Mai 60b76f53cf Merge pull request #2381 from Sonicadvance1/code_data_header
JIT: Adds a JIT data header and tail.
2023-02-06 18:07:20 -05:00
Mai 5da90aac46 Merge pull request #2378 from Sonicadvance1/fix_emitter_warnings
ARMEmitter: Fixes some warnings that cropped up.
2023-02-06 18:05:56 -05:00
Ryan Houdek ba5ad72ca2 JIT: Adds a JIT data header and tail.
This will be used to store various bits of data about the code going
forward.

Currently unused but that will change as we move forward.
2023-02-06 14:06:37 -08:00
Mai c7c47a827a Merge pull request #2377 from Sonicadvance1/code_data_support
Core: Support Data in JIT buffer header
2023-02-06 16:54:01 -05:00
Mai c4b66b41cd Merge pull request #2379 from Sonicadvance1/rename_fstatat64
Syscalls: Renamed fstatat64 to fstatat_64
2023-02-06 16:47:02 -05:00
Mai 6047ca9fe2 Merge pull request #2376 from Sonicadvance1/minor_flag_opt
Dispatcher: Minor flags optimization
2023-02-06 16:45:45 -05:00
Mai a5762b6faa Merge pull request #2375 from Sonicadvance1/inject_libsegfault
ELFCodeLoader: Adds an option to inject libSegFault
2023-02-06 16:44:52 -05:00
Ryan Houdek 3bc722ca69 Merge pull request #2380 from Joshua-Ashton/directfb_fix
Fix SDL2 directfb includes under Alpine Linux
2023-02-05 18:37:06 -08:00
Joshua Ashton d7d8a4e28a Fix SDL2 directfb includes under Alpine Linux 2023-02-06 02:14:36 +00:00
Ryan Houdek c54c568fef Syscalls: Renamed fstatat64 to fstatat_64
Similar to our other syscall conflicts, musl/Alpine Linux has a global
define that is conflicting with our name here
2023-02-05 18:13:45 -08:00
Ryan Houdek 37421d36e6 ARMEmitter: Fixes some warnings that cropped up. 2023-02-05 18:06:29 -08:00
Ryan Houdek bd86deb9ba Core: Support Data in JIT buffer header
Currently unused (The full data gets thrown away after CompileCode is
called), but allows us to separate code and data in what `CompileCode`
returns.

This will allow us put a header on JIT blocks which will fix a long
outstanding bug where RIP isn't always synchronized on block entry, but
since it only needs to synchronize on signal we can rebuild in the
handler. This future task will remove the `86dec_SynchronizeRIPOnAllBlocks`
config option, but the data will also end up being used for more things
in the future.
2023-02-05 17:55:31 -08:00
Ryan Houdek 2e701fc9e6 Dispatcher: Minor flags optimization
SelectCC shift wasn't necessary since we just need to ensure the final
result is zero when or'd together.

Also operations calculating SF can just use a BFE instead of a shifts
with a constant. BFE by immediate is more efficiently encoded in our IR.
2023-02-04 17:50:36 -08:00
Ryan Houdek 5b97e7f1a0 ELFCodeLoader: Adds an option to inject libSegFault
When used in conjuction with #2345 this is a useful way to enable
libSegFault in applications using application profiles.

Very useful for applications and games that use launcher scripts that
set LD_PRELOAD to nothing prior to launch.

A user was wanting this.
2023-02-04 11:17:26 -08:00
Ryan Houdek 844e27e9ad X86HelperGen: Support fallback sigreturn helpers
For the case that the 32-bit VDSO thunk library isn't available, have a
fallback that can work as well.
Otherwise 32-bit applications will just straight up crash on signal
return.
2023-02-04 10:54:30 -08:00
Ryan Houdek cf147e8ab2 github: Move install step to after the build
Also enable on all builders.
Some tests now rely on 32-bit thunks existing because we need VDSO.
2023-02-04 10:35:07 -08:00
Ryan Houdek e61132b481 VDSOEmu: Handle errors in VDSO
VDSO behaves like a raw syscall which doesn't set errno.
posix tests are testing that errno is set correctly.

Our VDSO handlers weren't wired up to return errors from VDSO correctly.
To handle this we need to have different handlers depending on if the
syscall being used comes from glibc or true VDSO.

This wasn't being uncovered previously since CI wasn't running with VDSO
thunks enabled, but now that it is this needs to be handled or CI will
fail.
2023-02-04 10:35:07 -08:00
Ryan Houdek c58e7a732e X86HelperGen: Remove now unused sigret codegen
This is no longer used so doesn't need to exist.
2023-02-04 10:35:07 -08:00
Ryan Houdek 0538574dd0 Dispatcher: Supports user provided signal restorer
This is required for backtrace to work correctly.
If we are using our custom instruction for returning from a signal, then
backtrace tries to read PC for the sigreturn code and finds our code,
breaking it.

Instead we now /correctly/ support using rt_sigreturn/sigreturn and the
restorer provided from the user.
To facilitate this, we now store a single 64-bit value on the stack to
return our host stack pointer to the correct location from before the
signal.
With cookie checking in place, we can know if an application betrays our
expectations and tries to pass its own signal frames.
If an application in the future /does/ try to pass its own signal
frames, that's unsafe and we cna deal with it then.
2023-02-04 10:35:07 -08:00
Ryan Houdek 1ed546d48f SignalDelegator: Reemit the default signal if it was caught
This fixes a bug where we are falling back down the default signal
delegator after a fatal error.
We need to reraise the event in the case that it didn't come from the
kernel.

Fixes backtrace crashing with incorrect signal when it tries to reraise
the signal that it handled using tgkill.
2023-02-04 10:35:07 -08:00
Ryan Houdek 5ba0053edc VDSOEmulation: Support parsing the 32-bit VDSO symbols
We need to extract the sigreturn handlers and pass them to the FEXCore
signal dispatcher.
2023-02-04 10:35:07 -08:00
Ryan Houdek abb8de0966 VDSO: Add sigreturn functions to VDSO
These need to be bit-exact following exactly what is shown in the
assembly.

libunwind parses where EIP is to see if it is in a stack frame.
Also needsto live in VDSO otherwise backtrace doesn't work.
2023-02-04 10:35:06 -08:00
Ryan Houdek abc596c634 IR: Removes SignalReturn op
This will no longer be used as we are swithing over to using the Linux
system call directly.
2023-02-04 10:35:06 -08:00
Ryan Houdek d107bc9a24 FEXCore: Adds handlers for signal handler returns
Lets the frontend syscall handlers for signal return call the JIT return
handlers directly.
2023-02-04 10:35:06 -08:00
Ryan Houdek a45047bc1e OpDispatcher: Removes SIGRET x86 instruction
We are switching over to syscalls.
2023-02-04 10:35:06 -08:00
Ryan Houdek 1089987a29 Merge pull request #2374 from lioncash/mul
ARMEmitter: Handle SVE SQDMULH/SQRDMULH (vector)
2023-02-04 02:23:59 -08:00
Lioncache 6522d3d6e4 ARMEmitter: Move 128-bit check into SVE2IntegerMultiplyVectors
Simplifies the amount of code needed. Also we can remove some
unnecessary namespacing to make these a little faster to grok when
looking at them.
2023-02-04 05:08:43 -05:00
Lioncache d65fcf7bb8 ARMEmitter: Handle SVE SQRDMULH (vectors) 2023-02-04 05:06:56 -05:00
Lioncache 16e0f628cd ARMEmitter: Handle SVE SQDMULH (vectors) 2023-02-04 05:05:27 -05:00
Ryan Houdek d81097482d Merge pull request #2373 from lioncash/vl
ARMEmitter: Handle ADDVL/ADDPL and RDVL
2023-02-04 01:34:14 -08:00
Lioncache 75bc997ab7 ARMEmitter: Handle RDVL 2023-02-04 01:11:57 -05:00
Lioncache 600e8749d7 ARMEmitter: Handle ADDPL 2023-02-04 01:05:25 -05:00
Lioncache fc3863f444 ARMEmitter: Handle ADDVL 2023-02-04 01:03:24 -05:00
Ryan Houdek 347abf09ef Merge pull request #2372 from lioncash/mla
ARMEmitter: Handle MLA/MLS (vector) and MAD/MSB
2023-02-03 21:09:26 -08:00
Lioncache 7442ef3a83 ARMEmitter: Handle SVE MSB 2023-02-03 23:14:10 -05:00
Lioncache 3783ad8dd1 ARMEmitter: Handle SVE MAD 2023-02-03 23:13:16 -05:00
Lioncache 65ad916984 ARMEmitter: Handle SVE MLS (vectors) 2023-02-03 23:06:42 -05:00
Lioncache d491ce7125 ARMEmitter: Handle SVE MLA (vectors) 2023-02-03 23:05:21 -05:00
Ryan Houdek c0bc5d9748 Merge pull request #2371 from lioncash/mul
ARMEmitter: Handle SVE predicated mul/div and finish off integer reduction category
2023-02-03 19:36:57 -08:00
Lioncache 3f6edf7b5a ARMEmitter: Clarify SVEReductionOperation as working on integer ops 2023-02-03 22:01:34 -05:00
Lioncache 5563a51b84 ARMEmitter: Allow 64-bit variants of min/max reduction
The instructions allow specifying 64-bit element sizes.

With this, we can also completely remove the size checking from the
functions, since the general SVE integer reduction operation already
checks for invalid sizes for us.
2023-02-03 22:00:56 -05:00
Lioncache dac075b871 ARMEmitter: Move min/max reduction over to generic reduction helper
Also enforces the use of a VRegister for the destination argument like
the manual.
2023-02-03 21:45:58 -05:00
Lioncache db8317caf8 ARMEmitter: Handle SVE ANDV (predicated) 2023-02-03 21:31:41 -05:00
Lioncache 3508f7a667 ARMEmitter: Handle SVE EORV (predicated) 2023-02-03 21:30:48 -05:00
Lioncache 02861f41eb ARMEmitter: Handle SVE ORV (predicated) 2023-02-03 21:26:06 -05:00
Lioncache 51c9f70904 ARMEmitter: Handle SVE UADDV (predicated) 2023-02-03 21:05:00 -05:00
Lioncache 4f9530cec3 ARMEmitter: Handle SVE SADDV (predicated) 2023-02-03 21:02:45 -05:00
Lioncache ecd711e691 ARMEmitter: Handle SVE UDIVR (predicated) 2023-02-03 20:49:57 -05:00
Lioncache 870115dd5d ARMEmitter: Handle SVE SDIVR (predicated) 2023-02-03 20:49:57 -05:00
Lioncache 848e5561ce ARMEmitter: Handle SVE UDIV (predicated) 2023-02-03 20:49:57 -05:00
Lioncache db8a9bb5cf ARMEmitter: Handle SVE SDIV (predicated) 2023-02-03 20:49:54 -05:00
Lioncache f8a1c43c06 ARMEmitter: Handle SVE UMULH (predicated) 2023-02-03 20:27:48 -05:00
Lioncache 3a98190119 ARMEmitter: Handle SVE SMULH (predicated) 2023-02-03 20:25:44 -05:00
Lioncache 19ad19193e ARMEmitter: Handle SVE MUL (predicated) 2023-02-03 20:16:24 -05:00
587 changed files with 32423 additions and 19283 deletions

No files matched your search

+14 -8
View File
@@ -14,6 +14,7 @@ env:
CC: clang
CXX: clang++
FEX_FORCE32BITALLOCATOR: 1
FEX_ENABLEAVX: 1
jobs:
build:
@@ -24,7 +25,7 @@ jobs:
fail-fast: false
steps:
- uses: actions/checkout@v2
- uses: actions/checkout@v3
- name: Set runner label
run: echo "runner_label=${{ matrix.arch[1] }}" >> $GITHUB_ENV
@@ -72,6 +73,11 @@ jobs:
# Execute the build. You can specify a specific target with "--target <NAME>"
run: cmake --build . --config $BUILD_TYPE
- name: Install
working-directory: ${{runner.workspace}}/build
shell: bash
run: cmake --build . --config $BUILD_TYPE --target install
- name: ASM Tests
working-directory: ${{runner.workspace}}/build
shell: bash
@@ -199,12 +205,6 @@ jobs:
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_ThunkgenTests.log || true
- name: Install
if: matrix.arch[1] == 'x64'
working-directory: ${{runner.workspace}}/build
shell: bash
run: cmake --build . --config $BUILD_TYPE --target install
- name: Test GL No-Thunks
if: matrix.arch[1] == 'x64'
working-directory: ${{runner.workspace}}/build
@@ -241,13 +241,19 @@ jobs:
# ASM tests get quite close to 10MB
run: truncate --size=<20M ${{runner.workspace}}/build/Testing/Temporary/LastTest_*.log || true
- name: Remove old SHM regions
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: cmake --build . --config $BUILD_TYPE --target remove_old_shm_regions
- name: Set runner name
if: ${{ always() }}
run: echo "runner_name=$(hostname)" >> $GITHUB_ENV
- name: Upload results
if: ${{ always() }}
uses: 'actions/upload-artifact@v2'
uses: 'actions/upload-artifact@v3'
with:
name: Results-${{ env.runner_name }}
path: ${{runner.workspace}}/build/Testing/Temporary/LastTest_*.log
+196
View File
@@ -0,0 +1,196 @@
name: GLIBC fault test
# This workflow file is the same as the `Build + Test` with some key differences
# - Runs on any x86 and ARM64 runner
# - Disables the glibc jemalloc compile option
# - Enables the glibc allocator fault option
# - Disables gvisor tests to reduce stress on CI machines (tmp/shm tests overwhelm them)
# - Disables thunk tests since they are incompatible with glibc fault allocator
# - Disables ARMEmitter tests (We don't want to fault test vixl's disassembler)
on:
push:
branches:
- main
pull_request:
branches:
- main
env:
# Customize the CMake build type here (Release, Debug, RelWithDebInfo, etc.)
BUILD_TYPE: Release
CC: clang
CXX: clang++
FEX_FORCE32BITALLOCATOR: 1
FEX_ENABLEAVX: 1
jobs:
build:
runs-on: ${{ matrix.arch }}
strategy:
matrix:
# Run on an x86 device and any ARM runner.
arch: [[self-hosted, x64], [self-hosted, ARM64]]
fail-fast: false
steps:
- uses: actions/checkout@v3
- name: Set runner label
run: echo "runner_label=${{ matrix.arch[1] }}" >> $GITHUB_ENV
- name: Set rootfs paths
run: |
echo "FEX_ROOTFS_MOUNT=/mnt/AutoNFS/rootfs/" >> $GITHUB_ENV
echo "FEX_ROOTFS_PATH=$HOME/Rootfs/" >> $GITHUB_ENV
echo "FEX_ROOTFS=$HOME/Rootfs/" >> $GITHUB_ENV
echo "ROOTFS=$HOME/Rootfs/" >> $GITHUB_ENV
- name: Update RootFS cache
# Use a bash shell so we can use the same syntax for environment variable
# access regardless of the host operating system
shell: bash
run: $GITHUB_WORKSPACE/Scripts/CI_FetchRootFS.py
- name : submodule checkout
# Need to update submodules
run: |
git submodule sync --recursive
git submodule update --init --depth 1
- name: Clean Build Environment
run: rm -Rf ${{runner.workspace}}/build
- name: Create Build Environment
# Some projects don't allow in-source building, so create a separate build directory
# We'll use this as our working directory for all subsequent commands
run: cmake -E make_directory ${{runner.workspace}}/build
- name: Configure CMake
# Use a bash shell so we can use the same syntax for environment variable
# access regardless of the host operating system
shell: bash
working-directory: ${{runner.workspace}}/build
# Note the current convention is to use the -S and -B options here to specify source
# and build directories, but this is only available with CMake 3.13 and higher.
# The CMake binaries on the Github Actions machines are (as of this writing) 3.12
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True -DENABLE_INTERPRETER=True -DBUILD_FEX_LINUX_TESTS=True -DENABLE_GLIBC_ALLOCATOR_HOOK_FAULT=True -DENABLE_JEMALLOC_GLIBC_ALLOC=False -DCMAKE_INSTALL_PREFIX=${{runner.workspace}}/build/install
- name: Build
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the build. You can specify a specific target with "--target <NAME>"
run: cmake --build . --config $BUILD_TYPE
- name: Install
working-directory: ${{runner.workspace}}/build
shell: bash
run: cmake --build . --config $BUILD_TYPE --target install
- name: ASM Tests
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the unit tests
run: cmake --build . --config $BUILD_TYPE --target asm_tests
- name: ASM Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_ASM.log || true
- name: IR Tests
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the unit tests
run: cmake --build . --config $BUILD_TYPE --target ir_tests
- name: IR Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_IR.log || true
- name: Posix Tests
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the posixtest
run: cmake --build . --config $BUILD_TYPE --target posix_tests
- name: Posix Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_Posix.log || true
- name: gcc target tests 64
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the gvisor tests
run: cmake --build . --config $BUILD_TYPE --target gcc_target_tests_64
- name: GCC64 Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_GCC64.log || true
- name: gcc target tests 32
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the gvisor tests
run: cmake --build . --config $BUILD_TYPE --target gcc_target_tests_32
- name: GCC32 Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_GCC32.log || true
- name: APITest tests
working-directory: ${{runner.workspace}}/build
shell: bash
run: cmake --build . --config $BUILD_TYPE --target api_tests
- name: APITest Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_APITests.log || true
- name: FEXLinuxTests
working-directory: ${{runner.workspace}}/build
shell: bash
run: cmake --build . --config $BUILD_TYPE --target fex_linux_tests_all
- name: FEXLinuxTests Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_FEXLinuxTests.log || true
- name: Truncate test results
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
# Cap out the log files at 20M in case something crash spins and dumps fault text
# ASM tests get quite close to 10MB
run: truncate --size=<20M ${{runner.workspace}}/build/Testing/Temporary/LastTest_*.log || true
- name: Remove old SHM regions
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: cmake --build . --config $BUILD_TYPE --target remove_old_shm_regions
- name: Set runner name
if: ${{ always() }}
run: echo "runner_name=$(hostname)" >> $GITHUB_ENV
- name: Upload results
if: ${{ always() }}
uses: 'actions/upload-artifact@v3'
with:
name: Results-${{ env.runner_name }}
path: ${{runner.workspace}}/build/Testing/Temporary/LastTest_*.log
retention-days: 3
+3 -2
View File
@@ -13,6 +13,7 @@ env:
BUILD_TYPE: Release
CC: clang
CXX: clang++
FEX_ENABLEAVX: 1
jobs:
build:
@@ -24,7 +25,7 @@ jobs:
fail-fast: false
steps:
- uses: actions/checkout@v2
- uses: actions/checkout@v3
- name: Set runner label
run: echo "runner_label=${{ matrix.arch[1] }}" >> $GITHUB_ENV
@@ -110,7 +111,7 @@ jobs:
- name: Upload results
if: ${{ always() }}
uses: 'actions/upload-artifact@v2'
uses: 'actions/upload-artifact@v3'
with:
name: Results-${{ env.runner_name }}
path: ${{runner.workspace}}/build/Testing/Temporary/LastTest_*.log
+5 -2
View File
@@ -3,7 +3,7 @@
path = External/vixl
url = https://github.com/FEX-Emu/vixl.git
[submodule "External/cpp-optparse"]
path = External/cpp-optparse
path = Source/Common/cpp-optparse
url = https://github.com/Sonicadvance1/cpp-optparse
[submodule "External/imgui"]
path = External/imgui
@@ -48,8 +48,11 @@
[submodule "External/robin-map"]
shallow = true
path = External/robin-map
url = https://github.com/Tessil/robin-map.git
url = https://github.com/FEX-Emu/robin-map.git
[submodule "External/Vulkan-Headers"]
shallow = true
path = External/Vulkan-Headers
url = https://github.com/KhronosGroup/Vulkan-Headers.git
[submodule "External/jemalloc_glibc"]
path = External/jemalloc_glibc
url = https://github.com/FEX-Emu/jemalloc.git
+32 -4
View File
@@ -18,10 +18,10 @@ option(ENABLE_ASAN "Enables Clang ASAN" FALSE)
option(ENABLE_TSAN "Enables Clang TSAN" FALSE)
option(ENABLE_ASSERTIONS "Enables assertions in build" FALSE)
option(ENABLE_GDB_SYMBOLS "Enables GDBSymbols integration support" ${HAVE_GDB_JIT_READER_H})
option(ENABLE_VISUAL_DEBUGGER "Enables the visual debugger for compiling" FALSE)
option(ENABLE_STRICT_WERROR "Enables stricter -Werror for CI" FALSE)
option(ENABLE_WERROR "Enables -Werror" FALSE)
option(ENABLE_JEMALLOC "Enables jemalloc allocator" TRUE)
option(ENABLE_JEMALLOC_GLIBC_ALLOC "Enables jemalloc glibc allocator" TRUE)
option(ENABLE_OFFLINE_TELEMETRY "Enables FEX offline telemetry" TRUE)
option(ENABLE_COMPILE_TIME_TRACE "Enables time trace compile option" FALSE)
option(ENABLE_LIBCXX "Enables LLVM libc++" FALSE)
@@ -33,11 +33,19 @@ option(ENABLE_VIXL_DISASSEMBLER "Enables debug disassembler output with VIXL" FA
option(COMPILE_VIXL_DISASSEMBLER "Compiles the vixl disassembler in to vixl" FALSE)
option(ENABLE_FEXCORE_PROFILER "Enables use of the FEXCore timeline profiling capabilities" FALSE)
set (FEXCORE_PROFILER_BACKEND "gpuvis" CACHE STRING "Set which backend you want to use for the FEXCore profiler")
option(ENABLE_GLIBC_ALLOCATOR_HOOK_FAULT "Enables glibc memory allocation hooking with fault for CI testing")
set (X86_32_TOOLCHAIN_FILE "${CMAKE_CURRENT_SOURCE_DIR}/toolchain_x86_32.cmake" CACHE FILEPATH "Toolchain file for the (cross-)compiler targeting i686")
set (X86_64_TOOLCHAIN_FILE "${CMAKE_CURRENT_SOURCE_DIR}/toolchain_x86_64.cmake" CACHE FILEPATH "Toolchain file for the (cross-)compiler targeting x86_64")
set (DATA_DIRECTORY "${CMAKE_INSTALL_PREFIX}/share/fex-emu" CACHE PATH "global data directory")
string(FIND ${CMAKE_BASE_NAME} mingw CONTAINS_MINGW)
if (NOT CONTAINS_MINGW EQUAL -1)
message (STATUS "Mingw build")
set (MINGW_BUILD TRUE)
set (ENABLE_JEMALLOC FALSE)
endif()
if (ENABLE_FEXCORE_PROFILER)
add_definitions(-DENABLE_FEXCORE_PROFILER=1)
string(TOUPPER "${FEXCORE_PROFILER_BACKEND}" FEXCORE_PROFILER_BACKEND)
@@ -49,6 +57,14 @@ if (ENABLE_FEXCORE_PROFILER)
endif()
endif()
if (ENABLE_JEMALLOC_GLIBC_ALLOC AND ENABLE_GLIBC_ALLOCATOR_HOOK_FAULT)
message(FATAL_ERROR "Can't have both glibc fault allocator and jemalloc glibc allocator enabled at the same time")
endif()
if (ENABLE_GLIBC_ALLOCATOR_HOOK_FAULT)
add_definitions(-DGLIBC_ALLOCATOR_FAULT=1)
endif()
# uninstall target
if(NOT TARGET uninstall)
configure_file(
@@ -178,7 +194,22 @@ if (ENABLE_TSAN)
link_libraries(-fno-omit-frame-pointer -fsanitize=thread)
endif()
if (ENABLE_JEMALLOC_GLIBC_ALLOC)
# The glibc jemalloc subproject which hooks the glibc allocator.
# Required for thunks to work.
# All host native libraries will use this allocator, while *most* other FEX internal allocations will use the other jemalloc allocator.
add_definitions(-DENABLE_JEMALLOC_GLIBC=1)
add_subdirectory(External/jemalloc_glibc/)
else()
message (STATUS
" jemalloc glibc allocator disabled!\n"
" This is not a recommended configuration!\n"
" This will very explicitly break thunk execution!\n"
" Use at your own risk!")
endif()
if (ENABLE_JEMALLOC)
# The jemalloc subproject that all FEXCore fextl objects allocate through.
add_definitions(-DENABLE_JEMALLOC=1)
add_subdirectory(External/jemalloc/)
include_directories(External/jemalloc/pregen/include/)
@@ -230,9 +261,6 @@ if (BUILD_TESTS)
include(Catch)
endif()
add_subdirectory(External/cpp-optparse/)
include_directories(External/cpp-optparse/)
add_subdirectory(External/fmt/)
add_subdirectory(External/imgui/)
-5
View File
@@ -1,5 +0,0 @@
{
"Config": {
"x86dec_SynchronizeRIPOnAllBlocks": "1"
}
}
-5
View File
@@ -1,5 +0,0 @@
{
"Config": {
"x86dec_SynchronizeRIPOnAllBlocks": "1"
}
}
+5
View File
@@ -0,0 +1,5 @@
{
"Config": {
"HideHypervisorBit": "1"
}
}
-5
View File
@@ -1,5 +0,0 @@
{
"Config": {
"x86dec_SynchronizeRIPOnAllBlocks": "1"
}
}
-5
View File
@@ -1,5 +0,0 @@
{
"Config": {
"x86dec_SynchronizeRIPOnAllBlocks": "1"
}
}
+8
View File
@@ -173,6 +173,14 @@
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libOpenCL.so.1.0.0"
]
},
"WaylandClient": {
"Library" : "libwayland-client-guest.so",
"Overlay": [
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libwayland-client.so",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libwayland-client.so.0",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libwayland-client.so.0.20.0"
]
},
"":{}
}
}
+1
View File
@@ -77,6 +77,7 @@ configure_file(
include_directories(${CMAKE_BINARY_DIR}/generated)
add_compile_options(-fno-exceptions)
add_subdirectory(Source/)
install (DIRECTORY include/FEXCore ${CMAKE_BINARY_DIR}/include/FEXCore
+10 -7
View File
@@ -22,10 +22,10 @@ def print_header():
#define OPT_UINT64(group, enum, json, default) OPT_BASE(uint64_t, group, enum, json, default)
#endif
#ifndef OPT_STR
#define OPT_STR(group, enum, json, default) OPT_BASE(std::string, group, enum, json, default)
#define OPT_STR(group, enum, json, default) OPT_BASE(fextl::string, group, enum, json, default)
#endif
#ifndef OPT_STRARRAY
#define OPT_STRARRAY(group, enum, json, default) OPT_BASE(std::string, group, enum, json, default)
#define OPT_STRARRAY(group, enum, json, default) OPT_BASE(fextl::string, group, enum, json, default)
#endif
'''
@@ -371,13 +371,16 @@ def print_parse_argloader_options(options):
value_type = op_vals["Type"]
NeedsString = False
conversion_func = "std::to_string"
conversion_func = "fextl::fmt::format(\"{}\", "
if ("ArgumentHandler" in op_vals):
NeedsString = True
conversion_func = "FEXCore::Config::Handler::{0}".format(op_vals["ArgumentHandler"])
conversion_func = "FEXCore::Config::Handler::{0}(".format(op_vals["ArgumentHandler"])
if (value_type == "str"):
NeedsString = True
conversion_func = ""
conversion_func = "("
if (value_type == "bool"):
# boolean values need a decimal specifier. Otherwise fmt prints strings.
conversion_func = "fextl::fmt::format(\"{:d}\", "
if (value_type == "strarray"):
# these need a bit more help
@@ -387,11 +390,11 @@ def print_parse_argloader_options(options):
output_argloader.write("\t}\n")
else:
if (NeedsString):
output_argloader.write("\tstd::string UserValue = Options[\"{0}\"];\n".format(op_key))
output_argloader.write("\tfextl::string UserValue = Options[\"{0}\"];\n".format(op_key))
else:
output_argloader.write("\t{0} UserValue = Options.get(\"{1}\");\n".format(value_type, op_key))
output_argloader.write("\tSet(FEXCore::Config::ConfigOption::CONFIG_{0}, {1}(UserValue));\n".format(op_key.upper(), conversion_func))
output_argloader.write("\tSet(FEXCore::Config::ConfigOption::CONFIG_{0}, {1}UserValue));\n".format(op_key.upper(), conversion_func))
output_argloader.write("}\n")
output_argloader.write("#endif\n")
+51 -20
View File
@@ -1,13 +1,19 @@
set (MAN_DIR ${CMAKE_INSTALL_PREFIX}/share/man CACHE PATH "MAN_DIR")
set (FEXCORE_BASE_SRCS
Common/Paths.cpp
Interface/Config/Config.cpp
Utils/Allocator.cpp
Utils/CPUInfo.cpp
Utils/FileLoading.cpp
Utils/ForcedAssert.cpp
Utils/LogManager.cpp
)
if (NOT MINGW_BUILD)
list(APPEND FEXCORE_BASE_SRCS
Utils/Allocator/64BitAllocator.cpp)
endif()
set (SRCS
Common/JitSymbols.cpp
Common/SoftFloat-3e/extF80_add.c
@@ -99,12 +105,11 @@ set (SRCS
Interface/Core/X86Tables.cpp
Interface/Core/X86DebugInfo.cpp
Interface/Core/X86HelperGen.cpp
Interface/Core/ArchHelpers/Arm64_stubs.cpp
Interface/Core/ArchHelpers/Arm64Emitter.cpp
Interface/Core/Dispatcher/Dispatcher.cpp
Interface/Core/Dispatcher/X86Dispatcher.cpp
Interface/Core/Dispatcher/Arm64Dispatcher.cpp
Interface/Core/Interpreter/InterpreterFallbacks.cpp
Interface/Core/Interpreter/Fallbacks/InterpreterFallbacks.cpp
Interface/Core/X86Tables/BaseTables.cpp
Interface/Core/X86Tables/DDDTables.cpp
Interface/Core/X86Tables/EVEXTables.cpp
@@ -137,14 +142,23 @@ set (SRCS
Interface/IR/Passes/DeadStoreElimination.cpp
Interface/IR/Passes/RegisterAllocationPass.cpp
Interface/IR/Passes/SyscallOptimization.cpp
Utils/Allocator.cpp
Utils/Allocator/64BitAllocator.cpp
Utils/NetStream.cpp
Utils/Telemetry.cpp
Utils/Threads.cpp
Utils/Profiler.cpp
)
if (_M_ARM_64)
list(APPEND SRCS Utils/ArchHelpers/Arm64.cpp)
else()
list(APPEND SRCS Utils/ArchHelpers/Arm64_stubs.cpp)
endif()
if (ENABLE_GLIBC_ALLOCATOR_HOOK_FAULT)
list(APPEND FEXCORE_BASE_SRCS
Utils/AllocatorOverride.cpp)
endif()
if (ENABLE_INTERPRETER)
list(APPEND SRCS
Interface/Core/Interpreter/InterpreterCore.cpp
@@ -162,11 +176,6 @@ if (ENABLE_INTERPRETER)
Interface/Core/Interpreter/VectorOps.cpp)
endif()
if(_M_ARM_64)
list(APPEND SRCS
Interface/Core/ArchHelpers/Arm64.cpp)
endif()
set(DEFINES -DTHREAD_LOCAL=_Thread_local)
if (_M_X86_64)
@@ -222,11 +231,22 @@ if (ENABLE_JIT_ARM64)
)
endif()
set (LIBS fmt::fmt vixl dl xxhash tiny-json FEXHeaderUtils)
set (LIBS fmt::fmt vixl xxhash tiny-json FEXHeaderUtils)
if (NOT MINGW_BUILD)
list (APPEND LIBS dl)
else()
list (APPEND LIBS synchronization)
endif()
if (ENABLE_JEMALLOC)
list (APPEND LIBS FEX_jemalloc)
endif()
if (ENABLE_JEMALLOC_GLIBC_ALLOC)
list (APPEND LIBS FEX_jemalloc_glibc)
endif()
# Generate config
configure_file(
${CMAKE_CURRENT_SOURCE_DIR}/Interface/Config/Config.json.in
@@ -331,7 +351,7 @@ function(AddDefaultOptionsToTarget Name)
target_include_directories(${Name} PUBLIC "${CMAKE_BINARY_DIR}/include/")
target_compile_definitions(${Name} PRIVATE ${DEFINES})
add_dependencies(${Name} CONFIG_INC)
add_dependencies(${Name} CONFIG_INC IR_INC)
target_compile_options(${Name}
PRIVATE
@@ -373,7 +393,6 @@ AddDefaultOptionsToTarget(FEXCore_Base)
function(AddObject Name Type)
add_library(${Name} ${Type} ${SRCS})
add_dependencies(${Name} IR_INC)
target_link_libraries(${Name} FEXCore_Base)
AddDefaultOptionsToTarget(${Name})
@@ -385,6 +404,16 @@ function(AddLibrary Name Type)
add_library(${Name} ${Type} $<TARGET_OBJECTS:${PROJECT_NAME}_object>)
target_link_libraries(${Name} FEXCore_Base)
set_target_properties(${Name} PROPERTIES OUTPUT_NAME FEXCore)
if (MINGW_BUILD)
# Mingw build isn't building a linux shared library, so it can't have a SONAME.
set_target_properties(${Name} PROPERTIES NO_SONAME ON)
# Change the suffixes otherwise cmake continues using .a and .so
if (${Type} STREQUAL SHARED)
set_target_properties(${Name} PROPERTIES SUFFIX ".dll")
elseif(${Type} STREQUAL STATIC)
set_target_properties(${Name} PROPERTIES SUFFIX ".lib")
endif()
endif()
AddDefaultOptionsToTarget(${Name})
endfunction()
@@ -393,10 +422,12 @@ AddObject(${PROJECT_NAME}_object OBJECT)
AddLibrary(${PROJECT_NAME} STATIC)
AddLibrary(${PROJECT_NAME}_shared SHARED)
install(TARGETS ${PROJECT_NAME} ${PROJECT_NAME}_shared
LIBRARY
DESTINATION lib
COMPONENT Libraries
ARCHIVE
DESTINATION lib
COMPONENT Libraries)
if (NOT MINGW_BUILD)
install(TARGETS ${PROJECT_NAME} ${PROJECT_NAME}_shared
LIBRARY
DESTINATION lib
COMPONENT Libraries
ARCHIVE
DESTINATION lib
COMPONENT Libraries)
endif()
+8 -9
View File
@@ -1,11 +1,10 @@
#include <FEXCore/fextl/fmt.h>
#include "Common/JitSymbols.h"
#include <fcntl.h>
#include <string>
#include <unistd.h>
#include <fmt/format.h>
namespace FEXCore {
JITSymbols::JITSymbols() {
}
@@ -18,7 +17,7 @@ namespace FEXCore {
void JITSymbols::InitFile() {
// We can't use FILE here since we must be robust against forking processes closing our FD from under us.
const auto PerfMap = fmt::format("/tmp/perf-{}.map", getpid());
const auto PerfMap = fextl::fmt::format("/tmp/perf-{}.map", getpid());
fd = open(PerfMap.c_str(), O_CREAT | O_TRUNC | O_WRONLY | O_APPEND, 0644);
}
@@ -28,7 +27,7 @@ namespace FEXCore {
// Linux perf format is very straightforward
// `<HostPtr> <Size> <Name>\n`
const auto Buffer = fmt::format("{} {:x} JIT_0x{:x}_{}\n", HostAddr, CodeSize, GuestAddr, HostAddr);
const auto Buffer = fextl::fmt::format("{} {:x} JIT_0x{:x}_{}\n", HostAddr, CodeSize, GuestAddr, HostAddr);
auto Result = write(fd, Buffer.c_str(), Buffer.size());
if (Result == -1 && errno == EBADF) {
fd = -1;
@@ -40,7 +39,7 @@ namespace FEXCore {
// Linux perf format is very straightforward
// `<HostPtr> <Size> <Name>\n`
const auto Buffer = fmt::format("{} {:x} {}_{}\n", HostAddr, CodeSize, Name, HostAddr);
const auto Buffer = fextl::fmt::format("{} {:x} {}_{}\n", HostAddr, CodeSize, Name, HostAddr);
auto Result = write(fd, Buffer.c_str(), Buffer.size());
if (Result == -1 && errno == EBADF) {
fd = -1;
@@ -52,7 +51,7 @@ namespace FEXCore {
// Linux perf format is very straightforward
// `<HostPtr> <Size> <Name>\n`
const auto Buffer = fmt::format("{} {:x} {}+0x{:x} ({})\n", HostAddr, CodeSize, Name, Offset, HostAddr);
const auto Buffer = fextl::fmt::format("{} {:x} {}+0x{:x} ({})\n", HostAddr, CodeSize, Name, Offset, HostAddr);
auto Result = write(fd, Buffer.c_str(), Buffer.size());
if (Result == -1 && errno == EBADF) {
fd = -1;
@@ -64,7 +63,7 @@ namespace FEXCore {
// Linux perf format is very straightforward
// `<HostPtr> <Size> <Name>\n`
const auto Buffer = fmt::format("{} {:x} {}\n", HostAddr, CodeSize, Name);
const auto Buffer = fextl::fmt::format("{} {:x} {}\n", HostAddr, CodeSize, Name);
auto Result = write(fd, Buffer.c_str(), Buffer.size());
if (Result == -1 && errno == EBADF) {
fd = -1;
@@ -76,7 +75,7 @@ namespace FEXCore {
// Linux perf format is very straightforward
// `<HostPtr> <Size> <Name>\n`
const auto Buffer = fmt::format("{} {:x} FEXJIT\n", HostAddr, CodeSize);
const auto Buffer = fextl::fmt::format("{} {:x} FEXJIT\n", HostAddr, CodeSize);
auto Result = write(fd, Buffer.c_str(), Buffer.size());
if (Result == -1 && errno == EBADF) {
fd = -1;
-91
View File
@@ -1,91 +0,0 @@
#include "Common/Paths.h"
#include <FEXCore/Utils/LogManager.h>
#include <cstdlib>
#include <filesystem>
#include <memory>
#include <pwd.h>
#include <system_error>
#include <unistd.h>
namespace FEXCore::Paths {
std::unique_ptr<std::string> CachePath;
std::unique_ptr<std::string> EntryCache;
char const* FindUserHomeThroughUID() {
auto passwd = getpwuid(geteuid());
if (passwd) {
return passwd->pw_dir;
}
return nullptr;
}
const char *GetHomeDirectory() {
char const *HomeDir = getenv("HOME");
// Try to get home directory from uid
if (!HomeDir) {
HomeDir = FindUserHomeThroughUID();
}
// try the PWD
if (!HomeDir) {
HomeDir = getenv("PWD");
}
// Still doesn't exit? You get local
if (!HomeDir) {
HomeDir = ".";
}
return HomeDir;
}
void InitializePaths() {
CachePath = std::make_unique<std::string>();
EntryCache = std::make_unique<std::string>();
char const *HomeDir = getenv("HOME");
if (!HomeDir) {
HomeDir = getenv("PWD");
}
if (!HomeDir) {
HomeDir = ".";
}
char *XDGDataDir = getenv("XDG_DATA_DIR");
if (XDGDataDir) {
*CachePath = XDGDataDir;
}
else {
if (HomeDir) {
*CachePath = HomeDir;
}
}
*CachePath += "/.fex-emu/";
*EntryCache = *CachePath + "/EntryCache/";
std::error_code ec{};
// Ensure the folder structure is created for our Data
if (!std::filesystem::exists(*EntryCache, ec) &&
!std::filesystem::create_directories(*EntryCache, ec)) {
LogMan::Msg::DFmt("Couldn't create EntryCache directory: '{}'", *EntryCache);
}
}
void ShutdownPaths() {
CachePath.reset();
EntryCache.reset();
}
std::string GetCachePath() {
return *CachePath;
}
std::string GetEntryCachePath() {
return *EntryCache;
}
}
-12
View File
@@ -1,12 +0,0 @@
#pragma once
#include <string>
namespace FEXCore::Paths {
void InitializePaths();
void ShutdownPaths();
const char *GetHomeDirectory();
std::string GetCachePath();
std::string GetEntryCachePath();
}
+38 -23
View File
@@ -2,19 +2,19 @@
#include <FEXCore/Utils/BitUtils.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/fextl/sstream.h>
#include <FEXCore/fextl/string.h>
#include <cmath>
#include <cstring>
#include <stdint.h>
#include <string>
#include <sstream>
extern "C" {
#include "SoftFloat-3e/platform.h"
#include "SoftFloat-3e/softfloat.h"
}
struct X80SoftFloat {
struct FEX_PACKED X80SoftFloat {
#ifdef _M_X86_64
// Define this to push some operations to x87
// Only useful to see if precision loss is killing something
@@ -32,22 +32,28 @@ struct X80SoftFloat {
#else
#error No 128bit float for this target!
#endif
struct __attribute__((packed)) {
uint64_t Significand : 64;
uint16_t Exponent : 15;
unsigned Sign : 1;
};
#ifndef _WIN32
#define LIBRARY_PRECISION BIGFLOAT
#else
// Mingw Win32 libraries don't have `__float128` helpers. Needs to use a lower precision.
#define LIBRARY_PRECISION double
#endif
uint64_t Significand : 64;
uint16_t Exponent : 15;
uint16_t Sign : 1;
X80SoftFloat() { memset(this, 0, sizeof(*this)); }
X80SoftFloat(unsigned _Sign, uint16_t _Exponent, uint64_t _Significand)
X80SoftFloat(uint16_t _Sign, uint16_t _Exponent, uint64_t _Significand)
: Significand {_Significand}
, Exponent {_Exponent}
, Sign {_Sign}
{
}
std::string str() const {
std::ostringstream string;
fextl::string str() const {
fextl::ostringstream string;
string << std::hex << Sign;
string << "_" << Exponent;
string << "_" << (Significand >> 63);
@@ -262,7 +268,7 @@ struct X80SoftFloat {
return Result;
#else
X80SoftFloat Int = FRNDINT(rhs, softfloat_round_minMag);
BIGFLOAT Src2_d = Int;
LIBRARY_PRECISION Src2_d = Int;
Src2_d = exp2l(Src2_d);
X80SoftFloat Src2_X80 = Src2_d;
X80SoftFloat Result = extF80_mul(lhs, Src2_X80);
@@ -286,8 +292,8 @@ struct X80SoftFloat {
return Result;
#else
BIGFLOAT Src1_d = lhs;
BIGFLOAT Result = exp2l(Src1_d);
LIBRARY_PRECISION Src1_d = lhs;
LIBRARY_PRECISION Result = exp2l(Src1_d);
Result -= 1.0;
return Result;
#endif
@@ -311,9 +317,9 @@ struct X80SoftFloat {
return Result;
#else
BIGFLOAT Src1_d = lhs;
BIGFLOAT Src2_d = rhs;
BIGFLOAT Tmp = Src2_d * log2l(Src1_d);
LIBRARY_PRECISION Src1_d = lhs;
LIBRARY_PRECISION Src2_d = rhs;
LIBRARY_PRECISION Tmp = Src2_d * log2l(Src1_d);
return Tmp;
#endif
}
@@ -336,9 +342,9 @@ struct X80SoftFloat {
return Result;
#else
BIGFLOAT Src1_d = lhs;
BIGFLOAT Src2_d = rhs;
BIGFLOAT Tmp = atan2l(Src1_d, Src2_d);
LIBRARY_PRECISION Src1_d = lhs;
LIBRARY_PRECISION Src2_d = rhs;
LIBRARY_PRECISION Tmp = atan2l(Src1_d, Src2_d);
return Tmp;
#endif
}
@@ -360,7 +366,7 @@ struct X80SoftFloat {
return Result;
#else
BIGFLOAT Src_d = lhs;
LIBRARY_PRECISION Src_d = lhs;
Src_d = tanl(Src_d);
return Src_d;
#endif
@@ -382,7 +388,7 @@ struct X80SoftFloat {
return Result;
#else
BIGFLOAT Src_d = lhs;
LIBRARY_PRECISION Src_d = lhs;
Src_d = sinl(Src_d);
return Src_d;
#endif
@@ -404,7 +410,7 @@ struct X80SoftFloat {
return Result;
#else
BIGFLOAT Src_d = lhs;
LIBRARY_PRECISION Src_d = lhs;
Src_d = cosl(Src_d);
return Src_d;
#endif
@@ -439,6 +445,7 @@ struct X80SoftFloat {
return FEXCore::BitCast<double>(Result);
}
#ifndef _WIN32
operator BIGFLOAT() const {
#if BIGFLOATSIZE == 16
const float128_t Result = extF80_to_f128(*this);
@@ -449,6 +456,7 @@ struct X80SoftFloat {
return result;
#endif
}
#endif
operator int16_t() const {
auto rv = extF80_to_i32(*this, softfloat_roundingMode, false);
@@ -517,6 +525,7 @@ struct X80SoftFloat {
*this = f64_to_extF80(FEXCore::BitCast<float64_t>(rhs));
}
#ifndef _WIN32
X80SoftFloat(BIGFLOAT rhs) {
#if BIGFLOATSIZE == 16
*this = f128_to_extF80(FEXCore::BitCast<float128_t>(rhs));
@@ -524,6 +533,7 @@ struct X80SoftFloat {
*this = FEXCore::BitCast<long double>(rhs);
#endif
}
#endif
X80SoftFloat(const int16_t rhs) {
*this = i32_to_extF80(rhs);
@@ -562,4 +572,9 @@ private:
static constexpr uint32_t ExponentBias = 16383;
};
#ifndef _WIN32
static_assert(sizeof(X80SoftFloat) == 10, "tword must be 10bytes in size");
#else
// Padding on this extends to 16-bytes rather than 10-bytes on WIN32.
static_assert(sizeof(X80SoftFloat) == 16, "tword must be 16bytes in size");
#endif
+13 -12
View File
@@ -1,48 +1,49 @@
#pragma once
#include <FEXCore/fextl/string.h>
#include <cstdint>
#include <string>
#include <string_view>
#include <optional>
namespace FEXCore::StrConv {
[[maybe_unused]] static bool Conv(std::string_view Value, bool *Result) {
*Result = std::stoi(std::string(Value), nullptr, 0);
*Result = std::strtoull(Value.data(), nullptr, 0);
return true;
}
[[maybe_unused]] static bool Conv(std::string_view Value, uint8_t *Result) {
*Result = std::stoi(std::string(Value), nullptr, 0);
*Result = std::strtoul(Value.data(), nullptr, 0);
return true;
}
[[maybe_unused]] static bool Conv(std::string_view Value, uint16_t *Result) {
*Result = std::stoi(std::string(Value), nullptr, 0);
*Result = std::strtoul(Value.data(), nullptr, 0);
return true;
}
[[maybe_unused]] static bool Conv(std::string_view Value, uint32_t *Result) {
*Result = std::stoi(std::string(Value), nullptr, 0);
*Result = std::strtoul(Value.data(), nullptr, 0);
return true;
}
[[maybe_unused]] static bool Conv(std::string_view Value, int32_t *Result) {
*Result = std::stoi(std::string(Value), nullptr, 0);
*Result = std::strtol(Value.data(), nullptr, 0);
return true;
}
[[maybe_unused]] static bool Conv(std::string_view Value, uint64_t *Result) {
*Result = std::stoull(std::string(Value), nullptr, 0);
return true;
}
[[maybe_unused]] static bool Conv(std::string_view Value, std::string *Result) {
*Result = Value;
*Result = std::strtoull(Value.data(), nullptr, 0);
return true;
}
template <typename T,
typename = std::enable_if<std::is_enum<T>::value, T>>
[[maybe_unused]] static bool Conv(std::string_view Value, T *Result) {
*Result = static_cast<T>(std::stoull(std::string(Value), nullptr, 0));
*Result = static_cast<T>(std::stoull(Value.data(), nullptr, 0));
return true;
}
[[maybe_unused]] static bool Conv(std::string_view Value, fextl::string *Result) {
*Result = Value;
return true;
}
}
+10 -9
View File
@@ -1,11 +1,11 @@
#pragma once
#include <string>
#include <FEXCore/fextl/string.h>
namespace FEXCore::StringUtils {
// Trim the left side of the string of whitespace and new lines
[[maybe_unused]] static std::string LeftTrim(std::string String, std::string TrimTokens = " \t\n\r") {
size_t pos = std::string::npos;
if ((pos = String.find_first_not_of(TrimTokens)) != std::string::npos) {
[[maybe_unused]] static fextl::string LeftTrim(fextl::string String, std::string_view TrimTokens = " \t\n\r") {
size_t pos = fextl::string::npos;
if ((pos = String.find_first_not_of(TrimTokens)) != fextl::string::npos) {
String.erase(0, pos);
}
@@ -13,9 +13,9 @@ namespace FEXCore::StringUtils {
}
// Trim the right side of the string of whitespace and new lines
[[maybe_unused]] static std::string RightTrim(std::string String, std::string TrimTokens = " \t\n\r") {
size_t pos = std::string::npos;
if ((pos = String.find_last_not_of(TrimTokens)) != std::string::npos) {
[[maybe_unused]] static fextl::string RightTrim(fextl::string String, std::string_view TrimTokens = " \t\n\r") {
size_t pos = fextl::string::npos;
if ((pos = String.find_last_not_of(TrimTokens)) != fextl::string::npos) {
String.erase(String.begin() + pos + 1, String.end());
}
@@ -23,7 +23,8 @@ namespace FEXCore::StringUtils {
}
// Trim both the left and right of the string of whitespace and new lines
[[maybe_unused]] static std::string Trim(std::string String, std::string TrimTokens = " \t\n\r") {
return RightTrim(LeftTrim(String, TrimTokens), TrimTokens);
[[maybe_unused]] static fextl::string Trim(fextl::string String, std::string_view TrimTokens = " \t\n\r") {
return RightTrim(LeftTrim(std::move(String), TrimTokens), TrimTokens);
}
}
+122 -155
View File
@@ -1,36 +1,36 @@
#include "Common/StringConv.h"
#include "Common/StringUtils.h"
#include "Common/Paths.h"
#include "Utils/FileLoading.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/Utils/Allocator.h>
#include <FEXCore/Utils/CPUInfo.h>
#include <FEXCore/Utils/FileLoading.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/fextl/fmt.h>
#include <FEXCore/fextl/list.h>
#include <FEXCore/fextl/map.h>
#include <FEXCore/fextl/memory.h>
#include <FEXCore/fextl/string.h>
#include <FEXCore/fextl/unordered_map.h>
#include <FEXCore/fextl/vector.h>
#include <FEXHeaderUtils/Filesystem.h>
#include <array>
#include <assert.h>
#include <cstdlib>
#include <filesystem>
#include <fstream>
#include <functional>
#include <map>
#include <memory>
#include <list>
#include <optional>
#include <stddef.h>
#include <stdint.h>
#include <string>
#include <string_view>
#include <sys/sysinfo.h>
#include <system_error>
#include <type_traits>
#include <unordered_map>
#include <utility>
#include <vector>
#include <tiny-json.h>
namespace FEXCore::Context {
struct Context;
class Context;
}
namespace FEXCore::Config {
@@ -45,13 +45,13 @@ namespace DefaultValues {
namespace JSON {
struct JsonAllocator {
jsonPool_t PoolObject;
std::unique_ptr<std::list<json_t>> json_objects;
fextl::unique_ptr<fextl::list<json_t>> json_objects;
};
static_assert(offsetof(JsonAllocator, PoolObject) == 0, "This needs to be at offset zero");
json_t* PoolInit(jsonPool_t* Pool) {
JsonAllocator* alloc = reinterpret_cast<JsonAllocator*>(Pool);
alloc->json_objects = std::make_unique<std::list<json_t>>();
alloc->json_objects = fextl::make_unique<fextl::list<json_t>>();
return &*alloc->json_objects->emplace(alloc->json_objects->end());
}
@@ -60,8 +60,8 @@ namespace JSON {
return &*alloc->json_objects->emplace(alloc->json_objects->end());
}
static void LoadJSonConfig(const std::string &Config, std::function<void(const char *Name, const char *ConfigSring)> Func) {
std::vector<char> Data;
static void LoadJSonConfig(const fextl::string &Config, std::function<void(const char *Name, const char *ConfigSring)> Func) {
fextl::vector<char> Data;
if (!FEXCore::FileLoading::LoadFile(Data, Config)) {
return;
}
@@ -107,108 +107,75 @@ namespace JSON {
}
}
std::string GetDataDirectory() {
std::string DataDir{};
enum Paths {
PATH_DATA_DIR = 0,
PATH_CONFIG_DIR_LOCAL,
PATH_CONFIG_DIR_GLOBAL,
PATH_CONFIG_FILE_LOCAL,
PATH_CONFIG_FILE_GLOBAL,
PATH_LAST,
};
static std::array<fextl::string, Paths::PATH_LAST> Paths;
char const *HomeDir = Paths::GetHomeDirectory();
char const *DataXDG = getenv("XDG_DATA_HOME");
char const *DataOverride = getenv("FEX_APP_DATA_LOCATION");
if (DataOverride) {
// Data override will override the complete directory
DataDir = DataOverride;
}
else {
DataDir = DataXDG ?: HomeDir;
DataDir += "/.fex-emu/";
}
return DataDir;
void SetDataDirectory(const std::string_view Path) {
Paths[PATH_DATA_DIR] = Path;
}
std::string GetConfigDirectory(bool Global) {
std::string ConfigDir;
if (Global) {
ConfigDir = GLOBAL_DATA_DIRECTORY;
}
else {
char const *HomeDir = Paths::GetHomeDirectory();
char const *ConfigXDG = getenv("XDG_CONFIG_HOME");
char const *ConfigOverride = getenv("FEX_APP_CONFIG_LOCATION");
if (ConfigOverride) {
// Config override completely overrides the config directory
ConfigDir = ConfigOverride;
}
else {
ConfigDir = ConfigXDG ? ConfigXDG : HomeDir;
ConfigDir += "/.fex-emu/";
}
// Ensure the folder structure is created for our configuration
std::error_code ec{};
if (!std::filesystem::exists(ConfigDir, ec) &&
!std::filesystem::create_directories(ConfigDir, ec)) {
// Let's go local in this case
return "./";
}
}
return ConfigDir;
void SetConfigDirectory(const std::string_view Path, bool Global) {
Paths[PATH_CONFIG_DIR_LOCAL + Global] = Path;
}
std::string GetConfigFileLocation(bool Global) {
std::string ConfigFile{};
if (Global) {
ConfigFile = GetConfigDirectory(true) + "Config.json";
}
else {
const char *AppConfig = getenv("FEX_APP_CONFIG");
if (AppConfig) {
// App config environment variable overwrites only the config file
ConfigFile = AppConfig;
}
else {
ConfigFile = GetConfigDirectory(false) + "Config.json";
}
}
return ConfigFile;
void SetConfigFileLocation(const std::string_view Path, bool Global) {
Paths[PATH_CONFIG_FILE_LOCAL + Global] = Path;
}
std::string GetApplicationConfig(const std::string &Filename, bool Global) {
std::string ConfigFile = GetConfigDirectory(Global);
fextl::string const& GetDataDirectory() {
return Paths[PATH_DATA_DIR];
}
fextl::string const& GetConfigDirectory(bool Global) {
return Paths[PATH_CONFIG_DIR_LOCAL + Global];
}
fextl::string const& GetConfigFileLocation(bool Global) {
return Paths[PATH_CONFIG_FILE_LOCAL + Global];
}
fextl::string GetApplicationConfig(const std::string_view Program, bool Global) {
fextl::string ConfigFile = GetConfigDirectory(Global);
std::error_code ec{};
if (!Global &&
!std::filesystem::exists(ConfigFile, ec) &&
!std::filesystem::create_directories(ConfigFile, ec)) {
!FHU::Filesystem::Exists(ConfigFile) &&
!FHU::Filesystem::CreateDirectories(ConfigFile)) {
LogMan::Msg::DFmt("Couldn't create config directory: '{}'", ConfigFile);
// Let's go local in this case
return "./" + Filename + ".json";
return fextl::fmt::format("./{}.json", Program);
}
ConfigFile += "AppConfig/";
// Attempt to create the local folder if it doesn't exist
if (!Global &&
!std::filesystem::exists(ConfigFile, ec) &&
!std::filesystem::create_directories(ConfigFile, ec)) {
!FHU::Filesystem::Exists(ConfigFile) &&
!FHU::Filesystem::CreateDirectories(ConfigFile)) {
// Let's go local in this case
return "./" + Filename + ".json";
return fextl::fmt::format("./{}.json", Program);
}
ConfigFile += Filename + ".json";
return ConfigFile;
return fextl::fmt::format("{}{}.json", ConfigFile, Program);
}
void SetConfig(FEXCore::Context::Context *CTX, ConfigOption Option, uint64_t Config) {
}
void SetConfig(FEXCore::Context::Context *CTX, ConfigOption Option, std::string const &Config) {
void SetConfig(FEXCore::Context::Context *CTX, ConfigOption Option, fextl::string const &Config) {
}
uint64_t GetConfig(FEXCore::Context::Context *CTX, ConfigOption Option) {
return 0;
}
static std::map<FEXCore::Config::LayerType, std::unique_ptr<FEXCore::Config::Layer>> ConfigLayers;
static fextl::map<FEXCore::Config::LayerType, fextl::unique_ptr<FEXCore::Config::Layer>> ConfigLayers;
static FEXCore::Config::Layer *Meta{};
constexpr std::array<FEXCore::Config::LayerType, 9> LoadOrder = {
@@ -268,17 +235,17 @@ namespace JSON {
}
// If an environment variable exists in both current meta and in the incoming layer then the meta layer value is overwritten
std::unordered_map<std::string, std::string> LookupMap;
fextl::unordered_map<fextl::string, fextl::string> LookupMap;
const auto AddToMap = [&LookupMap](FEXCore::Config::LayerValue const &Value) {
for (const auto &EnvVar : Value) {
const auto ItEq = EnvVar.find_first_of('=');
if (ItEq == std::string::npos) {
if (ItEq == fextl::string::npos) {
// Broken environment variable
// Skip
continue;
}
auto Key = std::string(EnvVar.begin(), EnvVar.begin() + ItEq);
auto Value = std::string(EnvVar.begin() + ItEq + 1, EnvVar.end());
auto Key = fextl::string(EnvVar.begin(), EnvVar.begin() + ItEq);
auto Value = fextl::string(EnvVar.begin() + ItEq + 1, EnvVar.end());
// Add the key to the map, overwriting whatever previous value was there
LookupMap.insert_or_assign(std::move(Key), std::move(Value));
@@ -311,7 +278,7 @@ namespace JSON {
}
void Initialize() {
AddLayer(std::make_unique<MetaLayer>(FEXCore::Config::LayerType::LAYER_TOP));
AddLayer(fextl::make_unique<MetaLayer>(FEXCore::Config::LayerType::LAYER_TOP));
Meta = ConfigLayers.begin()->second.get();
}
@@ -329,16 +296,15 @@ namespace JSON {
}
}
std::string ExpandPath(std::string const &ContainerPrefix, std::string PathName) {
fextl::string ExpandPath(fextl::string const &ContainerPrefix, fextl::string PathName) {
if (PathName.empty()) {
return {};
}
std::filesystem::path Path{PathName};
// Expand home if it exists
if (Path.is_relative()) {
std::string Home = getenv("HOME") ?: "";
if (FHU::Filesystem::IsRelative(PathName)) {
fextl::string Home = getenv("HOME") ?: "";
// Home expansion only works if it is the first character
// This matches bash behaviour
if (PathName.at(0) == '~') {
@@ -347,12 +313,15 @@ namespace JSON {
}
// Expand relative path to absolute
Path = std::filesystem::absolute(Path);
char ExistsTempPath[PATH_MAX];
char *RealPath = FHU::Filesystem::Absolute(PathName.c_str(), ExistsTempPath);
if (RealPath) {
PathName = RealPath;
}
// Only return if it exists
std::error_code ec{};
if (std::filesystem::exists(Path, ec)) {
return Path;
if (FHU::Filesystem::Exists(PathName)) {
return PathName;
}
}
else {
@@ -368,9 +337,9 @@ namespace JSON {
// HostThunks: $CMAKE_INSTALL_PREFIX/lib/fex-emu/HostThunks/
// GuestThunks: $CMAKE_INSTALL_PREFIX/share/fex-emu/GuestThunks/
if (!ContainerPrefix.empty() && !PathName.empty()) {
if (!std::filesystem::exists(PathName)) {
if (!FHU::Filesystem::Exists(PathName)) {
auto ContainerPath = ContainerPrefix + PathName;
if (std::filesystem::exists(ContainerPath)) {
if (FHU::Filesystem::Exists(ContainerPath)) {
return ContainerPath;
}
}
@@ -379,15 +348,15 @@ namespace JSON {
return {};
}
constexpr char ContainerManager[] = "/run/host/container-manager";
std::string FindContainer() {
fextl::string FindContainer() {
// We only support pressure-vessel at the moment
const static std::string ContainerManager = "/run/host/container-manager";
if (std::filesystem::exists(ContainerManager)) {
std::vector<char> Manager{};
if (FHU::Filesystem::Exists(ContainerManager)) {
fextl::vector<char> Manager{};
if (FEXCore::FileLoading::LoadFile(Manager, ContainerManager)) {
// Trim the whitespace, may contain a newline
std::string ManagerStr = Manager.data();
fextl::string ManagerStr = Manager.data();
ManagerStr = FEXCore::StringUtils::Trim(ManagerStr);
return ManagerStr;
}
@@ -395,14 +364,13 @@ namespace JSON {
return {};
}
std::string FindContainerPrefix() {
fextl::string FindContainerPrefix() {
// We only support pressure-vessel at the moment
const static std::string ContainerManager = "/run/host/container-manager";
if (std::filesystem::exists(ContainerManager)) {
std::vector<char> Manager{};
if (FHU::Filesystem::Exists(ContainerManager)) {
fextl::vector<char> Manager{};
if (FEXCore::FileLoading::LoadFile(Manager, ContainerManager)) {
// Trim the whitespace, may contain a newline
std::string ManagerStr = Manager.data();
fextl::string ManagerStr = Manager.data();
ManagerStr = FEXCore::StringUtils::Trim(ManagerStr);
if (strncmp(ManagerStr.data(), "pressure-vessel", Manager.size()) == 0) {
// We are running inside of pressure vessel
@@ -424,7 +392,7 @@ namespace JSON {
FEX_CONFIG_OPT(Cores, THREADS);
if (Cores == 0) {
// When the number of emulated CPU cores is zero then auto detect
FEXCore::Config::EraseSet(FEXCore::Config::CONFIG_THREADS, std::to_string(get_nprocs_conf()));
FEXCore::Config::EraseSet(FEXCore::Config::CONFIG_THREADS, fextl::fmt::format("{}", FEXCore::CPUInfo::CalculateNumberOfCPUs()));
}
}
@@ -443,7 +411,7 @@ namespace JSON {
#endif
if (Core > MaxCoreNumber || Core < MinCoreNumber) {
// Sanitize the core option by setting the core to the JIT if invalid
FEXCore::Config::EraseSet(FEXCore::Config::CONFIG_CORE, std::to_string(FEXCore::Config::CONFIG_IRJIT));
FEXCore::Config::EraseSet(FEXCore::Config::CONFIG_CORE, fextl::fmt::format("{}", static_cast<uint32_t>(FEXCore::Config::CONFIG_IRJIT)));
}
}
@@ -457,8 +425,8 @@ namespace JSON {
}
}
std::string ContainerPrefix { FindContainerPrefix() };
auto ExpandPathIfExists = [&ContainerPrefix](FEXCore::Config::ConfigOption Config, std::string PathName) {
fextl::string ContainerPrefix { FindContainerPrefix() };
auto ExpandPathIfExists = [&ContainerPrefix](FEXCore::Config::ConfigOption Config, fextl::string PathName) {
auto NewPath = ExpandPath(ContainerPrefix, PathName);
if (!NewPath.empty()) {
FEXCore::Config::EraseSet(Config, NewPath);
@@ -467,16 +435,15 @@ namespace JSON {
if (FEXCore::Config::Exists(FEXCore::Config::CONFIG_ROOTFS)) {
FEX_CONFIG_OPT(PathName, ROOTFS);
auto ExpandedString = ExpandPath(ContainerPrefix, PathName());
auto ExpandedString = ExpandPath(ContainerPrefix,PathName());
if (!ExpandedString.empty()) {
// Adjust the path if it ended up being relative
FEXCore::Config::EraseSet(FEXCore::Config::CONFIG_ROOTFS, ExpandedString);
}
else if (!PathName().empty()) {
// If the filesystem doesn't exist then let's see if it exists in the fex-emu folder
std::string NamedRootFS = GetDataDirectory() + "RootFS/" + PathName();
std::error_code ec{};
if (std::filesystem::exists(NamedRootFS, ec)) {
fextl::string NamedRootFS = GetDataDirectory() + "RootFS/" + PathName();
if (FHU::Filesystem::Exists(NamedRootFS)) {
FEXCore::Config::EraseSet(FEXCore::Config::CONFIG_ROOTFS, NamedRootFS);
}
}
@@ -498,9 +465,8 @@ namespace JSON {
}
else if (!PathName().empty()) {
// If the filesystem doesn't exist then let's see if it exists in the fex-emu folder
std::string NamedConfig = GetDataDirectory() + "ThunkConfigs/" + PathName();
std::error_code ec{};
if (std::filesystem::exists(NamedConfig, ec)) {
fextl::string NamedConfig = GetDataDirectory() + "ThunkConfigs/" + PathName();
if (FHU::Filesystem::Exists(NamedConfig)) {
FEXCore::Config::EraseSet(FEXCore::Config::CONFIG_THUNKCONFIG, NamedConfig);
}
}
@@ -514,11 +480,11 @@ namespace JSON {
if (FEXCore::Config::Exists(FEXCore::Config::CONFIG_SINGLESTEP)) {
// Single stepping also enforces single instruction size blocks
Set(FEXCore::Config::ConfigOption::CONFIG_MAXINST, std::to_string(1u));
Set(FEXCore::Config::ConfigOption::CONFIG_MAXINST, "1");
}
}
void AddLayer(std::unique_ptr<FEXCore::Config::Layer> _Layer) {
void AddLayer(fextl::unique_ptr<FEXCore::Config::Layer> _Layer) {
ConfigLayers.emplace(_Layer->GetLayerType(), std::move(_Layer));
}
@@ -530,7 +496,7 @@ namespace JSON {
return Meta->All(Option);
}
std::optional<std::string*> Get(ConfigOption Option) {
std::optional<fextl::string*> Get(ConfigOption Option) {
return Meta->Get(Option);
}
@@ -571,7 +537,7 @@ namespace JSON {
}
template<>
std::string Value<std::string>::GetIfExists(FEXCore::Config::ConfigOption Option, std::string Default) {
fextl::string Value<fextl::string>::GetIfExists(FEXCore::Config::ConfigOption Option, fextl::string Default) {
auto Value = FEXCore::Config::Get(Option);
if (Value) {
return **Value;
@@ -582,13 +548,13 @@ namespace JSON {
}
template<>
std::string Value<std::string>::GetIfExists(FEXCore::Config::ConfigOption Option, std::string_view Default) {
fextl::string Value<fextl::string>::GetIfExists(FEXCore::Config::ConfigOption Option, std::string_view Default) {
auto Value = FEXCore::Config::Get(Option);
if (Value) {
return **Value;
}
else {
return std::string(Default);
return fextl::string(Default);
}
}
@@ -603,39 +569,39 @@ namespace JSON {
template uint64_t Value<uint64_t>::GetIfExists(FEXCore::Config::ConfigOption Option, uint64_t Default);
// Constructor
template Value<std::string>::Value(FEXCore::Config::ConfigOption _Option, std::string Default);
template Value<fextl::string>::Value(FEXCore::Config::ConfigOption _Option, fextl::string Default);
template Value<bool>::Value(FEXCore::Config::ConfigOption _Option, bool Default);
template Value<uint8_t>::Value(FEXCore::Config::ConfigOption _Option, uint8_t Default);
template Value<uint64_t>::Value(FEXCore::Config::ConfigOption _Option, uint64_t Default);
template<typename T>
void Value<T>::GetListIfExists(FEXCore::Config::ConfigOption Option, std::list<std::string> *List) {
void Value<T>::GetListIfExists(FEXCore::Config::ConfigOption Option, fextl::list<fextl::string> *List) {
auto Value = FEXCore::Config::All(Option);
List->clear();
if (Value) {
*List = **Value;
}
}
template void Value<std::string>::GetListIfExists(FEXCore::Config::ConfigOption Option, std::list<std::string> *List);
template void Value<fextl::string>::GetListIfExists(FEXCore::Config::ConfigOption Option, fextl::list<fextl::string> *List);
// Application loaders
class MainLoader final : public FEXCore::Config::OptionMapper {
public:
explicit MainLoader(FEXCore::Config::LayerType Type);
explicit MainLoader(std::string ConfigFile);
explicit MainLoader(fextl::string ConfigFile);
void Load() override;
private:
std::string Config;
fextl::string Config;
};
class AppLoader final : public FEXCore::Config::OptionMapper {
public:
explicit AppLoader(const std::string& Filename, FEXCore::Config::LayerType Type);
explicit AppLoader(const fextl::string& Filename, FEXCore::Config::LayerType Type);
void Load();
private:
std::string Config;
fextl::string Config;
};
class EnvLoader final : public FEXCore::Config::Layer {
@@ -647,11 +613,11 @@ namespace JSON {
char *const *envp;
};
static const std::map<std::string, FEXCore::Config::ConfigOption, std::less<>> ConfigLookup = {{
static const fextl::map<fextl::string, FEXCore::Config::ConfigOption, std::less<>> ConfigLookup = {{
#define OPT_BASE(type, group, enum, json, default) {#json, FEXCore::Config::ConfigOption::CONFIG_##enum},
#include <FEXCore/Config/ConfigValues.inl>
}};
static const std::vector<std::pair<const char*, FEXCore::Config::ConfigOption>> EnvConfigLookup = {{
static const fextl::vector<std::pair<const char*, FEXCore::Config::ConfigOption>> EnvConfigLookup = {{
#define OPT_BASE(type, group, enum, json, default) {"FEX_" #enum, FEXCore::Config::ConfigOption::CONFIG_##enum},
#include <FEXCore/Config/ConfigValues.inl>
}};
@@ -672,7 +638,7 @@ namespace JSON {
, Config{FEXCore::Config::GetConfigFileLocation(Type == FEXCore::Config::LayerType::LAYER_GLOBAL_MAIN)} {
}
MainLoader::MainLoader(std::string ConfigFile)
MainLoader::MainLoader(fextl::string ConfigFile)
: FEXCore::Config::OptionMapper(FEXCore::Config::LayerType::LAYER_MAIN)
, Config{std::move(ConfigFile)} {
}
@@ -683,7 +649,7 @@ namespace JSON {
});
}
AppLoader::AppLoader(const std::string& Filename, FEXCore::Config::LayerType Type)
AppLoader::AppLoader(const fextl::string& Filename, FEXCore::Config::LayerType Type)
: FEXCore::Config::OptionMapper(Type) {
const bool Global = Type == FEXCore::Config::LayerType::LAYER_GLOBAL_STEAM_APP ||
Type == FEXCore::Config::LayerType::LAYER_GLOBAL_APP;
@@ -705,12 +671,13 @@ namespace JSON {
}
void EnvLoader::Load() {
std::unordered_map<std::string_view, std::string_view> EnvMap;
using EnvMapType = fextl::unordered_map<std::string_view, std::string_view>;
EnvMapType EnvMap;
for(const char *const *pvar=envp; pvar && *pvar; pvar++) {
std::string_view Var(*pvar);
size_t pos = Var.rfind('=');
if (std::string::npos == pos)
if (fextl::string::npos == pos)
continue;
std::string_view Key = Var.substr(0,pos);
@@ -719,10 +686,10 @@ namespace JSON {
#define ENVLOADER
#include <FEXCore/Config/ConfigOptions.inl>
EnvMap[Key]=Value;
EnvMap[Key] = Value;
}
std::function GetVar = [=](const std::string_view id) -> std::optional<std::string_view> {
auto GetVar = [](EnvMapType &EnvMap, const std::string_view id) -> std::optional<std::string_view> {
if (EnvMap.find(id) != EnvMap.end())
return EnvMap.at(id);
@@ -739,31 +706,31 @@ namespace JSON {
std::optional<std::string_view> Value;
for (auto &it : EnvConfigLookup) {
if ((Value = GetVar(it.first)).has_value()) {
Set(it.second, std::string(*Value));
if ((Value = GetVar(EnvMap, it.first)).has_value()) {
Set(it.second, fextl::string(*Value));
}
}
}
std::unique_ptr<FEXCore::Config::Layer> CreateGlobalMainLayer() {
return std::make_unique<FEXCore::Config::MainLoader>(FEXCore::Config::LayerType::LAYER_GLOBAL_MAIN);
fextl::unique_ptr<FEXCore::Config::Layer> CreateGlobalMainLayer() {
return fextl::make_unique<FEXCore::Config::MainLoader>(FEXCore::Config::LayerType::LAYER_GLOBAL_MAIN);
}
std::unique_ptr<FEXCore::Config::Layer> CreateMainLayer(std::string const *File) {
fextl::unique_ptr<FEXCore::Config::Layer> CreateMainLayer(fextl::string const *File) {
if (File) {
return std::make_unique<FEXCore::Config::MainLoader>(*File);
return fextl::make_unique<FEXCore::Config::MainLoader>(*File);
}
else {
return std::make_unique<FEXCore::Config::MainLoader>(FEXCore::Config::LayerType::LAYER_MAIN);
return fextl::make_unique<FEXCore::Config::MainLoader>(FEXCore::Config::LayerType::LAYER_MAIN);
}
}
std::unique_ptr<FEXCore::Config::Layer> CreateAppLayer(const std::string& Filename, FEXCore::Config::LayerType Type) {
return std::make_unique<FEXCore::Config::AppLoader>(Filename, Type);
fextl::unique_ptr<FEXCore::Config::Layer> CreateAppLayer(const fextl::string& Filename, FEXCore::Config::LayerType Type) {
return fextl::make_unique<FEXCore::Config::AppLoader>(Filename, Type);
}
std::unique_ptr<FEXCore::Config::Layer> CreateEnvironmentLayer(char *const _envp[]) {
return std::make_unique<FEXCore::Config::EnvLoader>(_envp);
fextl::unique_ptr<FEXCore::Config::Layer> CreateEnvironmentLayer(char *const _envp[]) {
return fextl::make_unique<FEXCore::Config::EnvLoader>(_envp);
}
}
+15 -6
View File
@@ -53,7 +53,7 @@
},
"EnableAVX": {
"Type": "bool",
"Default": "true",
"Default": "false",
"Desc": [
"Determines whether or not we use the expanded register file for AVX or not"
]
@@ -240,6 +240,17 @@
"Also needs x86_64-linux-gnu-objdump in PATH.",
"Can be very slow."
]
},
"InjectLibSegFault": {
"Type": "bool",
"Default": "false",
"Desc": [
"Sets the environment variable LD_PRELOAD=libSegFault.so",
"This allows the user to very easily enable libSegFault without dealing with environment variables",
"Very useful for applications that have launch scripts that set the variable to nothing at launch",
"Set this in an application configuration for injecting in to only specific applications.",
"\tNote: If x86/x86_64 libSegFault.so isn't installed then this option won't work."
]
}
},
"Logging": {
@@ -332,14 +343,12 @@
"Useful for a process that keeps restarting and doesn't work"
]
},
"x86dec_SynchronizeRIPOnAllBlocks": {
"HideHypervisorBit": {
"Type": "bool",
"Default": "false",
"Desc": [
"An application that uses try-catch or longjump extensively needs the ability to do context aware state flushing",
"In the case of FEX's block-linking, it won't always ensure that RIP is synchronized.",
"If an exception occurs and RIP isn't synchronized, then FEX's exception stack restore may not long jump as expected",
"Can be useful for Wine applications that rely on stack unwinding"
"Hides the hypervisor CPUID bit when set.",
"Should only be used for applications that have issues with this set."
]
}
},
+36 -178
View File
@@ -1,4 +1,3 @@
#include "Common/Paths.h"
#include "Interface/Context/Context.h"
#include "Interface/Core/Core.h"
#include "Interface/Core/OpcodeDispatcher.h"
@@ -19,221 +18,80 @@ namespace FEXCore::HLE {
namespace FEXCore::Context {
void InitializeStaticTables(OperatingMode Mode) {
FEXCore::Paths::InitializePaths();
X86Tables::InitializeInfoTables(Mode);
IR::InstallOpcodeHandlers(Mode);
}
void ShutdownStaticTables() {
FEXCore::Paths::ShutdownPaths();
fextl::unique_ptr<FEXCore::Context::Context> FEXCore::Context::Context::CreateNewContext() {
return fextl::make_unique<FEXCore::Context::ContextImpl>();
}
FEXCore::Context::Context *CreateNewContext() {
return new FEXCore::Context::Context{};
bool FEXCore::Context::ContextImpl::InitializeContext() {
return FEXCore::CPU::CreateCPUCore(this);
}
bool InitializeContext(FEXCore::Context::Context *CTX) {
return FEXCore::CPU::CreateCPUCore(CTX);
void FEXCore::Context::ContextImpl::SetExitHandler(ExitHandler handler) {
CustomExitHandler = std::move(handler);
}
void DestroyContext(FEXCore::Context::Context *CTX) {
if (CTX->ParentThread) {
CTX->DestroyThread(CTX->ParentThread);
}
delete CTX;
ExitHandler FEXCore::Context::ContextImpl::GetExitHandler() const {
return CustomExitHandler;
}
FEXCore::Core::InternalThreadState* InitCore(FEXCore::Context::Context *CTX, uint64_t InitialRIP, uint64_t StackPointer) {
return CTX->InitCore(InitialRIP, StackPointer);
void FEXCore::Context::ContextImpl::Stop() {
Stop(false);
}
void SetExitHandler(FEXCore::Context::Context *CTX, ExitHandler handler) {
CTX->CustomExitHandler = std::move(handler);
void FEXCore::Context::ContextImpl::CompileRIP(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP) {
CompileBlock(Thread->CurrentFrame, GuestRIP);
}
ExitHandler GetExitHandler(const FEXCore::Context::Context *CTX) {
return CTX->CustomExitHandler;
FEXCore::Context::ExitReason FEXCore::Context::ContextImpl::GetExitReason() {
return ParentThread->ExitReason;
}
void Run(FEXCore::Context::Context *CTX) {
CTX->Run();
bool FEXCore::Context::ContextImpl::IsDone() const {
return IsPaused();
}
void Step(FEXCore::Context::Context *CTX) {
CTX->Step();
void FEXCore::Context::ContextImpl::GetCPUState(FEXCore::Core::CPUState *State) const {
memcpy(State, ParentThread->CurrentFrame, sizeof(FEXCore::Core::CPUState));
}
void CompileRIP(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP) {
Thread->CTX->CompileBlock(Thread->CurrentFrame, GuestRIP);
void FEXCore::Context::ContextImpl::SetCPUState(const FEXCore::Core::CPUState *State) {
memcpy(ParentThread->CurrentFrame, State, sizeof(FEXCore::Core::CPUState));
}
FEXCore::Context::ExitReason RunUntilExit(FEXCore::Context::Context *CTX) {
return CTX->RunUntilExit();
void FEXCore::Context::ContextImpl::SetCustomCPUBackendFactory(CustomCPUFactoryType Factory) {
CustomCPUFactory = std::move(Factory);
}
int GetProgramStatus(const FEXCore::Context::Context *CTX) {
return CTX->GetProgramStatus();
}
FEXCore::Context::ExitReason GetExitReason(const FEXCore::Context::Context *CTX) {
return CTX->ParentThread->ExitReason;
}
bool IsDone(const FEXCore::Context::Context *CTX) {
return CTX->IsPaused();
}
void GetCPUState(const FEXCore::Context::Context *CTX, FEXCore::Core::CPUState *State) {
memcpy(State, CTX->ParentThread->CurrentFrame, sizeof(FEXCore::Core::CPUState));
}
void SetCPUState(FEXCore::Context::Context *CTX, const FEXCore::Core::CPUState *State) {
memcpy(CTX->ParentThread->CurrentFrame, State, sizeof(FEXCore::Core::CPUState));
}
void Pause(FEXCore::Context::Context *CTX) {
CTX->Pause();
}
void Stop(FEXCore::Context::Context *CTX) {
CTX->Stop(false);
}
void SetCustomCPUBackendFactory(FEXCore::Context::Context *CTX, CustomCPUFactoryType Factory) {
CTX->CustomCPUFactory = std::move(Factory);
}
bool AddVirtualMemoryMapping([[maybe_unused]] FEXCore::Context::Context *CTX, [[maybe_unused]] uint64_t VirtualAddress, [[maybe_unused]] uint64_t PhysicalAddress, [[maybe_unused]] uint64_t Size) {
bool FEXCore::Context::ContextImpl::AddVirtualMemoryMapping([[maybe_unused]] uint64_t VirtualAddress, [[maybe_unused]] uint64_t PhysicalAddress, [[maybe_unused]] uint64_t Size) {
return false;
}
void RegisterExternalSyscallVisitor(FEXCore::Context::Context *CTX, [[maybe_unused]] uint64_t Syscall, [[maybe_unused]] FEXCore::HLE::SyscallVisitor *Visitor) {
HostFeatures FEXCore::Context::ContextImpl::GetHostFeatures() const {
return HostFeatures;
}
HostFeatures GetHostFeatures(const FEXCore::Context::Context *CTX) {
return CTX->HostFeatures;
void FEXCore::Context::ContextImpl::SetSignalDelegator(FEXCore::SignalDelegator *_SignalDelegation) {
SignalDelegation = _SignalDelegation;
}
void HandleCallback(FEXCore::Context::Context *CTX, FEXCore::Core::InternalThreadState *Thread, uint64_t RIP) {
CTX->HandleCallback(Thread, RIP);
void FEXCore::Context::ContextImpl::SetSyscallHandler(FEXCore::HLE::SyscallHandler *Handler) {
SyscallHandler = Handler;
SourcecodeResolver = Handler->GetSourcecodeResolver();
}
void RegisterHostSignalHandler(FEXCore::Context::Context *CTX, int Signal, HostSignalDelegatorFunction Func, bool Required) {
CTX->RegisterHostSignalHandler(Signal, std::move(Func), Required);
FEXCore::CPUID::FunctionResults FEXCore::Context::ContextImpl::RunCPUIDFunction(uint32_t Function, uint32_t Leaf) {
return CPUID.RunFunction(Function, Leaf);
}
void RegisterFrontendHostSignalHandler(FEXCore::Context::Context *CTX, int Signal, HostSignalDelegatorFunction Func, bool Required) {
CTX->RegisterFrontendHostSignalHandler(Signal, std::move(Func), Required);
FEXCore::CPUID::XCRResults FEXCore::Context::ContextImpl::RunXCRFunction(uint32_t Function) {
return CPUID.RunXCRFunction(Function);
}
FEXCore::Core::InternalThreadState* CreateThread(FEXCore::Context::Context *CTX, FEXCore::Core::CPUState *NewThreadState, uint64_t ParentTID) {
return CTX->CreateThread(NewThreadState, ParentTID);
FEXCore::CPUID::FunctionResults FEXCore::Context::ContextImpl::RunCPUIDFunctionName(uint32_t Function, uint32_t Leaf, uint32_t CPU) {
return CPUID.RunFunctionName(Function, Leaf, CPU);
}
void ExecutionThread(FEXCore::Context::Context *CTX, FEXCore::Core::InternalThreadState *Thread) {
return CTX->ExecutionThread(Thread);
}
void InitializeThread(FEXCore::Context::Context *CTX, FEXCore::Core::InternalThreadState *Thread) {
return CTX->InitializeThread(Thread);
}
void RunThread(FEXCore::Context::Context *CTX, FEXCore::Core::InternalThreadState *Thread) {
CTX->RunThread(Thread);
}
void StopThread(FEXCore::Context::Context *CTX, FEXCore::Core::InternalThreadState *Thread) {
CTX->StopThread(Thread);
}
void DestroyThread(FEXCore::Context::Context *CTX, FEXCore::Core::InternalThreadState *Thread) {
CTX->DestroyThread(Thread);
}
void CleanupAfterFork(FEXCore::Context::Context *CTX, FEXCore::Core::InternalThreadState *Thread) {
CTX->CleanupAfterFork(Thread);
}
void SetSignalDelegator(FEXCore::Context::Context *CTX, FEXCore::SignalDelegator *SignalDelegation) {
CTX->SignalDelegation = SignalDelegation;
}
void SetSyscallHandler(FEXCore::Context::Context *CTX, FEXCore::HLE::SyscallHandler *Handler) {
CTX->SyscallHandler = Handler;
CTX->SourcecodeResolver = Handler->GetSourcecodeResolver();
}
FEXCore::CPUID::FunctionResults RunCPUIDFunction(FEXCore::Context::Context *CTX, uint32_t Function, uint32_t Leaf) {
return CTX->CPUID.RunFunction(Function, Leaf);
}
FEX_DEFAULT_VISIBILITY FEXCore::CPUID::FunctionResults RunCPUIDFunctionName(FEXCore::Context::Context *CTX, uint32_t Function, uint32_t Leaf, uint32_t CPU) {
return CTX->CPUID.RunFunctionName(Function, Leaf, CPU);
}
void SetAOTIRLoader(FEXCore::Context::Context *CTX, std::function<int(const std::string&)> CacheReader) {
CTX->SetAOTIRLoader(CacheReader);
}
void SetAOTIRWriter(FEXCore::Context::Context *CTX, std::function<std::unique_ptr<std::ofstream>(const std::string&)> CacheWriter) {
CTX->SetAOTIRWriter(CacheWriter);
}
void SetAOTIRRenamer(FEXCore::Context::Context *CTX, std::function<void(const std::string&)> CacheRenamer) {
CTX->SetAOTIRRenamer(CacheRenamer);
}
void FinalizeAOTIRCache(FEXCore::Context::Context *CTX) {
CTX->FinalizeAOTIRCache();
}
void WriteFilesWithCode(FEXCore::Context::Context *CTX, std::function<void(const std::string& fileid, const std::string& filename)> Writer) {
CTX->WriteFilesWithCode(Writer);
}
IR::AOTIRCacheEntry *LoadAOTIRCacheEntry(FEXCore::Context::Context *CTX, const std::string &Name) {
return CTX->LoadAOTIRCacheEntry(Name);
}
void UnloadAOTIRCacheEntry(FEXCore::Context::Context *CTX, IR::AOTIRCacheEntry *Entry) {
return CTX->UnloadAOTIRCacheEntry(Entry);
}
CustomIRResult AddCustomIREntrypoint(FEXCore::Context::Context *CTX, uintptr_t Entrypoint, std::function<void(uintptr_t Entrypoint, FEXCore::IR::IREmitter *)> Handler, void *Creator, void *Data) {
return CTX->AddCustomIREntrypoint(Entrypoint, Handler, Creator, Data);
}
void AppendThunkDefinitions(FEXCore::Context::Context *CTX, std::vector<FEXCore::IR::ThunkDefinition> const& Definitions) {
CTX->AppendThunkDefinitions(Definitions);
}
namespace Debug {
void CompileRIP(FEXCore::Context::Context *CTX, uint64_t RIP) {
CTX->CompileRIP(CTX->ParentThread, RIP);
}
uint64_t GetThreadCount(FEXCore::Context::Context *CTX) {
return CTX->GetThreadCount();
}
FEXCore::Core::RuntimeStats *GetRuntimeStatsForThread(FEXCore::Context::Context *CTX, uint64_t Thread) {
return CTX->GetRuntimeStatsForThread(Thread);
}
bool GetDebugDataForRIP(FEXCore::Context::Context *CTX, uint64_t RIP, FEXCore::Core::DebugData *Data) {
return CTX->GetDebugDataForRIP(RIP, Data);
}
bool FindHostCodeForRIP(FEXCore::Context::Context *CTX, uint64_t RIP, uint8_t **Code) {
return CTX->FindHostCodeForRIP(RIP, Code);
}
// XXX:
// bool FindIRForRIP(FEXCore::Context::Context *CTX, uint64_t RIP, FEXCore::IR::IntrusiveIRList **ir) {
// return CTX->FindIRForRIP(RIP, ir);
// }
// void SetIRForRIP(FEXCore::Context::Context *CTX, uint64_t RIP, FEXCore::IR::IntrusiveIRList *const ir) {
// CTX->SetIRForRIP(RIP, ir);
// }
}
}
+176 -115
View File
@@ -1,7 +1,6 @@
#pragma once
#include "Common/JitSymbols.h"
#include "FEXHeaderUtils/ScopedSignalMask.h"
#include "Interface/Core/CPUID.h"
#include "Interface/Core/X86HelperGen.h"
#include "Interface/Core/ObjectCache/ObjectCacheService.h"
@@ -14,24 +13,25 @@
#include <FEXCore/Core/SignalDelegator.h>
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXCore/Utils/CompilerDefs.h>
#include <FEXCore/Utils/DeferredSignalMutex.h>
#include <FEXCore/Utils/Event.h>
#include <FEXCore/fextl/memory.h>
#include <FEXCore/fextl/set.h>
#include <FEXCore/fextl/string.h>
#include <FEXCore/fextl/unordered_map.h>
#include <FEXCore/fextl/vector.h>
#include <FEXHeaderUtils/Syscalls.h>
#include <stdint.h>
#include <atomic>
#include <condition_variable>
#include <functional>
#include <istream>
#include <map>
#include <memory>
#include <mutex>
#include <shared_mutex>
#include <stddef.h>
#include <string>
#include <unordered_map>
#include <queue>
#include <vector>
namespace FEXCore {
class CodeLoader;
@@ -70,7 +70,128 @@ namespace FEXCore::Context {
MODE_SINGLESTEP = 1,
};
struct Context {
class ContextImpl final : public FEXCore::Context::Context {
public:
// Context base class implementation.
bool InitializeContext() override;
FEXCore::Core::InternalThreadState* InitCore(uint64_t InitialRIP, uint64_t StackPointer) override;
void SetExitHandler(ExitHandler handler) override;
ExitHandler GetExitHandler() const override;
void Pause() override;
void Run() override;
void Stop() override;
void Step() override;
ExitReason RunUntilExit() override;
void CompileRIP(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP) override;
int GetProgramStatus() const override;
ExitReason GetExitReason() override;
bool IsDone() const override;
void GetCPUState(FEXCore::Core::CPUState *State) const override;
void SetCPUState(const FEXCore::Core::CPUState *State) override;
void SetCustomCPUBackendFactory(CustomCPUFactoryType Factory) override;
bool AddVirtualMemoryMapping(uint64_t VirtualAddress, uint64_t PhysicalAddress, uint64_t Size) override;
HostFeatures GetHostFeatures() const override;
void HandleCallback(FEXCore::Core::InternalThreadState *Thread, uint64_t RIP) override;
uint64_t RestoreRIPFromHostPC(FEXCore::Core::InternalThreadState *Thread, uint64_t HostPC) override;
/**
* @brief Used to create FEX thread objects in preparation for creating a true OS thread. Does set a TID or PID.
*
* @param NewThreadState The initial thread state to setup for our state
* @param ParentTID The PID that was the parent thread that created this
*
* @return The InternalThreadState object that tracks all of the emulated thread's state
*
* Usecases:
* OS thread Creation:
* - Thread = CreateThread(NewState, PPID);
* - InitializeThread(Thread);
* OS fork (New thread created with a clone of thread state):
* - clone{2, 3}
* - Thread = CreateThread(CopyOfThreadState, PPID);
* - ExecutionThread(Thread); // Starts executing without creating another host thread
* Thunk callback executing guest code from native host thread
* - Thread = CreateThread(NewState, PPID);
* - InitializeThreadTLSData(Thread);
* - HandleCallback(Thread, RIP);
*/
FEXCore::Core::InternalThreadState* CreateThread(FEXCore::Core::CPUState *NewThreadState, uint64_t ParentTID) override;
// Public for threading
void ExecutionThread(FEXCore::Core::InternalThreadState *Thread) override;
/**
* @brief Initializes the OS thread object and prepares to start executing on that new OS thread
*
* @param Thread The internal FEX thread state object
*
* The OS thread will wait until RunThread is executed
*/
void InitializeThread(FEXCore::Core::InternalThreadState *Thread) override;
/**
* @brief Starts the OS thread object to start executing guest code
*
* @param Thread The internal FEX thread state object
*/
void RunThread(FEXCore::Core::InternalThreadState *Thread) override;
void StopThread(FEXCore::Core::InternalThreadState *Thread) override;
/**
* @brief Destroys this FEX thread object and stops tracking it internally
*
* @param Thread The internal FEX thread state object
*/
void DestroyThread(FEXCore::Core::InternalThreadState *Thread) override;
void CleanupAfterFork(FEXCore::Core::InternalThreadState *Thread) override;
void SetSignalDelegator(FEXCore::SignalDelegator *SignalDelegation) override;
void SetSyscallHandler(FEXCore::HLE::SyscallHandler *Handler) override;
FEXCore::CPUID::FunctionResults RunCPUIDFunction(uint32_t Function, uint32_t Leaf) override;
FEXCore::CPUID::XCRResults RunXCRFunction(uint32_t Function) override;
FEXCore::CPUID::FunctionResults RunCPUIDFunctionName(uint32_t Function, uint32_t Leaf, uint32_t CPU) override;
FEXCore::IR::AOTIRCacheEntry *LoadAOTIRCacheEntry(const fextl::string& Name) override;
void UnloadAOTIRCacheEntry(FEXCore::IR::AOTIRCacheEntry *Entry) override;
void SetAOTIRLoader(std::function<int(const fextl::string&)> CacheReader) override {
IRCaptureCache.SetAOTIRLoader(CacheReader);
}
void SetAOTIRWriter(std::function<fextl::unique_ptr<AOTIRWriter>(const fextl::string&)> CacheWriter) override {
IRCaptureCache.SetAOTIRWriter(CacheWriter);
}
void SetAOTIRRenamer(std::function<void(const fextl::string&)> CacheRenamer) override {
IRCaptureCache.SetAOTIRRenamer(CacheRenamer);
}
void FinalizeAOTIRCache() override {
IRCaptureCache.FinalizeAOTIRCache();
}
void WriteFilesWithCode(std::function<void(const fextl::string& fileid, const fextl::string& filename)> Writer) override {
IRCaptureCache.WriteFilesWithCode(Writer);
}
void InvalidateGuestCodeRange(FEXCore::Core::InternalThreadState *Thread, uint64_t Start, uint64_t Length) override;
void InvalidateGuestCodeRange(FEXCore::Core::InternalThreadState *Thread, uint64_t Start, uint64_t Length, std::function<void(uint64_t start, uint64_t Length)> callback) override;
void MarkMemoryShared() override;
void ConfigureAOTGen(FEXCore::Core::InternalThreadState *Thread, fextl::set<uint64_t> *ExternalBranches, uint64_t SectionMaxAddress) override;
// returns false if a handler was already registered
CustomIRResult AddCustomIREntrypoint(uintptr_t Entrypoint, std::function<void(uintptr_t Entrypoint, FEXCore::IR::IREmitter *)> Handler, void *Creator = nullptr, void *Data = nullptr) override;
void AppendThunkDefinitions(fextl::vector<FEXCore::IR::ThunkDefinition> const& Definitions) override;
public:
friend class FEXCore::HLE::SyscallHandler;
#ifdef JIT_ARM64
friend class FEXCore::CPU::Arm64JITCore;
@@ -116,15 +237,13 @@ namespace FEXCore::Context {
FEX_CONFIG_OPT(ParanoidTSO, PARANOIDTSO);
FEX_CONFIG_OPT(CacheObjectCodeCompilation, CACHEOBJECTCODECOMPILATION);
FEX_CONFIG_OPT(x87ReducedPrecision, X87REDUCEDPRECISION);
FEX_CONFIG_OPT(x86dec_SynchronizeRIPOnAllBlocks, X86DEC_SYNCHRONIZERIPONALLBLOCKS);
FEX_CONFIG_OPT(EnableAVX, ENABLEAVX);
} Config;
FEXCore::HostFeatures HostFeatures;
std::mutex ThreadCreationMutex;
FEXCore::Core::InternalThreadState* ParentThread{};
std::vector<FEXCore::Core::InternalThreadState*> Threads;
fextl::vector<FEXCore::Core::InternalThreadState*> Threads;
std::atomic_bool CoreShuttingDown{false};
bool NeedToCheckXID{true};
@@ -140,48 +259,39 @@ namespace FEXCore::Context {
FEXCore::CPUIDEmu CPUID;
FEXCore::HLE::SyscallHandler *SyscallHandler{};
FEXCore::HLE::SourcecodeResolver *SourcecodeResolver{};
std::unique_ptr<FEXCore::ThunkHandler> ThunkHandler;
std::unique_ptr<FEXCore::CPU::Dispatcher> Dispatcher;
fextl::unique_ptr<FEXCore::ThunkHandler> ThunkHandler;
fextl::unique_ptr<FEXCore::CPU::Dispatcher> Dispatcher;
CustomCPUFactoryType CustomCPUFactory;
FEXCore::Context::ExitHandler CustomExitHandler;
#ifdef BLOCKSTATS
std::unique_ptr<FEXCore::BlockSamplingData> BlockData;
fextl::unique_ptr<FEXCore::BlockSamplingData> BlockData;
#endif
SignalDelegator *SignalDelegation{};
X86GeneratedCode X86CodeGen;
Context();
~Context();
ContextImpl();
~ContextImpl();
FEXCore::Core::InternalThreadState* InitCore(uint64_t InitialRIP, uint64_t StackPointer);
FEXCore::Context::ExitReason RunUntilExit();
int GetProgramStatus() const;
bool IsPaused() const { return !Running; }
void Pause();
void Run();
void WaitForThreadsToRun();
void Step();
void Stop(bool IgnoreCurrentThread);
void WaitForIdle();
void StopThread(FEXCore::Core::InternalThreadState *Thread);
void SignalThread(FEXCore::Core::InternalThreadState *Thread, FEXCore::Core::SignalEvent Event);
bool GetGdbServerStatus() const { return DebugServer != nullptr; }
void StartGdbServer();
void StopGdbServer();
void HandleCallback(FEXCore::Core::InternalThreadState *Thread, uint64_t RIP);
void RegisterHostSignalHandler(int Signal, HostSignalDelegatorFunction Func, bool Required);
void RegisterFrontendHostSignalHandler(int Signal, HostSignalDelegatorFunction Func, bool Required);
static void ThreadRemoveCodeEntry(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP);
static void ThreadAddBlockLink(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestDestination, uintptr_t HostLink, const std::function<void()> &delinker);
template<auto Fn>
static uint64_t ThreadExitFunctionLink(FEXCore::Core::CpuStateFrame *Frame, uint64_t *record) {
FHU::ScopedSignalMaskWithSharedLock lk(Frame->Thread->CTX->CodeInvalidationMutex);
auto Thread = Frame->Thread;
ScopedDeferredSignalWithSharedLock lk(static_cast<ContextImpl*>(Thread->CTX)->CodeInvalidationMutex, Thread);
return Fn(Frame, record);
}
@@ -190,26 +300,16 @@ namespace FEXCore::Context {
// Must be called from owning thread
static void ThreadRemoveCodeEntryFromJit(FEXCore::Core::CpuStateFrame *Frame, uint64_t GuestRIP) {
auto Thread = Frame->Thread;
LogMan::Throw::AFmt(Thread->ThreadManager.GetTID() == FHU::Syscalls::gettid(), "Must be called from owning thread {}, not {}", Thread->ThreadManager.GetTID(), FHU::Syscalls::gettid());
FHU::ScopedSignalMaskWithUniqueLock lk(Thread->CTX->CodeInvalidationMutex);
ScopedDeferredSignalWithUniqueLock lk(static_cast<ContextImpl*>(Thread->CTX)->CodeInvalidationMutex, Thread);
ThreadRemoveCodeEntry(Thread, GuestRIP);
}
// returns false if a handler was already registered
CustomIRResult AddCustomIREntrypoint(uintptr_t Entrypoint, std::function<void(uintptr_t Entrypoint, FEXCore::IR::IREmitter *)> Handler, void *Creator, void *Data);
void RemoveCustomIREntrypoint(uintptr_t Entrypoint);
// Debugger interface
void CompileRIP(FEXCore::Core::InternalThreadState *Thread, uint64_t RIP);
uint64_t GetThreadCount() const;
FEXCore::Core::RuntimeStats *GetRuntimeStatsForThread(uint64_t Thread);
bool GetDebugDataForRIP(uint64_t RIP, FEXCore::Core::DebugData *Data);
bool FindHostCodeForRIP(uint64_t RIP, uint8_t **Code);
struct GenerateIRResult {
FEXCore::IR::IRListView* IRList;
FEXCore::IR::RegisterAllocationData::UniquePtr RAData;
@@ -236,29 +336,6 @@ namespace FEXCore::Context {
void CompileBlockJit(FEXCore::Core::CpuStateFrame *Frame, uint64_t GuestRIP);
// Used for thread creation from syscalls
/**
* @brief Used to create FEX thread objects in preparation for creating a true OS thread. Does set a TID or PID.
*
* @param NewThreadState The initial thread state to setup for our state
* @param ParentTID The PID that was the parent thread that created this
*
* @return The InternalThreadState object that tracks all of the emulated thread's state
*
* Usecases:
* OS thread Creation:
* - Thread = CreateThread(NewState, PPID);
* - InitializeThread(Thread);
* OS fork (New thread created with a clone of thread state):
* - clone{2, 3}
* - Thread = CreateThread(CopyOfThreadState, PPID);
* - ExecutionThread(Thread); // Starts executing without creating another host thread
* Thunk callback executing guest code from native host thread
* - Thread = CreateThread(NewState, PPID);
* - InitializeThreadTLSData(Thread);
* - HandleCallback(Thread, RIP);
*/
FEXCore::Core::InternalThreadState* CreateThread(FEXCore::Core::CPUState *NewThreadState, uint64_t ParentTID);
/**
* @brief Initializes TID, PID and TLS data for a thread
*
@@ -266,76 +343,57 @@ namespace FEXCore::Context {
*/
void InitializeThreadTLSData(FEXCore::Core::InternalThreadState *Thread);
/**
* @brief Initializes the OS thread object and prepares to start executing on that new OS thread
*
* @param Thread The internal FEX thread state object
*
* The OS thread will wait until RunThread is executed
*/
void InitializeThread(FEXCore::Core::InternalThreadState *Thread);
/**
* @brief Starts the OS thread object to start executing guest code
*
* @param Thread The internal FEX thread state object
*/
void RunThread(FEXCore::Core::InternalThreadState *Thread);
/**
* @brief Destroys this FEX thread object and stops tracking it internally
*
* @param Thread The internal FEX thread state object
*/
void DestroyThread(FEXCore::Core::InternalThreadState *Thread);
void CopyMemoryMapping(FEXCore::Core::InternalThreadState *ParentThread, FEXCore::Core::InternalThreadState *ChildThread);
void CleanupAfterFork(FEXCore::Core::InternalThreadState *ExceptForThread);
std::vector<FEXCore::Core::InternalThreadState*>* GetThreads() { return &Threads; }
fextl::vector<FEXCore::Core::InternalThreadState*>* GetThreads() { return &Threads; }
uint8_t GetGPRSize() const { return Config.Is64BitMode ? 8 : 4; }
IR::AOTIRCacheEntry *LoadAOTIRCacheEntry(const std::string &filename);
void UnloadAOTIRCacheEntry(IR::AOTIRCacheEntry *Entry);
FEXCore::JITSymbols Symbols;
// Public for threading
void ExecutionThread(FEXCore::Core::InternalThreadState *Thread);
void GetVDSOSigReturn(VDSOSigReturn *VDSOPointers) override {
if (VDSOPointers->VDSO_kernel_sigreturn == nullptr) {
VDSOPointers->VDSO_kernel_sigreturn = reinterpret_cast<void*>(X86CodeGen.sigreturn_32);
}
void FinalizeAOTIRCache() {
IRCaptureCache.FinalizeAOTIRCache();
if (VDSOPointers->VDSO_kernel_rt_sigreturn == nullptr) {
VDSOPointers->VDSO_kernel_rt_sigreturn = reinterpret_cast<void*>(X86CodeGen.rt_sigreturn_32);
}
}
void WriteFilesWithCode(std::function<void(const std::string& fileid, const std::string& filename)> Writer) {
IRCaptureCache.WriteFilesWithCode(Writer);
void IncrementIdleRefCount() override {
++IdleWaitRefCount;
}
void SetAOTIRLoader(std::function<int(const std::string&)> CacheReader) {
IRCaptureCache.SetAOTIRLoader(CacheReader);
FEXCore::Utils::PooledAllocatorVirtual OpDispatcherAllocator;
FEXCore::Utils::PooledAllocatorVirtual FrontendAllocator;
// If Atomic-based TSO emulation is enabled or not.
bool IsAtomicTSOEnabled() const { return AtomicTSOEmulationEnabled; }
void SetHardwareTSOSupport(bool HardwareTSOSupported) override {
SupportsHardwareTSO = HardwareTSOSupported;
UpdateAtomicTSOEmulationConfig();
}
void SetAOTIRWriter(std::function<std::unique_ptr<std::ofstream>(const std::string&)> CacheWriter) {
IRCaptureCache.SetAOTIRWriter(CacheWriter);
}
void EnableExitOnHLT() override { ExitOnHLT = true; }
void SetAOTIRRenamer(std::function<void(const std::string&)> CacheRenamer) {
IRCaptureCache.SetAOTIRRenamer(CacheRenamer);
}
void AppendThunkDefinitions(std::vector<FEXCore::IR::ThunkDefinition> const& Definitions);
FEXCore::Utils::PooledAllocatorMMap OpDispatcherAllocator;
FEXCore::Utils::PooledAllocatorMMap FrontendAllocator;
void MarkMemoryShared();
bool IsTSOEnabled() { return (IsMemoryShared || !Config.TSOAutoMigration) && Config.TSOEnabled; }
bool ExitOnHLTEnabled() const { return ExitOnHLT; }
protected:
void ClearCodeCache(FEXCore::Core::InternalThreadState *Thread);
void UpdateAtomicTSOEmulationConfig() {
if (SupportsHardwareTSO) {
// If the hardware supports TSO then we don't need to emulate it through atomics.
AtomicTSOEmulationEnabled = false;
}
else {
// Atomic TSO emulation only enabled if the config option is enabled.
AtomicTSOEmulationEnabled = (IsMemoryShared || !Config.TSOAutoMigration) && Config.TSOEnabled;
}
}
private:
/**
* @brief Does some final thread initialization
@@ -363,17 +421,20 @@ namespace FEXCore::Context {
// Entry Cache
std::mutex ExitMutex;
std::unique_ptr<GdbServer> DebugServer;
fextl::unique_ptr<GdbServer> DebugServer;
IR::AOTIRCaptureCache IRCaptureCache;
std::unique_ptr<FEXCore::CodeSerialize::CodeObjectSerializeService> CodeObjectCacheService;
fextl::unique_ptr<FEXCore::CodeSerialize::CodeObjectSerializeService> CodeObjectCacheService;
bool StartPaused = false;
bool IsMemoryShared = false;
bool SupportsHardwareTSO = false;
bool AtomicTSOEmulationEnabled = true;
bool ExitOnHLT = false;
FEX_CONFIG_OPT(AppFilename, APP_FILENAME);
std::shared_mutex CustomIRMutex;
std::unordered_map<uint64_t, std::tuple<std::function<void(uintptr_t Entrypoint, FEXCore::IR::IREmitter *)>, void *, void *>> CustomIRHandlers;
fextl::unordered_map<uint64_t, std::tuple<std::function<void(uintptr_t Entrypoint, FEXCore::IR::IREmitter *)>, void *, void *>> CustomIRHandlers;
FEXCore::CPU::CPUBackendFeatures BackendFeatures;
FEXCore::CPU::DispatcherConfig DispatcherConfig;
};
@@ -1,10 +1,12 @@
#include "Interface/Core/ArchHelpers/Arm64Emitter.h"
#include "FEXCore/Utils/AllocatorHooks.h"
#include "Interface/Core/ArchHelpers/CodeEmitter/Emitter.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Context/Context.h"
#include "Interface/HLE/Thunks/Thunks.h"
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Utils/BitUtils.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/MathUtils.h>
@@ -20,16 +22,36 @@
namespace FEXCore::CPU {
// We want vixl to not allocate a default buffer. Jit and dispatcher will manually create one.
Arm64Emitter::Arm64Emitter(FEXCore::Context::Context *ctx, size_t size)
: Emitter(size ? (uint8_t*)FEXCore::Allocator::mmap(nullptr, size, PROT_READ | PROT_WRITE | PROT_EXEC, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0) : nullptr, size)
Arm64Emitter::Arm64Emitter(FEXCore::Context::ContextImpl *ctx, size_t size)
: Emitter(size ? (uint8_t*)FEXCore::Allocator::VirtualAlloc(size, true) : nullptr, size)
, EmitterCTX {ctx} {
CPU.SetUp();
// Number of register available is dependent on what operating mode the proccess is in.
if (EmitterCTX->Config.Is64BitMode()) {
ConfiguredGPRs = NumGPRs64;
ConfiguredSRAGPRs = NumSRAGPRs64;
ConfiguredGPRPairs = NumGPRPairs64;
ConfiguredFPRs = NumFPRs64;
ConfiguredSRAFPRs = NumSRAFPRs64;
ConfiguredDynamicGPRs = NumGPRs64 - NumGPRs64; // Will be zero, just to be consistent with 32-bit side
ConfiguredDynamicRegisterBase = nullptr;
}
else {
ConfiguredGPRs = NumGPRs32;
ConfiguredSRAGPRs = NumSRAGPRs32;
ConfiguredGPRPairs = NumGPRPairs32;
ConfiguredFPRs = NumFPRs32;
ConfiguredSRAFPRs = NumSRAFPRs32;
ConfiguredDynamicGPRs = NumGPRs32 - NumGPRs64; // Will be 8
ConfiguredDynamicRegisterBase = &RA64[9];
}
}
Arm64Emitter::~Arm64Emitter() {
auto BufferSize = GetBufferSize();
if (BufferSize) {
FEXCore::Allocator::munmap(GetBufferBase(), BufferSize);
FEXCore::Allocator::VirtualFree(GetBufferBase(), BufferSize);
}
}
@@ -111,7 +133,11 @@ void Arm64Emitter::LoadConstant(ARMEmitter::Size s, ARMEmitter::Register Reg, ui
void Arm64Emitter::PushCalleeSavedRegisters() {
// We need to save pairs of registers
// We save r19-r30
const std::array<std::pair<ARMEmitter::XRegister, ARMEmitter::XRegister>, 6> CalleeSaved = {{
const fextl::vector<std::pair<ARMEmitter::XRegister, ARMEmitter::XRegister>> CalleeSaved = {{
#ifdef _WIN32
// Platform register, Just save it twice to make logic easy.
{ARMEmitter::XReg::x18, ARMEmitter::XReg::x18},
#endif
{ARMEmitter::XReg::x19, ARMEmitter::XReg::x20},
{ARMEmitter::XReg::x21, ARMEmitter::XReg::x22},
{ARMEmitter::XReg::x23, ARMEmitter::XReg::x24},
@@ -175,13 +201,17 @@ void Arm64Emitter::PopCalleeSavedRegisters() {
32);
}
const std::array<std::pair<ARMEmitter::XRegister, ARMEmitter::XRegister>, 6> CalleeSaved = {{
const fextl::vector<std::pair<ARMEmitter::XRegister, ARMEmitter::XRegister>> CalleeSaved = {{
{ARMEmitter::XReg::x29, ARMEmitter::XReg::x30},
{ARMEmitter::XReg::x27, ARMEmitter::XReg::x28},
{ARMEmitter::XReg::x25, ARMEmitter::XReg::x26},
{ARMEmitter::XReg::x23, ARMEmitter::XReg::x24},
{ARMEmitter::XReg::x21, ARMEmitter::XReg::x22},
{ARMEmitter::XReg::x19, ARMEmitter::XReg::x20},
#ifdef _WIN32
// Platform register.
{ARMEmitter::XReg::x18, ARMEmitter::XReg::zr},
#endif
}};
for (auto &RegPair : CalleeSaved) {
@@ -189,12 +219,12 @@ void Arm64Emitter::PopCalleeSavedRegisters() {
}
}
void Arm64Emitter::SpillStaticRegs(bool FPRs, uint32_t GPRSpillMask, uint32_t FPRSpillMask) {
void Arm64Emitter::SpillStaticRegs(FEXCore::ARMEmitter::Register TmpReg, bool FPRs, uint32_t GPRSpillMask, uint32_t FPRSpillMask) {
if (!StaticRegisterAllocation()) {
return;
}
for (size_t i = 0; i < SRA64.size(); i+=2) {
for (size_t i = 0; i < ConfiguredSRAGPRs; i+=2) {
auto Reg1 = SRA64[i];
auto Reg2 = SRA64[i+1];
if (((1U << Reg1.Idx()) & GPRSpillMask) &&
@@ -211,22 +241,20 @@ void Arm64Emitter::SpillStaticRegs(bool FPRs, uint32_t GPRSpillMask, uint32_t FP
if (FPRs) {
if (EmitterCTX->HostFeatures.SupportsAVX) {
for (size_t i = 0; i < SRAFPR.size(); i++) {
for (size_t i = 0; i < ConfiguredSRAFPRs; i++) {
const auto Reg = SRAFPR[i];
if (((1U << Reg.Idx()) & FPRSpillMask) != 0) {
mov(ARMEmitter::Size::i64Bit, TMP4.R(), offsetof(Core::CpuStateFrame, State.xmm.avx.data[i][0]));
st1b<ARMEmitter::SubRegSize::i8Bit>(Reg, PRED_TMP_32B, STATE.R(), TMP4.R());
st1b<ARMEmitter::SubRegSize::i8Bit>(Reg.Z(), PRED_TMP_32B, STATE.R(), TMP4.R());
}
}
} else {
if (GPRSpillMask && FPRSpillMask == ~0U) {
// Optimize the common case where we can spill four registers per instruction
auto TmpReg = SRA64[__builtin_ffs(GPRSpillMask)];
// Load the sse offset in to the temporary register
add(ARMEmitter::Size::i64Bit, TmpReg, STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[0][0]));
for (size_t i = 0; i < SRAFPR.size(); i += 4) {
for (size_t i = 0; i < ConfiguredSRAFPRs; i += 4) {
const auto Reg1 = SRAFPR[i];
const auto Reg2 = SRAFPR[i + 1];
const auto Reg3 = SRAFPR[i + 2];
@@ -235,7 +263,7 @@ void Arm64Emitter::SpillStaticRegs(bool FPRs, uint32_t GPRSpillMask, uint32_t FP
}
}
else {
for (size_t i = 0; i < SRAFPR.size(); i += 2) {
for (size_t i = 0; i < ConfiguredSRAFPRs; i += 2) {
const auto Reg1 = SRAFPR[i];
const auto Reg2 = SRAFPR[i + 1];
@@ -269,22 +297,22 @@ void Arm64Emitter::FillStaticRegs(bool FPRs, uint32_t GPRFillMask, uint32_t FPRF
ptrue<ARMEmitter::SubRegSize::i8Bit>(PRED_TMP_16B, ARMEmitter::PredicatePattern::SVE_VL16);
ptrue<ARMEmitter::SubRegSize::i8Bit>(PRED_TMP_32B, ARMEmitter::PredicatePattern::SVE_VL32);
for (size_t i = 0; i < SRAFPR.size(); i++) {
for (size_t i = 0; i < ConfiguredSRAFPRs; i++) {
const auto Reg = SRAFPR[i];
if (((1U << Reg.Idx()) & FPRFillMask) != 0) {
mov(ARMEmitter::Size::i64Bit, TMP4.R(), offsetof(Core::CpuStateFrame, State.xmm.avx.data[i][0]));
ld1b<ARMEmitter::SubRegSize::i8Bit>(Reg, PRED_TMP_32B, STATE.R(), TMP4.R());
ld1b<ARMEmitter::SubRegSize::i8Bit>(Reg.Z(), PRED_TMP_32B.Zeroing(), STATE.R(), TMP4.R());
}
}
} else {
if (GPRFillMask && FPRFillMask == ~0U) {
// Optimize the common case where we can fill four registers per instruction.
// Use one of the filling static registers before we fill it.
auto TmpReg = SRA64[__builtin_ffs(GPRFillMask)];
auto TmpReg = SRA64[FindFirstSetBit(GPRFillMask)];
// Load the sse offset in to the temporary register
add(ARMEmitter::Size::i64Bit, TmpReg, STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[0][0]));
for (size_t i = 0; i < SRAFPR.size(); i += 4) {
for (size_t i = 0; i < ConfiguredSRAFPRs; i += 4) {
const auto Reg1 = SRAFPR[i];
const auto Reg2 = SRAFPR[i + 1];
const auto Reg3 = SRAFPR[i + 2];
@@ -293,7 +321,7 @@ void Arm64Emitter::FillStaticRegs(bool FPRs, uint32_t GPRFillMask, uint32_t FPRF
}
}
else {
for (size_t i = 0; i < SRAFPR.size(); i += 2) {
for (size_t i = 0; i < ConfiguredSRAFPRs; i += 2) {
const auto Reg1 = SRAFPR[i];
const auto Reg2 = SRAFPR[i + 1];
@@ -312,7 +340,7 @@ void Arm64Emitter::FillStaticRegs(bool FPRs, uint32_t GPRFillMask, uint32_t FPRF
}
}
for (size_t i = 0; i < SRA64.size(); i+=2) {
for (size_t i = 0; i < ConfiguredSRAGPRs; i+=2) {
auto Reg1 = SRA64[i];
auto Reg2 = SRA64[i+1];
if (((1U << Reg1.Idx()) & GPRFillMask) &&
@@ -330,10 +358,10 @@ void Arm64Emitter::FillStaticRegs(bool FPRs, uint32_t GPRFillMask, uint32_t FPRF
void Arm64Emitter::PushDynamicRegsAndLR(FEXCore::ARMEmitter::Register TmpReg) {
const auto CanUseSVE = EmitterCTX->HostFeatures.SupportsAVX;
const auto GPRSize = 1 * Core::CPUState::GPR_REG_SIZE;
const auto GPRSize = (ConfiguredDynamicGPRs + 1) * Core::CPUState::GPR_REG_SIZE;
const auto FPRRegSize = CanUseSVE ? Core::CPUState::XMM_AVX_REG_SIZE
: Core::CPUState::XMM_SSE_REG_SIZE;
const auto FPRSize = RAFPR.size() * FPRRegSize;
const auto FPRSize = ConfiguredFPRs * FPRRegSize;
const uint64_t SPOffset = AlignUp(GPRSize + FPRSize, 16);
sub(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, ARMEmitter::Reg::rsp, SPOffset);
@@ -342,17 +370,17 @@ void Arm64Emitter::PushDynamicRegsAndLR(FEXCore::ARMEmitter::Register TmpReg) {
add(ARMEmitter::Size::i64Bit, TmpReg, ARMEmitter::Reg::rsp, 0);
if (CanUseSVE) {
for (size_t i = 0; i < RAFPR.size(); i += 4) {
for (size_t i = 0; i < ConfiguredFPRs; i += 4) {
const auto Reg1 = RAFPR[i];
const auto Reg2 = RAFPR[i + 1];
const auto Reg3 = RAFPR[i + 2];
const auto Reg4 = RAFPR[i + 3];
st4b(Reg1, Reg2, Reg3, Reg4, PRED_TMP_32B, TmpReg, 0);
st4b(Reg1.Z(), Reg2.Z(), Reg3.Z(), Reg4.Z(), PRED_TMP_32B, TmpReg, 0);
add(ARMEmitter::Size::i64Bit, TmpReg, TmpReg, 32 * 4);
}
} else {
static_assert(RAFPR.size() % 4 == 0, "Needs to have multiple of 4 FPRs for RA");
for (size_t i = 0; i < RAFPR.size(); i += 4) {
LOGMAN_THROW_AA_FMT(ConfiguredFPRs % 4 == 0, "Needs to have multiple of 4 FPRs for RA");
for (size_t i = 0; i < ConfiguredFPRs; i += 4) {
const auto Reg1 = RAFPR[i];
const auto Reg2 = RAFPR[i + 1];
const auto Reg3 = RAFPR[i + 2];
@@ -361,6 +389,14 @@ void Arm64Emitter::PushDynamicRegsAndLR(FEXCore::ARMEmitter::Register TmpReg) {
}
}
if (ConfiguredDynamicRegisterBase) {
for (size_t i = 0; i < ConfiguredDynamicGPRs; i += 2) {
const auto Reg1 = ConfiguredDynamicRegisterBase[i];
const auto Reg2 = ConfiguredDynamicRegisterBase[i + 1];
stp<ARMEmitter::IndexType::POST>(Reg1.X(), Reg2.X(), TmpReg, 16);
}
}
str(ARMEmitter::XReg::lr, TmpReg, 0);
}
@@ -368,16 +404,16 @@ void Arm64Emitter::PopDynamicRegsAndLR() {
const auto CanUseSVE = EmitterCTX->HostFeatures.SupportsAVX;
if (CanUseSVE) {
for (size_t i = 0; i < RAFPR.size(); i += 4) {
for (size_t i = 0; i < ConfiguredFPRs; i += 4) {
const auto Reg1 = RAFPR[i];
const auto Reg2 = RAFPR[i + 1];
const auto Reg3 = RAFPR[i + 2];
const auto Reg4 = RAFPR[i + 3];
ld4b(Reg1, Reg2, Reg3, Reg4, PRED_TMP_32B, ARMEmitter::Reg::rsp);
ld4b(Reg1.Z(), Reg2.Z(), Reg3.Z(), Reg4.Z(), PRED_TMP_32B.Zeroing(), ARMEmitter::Reg::rsp);
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, ARMEmitter::Reg::rsp, 32 * 4);
}
} else {
for (size_t i = 0; i < RAFPR.size(); i += 4) {
for (size_t i = 0; i < ConfiguredFPRs; i += 4) {
const auto Reg1 = RAFPR[i];
const auto Reg2 = RAFPR[i + 1];
const auto Reg3 = RAFPR[i + 2];
@@ -386,6 +422,14 @@ void Arm64Emitter::PopDynamicRegsAndLR() {
}
}
if (ConfiguredDynamicRegisterBase) {
for (size_t i = 0; i < ConfiguredDynamicGPRs; i += 2) {
const auto Reg1 = ConfiguredDynamicRegisterBase[i];
const auto Reg2 = ConfiguredDynamicRegisterBase[i + 1];
ldp<ARMEmitter::IndexType::POST>(Reg1.X(), Reg2.X(), ARMEmitter::Reg::rsp, 16);
}
}
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
}
@@ -28,35 +28,83 @@
#include <utility>
namespace FEXCore::CPU {
// All but x29 are caller saved
// Register x18 is unused in the current configuration.
// This is due to it being a platform register on wine platforms.
// TODO: Allow x18 register allocation in the future to gain one more register.
// All but x19 and x29 are caller saved
constexpr std::array<FEXCore::ARMEmitter::Register, 16> SRA64 = {
FEXCore::ARMEmitter::Reg::r4, FEXCore::ARMEmitter::Reg::r5, FEXCore::ARMEmitter::Reg::r6, FEXCore::ARMEmitter::Reg::r7, FEXCore::ARMEmitter::Reg::r8, FEXCore::ARMEmitter::Reg::r9, FEXCore::ARMEmitter::Reg::r10, FEXCore::ARMEmitter::Reg::r11,
FEXCore::ARMEmitter::Reg::r12, FEXCore::ARMEmitter::Reg::r18, FEXCore::ARMEmitter::Reg::r17, FEXCore::ARMEmitter::Reg::r16, FEXCore::ARMEmitter::Reg::r15, FEXCore::ARMEmitter::Reg::r14, FEXCore::ARMEmitter::Reg::r13, FEXCore::ARMEmitter::Reg::r29
FEXCore::ARMEmitter::Reg::r4, FEXCore::ARMEmitter::Reg::r5,
FEXCore::ARMEmitter::Reg::r6, FEXCore::ARMEmitter::Reg::r7,
FEXCore::ARMEmitter::Reg::r8, FEXCore::ARMEmitter::Reg::r9,
FEXCore::ARMEmitter::Reg::r10, FEXCore::ARMEmitter::Reg::r11,
// Registers that don't exist on 32-bit
FEXCore::ARMEmitter::Reg::r12, FEXCore::ARMEmitter::Reg::r13,
FEXCore::ARMEmitter::Reg::r14, FEXCore::ARMEmitter::Reg::r15,
FEXCore::ARMEmitter::Reg::r16, FEXCore::ARMEmitter::Reg::r17,
FEXCore::ARMEmitter::Reg::r19, FEXCore::ARMEmitter::Reg::r29
};
// All are callee saved
constexpr std::array<FEXCore::ARMEmitter::Register, 9> RA64 = {
FEXCore::ARMEmitter::Reg::r20, FEXCore::ARMEmitter::Reg::r21, FEXCore::ARMEmitter::Reg::r22, FEXCore::ARMEmitter::Reg::r23, FEXCore::ARMEmitter::Reg::r24, FEXCore::ARMEmitter::Reg::r25, FEXCore::ARMEmitter::Reg::r26, FEXCore::ARMEmitter::Reg::r27,
FEXCore::ARMEmitter::Reg::r19
constexpr std::array<FEXCore::ARMEmitter::Register, 9 + 8> RA64 = {
// All these callee saved
FEXCore::ARMEmitter::Reg::r20, FEXCore::ARMEmitter::Reg::r21,
FEXCore::ARMEmitter::Reg::r22, FEXCore::ARMEmitter::Reg::r23,
FEXCore::ARMEmitter::Reg::r24, FEXCore::ARMEmitter::Reg::r25,
FEXCore::ARMEmitter::Reg::r26, FEXCore::ARMEmitter::Reg::r27,
FEXCore::ARMEmitter::Reg::r30,
// Registers only available on 32-bit
// All these are caller saved (except for r19).
FEXCore::ARMEmitter::Reg::r12, FEXCore::ARMEmitter::Reg::r13,
FEXCore::ARMEmitter::Reg::r14, FEXCore::ARMEmitter::Reg::r15,
FEXCore::ARMEmitter::Reg::r16, FEXCore::ARMEmitter::Reg::r17,
FEXCore::ARMEmitter::Reg::r19, FEXCore::ARMEmitter::Reg::r29
};
constexpr std::array<std::pair<FEXCore::ARMEmitter::Register, FEXCore::ARMEmitter::Register>, 4> RA64Pair = {{
constexpr std::array<std::pair<FEXCore::ARMEmitter::Register, FEXCore::ARMEmitter::Register>, 4 + 3> RA64Pair = {{
{FEXCore::ARMEmitter::Reg::r20, FEXCore::ARMEmitter::Reg::r21},
{FEXCore::ARMEmitter::Reg::r22, FEXCore::ARMEmitter::Reg::r23},
{FEXCore::ARMEmitter::Reg::r24, FEXCore::ARMEmitter::Reg::r25},
{FEXCore::ARMEmitter::Reg::r26, FEXCore::ARMEmitter::Reg::r27},
// Registers only available on 32-bit
{FEXCore::ARMEmitter::Reg::r12, FEXCore::ARMEmitter::Reg::r13},
{FEXCore::ARMEmitter::Reg::r14, FEXCore::ARMEmitter::Reg::r15},
{FEXCore::ARMEmitter::Reg::r16, FEXCore::ARMEmitter::Reg::r17}
}};
// All are caller saved
constexpr std::array<FEXCore::ARMEmitter::VRegister, 16> SRAFPR = {
FEXCore::ARMEmitter::VReg::v16, FEXCore::ARMEmitter::VReg::v17, FEXCore::ARMEmitter::VReg::v18, FEXCore::ARMEmitter::VReg::v19, FEXCore::ARMEmitter::VReg::v20, FEXCore::ARMEmitter::VReg::v21, FEXCore::ARMEmitter::VReg::v22, FEXCore::ARMEmitter::VReg::v23,
FEXCore::ARMEmitter::VReg::v24, FEXCore::ARMEmitter::VReg::v25, FEXCore::ARMEmitter::VReg::v26, FEXCore::ARMEmitter::VReg::v27, FEXCore::ARMEmitter::VReg::v28, FEXCore::ARMEmitter::VReg::v29, FEXCore::ARMEmitter::VReg::v30, FEXCore::ARMEmitter::VReg::v31
FEXCore::ARMEmitter::VReg::v16, FEXCore::ARMEmitter::VReg::v17,
FEXCore::ARMEmitter::VReg::v18, FEXCore::ARMEmitter::VReg::v19,
FEXCore::ARMEmitter::VReg::v20, FEXCore::ARMEmitter::VReg::v21,
FEXCore::ARMEmitter::VReg::v22, FEXCore::ARMEmitter::VReg::v23,
// Registers that don't exist on 32-bit
FEXCore::ARMEmitter::VReg::v24, FEXCore::ARMEmitter::VReg::v25,
FEXCore::ARMEmitter::VReg::v26, FEXCore::ARMEmitter::VReg::v27,
FEXCore::ARMEmitter::VReg::v28, FEXCore::ARMEmitter::VReg::v29,
FEXCore::ARMEmitter::VReg::v30, FEXCore::ARMEmitter::VReg::v31
};
// v8..v15 = (lower 64bits) Callee saved
constexpr std::array<FEXCore::ARMEmitter::VRegister, 12> RAFPR = {
/*FEXCore::ARMEmitter::VReg::v0, FEXCore::ARMEmitter::VReg::v1, FEXCore::ARMEmitter::VReg::v2, FEXCore::ARMEmitter::VReg::v3,*/FEXCore::ARMEmitter::VReg::v4, FEXCore::ARMEmitter::VReg::v5, FEXCore::ARMEmitter::VReg::v6, FEXCore::ARMEmitter::VReg::v7, // FEXCore::ARMEmitter::VReg::v0 ~ FEXCore::ARMEmitter::VReg::v3 are used as temps
FEXCore::ARMEmitter::VReg::v8, FEXCore::ARMEmitter::VReg::v9, FEXCore::ARMEmitter::VReg::v10, FEXCore::ARMEmitter::VReg::v11, FEXCore::ARMEmitter::VReg::v12, FEXCore::ARMEmitter::VReg::v13, FEXCore::ARMEmitter::VReg::v14, FEXCore::ARMEmitter::VReg::v15
constexpr std::array<FEXCore::ARMEmitter::VRegister, 12 + 8> RAFPR = {
// v0 ~ v3 are used as temps.
// FEXCore::ARMEmitter::VReg::v0, FEXCore::ARMEmitter::VReg::v1,
// FEXCore::ARMEmitter::VReg::v2, FEXCore::ARMEmitter::VReg::v3,
FEXCore::ARMEmitter::VReg::v4, FEXCore::ARMEmitter::VReg::v5,
FEXCore::ARMEmitter::VReg::v6, FEXCore::ARMEmitter::VReg::v7,
FEXCore::ARMEmitter::VReg::v8, FEXCore::ARMEmitter::VReg::v9,
FEXCore::ARMEmitter::VReg::v10, FEXCore::ARMEmitter::VReg::v11,
FEXCore::ARMEmitter::VReg::v12, FEXCore::ARMEmitter::VReg::v13,
FEXCore::ARMEmitter::VReg::v14, FEXCore::ARMEmitter::VReg::v15,
// Registers only available on 32-bit
FEXCore::ARMEmitter::VReg::v24, FEXCore::ARMEmitter::VReg::v25,
FEXCore::ARMEmitter::VReg::v26, FEXCore::ARMEmitter::VReg::v27,
FEXCore::ARMEmitter::VReg::v28, FEXCore::ARMEmitter::VReg::v29,
FEXCore::ARMEmitter::VReg::v30, FEXCore::ARMEmitter::VReg::v31
};
// Contains the address to the currently available CPU state
@@ -85,21 +133,61 @@ constexpr FEXCore::ARMEmitter::PRegister PRED_TMP_32B = FEXCore::ARMEmitter::PRe
// be used by both Arm64 JIT and ARM64 Dispatcher
class Arm64Emitter : public FEXCore::ARMEmitter::Emitter {
protected:
Arm64Emitter(FEXCore::Context::Context *ctx, size_t size);
Arm64Emitter(FEXCore::Context::ContextImpl *ctx, size_t size);
~Arm64Emitter();
FEXCore::Context::Context *EmitterCTX;
FEXCore::Context::ContextImpl *EmitterCTX;
vixl::aarch64::CPU CPU;
uint32_t ConfiguredGPRs;
uint32_t ConfiguredSRAGPRs;
uint32_t ConfiguredGPRPairs;
uint32_t ConfiguredFPRs;
uint32_t ConfiguredSRAFPRs;
uint32_t ConfiguredDynamicGPRs;
const FEXCore::ARMEmitter::Register *ConfiguredDynamicRegisterBase{};
/**
* @name Register Allocation
* @{ */
// 64-bit gets removal of additional pairs
constexpr static uint32_t NumGPRs64 = RA64.size() - 8;
constexpr static uint32_t NumSRAGPRs64 = SRA64.size();
constexpr static uint32_t NumFPRs64 = RAFPR.size() - 8;
constexpr static uint32_t NumSRAFPRs64 = SRAFPR.size();
constexpr static uint32_t NumGPRPairs64 = RA64Pair.size() - 3;
// 32-bit gets full array of GPR registers
// SRA registers remove the additional 8
constexpr static uint32_t NumGPRs32 = RA64.size();
constexpr static uint32_t NumSRAGPRs32 = SRA64.size() - 8;
constexpr static uint32_t NumFPRs32 = RAFPR.size();
constexpr static uint32_t NumSRAFPRs32 = SRAFPR.size() - 8;
constexpr static uint32_t NumGPRPairs32 = RA64Pair.size();
constexpr static uint32_t RegisterClasses = 6;
constexpr static uint64_t GPRBase = (0ULL << 32);
constexpr static uint64_t FPRBase = (1ULL << 32);
constexpr static uint64_t GPRPairBase = (2ULL << 32);
/** @} */
constexpr static uint8_t RA_32 = 0;
constexpr static uint8_t RA_64 = 1;
constexpr static uint8_t RA_FPR = 2;
void LoadConstant(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register Reg, uint64_t Constant, bool NOPPad = false);
// NOTE: These functions WILL clobber the register TMP4 if AVX support is enabled
// and FPRs are being spilled or filled. If only GPRs are spilled/filled, then
// TMP4 is left alone.
void SpillStaticRegs(bool FPRs = true, uint32_t GPRSpillMask = ~0U, uint32_t FPRSpillMask = ~0U);
void SpillStaticRegs(FEXCore::ARMEmitter::Register TmpReg, bool FPRs = true, uint32_t GPRSpillMask = ~0U, uint32_t FPRSpillMask = ~0U);
void FillStaticRegs(bool FPRs = true, uint32_t GPRFillMask = ~0U, uint32_t FPRFillMask = ~0U);
static constexpr uint32_t CALLER_GPR_MASK = 0b0011'1111'1111'1111'1111;
// Register 0-18 + 29 + 30 are caller saved
static constexpr uint32_t CALLER_GPR_MASK = 0b0110'0000'0000'0111'1111'1111'1111'1111U;
// This isn't technically true because the lower 64-bits of v8..v15 are callee saved
// We can't guarantee only the lower 64bits are used so flush everything
@@ -257,8 +257,8 @@ public:
void sxth(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn) {
sbfm(s, rd, rn, 0, 15);
}
void sxtw(FEXCore::ARMEmitter::XRegister rd, FEXCore::ARMEmitter::XRegister rn) {
sbfm(ARMEmitter::Size::i64Bit, rd, rn, 0, 31);
void sxtw(FEXCore::ARMEmitter::XRegister rd, FEXCore::ARMEmitter::WRegister rn) {
sbfm(ARMEmitter::Size::i64Bit, rd, rn.X(), 0, 31);
}
void sbfx(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, uint32_t lsb, uint32_t width) {
LOGMAN_THROW_A_FMT(width > 0, "sbfx needs width > 0");
@@ -287,12 +287,12 @@ public:
void lsl(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, uint32_t shift) {
const auto RegSize = RegSizeInBits(s);
LOGMAN_THROW_A_FMT(shift < RegSize, "Tried to asr a region larger than the register");
LOGMAN_THROW_A_FMT(shift < RegSize, "Tried to lsl a region larger than the register");
ubfm(s, rd, rn, (RegSize - shift) % RegSize, RegSize - shift - 1);
}
void lsr(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, uint32_t shift) {
const auto RegSize = RegSizeInBits(s);
LOGMAN_THROW_A_FMT(shift < RegSize, "Tried to asr a region larger than the register");
LOGMAN_THROW_A_FMT(shift < RegSize, "Tried to lsr a region larger than the register");
ubfm(s, rd, rn, shift, RegSize - 1);
}
void ubfx(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, uint32_t lsb, uint32_t width) {
@@ -303,8 +303,8 @@ public:
void bfi(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, uint32_t lsb, uint32_t width) {
const auto RegSize = RegSizeInBits(s);
LOGMAN_THROW_A_FMT(width > 0, "sbfx needs width > 0");
LOGMAN_THROW_A_FMT((lsb + width) <= RegSize, "Tried to sbfx a region larger than the register");
LOGMAN_THROW_A_FMT(width > 0, "bfi needs width > 0");
LOGMAN_THROW_A_FMT((lsb + width) <= RegSize, "Tried to bfi a region larger than the register");
bfm(s, rd, rn, (RegSize - lsb) & (RegSize - 1), width - 1);
}
@@ -316,7 +316,6 @@ public:
}
void ror(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, uint32_t Imm) {
LOGMAN_THROW_A_FMT(Imm < RegSizeInBits(s), "Tried to extr a region larger than the register");
extr(s, rd, rn, rn, Imm);
}
@@ -711,28 +710,28 @@ public:
DataProcessing_3Source(Op, 0, s, rd, rn, rm, ra);
}
void mul(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, FEXCore::ARMEmitter::Register rm) {
madd(s, rd, rn, rm, FEXCore::ARMEmitter::Reg::zr);
madd(s, rd, rn, rm, XReg::zr);
}
void msub(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, FEXCore::ARMEmitter::Register rm, FEXCore::ARMEmitter::Register ra) {
constexpr uint32_t Op = 0b001'1011'000U << 21;
DataProcessing_3Source(Op, 1, s, rd, rn, rm, ra);
}
void mneg(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::Register rn, FEXCore::ARMEmitter::Register rm) {
msub(s, rd, rn, rm, FEXCore::ARMEmitter::Reg::zr);
msub(s, rd, rn, rm, XReg::zr);
}
void smaddl(FEXCore::ARMEmitter::XRegister rd, FEXCore::ARMEmitter::WRegister rn, FEXCore::ARMEmitter::WRegister rm, FEXCore::ARMEmitter::XRegister ra) {
constexpr uint32_t Op = 0b001'1011'001U << 21;
DataProcessing_3Source(Op, 0, FEXCore::ARMEmitter::Size::i64Bit, rd, rn, rm, ra);
}
void smull(FEXCore::ARMEmitter::XRegister rd, FEXCore::ARMEmitter::WRegister rn, FEXCore::ARMEmitter::WRegister rm) {
smaddl(rd, rn, rm, FEXCore::ARMEmitter::Reg::zr);
smaddl(rd, rn, rm, XReg::zr);
}
void smsubl(FEXCore::ARMEmitter::XRegister rd, FEXCore::ARMEmitter::WRegister rn, FEXCore::ARMEmitter::WRegister rm, FEXCore::ARMEmitter::XRegister ra) {
constexpr uint32_t Op = 0b001'1011'001U << 21;
DataProcessing_3Source(Op, 1, FEXCore::ARMEmitter::Size::i64Bit, rd, rn, rm, ra);
}
void smnegl(FEXCore::ARMEmitter::XRegister rd, FEXCore::ARMEmitter::WRegister rn, FEXCore::ARMEmitter::WRegister rm) {
smsubl(rd, rn, rm, FEXCore::ARMEmitter::Reg::zr);
smsubl(rd, rn, rm, XReg::zr);
}
void smulh(FEXCore::ARMEmitter::XRegister rd, FEXCore::ARMEmitter::XRegister rn, FEXCore::ARMEmitter::XRegister rm) {
constexpr uint32_t Op = 0b001'1011'010U << 21;
@@ -743,14 +742,14 @@ public:
DataProcessing_3Source(Op, 0, FEXCore::ARMEmitter::Size::i64Bit, rd, rn, rm, ra);
}
void umull(FEXCore::ARMEmitter::XRegister rd, FEXCore::ARMEmitter::WRegister rn, FEXCore::ARMEmitter::WRegister rm) {
umaddl(rd, rn, rm, FEXCore::ARMEmitter::Reg::zr);
umaddl(rd, rn, rm, XReg::zr);
}
void umsubl(FEXCore::ARMEmitter::XRegister rd, FEXCore::ARMEmitter::WRegister rn, FEXCore::ARMEmitter::WRegister rm, FEXCore::ARMEmitter::XRegister ra) {
constexpr uint32_t Op = 0b001'1011'101U << 21;
DataProcessing_3Source(Op, 1, FEXCore::ARMEmitter::Size::i64Bit, rd, rn, rm, ra);
}
void umnegl(FEXCore::ARMEmitter::XRegister rd, FEXCore::ARMEmitter::WRegister rn, FEXCore::ARMEmitter::WRegister rm) {
umsubl(rd, rn, rm, FEXCore::ARMEmitter::Reg::zr);
umsubl(rd, rn, rm, XReg::zr);
}
void umulh(FEXCore::ARMEmitter::XRegister rd, FEXCore::ARMEmitter::XRegister rn, FEXCore::ARMEmitter::XRegister rm) {
constexpr uint32_t Op = 0b001'1011'110U << 21;
File diff suppressed because it is too large. Load diff
@@ -6,13 +6,13 @@
#include <FEXCore/Utils/CompilerDefs.h>
#include <FEXCore/Utils/EnumUtils.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/fextl/vector.h>
#include <aarch64/assembler-aarch64.h>
#include <cstdint>
#include <utility>
#include <type_traits>
#include <vector>
/*
* Welcome to FEX-Emu's custom AArch64 emitter.
@@ -62,20 +62,9 @@ namespace FEXCore::ARMEmitter {
};
// This allows us to get the `Size` enum in bits.
template<Size size>
constexpr size_t RegSizeInBits() {
constexpr size_t RegSize[] = {
32, 64, 128,
};
return RegSize[FEXCore::ToUnderlying(size)];
}
[[maybe_unused]]
static inline size_t RegSizeInBits(Size size) {
constexpr size_t RegSize[] = {
32, 64, 128,
};
return RegSize[FEXCore::ToUnderlying(size)];
[[nodiscard]]
constexpr size_t RegSizeInBits(Size size) {
return size_t{32} << FEXCore::ToUnderlying(size);
}
/* This `SubRegSize` enum is used for most ASIMD operations.
@@ -90,14 +79,9 @@ namespace FEXCore::ARMEmitter {
};
// This allows us to get the `SubRegSize` in bits.
template<SubRegSize size>
constexpr size_t SubRegSizeInBits() {
return (1 << FEXCore::ToUnderlying(size)) * 8;
}
[[maybe_unused]]
static inline size_t SubRegSizeInBits(SubRegSize size) {
return (1 << FEXCore::ToUnderlying(size)) * 8;
[[nodiscard]]
constexpr size_t SubRegSizeInBits(SubRegSize size) {
return size_t{8} << FEXCore::ToUnderlying(size);
}
/* This `ScalarRegSize` enum is used for most scalar float
@@ -117,14 +101,9 @@ namespace FEXCore::ARMEmitter {
};
// This allows us to get the `ScalarRegSize` in bits.
template<ScalarRegSize size>
constexpr size_t ScalarRegSizeInBits() {
return (1 << FEXCore::ToUnderlying(size)) * 8;
}
[[maybe_unused]]
static inline size_t ScalarRegSizeInBits(ScalarRegSize size) {
return (1 << FEXCore::ToUnderlying(size)) * 8;
[[nodiscard]]
constexpr size_t ScalarRegSizeInBits(ScalarRegSize size) {
return size_t{8} << FEXCore::ToUnderlying(size);
}
/* This `VectorRegSizePair` union allows us to have an overlapping type
@@ -140,12 +119,12 @@ namespace FEXCore::ARMEmitter {
};
// This allows us to create a `VectorRegSizePair` union.
[[maybe_unused]]
static inline VectorRegSizePair ToVectorSizePair(SubRegSize size) {
[[nodiscard]]
constexpr VectorRegSizePair ToVectorSizePair(SubRegSize size) {
return VectorRegSizePair {.Vector = size};
}
[[maybe_unused]]
static inline VectorRegSizePair ToVectorSizePair(ScalarRegSize size) {
[[nodiscard]]
constexpr VectorRegSizePair ToVectorSizePair(ScalarRegSize size) {
return VectorRegSizePair {.Scalar = size};
}
@@ -524,7 +503,7 @@ namespace FEXCore::ARMEmitter {
uint8_t *Location{};
InstType Type;
};
std::vector<Instructions> Insts{};
fextl::vector<Instructions> Insts{};
};
/* This `BiDirectionalLabel` struct used for retaining a location for PC-Relative instructions.
@@ -544,6 +523,36 @@ namespace FEXCore::ARMEmitter {
ROTATE_270 = 0b11,
};
// Concept for contraining some instructions to accept only an XRegister or WRegister.
// Particularly for operations that differ encodings depending on which one is used.
template <typename T>
concept IsXOrWRegister = std::is_same_v<T, XRegister> || std::is_same_v<T, WRegister>;
// Whether or not a given set of vector registers are sequential
// in increasing order as far as the register file is concerned (modulo its size)
//
// For example, a set of registers like:
//
// v1, v2, v3 and
// v31, v0, v1
//
// would both be considered sequential sequences, and some instructions in particular
// limit register lists to these kind of sequences.
//
template <typename T, typename... Args>
constexpr bool AreVectorsSequential(T first, const Args&... args) {
// Ensure we always have a pair of registers to compare against.
static_assert(sizeof...(args) >= 1, "Number of arguments must be greater than 1");
const auto fn = [](auto& lhs, const auto& rhs) {
const auto result = ((lhs.Idx() + 1) % 32) == rhs.Idx();
lhs = rhs;
return result;
};
return (fn(first, args) && ...);
}
// This is an emitter that is designed around the smallest code bloat as possible.
// Eschewing most developer convenience in order to keep code as small as possible.
File diff suppressed because it is too large. Load diff
@@ -1,5 +1,8 @@
#pragma once
#include <FEXCore/Utils/EnumUtils.h>
#include <compare>
#include <cstdint>
namespace FEXCore::ARMEmitter {
@@ -15,13 +18,12 @@ namespace FEXCore::ARMEmitter {
constexpr explicit Register(uint32_t Idx)
: Index {Idx} {}
friend constexpr auto operator<=>(const Register&, const Register&) = default;
uint32_t Idx() const {
return Index;
}
operator WRegister() const;
operator XRegister() const;
WRegister W() const;
XRegister X() const;
@@ -41,9 +43,7 @@ namespace FEXCore::ARMEmitter {
constexpr explicit WRegister(uint32_t Idx)
: Index {Idx} {}
bool operator==(const WRegister &rhs) {
return Idx() == rhs.Idx();
}
friend constexpr auto operator<=>(const WRegister&, const WRegister&) = default;
uint32_t Idx() const {
return Index;
@@ -53,10 +53,7 @@ namespace FEXCore::ARMEmitter {
return Register(Index);
}
operator XRegister() const;
XRegister X() const;
Register R() const;
private:
@@ -75,9 +72,7 @@ namespace FEXCore::ARMEmitter {
constexpr explicit XRegister(uint32_t Idx)
: Index {Idx} {}
bool operator==(const XRegister &rhs) {
return Idx() == rhs.Idx();
}
friend constexpr auto operator<=>(const XRegister&, const XRegister&) = default;
uint32_t Idx() const {
return Index;
@@ -87,10 +82,7 @@ namespace FEXCore::ARMEmitter {
return Register(Index);
}
operator WRegister() const;
WRegister W() const;
Register R() const;
private:
@@ -101,45 +93,29 @@ namespace FEXCore::ARMEmitter {
static_assert(std::is_standard_layout_v<Register>, "Needs to be standard");
inline WRegister Register::W() const {
return *this;
return WRegister{Index};
}
inline XRegister Register::X() const {
return *this;
}
inline Register::operator WRegister () const {
return WRegister(Index);
}
inline Register::operator XRegister () const {
return XRegister(Index);
return XRegister{Index};
}
inline XRegister WRegister::X() const {
return *this;
return XRegister{Index};
}
inline Register WRegister::R() const {
return *this;
}
inline WRegister::operator XRegister () const {
return XRegister(Index);
}
inline WRegister XRegister::W() const {
return *this;
return WRegister{Index};
}
inline Register XRegister::R() const {
return *this;
}
inline XRegister::operator WRegister () const {
return WRegister(Index);
}
// Namespace containing all unsized GPR register objects.
namespace Reg {
constexpr static Register r0(0);
@@ -291,20 +267,15 @@ namespace FEXCore::ARMEmitter {
class VRegister {
public:
VRegister() = delete;
constexpr VRegister(uint32_t Idx)
constexpr explicit VRegister(uint32_t Idx)
: Index {Idx} {}
friend constexpr auto operator<=>(const VRegister&, const VRegister&) = default;
uint32_t Idx() const {
return Index;
}
operator BRegister() const;
operator HRegister() const;
operator SRegister() const;
operator DRegister() const;
operator QRegister() const;
operator ZRegister() const;
BRegister B() const;
HRegister H() const;
SRegister S() const;
@@ -328,16 +299,15 @@ namespace FEXCore::ARMEmitter {
constexpr explicit BRegister(uint32_t Idx)
: Index {Idx} {}
friend constexpr auto operator<=>(const BRegister&, const BRegister&) = default;
uint32_t Idx() const {
return Index;
}
operator VRegister() const;
operator HRegister() const;
operator SRegister() const;
operator DRegister() const;
operator QRegister() const;
operator ZRegister() const;
operator VRegister () const {
return VRegister(Index);
}
BRegister V() const;
HRegister H() const;
@@ -362,16 +332,15 @@ namespace FEXCore::ARMEmitter {
constexpr explicit HRegister(uint32_t Idx)
: Index {Idx} {}
friend constexpr auto operator<=>(const HRegister&, const HRegister&) = default;
uint32_t Idx() const {
return Index;
}
operator VRegister() const;
operator BRegister() const;
operator SRegister() const;
operator DRegister() const;
operator QRegister() const;
operator ZRegister() const;
operator VRegister() const {
return VRegister(Index);
}
HRegister V() const;
BRegister B() const;
@@ -396,16 +365,15 @@ namespace FEXCore::ARMEmitter {
constexpr explicit SRegister(uint32_t Idx)
: Index {Idx} {}
friend constexpr auto operator<=>(const SRegister&, const SRegister&) = default;
uint32_t Idx() const {
return Index;
}
operator VRegister() const;
operator BRegister() const;
operator HRegister() const;
operator DRegister() const;
operator QRegister() const;
operator ZRegister() const;
operator VRegister() const {
return VRegister(Index);
}
SRegister V() const;
BRegister B() const;
@@ -431,16 +399,15 @@ namespace FEXCore::ARMEmitter {
constexpr explicit DRegister(uint32_t Idx)
: Index {Idx} {}
friend constexpr auto operator<=>(const DRegister&, const DRegister&) = default;
uint32_t Idx() const {
return Index;
}
operator VRegister() const;
operator BRegister() const;
operator HRegister() const;
operator SRegister() const;
operator QRegister() const;
operator ZRegister() const;
operator VRegister() const {
return VRegister(Index);
}
DRegister V() const;
BRegister B() const;
@@ -466,16 +433,15 @@ namespace FEXCore::ARMEmitter {
constexpr explicit QRegister(uint32_t Idx)
: Index {Idx} {}
friend constexpr auto operator<=>(const QRegister&, const QRegister&) = default;
uint32_t Idx() const {
return Index;
}
operator VRegister() const;
operator BRegister() const;
operator HRegister() const;
operator SRegister() const;
operator DRegister() const;
operator ZRegister() const;
operator VRegister () const {
return VRegister(Index);
}
QRegister V() const;
BRegister B() const;
@@ -500,6 +466,8 @@ namespace FEXCore::ARMEmitter {
constexpr explicit ZRegister(uint32_t Idx)
: Index {Idx} {}
friend constexpr auto operator<=>(const ZRegister&, const ZRegister&) = default;
uint32_t Idx() const {
return Index;
}
@@ -520,41 +488,22 @@ namespace FEXCore::ARMEmitter {
// VRegister
inline BRegister VRegister::B() const {
return *this;
return BRegister{Index};
}
inline HRegister VRegister::H() const {
return *this;
return HRegister{Index};
}
inline SRegister VRegister::S() const {
return *this;
return SRegister{Index};
}
inline DRegister VRegister::D() const {
return *this;
return DRegister{Index};
}
inline QRegister VRegister::Q() const {
return *this;
return QRegister{Index};
}
inline ZRegister VRegister::Z() const {
return *this;
}
inline VRegister::operator BRegister () const {
return BRegister(Index);
}
inline VRegister::operator HRegister () const {
return HRegister(Index);
}
inline VRegister::operator SRegister () const {
return SRegister(Index);
}
inline VRegister::operator DRegister () const {
return DRegister(Index);
}
inline VRegister::operator QRegister () const {
return QRegister(Index);
}
inline VRegister::operator ZRegister () const {
return ZRegister(Index);
return ZRegister{Index};
}
// BRegister
@@ -562,38 +511,19 @@ namespace FEXCore::ARMEmitter {
return *this;
}
inline HRegister BRegister::H() const {
return *this;
return HRegister{Index};
}
inline SRegister BRegister::S() const {
return *this;
return SRegister{Index};
}
inline DRegister BRegister::D() const {
return *this;
return DRegister{Index};
}
inline QRegister BRegister::Q() const {
return *this;
return QRegister{Index};
}
inline ZRegister BRegister::Z() const {
return *this;
}
inline BRegister::operator VRegister () const {
return VRegister(Index);
}
inline BRegister::operator HRegister () const {
return HRegister(Index);
}
inline BRegister::operator SRegister () const {
return SRegister(Index);
}
inline BRegister::operator DRegister () const {
return DRegister(Index);
}
inline BRegister::operator QRegister () const {
return QRegister(Index);
}
inline BRegister::operator ZRegister () const {
return ZRegister(Index);
return ZRegister{Index};
}
// HRegister
@@ -601,38 +531,19 @@ namespace FEXCore::ARMEmitter {
return *this;
}
inline BRegister HRegister::B() const {
return *this;
return BRegister{Index};
}
inline SRegister HRegister::S() const {
return *this;
return SRegister{Index};
}
inline DRegister HRegister::D() const {
return *this;
return DRegister{Index};
}
inline QRegister HRegister::Q() const {
return *this;
return QRegister{Index};
}
inline ZRegister HRegister::Z() const {
return *this;
}
inline HRegister::operator VRegister () const {
return VRegister(Index);
}
inline HRegister::operator BRegister () const {
return BRegister(Index);
}
inline HRegister::operator SRegister () const {
return SRegister(Index);
}
inline HRegister::operator DRegister () const {
return DRegister(Index);
}
inline HRegister::operator QRegister () const {
return QRegister(Index);
}
inline HRegister::operator ZRegister () const {
return ZRegister(Index);
return ZRegister{Index};
}
// SRegister
@@ -640,77 +551,39 @@ namespace FEXCore::ARMEmitter {
return *this;
}
inline BRegister SRegister::B() const {
return *this;
return BRegister{Index};
}
inline HRegister SRegister::H() const {
return *this;
return HRegister{Index};
}
inline DRegister SRegister::D() const {
return *this;
return DRegister{Index};
}
inline QRegister SRegister::Q() const {
return *this;
return QRegister{Index};
}
inline ZRegister SRegister::Z() const {
return *this;
}
inline SRegister::operator VRegister () const {
return VRegister(Index);
}
inline SRegister::operator BRegister () const {
return BRegister(Index);
}
inline SRegister::operator HRegister () const {
return HRegister(Index);
}
inline SRegister::operator DRegister () const {
return DRegister(Index);
}
inline SRegister::operator QRegister () const {
return QRegister(Index);
}
inline SRegister::operator ZRegister () const {
return ZRegister(Index);
return ZRegister{Index};
}
// DRegister
inline DRegister DRegister::V() const {
return *this;
return DRegister{Index};
}
inline BRegister DRegister::B() const {
return *this;
return BRegister{Index};
}
inline HRegister DRegister::H() const {
return *this;
return HRegister{Index};
}
inline SRegister DRegister::S() const {
return *this;
return SRegister{Index};
}
inline QRegister DRegister::Q() const {
return *this;
return QRegister{Index};
}
inline ZRegister DRegister::Z() const {
return *this;
}
inline DRegister::operator VRegister () const {
return VRegister(Index);
}
inline DRegister::operator BRegister () const {
return BRegister(Index);
}
inline DRegister::operator HRegister () const {
return HRegister(Index);
}
inline DRegister::operator SRegister () const {
return SRegister(Index);
}
inline DRegister::operator QRegister () const {
return QRegister(Index);
}
inline DRegister::operator ZRegister () const {
return ZRegister(Index);
return ZRegister{Index};
}
// QRegister
@@ -718,38 +591,19 @@ namespace FEXCore::ARMEmitter {
return *this;
}
inline BRegister QRegister::B() const {
return *this;
return BRegister{Index};
}
inline HRegister QRegister::H() const {
return *this;
return HRegister{Index};
}
inline SRegister QRegister::S() const {
return *this;
return SRegister{Index};
}
inline DRegister QRegister::D() const {
return *this;
return DRegister{Index};
}
inline ZRegister QRegister::Z() const {
return *this;
}
inline QRegister::operator VRegister () const {
return VRegister(Index);
}
inline QRegister::operator BRegister () const {
return BRegister(Index);
}
inline QRegister::operator HRegister () const {
return HRegister(Index);
}
inline QRegister::operator SRegister () const {
return SRegister(Index);
}
inline QRegister::operator DRegister () const {
return DRegister(Index);
}
inline QRegister::operator ZRegister () const {
return ZRegister(Index);
return ZRegister{Index};
}
// ZRegister
@@ -1069,17 +923,12 @@ namespace FEXCore::ARMEmitter {
constexpr PRegister(uint32_t Idx)
: Index {Idx} {}
operator uint32_t() const {
return Index;
}
friend constexpr auto operator<=>(const PRegister&, const PRegister&) = default;
uint32_t Idx() const {
return Index;
}
operator PRegisterZero() const;
operator PRegisterMerge() const;
PRegisterZero Zeroing() const;
PRegisterMerge Merging() const;
@@ -1097,16 +946,13 @@ namespace FEXCore::ARMEmitter {
constexpr PRegisterZero(uint32_t Idx)
: Index {Idx} {}
operator uint32_t() const {
return Index;
}
friend constexpr auto operator<=>(const PRegisterZero&, const PRegisterZero&) = default;
uint32_t Idx() const {
return Index;
}
operator PRegister() const;
operator PRegisterMerge() const;
PRegister P() const;
PRegisterMerge Merging() const;
@@ -1125,16 +971,13 @@ namespace FEXCore::ARMEmitter {
constexpr PRegisterMerge(uint32_t Idx)
: Index {Idx} {}
operator uint32_t() const {
return Index;
}
friend constexpr auto operator<=>(const PRegisterMerge&, const PRegisterMerge&) = default;
uint32_t Idx() const {
return Index;
}
operator PRegister() const;
operator PRegisterZero() const;
PRegister P() const;
PRegisterZero Zeroing() const;
@@ -1148,14 +991,6 @@ namespace FEXCore::ARMEmitter {
// PRegister
inline PRegister::operator PRegisterZero() const {
return PRegisterZero(Index);
}
inline PRegister::operator PRegisterMerge() const {
return PRegisterMerge(Index);
}
inline PRegisterZero PRegister::Zeroing() const {
return PRegisterZero(Idx());
}
@@ -1169,10 +1004,6 @@ namespace FEXCore::ARMEmitter {
return PRegister(Index);
}
inline PRegisterZero::operator PRegisterMerge() const {
return PRegisterMerge(Index);
}
inline PRegister PRegisterZero::P() const {
return PRegister(Idx());
}
@@ -1186,10 +1017,6 @@ namespace FEXCore::ARMEmitter {
return PRegisterZero(Index);
}
inline PRegisterMerge::operator PRegisterZero() const {
return PRegisterZero(Index);
}
inline PRegister PRegisterMerge::P() const {
return PRegister(Idx());
}
File diff suppressed because it is too large. Load diff
File diff suppressed because it is too large. Load diff
+5 -4
View File
@@ -1,3 +1,4 @@
#include "FEXCore/Utils/AllocatorHooks.h"
#include "Interface/Context/Context.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include <FEXCore/Core/CPUBackend.h>
@@ -57,17 +58,17 @@ auto CPUBackend::AllocateNewCodeBuffer(size_t Size) -> CodeBuffer {
CodeBuffer Buffer;
Buffer.Size = Size;
Buffer.Ptr = static_cast<uint8_t *>(
FEXCore::Allocator::mmap(nullptr, Buffer.Size, PROT_READ | PROT_WRITE | PROT_EXEC, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0));
FEXCore::Allocator::VirtualAlloc(Buffer.Size, true));
LOGMAN_THROW_AA_FMT(!!Buffer.Ptr, "Couldn't allocate code buffer");
if (ThreadState->CTX->Config.GlobalJITNaming()) {
ThreadState->CTX->Symbols.RegisterJITSpace(Buffer.Ptr, Buffer.Size);
if (static_cast<Context::ContextImpl*>(ThreadState->CTX)->Config.GlobalJITNaming()) {
static_cast<Context::ContextImpl*>(ThreadState->CTX)->Symbols.RegisterJITSpace(Buffer.Ptr, Buffer.Size);
}
return Buffer;
}
void CPUBackend::FreeCodeBuffer(CodeBuffer Buffer) {
FEXCore::Allocator::munmap(Buffer.Ptr, Buffer.Size);
FEXCore::Allocator::VirtualFree(Buffer.Ptr, Buffer.Size);
}
bool CPUBackend::IsAddressInCodeBuffer(uintptr_t Address) const {
+58 -45
View File
@@ -8,18 +8,20 @@ $end_info$
#include "Common/StringConv.h"
#include "Interface/Context/Context.h"
#include "Interface/Core/CPUID.h"
#include "Utils/FileLoading.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/CPUID.h>
#include <FEXCore/Core/HostFeatures.h>
#include <FEXCore/Utils/CPUInfo.h>
#include <FEXCore/Utils/FileLoading.h>
#include <FEXCore/fextl/string.h>
#include <FEXHeaderUtils/Syscalls.h>
#include "git_version.h"
#include <cstring>
#ifdef _M_X86_64
#include <cpuid.h>
#include "Interface/Core/Dispatcher/X86Dispatcher.h"
#endif
namespace FEXCore {
@@ -74,20 +76,6 @@ static uint32_t GetCPUID() {
return CPU;
}
static uint32_t CalculateNumberOfCPUs() {
size_t CPUs = 1;
while(std::filesystem::exists("/sys/devices/system/cpu/cpu" + std::to_string(CPUs))) {
CPUs++;
}
return CPUs;
}
// TODO: Replace usages with CTX->HostFeatures.EnableAVX
// when AVX implementations are further along.
constexpr uint32_t SUPPORTS_AVX = 0;
#ifdef CPUID_AMD
constexpr uint32_t FAMILY_IDENTIFIER =
0 | // Stepping
@@ -115,13 +103,13 @@ static uint32_t GetCycleCounterFrequency() {
}
void CPUIDEmu::SetupHostHybridFlag() {
size_t CPUs = CalculateNumberOfCPUs();
size_t CPUs = FEXCore::CPUInfo::CalculateNumberOfCPUs();
PerCPUData.resize(CPUs);
uint64_t MIDR{};
for (size_t i = 0; i < CPUs; ++i) {
std::error_code ec{};
std::string MIDRPath = fmt::format("/sys/devices/system/cpu/cpu{}/regs/identification/midr_el1", i);
fextl::string MIDRPath = fextl::fmt::format("/sys/devices/system/cpu/cpu{}/regs/identification/midr_el1", i);
std::array<char, 18> Data;
// Needs to be a fixed size since depending on kernel it will try to read a full page of data and fail
@@ -217,8 +205,8 @@ void CPUIDEmu::SetupHostHybridFlag() {
if (Hybrid) {
// Walk the MIDRs and calculate big little designs
std::vector<const CPUMIDR*> BigCores;
std::vector<const CPUMIDR*> LittleCores;
fextl::vector<const CPUMIDR*> BigCores;
fextl::vector<const CPUMIDR*> LittleCores;
// Separate CPU cores out to big or little selected
for (size_t i = 0; i < CPUs; ++i) {
@@ -354,28 +342,28 @@ void CPUIDEmu::SetupHostHybridFlag() {
#else
static uint32_t GetCycleCounterFrequency() {
uint32_t eax, ebx, ecx, edx;
__cpuid(0, eax, ebx, ecx, edx);
if (eax >= 0x15) {
__cpuid(0x15, eax, ebx, ecx, edx);
uint32_t data[4];
Xbyak::util::Cpu::getCpuid(0, data);
if (data[0] >= 0x15) {
Xbyak::util::Cpu::getCpuid(0x15, data);
if (eax && ebx && ecx) {
return ecx * ebx / eax;
if (data[0] && data[1] && data[2]) {
return data[2] * data[1] / data[0];
}
}
return 0;
}
void CPUIDEmu::SetupHostHybridFlag() {
uint32_t eax, ebx, ecx, edx;
__cpuid(0, eax, ebx, ecx, edx);
if (eax >= 0x7) {
__cpuid(0x7, eax, ebx, ecx, edx);
uint32_t data[4];
Xbyak::util::Cpu::getCpuid(0, data);
if (data[0] >= 0x7) {
Xbyak::util::Cpu::getCpuid(0x7, data);
// Bit 15 of edx claims hybrid CPU
Hybrid = (edx & (1U << 15)) != 0;
Hybrid = (data[3] & (1U << 15)) != 0;
}
size_t CPUs = CalculateNumberOfCPUs();
size_t CPUs = FEXCore::CPUInfo::CalculateNumberOfCPUs();
PerCPUData.resize(CPUs);
for (size_t i = 0; i < CPUs; ++i) {
PerCPUData[i].IsBig = true;
@@ -410,6 +398,9 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_01h(uint32_t Leaf) {
// XXX: Enable once the rest of the SSE4.2 instructions are emulated
uint32_t SupportsSSE42 = CTX->HostFeatures.SupportsCRC && false ? 1 : 0;
// Hypervisor bit is normally set but some applications have issues with it.
uint32_t Hypervisor = HideHypervisorBit() ? 0 : 1;
Res.eax = FAMILY_IDENTIFIER;
Res.ebx = 0 | // Brand index
@@ -446,10 +437,10 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_01h(uint32_t Leaf) {
(CTX->HostFeatures.SupportsAES << 25) | // AES
(0 << 26) | // XSAVE
(0 << 27) | // OSXSAVE
(SUPPORTS_AVX << 28) | // AVX
(SupportsAVX() << 28) | // AVX
(0 << 29) | // F16C
(CTX->HostFeatures.SupportsRAND << 30) | // RDRAND
(1 << 31); // Hypervisor always returns one
(Hypervisor << 31);
Res.edx =
(1 << 0) | // FPU
@@ -635,12 +626,12 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_07h(uint32_t Leaf) {
(1 << 0) | // FS/GS support
(0 << 1) | // TSC adjust MSR
(0 << 2) | // SGX
(1 << 3) | // BMI1
(SupportsAVX() << 3) | // BMI1
(0 << 4) | // Intel Hardware Lock Elison
(0 << 5) | // AVX2 support
(1 << 6) | // FPU data pointer updated only on exception
(1 << 7) | // SMEP support
(1 << 8) | // BMI2
(SupportsAVX() << 8) | // BMI2
(0 << 9) | // Enhanced REP MOVSB/STOSB
(1 << 10) | // INVPCID for system software control of process-context
(0 << 11) | // Restricted transactional memory
@@ -704,7 +695,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_07h(uint32_t Leaf) {
(0 << 1) | // Reserved
(0 << 2) | // AVX512_4VNNIW
(0 << 3) | // AVX512_4FMAPS
(0 << 4) | // Fast Short Rep Mov
(1 << 4) | // Fast Short Rep Mov
(0 << 5) | // Reserved
(0 << 6) | // Reserved
(0 << 7) | // Reserved
@@ -741,13 +732,13 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_0Dh(uint32_t Leaf) {
// Leaf 0
FEXCore::CPUID::FunctionResults Res{};
uint32_t XFeatureSupportedSizeMax = SUPPORTS_AVX ? 0x0000'0340 : 0x0000'0240; // XFeatureEnabledSizeMax: Legacy Header + FPU/SSE + AVX
uint32_t XFeatureSupportedSizeMax = SupportsAVX() ? 0x0000'0340 : 0x0000'0240; // XFeatureEnabledSizeMax: Legacy Header + FPU/SSE + AVX
if (Leaf == 0) {
// XFeatureSupportedMask[31:0]
Res.eax =
(1 << 0) | // X87 support
(1 << 1) | // 128-bit SSE support
(SUPPORTS_AVX << 2) | // 256-bit AVX support
(SupportsAVX() << 2) | // 256-bit AVX support
(0b00 << 3) | // MPX State
(0b000 << 5) | // AVX-512 state
(0 << 8) | // "Used for IA32_XSS" ... Used for what?
@@ -781,8 +772,8 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_0Dh(uint32_t Leaf) {
Res.edx = 0;
}
else if (Leaf == 2) {
Res.eax = SUPPORTS_AVX ? 0x0000'0100 : 0; // YmmSaveStateSize
Res.ebx = SUPPORTS_AVX ? 0x0000'0240 : 0; // YmmSaveStateOffset
Res.eax = SupportsAVX() ? 0x0000'0100 : 0; // YmmSaveStateSize
Res.ebx = SupportsAVX() ? 0x0000'0240 : 0; // YmmSaveStateOffset
// Reserved
Res.ecx = 0;
@@ -877,6 +868,13 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0000h(uint32_t Leaf) {
// Extended processor and feature bits
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0001h(uint32_t Leaf) {
// RDTSCP is disabled on WIN32/Wine because there is no sane way to query processor ID.
#ifndef _WIN32
constexpr uint32_t SUPPORTS_RDTSCP = 0;
#else
constexpr uint32_t SUPPORTS_RDTSCP = 1;
#endif
FEXCore::CPUID::FunctionResults Res{};
Res.eax = FAMILY_IDENTIFIER;
@@ -943,7 +941,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0001h(uint32_t Leaf) {
(1 << 24) | // FXSAVE/FXRSTOR
(1 << 25) | // FXSAVE/FXRSTOR Optimizations
(0 << 26) | // 1 gigabit pages
(1 << 27) | // RDTSCP
(SUPPORTS_RDTSCP << 27) | // RDTSCP
(0 << 28) | // Reserved
(1 << 29) | // Long Mode
(1 << 30) | // 3DNow! Extensions
@@ -974,14 +972,14 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0004h(uint32_t Leaf) {
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0002h(uint32_t Leaf, uint32_t CPU) {
FEXCore::CPUID::FunctionResults Res{};
memset(&Res, ' ', sizeof(FEXCore::CPUID::FunctionResults));
memcpy(&Res, &ProcessorBrand[0], std::min(16L, DESCRIBE_STR_SIZE));
memcpy(&Res, &ProcessorBrand[0], std::min(ssize_t{16L}, DESCRIBE_STR_SIZE));
return Res;
}
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0003h(uint32_t Leaf, uint32_t CPU) {
FEXCore::CPUID::FunctionResults Res{};
memset(&Res, ' ', sizeof(FEXCore::CPUID::FunctionResults));
memcpy(&Res, &ProcessorBrand[16], std::max(0L, DESCRIBE_STR_SIZE - 16));
memcpy(&Res, &ProcessorBrand[16], std::max(ssize_t{0L}, DESCRIBE_STR_SIZE - 16));
return Res;
}
@@ -1210,11 +1208,26 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_Reserved(uint32_t Leaf) {
return Res;
}
void CPUIDEmu::Init(FEXCore::Context::Context *ctx) {
FEXCore::CPUID::XCRResults CPUIDEmu::XCRFunction_0h() {
// This just returns XCR0
FEXCore::CPUID::XCRResults Res{
.eax = static_cast<uint32_t>(XCR0),
.edx = static_cast<uint32_t>(XCR0 >> 32),
};
return Res;
}
void CPUIDEmu::Init(FEXCore::Context::ContextImpl *ctx) {
CTX = ctx;
// Setup some state tracking
SetupHostHybridFlag();
// TODO: Enable once AVX is supported.
if (false && CTX->HostFeatures.SupportsAVX) {
XCR0 |= XCR0_AVX;
}
}
}
+50 -8
View File
@@ -1,16 +1,16 @@
#pragma once
#include <FEXCore/Core/CPUID.h>
#include <FEXCore/Config/Config.h>
#include <FEXCore/fextl/vector.h>
#include <cstdint>
#include <unordered_map>
#include <utility>
#include <vector>
#include <FEXCore/Core/CPUID.h>
#include <FEXCore/Config/Config.h>
namespace FEXCore {
namespace Context {
struct Context;
class ContextImpl;
}
// Debugging define to switch what family of CPU we execute as.
@@ -31,7 +31,7 @@ public:
// if we report anything differently then applications are likely to break
constexpr static uint64_t CACHELINE_SIZE = 64;
void Init(FEXCore::Context::Context *ctx);
void Init(FEXCore::Context::ContextImpl *ctx);
FEXCore::CPUID::FunctionResults RunFunction(uint32_t Function, uint32_t Leaf) {
if (Function < Primary.size()) {
@@ -63,12 +63,52 @@ public:
return Function_8000_0004h(Leaf, CPU % PerCPUData.size());
}
FEXCore::CPUID::XCRResults RunXCRFunction(uint32_t Function) {
if (Function >= 1) {
// XCR function 1 is not yet supported.
return {};
}
return XCRFunction_0h();
}
private:
FEXCore::Context::Context *CTX;
FEXCore::Context::ContextImpl *CTX;
bool Hybrid{};
FEX_CONFIG_OPT(Cores, THREADS);
FEX_CONFIG_OPT(HideHypervisorBit, HIDEHYPERVISORBIT);
// XFEATURE_ENABLED_MASK
// Mask that configures what features are enabled on the CPU.
// Affects XSAVE and XRSTOR when modified.
// Bit layout is as follows.
// [0] - x87 enabled
// [1] - SSE enabled
// [2] - YMM enabled (256-bit SSE)
// [8:3] - Reserved. MBZ.
// [9] - MPK
// [10] - Reserved. MBZ.
// [11] - CET_U
// [12] - CET_S
// [61:13] - Reserved. MBZ.
// [62] - LWP (Lightweight profiling)
// [63] - Reserved for XCR bit vector expansion. MBZ.
// Always enable x87 and SSE by default.
constexpr static uint64_t XCR0_X87 = 1ULL << 0;
constexpr static uint64_t XCR0_SSE = 1ULL << 1;
constexpr static uint64_t XCR0_AVX = 1ULL << 2;
uint64_t XCR0 {
XCR0_X87 |
XCR0_SSE
};
uint32_t SupportsAVX() const {
return (XCR0 & XCR0_AVX) ? 1 : 0;
}
using FunctionHandler = FEXCore::CPUID::FunctionResults (CPUIDEmu::*)(uint32_t Leaf);
struct CPUData {
const char *ProductName{};
#ifdef _M_ARM_64
@@ -76,7 +116,7 @@ private:
#endif
bool IsBig{};
};
std::vector<CPUData> PerCPUData{};
fextl::vector<CPUData> PerCPUData{};
// Functions
FEXCore::CPUID::FunctionResults Function_0h(uint32_t Leaf);
@@ -108,6 +148,8 @@ private:
FEXCore::CPUID::FunctionResults Function_8000_001Dh(uint32_t Leaf);
FEXCore::CPUID::FunctionResults Function_Reserved(uint32_t Leaf);
FEXCore::CPUID::XCRResults XCRFunction_0h();
void SetupHostHybridFlag();
static constexpr std::array<FunctionHandler, 27> Primary = {
// 0: Highest function parameter and ID
+189 -224
View File
@@ -19,10 +19,12 @@ $end_info$
#include "Interface/Core/Interpreter/InterpreterCore.h"
#include "Interface/Core/JIT/JITCore.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Core/X86Tables/X86Tables.h"
#include "Interface/HLE/Thunks/Thunks.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
#include "Interface/IR/Passes.h"
#include "Interface/IR/PassManager.h"
#include "Utils/Allocator/HostAllocator.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/CodeLoader.h>
@@ -32,7 +34,6 @@ $end_info$
#include <FEXCore/Core/SignalDelegator.h>
#include <FEXCore/Core/X86Enums.h>
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXCore/Debug/X86Tables.h>
#include <FEXCore/HLE/SyscallHandler.h>
#include <FEXCore/HLE/SourcecodeResolver.h>
#include <FEXCore/HLE/Linux/ThreadManagement.h>
@@ -42,9 +43,15 @@ $end_info$
#include <FEXCore/IR/RegisterAllocationData.h>
#include <FEXCore/Utils/Allocator.h>
#include <FEXCore/Utils/Event.h>
#include <FEXCore/Utils/File.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/Threads.h>
#include <FEXCore/Utils/Profiler.h>
#include <FEXCore/fextl/fmt.h>
#include <FEXCore/fextl/memory.h>
#include <FEXCore/fextl/set.h>
#include <FEXCore/fextl/sstream.h>
#include <FEXCore/fextl/vector.h>
#include <FEXHeaderUtils/Syscalls.h>
#include <FEXHeaderUtils/TodoDefines.h>
@@ -53,33 +60,24 @@ $end_info$
#include <atomic>
#include <chrono>
#include <condition_variable>
#include <filesystem>
#include <fcntl.h>
#include <functional>
#include <fstream>
#include <map>
#include <memory>
#include <mutex>
#include <queue>
#include <set>
#include <shared_mutex>
#include <signal.h>
#include <stdio.h>
#include <string.h>
#include <string>
#include <string_view>
#include <sys/mman.h>
#include <sys/stat.h>
#include <sys/syscall.h>
#include <type_traits>
#include <unistd.h>
#include <unordered_map>
#include <utility>
#include <vector>
#include <xxhash.h>
namespace FEXCore::CPU {
bool CreateCPUCore(FEXCore::Context::Context *CTX) {
bool CreateCPUCore(Context::ContextImpl *CTX) {
// This should be used for generating things that are shared between threads
CTX->CPUID.Init(CTX);
return true;
@@ -147,18 +145,14 @@ std::string_view const& GetGRegName(unsigned Reg) {
} // namespace FEXCore::Core
namespace FEXCore::Context {
Context::Context()
ContextImpl::ContextImpl()
: IRCaptureCache {this} {
#ifdef BLOCKSTATS
BlockData = std::make_unique<FEXCore::BlockSamplingData>();
#endif
if (Config.CacheObjectCodeCompilation() != FEXCore::Config::ConfigObjectCodeHandler::CONFIG_NONE) {
CodeObjectCacheService = std::make_unique<FEXCore::CodeSerialize::CodeObjectSerializeService>(this);
CodeObjectCacheService = fextl::make_unique<FEXCore::CodeSerialize::CodeObjectSerializeService>(this);
}
if (!Config.EnableAVX) {
HostFeatures.SupportsAVX = false;
}
if (!Config.Is64BitMode()) {
// When operating in 32-bit mode, the virtual memory we care about is only the lower 32-bits.
Config.VirtualMemSize = 1ULL << 32;
@@ -170,9 +164,16 @@ namespace FEXCore::Context {
// Only initialize symbols file if enabled. Ensures we don't pollute /tmp with empty files.
Symbols.InitFile();
}
// Track atomic TSO emulation configuration.
UpdateAtomicTSOEmulationConfig();
}
Context::~Context() {
ContextImpl::~ContextImpl() {
if (ParentThread) {
DestroyThread(ParentThread);
}
{
if (CodeObjectCacheService) {
CodeObjectCacheService->Shutdown();
@@ -191,44 +192,39 @@ namespace FEXCore::Context {
}
}
static FEXCore::Core::CPUState CreateDefaultCPUState() {
FEXCore::Core::CPUState NewThreadState{};
uint64_t ContextImpl::RestoreRIPFromHostPC(FEXCore::Core::InternalThreadState *Thread, uint64_t HostPC) {
const auto Frame = Thread->CurrentFrame;
const uint64_t BlockBegin = Frame->State.InlineJITBlockHeader;
const CPU::CPUBackend::JITCodeHeader *InlineHeader = reinterpret_cast<const CPU::CPUBackend::JITCodeHeader *>(BlockBegin);
// Initialize default CPU state
NewThreadState.rip = ~0ULL;
for (auto& greg : NewThreadState.gregs) {
greg = 0;
if (InlineHeader) {
const CPU::CPUBackend::JITCodeTail *InlineTail = reinterpret_cast<const CPU::CPUBackend::JITCodeTail *>(Frame->State.InlineJITBlockHeader + InlineHeader->OffsetToBlockTail);
// Check if the host PC is currently within a code block.
// If it is then RIP can be reconstructed from the beginning of the code block.
// This is currently as close as FEX can get RIP reconstructions.
if (HostPC >= reinterpret_cast<uint64_t>(BlockBegin) &&
HostPC < reinterpret_cast<uint64_t>(BlockBegin + InlineTail->Size)) {
return InlineTail->RIP;
}
}
for (auto& xmm : NewThreadState.xmm.avx.data) {
xmm[0] = 0xDEADBEEFULL;
xmm[1] = 0xBAD0DAD1ULL;
xmm[2] = 0xDEADCAFEULL;
xmm[3] = 0xBAD2CAD3ULL;
}
memset(NewThreadState.flags, 0, Core::CPUState::NUM_EFLAG_BITS);
NewThreadState.flags[1] = 1;
NewThreadState.flags[9] = 1;
NewThreadState.FCW = 0x37F;
NewThreadState.FTW = 0xFFFF;
return NewThreadState;
// Fallback to what is stored in the RIP currently.
return Frame->State.rip;
}
FEXCore::Core::InternalThreadState* Context::InitCore(uint64_t InitialRIP, uint64_t StackPointer) {
FEXCore::Core::InternalThreadState* ContextImpl::InitCore(uint64_t InitialRIP, uint64_t StackPointer) {
// Initialize the CPU core signal handlers & DispatcherConfig
switch (Config.Core) {
#ifdef INTERPRETER_ENABLED
case FEXCore::Config::CONFIG_INTERPRETER:
FEXCore::CPU::InitializeInterpreterSignalHandlers(this);
BackendFeatures = FEXCore::CPU::GetInterpreterBackendFeatures();
break;
#endif
case FEXCore::Config::CONFIG_IRJIT:
#if (_M_X86_64 && JIT_X86_64)
FEXCore::CPU::InitializeX86JITSignalHandlers(this);
BackendFeatures = FEXCore::CPU::GetX86JITBackendFeatures();
#elif (_M_ARM_64 && JIT_ARM64) || defined(VIXL_SIMULATOR)
FEXCore::CPU::InitializeArm64JITSignalHandlers(this);
BackendFeatures = FEXCore::CPU::GetArm64JITBackendFeatures();
#else
ERROR_AND_DIE_FMT("FEXCore has been compiled without a viable JIT core");
@@ -252,24 +248,37 @@ namespace FEXCore::Context {
ERROR_AND_DIE_FMT("FEXCore has been compiled with an unknown target");
#endif
// Initialize common signal handlers
// Set up the SignalDelegator config since core is initialized.
FEXCore::SignalDelegator::SignalDelegatorConfig SignalConfig {
.StaticRegisterAllocation = DispatcherConfig.StaticRegisterAllocation,
.SupportsAVX = HostFeatures.SupportsAVX,
auto PauseHandler = [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
return Thread->CTX->Dispatcher->HandleSignalPause(Thread, Signal, info, ucontext);
.DispatcherBegin = Dispatcher->Start,
.DispatcherEnd = Dispatcher->End,
.AbsoluteLoopTopAddressFillSRA = Dispatcher->AbsoluteLoopTopAddressFillSRA,
.SignalHandlerReturnAddress = Dispatcher->SignalHandlerReturnAddress,
.SignalHandlerReturnAddressRT = Dispatcher->SignalHandlerReturnAddressRT,
.PauseReturnInstruction = Dispatcher->PauseReturnInstruction,
.ThreadPauseHandlerAddressSpillSRA = Dispatcher->ThreadPauseHandlerAddressSpillSRA,
.ThreadPauseHandlerAddress = Dispatcher->ThreadPauseHandlerAddress,
// Stop handlers.
.ThreadStopHandlerAddressSpillSRA = Dispatcher->ThreadStopHandlerAddressSpillSRA,
.ThreadStopHandlerAddress = Dispatcher->ThreadStopHandlerAddress,
// SRA information.
.SRAGPRCount = Dispatcher->GetSRAGPRCount(),
.SRAFPRCount = Dispatcher->GetSRAFPRCount(),
};
SignalDelegation->RegisterHostSignalHandler(SignalDelegator::SIGNAL_FOR_PAUSE, PauseHandler, true);
Dispatcher->GetSRAGPRMapping(SignalConfig.SRAGPRMapping);
Dispatcher->GetSRAFPRMapping(SignalConfig.SRAFPRMapping);
auto GuestSignalHandler = [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack) -> bool {
return Thread->CTX->Dispatcher->HandleGuestSignal(Thread, Signal, info, ucontext, GuestAction, GuestStack);
};
// Give this configuration to the SignalDelegator.
SignalDelegation->SetConfig(SignalConfig);
for (uint32_t Signal = 0; Signal <= SignalDelegator::MAX_SIGNALS; ++Signal) {
SignalDelegation->RegisterHostSignalHandlerForGuest(Signal, GuestSignalHandler);
}
// Initialize GDBServer after the signal handlers are installed
// It may install its own handlers that need to be executed AFTER the CPU cores
if (Config.GdbServer) {
StartGdbServer();
}
@@ -277,12 +286,13 @@ namespace FEXCore::Context {
StopGdbServer();
}
ThunkHandler.reset(FEXCore::ThunkHandler::Create());
#ifndef _WIN32
ThunkHandler = FEXCore::ThunkHandler::Create();
#endif
using namespace FEXCore::Core;
FEXCore::Core::CPUState NewThreadState = CreateDefaultCPUState();
FEXCore::Core::InternalThreadState *Thread = CreateThread(&NewThreadState, 0);
FEXCore::Core::InternalThreadState *Thread = CreateThread(nullptr, 0);
// We are the parent thread
ParentThread = Thread;
@@ -295,30 +305,26 @@ namespace FEXCore::Context {
return Thread;
}
void Context::StartGdbServer() {
void ContextImpl::StartGdbServer() {
#ifndef _WIN32
if (!DebugServer) {
DebugServer = std::make_unique<GdbServer>(this);
DebugServer = fextl::make_unique<GdbServer>(this);
StartPaused = true;
}
#endif
}
void Context::StopGdbServer() {
void ContextImpl::StopGdbServer() {
#ifndef _WIN32
DebugServer.reset();
#endif
}
void Context::HandleCallback(FEXCore::Core::InternalThreadState *Thread, uint64_t RIP) {
Thread->CTX->Dispatcher->ExecuteJITCallback(Thread->CurrentFrame, RIP);
void ContextImpl::HandleCallback(FEXCore::Core::InternalThreadState *Thread, uint64_t RIP) {
static_cast<ContextImpl*>(Thread->CTX)->Dispatcher->ExecuteJITCallback(Thread->CurrentFrame, RIP);
}
void Context::RegisterHostSignalHandler(int Signal, HostSignalDelegatorFunction Func, bool Required) {
SignalDelegation->RegisterHostSignalHandler(Signal, Func, Required);
}
void Context::RegisterFrontendHostSignalHandler(int Signal, HostSignalDelegatorFunction Func, bool Required) {
SignalDelegation->RegisterFrontendHostSignalHandler(Signal, Func, Required);
}
void Context::WaitForIdle() {
void ContextImpl::WaitForIdle() {
std::unique_lock<std::mutex> lk(IdleWaitMutex);
IdleWaitCV.wait(lk, [this] {
return IdleWaitRefCount.load() == 0;
@@ -327,7 +333,7 @@ namespace FEXCore::Context {
Running = false;
}
void Context::WaitForIdleWithTimeout() {
void ContextImpl::WaitForIdleWithTimeout() {
std::unique_lock<std::mutex> lk(IdleWaitMutex);
bool WaitResult = IdleWaitCV.wait_for(lk, std::chrono::milliseconds(1500),
[this] {
@@ -345,20 +351,16 @@ namespace FEXCore::Context {
WaitForIdle();
}
void Context::NotifyPause() {
void ContextImpl::NotifyPause() {
// Tell all the threads that they should pause
std::lock_guard<std::mutex> lk(ThreadCreationMutex);
for (auto &Thread : Threads) {
Thread->SignalReason.store(FEXCore::Core::SignalEvent::Pause);
if (Thread->RunningEvents.Running.load()) {
// Only attempt to stop this thread if it is running
FHU::Syscalls::tgkill(Thread->ThreadManager.PID, Thread->ThreadManager.TID, SignalDelegator::SIGNAL_FOR_PAUSE);
}
SignalDelegation->SignalThread(Thread, FEXCore::Core::SignalEvent::Pause);
}
}
void Context::Pause() {
void ContextImpl::Pause() {
// If we aren't running, WaitForIdle will never compete.
if (Running) {
NotifyPause();
@@ -367,7 +369,7 @@ namespace FEXCore::Context {
}
}
void Context::Run() {
void ContextImpl::Run() {
// Spin up all the threads
std::lock_guard<std::mutex> lk(ThreadCreationMutex);
for (auto &Thread : Threads) {
@@ -379,7 +381,7 @@ namespace FEXCore::Context {
}
}
void Context::WaitForThreadsToRun() {
void ContextImpl::WaitForThreadsToRun() {
size_t NumThreads{};
{
std::lock_guard<std::mutex> lk(ThreadCreationMutex);
@@ -395,7 +397,7 @@ namespace FEXCore::Context {
Running = true;
}
void Context::Step() {
void ContextImpl::Step() {
{
std::lock_guard<std::mutex> lk(ThreadCreationMutex);
// Walk the threads and tell them to clear their caches
@@ -415,7 +417,7 @@ namespace FEXCore::Context {
this->Config.MaxInstPerBlock = PreviousMaxIntPerBlock;
}
void Context::Stop(bool IgnoreCurrentThread) {
void ContextImpl::Stop(bool IgnoreCurrentThread) {
pid_t tid = FHU::Syscalls::gettid();
FEXCore::Core::InternalThreadState* CurrentThread{};
@@ -453,21 +455,19 @@ namespace FEXCore::Context {
}
}
void Context::StopThread(FEXCore::Core::InternalThreadState *Thread) {
void ContextImpl::StopThread(FEXCore::Core::InternalThreadState *Thread) {
if (Thread->RunningEvents.Running.exchange(false)) {
Thread->SignalReason.store(FEXCore::Core::SignalEvent::Stop);
FHU::Syscalls::tgkill(Thread->ThreadManager.PID, Thread->ThreadManager.TID, SignalDelegator::SIGNAL_FOR_PAUSE);
SignalDelegation->SignalThread(Thread, FEXCore::Core::SignalEvent::Stop);
}
}
void Context::SignalThread(FEXCore::Core::InternalThreadState *Thread, FEXCore::Core::SignalEvent Event) {
void ContextImpl::SignalThread(FEXCore::Core::InternalThreadState *Thread, FEXCore::Core::SignalEvent Event) {
if (Thread->RunningEvents.Running.load()) {
Thread->SignalReason.store(Event);
FHU::Syscalls::tgkill(Thread->ThreadManager.PID, Thread->ThreadManager.TID, SignalDelegator::SIGNAL_FOR_PAUSE);
SignalDelegation->SignalThread(Thread, Event);
}
}
FEXCore::Context::ExitReason Context::RunUntilExit() {
FEXCore::Context::ExitReason ContextImpl::RunUntilExit() {
if(!StartPaused) {
// We will only have one thread at this point, but just in case run notify everything
std::lock_guard lk(ThreadCreationMutex);
@@ -488,16 +488,16 @@ namespace FEXCore::Context {
}
}
int Context::GetProgramStatus() const {
int ContextImpl::GetProgramStatus() const {
return ParentThread->StatusCode;
}
void Context::InitializeThreadData(FEXCore::Core::InternalThreadState *Thread) {
void ContextImpl::InitializeThreadData(FEXCore::Core::InternalThreadState *Thread) {
Thread->CPUBackend->Initialize();
}
struct ExecutionThreadHandler {
FEXCore::Context::Context *This;
ContextImpl *This;
FEXCore::Core::InternalThreadState *Thread;
};
@@ -508,7 +508,7 @@ namespace FEXCore::Context {
return nullptr;
}
void Context::InitializeThread(FEXCore::Core::InternalThreadState *Thread) {
void ContextImpl::InitializeThread(FEXCore::Core::InternalThreadState *Thread) {
// This will create the execution thread but it won't actually start executing
ExecutionThreadHandler *Arg = reinterpret_cast<ExecutionThreadHandler*>(FEXCore::Allocator::malloc(sizeof(ExecutionThreadHandler)));
Arg->This = this;
@@ -530,25 +530,27 @@ namespace FEXCore::Context {
}
}
void Context::InitializeThreadTLSData(FEXCore::Core::InternalThreadState *Thread) {
void ContextImpl::InitializeThreadTLSData(FEXCore::Core::InternalThreadState *Thread) {
// Let's do some initial bookkeeping here
Thread->ThreadManager.TID = FHU::Syscalls::gettid();
Thread->ThreadManager.PID = ::getpid();
SignalDelegation->RegisterTLSState(Thread);
ThunkHandler->RegisterTLSState(Thread);
if (ThunkHandler) {
ThunkHandler->RegisterTLSState(Thread);
}
}
void Context::RunThread(FEXCore::Core::InternalThreadState *Thread) {
void ContextImpl::RunThread(FEXCore::Core::InternalThreadState *Thread) {
// Tell the thread to start executing
Thread->StartRunning.NotifyAll();
}
void Context::InitializeCompiler(FEXCore::Core::InternalThreadState* Thread) {
Thread->OpDispatcher = std::make_unique<FEXCore::IR::OpDispatchBuilder>(this);
void ContextImpl::InitializeCompiler(FEXCore::Core::InternalThreadState* Thread) {
Thread->OpDispatcher = fextl::make_unique<FEXCore::IR::OpDispatchBuilder>(this);
Thread->OpDispatcher->SetMultiblock(Config.Multiblock);
Thread->LookupCache = std::make_unique<FEXCore::LookupCache>(this);
Thread->FrontendDecoder = std::make_unique<FEXCore::Frontend::Decoder>(this);
Thread->PassManager = std::make_unique<FEXCore::IR::PassManager>();
Thread->LookupCache = fextl::make_unique<FEXCore::LookupCache>(this);
Thread->FrontendDecoder = fextl::make_unique<FEXCore::Frontend::Decoder>(this);
Thread->PassManager = fextl::make_unique<FEXCore::IR::PassManager>();
Thread->PassManager->RegisterExitHandler([this]() {
Stop(false /* Ignore current thread */);
});
@@ -595,11 +597,13 @@ namespace FEXCore::Context {
}
}
FEXCore::Core::InternalThreadState* Context::CreateThread(FEXCore::Core::CPUState *NewThreadState, uint64_t ParentTID) {
FEXCore::Core::InternalThreadState* ContextImpl::CreateThread(FEXCore::Core::CPUState *NewThreadState, uint64_t ParentTID) {
FEXCore::Core::InternalThreadState *Thread = new FEXCore::Core::InternalThreadState{};
// Copy over the new thread state to the new object
memcpy(Thread->CurrentFrame, NewThreadState, sizeof(FEXCore::Core::CPUState));
if (NewThreadState) {
memcpy(Thread->CurrentFrame, NewThreadState, sizeof(FEXCore::Core::CPUState));
}
Thread->CurrentFrame->Thread = Thread;
// Set up the thread manager state
@@ -608,6 +612,9 @@ namespace FEXCore::Context {
InitializeCompiler(Thread);
InitializeThreadData(Thread);
Thread->CurrentFrame->State.DeferredSignalRefCount.Store(0);
Thread->CurrentFrame->State.DeferredSignalFaultAddress = reinterpret_cast<Core::NonAtomicRefCounter<uint64_t>*>(FEXCore::Allocator::VirtualAlloc(4096));
// Insert after the Thread object has been fully initialized
{
std::lock_guard lk(ThreadCreationMutex);
@@ -617,7 +624,7 @@ namespace FEXCore::Context {
return Thread;
}
void Context::DestroyThread(FEXCore::Core::InternalThreadState *Thread) {
void ContextImpl::DestroyThread(FEXCore::Core::InternalThreadState *Thread) {
// remove new thread object
{
std::lock_guard lk(ThreadCreationMutex);
@@ -633,10 +640,12 @@ namespace FEXCore::Context {
// To be able to delete a thread from itself, we need to detached the std::thread object
Thread->ExecutionThread->detach();
}
FEXCore::Allocator::VirtualFree(reinterpret_cast<void*>(Thread->CurrentFrame->State.DeferredSignalFaultAddress), 4096);
delete Thread;
}
void Context::CleanupAfterFork(FEXCore::Core::InternalThreadState *LiveThread) {
void ContextImpl::CleanupAfterFork(FEXCore::Core::InternalThreadState *LiveThread) {
// This function is called after fork
// We need to cleanup some of the thread data that is dead
for (auto &DeadThread : Threads) {
@@ -675,11 +684,11 @@ namespace FEXCore::Context {
FEXCore::Threads::Thread::CleanupAfterFork();
}
void Context::AddBlockMapping(FEXCore::Core::InternalThreadState *Thread, uint64_t Address, void *Ptr) {
void ContextImpl::AddBlockMapping(FEXCore::Core::InternalThreadState *Thread, uint64_t Address, void *Ptr) {
Thread->LookupCache->AddBlockMapping(Address, Ptr);
}
void Context::ClearCodeCache(FEXCore::Core::InternalThreadState *Thread) {
void ContextImpl::ClearCodeCache(FEXCore::Core::InternalThreadState *Thread) {
FEXCORE_PROFILE_INSTANT("ClearCodeCache");
{
@@ -695,49 +704,46 @@ namespace FEXCore::Context {
}
static void IRDumper(FEXCore::Core::InternalThreadState *Thread, IR::IREmitter *IREmitter, uint64_t GuestRIP, IR::RegisterAllocationData* RA) {
FILE* f = nullptr;
bool CloseAfter = false;
const auto DumpIRStr = Thread->CTX->Config.DumpIR();
FEXCore::File::File FD;
const auto DumpIRStr = static_cast<ContextImpl*>(Thread->CTX)->Config.DumpIR();
// DumpIRStr might be no if not dumping but ShouldDump is set in OpDisp
if (DumpIRStr =="stderr" || DumpIRStr =="no") {
f = stderr;
FD = FEXCore::File::File::GetStdERR();
}
else if (DumpIRStr =="stdout") {
f = stdout;
FD = FEXCore::File::File::GetStdOUT();
}
else {
const auto fileName = fmt::format("{}/{:x}{}", DumpIRStr, GuestRIP, RA ? "-post.ir" : "-pre.ir");
f = fopen(fileName.c_str(), "w");
CloseAfter = true;
const auto fileName = fextl::fmt::format("{}/{:x}{}", DumpIRStr, GuestRIP, RA ? "-post.ir" : "-pre.ir");
FD = FEXCore::File::File(fileName.c_str(),
FEXCore::File::FileModes::WRITE |
FEXCore::File::FileModes::CREATE |
FEXCore::File::FileModes::TRUNCATE);
}
if (f) {
std::stringstream out;
if (FD.IsValid()) {
fextl::stringstream out;
auto NewIR = IREmitter->ViewIR();
FEXCore::IR::Dump(&out, &NewIR, RA);
fmt::print(f,"IR-{} 0x{:x}:\n{}\n@@@@@\n", RA ? "post" : "pre", GuestRIP, out.str());
if (CloseAfter) {
fclose(f);
}
fextl::fmt::print(FD, "IR-{} 0x{:x}:\n{}\n@@@@@\n", RA ? "post" : "pre", GuestRIP, out.str());
}
};
static void ValidateIR(FEXCore::Context::Context *ctx, IR::IREmitter *IREmitter) {
static void ValidateIR(ContextImpl *ctx, IR::IREmitter *IREmitter) {
// Convert to text, Parse, Convert to text again and make sure the texts match
std::stringstream out;
fextl::stringstream out;
static auto compaction = IR::CreateIRCompaction(ctx->OpDispatcherAllocator);
compaction->Run(IREmitter);
auto NewIR = IREmitter->ViewIR();
Dump(&out, &NewIR, nullptr);
out.seekg(0);
FEXCore::Utils::PooledAllocatorMalloc Allocator;
auto reparsed = IR::Parse(Allocator, &out);
auto reparsed = IR::Parse(Allocator, out);
if (reparsed == nullptr) {
LOGMAN_MSG_A_FMT("Failed to parse IR\n");
} else {
std::stringstream out2;
fextl::stringstream out2;
auto NewIR2 = reparsed->ViewIR();
Dump(&out2, &NewIR2, nullptr);
if (out.str() != out2.str()) {
@@ -748,7 +754,7 @@ namespace FEXCore::Context {
}
}
Context::GenerateIRResult Context::GenerateIR(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP, bool ExtendedDebugInfo) {
ContextImpl::GenerateIRResult ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP, bool ExtendedDebugInfo) {
FEXCORE_PROFILE_SCOPED("GenerateIR");
Thread->OpDispatcher->ReownOrClaimBuffer();
@@ -775,7 +781,7 @@ namespace FEXCore::Context {
Thread->FrontendDecoder->DecodeInstructionsAtEntry(GuestCode, GuestRIP, [Thread](uint64_t BlockEntry, uint64_t Start, uint64_t Length) {
if (Thread->LookupCache->AddBlockExecutableRange(BlockEntry, Start, Length)) {
Thread->CTX->SyscallHandler->MarkGuestExecutableRange(Start, Length);
static_cast<ContextImpl*>(Thread->CTX)->SyscallHandler->MarkGuestExecutableRange(Thread, Start, Length);
}
});
@@ -795,13 +801,6 @@ namespace FEXCore::Context {
// Reset any block-specific state
Thread->OpDispatcher->StartNewBlock();
if (Config.x86dec_SynchronizeRIPOnAllBlocks) {
// Ensure the RIP is synchronized to the context on block entry.
// In the case of block linking, the RIP may not have synchronized.
auto NewRIP = Thread->OpDispatcher->_EntrypointOffset(Block.Entry - GuestRIP, GPRSize);
Thread->OpDispatcher->_StoreContext(GPRSize, IR::GPRClass, NewRIP, offsetof(FEXCore::Core::CPUState, rip));
}
uint64_t InstsInBlock = Block.NumInstructions;
for (size_t i = 0; i < InstsInBlock; ++i) {
@@ -856,6 +855,9 @@ namespace FEXCore::Context {
}
}
else {
if (TableInfo) {
LogMan::Msg::EFmt("Invalid or Unknown instruction: {} 0x{:x}", TableInfo->Name ?: "UND", Block.Entry - GuestRIP);
}
// Invalid instruction
Thread->OpDispatcher->InvalidOp(DecodedInfo);
Thread->OpDispatcher->_ExitFunction(Thread->OpDispatcher->_EntrypointOffset(Block.Entry - GuestRIP, GPRSize));
@@ -893,14 +895,14 @@ namespace FEXCore::Context {
IR::IREmitter *IREmitter = Thread->OpDispatcher.get();
auto ShouldDump = Thread->CTX->Config.DumpIR() != "no" || Thread->OpDispatcher->ShouldDump;
auto ShouldDump = static_cast<ContextImpl*>(Thread->CTX)->Config.DumpIR() != "no" || Thread->OpDispatcher->ShouldDump;
// Debug
{
if (ShouldDump) {
IRDumper(Thread, IREmitter, GuestRIP, nullptr);
}
if (Thread->CTX->Config.ValidateIRarser) {
if (static_cast<ContextImpl*>(Thread->CTX)->Config.ValidateIRarser) {
ValidateIR(this, IREmitter);
}
}
@@ -930,7 +932,7 @@ namespace FEXCore::Context {
};
}
Context::CompileCodeResult Context::CompileCode(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP) {
ContextImpl::CompileCodeResult ContextImpl::CompileCode(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP) {
FEXCore::IR::IRListView *IRList {};
FEXCore::Core::DebugData *DebugData {};
FEXCore::IR::RegisterAllocationData::UniquePtr RAData {};
@@ -958,7 +960,7 @@ namespace FEXCore::Context {
}
if (SourcecodeResolver && Config.GDBSymbols()) {
auto AOTIRCacheEntry = SyscallHandler->LookupAOTIRCacheEntry(GuestRIP);
auto AOTIRCacheEntry = SyscallHandler->LookupAOTIRCacheEntry(Thread, GuestRIP);
if (AOTIRCacheEntry.Entry && !AOTIRCacheEntry.Entry->ContainsCode) {
AOTIRCacheEntry.Entry->SourcecodeMap =
SourcecodeResolver->GenerateMap(AOTIRCacheEntry.Entry->Filename, AOTIRCacheEntry.Entry->FileId);
@@ -967,7 +969,7 @@ namespace FEXCore::Context {
// AOT IR bookkeeping and cache
{
auto [IRCopy, RACopy, DebugDataCopy, _StartAddr, _Length, _GeneratedIR] = IRCaptureCache.PreGenerateIRFetch(GuestRIP, IRList);
auto [IRCopy, RACopy, DebugDataCopy, _StartAddr, _Length, _GeneratedIR] = IRCaptureCache.PreGenerateIRFetch(Thread, GuestRIP, IRList);
if (_GeneratedIR) {
// Setup pointers to internal structures
IRList = IRCopy;
@@ -990,9 +992,6 @@ namespace FEXCore::Context {
StartAddr = _StartAddr;
Length = _Length;
// Increment stats
Thread->Stats.BlocksCompiled.fetch_add(1);
// These blocks aren't already in the cache
GeneratedIR = true;
}
@@ -1002,7 +1001,10 @@ namespace FEXCore::Context {
}
// Attempt to get the CPU backend to compile this code
return {
.CompiledCode = Thread->CPUBackend->CompileCode(GuestRIP, IRList, DebugData, RAData.get(), GetGdbServerStatus()),
// FEX currently throws away the CPUBackend::CompiledCode object other than the entrypoint
// In the future with code caching getting wired up, we will pass the rest of the data forward.
// TODO: Pass the data forward when code caching is wired up to this.
.CompiledCode = Thread->CPUBackend->CompileCode(GuestRIP, IRList, DebugData, RAData.get(), GetGdbServerStatus()).BlockEntry,
.IRData = IRList,
.DebugData = DebugData,
.RAData = std::move(RAData),
@@ -1012,7 +1014,7 @@ namespace FEXCore::Context {
};
}
void Context::CompileBlockJit(FEXCore::Core::CpuStateFrame *Frame, uint64_t GuestRIP) {
void ContextImpl::CompileBlockJit(FEXCore::Core::CpuStateFrame *Frame, uint64_t GuestRIP) {
auto NewBlock = CompileBlock(Frame, GuestRIP);
if (NewBlock == 0) {
@@ -1023,7 +1025,7 @@ namespace FEXCore::Context {
}
}
uintptr_t Context::CompileBlock(FEXCore::Core::CpuStateFrame *Frame, uint64_t GuestRIP) {
uintptr_t ContextImpl::CompileBlock(FEXCore::Core::CpuStateFrame *Frame, uint64_t GuestRIP) {
FEXCORE_PROFILE_SCOPED("CompileBlock");
auto Thread = Frame->Thread;
@@ -1060,7 +1062,7 @@ namespace FEXCore::Context {
auto FragmentBasePtr = reinterpret_cast<uint8_t *>(CodePtr);
if (DebugData) {
auto GuestRIPLookup = SyscallHandler->LookupAOTIRCacheEntry(GuestRIP);
auto GuestRIPLookup = SyscallHandler->LookupAOTIRCacheEntry(Thread, GuestRIP);
if (DebugData->Subblocks.size()) {
for (auto& Subblock: DebugData->Subblocks) {
@@ -1085,7 +1087,7 @@ namespace FEXCore::Context {
if (CodeObjectCacheService &&
Config.CacheObjectCodeCompilation == FEXCore::Config::ConfigObjectCodeHandler::CONFIG_READWRITE &&
DebugData) {
CodeObjectCacheService->AsyncAddSerializationJob(std::make_unique<CodeSerialize::AsyncJobHandler::SerializationJobData>(
CodeObjectCacheService->AsyncAddSerializationJob(fextl::make_unique<CodeSerialize::AsyncJobHandler::SerializationJobData>(
CodeSerialize::AsyncJobHandler::SerializationJobData {
.GuestRIP = GuestRIP,
.GuestCodeLength = Length,
@@ -1123,18 +1125,19 @@ namespace FEXCore::Context {
return (uintptr_t)CodePtr;
}
void Context::ExecutionThread(FEXCore::Core::InternalThreadState *Thread) {
void ContextImpl::ExecutionThread(FEXCore::Core::InternalThreadState *Thread) {
Core::ThreadData.Thread = Thread;
Thread->ExitReason = FEXCore::Context::ExitReason::EXIT_WAITING;
InitializeThreadTLSData(Thread);
Alloc::OSAllocator::RegisterTLSData(Thread);
++IdleWaitRefCount;
// Now notify the thread that we are initialized
Thread->ThreadWaiting.NotifyAll();
if (Thread != Thread->CTX->ParentThread || StartPaused || Thread->StartPaused) {
if (Thread != static_cast<ContextImpl*>(Thread->CTX)->ParentThread || StartPaused || Thread->StartPaused) {
// Parent thread doesn't need to wait to run
Thread->StartRunning.Wait();
}
@@ -1146,7 +1149,7 @@ namespace FEXCore::Context {
Thread->RunningEvents.Running = true;
Thread->CTX->Dispatcher->ExecuteDispatch(Thread->CurrentFrame);
static_cast<ContextImpl*>(Thread->CTX)->Dispatcher->ExecuteDispatch(Thread->CurrentFrame);
Thread->RunningEvents.Running = false;
}
@@ -1172,10 +1175,11 @@ namespace FEXCore::Context {
--IdleWaitRefCount;
IdleWaitCV.notify_all();
Alloc::OSAllocator::UninstallTLSData(Thread);
SignalDelegation->UninstallTLSState(Thread);
// If the parent thread is waiting to join, then we can't destroy our thread object
if (!Thread->DestroyedByParent && Thread != Thread->CTX->ParentThread) {
if (!Thread->DestroyedByParent && Thread != static_cast<ContextImpl*>(Thread->CTX)->ParentThread) {
Thread->CTX->DestroyThread(Thread);
}
}
@@ -1188,36 +1192,43 @@ namespace FEXCore::Context {
for (auto it = lower; it != upper; it++) {
for (auto Address: it->second) {
Context::ThreadRemoveCodeEntry(Thread, Address);
ContextImpl::ThreadRemoveCodeEntry(Thread, Address);
}
it->second.clear();
}
}
static void InvalidateGuestCodeRangeInternal(FEXCore::Context::Context *CTX, uint64_t Start, uint64_t Length) {
std::lock_guard lk(CTX->ThreadCreationMutex);
static void InvalidateGuestCodeRangeInternal(ContextImpl *CTX, uint64_t Start, uint64_t Length) {
std::lock_guard lk(static_cast<ContextImpl*>(CTX)->ThreadCreationMutex);
for (auto &Thread : CTX->Threads) {
for (auto &Thread : static_cast<ContextImpl*>(CTX)->Threads) {
InvalidateGuestThreadCodeRange(Thread, Start, Length);
}
}
void InvalidateGuestCodeRange(FEXCore::Context::Context *CTX, uint64_t Start, uint64_t Length) {
FHU::ScopedSignalMaskWithUniqueLock CodeInvalidationLock(CTX->CodeInvalidationMutex);
void ContextImpl::InvalidateGuestCodeRange(FEXCore::Core::InternalThreadState *Thread, uint64_t Start, uint64_t Length) {
// Potential deferred since Thread might not be valid.
// Thread object isn't valid very early in frontend's initialization.
// To be more optimal the frontend should provide this code with a valid Thread object earlier.
ScopedPotentialDeferredSignalWithUniqueLock CodeInvalidationLock(CodeInvalidationMutex, Thread);
InvalidateGuestCodeRangeInternal(CTX, Start, Length);
InvalidateGuestCodeRangeInternal(this, Start, Length);
}
void InvalidateGuestCodeRange(FEXCore::Context::Context *CTX, uint64_t Start, uint64_t Length, std::function<void(uint64_t start, uint64_t Length)> CallAfter) {
FHU::ScopedSignalMaskWithUniqueLock CodeInvalidationLock(CTX->CodeInvalidationMutex);
void ContextImpl::InvalidateGuestCodeRange(FEXCore::Core::InternalThreadState *Thread, uint64_t Start, uint64_t Length, std::function<void(uint64_t start, uint64_t Length)> CallAfter) {
// Potential deferred since Thread might not be valid.
// Thread object isn't valid very early in frontend's initialization.
// To be more optimal the frontend should provide this code with a valid Thread object earlier.
ScopedPotentialDeferredSignalWithUniqueLock CodeInvalidationLock(CodeInvalidationMutex, Thread);
InvalidateGuestCodeRangeInternal(CTX, Start, Length);
InvalidateGuestCodeRangeInternal(this, Start, Length);
CallAfter(Start, Length);
}
void Context::MarkMemoryShared() {
void ContextImpl::MarkMemoryShared() {
if (!IsMemoryShared) {
IsMemoryShared = true;
UpdateAtomicTSOEmulationConfig();
if (Config.TSOAutoMigration) {
std::lock_guard<std::mutex> lkThreads(ThreadCreationMutex);
@@ -1235,18 +1246,14 @@ namespace FEXCore::Context {
}
}
void MarkMemoryShared(FEXCore::Context::Context *CTX) {
CTX->MarkMemoryShared();
}
void Context::ThreadAddBlockLink(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestDestination, uintptr_t HostLink, const std::function<void()> &delinker) {
std::shared_lock lk(Thread->CTX->CodeInvalidationMutex);
void ContextImpl::ThreadAddBlockLink(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestDestination, uintptr_t HostLink, const std::function<void()> &delinker) {
std::shared_lock lk(static_cast<ContextImpl*>(Thread->CTX)->CodeInvalidationMutex);
Thread->LookupCache->AddBlockLink(GuestDestination, HostLink, delinker);
}
void Context::ThreadRemoveCodeEntry(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP) {
LogMan::Throw::AFmt(Thread->CTX->CodeInvalidationMutex.try_lock() == false, "CodeInvalidationMutex needs to be unique_locked here");
void ContextImpl::ThreadRemoveCodeEntry(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP) {
LogMan::Throw::AFmt(static_cast<ContextImpl*>(Thread->CTX)->CodeInvalidationMutex.try_lock() == false, "CodeInvalidationMutex needs to be unique_locked here");
std::lock_guard<std::recursive_mutex> lk(Thread->LookupCache->WriteLock);
@@ -1254,7 +1261,7 @@ namespace FEXCore::Context {
Thread->LookupCache->Erase(GuestRIP);
}
CustomIRResult Context::AddCustomIREntrypoint(uintptr_t Entrypoint, std::function<void(uintptr_t Entrypoint, FEXCore::IR::IREmitter *)> Handler, void *Creator, void *Data) {
CustomIRResult ContextImpl::AddCustomIREntrypoint(uintptr_t Entrypoint, std::function<void(uintptr_t Entrypoint, FEXCore::IR::IREmitter *)> Handler, void *Creator, void *Data) {
LOGMAN_THROW_A_FMT(Config.Is64BitMode || !(Entrypoint >> 32), "64-bit Entrypoint in 32-bit mode {:x}", Entrypoint);
std::unique_lock lk(CustomIRMutex);
@@ -1270,66 +1277,23 @@ namespace FEXCore::Context {
}
}
void Context::RemoveCustomIREntrypoint(uintptr_t Entrypoint) {
void ContextImpl::RemoveCustomIREntrypoint(uintptr_t Entrypoint) {
LOGMAN_THROW_A_FMT(Config.Is64BitMode || !(Entrypoint >> 32), "64-bit Entrypoint in 32-bit mode {:x}", Entrypoint);
std::scoped_lock lk(CustomIRMutex);
InvalidateGuestCodeRange(this, Entrypoint, 1, [this](uint64_t Entrypoint, uint64_t) {
InvalidateGuestCodeRange(nullptr, Entrypoint, 1, [this](uint64_t Entrypoint, uint64_t) {
CustomIRHandlers.erase(Entrypoint);
});
}
// Debug interface
void Context::CompileRIP(FEXCore::Core::InternalThreadState *Thread, uint64_t RIP) {
uint64_t RIPBackup = Thread->CurrentFrame->State.rip;
Thread->CurrentFrame->State.rip = RIP;
// Erase the RIP from all the storage backings if it exists
ThreadRemoveCodeEntry(Thread, RIP);
// We don't care if compilation passes or not
CompileBlock(Thread->CurrentFrame, RIP);
Thread->CurrentFrame->State.rip = RIPBackup;
}
uint64_t Context::GetThreadCount() const {
return Threads.size();
}
FEXCore::Core::RuntimeStats *Context::GetRuntimeStatsForThread(uint64_t Thread) {
return &Threads[Thread]->Stats;
}
bool Context::GetDebugDataForRIP(uint64_t RIP, FEXCore::Core::DebugData *Data) {
std::lock_guard<std::recursive_mutex> lk(ParentThread->LookupCache->WriteLock);
auto it = ParentThread->DebugStore.find(RIP);
if (it == ParentThread->DebugStore.end()) {
return false;
}
memcpy(Data, it->second.DebugData.get(), sizeof(FEXCore::Core::DebugData));
return true;
}
bool Context::FindHostCodeForRIP(uint64_t RIP, uint8_t **Code) {
uintptr_t HostCode = ParentThread->LookupCache->FindBlock(RIP);
if (!HostCode) {
return false;
}
*Code = reinterpret_cast<uint8_t*>(HostCode);
return true;
}
uint64_t HandleSyscall(FEXCore::HLE::SyscallHandler *Handler, FEXCore::Core::CpuStateFrame *Frame, FEXCore::HLE::SyscallArguments *Args) {
uint64_t Result{};
Result = Handler->HandleSyscall(Frame, Args);
return Result;
}
IR::AOTIRCacheEntry *Context::LoadAOTIRCacheEntry(const std::string &filename) {
IR::AOTIRCacheEntry *ContextImpl::LoadAOTIRCacheEntry(const fextl::string &filename) {
auto rv = IRCaptureCache.LoadAOTIRCacheEntry(filename);
if (DebugServer) {
DebugServer->AlertLibrariesChanged();
@@ -1337,19 +1301,20 @@ namespace FEXCore::Context {
return rv;
}
void Context::UnloadAOTIRCacheEntry(IR::AOTIRCacheEntry *Entry) {
void ContextImpl::UnloadAOTIRCacheEntry(IR::AOTIRCacheEntry *Entry) {
IRCaptureCache.UnloadAOTIRCacheEntry(Entry);
if (DebugServer) {
DebugServer->AlertLibrariesChanged();
}
}
void Context::AppendThunkDefinitions(std::vector<FEXCore::IR::ThunkDefinition> const& Definitions) {
ThunkHandler->AppendThunkDefinitions(Definitions);
void ContextImpl::AppendThunkDefinitions(fextl::vector<FEXCore::IR::ThunkDefinition> const& Definitions) {
if (ThunkHandler) {
ThunkHandler->AppendThunkDefinitions(Definitions);
}
}
void ConfigureAOTGen(FEXCore::Core::InternalThreadState *Thread, std::set<uint64_t> *ExternalBranches, uint64_t SectionMaxAddress) {
void ContextImpl::ConfigureAOTGen(FEXCore::Core::InternalThreadState *Thread, fextl::set<uint64_t> *ExternalBranches, uint64_t SectionMaxAddress) {
Thread->FrontendDecoder->SetExternalBranches(ExternalBranches);
Thread->FrontendDecoder->SetSectionMaxAddress(SectionMaxAddress);
}
+3 -3
View File
@@ -5,7 +5,7 @@ namespace FEXCore {
}
namespace FEXCore::Context {
struct Context;
class ContextImpl;
}
namespace FEXCore::CPU {
@@ -17,7 +17,7 @@ namespace FEXCore::CPU {
*
* @return true if core was able to be create
*/
bool CreateCPUCore(FEXCore::Context::Context *CTX);
bool CreateCPUCore(FEXCore::Context::ContextImpl *CTX);
bool LoadCode(FEXCore::Context::Context *CTX, FEXCore::CodeLoader *Loader);
bool LoadCode(FEXCore::Context::ContextImpl *CTX, FEXCore::CodeLoader *Loader);
}
@@ -1,7 +1,6 @@
#include "Interface/Core/ArchHelpers/CodeEmitter/Emitter.h"
#include "Interface/Core/LookupCache.h"
#include "Interface/Core/ArchHelpers/MContext.h"
#include "Interface/Core/Dispatcher/Arm64Dispatcher.h"
#include "Interface/Core/Interpreter/InterpreterClass.h"
@@ -12,6 +11,7 @@
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Core/X86Enums.h>
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXCore/fextl/memory.h>
#include <FEXHeaderUtils/Syscalls.h>
#include <array>
@@ -28,14 +28,13 @@
#include <code-buffer-vixl.h>
#include <platform-vixl.h>
#include <sys/syscall.h>
#include <unistd.h>
namespace FEXCore::CPU {
constexpr size_t MAX_DISPATCHER_CODE_SIZE = 4096;
Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, const DispatcherConfig &config)
Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::ContextImpl *ctx, const DispatcherConfig &config)
: FEXCore::CPU::Dispatcher(ctx, config), Arm64Emitter(ctx, MAX_DISPATCHER_CODE_SIZE)
#ifdef VIXL_SIMULATOR
, Simulator {&Decoder}
@@ -130,13 +129,6 @@ void Arm64Dispatcher::EmitDispatcher() {
and_(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r3, RipReg.R(), ARMEmitter::Reg::r3);
}
#ifdef VIXL_SIMULATOR
// VIXL simulator can't run syscalls.
constexpr bool SignalSafeCompile = false;
#else
constexpr bool SignalSafeCompile = true;
#endif
ARMEmitter::ForwardLabel NoBlock;
{
@@ -184,7 +176,7 @@ void Arm64Dispatcher::EmitDispatcher() {
{
ThreadStopHandlerAddressSpillSRA = GetCursorAddress<uint64_t>();
if (config.StaticRegisterAllocation)
SpillStaticRegs();
SpillStaticRegs(TMP1);
ThreadStopHandlerAddress = GetCursorAddress<uint64_t>();
@@ -198,26 +190,11 @@ void Arm64Dispatcher::EmitDispatcher() {
{
ExitFunctionLinkerAddress = GetCursorAddress<uint64_t>();
if (config.StaticRegisterAllocation)
SpillStaticRegs();
SpillStaticRegs(TMP1);
if (SignalSafeCompile) {
// When compiling code, mask all signals to reduce the chance of reentrant allocations
// Args:
// X0: SETMASK
// X1: Pointer to mask value (uint64_t)
// X2: Pointer to old mask value (uint64_t)
// X3: Size of mask, sizeof(uint64_t)
// X8: Syscall
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, ~0ULL);
stp<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::x0, ARMEmitter::XReg::x0, ARMEmitter::Reg::rsp, -16);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, SIG_SETMASK);
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r1, ARMEmitter::Reg::rsp, 0);
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r2, ARMEmitter::Reg::rsp, 0);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r3, 8);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r8, SYS_rt_sigprocmask);
svc(0);
}
ldr(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
add(ARMEmitter::Size::i64Bit, ARMEmitter::XReg::x0, ARMEmitter::XReg::x0, 1);
str(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
mov(ARMEmitter::XReg::x0, STATE);
mov(ARMEmitter::XReg::x1, ARMEmitter::XReg::lr);
@@ -229,26 +206,17 @@ void Arm64Dispatcher::EmitDispatcher() {
blr(ARMEmitter::Reg::r2);
#endif
if (SignalSafeCompile) {
// Now restore the signal mask
// Living in the same location
mov(ARMEmitter::XReg::x4, ARMEmitter::XReg::x0);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, SIG_SETMASK);
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r1, ARMEmitter::Reg::rsp, 0);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r2, 0);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r3, 8);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r8, SYS_rt_sigprocmask);
svc(0);
// Bring stack back
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, ARMEmitter::Reg::rsp, 16);
mov(ARMEmitter::XReg::x0, ARMEmitter::XReg::x4);
}
if (config.StaticRegisterAllocation)
FillStaticRegs();
ldr(ARMEmitter::XReg::x1, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
subs(ARMEmitter::Size::i64Bit, ARMEmitter::XReg::x1, ARMEmitter::XReg::x1, 1);
str(ARMEmitter::XReg::x1, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
// Trigger segfault if any deferred signals are pending
ldr(ARMEmitter::XReg::x1, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalFaultAddress));
str(ARMEmitter::XReg::zr, ARMEmitter::XReg::x1, 0);
br(ARMEmitter::Reg::r0);
}
@@ -257,29 +225,11 @@ void Arm64Dispatcher::EmitDispatcher() {
Bind(&NoBlock);
if (config.StaticRegisterAllocation)
SpillStaticRegs();
SpillStaticRegs(TMP1);
if (SignalSafeCompile) {
// When compiling code, mask all signals to reduce the chance of reentrant allocations
// Args:
// X0: SETMASK
// X1: Pointer to mask value (uint64_t)
// X2: Pointer to old mask value (uint64_t)
// X3: Size of mask, sizeof(uint64_t)
// X8: Syscall
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, ~0ULL);
stp<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::x0, ARMEmitter::XReg::x2, ARMEmitter::Reg::rsp, -16);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, SIG_SETMASK);
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r1, ARMEmitter::Reg::rsp, 0);
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r2, ARMEmitter::Reg::rsp, 0);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r3, 8);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r8, SYS_rt_sigprocmask);
svc(0);
// Reload x2 to bring back RIP
ldr(ARMEmitter::XReg::x2, ARMEmitter::Reg::rsp, 8);
}
ldr(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
add(ARMEmitter::Size::i64Bit, ARMEmitter::XReg::x0, ARMEmitter::XReg::x0, 1);
str(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
ldr(ARMEmitter::XReg::x0, &l_CTX);
mov(ARMEmitter::XReg::x1, STATE);
@@ -292,23 +242,17 @@ void Arm64Dispatcher::EmitDispatcher() {
blr(ARMEmitter::Reg::r3); // { CTX, Frame, RIP}
#endif
if (SignalSafeCompile) {
// Now restore the signal mask
// Living in the same location
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, SIG_SETMASK);
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r1, ARMEmitter::Reg::rsp, 0);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r2, 0);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r3, 8);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r8, SYS_rt_sigprocmask);
svc(0);
// Bring stack back
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, ARMEmitter::Reg::rsp, 16);
}
if (config.StaticRegisterAllocation)
FillStaticRegs();
ldr(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
subs(ARMEmitter::Size::i64Bit, ARMEmitter::XReg::x0, ARMEmitter::XReg::x0, 1);
str(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
// Trigger segfault if any deferred signals are pending
ldr(TMP1, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalFaultAddress));
str(ARMEmitter::XReg::zr, TMP1, 0);
b(&LoopTop);
}
@@ -334,7 +278,7 @@ void Arm64Dispatcher::EmitDispatcher() {
GuestSignal_SIGILL = GetCursorAddress<uint64_t>();
if (config.StaticRegisterAllocation)
SpillStaticRegs();
SpillStaticRegs(TMP1);
hlt(0);
}
@@ -345,7 +289,7 @@ void Arm64Dispatcher::EmitDispatcher() {
GuestSignal_SIGTRAP = GetCursorAddress<uint64_t>();
if (config.StaticRegisterAllocation)
SpillStaticRegs();
SpillStaticRegs(TMP1);
brk(0);
}
@@ -356,20 +300,28 @@ void Arm64Dispatcher::EmitDispatcher() {
GuestSignal_SIGSEGV = GetCursorAddress<uint64_t>();
if (config.StaticRegisterAllocation)
SpillStaticRegs();
SpillStaticRegs(TMP1);
// hlt/udf = SIGILL
// brk = SIGTRAP
// ??? = SIGSEGV
// Force a SIGSEGV by loading zero
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r1, 0);
ldr(ARMEmitter::XReg::x1, ARMEmitter::Reg::r1);
if (CTX->ExitOnHLTEnabled()) {
ldr(ARMEmitter::XReg::x0, STATE_PTR(CpuStateFrame, ReturningStackLocation));
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, ARMEmitter::Reg::r0, 0);
PopCalleeSavedRegisters();
ret();
}
else {
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r1, 0);
ldr(ARMEmitter::XReg::x1, ARMEmitter::Reg::r1);
}
}
{
ThreadPauseHandlerAddressSpillSRA = GetCursorAddress<uint64_t>();
if (config.StaticRegisterAllocation)
SpillStaticRegs();
SpillStaticRegs(TMP1);
ThreadPauseHandlerAddress = GetCursorAddress<uint64_t>();
// We are pausing, this means the frontend should be waiting for this thread to idle
@@ -446,7 +398,7 @@ void Arm64Dispatcher::EmitDispatcher() {
LUDIVHandlerAddress = GetCursorAddress<uint64_t>();
PushDynamicRegsAndLR(ARMEmitter::Reg::r3);
SpillStaticRegs();
SpillStaticRegs(ARMEmitter::Reg::r3);
ldr(ARMEmitter::XReg::x3, STATE_PTR(CpuStateFrame, Pointers.AArch64.LUDIV));
#ifdef VIXL_SIMULATOR
@@ -468,7 +420,7 @@ void Arm64Dispatcher::EmitDispatcher() {
LDIVHandlerAddress = GetCursorAddress<uint64_t>();
PushDynamicRegsAndLR(ARMEmitter::Reg::r3);
SpillStaticRegs();
SpillStaticRegs(ARMEmitter::Reg::r3);
ldr(ARMEmitter::XReg::x3, STATE_PTR(CpuStateFrame, Pointers.AArch64.LDIV));
#ifdef VIXL_SIMULATOR
@@ -490,7 +442,7 @@ void Arm64Dispatcher::EmitDispatcher() {
LUREMHandlerAddress = GetCursorAddress<uint64_t>();
PushDynamicRegsAndLR(ARMEmitter::Reg::r3);
SpillStaticRegs();
SpillStaticRegs(ARMEmitter::Reg::r3);
ldr(ARMEmitter::XReg::x3, STATE_PTR(CpuStateFrame, Pointers.AArch64.LUREM));
#ifdef VIXL_SIMULATOR
@@ -512,7 +464,7 @@ void Arm64Dispatcher::EmitDispatcher() {
LREMHandlerAddress = GetCursorAddress<uint64_t>();
PushDynamicRegsAndLR(ARMEmitter::Reg::r3);
SpillStaticRegs();
SpillStaticRegs(ARMEmitter::Reg::r3);
ldr(ARMEmitter::XReg::x3, STATE_PTR(CpuStateFrame, Pointers.AArch64.LREM));
@@ -543,7 +495,7 @@ void Arm64Dispatcher::EmitDispatcher() {
ClearICache(reinterpret_cast<void*>(DispatchPtr), End - reinterpret_cast<uint64_t>(DispatchPtr));
if (CTX->Config.BlockJITNaming()) {
std::string Name = "Dispatch_" + std::to_string(FHU::Syscalls::gettid());
fextl::string Name = fextl::fmt::format("Dispatch_{}", FHU::Syscalls::gettid());
CTX->Symbols.Register(reinterpret_cast<void*>(DispatchPtr), End - reinterpret_cast<uint64_t>(DispatchPtr), Name);
}
if (CTX->Config.GlobalJITNaming()) {
@@ -578,10 +530,10 @@ size_t Arm64Dispatcher::GenerateGDBPauseCheck(uint8_t *CodeBuffer, uint64_t Gues
// If we have a gdb server running then run in a less efficient mode that checks if we need to exit
// This happens when single stepping
static_assert(sizeof(FEXCore::Context::Context::Config.RunningMode) == 4, "This is expected to be size of 4");
static_assert(sizeof(FEXCore::Context::ContextImpl::Config.RunningMode) == 4, "This is expected to be size of 4");
emit.ldr(ARMEmitter::XReg::x0, STATE_PTR(CpuStateFrame, Thread));
emit.ldr(ARMEmitter::XReg::x0, ARMEmitter::Reg::r0, offsetof(FEXCore::Core::InternalThreadState, CTX)); // Get Context
emit.ldr(ARMEmitter::WReg::w0, ARMEmitter::Reg::r0, offsetof(FEXCore::Context::Context, Config.RunningMode));
emit.ldr(ARMEmitter::WReg::w0, ARMEmitter::Reg::r0, offsetof(FEXCore::Context::ContextImpl, Config.RunningMode));
// If the value == 0 then we don't need to stop
emit.cbz(ARMEmitter::Size::i32Bit, ARMEmitter::Reg::r0, &RunBlock);
@@ -626,28 +578,6 @@ size_t Arm64Dispatcher::GenerateInterpreterTrampoline(uint8_t *CodeBuffer) {
return UsedBytes;
}
void Arm64Dispatcher::SpillSRA(FEXCore::Core::InternalThreadState *Thread, void *ucontext, uint32_t IgnoreMask) {
for (size_t i = 0; i < SRA64.size(); i++) {
if (IgnoreMask & (1U << SRA64[i].Idx())) {
// Skip this one, it's already spilled
continue;
}
Thread->CurrentFrame->State.gregs[i] = ArchHelpers::Context::GetArmReg(ucontext, SRA64[i].Idx());
}
if (EmitterCTX->HostFeatures.SupportsAVX) {
for (size_t i = 0; i < SRAFPR.size(); i++) {
auto FPR = ArchHelpers::Context::GetArmFPR(ucontext, SRAFPR[i].Idx());
memcpy(&Thread->CurrentFrame->State.xmm.avx.data[i][0], &FPR, sizeof(__uint128_t));
}
} else {
for (size_t i = 0; i < SRAFPR.size(); i++) {
auto FPR = ArchHelpers::Context::GetArmFPR(ucontext, SRAFPR[i].Idx());
memcpy(&Thread->CurrentFrame->State.xmm.sse.data[i][0], &FPR, sizeof(__uint128_t));
}
}
}
void Arm64Dispatcher::InitThreadPointers(FEXCore::Core::InternalThreadState *Thread) {
// Setup dispatcher specific pointers that need to be accessed from JIT code
{
@@ -672,8 +602,8 @@ void Arm64Dispatcher::InitThreadPointers(FEXCore::Core::InternalThreadState *Thr
}
}
std::unique_ptr<Dispatcher> Dispatcher::CreateArm64(FEXCore::Context::Context *CTX, const DispatcherConfig &Config) {
return std::make_unique<Arm64Dispatcher>(CTX, Config);
fextl::unique_ptr<Dispatcher> Dispatcher::CreateArm64(FEXCore::Context::ContextImpl *CTX, const DispatcherConfig &Config) {
return fextl::make_unique<Arm64Dispatcher>(CTX, Config);
}
}
@@ -7,10 +7,6 @@
#include <aarch64/simulator-aarch64.h>
#endif
namespace FEXCore::Context {
struct Context;
}
namespace FEXCore::Core {
struct InternalThreadState;
}
@@ -22,7 +18,7 @@ namespace FEXCore::CPU {
class Arm64Dispatcher final : public Dispatcher, public Arm64Emitter {
public:
Arm64Dispatcher(FEXCore::Context::Context *ctx, const DispatcherConfig &config);
Arm64Dispatcher(FEXCore::Context::ContextImpl *ctx, const DispatcherConfig &config);
void InitThreadPointers(FEXCore::Core::InternalThreadState *Thread) override;
size_t GenerateGDBPauseCheck(uint8_t *CodeBuffer, uint64_t GuestRIP) override;
size_t GenerateInterpreterTrampoline(uint8_t *CodeBuffer) override;
@@ -34,8 +30,25 @@ class Arm64Dispatcher final : public Dispatcher, public Arm64Emitter {
void EmitDispatcher();
protected:
void SpillSRA(FEXCore::Core::InternalThreadState *Thread, void *ucontext, uint32_t IgnoreMask) override;
uint16_t GetSRAGPRCount() const override {
return SRA64.size();
}
uint16_t GetSRAFPRCount() const override {
return SRAFPR.size();
}
void GetSRAGPRMapping(uint8_t Mapping[16]) const override {
for (size_t i = 0; i < SRA64.size(); ++i) {
Mapping[i] = SRA64[i].Idx();
}
}
void GetSRAFPRMapping(uint8_t Mapping[16]) const override {
for (size_t i = 0; i < SRAFPR.size(); ++i) {
Mapping[i] = SRAFPR[i].Idx();
}
}
private:
// Long division helpers
File diff suppressed because it is too large. Load diff
@@ -1,14 +1,13 @@
#pragma once
#include <FEXCore/Core/CPUBackend.h>
#include "Interface/Core/ArchHelpers/MContext.h"
#include <FEXCore/fextl/memory.h>
#include <cstdint>
#include <signal.h>
#include <stddef.h>
#include <stack>
#include <tuple>
#include <vector>
namespace FEXCore {
struct GuestSigAction;
@@ -20,7 +19,7 @@ struct InternalThreadState;
}
namespace FEXCore::Context {
struct Context;
class ContextImpl;
}
namespace FEXCore::CPU {
@@ -57,14 +56,6 @@ public:
uint64_t Start{};
uint64_t End{};
bool HandleGuestSignal(FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack);
bool HandleSIGILL(FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext);
bool HandleSignalPause(FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext);
bool IsAddressInDispatcher(uint64_t Address) const {
return Address >= Start && Address < End;
}
virtual void InitThreadPointers(FEXCore::Core::InternalThreadState *Thread) = 0;
// These are across all arches for now
@@ -74,8 +65,8 @@ public:
virtual size_t GenerateGDBPauseCheck(uint8_t *CodeBuffer, uint64_t GuestRIP) = 0;
virtual size_t GenerateInterpreterTrampoline(uint8_t *CodeBuffer) = 0;
static std::unique_ptr<Dispatcher> CreateX86(FEXCore::Context::Context *CTX, const DispatcherConfig &Config);
static std::unique_ptr<Dispatcher> CreateArm64(FEXCore::Context::Context *CTX, const DispatcherConfig &Config);
static fextl::unique_ptr<Dispatcher> CreateX86(FEXCore::Context::ContextImpl *CTX, const DispatcherConfig &Config);
static fextl::unique_ptr<Dispatcher> CreateArm64(FEXCore::Context::ContextImpl *CTX, const DispatcherConfig &Config);
virtual void ExecuteDispatch(FEXCore::Core::CpuStateFrame *Frame) {
DispatchPtr(Frame);
@@ -85,85 +76,32 @@ public:
CallbackPtr(Frame, RIP);
}
virtual uint16_t GetSRAGPRCount() const {
return 0U;
}
virtual uint16_t GetSRAFPRCount() const {
return 0U;
}
virtual void GetSRAGPRMapping(uint8_t Mapping[16]) const {
}
virtual void GetSRAFPRMapping(uint8_t Mapping[16]) const {
}
const DispatcherConfig& GetConfig() const { return config; }
protected:
Dispatcher(FEXCore::Context::Context *ctx, const DispatcherConfig &Config)
Dispatcher(FEXCore::Context::ContextImpl *ctx, const DispatcherConfig &Config)
: CTX {ctx}
, config {Config}
{}
void RestoreFrame_x64(ArchHelpers::Context::ContextBackup* Context, FEXCore::Core::CpuStateFrame *Frame, void *ucontext);
void RestoreFrame_ia32(ArchHelpers::Context::ContextBackup* Context, FEXCore::Core::CpuStateFrame *Frame, void *ucontext);
void RestoreRTFrame_ia32(ArchHelpers::Context::ContextBackup* Context, FEXCore::Core::CpuStateFrame *Frame, void *ucontext);
const bool incomplete_guest_restorer_support = false;
///< Setup the signal frame for x64.
uint64_t SetupFrame_x64(FEXCore::Core::InternalThreadState *Thread, ArchHelpers::Context::ContextBackup* ContextBackup, FEXCore::Core::CpuStateFrame *Frame,
int Signal, siginfo_t *HostSigInfo, void *ucontext,
GuestSigAction *GuestAction, stack_t *GuestStack,
uint64_t NewGuestSP, const uint32_t eflags);
///< Setup the signal frame for a 32-bit signal without SA_SIGINFO.
uint64_t SetupFrame_ia32(ArchHelpers::Context::ContextBackup* ContextBackup, FEXCore::Core::CpuStateFrame *Frame,
int Signal, siginfo_t *HostSigInfo, void *ucontext,
GuestSigAction *GuestAction, stack_t *GuestStack,
uint64_t NewGuestSP, const uint32_t eflags);
///< Setup the signal frame for a 32-bit signal with SA_SIGINFO.
uint64_t SetupRTFrame_ia32(ArchHelpers::Context::ContextBackup* ContextBackup, FEXCore::Core::CpuStateFrame *Frame,
int Signal, siginfo_t *HostSigInfo, void *ucontext,
GuestSigAction *GuestAction, stack_t *GuestStack,
uint64_t NewGuestSP, const uint32_t eflags);
ArchHelpers::Context::ContextBackup* StoreThreadState(FEXCore::Core::InternalThreadState *Thread, int Signal, void *ucontext);
enum class RestoreType {
TYPE_REALTIME, ///< Signal restore type is from a `realtime` signal.
TYPE_NONREALTIME, ///< Signal restore type is from a `non-realtime` signal.
TYPE_PAUSE, ///< Signal restore type is from a GDB pause event.
};
/*
* Signal frames on 32-bit architecture needs to match exactly how the kernel generates the frame.
* This is because large parts of the signal frame definition is part of the UAPI.
* This means that when FEX sets up the signal frame, it needs to match the UAPI stack setup.
*
* The two signal stack frame types below describe the two different 32-bit frame types.
*/
// The 32-bit non-realtime signal frame.
// This frame type is used when the guest signal is used without the `SA_SIGINFO` flag.
struct SigFrame_i32 {
uint32_t pretcode; ///< sigreturn return branch point.
int32_t Signal; ///< The signal hit.
FEXCore::x86::sigcontext sc; ///< The signal context.
x86::_libc_fpstate fpstate_unused; ///< Unused fpstate. Retained for backwards compatibility.
uint32_t extramask[1]; ///< Upper 32-bits of the signal mask. Lower 32-bits is in the sigcontext.
char retcode[8]; ///< Unused but needs to be filled. GDB seemingly uses as a debug marker.
///< FP state now follows after this.
};
// The 32-bit realtime signal frame.
// This frame type is used when the guest signal is used with the `SA_SIGINFO` flag.
struct RTSigFrame_i32 {
uint32_t pretcode; ///< sigreturn return branch point.
int32_t Signal; ///< The signal hit.
uint32_t pinfo; ///< Pointer to siginfo_t
uint32_t puc; ///< Pointer to ucontext_t
FEXCore::x86::siginfo_t info;
FEXCore::x86::ucontext_t uc;
char retcode[8]; ///< Unused but needs to be filled. GDB seemingly uses as a debug marker.
///< FP state now follows after this.
};
void RestoreThreadState(FEXCore::Core::InternalThreadState *Thread, void *ucontext, RestoreType Type);
std::stack<uint64_t, std::vector<uint64_t>> SignalFrames;
virtual void SpillSRA(FEXCore::Core::InternalThreadState *Thread, void *ucontext, uint32_t IgnoreMask) {}
FEXCore::Context::Context *CTX;
FEXCore::Context::ContextImpl *CTX;
DispatcherConfig config;
static void SleepThread(FEXCore::Context::Context *ctx, FEXCore::Core::CpuStateFrame *Frame);
static void SleepThread(FEXCore::Context::ContextImpl *ctx, FEXCore::Core::CpuStateFrame *Frame);
static uint64_t GetCompileBlockPtr();
@@ -1,3 +1,4 @@
#include "FEXCore/Utils/AllocatorHooks.h"
#include "Interface/Core/LookupCache.h"
#include "Interface/Core/Dispatcher/X86Dispatcher.h"
@@ -11,14 +12,14 @@
#include <FEXCore/Core/CPUBackend.h>
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXCore/Utils/Allocator.h>
#include <FEXCore/fextl/memory.h>
#include <FEXCore/fextl/string.h>
#include <FEXHeaderUtils/Syscalls.h>
#include <cmath>
#include <memory>
#include <stddef.h>
#include <stdint.h>
#include <sys/mman.h>
#include <xbyak/xbyak.h>
#define STATE_PTR(STATE_TYPE, FIELD) \
[STATE + offsetof(FEXCore::Core::STATE_TYPE, FIELD)]
@@ -27,10 +28,10 @@ namespace FEXCore::CPU {
static constexpr size_t MAX_DISPATCHER_CODE_SIZE = 4096;
#define STATE r14
X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, const DispatcherConfig &config)
X86Dispatcher::X86Dispatcher(FEXCore::Context::ContextImpl *ctx, const DispatcherConfig &config)
: Dispatcher(ctx, config)
, Xbyak::CodeGenerator(MAX_DISPATCHER_CODE_SIZE,
FEXCore::Allocator::mmap(nullptr, MAX_DISPATCHER_CODE_SIZE, PROT_READ | PROT_WRITE | PROT_EXEC, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0),
FEXCore::Allocator::VirtualAlloc(MAX_DISPATCHER_CODE_SIZE, true),
nullptr) {
LOGMAN_THROW_AA_FMT(!config.StaticRegisterAllocation, "X86 dispatcher does not support SRA");
@@ -169,36 +170,11 @@ X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, const DispatcherCon
ret();
}
constexpr bool SignalSafeCompile = true;
// Block creation
{
L(NoBlock);
if (SignalSafeCompile) {
// When compiling code, mask all signals to reduce the chance of reentrant allocations
// RDI: SETMASK
// RSI: Pointer to mask value (uint64_t)
// RDX: Pointer to old mask value (uint64_t)
// R10: Size of mask, sizeof(uint64_t)
// RAX: Syscall
// Backup rdx
mov(r9, rdx);
mov(rdi, ~0ULL);
sub(rsp, 16);
mov(qword [rsp], rdi);
mov(qword [rsp + 8], rdi);
mov(rdi, SIG_SETMASK);
mov(rsi, rsp);
mov(rdx, rsp);
mov(r10, 8);
mov(rax, SYS_rt_sigprocmask);
syscall();
mov(rdx, r9);
}
inc(qword [STATE + offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount)]);
// {rdi, rsi, rdx}
mov(rdi, reinterpret_cast<uint64_t>(CTX));
@@ -207,24 +183,15 @@ X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, const DispatcherCon
call(rax);
if (SignalSafeCompile) {
// Now restore the signal mask
// Living in the same location
// Backup rdx
mov(r9, rdx);
dec(qword [STATE + offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount)]);
mov(rdi, SIG_SETMASK);
mov(rsi, rsp);
mov(rdx, 0); // Don't care about result
mov(r10, 8);
mov(rax, SYS_rt_sigprocmask);
syscall();
Label AfterStore;
// Skip the deferred fault address if the refcount isn't zero
jne(AfterStore);
mov(rax, qword [STATE + offsetof(FEXCore::Core::CPUState, DeferredSignalFaultAddress)]);
mov(rax, qword [rax]);
// Bring stack back
add(rsp, 16);
mov(rdx, r9);
}
L(AfterStore);
// rdx already contains RIP here
jmp(LoopTop);
@@ -232,31 +199,7 @@ X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, const DispatcherCon
{
ExitFunctionLinkerAddress = getCurr<uint64_t>();
if (SignalSafeCompile) {
// When compiling code, mask all signals to reduce the chance of reentrant allocations
// RDI: SETMASK
// RSI: Pointer to mask value (uint64_t)
// RDX: Pointer to old mask value (uint64_t)
// R10: Size of mask, sizeof(uint64_t)
// RAX: Syscall
// Backup rax
mov(r9, rax);
mov(rdi, ~0ULL);
sub(rsp, 16);
mov(qword [rsp], rdi);
mov(qword [rsp + 8], rdi);
mov(rdi, SIG_SETMASK);
mov(rsi, rsp);
mov(rdx, rsp);
mov(r10, 8);
mov(rax, SYS_rt_sigprocmask);
syscall();
mov(rax, r9);
}
inc(qword [STATE + offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount)]);
// {rdi, rsi}
mov(rdi, STATE);
@@ -264,27 +207,17 @@ X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, const DispatcherCon
call(qword STATE_PTR(CpuStateFrame, Pointers.Common.ExitFunctionLink));
if (SignalSafeCompile) {
// Now restore the signal mask
// Living in the same location
// Backup rax
mov(r9, rax);
dec(qword [STATE + offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount)]);
mov(rdi, SIG_SETMASK);
mov(rsi, rsp);
mov(rdx, 0); // Don't care about result
mov(r10, 8);
mov(rax, SYS_rt_sigprocmask);
syscall();
Label AfterStore;
// Skip the deferred fault address if the refcount isn't zero
jne(AfterStore);
mov(rbx, qword [STATE + offsetof(FEXCore::Core::CPUState, DeferredSignalFaultAddress)]);
mov(qword [rbx], rbx);
// Bring stack back
add(rsp, 16);
L(AfterStore);
jmp(r9);
}
else {
jmp(rax);
}
jmp(rax);
}
{
@@ -407,7 +340,7 @@ X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, const DispatcherCon
End = Start + getSize();
if (CTX->Config.BlockJITNaming()) {
std::string Name = "Dispatch_" + std::to_string(FHU::Syscalls::gettid());
fextl::string Name = fextl::fmt::format("Dispatch_{}", FHU::Syscalls::gettid());
CTX->Symbols.Register(reinterpret_cast<void*>(Start), End-Start, Name);
}
if (CTX->Config.GlobalJITNaming()) {
@@ -416,15 +349,13 @@ X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, const DispatcherCon
}
// Used by GenerateGDBPauseCheck, GenerateInterpreterTrampoline
static thread_local Xbyak::CodeGenerator emit(1, &emit); // actual emit target set with setNewBuffer
size_t X86Dispatcher::GenerateGDBPauseCheck(uint8_t *CodeBuffer, uint64_t GuestRIP) {
using namespace Xbyak;
using namespace Xbyak::util;
Xbyak::CodeGenerator emit(1, &emit); // actual emit target set with setNewBuffer
emit.setNewBuffer(CodeBuffer, MaxGDBPauseCheckSize);
Label RunBlock;
// If we have a gdb server running then run in a less efficient mode that checks if we need to exit
@@ -433,7 +364,7 @@ size_t X86Dispatcher::GenerateGDBPauseCheck(uint8_t *CodeBuffer, uint64_t GuestR
emit.mov(rax, reinterpret_cast<uint64_t>(CTX));
// If the value == 0 then we don't need to stop
emit.cmp(dword [rax + (offsetof(FEXCore::Context::Context, Config.RunningMode))], 0);
emit.cmp(dword [rax + (offsetof(FEXCore::Context::ContextImpl, Config.RunningMode))], 0);
emit.je(RunBlock);
{
// Make sure RIP is syncronized to the context
@@ -457,10 +388,11 @@ size_t X86Dispatcher::GenerateInterpreterTrampoline(uint8_t *CodeBuffer) {
using namespace Xbyak;
using namespace Xbyak::util;
Xbyak::CodeGenerator emit(1, &emit); // actual emit target set with setNewBuffer
emit.setNewBuffer(CodeBuffer, MaxInterpreterTrampolineSize);
Label InlineIRData;
emit.mov(rdi, STATE);
emit.lea(rsi, ptr[rip + InlineIRData]);
emit.call(qword STATE_PTR(CpuStateFrame, Pointers.Interpreter.FragmentExecuter));
@@ -475,7 +407,7 @@ size_t X86Dispatcher::GenerateInterpreterTrampoline(uint8_t *CodeBuffer) {
}
X86Dispatcher::~X86Dispatcher() {
FEXCore::Allocator::munmap(top_, MAX_DISPATCHER_CODE_SIZE);
FEXCore::Allocator::VirtualFree(top_, MAX_DISPATCHER_CODE_SIZE);
}
void X86Dispatcher::InitThreadPointers(FEXCore::Core::InternalThreadState *Thread) {
@@ -499,8 +431,8 @@ void X86Dispatcher::InitThreadPointers(FEXCore::Core::InternalThreadState *Threa
}
}
std::unique_ptr<Dispatcher> Dispatcher::CreateX86(FEXCore::Context::Context *CTX, const DispatcherConfig &Config) {
return std::make_unique<X86Dispatcher>(CTX, Config);
fextl::unique_ptr<Dispatcher> Dispatcher::CreateX86(FEXCore::Context::ContextImpl *CTX, const DispatcherConfig &Config) {
return fextl::make_unique<X86Dispatcher>(CTX, Config);
}
}
@@ -1,13 +1,24 @@
#pragma once
#include <FEXCore/fextl/list.h>
#include <FEXCore/fextl/unordered_map.h>
#include <FEXCore/fextl/unordered_set.h>
#include "Interface/Core/Dispatcher/Dispatcher.h"
#define XBYAK64
#include <xbyak/xbyak.h>
#define XBYAK_CUSTOM_ALLOC
#define XBYAK_CUSTOM_MALLOC FEXCore::Allocator::malloc
#define XBYAK_CUSTOM_FREE FEXCore::Allocator::free
#define XBYAK_CUSTOM_SETS
#define XBYAK_STD_UNORDERED_SET fextl::unordered_set
#define XBYAK_STD_UNORDERED_MAP fextl::unordered_map
#define XBYAK_STD_UNORDERED_MULTIMAP fextl::unordered_multimap
#define XBYAK_STD_LIST fextl::list
#define XBYAK_NO_EXCEPTION
namespace FEXCore::Context {
struct Context;
}
#include <xbyak/xbyak.h>
#include <xbyak/xbyak_util.h>
namespace FEXCore::Core {
struct InternalThreadState;
@@ -17,7 +28,7 @@ namespace FEXCore::CPU {
class X86Dispatcher final : public Dispatcher, public Xbyak::CodeGenerator {
public:
X86Dispatcher(FEXCore::Context::Context *ctx, const DispatcherConfig &config);
X86Dispatcher(FEXCore::Context::ContextImpl *ctx, const DispatcherConfig &config);
void InitThreadPointers(FEXCore::Core::InternalThreadState *Thread) override;
size_t GenerateGDBPauseCheck(uint8_t *CodeBuffer, uint64_t GuestRIP) override;
size_t GenerateInterpreterTrampoline(uint8_t *CodeBuffer) override;
+26 -22
View File
@@ -7,6 +7,7 @@ $end_info$
#include "Interface/Context/Context.h"
#include "Interface/Core/Frontend.h"
#include "Interface/Core/X86Tables/X86Tables.h"
#include <array>
#include <assert.h>
@@ -14,15 +15,13 @@ $end_info$
#include <cstring>
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/X86Enums.h>
#include <FEXCore/Debug/X86Tables.h>
#include <FEXCore/HLE/SyscallHandler.h>
#include <FEXCore/Utils/Allocator.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/Profiler.h>
#include <FEXCore/Utils/Telemetry.h>
#include <FEXCore/fextl/set.h>
#include <FEXHeaderUtils/TypeDefines.h>
#include <set>
#include <sys/mman.h>
namespace FEXCore::Frontend {
#include "Interface/Core/VSyscall/VSyscall.inc"
@@ -79,7 +78,7 @@ static uint32_t MapVEXToReg(uint8_t vvvv, bool HasXMM) {
}
}
Decoder::Decoder(FEXCore::Context::Context *ctx)
Decoder::Decoder(FEXCore::Context::ContextImpl *ctx)
: CTX {ctx}
, OSABI { ctx->SyscallHandler ? ctx->SyscallHandler->GetOSABI() : FEXCore::HLE::SyscallOSABI::OS_UNKNOWN }
, PoolObject {ctx->FrontendAllocator, sizeof(FEXCore::X86Tables::DecodedInst) * DefaultDecodedBufferSize} {
@@ -304,24 +303,24 @@ bool Decoder::NormalOp(FEXCore::X86Tables::X86InstInfo const *Info, uint16_t Op,
(Options.w && CTX->Config.Is64BitMode);
const bool HasNarrowingDisplacement = (FEXCore::X86Tables::DecodeFlags::GetOpAddr(DecodeInst->Flags, 0) & FEXCore::X86Tables::DecodeFlags::FLAG_OPERAND_SIZE_LAST) != 0;
bool HasXMMSrc = !!(Info->Flags & FEXCore::X86Tables::InstFlags::FLAGS_XMM_FLAGS) &&
!HAS_XMM_SUBFLAG(Info->Flags, FEXCore::X86Tables::InstFlags::FLAGS_SF_SRC_GPR) &&
!HAS_XMM_SUBFLAG(Info->Flags, FEXCore::X86Tables::InstFlags::FLAGS_SF_MMX_SRC);
bool HasXMMDst = !!(Info->Flags & FEXCore::X86Tables::InstFlags::FLAGS_XMM_FLAGS) &&
!HAS_XMM_SUBFLAG(Info->Flags, FEXCore::X86Tables::InstFlags::FLAGS_SF_DST_GPR) &&
!HAS_XMM_SUBFLAG(Info->Flags, FEXCore::X86Tables::InstFlags::FLAGS_SF_MMX_DST);
bool HasMMSrc = !!(Info->Flags & FEXCore::X86Tables::InstFlags::FLAGS_XMM_FLAGS) &&
!HAS_XMM_SUBFLAG(Info->Flags, FEXCore::X86Tables::InstFlags::FLAGS_SF_SRC_GPR) &&
HAS_XMM_SUBFLAG(Info->Flags, FEXCore::X86Tables::InstFlags::FLAGS_SF_MMX_SRC);
bool HasMMDst = !!(Info->Flags & FEXCore::X86Tables::InstFlags::FLAGS_XMM_FLAGS) &&
!HAS_XMM_SUBFLAG(Info->Flags, FEXCore::X86Tables::InstFlags::FLAGS_SF_DST_GPR) &&
HAS_XMM_SUBFLAG(Info->Flags, FEXCore::X86Tables::InstFlags::FLAGS_SF_MMX_DST);
const bool HasXMMFlags = (Info->Flags & InstFlags::FLAGS_XMM_FLAGS) != 0;
bool HasXMMSrc = HasXMMFlags &&
!HAS_XMM_SUBFLAG(Info->Flags, InstFlags::FLAGS_SF_SRC_GPR) &&
!HAS_XMM_SUBFLAG(Info->Flags, InstFlags::FLAGS_SF_MMX_SRC);
bool HasXMMDst = HasXMMFlags &&
!HAS_XMM_SUBFLAG(Info->Flags, InstFlags::FLAGS_SF_DST_GPR) &&
!HAS_XMM_SUBFLAG(Info->Flags, InstFlags::FLAGS_SF_MMX_DST);
bool HasMMSrc = HasXMMFlags &&
!HAS_XMM_SUBFLAG(Info->Flags, InstFlags::FLAGS_SF_SRC_GPR) &&
HAS_XMM_SUBFLAG(Info->Flags, InstFlags::FLAGS_SF_MMX_SRC);
bool HasMMDst = HasXMMFlags &&
!HAS_XMM_SUBFLAG(Info->Flags, InstFlags::FLAGS_SF_DST_GPR) &&
HAS_XMM_SUBFLAG(Info->Flags, InstFlags::FLAGS_SF_MMX_DST);
// Is ModRM present via explicit instruction encoded or REX?
const bool HasMODRM = !!(Info->Flags & FEXCore::X86Tables::InstFlags::FLAGS_MODRM);
const bool HasREX = !!(DecodeInst->Flags & DecodeFlags::FLAG_REX_PREFIX);
const bool HasHighXMM = HAS_XMM_SUBFLAG(Info->Flags, FEXCore::X86Tables::InstFlags::FLAGS_SF_HIGH_XMM_REG);
const bool Has16BitAddressing = !CTX->Config.Is64BitMode &&
DecodeInst->Flags & DecodeFlags::FLAG_ADDRESS_SIZE;
@@ -444,7 +443,7 @@ bool Decoder::NormalOp(FEXCore::X86Tables::X86InstInfo const *Info, uint16_t Op,
// ADDITIONALLY:
// If there is a REX prefix then that allows extended GPR usage
CurrentDest->Type = DecodedOperand::OpType::GPR;
DecodeInst->Dest.Data.GPR.HighBits = (Is8BitDest && !HasREX && (Op & 0b111) >= 0b100) || HasHighXMM;
DecodeInst->Dest.Data.GPR.HighBits = (Is8BitDest && !HasREX && (Op & 0b111) >= 0b100);
CurrentDest->Data.GPR.GPR = MapModRMToReg(DecodeInst->Flags & DecodeFlags::FLAG_REX_XGPR_B ? 1 : 0, Op & 0b111, Is8BitDest, HasREX, false, false);
if (CurrentDest->Data.GPR.GPR == FEXCore::X86State::REG_INVALID)
@@ -472,7 +471,7 @@ bool Decoder::NormalOp(FEXCore::X86Tables::X86InstInfo const *Info, uint16_t Op,
// Decode the GPR source first
GPR.Type = DecodedOperand::OpType::GPR;
GPR.Data.GPR.HighBits = (GPR8Bit && ModRM.reg >= 0b100 && !HasREX) || HasHighXMM;
GPR.Data.GPR.HighBits = (GPR8Bit && ModRM.reg >= 0b100 && !HasREX);
GPR.Data.GPR.GPR = MapModRMToReg(DecodeInst->Flags & DecodeFlags::FLAG_REX_XGPR_R ? 1 : 0, ModRM.reg, GPR8Bit, HasREX, HasXMMGPR, HasMMGPR);
if (GPR.Data.GPR.GPR == FEXCore::X86State::REG_INVALID)
@@ -482,7 +481,7 @@ bool Decoder::NormalOp(FEXCore::X86Tables::X86InstInfo const *Info, uint16_t Op,
// ModRM.Mod != 0b11 == Register-direct addressing
if (ModRM.mod == 0b11) {
NonGPR.Type = DecodedOperand::OpType::GPR;
NonGPR.Data.GPR.HighBits = (NonGPR8Bit && ModRM.rm >= 0b100 && !HasREX) || HasHighXMM;
NonGPR.Data.GPR.HighBits = (NonGPR8Bit && ModRM.rm >= 0b100 && !HasREX);
NonGPR.Data.GPR.GPR = MapModRMToReg(DecodeInst->Flags & DecodeFlags::FLAG_REX_XGPR_B ? 1 : 0, ModRM.rm, NonGPR8Bit, HasREX, HasXMMNonGPR, HasMMNonGPR);
if (NonGPR.Data.GPR.GPR == FEXCore::X86State::REG_INVALID)
return false;
@@ -503,7 +502,12 @@ bool Decoder::NormalOp(FEXCore::X86Tables::X86InstInfo const *Info, uint16_t Op,
if ((Info->Flags & FEXCore::X86Tables::InstFlags::FLAGS_VEX_1ST_SRC) != 0) {
DecodeInst->Src[CurrentSrc].Type = DecodedOperand::OpType::GPR;
DecodeInst->Src[CurrentSrc].Data.GPR.HighBits = false;
DecodeInst->Src[CurrentSrc].Data.GPR.GPR = MapVEXToReg(Options.vvvv, HasXMMSrc);
// If we have XMM flags at all, then SRC 1 cannot be a GPR. The only case where
// this is possible is with BMI1 and BMI2 instructions (which are all GPR-based
// and don't use XMM flags)
DecodeInst->Src[CurrentSrc].Data.GPR.GPR = MapVEXToReg(Options.vvvv, HasXMMFlags);
++CurrentSrc;
}
@@ -1109,7 +1113,7 @@ void Decoder::DecodeInstructionsAtEntry(uint8_t const* _InstStream, uint64_t PC,
uint64_t CurrentCodePage = PC & FHU::FEX_PAGE_MASK;
std::set<uint64_t> CodePages = { CurrentCodePage };
fextl::set<uint64_t> CodePages = { CurrentCodePage };
AddContainedCodePage(PC, CurrentCodePage, FHU::FEX_PAGE_SIZE);
+13 -12
View File
@@ -1,17 +1,18 @@
#pragma once
#include <FEXCore/Debug/X86Tables.h>
#include "Interface/Core/X86Tables/X86Tables.h"
#include <FEXCore/HLE/SyscallHandler.h>
#include <FEXCore/Utils/Telemetry.h>
#include <FEXCore/fextl/set.h>
#include <FEXCore/fextl/vector.h>
#include <array>
#include <cstdint>
#include <set>
#include <stddef.h>
#include <vector>
namespace FEXCore::Context {
struct Context;
class ContextImpl;
}
namespace FEXCore::Frontend {
@@ -25,11 +26,11 @@ public:
bool HasInvalidInstruction{};
};
Decoder(FEXCore::Context::Context *ctx);
Decoder(FEXCore::Context::ContextImpl *ctx);
~Decoder();
void DecodeInstructionsAtEntry(uint8_t const* InstStream, uint64_t PC, std::function<void(uint64_t BlockEntry, uint64_t Start, uint64_t Length)> AddContainedCodePage);
std::vector<DecodedBlocks> const *GetDecodedBlocks() const {
fextl::vector<DecodedBlocks> const *GetDecodedBlocks() const {
return &Blocks;
}
@@ -37,7 +38,7 @@ public:
uint64_t DecodedMaxAddress {~0ULL};
void SetSectionMaxAddress(uint64_t v) { SectionMaxAddress = v; }
void SetExternalBranches(std::set<uint64_t> *v) { ExternalBranches = v; }
void SetExternalBranches(fextl::set<uint64_t> *v) { ExternalBranches = v; }
void DelayedDisownBuffer() {
PoolObject.DelayedDisownBuffer();
@@ -52,7 +53,7 @@ private:
bool L; // VEX.L bit (if set then 256 bit operation, if unset then scalar or 128-bit operation)
};
FEXCore::Context::Context *CTX;
FEXCore::Context::ContextImpl *CTX;
const FEXCore::HLE::SyscallOSABI OSABI{};
bool DecodeInstruction(uint64_t PC);
@@ -89,10 +90,10 @@ private:
uint64_t SymbolMinAddress {~0ULL};
uint64_t SectionMaxAddress {~0ULL};
std::vector<DecodedBlocks> Blocks;
std::set<uint64_t> BlocksToDecode;
std::set<uint64_t> HasBlocks;
std::set<uint64_t> *ExternalBranches {nullptr};
fextl::vector<DecodedBlocks> Blocks;
fextl::set<uint64_t> BlocksToDecode;
fextl::set<uint64_t> HasBlocks;
fextl::set<uint64_t> *ExternalBranches {nullptr};
// ModRM rm decoding
using DecodeModRMPtr = void (FEXCore::Frontend::Decoder::*)(X86Tables::DecodedOperand *Operand, X86Tables::ModRMDecoded ModRM);
+125 -119
View File
@@ -8,11 +8,8 @@ $end_info$
#include <cstdlib>
#include <cstdio>
#include <iomanip>
#include <sstream>
#include <string>
#include <memory>
#include <optional>
#include <vector>
#include "Common/SoftFloat.h"
#include "Common/StringUtils.h"
#include "Interface/Context/Context.h"
@@ -27,39 +24,45 @@ $end_info$
#include <FEXCore/HLE/Linux/ThreadManagement.h>
#include <FEXCore/HLE/SyscallHandler.h>
#include <FEXCore/Utils/CompilerDefs.h>
#include <FEXCore/Utils/FileLoading.h>
#include <FEXCore/Utils/NetStream.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/Threads.h>
#include <FEXCore/fextl/fmt.h>
#include <FEXCore/fextl/map.h>
#include <FEXCore/fextl/sstream.h>
#include <FEXCore/fextl/string.h>
#include <FEXCore/fextl/vector.h>
#include <atomic>
#include <cstring>
#ifndef _WIN32
#include <elf.h>
#include <netdb.h>
#include <sys/socket.h>
#endif
#include <errno.h>
#include <fcntl.h>
#include <fstream>
#include <fmt/format.h>
#include <netdb.h>
#include <signal.h>
#include <stddef.h>
#include <string_view>
#include <sys/socket.h>
#include <sys/stat.h>
#include <unistd.h>
#include <utility>
#include <vector>
#include "GdbServer.h"
namespace FEXCore
{
#ifndef _WIN32
void GdbServer::Break(int signal) {
std::lock_guard lk(sendMutex);
if (!CommsStream) {
return;
}
const auto str = fmt::format("S{:02x}", signal);
const fextl::string str = fextl::fmt::format("S{:02x}", signal);
SendPacket(*CommsStream, str);
}
@@ -68,11 +71,11 @@ void GdbServer::WaitForThreadWakeup() {
ThreadBreakEvent.Wait();
}
GdbServer::GdbServer(FEXCore::Context::Context *ctx) : CTX(ctx) {
GdbServer::GdbServer(FEXCore::Context::ContextImpl *ctx) : CTX(ctx) {
// Pass all signals by default
std::fill(PassSignals.begin(), PassSignals.end(), true);
Context::SetExitHandler(ctx, [this](uint64_t ThreadId, FEXCore::Context::ExitReason ExitReason) {
ctx->SetExitHandler([this](uint64_t ThreadId, FEXCore::Context::ExitReason ExitReason) {
if (ExitReason == FEXCore::Context::ExitReason::EXIT_DEBUG) {
this->Break(SIGTRAP);
}
@@ -101,7 +104,7 @@ GdbServer::GdbServer(FEXCore::Context::Context *ctx) : CTX(ctx) {
StartThread();
}
static int calculateChecksum(const std::string &packet) {
static int calculateChecksum(const fextl::string &packet) {
unsigned char checksum = 0;
for (const char &c : packet) {
checksum += c;
@@ -109,8 +112,8 @@ static int calculateChecksum(const std::string &packet) {
return checksum;
}
static std::string hexstring(std::istringstream &ss, int delm) {
std::string ret;
static fextl::string hexstring(fextl::istringstream &ss, int delm) {
fextl::string ret;
char hexString[3] = {0, 0, 0};
while (ss.peek() != delm) {
@@ -125,8 +128,8 @@ static std::string hexstring(std::istringstream &ss, int delm) {
return ret;
}
static std::string encodeHex(const unsigned char *data, size_t length) {
std::ostringstream ss;
static fextl::string encodeHex(const unsigned char *data, size_t length) {
fextl::ostringstream ss;
for (size_t i=0; i < length; i++) {
ss << std::setfill('0') << std::setw(2) << std::hex << int(data[i]);
@@ -134,26 +137,19 @@ static std::string encodeHex(const unsigned char *data, size_t length) {
return ss.str();
}
static std::string getThreadName(uint32_t ThreadID) {
const auto ThreadFile = fmt::format("/proc/{}/task/{}/comm", getpid(), ThreadID);
std::fstream fs(ThreadFile, std::fstream::in | std::fstream::binary);
if (fs.is_open()) {
std::string ThreadName;
fs >> ThreadName;
fs.close();
return ThreadName;
}
return "<No Name>";
static fextl::string getThreadName(uint32_t ThreadID) {
const auto ThreadFile = fextl::fmt::format("/proc/{}/task/{}/comm", getpid(), ThreadID);
fextl::string ThreadName {"<No Name>"};
FEXCore::FileLoading::LoadFile(ThreadName, ThreadFile);
return ThreadName;
}
// Packet parser
// Takes a serial stream and reads a single packet
// Un-escapes chars, checks the checksum and request a retransmit if it fails.
// Once the checksum is validated, it acknowledges and returns the packet in a string
std::string GdbServer::ReadPacket(std::iostream &stream) {
std::string packet{};
fextl::string GdbServer::ReadPacket(std::iostream &stream) {
fextl::string packet{};
// The GDB "Remote Serial Protocal" was originally 7bit clean for use on serial ports.
// Binary data is useally hex encoded. However some later extentions just put
@@ -172,7 +168,7 @@ std::string GdbServer::ReadPacket(std::iostream &stream) {
LogMan::Msg::EFmt("Dropping unexpected data: \"{}\"", packet);
// clear any existing data, must have been a mistake.
packet = std::string();
packet = fextl::string();
break;
case '}': // escape char
{
@@ -203,8 +199,8 @@ std::string GdbServer::ReadPacket(std::iostream &stream) {
return "";
}
static std::string escapePacket(const std::string& packet) {
std::ostringstream ss;
static fextl::string escapePacket(const fextl::string& packet) {
fextl::ostringstream ss;
for(const auto &c : packet) {
switch (c) {
@@ -225,9 +221,9 @@ static std::string escapePacket(const std::string& packet) {
return ss.str();
}
void GdbServer::SendPacket(std::ostream &stream, const std::string& packet) {
void GdbServer::SendPacket(std::ostream &stream, const fextl::string& packet) {
const auto escaped = escapePacket(packet);
const auto str = fmt::format("${}#{:02x}", escaped, calculateChecksum(escaped));
const auto str = fextl::fmt::format("${}#{:02x}", escaped, calculateChecksum(escaped));
stream << str << std::flush;
}
@@ -263,7 +259,7 @@ struct FEX_PACKED GDBContextDefinition {
uint32_t mxcsr;
};
std::string GdbServer::readRegs() {
fextl::string GdbServer::readRegs() {
GDBContextDefinition GDB{};
FEXCore::Core::CPUState state{};
@@ -311,11 +307,11 @@ std::string GdbServer::readRegs() {
return encodeHex((unsigned char *)&GDB, sizeof(GDBContextDefinition));
}
GdbServer::HandledPacketType GdbServer::readReg(const std::string& packet) {
size_t addr;
auto ss = std::istringstream(packet);
ss.get(); // Drop first letter
ss >> std::hex >> addr;
GdbServer::HandledPacketType GdbServer::readReg(const fextl::string& packet) {
size_t addr;
auto ss = fextl::istringstream(packet);
ss.get(); // Drop first letter
ss >> std::hex >> addr;
FEXCore::Core::CPUState state{};
@@ -395,8 +391,8 @@ GdbServer::HandledPacketType GdbServer::readReg(const std::string& packet) {
return {"E00", HandledPacketType::TYPE_ACK};
}
std::string buildTargetXML() {
std::ostringstream xml;
fextl::string buildTargetXML() {
fextl::ostringstream xml;
xml << "<?xml version='1.0'?>\n";
xml << "<!DOCTYPE target SYSTEM 'gdb-target.dtd'>\n";
@@ -448,7 +444,7 @@ std::string buildTargetXML() {
// x87 stack
for (int i=0; i < 8; i++) {
reg("st" + std::to_string(i), "i387_ext", 80);
reg(fextl::fmt::format("st{}", i), "i387_ext", 80);
}
// x87 control
@@ -484,7 +480,7 @@ std::string buildTargetXML() {
// SSE regs
for (size_t i = 0; i < Core::CPUState::NUM_XMMS; i++) {
reg("xmm" + std::to_string(i), "vec128", 128);
reg(fextl::fmt::format("xmm{}", i), "vec128", 128);
}
reg("mxcsr", "int", 32);
@@ -520,8 +516,8 @@ std::string buildTargetXML() {
return xml.str();
}
std::string buildOSData() {
std::ostringstream xml;
fextl::string buildOSData() {
fextl::ostringstream xml;
xml << "<?xml version='1.0'?>\n";
@@ -541,24 +537,27 @@ void GdbServer::buildLibraryMap() {
return;
}
std::ostringstream xml;
fextl::ostringstream xml;
std::fstream fs("/proc/self/maps", std::fstream::in | std::fstream::binary);
std::string Line;
fextl::string MapsFile;
FEXCore::FileLoading::LoadFile(MapsFile, "/proc/self/maps");
fextl::istringstream MapsStream(MapsFile);
fextl::string Line;
struct FileData {
uint64_t Begin;
};
std::map<std::string, std::vector<FileData>> SegmentMaps;
fextl::map<fextl::string, fextl::vector<FileData>> SegmentMaps;
// 7ff5dd6d2000-7ff5dd6d3000 rw-p 0000a000 103:0b 1881447 /usr/lib/x86_64-linux-gnu/libnss_compat.so.2
std::string const &RuntimeExecutable = Filename();
while (std::getline(fs, Line)) {
auto ss = std::istringstream(Line);
std::string Tmp;
std::string Begin;
std::string Name;
fextl::string const &RuntimeExecutable = Filename();
while (std::getline(MapsStream, Line)) {
auto ss = fextl::istringstream(Line);
fextl::string Tmp;
fextl::string Begin;
fextl::string Name;
std::getline(ss, Begin, '-');
std::getline(ss, Tmp, ' '); // End
std::getline(ss, Tmp, ' '); // Perm
@@ -609,18 +608,18 @@ void GdbServer::buildLibraryMap() {
LibraryMapChanged = false;
}
GdbServer::HandledPacketType GdbServer::handleXfer(const std::string &packet) {
std::string object;
std::string rw;
std::string annex;
GdbServer::HandledPacketType GdbServer::handleXfer(const fextl::string &packet) {
fextl::string object;
fextl::string rw;
fextl::string annex;
int annex_pid;
int offset;
int length;
// Parse Xfer message
{
auto ss = std::istringstream(packet);
std::string expectXfer;
auto ss = fextl::istringstream(packet);
fextl::string expectXfer;
char expectComma;
std::getline(ss, expectXfer, ':');
@@ -631,7 +630,7 @@ GdbServer::HandledPacketType GdbServer::handleXfer(const std::string &packet) {
annex_pid = getpid();
}
else {
auto ss_pid = std::istringstream(annex);
auto ss_pid = fextl::istringstream(annex);
ss_pid >> std::hex >> annex_pid;
}
ss >> std::hex >> offset;
@@ -644,7 +643,7 @@ GdbServer::HandledPacketType GdbServer::handleXfer(const std::string &packet) {
}
// Lambda to correctly encode any reply
auto encode = [&](std::string data) -> std::string {
auto encode = [&](fextl::string data) -> fextl::string {
if (offset == data.size())
return "l";
if (offset >= data.size())
@@ -674,7 +673,7 @@ GdbServer::HandledPacketType GdbServer::handleXfer(const std::string &packet) {
auto Threads = CTX->GetThreads();
ThreadString.clear();
std::ostringstream ss;
fextl::ostringstream ss;
ss << "<?xml version=\"1.0\"?>\n";
ss << "<threads>\n";
for (auto &Thread : *Threads) {
@@ -710,7 +709,7 @@ GdbServer::HandledPacketType GdbServer::handleXfer(const std::string &packet) {
auto CodeLoader = CTX->SyscallHandler->GetCodeLoader();
uint64_t auxv_ptr, auxv_size;
CodeLoader->GetAuxv(auxv_ptr, auxv_size);
std::string data;
fextl::string data;
if (CTX->Config.Is64BitMode) {
data.resize(auxv_size);
memcpy(data.data(), reinterpret_cast<void*>(auxv_ptr), data.size());
@@ -736,11 +735,14 @@ GdbServer::HandledPacketType GdbServer::handleXfer(const std::string &packet) {
static size_t CheckMemMapping(uint64_t Address, size_t Size) {
uint64_t AddressEnd = Address + Size;
std::fstream fs("/proc/self/maps", std::fstream::in | std::fstream::binary);
std::string Line;
fextl::string MapsFile;
FEXCore::FileLoading::LoadFile(MapsFile, "/proc/self/maps");
fextl::istringstream MapsStream(MapsFile);
while (std::getline(fs, Line)) {
if (fs.eof()) break;
fextl::string Line;
while (std::getline(MapsStream, Line)) {
if (MapsStream.eof()) break;
uint64_t Begin, End;
char r,w,x,p;
if (sscanf(Line.c_str(), "%lx-%lx %c%c%c%c", &Begin, &End, &r, &w, &x, &p) == 6) {
@@ -761,17 +763,17 @@ static size_t CheckMemMapping(uint64_t Address, size_t Size) {
GdbServer::HandledPacketType GdbServer::handleProgramOffsets() {
auto CodeLoader = CTX->SyscallHandler->GetCodeLoader();
uint64_t BaseOffset = CodeLoader->GetBaseOffset();
auto str = fmt::format("Text={:x};Data={:x};Bss={:x}", BaseOffset, BaseOffset, BaseOffset);
fextl::string str = fextl::fmt::format("Text={:x};Data={:x};Bss={:x}", BaseOffset, BaseOffset, BaseOffset);
return {std::move(str), HandledPacketType::TYPE_ACK};
}
GdbServer::HandledPacketType GdbServer::handleMemory(const std::string &packet) {
GdbServer::HandledPacketType GdbServer::handleMemory(const fextl::string &packet) {
bool write;
size_t addr;
size_t length;
std::string data;
fextl::string data;
auto ss = std::istringstream(packet);
auto ss = fextl::istringstream(packet);
write = ss.get() == 'M';
ss >> std::hex >> addr;
ss.get(); // discard comma
@@ -806,22 +808,22 @@ GdbServer::HandledPacketType GdbServer::handleMemory(const std::string &packet)
}
GdbServer::HandledPacketType GdbServer::handleQuery(const std::string &packet) {
GdbServer::HandledPacketType GdbServer::handleQuery(const fextl::string &packet) {
const auto match = [&](const char *str) -> bool { return packet.rfind(str, 0) == 0; };
const auto MatchStr = [](const std::string &Str, const char *str) -> bool { return Str.rfind(str, 0) == 0; };
const auto MatchStr = [](const fextl::string &Str, const char *str) -> bool { return Str.rfind(str, 0) == 0; };
const auto split = [](const std::string &Str, char deliminator) -> std::vector<std::string> {
std::vector<std::string> Elements;
std::istringstream Input(Str);
for (std::string line;
const auto split = [](const fextl::string &Str, char deliminator) -> fextl::vector<fextl::string> {
fextl::vector<fextl::string> Elements;
fextl::istringstream Input(Str);
for (fextl::string line;
std::getline(Input, line);
Elements.emplace_back(line));
return Elements;
};
if (match("QNonStop:")) {
auto ss = std::istringstream(packet);
ss.seekg(std::string("QNonStop:").size());
auto ss = fextl::istringstream(packet);
ss.seekg(fextl::string("QNonStop:").size());
ss.get(); // discard colon
ss >> NonStopMode;
return {"OK", HandledPacketType::TYPE_ACK};
@@ -832,7 +834,7 @@ GdbServer::HandledPacketType GdbServer::handleQuery(const std::string &packet) {
// For feature documentation
// https://sourceware.org/gdb/current/onlinedocs/gdb/General-Query-Packets.html#qSupported
std::string SupportedFeatures{};
fextl::string SupportedFeatures{};
// Required features
SupportedFeatures += "PacketSize=32768;";
@@ -901,7 +903,7 @@ GdbServer::HandledPacketType GdbServer::handleQuery(const std::string &packet) {
if (match("qfThreadInfo")) {
auto Threads = CTX->GetThreads();
std::ostringstream ss;
fextl::ostringstream ss;
ss << "m";
for (size_t i = 0; i < Threads->size(); ++i) {
auto Thread = Threads->at(i);
@@ -916,8 +918,8 @@ GdbServer::HandledPacketType GdbServer::handleQuery(const std::string &packet) {
return {"l", HandledPacketType::TYPE_ACK};
}
if (match("qThreadExtraInfo")) {
auto ss = std::istringstream(packet);
ss.seekg(std::string("qThreadExtraInfo").size());
auto ss = fextl::istringstream(packet);
ss.seekg(fextl::string("qThreadExtraInfo").size());
ss.get(); // discard comma
uint32_t ThreadID;
ss >> std::hex >> ThreadID;
@@ -926,7 +928,7 @@ GdbServer::HandledPacketType GdbServer::handleQuery(const std::string &packet) {
}
if (match("qC")) {
// Returns the current Thread ID
std::ostringstream ss;
fextl::ostringstream ss;
ss << "m" << std::hex << CTX->ParentThread->ThreadManager.TID;
return {ss.str(), HandledPacketType::TYPE_ACK};
}
@@ -935,10 +937,10 @@ GdbServer::HandledPacketType GdbServer::handleQuery(const std::string &packet) {
return {"OK", HandledPacketType::TYPE_ACK};
}
if (match("qSymbol")) {
auto ss = std::istringstream(packet);
ss.seekg(std::string("qSymbol").size());
auto ss = fextl::istringstream(packet);
ss.seekg(fextl::string("qSymbol").size());
ss.get(); // discard colon
std::string Symbol_Val, Symbol_name;
fextl::string Symbol_Val, Symbol_name;
std::getline(ss, Symbol_Val, ':');
std::getline(ss, Symbol_name, ':');
@@ -955,13 +957,13 @@ GdbServer::HandledPacketType GdbServer::handleQuery(const std::string &packet) {
std::fill(PassSignals.begin(), PassSignals.end(), false);
// eg: QPassSignals:e;10;14;17;1a;1b;1c;21;24;25;2c;4c;97;
auto ss = std::istringstream(packet);
ss.seekg(std::string("QPassSignals").size());
auto ss = fextl::istringstream(packet);
ss.seekg(fextl::string("QPassSignals").size());
ss.get(); // discard colon
// We now have a semi-colon deliminated list of signals to pass to the guest process
for (std::string tmp; std::getline(ss, tmp, ';'); ) {
uint32_t Signal = std::stoi(tmp, nullptr, 16);
for (fextl::string tmp; std::getline(ss, tmp, ';'); ) {
uint32_t Signal = std::stoi(tmp.c_str(), nullptr, 16);
if (Signal < SignalDelegator::MAX_SIGNALS) {
PassSignals[Signal] = true;
}
@@ -984,7 +986,7 @@ GdbServer::HandledPacketType GdbServer::ThreadAction(char action, uint32_t tid)
case 's': {
CTX->Step();
SendPacketPair({"OK", HandledPacketType::TYPE_ACK});
auto str = fmt::format("T05thread:{:02x};", getpid());
fextl::string str = fextl::fmt::format("T05thread:{:02x};", getpid());
if (LibraryMapChanged) {
// If libraries have changed then let gdb know
str += "library:1;";
@@ -1002,26 +1004,26 @@ GdbServer::HandledPacketType GdbServer::ThreadAction(char action, uint32_t tid)
}
}
GdbServer::HandledPacketType GdbServer::handleV(const std::string& packet) {
const auto match = [&](const std::string& str) -> std::optional<std::istringstream> {
GdbServer::HandledPacketType GdbServer::handleV(const fextl::string& packet) {
const auto match = [&](const fextl::string& str) -> std::optional<fextl::istringstream> {
if (packet.rfind(str, 0) == 0) {
auto ss = std::istringstream(packet);
auto ss = fextl::istringstream(packet);
ss.seekg(str.size());
return ss;
}
return std::nullopt;
};
const auto F = [](int result) { return fmt::format("F{:x}", result); };
const auto F_error = [] { return fmt::format("F-1,{:x}", errno); };
const auto F_data = [](int result, const std::string& data) {
const auto F = [](int result) -> fextl::string { return fextl::fmt::format("F{:x}", result); };
const auto F_error = []() -> fextl::string { return fextl::fmt::format("F-1,{:x}", errno); };
const auto F_data = [](int result, const fextl::string& data) -> fextl::string {
// Binary encoded data is raw appended to the end
return fmt::format("F{:#x};", result) + data;
return fextl::fmt::format("F{:#x};", result) + data;
};
std::optional<std::istringstream> ss;
std::optional<fextl::istringstream> ss;
if((ss = match("vFile:open:"))) {
std::string filename;
fextl::string filename;
int flags;
int mode;
@@ -1053,7 +1055,7 @@ GdbServer::HandledPacketType GdbServer::handleV(const std::string& packet) {
ss->get(); // discard comma
*ss >> std::hex >> offset;
std::string data(count, '\0');
fextl::string data(count, '\0');
if (lseek(fd, offset, SEEK_SET) < 0) {
return {F_error(), HandledPacketType::TYPE_ACK};
}
@@ -1093,14 +1095,14 @@ GdbServer::HandledPacketType GdbServer::handleV(const std::string& packet) {
return {"", HandledPacketType::TYPE_ACK};
}
GdbServer::HandledPacketType GdbServer::handleThreadOp(const std::string &packet) {
GdbServer::HandledPacketType GdbServer::handleThreadOp(const fextl::string &packet) {
const auto match = [&](const char *str) -> bool { return packet.rfind(str, 0) == 0; };
if (match("Hc")) {
// Sets thread to this ID for stepping
// This is deprecated and vCont should be used instead
auto ss = std::istringstream(packet);
ss.seekg(std::string("Hc").size());
auto ss = fextl::istringstream(packet);
ss.seekg(fextl::string("Hc").size());
ss >> std::hex >> CurrentDebuggingThread;
CTX->Pause();
@@ -1109,7 +1111,7 @@ GdbServer::HandledPacketType GdbServer::handleThreadOp(const std::string &packet
if (match("Hg")) {
// Sets thread for "other" operations
auto ss = std::istringstream(packet);
auto ss = fextl::istringstream(packet);
ss.seekg(std::string_view("Hg").size());
ss >> std::hex >> CurrentDebuggingThread;
@@ -1121,8 +1123,8 @@ GdbServer::HandledPacketType GdbServer::handleThreadOp(const std::string &packet
return {"", HandledPacketType::TYPE_UNKNOWN};
}
GdbServer::HandledPacketType GdbServer::handleBreakpoint(const std::string &packet) {
auto ss = std::istringstream(packet);
GdbServer::HandledPacketType GdbServer::handleBreakpoint(const fextl::string &packet) {
auto ss = fextl::istringstream(packet);
// Don't do anything with set breakpoints yet
[[maybe_unused]] bool Set{};
@@ -1138,13 +1140,13 @@ GdbServer::HandledPacketType GdbServer::handleBreakpoint(const std::string &pack
return {"OK", HandledPacketType::TYPE_ACK};
}
GdbServer::HandledPacketType GdbServer::ProcessPacket(const std::string &packet) {
GdbServer::HandledPacketType GdbServer::ProcessPacket(const fextl::string &packet) {
switch (packet[0]) {
case '?': {
// Indicates the reason that the thread has stopped
// Behaviour changes if the target is in non-stop mode
// Binja doesn't support S response here
auto str = fmt::format("T00thread:{:x};", getpid());
fextl::string str = fextl::fmt::format("T00thread:{:x};", getpid());
return {std::move(str), HandledPacketType::TYPE_ACK};
}
case 'c':
@@ -1225,7 +1227,7 @@ void GdbServer::GdbServerLoop() {
while ((c = CommsStream->get()) >= 0 ) {
switch (c) {
case '$': {
std::string packet = ReadPacket(*CommsStream);
auto packet = ReadPacket(*CommsStream);
response = ProcessPacket(packet);
SendPacketPair(response);
if (response.TypeResponse == HandledPacketType::TYPE_UNKNOWN) {
@@ -1245,7 +1247,7 @@ void GdbServer::GdbServerLoop() {
break;
case '\x03': { // ASCII EOT
CTX->Pause();
auto str = fmt::format("T02thread:{:02x};", getpid());
fextl::string str = fextl::fmt::format("T02thread:{:02x};", getpid());
if (LibraryMapChanged) {
// If libraries have changed then let gdb know
str += "library:1;";
@@ -1279,7 +1281,8 @@ void GdbServer::StartThread() {
}
void GdbServer::OpenListenSocket() {
// open socket
// getaddrinfo allocates memory that can't be removed.
FEXCore::Allocator::YesIKnowImNotSupposedToUseTheGlibcAllocator glibc;
struct addrinfo hints, *res;
memset(&hints, 0, sizeof(hints));
@@ -1308,9 +1311,11 @@ void GdbServer::OpenListenSocket() {
}
listen(ListenSocket, 1);
freeaddrinfo(res);
}
std::unique_ptr<std::iostream> GdbServer::OpenSocket() {
fextl::unique_ptr<std::iostream> GdbServer::OpenSocket() {
// Block until a connection arrives
struct sockaddr_storage their_addr{};
socklen_t addr_size{};
@@ -1318,7 +1323,8 @@ std::unique_ptr<std::iostream> GdbServer::OpenSocket() {
LogMan::Msg::IFmt("GdbServer, waiting for connection on localhost:8086");
int new_fd = accept(ListenSocket, (struct sockaddr *)&their_addr, &addr_size);
return std::make_unique<FEXCore::Utils::NetStream>(new_fd);
return fextl::make_unique<FEXCore::Utils::NetStream>(new_fd);
}
#endif
} // namespace FEXCore
+23 -22
View File
@@ -8,23 +8,24 @@ $end_info$
#include <FEXCore/Config/Config.h>
#include <FEXCore/Utils/Event.h>
#include <FEXCore/Utils/Threads.h>
#include <FEXCore/fextl/memory.h>
#include <FEXCore/fextl/string.h>
#include <atomic>
#include <istream>
#include <memory>
#include <mutex>
#include <stdint.h>
#include <string>
namespace FEXCore {
namespace Context {
struct Context;
class ContextImpl;
}
class GdbServer {
public:
GdbServer(FEXCore::Context::Context *ctx);
GdbServer(FEXCore::Context::ContextImpl *ctx);
// Public for threading
void GdbServerLoop();
@@ -37,10 +38,10 @@ private:
void Break(int signal);
void OpenListenSocket();
std::unique_ptr<std::iostream> OpenSocket();
fextl::unique_ptr<std::iostream> OpenSocket();
void StartThread();
std::string ReadPacket(std::iostream &stream);
void SendPacket(std::ostream &stream, const std::string& packet);
fextl::string ReadPacket(std::iostream &stream);
void SendPacket(std::ostream &stream, const fextl::string& packet);
void SendACK(std::ostream &stream, bool NACK);
@@ -48,7 +49,7 @@ private:
void WaitForThreadWakeup();
struct HandledPacketType {
std::string Response{};
fextl::string Response{};
enum ResponseType {
TYPE_NONE,
TYPE_UNKNOWN,
@@ -61,32 +62,32 @@ private:
};
void SendPacketPair(const HandledPacketType& packetPair);
HandledPacketType ProcessPacket(const std::string &packet);
HandledPacketType handleQuery(const std::string &packet);
HandledPacketType handleXfer(const std::string &packet);
HandledPacketType handleMemory(const std::string &packet);
HandledPacketType handleV(const std::string& packet);
HandledPacketType handleThreadOp(const std::string &packet);
HandledPacketType handleBreakpoint(const std::string &packet);
HandledPacketType ProcessPacket(const fextl::string &packet);
HandledPacketType handleQuery(const fextl::string &packet);
HandledPacketType handleXfer(const fextl::string &packet);
HandledPacketType handleMemory(const fextl::string &packet);
HandledPacketType handleV(const fextl::string& packet);
HandledPacketType handleThreadOp(const fextl::string &packet);
HandledPacketType handleBreakpoint(const fextl::string &packet);
HandledPacketType handleProgramOffsets();
HandledPacketType ThreadAction(char action, uint32_t tid);
std::string readRegs();
HandledPacketType readReg(const std::string& packet);
fextl::string readRegs();
HandledPacketType readReg(const fextl::string& packet);
FEXCore::Context::Context *CTX;
std::unique_ptr<FEXCore::Threads::Thread> gdbServerThread;
std::unique_ptr<std::iostream> CommsStream;
FEXCore::Context::ContextImpl *CTX;
fextl::unique_ptr<FEXCore::Threads::Thread> gdbServerThread;
fextl::unique_ptr<std::iostream> CommsStream;
std::mutex sendMutex;
bool SettingNoAckMode{false};
bool NoAckMode{false};
bool NonStopMode{false};
std::string ThreadString{};
std::string OSDataString{};
fextl::string ThreadString{};
fextl::string OSDataString{};
void buildLibraryMap();
std::atomic<bool> LibraryMapChanged = true;
std::string LibraryMapString{};
fextl::string LibraryMapString{};
// Used to keep track of which signals to pass to the guest
std::array<bool, SignalDelegator::MAX_SIGNALS + 1> PassSignals{};
+21 -9
View File
@@ -9,7 +9,7 @@
#endif
#ifdef _M_X86_64
#include <xbyak/xbyak_util.h>
#include "Interface/Core/Dispatcher/X86Dispatcher.h"
#endif
namespace FEXCore {
@@ -17,12 +17,12 @@ namespace FEXCore {
// Data Zero Prohibited flag
// 0b0 = ZVA/GVA/GZVA permitted
// 0b1 = ZVA/GVA/GZVA prohibited
constexpr uint32_t DCZID_DZP_MASK = 0b1'0000;
[[maybe_unused]] constexpr uint32_t DCZID_DZP_MASK = 0b1'0000;
// Log2 of the blocksize in 32-bit words
constexpr uint32_t DCZID_BS_MASK = 0b0'1111;
[[maybe_unused]] constexpr uint32_t DCZID_BS_MASK = 0b0'1111;
#ifdef _M_ARM_64
static uint32_t GetDCZID() {
[[maybe_unused]] static uint32_t GetDCZID() {
uint64_t Result{};
__asm("mrs %[Res], DCZID_EL0"
: [Res] "=r" (Result));
@@ -54,7 +54,12 @@ HostFeatures::HostFeatures() {
#ifdef VIXL_SIMULATOR
auto Features = vixl::CPUFeatures::All();
#else
#ifndef _WIN32
auto Features = vixl::CPUFeatures::InferFromOS();
#else
// Need to use ID registers in WINE.
auto Features = vixl::CPUFeatures::InferFromIDRegisters();
#endif
#endif
SupportsAES = Features.Has(vixl::CPUFeatures::Feature::kAES);
SupportsCRC = Features.Has(vixl::CPUFeatures::Feature::kCRC32);
@@ -133,13 +138,14 @@ HostFeatures::HostFeatures() {
SupportsPMULL_128Bit = Features.has(Xbyak::util::Cpu::tPCLMULQDQ);
// xbyak doesn't know how to check for CLZero
uint32_t eax, ebx, ecx, edx;
// First ensure we support a new enough extended CPUID function range
__cpuid(0x8000'0000, eax, ebx, ecx, edx);
if (eax >= 0x8000'0008U) {
uint32_t data[4];
Xbyak::util::Cpu::getCpuid(0x8000'0000, data);
if (data[0] >= 0x8000'0008U) {
// CLZero defined in 8000_00008_EBX[bit 0]
__cpuid(0x8000'0008, eax, ebx, ecx, edx);
SupportsCLZERO = ebx & 1;
Xbyak::util::Cpu::getCpuid(0x8000'0008, data);
SupportsCLZERO = data[1] & 1;
}
SupportsFlushInputsToZero = true;
@@ -160,5 +166,11 @@ HostFeatures::HostFeatures() {
SupportsCLZERO = DCZID_Bytes == CPUIDEmu::CACHELINE_SIZE;
}
#endif
// Disable AVX if the configuration explicitly has disabled it.
FEX_CONFIG_OPT(EnableAVX, ENABLEAVX);
if (!EnableAVX) {
SupportsAVX = false;
}
}
}
@@ -17,22 +17,8 @@ $end_info$
#include <unistd.h>
namespace FEXCore::CPU {
[[noreturn]]
static void SignalReturn(FEXCore::Core::InternalThreadState *Thread, bool RT) {
Thread->CTX->SignalThread(Thread, RT ? FEXCore::Core::SignalEvent::ReturnRT : FEXCore::Core::SignalEvent::Return);
LOGMAN_MSG_A_FMT("unreachable");
FEX_UNREACHABLE;
}
#define DEF_OP(x) void InterpreterOps::Op_##x(IR::IROp_Header *IROp, IROpData *Data, IR::NodeID Node)
DEF_OP(SignalReturn) {
auto Op = IROp->C<IR::IROp_SignalReturn>();
SignalReturn(Data->State, Op->IsRT);
}
DEF_OP(CallbackReturn) {
Data->State->CurrentFrame->Pointers.Interpreter.CallbackReturn(Data->State, Data->StackEntry);
}
@@ -93,7 +79,7 @@ DEF_OP(Syscall) {
Args.Argument[j] = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[j]);
}
uint64_t Res = FEXCore::Context::HandleSyscall(Data->State->CTX->SyscallHandler, Data->State->CurrentFrame, &Args);
uint64_t Res = FEXCore::Context::HandleSyscall(static_cast<Context::ContextImpl*>(Data->State->CTX)->SyscallHandler, Data->State->CurrentFrame, &Args);
GD = Res;
}
@@ -128,7 +114,7 @@ DEF_OP(InlineSyscall) {
DEF_OP(Thunk) {
auto Op = IROp->C<IR::IROp_Thunk>();
auto thunkFn = Data->State->CTX->ThunkHandler->LookupThunk(Op->ThunkNameHash);
auto thunkFn = static_cast<Context::ContextImpl*>(Data->State->CTX)->ThunkHandler->LookupThunk(Op->ThunkNameHash);
thunkFn(*GetSrc<void**>(Data->SSAData, Op->ArgPtr));
}
@@ -144,7 +130,7 @@ DEF_OP(ValidateCode) {
}
DEF_OP(ThreadRemoveCodeEntry) {
Data->State->CTX->ThreadRemoveCodeEntryFromJit(Data->State->CurrentFrame, Data->CurrentEntry);
static_cast<Context::ContextImpl*>(Data->State->CTX)->ThreadRemoveCodeEntryFromJit(Data->State->CurrentFrame, Data->CurrentEntry);
}
DEF_OP(CPUID) {
@@ -153,10 +139,19 @@ DEF_OP(CPUID) {
const uint64_t Arg = *GetSrc<uint64_t*>(Data->SSAData, Op->Function);
const uint64_t Leaf = *GetSrc<uint64_t*>(Data->SSAData, Op->Leaf);
auto Results = Data->State->CTX->CPUID.RunFunction(Arg, Leaf);
auto Results = Data->State->CTX->RunCPUIDFunction(Arg, Leaf);
memcpy(DstPtr, &Results, sizeof(uint32_t) * 4);
}
DEF_OP(XGETBV) {
auto Op = IROp->C<IR::IROp_XGetBV>();
uint32_t *DstPtr = GetDest<uint32_t*>(Data->SSAData, Node);
const uint32_t Function = *GetSrc<uint32_t*>(Data->SSAData, Op->Function);
auto Results = Data->State->CTX->RunXCRFunction(Function);
memcpy(DstPtr, &Results, sizeof(uint32_t) * 2);
}
#undef DEF_OP
} // namespace FEXCore::CPU
@@ -62,6 +62,23 @@ DEF_OP(VCastFromGPR) {
memcpy(GDP, GetSrc<void*>(Data->SSAData, Op->Src), Op->Header.ElementSize);
}
DEF_OP(VDupFromGPR) {
const auto Op = IROp->C<IR::IROp_VDupFromGPR>();
const auto OpSize = IROp->Size;
const auto ElementSize = IROp->ElementSize;
const auto NumElements = OpSize / IROp->ElementSize;
uint8_t Tmp[Core::CPUState::XMM_AVX_REG_SIZE]{};
const auto *Src = GetSrc<void*>(Data->SSAData, Op->Src);
for (size_t i = 0; i < NumElements; i++) {
memcpy(Tmp + (i * ElementSize), Src, ElementSize);
}
memcpy(GDP, Tmp, sizeof(Tmp));
}
DEF_OP(Float_FromGPR_S) {
auto Op = IROp->C<IR::IROp_Float_FromGPR_S>();
@@ -8,7 +8,7 @@ $end_info$
#include "Interface/Core/Interpreter/InterpreterOps.h"
#include "Interface/Core/Interpreter/InterpreterDefines.h"
#include "F80Ops.h"
#include "Interface/Core/Interpreter/Fallbacks/F80Fallbacks.h"
#include <cstdint>
@@ -417,7 +417,6 @@ DEF_OP(F64SCALE) {
memcpy(GDP, &Tmp, sizeof(double));
}
#undef DEF_OP
} // namespace FEXCore::CPU
@@ -4,12 +4,9 @@
#include <FEXCore/IR/IR.h>
#include "Interface/Core/Interpreter/Fallbacks/FallbackOpHandler.h"
namespace FEXCore::CPU {
template<IR::IROps Op>
struct OpHandlers {
};
template<>
struct OpHandlers<IR::OP_F80CVTTO> {
static X80SoftFloat handle4(float src) {
@@ -395,5 +392,4 @@ struct OpHandlers<IR::OP_F80LOADFCW> {
}
};
}
} // namespace FEXCore::CPU
@@ -0,0 +1,75 @@
#pragma once
#include <cstdint>
namespace FEXCore::IR {
enum IROps : uint8_t;
}
namespace FEXCore::CPU {
// Base template for fallback handling.
//
// Registering and hooking up fallback is currently like so:
//
// 1. Go to InterpreterFallbacks.cpp and create a template specialization of
// the GetFallbackInfo member function.
//
// This member function should reasonably define what the fallback you're
// going to create will take as parameters and return as a result. For example:
//
// template<>
// FallbackInfo GetFallbackInfo(X80SoftFloat(*fn)(double), Core::FallbackHandlerIndex Index) {
// return {FABI_F80_F64, (void*)fn, Index};
// }
//
// Defines info about a fallback that takes a double as an argument and
// returns a X80SoftFloat instance.
//
// You will also want to define a new FallbackHandlerIndex enum member and use it
// to set up the new info handler into the Info array in FillFallbackIndexPointers.
//
// 1.1. (potentially optional). Define a new ABI element in the FallbackAPI enum.
// This ABI enum value will be used to tell the JITs how to handle the fallback
// properly. These enum values specify the return type followed by its argument types.
//
// So, FABI_I64_F80_F80, for example indicates that the function will behave like a
// function as if were defined as:
//
// uint64_t fn(X80SoftFloat, X80SoftFloat)
//
// 1.2. (potentially optional). If you needed to define a new enum ABI type like in 1.1, then
// you need to add the handling for it in the JITs, which can be found in the respective
// JIT's JIT.cpp file in a function called Op_Unhandled
//
// You need to add a new case to the ABI switch statement using the new ABI type
// and do the necessary moving of data from register-allocated JIT parameters
// into that platform's registers that respects the calling convention. After this is
// done, most of the necessary background boilerplate is finished.
//
// 2. Now, make a specialization of this class with a member function named 'handle()'
// that takes the same parameters as the ones described in the fallback info function
// specialization.
//
// For example, if you have the fallback info from the example in step 1, it would be:
//
// template <>
// struct OpHandlers<IR::CoolNewIROpcode> {
// static X80SoftFloat handle(double src) {
// return ...;
// }
// };
//
// 3. Fill out the behavior of the OpHandler specialization to perform what you would like
// the fallback to do.
//
// 4. Add an implementation of the IR op to the Interpreter that passes through to the
// OpHandler implementation.
//
// 5. Done.
//
template <IR::IROps Op>
struct OpHandlers {
};
} // namespace FEXCore::CPU
@@ -1,6 +1,8 @@
#include "FEXCore/Core/CoreState.h"
#include <FEXCore/Core/CoreState.h>
#include "Interface/Core/Interpreter/InterpreterOps.h"
#include "Interface/Core/Interpreter/F80Ops.h"
#include "Interface/Core/Interpreter/Fallbacks/F80Fallbacks.h"
#include "Interface/Core/Interpreter/Fallbacks/VectorFallbacks.h"
#include <cstddef>
#include <cstdint>
@@ -87,6 +89,16 @@ FallbackInfo GetFallbackInfo(X80SoftFloat(*fn)(X80SoftFloat, X80SoftFloat), FEXC
return {FABI_F80_F80_F80, (void*)fn, HandlerIndex};
}
template<>
FallbackInfo GetFallbackInfo(uint32_t(*fn)(uint64_t, uint64_t, __uint128_t, __uint128_t, uint16_t), FEXCore::Core::FallbackHandlerIndex HandlerIndex) {
return {FABI_I32_I64_I64_I128_I128_I16, (void*)fn, HandlerIndex};
}
template<>
FallbackInfo GetFallbackInfo(uint32_t(*fn)(__uint128_t, __uint128_t, uint16_t), FEXCore::Core::FallbackHandlerIndex HandlerIndex) {
return {FABI_I32_I128_I128_I16, (void*)fn, HandlerIndex};
}
void InterpreterOps::FillFallbackIndexPointers(uint64_t *Info) {
Info[Core::OPINDEX_F80LOADFCW] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80LOADFCW>::handle, Core::OPINDEX_F80LOADFCW).fn);
Info[Core::OPINDEX_F80CVTTO_4] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTO>::handle4, Core::OPINDEX_F80CVTTO_4).fn);
@@ -144,6 +156,9 @@ void InterpreterOps::FillFallbackIndexPointers(uint64_t *Info) {
Info[Core::OPINDEX_F64FPREM1] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F64FPREM1>::handle, Core::OPINDEX_F64FPREM1).fn);
Info[Core::OPINDEX_F64SCALE] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F64SCALE>::handle, Core::OPINDEX_F64SCALE).fn);
// SSE4.2 string instructions
Info[Core::OPINDEX_VPCMPESTRX] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_VPCMPESTRX>::handle, Core::OPINDEX_VPCMPESTRX).fn);
Info[Core::OPINDEX_VPCMPISTRX] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_VPCMPISTRX>::handle, Core::OPINDEX_VPCMPISTRX).fn);
}
bool InterpreterOps::GetFallbackHandler(IR::IROp_Header const *IROp, FallbackInfo *Info) {
@@ -302,6 +317,14 @@ bool InterpreterOps::GetFallbackHandler(IR::IROp_Header const *IROp, FallbackInf
COMMON_F64_OP(FPREM)
COMMON_F64_OP(SCALE)
// SSE4.2 Fallbacks
case IR::OP_VPCMPESTRX:
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_VPCMPESTRX>::handle, Core::OPINDEX_VPCMPESTRX);
return true;
case IR::OP_VPCMPISTRX:
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_VPCMPISTRX>::handle, Core::OPINDEX_VPCMPISTRX);
return true;
default:
break;
}
@@ -0,0 +1,408 @@
#pragma once
#include <algorithm>
#include <cstddef>
#include <cstdint>
#include <cstdlib>
#include <cstring>
#include <FEXCore/IR/IR.h>
#include "Interface/Core/Interpreter/Fallbacks/FallbackOpHandler.h"
namespace FEXCore::CPU {
template<>
struct OpHandlers<IR::OP_VPCMPESTRX> {
enum class AggregationOp {
EqualAny = 0b00,
Ranges = 0b01,
EqualEach = 0b10,
EqualOrdered = 0b11,
};
enum class SourceData {
U8,
U16,
S8,
S16,
};
enum class Polarity {
Positive,
Negative,
PositiveMasked,
NegativeMasked,
};
static uint32_t handle(uint64_t RAX, uint64_t RDX, __uint128_t lhs, __uint128_t rhs, uint16_t control) {
// Subtract by 1 in order to make validity limits 0-based
const auto valid_lhs = GetExplicitLength(RAX, control) - 1;
const auto valid_rhs = GetExplicitLength(RDX, control) - 1;
return MainBody(lhs, valid_lhs, rhs, valid_rhs, control);
}
// Main PCMPXSTRX algorithm body. Allows for reuse with both implicit and explicit length variants.
static uint32_t MainBody(const __uint128_t& lhs, int valid_lhs, const __uint128_t& rhs, int valid_rhs, uint16_t control) {
const uint32_t aggregation = PerformAggregation(lhs, valid_lhs, rhs, valid_rhs, control);
const uint32_t upper_limit = (16U >> (control & 1)) - 1;
// Bits are arranged as:
// Bit #: 3 2 1 0
// [OF | CF | SF | ZF]
uint32_t flags = 0;
flags |= (valid_rhs < upper_limit) ? 0b01 : 0b00;
flags |= (valid_lhs < upper_limit) ? 0b10 : 0b00;
const uint32_t result = HandlePolarity(aggregation, control, upper_limit, valid_rhs);
if (result != 0) {
flags |= 0b0100;
}
if ((result & 1) != 0) {
flags |= 0b1000;
}
// We tack the flags on top of the result to avoid needing to handle
// multiple return values in the JITs.
return result | (flags << 16);
}
static int32_t GetExplicitLength(uint64_t reg, uint16_t control) {
// Bit 8 controls whether or not the reg value is 64-bit or 32-bit.
int64_t value = 0;
if (((control >> 8) & 1) != 0) {
value = static_cast<int64_t>(reg);
} else {
// We need a sign extend in this case.
value = static_cast<int32_t>(reg);
}
// If control[0] is set, then we're dealing with words instead of bytes
const int64_t limit = (control & 1) != 0 ? 8 : 16;
// Length needs to saturate to 16 (if bytes) or 8 (if words)
// when the length value is greater than 16 (if bytes)/8 (if words)
// or if the length value is less than -16 (if bytes)/-8 (if words).
if (value < -limit || value > limit) {
return limit;
}
return std::abs(static_cast<int>(value));
}
static int32_t GetElement(const __uint128_t& vec, int32_t index, uint16_t control) {
const auto* vec_ptr = reinterpret_cast<const uint8_t*>(&vec);
// Control bits [1:0] define the data type being dealt with.
switch (static_cast<SourceData>(control & 0b11)) {
case SourceData::U8:
return static_cast<int32_t>(vec_ptr[index]);
case SourceData::U16: {
uint16_t value{};
std::memcpy(&value, vec_ptr + (sizeof(uint16_t) * static_cast<size_t>(index)), sizeof(value));
return value;
}
case SourceData::S8:
return static_cast<int8_t>(vec_ptr[index]);
case SourceData::S16:
default: {
int16_t value{};
std::memcpy(&value, vec_ptr + (sizeof(int16_t) * static_cast<size_t>(index)), sizeof(value));
return value;
}
}
}
static uint32_t PerformAggregation(const __uint128_t& lhs, int32_t valid_lhs,
const __uint128_t& rhs, int32_t valid_rhs,
uint16_t control) {
switch (static_cast<AggregationOp>((control >> 2) & 0b11)) {
case AggregationOp::EqualAny:
return HandleEqualAny(lhs, valid_lhs, rhs, valid_rhs, control);
case AggregationOp::Ranges:
return HandleRanges(lhs, valid_lhs, rhs, valid_rhs, control);
case AggregationOp::EqualEach:
return HandleEqualEach(lhs, valid_lhs, rhs, valid_rhs, control);
case AggregationOp::EqualOrdered:
default:
return HandleEqualOrdered(lhs, valid_lhs, rhs, valid_rhs, control);
}
}
static uint32_t HandlePolarity(uint32_t value, uint16_t control, int upper_limit, int valid_rhs) {
switch (static_cast<Polarity>((control >> 4) & 0b11)) {
case Polarity::Negative:
return value ^ ((2U << upper_limit) - 1);
case Polarity::NegativeMasked:
return value ^ ((1U << (valid_rhs + 1)) - 1);
case Polarity::Positive:
case Polarity::PositiveMasked:
default:
// Both positive masking and positive polarity are documented
// as both being equivalent to "IntRes2 = IntRes1", where IntRes1
// is our 'value' parameter, so we don't need to do anything in
// these cases except return the same value.
return value;
}
}
// Finds characters from an overall character set.
//
// Scans through RHS trying to find any characters contained in LHS.
// Think of this as a sort of vectorized version of strspn (kind of).
//
// e.g. Assume operating on two character vectors as unsigned words
//
// 0 1 2 3 4 5 6 7
// LHS -> [a, b, c, d, e, f, g, n]
// RHS -> [z, k, v, c, d, o, p, n]
//
// With both explicit lengths for each string being 8 (the max length for words),
// this would result in an intermediate result like:
//
// 0b1001'1000
// │ │ │
// 'n' match ───┘ │ │
// │ │
// 'd' match ──────┘ │
// │
// 'c' match ────────┘
//
static uint32_t HandleEqualAny(const __uint128_t& lhs, int32_t valid_lhs,
const __uint128_t& rhs, int32_t valid_rhs,
uint16_t control) {
uint32_t result = 0;
for (int j = valid_rhs; j >= 0; j--) {
result <<= 1;
const int rhs_value = GetElement(rhs, j, control);
for (int i = valid_lhs; i >= 0; i--) {
const int lhs_value = GetElement(lhs, i, control);
result |= static_cast<uint32_t>(rhs_value == lhs_value);
}
}
return result;
}
// Determines if a character falls within a limited range
//
// Scans through rhs using a range denoted by two elements
// in lhs and determines if the respective character in rhs
// falls within its range.
//
// i.e.
// lhs_upper_bound >= rhs_value && lhs_lower_bound <= rhs_value
//
// e.g. Assume operating on two character vectors as unsigned words
//
// 0 1 2 3 4 5 6 7
// LHS -> [a, z, A, Z, 0, 0, 0, 0]
// RHS -> [z, k, ., C, M, ;, \, ']
//
// With LHS's length being 4 and RHS's lenth being 8,
// this would result in an intermediate result like:
//
// 0b0001'1011
// │ │ ││
// 'z' >= 'M' && 'a' <= 'M' ─────┘ │ ││
// │ ││
// 'z' >= 'C' && 'a' <= 'C' ───────┘ ││
// ││
// 'Z' >= 'k' && 'A' <= 'k' ─────────┘│
// │
// 'Z' >= 'z' && 'A' <= 'z' ──────────┘
//
static uint32_t HandleRanges(const __uint128_t& lhs, int32_t valid_lhs,
const __uint128_t& rhs, int32_t valid_rhs,
uint16_t control) {
uint32_t result = 0;
for (int j = valid_rhs; j >= 0; j--) {
result <<= 1;
const int element = GetElement(rhs, j, control);
for (int i = (valid_lhs - 1) | 1; i >= 0; i -= 2) {
const int upper_bound = GetElement(lhs, i - 0, control);
const int lower_bound = GetElement(lhs, i - 1, control);
const bool ge = upper_bound >= element;
const bool le = lower_bound <= element;
result |= static_cast<uint32_t>(ge && le);
}
}
return result;
}
// Determines if each character is equal to one another (string compare)
//
// Essentially the PCMPXSTRX variant of memcmp/strcmp. Sets the bit of the
// resulting mask if both elements are equal to one another. Otherwise
// sets it to false.
//
// e.g. Assume operating on two character vectors as unsigned words
//
// 0 1 2 3 4 5 6 7
// LHS -> [a, b, c, d, e, f, g, n]
// RHS -> [a, b, c, d, e, f, e, x]
//
// With both explicit lengths for each string being 8 (the max length for words),
// this would result in an intermediate result like:
//
// 0b0011'1111
// ││ ││││
// 'f' == 'f' ────┘│ ││││
// │ ││││
// 'e' == 'e' ─────┘ ││││
// ││││
// 'd' == 'd' ───────┘│││
// │││
// 'c' == 'c' ────────┘││
// ││
// 'b' == 'b' ─────────┘│
// │
// 'a' == 'a' ──────────┘
//
static uint32_t HandleEqualEach(const __uint128_t& lhs, int32_t valid_lhs,
const __uint128_t& rhs, int32_t valid_rhs,
uint16_t control) {
const auto upper_limit = (16 >> (control & 1)) - 1;
const auto max_valid = std::max(valid_lhs, valid_rhs);
const auto min_valid = std::min(valid_lhs, valid_rhs);
// All values past the end of string must be forced to true.
// (See 4.1.6 Valid/Invalid Override of Comparisons in the Intel Software Development Manual)
// So we can calculate this part of the mask ahead of time and set all those to-be bits to true
// and then progressively shift them into place over the course of execution.
uint32_t result = (1U << (upper_limit - max_valid)) - 1;
result <<= (max_valid - min_valid);
for (int i = min_valid; i >= 0; i--) {
const int lhs_element = GetElement(lhs, i, control);
const int rhs_element = GetElement(rhs, i, control);
result <<= 1;
result |= static_cast<uint32_t>(lhs_element == rhs_element);
}
return result;
}
// Determines if a substring exists within an overall string
//
// Somewhat equivalent to the behavior of strstr.
//
// Sets the corresponding index in the result where a substring is found.
//
// e.g. Assume operating on two character vectors as unsigned words
//
// 0 1 2 3 4 5 6 7
// LHS -> [b, a, x, z, y, v, o, m]
// RHS -> [b, a, d, b, a, n, k, s]
//
// With the length of LHS being 2 and the length of RHS being 8, we have a composition like:
//
// Substring to look for
// ┌──┴──┐
// LHS -> [b, a, x, z, y, v, o, m]
// RHS -> [b, a, d, b, a, n, k, s]
// └───────────┬────────────┘
// Entire string to search
//
// And we end up with a result like:
//
// 0b0000'1001
// │ │
// At index 3 ───────┘ │
// │
// At index 0 ──────────┘
//
static uint32_t HandleEqualOrdered(const __uint128_t& lhs, int32_t valid_lhs,
const __uint128_t& rhs, int32_t valid_rhs,
uint16_t control) {
const auto upper_limit = (16 >> (control & 1)) - 1;
// Edge case!
// If we have *no* valid characters in our inner string, then
// we need to return the intermediate result as
// 0xFF (if operating on words) or 0xFFFF (if operating on bytes)
if (valid_lhs == -1) {
return (2U << upper_limit) - 1;
}
uint32_t result = 0;
const int initial = valid_rhs == upper_limit ? valid_rhs
: valid_rhs - valid_lhs;
for (int j = initial; j >= 0; j--) {
result <<= 1;
uint32_t value = 1;
const int start = std::min(valid_rhs - j, valid_lhs);
for (int i = start; i >= 0; i--) {
const int lhs_value = GetElement(lhs, i + 0, control);
const int rhs_value = GetElement(rhs, i + j, control);
value &= static_cast<uint32_t>(lhs_value == rhs_value);
}
result |= value;
}
return result;
}
};
template<>
struct OpHandlers<IR::OP_VPCMPISTRX> {
// Essentially the same in terms of behavior with VPCMPESTRX instructions,
// with the only difference being that the length of the string is encoded
// as part of the data vectors passed in.
//
// i.e. Length is determined by the presence of a NUL (all-zero) character
// within the data.
//
// If no NUL character exists, then the length of the strings are assumed
// to be the max length possible for the given character size specified
// in the control flags (16 characters for 8-bit, and 8 characters for 16-bit).
//
static uint32_t handle(__uint128_t lhs, __uint128_t rhs, uint16_t control) {
// Subtract by 1 in order to make validity limits 0-based
const auto valid_lhs = GetImplicitLength(lhs, control) - 1;
const auto valid_rhs = GetImplicitLength(rhs, control) - 1;
return OpHandlers<IR::OP_VPCMPESTRX>::MainBody(lhs, valid_lhs, rhs, valid_rhs, control);
}
static int32_t GetImplicitLength(const __uint128_t& data, uint16_t control) {
const auto* data_u8 = reinterpret_cast<const uint8_t*>(&data);
const auto is_using_words = (control & 1) != 0;
int32_t length = 0;
if (is_using_words) {
const auto get_word = [data_u8](int32_t index) {
const auto* src = data_u8 + (index * sizeof(uint16_t));
uint16_t element{};
std::memcpy(&element, src, sizeof(uint16_t));
return element;
};
while (length < 8 && get_word(length) != 0) {
length++;
}
} else {
while (length < 16 && data_u8[length] != 0) {
length++;
}
}
return length;
}
};
} // namespace FEXCore::CPU
@@ -6,27 +6,24 @@
#include <FEXCore/Core/CPUBackend.h>
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/fextl/string.h>
#include <FEXCore/fextl/vector.h>
namespace FEXCore::CPU {
class Dispatcher;
class X86DispatchGenerator;
class Arm64DispatchGenerator;
#define DESTMAP_AS_MAP 0
#if DESTMAP_AS_MAP
using DestMapType = std::unordered_map<uint32_t, uint32_t>;
#else
using DestMapType = std::vector<uint32_t>;
#endif
using DestMapType = fextl::vector<uint32_t>;
class InterpreterCore final : public CPUBackend {
public:
explicit InterpreterCore(Dispatcher *Dispatch,
FEXCore::Core::InternalThreadState *Thread);
[[nodiscard]] std::string GetName() override { return "Interpreter"; }
[[nodiscard]] fextl::string GetName() override { return "Interpreter"; }
[[nodiscard]] void *CompileCode(uint64_t Entry,
[[nodiscard]] CPUBackend::CompiledCode CompileCode(uint64_t Entry,
FEXCore::IR::IRListView const *IR,
FEXCore::Core::DebugData *DebugData,
FEXCore::IR::RegisterAllocationData *RAData, bool GDBEnabled) override;
@@ -35,8 +32,8 @@ public:
[[nodiscard]] bool NeedsOpDispatch() override { return true; }
static void InitializeSignalHandlers(FEXCore::Context::Context *CTX);
static void InitializeSignalHandlers(FEXCore::Context::ContextImpl *CTX);
void ClearCache() override;
private:
@@ -1,17 +1,14 @@
#include "Interface/Context/Context.h"
#include "Interface/Core/ArchHelpers/Arm64.h"
#include "Interface/Core/ArchHelpers/MContext.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Core/Interpreter/InterpreterClass.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Core/SignalDelegator.h>
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/MathUtils.h>
#include <FEXCore/fextl/memory.h>
#include <memory>
#include <signal.h>
#include <stdint.h>
#include <utility>
@@ -49,27 +46,21 @@ InterpreterCore::InterpreterCore(Dispatcher *Dispatcher, FEXCore::Core::Internal
ClearCache();
}
void InterpreterCore::InitializeSignalHandlers(FEXCore::Context::Context *CTX) {
#ifdef _M_ARM_64
CTX->SignalDelegation->RegisterHostSignalHandler(SIGBUS, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
return FEXCore::ArchHelpers::Arm64::HandleSIGBUS(true, Signal, info, ucontext);
}, true);
#endif
}
void *InterpreterCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IRListView const *IR, [[maybe_unused]] FEXCore::Core::DebugData *DebugData, FEXCore::IR::RegisterAllocationData *RAData, bool GDBEnabled) {
CPUBackend::CompiledCode InterpreterCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IRListView const *IR, [[maybe_unused]] FEXCore::Core::DebugData *DebugData, FEXCore::IR::RegisterAllocationData *RAData, bool GDBEnabled) {
const auto IRSize = AlignUp(IR->GetInlineSize(), 16);
const auto MaxSize = IRSize + Dispatcher::MaxInterpreterTrampolineSize + GDBEnabled * Dispatcher::MaxGDBPauseCheckSize;
if ((BufferUsed + MaxSize) > CurrentCodeBuffer->Size) {
ThreadState->CTX->ClearCodeCache(ThreadState);
static_cast<Context::ContextImpl*>(ThreadState->CTX)->ClearCodeCache(ThreadState);
}
const auto BufferStart = CurrentCodeBuffer->Ptr + BufferUsed;
CPUBackend::CompiledCode CodeData{};
auto DestBuffer = BufferStart;
const auto BufferStartOffset = BufferUsed;
CodeData.BlockBegin = CodeData.BlockEntry = CurrentCodeBuffer->Ptr + BufferStartOffset;
auto DestBuffer = CodeData.BlockBegin;
if (GDBEnabled) {
const auto GDBSize = Dispatch->GenerateGDBPauseCheck(DestBuffer, Entry);
@@ -86,7 +77,9 @@ void *InterpreterCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR:
DestBuffer += IRSize;
BufferUsed += IRSize;
return BufferStart;
CodeData.Size = BufferUsed - BufferStartOffset;
return CodeData;
}
void InterpreterCore::ClearCache() {
@@ -95,12 +88,8 @@ void InterpreterCore::ClearCache() {
BufferUsed = 0;
}
std::unique_ptr<CPUBackend> CreateInterpreterCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread) {
return std::make_unique<InterpreterCore>(ctx->Dispatcher.get(), Thread);
}
void InitializeInterpreterSignalHandlers(FEXCore::Context::Context *CTX) {
InterpreterCore::InitializeSignalHandlers(CTX);
fextl::unique_ptr<CPUBackend> CreateInterpreterCore(FEXCore::Context::ContextImpl *ctx, FEXCore::Core::InternalThreadState *Thread) {
return fextl::make_unique<InterpreterCore>(ctx->Dispatcher.get(), Thread);
}
CPUBackendFeatures GetInterpreterBackendFeatures() {
@@ -1,9 +1,10 @@
#pragma once
#include <memory>
#include <FEXCore/Core/CPUBackend.h>
#include <FEXCore/fextl/memory.h>
namespace FEXCore::Context {
struct Context;
class ContextImpl;
}
namespace FEXCore::Core {
@@ -14,9 +15,9 @@ namespace FEXCore::CPU {
class CPUBackend;
struct DispatcherConfig;
[[nodiscard]] std::unique_ptr<CPUBackend> CreateInterpreterCore(FEXCore::Context::Context *ctx,
[[nodiscard]] fextl::unique_ptr<CPUBackend> CreateInterpreterCore(FEXCore::Context::ContextImpl *ctx,
FEXCore::Core::InternalThreadState *Thread);
void InitializeInterpreterSignalHandlers(FEXCore::Context::Context *CTX);
void InitializeInterpreterSignalHandlers(FEXCore::Context::ContextImpl *CTX);
CPUBackendFeatures GetInterpreterBackendFeatures();
} // namespace FEXCore::CPU
@@ -2,11 +2,6 @@
#include "Interface/Core/CPUID.h"
#include "InterpreterDefines.h"
#include "InterpreterOps.h"
#include "F80Ops.h"
#ifdef _M_ARM_64
#include "Interface/Core/ArchHelpers/Arm64.h"
#endif
#include <FEXCore/Core/CPUBackend.h>
#include <FEXCore/Core/CoreState.h>
@@ -113,7 +108,6 @@ constexpr OpHandlerArray InterpreterOpHandlers = [] {
REGISTER_OP(ATOMICFETCHNEG, AtomicFetchNeg);
// Branch ops
REGISTER_OP(SIGNALRETURN, SignalReturn);
REGISTER_OP(CALLBACKRETURN, CallbackReturn);
REGISTER_OP(EXITFUNCTION, ExitFunction);
REGISTER_OP(JUMP, Jump);
@@ -124,10 +118,12 @@ constexpr OpHandlerArray InterpreterOpHandlers = [] {
REGISTER_OP(VALIDATECODE, ValidateCode);
REGISTER_OP(THREADREMOVECODEENTRY, ThreadRemoveCodeEntry);
REGISTER_OP(CPUID, CPUID);
REGISTER_OP(XGETBV, XGETBV);
// Conversion ops
REGISTER_OP(VINSGPR, VInsGPR);
REGISTER_OP(VCASTFROMGPR, VCastFromGPR);
REGISTER_OP(VDUPFROMGPR, VDupFromGPR);
REGISTER_OP(FLOAT_FROMGPR_S, Float_FromGPR_S);
REGISTER_OP(FLOAT_FTOF, Float_FToF);
REGISTER_OP(VECTOR_STOF, Vector_SToF);
@@ -154,6 +150,10 @@ constexpr OpHandlerArray InterpreterOpHandlers = [] {
REGISTER_OP(STOREMEM, StoreMem);
REGISTER_OP(LOADMEMTSO, LoadMem);
REGISTER_OP(STOREMEMTSO, StoreMem);
REGISTER_OP(VLOADVECTORMASKED, VLoadVectorMasked);
REGISTER_OP(VSTOREVECTORMASKED, VStoreVectorMasked);
REGISTER_OP(MEMSET, MemSet);
REGISTER_OP(MEMCPY, MemCpy);
REGISTER_OP(CACHELINECLEAR, CacheLineClear);
REGISTER_OP(CACHELINECLEAN, CacheLineClean);
REGISTER_OP(CACHELINEZERO, CacheLineZero);
@@ -222,6 +222,8 @@ constexpr OpHandlerArray InterpreterOpHandlers = [] {
REGISTER_OP(VZIP2, VZip);
REGISTER_OP(VUNZIP, VUnZip);
REGISTER_OP(VUNZIP2, VUnZip);
REGISTER_OP(VTRN, VTrn);
REGISTER_OP(VTRN2, VTrn);
REGISTER_OP(VBSL, VBSL);
REGISTER_OP(VCMPEQ, VCMPEQ);
REGISTER_OP(VCMPEQZ, VCMPEQZ);
@@ -266,6 +268,8 @@ constexpr OpHandlerArray InterpreterOpHandlers = [] {
REGISTER_OP(VUABDL, VUABDL);
REGISTER_OP(VTBL1, VTBL1);
REGISTER_OP(VREV64, VRev64);
REGISTER_OP(VPCMPESTRX, VPCMPESTRX);
REGISTER_OP(VPCMPISTRX, VPCMPISTRX);
// Encryption ops
REGISTER_OP(VAESIMC, AESImc);
@@ -36,6 +36,8 @@ namespace FEXCore::CPU {
FABI_I64_F80_F80,
FABI_F80_F80,
FABI_F80_F80_F80,
FABI_I32_I64_I64_I128_I128_I16,
FABI_I32_I128_I128_I16,
};
struct FallbackInfo {
@@ -142,7 +144,6 @@ namespace FEXCore::CPU {
DEF_OP(AtomicFetchNeg);
///< Branch ops
DEF_OP(SignalReturn);
DEF_OP(CallbackReturn);
DEF_OP(ExitFunction);
DEF_OP(Jump);
@@ -153,10 +154,12 @@ namespace FEXCore::CPU {
DEF_OP(ValidateCode);
DEF_OP(ThreadRemoveCodeEntry);
DEF_OP(CPUID);
DEF_OP(XGETBV);
///< Conversion ops
DEF_OP(VInsGPR);
DEF_OP(VCastFromGPR);
DEF_OP(VDupFromGPR);
DEF_OP(Float_FromGPR_S);
DEF_OP(Float_FToF);
DEF_OP(Vector_SToF);
@@ -181,6 +184,10 @@ namespace FEXCore::CPU {
DEF_OP(StoreFlag);
DEF_OP(LoadMem);
DEF_OP(StoreMem);
DEF_OP(VLoadVectorMasked);
DEF_OP(VStoreVectorMasked);
DEF_OP(MemSet);
DEF_OP(MemCpy);
DEF_OP(CacheLineClear);
DEF_OP(CacheLineClean);
DEF_OP(CacheLineZero);
@@ -242,6 +249,7 @@ namespace FEXCore::CPU {
DEF_OP(VSMax);
DEF_OP(VZip);
DEF_OP(VUnZip);
DEF_OP(VTrn);
DEF_OP(VBSL);
DEF_OP(VCMPEQ);
DEF_OP(VCMPEQZ);
@@ -286,6 +294,8 @@ namespace FEXCore::CPU {
DEF_OP(VUABDL);
DEF_OP(VTBL1);
DEF_OP(VRev64);
DEF_OP(VPCMPESTRX);
DEF_OP(VPCMPISTRX);
///< Encryption ops
DEF_OP(AESImc);
@@ -288,6 +288,366 @@ DEF_OP(StoreMem) {
}
}
DEF_OP(VLoadVectorMasked) {
const auto Op = IROp->C<IR::IROp_VLoadVectorMasked>();
const auto OpSize = IROp->Size;
const auto ElementSize = IROp->ElementSize;
const auto NumElements = OpSize / ElementSize;
const auto *MemData = *GetSrc<uint8_t const**>(Data->SSAData, Op->Addr);
const auto *Mask = GetSrc<uint8_t const*>(Data->SSAData, Op->Mask);
const auto SetElements = [NumElements]<typename T>(void* Dst, const T* MaskValues, const T* MemoryData) {
const auto SignBit = 1ULL << ((sizeof(T) * 8) - 1);
for (size_t i = 0; i < NumElements; i++) {
if ((MaskValues[i] & SignBit) != 0) {
std::memcpy(static_cast<uint8_t*>(Dst) + (i * sizeof(T)), MemoryData + i, sizeof(T));
}
}
};
if (!Op->Offset.IsInvalid()) {
auto Offset = *GetSrc<uintptr_t const*>(Data->SSAData, Op->Offset) * Op->OffsetScale;
switch(Op->OffsetType.Val) {
case IR::MEM_OFFSET_SXTX.Val: MemData += Offset; break;
case IR::MEM_OFFSET_UXTW.Val: MemData += (uint32_t)Offset; break;
case IR::MEM_OFFSET_SXTW.Val: MemData += (int32_t)Offset; break;
}
}
memset(GDP, 0, Core::CPUState::XMM_AVX_REG_SIZE);
switch (ElementSize) {
case 1: {
SetElements(GDP, Mask, MemData);
return;
}
case 2: {
SetElements(GDP,
reinterpret_cast<const uint16_t*>(Mask),
reinterpret_cast<const uint16_t*>(MemData));
return;
}
case 4: {
SetElements(GDP,
reinterpret_cast<const uint32_t*>(Mask),
reinterpret_cast<const uint32_t*>(MemData));
return;
}
case 8: {
SetElements(GDP,
reinterpret_cast<const uint64_t*>(Mask),
reinterpret_cast<const uint64_t*>(MemData));
return;
}
default:
LOGMAN_MSG_A_FMT("Unhandled VLoadVectorMasked element size: {}", ElementSize);
return;
}
}
DEF_OP(VStoreVectorMasked) {
const auto Op = IROp->C<IR::IROp_VStoreVectorMasked>();
const auto OpSize = IROp->Size;
const auto ElementSize = IROp->ElementSize;
const auto NumElements = OpSize / ElementSize;
auto *Dst = *GetSrc<uint8_t**>(Data->SSAData, Op->Addr);
const auto *RegData = GetSrc<uint8_t const*>(Data->SSAData, Op->Data);
const auto *Mask = GetSrc<uint8_t const*>(Data->SSAData, Op->Mask);
const auto SetElements = [NumElements]<typename T>(void* Dst, const T* MaskValues, const T* DataVals) {
const auto SignBit = 1ULL << ((sizeof(T) * 8) - 1);
for (size_t i = 0; i < NumElements; i++) {
if ((MaskValues[i] & SignBit) != 0) {
std::memcpy(static_cast<uint8_t*>(Dst) + (i * sizeof(T)), DataVals + i, sizeof(T));
}
}
};
if (!Op->Offset.IsInvalid()) {
auto Offset = *GetSrc<uintptr_t const*>(Data->SSAData, Op->Offset) * Op->OffsetScale;
switch(Op->OffsetType.Val) {
case IR::MEM_OFFSET_SXTX.Val: Dst += Offset; break;
case IR::MEM_OFFSET_UXTW.Val: Dst += (uint32_t)Offset; break;
case IR::MEM_OFFSET_SXTW.Val: Dst += (int32_t)Offset; break;
}
}
switch (ElementSize) {
case 1: {
SetElements(Dst, Mask, RegData);
return;
}
case 2: {
SetElements(Dst,
reinterpret_cast<const uint16_t*>(Mask),
reinterpret_cast<const uint16_t*>(RegData));
return;
}
case 4: {
SetElements(Dst,
reinterpret_cast<const uint32_t*>(Mask),
reinterpret_cast<const uint32_t*>(RegData));
return;
}
case 8: {
SetElements(Dst,
reinterpret_cast<const uint64_t*>(Mask),
reinterpret_cast<const uint64_t*>(RegData));
return;
}
default:
LOGMAN_MSG_A_FMT("Unhandled VStoreVectorMasked element size: {}", ElementSize);
return;
}
}
DEF_OP(MemSet) {
const auto Op = IROp->C<IR::IROp_MemSet>();
const int32_t Size = Op->Size;
char *MemData = *GetSrc<char **>(Data->SSAData, Op->Addr);
uint64_t MemPrefix{};
if (!Op->Prefix.IsInvalid()) {
MemPrefix = *GetSrc<uint64_t*>(Data->SSAData, Op->Prefix);
}
const auto Value = *GetSrc<uint64_t*>(Data->SSAData, Op->Value);
const auto Length = *GetSrc<uint64_t*>(Data->SSAData, Op->Length);
const auto Direction = *GetSrc<uint8_t*>(Data->SSAData, Op->Direction);
auto MemSetElements = [](auto* Memory, uint64_t Value, size_t Length) {
for (size_t i = 0; i < Length; ++i) {
Memory[i] = Value;
}
};
auto MemSetElementsInverse = [](auto* Memory, uint64_t Value, size_t Length) {
for (size_t i = 0; i < Length; ++i) {
Memory[-i] = Value;
}
};
if (Direction == 0) { // Forward
if (Op->IsAtomic) {
switch (Size) {
case 1:
MemSetElements(reinterpret_cast<std::atomic<uint8_t>*>(MemData + MemPrefix), Value, Length);
break;
case 2:
MemSetElements(reinterpret_cast<std::atomic<uint16_t>*>(MemData + MemPrefix), Value, Length);
break;
case 4:
MemSetElements(reinterpret_cast<std::atomic<uint32_t>*>(MemData + MemPrefix), Value, Length);
break;
case 8:
MemSetElements(reinterpret_cast<std::atomic<uint64_t>*>(MemData + MemPrefix), Value, Length);
break;
default:
LOGMAN_MSG_A_FMT("Unhandled {} size: {}", __func__, Size);
break;
}
}
else {
switch (Size) {
case 1:
MemSetElements(reinterpret_cast<uint8_t*>(MemData + MemPrefix), Value, Length);
break;
case 2:
MemSetElements(reinterpret_cast<uint16_t*>(MemData + MemPrefix), Value, Length);
break;
case 4:
MemSetElements(reinterpret_cast<uint32_t*>(MemData + MemPrefix), Value, Length);
break;
case 8:
MemSetElements(reinterpret_cast<uint64_t*>(MemData + MemPrefix), Value, Length);
break;
default:
LOGMAN_MSG_A_FMT("Unhandled {} size: {}", __func__, Size);
break;
}
}
GD = reinterpret_cast<uint64_t>(MemData + (Length * Size));
}
else { // Backward
if (Op->IsAtomic) {
switch (Size) {
case 1:
MemSetElementsInverse(reinterpret_cast<std::atomic<uint8_t>*>(MemData + MemPrefix), Value, Length);
break;
case 2:
MemSetElementsInverse(reinterpret_cast<std::atomic<uint16_t>*>(MemData + MemPrefix), Value, Length);
break;
case 4:
MemSetElementsInverse(reinterpret_cast<std::atomic<uint32_t>*>(MemData + MemPrefix), Value, Length);
break;
case 8:
MemSetElementsInverse(reinterpret_cast<std::atomic<uint64_t>*>(MemData + MemPrefix), Value, Length);
break;
default:
LOGMAN_MSG_A_FMT("Unhandled {} size: {}", __func__, Size);
break;
}
}
else {
switch (Size) {
case 1:
MemSetElementsInverse(reinterpret_cast<uint8_t*>(MemData + MemPrefix), Value, Length);
break;
case 2:
MemSetElementsInverse(reinterpret_cast<uint16_t*>(MemData + MemPrefix), Value, Length);
break;
case 4:
MemSetElementsInverse(reinterpret_cast<uint32_t*>(MemData + MemPrefix), Value, Length);
break;
case 8:
MemSetElementsInverse(reinterpret_cast<uint64_t*>(MemData + MemPrefix), Value, Length);
break;
default:
LOGMAN_MSG_A_FMT("Unhandled {} size: {}", __func__, Size);
break;
}
}
GD = reinterpret_cast<uint64_t>(MemData - (Length * Size));
}
}
DEF_OP(MemCpy) {
const auto Op = IROp->C<IR::IROp_MemCpy>();
const int32_t Size = Op->Size;
uint64_t *DstPtr = GetDest<uint64_t*>(Data->SSAData, Node);
char *MemDataDest = *GetSrc<char **>(Data->SSAData, Op->AddrDest);
char *MemDataSrc = *GetSrc<char **>(Data->SSAData, Op->AddrSrc);
uint64_t DestPrefix{};
uint64_t SrcPrefix{};
if (!Op->PrefixDest.IsInvalid()) {
DestPrefix = *GetSrc<uint64_t*>(Data->SSAData, Op->PrefixDest);
}
if (!Op->PrefixSrc.IsInvalid()) {
SrcPrefix = *GetSrc<uint64_t*>(Data->SSAData, Op->PrefixSrc);
}
const auto Length = *GetSrc<uint64_t*>(Data->SSAData, Op->Length);
const auto Direction = *GetSrc<uint8_t*>(Data->SSAData, Op->Direction);
auto MemSetElementsAtomic = [](auto* MemDst, auto* MemSrc, size_t Length) {
for (size_t i = 0; i < Length; ++i) {
MemDst[i].store(MemSrc[i].load());
}
};
auto MemSetElementsAtomicInverse = [](auto* MemDst, auto* MemSrc, size_t Length) {
for (size_t i = 0; i < Length; ++i) {
MemDst[-i].store(MemSrc[-i].load());
}
};
auto MemSetElements = [](auto* MemDst, auto* MemSrc, size_t Length) {
for (size_t i = 0; i < Length; ++i) {
MemDst[i] = MemSrc[i];
}
};
auto MemSetElementsInverse = [](auto* MemDst, auto* MemSrc, size_t Length) {
for (size_t i = 0; i < Length; ++i) {
MemDst[-i] = MemSrc[-i];
}
};
if (Direction == 0) { // Forward
if (Op->IsAtomic) {
switch (Size) {
case 1:
MemSetElementsAtomic(reinterpret_cast<std::atomic<uint8_t>*>(MemDataDest + DestPrefix), reinterpret_cast<std::atomic<uint8_t>*>(MemDataSrc + SrcPrefix), Length);
break;
case 2:
MemSetElementsAtomic(reinterpret_cast<std::atomic<uint16_t>*>(MemDataDest + DestPrefix), reinterpret_cast<std::atomic<uint16_t>*>(MemDataSrc + SrcPrefix), Length);
break;
case 4:
MemSetElementsAtomic(reinterpret_cast<std::atomic<uint32_t>*>(MemDataDest + DestPrefix), reinterpret_cast<std::atomic<uint32_t>*>(MemDataSrc + SrcPrefix), Length);
break;
case 8:
MemSetElementsAtomic(reinterpret_cast<std::atomic<uint64_t>*>(MemDataDest + DestPrefix), reinterpret_cast<std::atomic<uint64_t>*>(MemDataSrc + SrcPrefix), Length);
break;
default:
LOGMAN_MSG_A_FMT("Unhandled {} size: {}", __func__, Size);
break;
}
}
else {
switch (Size) {
case 1:
MemSetElements(reinterpret_cast<uint8_t*>(MemDataDest + DestPrefix), reinterpret_cast<uint8_t*>(MemDataSrc + SrcPrefix), Length);
break;
case 2:
MemSetElements(reinterpret_cast<uint16_t*>(MemDataDest + DestPrefix), reinterpret_cast<uint16_t*>(MemDataSrc + SrcPrefix), Length);
break;
case 4:
MemSetElements(reinterpret_cast<uint32_t*>(MemDataDest + DestPrefix), reinterpret_cast<uint32_t*>(MemDataSrc + SrcPrefix), Length);
break;
case 8:
MemSetElements(reinterpret_cast<uint64_t*>(MemDataDest + DestPrefix), reinterpret_cast<uint64_t*>(MemDataSrc + SrcPrefix), Length);
break;
default:
LOGMAN_MSG_A_FMT("Unhandled {} size: {}", __func__, Size);
break;
}
}
DstPtr[0] = reinterpret_cast<uint64_t>(MemDataDest + (Length * Size));
DstPtr[1] = reinterpret_cast<uint64_t>(MemDataSrc + (Length * Size));
}
else { // Backward
if (Op->IsAtomic) {
switch (Size) {
case 1:
MemSetElementsAtomicInverse(reinterpret_cast<std::atomic<uint8_t>*>(MemDataDest + DestPrefix), reinterpret_cast<std::atomic<uint8_t>*>(MemDataSrc + SrcPrefix), Length);
break;
case 2:
MemSetElementsAtomicInverse(reinterpret_cast<std::atomic<uint16_t>*>(MemDataDest + DestPrefix), reinterpret_cast<std::atomic<uint16_t>*>(MemDataSrc + SrcPrefix), Length);
break;
case 4:
MemSetElementsAtomicInverse(reinterpret_cast<std::atomic<uint32_t>*>(MemDataDest + DestPrefix), reinterpret_cast<std::atomic<uint32_t>*>(MemDataSrc + SrcPrefix), Length);
break;
case 8:
MemSetElementsAtomicInverse(reinterpret_cast<std::atomic<uint64_t>*>(MemDataDest + DestPrefix), reinterpret_cast<std::atomic<uint64_t>*>(MemDataSrc + SrcPrefix), Length);
break;
default:
LOGMAN_MSG_A_FMT("Unhandled {} size: {}", __func__, Size);
break;
}
}
else {
switch (Size) {
case 1:
MemSetElementsInverse(reinterpret_cast<uint8_t*>(MemDataDest + DestPrefix), reinterpret_cast<uint8_t*>(MemDataSrc + SrcPrefix), Length);
break;
case 2:
MemSetElementsInverse(reinterpret_cast<uint16_t*>(MemDataDest + DestPrefix), reinterpret_cast<uint16_t*>(MemDataSrc + SrcPrefix), Length);
break;
case 4:
MemSetElementsInverse(reinterpret_cast<uint32_t*>(MemDataDest + DestPrefix), reinterpret_cast<uint32_t*>(MemDataSrc + SrcPrefix), Length);
break;
case 8:
MemSetElementsInverse(reinterpret_cast<uint64_t*>(MemDataDest + DestPrefix), reinterpret_cast<uint64_t*>(MemDataSrc + SrcPrefix), Length);
break;
default:
LOGMAN_MSG_A_FMT("Unhandled {} size: {}", __func__, Size);
break;
}
}
DstPtr[0] = reinterpret_cast<uint64_t>(MemDataDest - (Length * Size));
DstPtr[1] = reinterpret_cast<uint64_t>(MemDataSrc - (Length * Size));
}
}
DEF_OP(CacheLineClear) {
auto Op = IROp->C<IR::IROp_CacheLineClear>();
@@ -8,6 +8,8 @@ $end_info$
#include "Interface/Core/Interpreter/InterpreterOps.h"
#include "Interface/Core/Interpreter/InterpreterDefines.h"
#include "Interface/Core/Interpreter/Fallbacks/VectorFallbacks.h"
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Utils/BitUtils.h>
@@ -902,6 +904,67 @@ DEF_OP(VZip) {
memcpy(GDP, Tmp, OpSize);
}
DEF_OP(VTrn) {
const auto Op = IROp->C<IR::IROp_VTrn>();
const uint8_t OpSize = IROp->Size;
void *Src1 = GetSrc<void*>(Data->SSAData, Op->VectorLower);
void *Src2 = GetSrc<void*>(Data->SSAData, Op->VectorUpper);
uint8_t Tmp[Core::CPUState::XMM_AVX_REG_SIZE]{};
const uint8_t ElementSize = Op->Header.ElementSize;
uint8_t Elements = OpSize / ElementSize;
const uint8_t BaseOffset = IROp->Op == IR::OP_VTRN2 ? 1 : 0;
Elements >>= 1;
switch (ElementSize) {
case 1: {
auto *Dst_d = reinterpret_cast<uint8_t*>(Tmp);
auto *Src1_d = reinterpret_cast<uint8_t*>(Src1);
auto *Src2_d = reinterpret_cast<uint8_t*>(Src2);
for (unsigned i = 0; i < Elements; ++i) {
Dst_d[i*2] = Src1_d[i*2 + BaseOffset];
Dst_d[i*2+1] = Src2_d[i*2 + BaseOffset];
}
break;
}
case 2: {
auto *Dst_d = reinterpret_cast<uint16_t*>(Tmp);
auto *Src1_d = reinterpret_cast<uint16_t*>(Src1);
auto *Src2_d = reinterpret_cast<uint16_t*>(Src2);
for (unsigned i = 0; i < Elements; ++i) {
Dst_d[i*2] = Src1_d[i*2 + BaseOffset];
Dst_d[i*2+1] = Src2_d[i*2 + BaseOffset];
}
break;
}
case 4: {
auto *Dst_d = reinterpret_cast<uint32_t*>(Tmp);
auto *Src1_d = reinterpret_cast<uint32_t*>(Src1);
auto *Src2_d = reinterpret_cast<uint32_t*>(Src2);
for (unsigned i = 0; i < Elements; ++i) {
Dst_d[i*2] = Src1_d[i*2 + BaseOffset];
Dst_d[i*2+1] = Src2_d[i*2 + BaseOffset];
}
break;
}
case 8: {
auto *Dst_d = reinterpret_cast<uint64_t*>(Tmp);
auto *Src1_d = reinterpret_cast<uint64_t*>(Src1);
auto *Src2_d = reinterpret_cast<uint64_t*>(Src2);
for (unsigned i = 0; i < Elements; ++i) {
Dst_d[i*2] = Src1_d[i*2 + BaseOffset];
Dst_d[i*2+1] = Src2_d[i*2 + BaseOffset];
}
break;
}
default:
LOGMAN_MSG_A_FMT("Unknown Element Size: {}", ElementSize);
break;
}
memcpy(GDP, Tmp, OpSize);
}
DEF_OP(VUnZip) {
const auto Op = IROp->C<IR::IROp_VUnZip>();
const uint8_t OpSize = IROp->Size;
@@ -964,7 +1027,9 @@ DEF_OP(VUnZip) {
}
DEF_OP(VBSL) {
auto Op = IROp->C<IR::IROp_VBSL>();
const auto Op = IROp->C<IR::IROp_VBSL>();
const auto OpSize = IROp->Size;
const auto Src1 = *GetSrc<InterpVector256*>(Data->SSAData, Op->VectorMask);
const auto Src2 = *GetSrc<InterpVector256*>(Data->SSAData, Op->VectorTrue);
const auto Src3 = *GetSrc<InterpVector256*>(Data->SSAData, Op->VectorFalse);
@@ -974,7 +1039,8 @@ DEF_OP(VBSL) {
.Upper = (Src2.Upper & Src1.Upper) | (Src3.Upper & ~Src1.Upper),
};
memcpy(GDP, &Tmp, sizeof(Tmp));
memset(GDP, 0, sizeof(InterpVector256));
memcpy(GDP, &Tmp, OpSize);
}
DEF_OP(VCMPEQ) {
@@ -2212,6 +2278,37 @@ DEF_OP(VRev64) {
memcpy(GDP, Tmp, OpSize);
}
DEF_OP(VPCMPESTRX) {
const auto Op = IROp->C<IR::IROp_VPCMPESTRX>();
const auto Is64Bit = Op->GPRSize == 8;
const auto RAX = *GetSrc<uint64_t*>(Data->SSAData, Op->RAX);
const auto RDX = *GetSrc<uint64_t*>(Data->SSAData, Op->RDX);
const auto LHS = *GetSrc<__uint128_t*>(Data->SSAData, Op->LHS);
const auto RHS = *GetSrc<__uint128_t*>(Data->SSAData, Op->RHS);
// We can be cheeky and encode the size at bit 8 to save a parameter
const auto Control = Op->Control | (uint16_t(Is64Bit) << 8);
const auto Result = OpHandlers<IR::OP_VPCMPESTRX>::handle(RAX, RDX, LHS, RHS, Control);
memset(GDP, 0, sizeof(uint64_t));
memcpy(GDP, &Result, sizeof(Result));
}
DEF_OP(VPCMPISTRX) {
const auto Op = IROp->C<IR::IROp_VPCMPISTRX>();
const auto LHS = *GetSrc<__uint128_t*>(Data->SSAData, Op->LHS);
const auto RHS = *GetSrc<__uint128_t*>(Data->SSAData, Op->RHS);
const auto Control = Op->Control;
const auto Result = OpHandlers<IR::OP_VPCMPISTRX>::handle(LHS, RHS, Control);
memset(GDP, 0, sizeof(uint64_t));
memcpy(GDP, &Result, sizeof(Result));
}
#undef DEF_OP
} // namespace FEXCore::CPU
+10 -10
View File
@@ -262,13 +262,13 @@ DEF_OP(MulH) {
const auto Src2 = GetReg(Op->Src2.ID());
if (OpSize == 4) {
sxtw(TMP1, Src1);
sxtw(TMP2, Src2);
sxtw(TMP1, Src1.W());
sxtw(TMP2, Src2.W());
mul(ARMEmitter::Size::i32Bit, Dst, TMP1, TMP2);
ubfx(ARMEmitter::Size::i32Bit, Dst, Dst, 32, 32);
}
else {
smulh(Dst, Src1, Src2);
smulh(Dst.X(), Src1.X(), Src2.X());
}
}
@@ -289,7 +289,7 @@ DEF_OP(UMulH) {
ubfx(ARMEmitter::Size::i64Bit, Dst, Dst, 32, 32);
}
else {
umulh(Dst, Src1, Src2);
umulh(Dst.X(), Src1.X(), Src2.X());
}
}
@@ -494,7 +494,7 @@ DEF_OP(PDep) {
// We sadly need to spill regs for this for the time being
// TODO: Remove when scratch registers can be allocated
// explicitly.
SpillStaticRegs(false, SpillCode);
SpillStaticRegs(TMP1, false, SpillCode);
mov(EmitSize, InputReg, Input);
@@ -558,7 +558,7 @@ DEF_OP(PExt) {
// We sadly need to spill a reg for this for the time being
// TODO: Remove when scratch registers can be allocated
// explicitly.
SpillStaticRegs(false, 1U << Mask.Idx());
SpillStaticRegs(TMP2, false, 1U << Mask.Idx());
mov(EmitSize, Mask, ZeroReg);
// Main loop
@@ -610,7 +610,7 @@ DEF_OP(LDiv) {
case 4: {
mov(EmitSize, TMP1, Lower);
bfi(EmitSize, TMP1, Upper, 32, 32);
sxtw(TMP2, Divisor);
sxtw(TMP2, Divisor.W());
sdiv(EmitSize, Dst, TMP1, TMP2);
break;
}
@@ -744,7 +744,7 @@ DEF_OP(LRem) {
case 4: {
mov(EmitSize, TMP1, Lower);
bfi(EmitSize, TMP1, Upper, 32, 32);
sxtw(TMP3, Divisor);
sxtw(TMP3, Divisor.W());
sdiv(EmitSize, TMP2, TMP1, TMP3);
msub(EmitSize, Dst, TMP2, TMP3, TMP1);
break;
@@ -1173,8 +1173,8 @@ DEF_OP(VExtractToGPR) {
// Inverting our dedicated predicate for 128-bit operations selects
// all of the top lanes. We can then compact those into a temporary.
const auto CompactPred = ARMEmitter::PReg::p0;
not_(CompactPred, PRED_TMP_32B, PRED_TMP_16B);
compact(ARMEmitter::SubRegSize::i64Bit, VTMP1, CompactPred, Vector);
not_(CompactPred, PRED_TMP_32B.Zeroing(), PRED_TMP_16B);
compact(ARMEmitter::SubRegSize::i64Bit, VTMP1.Z(), CompactPred, Vector.Z());
// Sanitize the zero-based index to work on the now-moved
// upper half of the vector.
@@ -27,7 +27,7 @@ void Arm64JITCore::InsertNamedThunkRelocation(ARMEmitter::Register Reg, const IR
MoveABI.NamedThunkMove.Header.Type = FEXCore::CPU::RelocationTypes::RELOC_NAMED_THUNK_MOVE;
// Offset is the offset from the entrypoint of the block
auto CurrentCursor = GetCursorAddress<uint8_t *>();
MoveABI.NamedThunkMove.Offset = CurrentCursor - GuestEntry;
MoveABI.NamedThunkMove.Offset = CurrentCursor - CodeData.BlockBegin;
MoveABI.NamedThunkMove.Symbol = Sum;
MoveABI.NamedThunkMove.RegisterIndex = Reg.Idx();
@@ -58,7 +58,7 @@ Arm64JITCore::NamedSymbolLiteralPair Arm64JITCore::InsertNamedSymbolLiteral(FEXC
void Arm64JITCore::PlaceNamedSymbolLiteral(NamedSymbolLiteralPair &Lit) {
// Offset is the offset from the entrypoint of the block
auto CurrentCursor = GetCursorAddress<uint8_t *>();
Lit.MoveABI.NamedSymbolLiteral.Offset = CurrentCursor - GuestEntry;
Lit.MoveABI.NamedSymbolLiteral.Offset = CurrentCursor - CodeData.BlockBegin;
Bind(&Lit.Loc);
dc64(Lit.Lit);
@@ -70,7 +70,7 @@ void Arm64JITCore::InsertGuestRIPMove(ARMEmitter::Register Reg, uint64_t Constan
MoveABI.GuestRIPMove.Header.Type = FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_MOVE;
// Offset is the offset from the entrypoint of the block
auto CurrentCursor = GetCursorAddress<uint8_t *>();
MoveABI.GuestRIPMove.Offset = CurrentCursor - GuestEntry;
MoveABI.GuestRIPMove.Offset = CurrentCursor - CodeData.BlockBegin;
MoveABI.GuestRIPMove.GuestRIP = Constant;
MoveABI.GuestRIPMove.RegisterIndex = Reg.Idx();
@@ -20,26 +20,9 @@ $end_info$
namespace FEXCore::CPU {
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const *IROp, IR::NodeID Node)
DEF_OP(SignalReturn) {
auto Op = IROp->C<IR::IROp_SignalReturn>();
// First we must reset the stack
ResetStack();
// Now branch to our signal return helper
// This can't be a direct branch since the code needs to live at a constant location
if (Op->IsRT) {
ldr(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.SignalReturnHandlerRT));
}
else {
ldr(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.SignalReturnHandler));
}
br(ARMEmitter::Reg::r0);
}
DEF_OP(CallbackReturn) {
// spill back to CTX
SpillStaticRegs();
SpillStaticRegs(TMP1);
// First we must reset the stack
ResetStack();
@@ -184,14 +167,23 @@ DEF_OP(Syscall) {
FEXCore::IR::SyscallFlags Flags = Op->Flags;
PushDynamicRegsAndLR(TMP1);
if ((Flags & FEXCore::IR::SyscallFlags::NOSYNCSTATEONENTRY) != FEXCore::IR::SyscallFlags::NOSYNCSTATEONENTRY) {
SpillStaticRegs();
}
else {
uint32_t GPRSpillMask = ~0U;
uint32_t FPRSpillMask = ~0U;
if ((Flags & FEXCore::IR::SyscallFlags::NOSYNCSTATEONENTRY) == FEXCore::IR::SyscallFlags::NOSYNCSTATEONENTRY) {
// Need to spill all caller saved registers still
SpillStaticRegs(true, CALLER_GPR_MASK, CALLER_FPR_MASK);
GPRSpillMask = CALLER_GPR_MASK;
FPRSpillMask = CALLER_FPR_MASK;
}
SpillStaticRegs(TMP1, true, GPRSpillMask, FPRSpillMask);
// Now that we are spilled, store in the state that we are in a syscall
// Still without overwriting registers that matter
// 16bit LoadConstant to be a single instruction
// This gives the signal handler a value to check to see if we are in a syscall at all
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, GPRSpillMask & 0xFFFF);
str(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CpuStateFrame, InSyscallInfo));
uint64_t SPOffset = AlignUp(FEXCore::HLE::SyscallArguments::MAX_ARGS * 8, 16);
sub(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, ARMEmitter::Reg::rsp, SPOffset);
for (uint32_t i = 0; i < FEXCore::HLE::SyscallArguments::MAX_ARGS; ++i) {
@@ -213,21 +205,22 @@ DEF_OP(Syscall) {
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, ARMEmitter::Reg::rsp, SPOffset);
if ((Flags & FEXCore::IR::SyscallFlags::NOSYNCSTATEONENTRY) != FEXCore::IR::SyscallFlags::NOSYNCSTATEONENTRY &&
(Flags & FEXCore::IR::SyscallFlags::NORETURN) != FEXCore::IR::SyscallFlags::NORETURN) {
FillStaticRegs();
}
else {
if ((Flags & FEXCore::IR::SyscallFlags::NORETURN) != FEXCore::IR::SyscallFlags::NORETURN) {
// Result is now in x0
// Fix the stack and any values that were stepped on
FillStaticRegs(true, CALLER_GPR_MASK, CALLER_FPR_MASK);
}
FillStaticRegs(true, GPRSpillMask, FPRSpillMask);
PopDynamicRegsAndLR();
// Now the registers we've spilled are back in their original host registers
// We can safely claim we are no longer in a syscall
str(ARMEmitter::XReg::zr, STATE, offsetof(FEXCore::Core::CpuStateFrame, InSyscallInfo));
if ((Flags & FEXCore::IR::SyscallFlags::NORETURN) != FEXCore::IR::SyscallFlags::NORETURN) {
// Move result to its destination register
mov(ARMEmitter::Size::i64Bit, GetReg(Node), ARMEmitter::Reg::r0);
PopDynamicRegsAndLR();
if ((Flags & FEXCore::IR::SyscallFlags::NORETURNEDRESULT) != FEXCore::IR::SyscallFlags::NORETURNEDRESULT) {
// Move result to its destination register.
// Only if `NORETURNEDRESULT` wasn't set, otherwise we might overwrite the CPUState refilled with `FillStaticRegs`
mov(ARMEmitter::Size::i64Bit, GetReg(Node), ARMEmitter::Reg::r0);
}
}
}
@@ -255,9 +248,9 @@ DEF_OP(InlineSyscall) {
if (Op->Header.Args[i].IsInvalid()) break;
auto Reg = GetReg(Op->Header.Args[i].ID());
if (Reg.Idx() == ARMEmitter::Reg::r8.Idx() ||
Reg.Idx() == ARMEmitter::Reg::r4.Idx() ||
Reg.Idx() == ARMEmitter::Reg::r5.Idx()) {
if (Reg == ARMEmitter::Reg::r8 ||
Reg == ARMEmitter::Reg::r4 ||
Reg == ARMEmitter::Reg::r5) {
SpillMask |= (1U << Reg.Idx());
Intersects = true;
@@ -267,7 +260,7 @@ DEF_OP(InlineSyscall) {
// Ordering is incredibly important here
// We must spill any overlapping registers first THEN claim we are in a syscall without invalidating state at all
// Only spill the registers that intersect with our usage
SpillStaticRegs(false, SpillMask);
SpillStaticRegs(TMP1, false, SpillMask);
// Now that we are spilled, store in the state that we are in a syscall
// Still without overwriting registers that matter
@@ -288,13 +281,13 @@ DEF_OP(InlineSyscall) {
// In the case of intersection with x4, x5, or x8 then these are currently SRA
// for registers RAX, RBX, and RSI. Which have just been spilled
// Just load back from the context. Could be slightly smarter but this is fairly uncommon
if (Reg.Idx() == FEXCore::ARMEmitter::Reg::r8.Idx()) {
if (Reg == ARMEmitter::Reg::r8) {
ldr(EmitSubSize, RegArgs[i].R(), STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RSI]));
}
else if (Reg.Idx() == FEXCore::ARMEmitter::Reg::r4.Idx()) {
else if (Reg == ARMEmitter::Reg::r4) {
ldr(EmitSubSize, RegArgs[i].R(), STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RAX]));
}
else if (Reg.Idx() == FEXCore::ARMEmitter::Reg::r5.Idx()) {
else if (Reg == ARMEmitter::Reg::r5) {
ldr(EmitSubSize, RegArgs[i].R(), STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RBX]));
}
else {
@@ -335,13 +328,13 @@ DEF_OP(Thunk) {
// X0: CTX
// X1: Args (from guest stack)
SpillStaticRegs(); // spill to ctx before ra64 spill
SpillStaticRegs(TMP1); // spill to ctx before ra64 spill
PushDynamicRegsAndLR(TMP1);
mov(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, GetReg(Op->ArgPtr.ID()));
auto thunkFn = ThreadState->CTX->ThunkHandler->LookupThunk(Op->ThunkNameHash);
auto thunkFn = static_cast<Context::ContextImpl*>(ThreadState->CTX)->ThunkHandler->LookupThunk(Op->ThunkNameHash);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r2, (uintptr_t)thunkFn);
#ifdef VIXL_SIMULATOR
GenerateIndirectRuntimeCall<void, void*, void*>(ARMEmitter::Reg::r2);
@@ -410,12 +403,12 @@ DEF_OP(ThreadRemoveCodeEntry) {
// X1: RIP
PushDynamicRegsAndLR(TMP1);
SpillStaticRegs(TMP1);
mov(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, STATE.R());
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r1, Entry);
ldr(ARMEmitter::XReg::x2, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.ThreadRemoveCodeEntryFromJIT));
SpillStaticRegs();
#ifdef VIXL_SIMULATOR
GenerateIndirectRuntimeCall<void, void*, void*>(ARMEmitter::Reg::r2);
#else
@@ -431,7 +424,7 @@ DEF_OP(CPUID) {
auto Op = IROp->C<IR::IROp_CPUID>();
PushDynamicRegsAndLR(TMP1);
SpillStaticRegs();
SpillStaticRegs(TMP1);
// x0 = CPUID Handler
// x1 = CPUID Function
@@ -457,6 +450,34 @@ DEF_OP(CPUID) {
mov(ARMEmitter::Size::i64Bit, Dst.second, ARMEmitter::Reg::r1);
}
DEF_OP(XGETBV) {
auto Op = IROp->C<IR::IROp_XGetBV>();
PushDynamicRegsAndLR(TMP1);
SpillStaticRegs(TMP1);
// x0 = CPUID Handler
// x1 = XCR Function
ldr(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.CPUIDObj));
ldr(ARMEmitter::XReg::x2, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.XCRFunction));
mov(ARMEmitter::Size::i32Bit, ARMEmitter::Reg::r1, GetReg(Op->Function.ID()));
#ifdef VIXL_SIMULATOR
GenerateIndirectRuntimeCall<uint64_t, void*, uint32_t>(ARMEmitter::Reg::r2);
#else
blr(ARMEmitter::Reg::r2);
#endif
FillStaticRegs();
PopDynamicRegsAndLR();
// Results are in x0
// Results want to be in a i32v2 vector
auto Dst = GetRegPair(Node);
mov(ARMEmitter::Size::i32Bit, Dst.first, ARMEmitter::Reg::r0);
lsr(ARMEmitter::Size::i64Bit, Dst.second, ARMEmitter::Reg::r0, 32);
}
#undef DEF_OP
}
@@ -55,7 +55,7 @@ DEF_OP(VInsGPR) {
// Move the upper lane down for the insertion.
const auto CompactPred = ARMEmitter::PReg::p0;
not_(CompactPred, PRED_TMP_32B.Zeroing(), PRED_TMP_16B);
compact(ARMEmitter::SubRegSize::i64Bit, VTMP1.Z(), CompactPred, DestVector);
compact(ARMEmitter::SubRegSize::i64Bit, VTMP1.Z(), CompactPred, DestVector.Z());
}
// Put data in place for destructive SPLICE below.
@@ -108,6 +108,32 @@ DEF_OP(VCastFromGPR) {
}
}
DEF_OP(VDupFromGPR) {
const auto Op = IROp->C<IR::IROp_VDupFromGPR>();
const auto OpSize = IROp->Size;
const auto Dst = GetVReg(Node);
const auto Src = GetReg(Op->Src.ID());
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto ElementSize = IROp->ElementSize;
LOGMAN_THROW_AA_FMT(ElementSize == 8 || ElementSize == 4 || ElementSize == 2 || ElementSize == 1,
"Unexpected {} element size: {}", __func__, ElementSize);
const auto SubEmitSize =
ElementSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
ElementSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
ElementSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
ElementSize == 1 ? ARMEmitter::SubRegSize::i8Bit : ARMEmitter::SubRegSize::i8Bit;
if (HostSupportsSVE && Is256Bit) {
dup(SubEmitSize, Dst.Z(), Src);
} else {
dup(SubEmitSize, Dst.Q(), Src);
}
}
DEF_OP(Float_FromGPR_S) {
const auto Op = IROp->C<IR::IROp_Float_FromGPR_S>();
@@ -199,7 +225,7 @@ DEF_OP(Vector_FToZS) {
const auto Vector = GetVReg(Op->Vector.ID());
if (HostSupportsSVE && Is256Bit) {
const auto Mask = PRED_TMP_32B;
fcvtzs(Dst, SubEmitSize, Mask.Merging(), Vector, SubEmitSize);
fcvtzs(Dst.Z(), SubEmitSize, Mask.Merging(), Vector.Z(), SubEmitSize);
} else {
fcvtzs(SubEmitSize, Dst.Q(), Vector.Q());
}
@@ -222,8 +248,8 @@ DEF_OP(Vector_FToS) {
if (HostSupportsSVE && Is256Bit) {
const auto Mask = PRED_TMP_32B;
frinti(SubEmitSize, Dst, Mask.Merging(), Vector);
fcvtzs(Dst, SubEmitSize, Mask.Merging(), Dst, SubEmitSize);
frinti(SubEmitSize, Dst.Z(), Mask.Merging(), Vector.Z());
fcvtzs(Dst.Z(), SubEmitSize, Mask.Merging(), Dst.Z(), SubEmitSize);
} else {
const auto Dst = GetVReg(Node);
const auto Vector = GetVReg(Op->Vector.ID());
@@ -276,12 +302,12 @@ DEF_OP(Vector_FToF) {
break;
}
case 0x0204: { // Half <- Float
fcvtnt(FEXCore::ARMEmitter::SubRegSize::i16Bit, Dst, Mask, Vector);
fcvtnt(FEXCore::ARMEmitter::SubRegSize::i16Bit, Dst.Z(), Mask, Vector.Z());
uzp2(FEXCore::ARMEmitter::SubRegSize::i16Bit, Dst.Z(), Dst.Z(), Dst.Z());
break;
}
case 0x0408: { // Float <- Double
fcvtnt(FEXCore::ARMEmitter::SubRegSize::i32Bit, Dst, Mask, Vector);
fcvtnt(FEXCore::ARMEmitter::SubRegSize::i32Bit, Dst.Z(), Mask, Vector.Z());
uzp2(FEXCore::ARMEmitter::SubRegSize::i32Bit, Dst.Z(), Dst.Z(), Dst.Z());
break;
}
@@ -17,37 +17,73 @@ DEF_OP(AESImc) {
}
DEF_OP(AESEnc) {
auto Op = IROp->C<IR::IROp_VAESEnc>();
const auto Op = IROp->C<IR::IROp_VAESEnc>();
const auto OpSize = IROp->Size;
const auto Dst = GetVReg(Node);
const auto Key = GetVReg(Op->Key.ID());
const auto State = GetVReg(Op->State.ID());
LOGMAN_THROW_AA_FMT(OpSize == Core::CPUState::XMM_SSE_REG_SIZE,
"Currently only supports 128-bit operations.");
eor(VTMP2.Q(), VTMP2.Q(), VTMP2.Q());
mov(VTMP1.Q(), GetVReg(Op->State.ID()).Q());
mov(VTMP1.Q(), State.Q());
aese(VTMP1, VTMP2);
aesmc(VTMP1, VTMP1);
eor(GetVReg(Node).Q(), VTMP1.Q(), GetVReg(Op->Key.ID()).Q());
eor(Dst.Q(), VTMP1.Q(), Key.Q());
}
DEF_OP(AESEncLast) {
auto Op = IROp->C<IR::IROp_VAESEncLast>();
const auto Op = IROp->C<IR::IROp_VAESEncLast>();
const auto OpSize = IROp->Size;
const auto Dst = GetVReg(Node);
const auto Key = GetVReg(Op->Key.ID());
const auto State = GetVReg(Op->State.ID());
LOGMAN_THROW_AA_FMT(OpSize == Core::CPUState::XMM_SSE_REG_SIZE,
"Currently only supports 128-bit operations.");
eor(VTMP2.Q(), VTMP2.Q(), VTMP2.Q());
mov(VTMP1.Q(), GetVReg(Op->State.ID()).Q());
mov(VTMP1.Q(), State.Q());
aese(VTMP1, VTMP2);
eor(GetVReg(Node).Q(), VTMP1.Q(), GetVReg(Op->Key.ID()).Q());
eor(Dst.Q(), VTMP1.Q(), Key.Q());
}
DEF_OP(AESDec) {
auto Op = IROp->C<IR::IROp_VAESDec>();
const auto Op = IROp->C<IR::IROp_VAESDec>();
const auto OpSize = IROp->Size;
const auto Dst = GetVReg(Node);
const auto Key = GetVReg(Op->Key.ID());
const auto State = GetVReg(Op->State.ID());
LOGMAN_THROW_AA_FMT(OpSize == Core::CPUState::XMM_SSE_REG_SIZE,
"Currently only supports 128-bit operations.");
eor(VTMP2.Q(), VTMP2.Q(), VTMP2.Q());
mov(VTMP1.Q(), GetVReg(Op->State.ID()).Q());
mov(VTMP1.Q(), State.Q());
aesd(VTMP1, VTMP2);
aesimc(VTMP1, VTMP1);
eor(GetVReg(Node).Q(), VTMP1.Q(), GetVReg(Op->Key.ID()).Q());
eor(Dst.Q(), VTMP1.Q(), Key.Q());
}
DEF_OP(AESDecLast) {
auto Op = IROp->C<IR::IROp_VAESDecLast>();
const auto Op = IROp->C<IR::IROp_VAESDecLast>();
const auto OpSize = IROp->Size;
const auto Dst = GetVReg(Node);
const auto Key = GetVReg(Op->Key.ID());
const auto State = GetVReg(Op->State.ID());
LOGMAN_THROW_AA_FMT(OpSize == Core::CPUState::XMM_SSE_REG_SIZE,
"Currently only supports 128-bit operations.");
eor(VTMP2.Q(), VTMP2.Q(), VTMP2.Q());
mov(VTMP1.Q(), GetVReg(Op->State.ID()).Q());
mov(VTMP1.Q(), State.Q());
aesd(VTMP1, VTMP2);
eor(GetVReg(Node).Q(), VTMP1.Q(), GetVReg(Op->Key.ID()).Q());
eor(Dst.Q(), VTMP1.Q(), Key.Q());
}
DEF_OP(AESKeyGenAssist) {
@@ -101,18 +137,22 @@ DEF_OP(CRC32) {
crc32cw(Dst.W(), Src1.W(), Src2.W());
break;
case 8:
crc32cx(Dst, Src1, Src2);
crc32cx(Dst.X(), Src1.X(), Src2.X());
break;
default: LOGMAN_MSG_A_FMT("Unknown CRC32 size: {}", Op->SrcSize);
}
}
DEF_OP(PCLMUL) {
auto Op = IROp->C<IR::IROp_PCLMUL>();
const auto Op = IROp->C<IR::IROp_PCLMUL>();
const auto OpSize = IROp->Size;
auto Dst = GetVReg(Node);
auto Src1 = GetVReg(Op->Src1.ID());
auto Src2 = GetVReg(Op->Src2.ID());
const auto Dst = GetVReg(Node);
const auto Src1 = GetVReg(Op->Src1.ID());
const auto Src2 = GetVReg(Op->Src2.ID());
LOGMAN_THROW_AA_FMT(OpSize == Core::CPUState::XMM_SSE_REG_SIZE,
"Currently only supports 128-bit operations.");
switch (Op->Selector) {
case 0b00000000:
+180 -90
View File
@@ -14,8 +14,6 @@ $end_info$
#include "Interface/Core/ArchHelpers/CodeEmitter/Emitter.h"
#include "Interface/Core/LookupCache.h"
#include "Interface/Core/ArchHelpers/Arm64.h"
#include "Interface/Core/ArchHelpers/MContext.h"
#include "Interface/Core/Dispatcher/Arm64Dispatcher.h"
#include "Interface/Core/JIT/Arm64/JITClass.h"
#include "Interface/Core/InternalThreadState.h"
@@ -25,7 +23,6 @@ $end_info$
#include "Utils/MemberFunctionToPointer.h"
#include <FEXCore/Core/X86Enums.h>
#include <FEXCore/Core/UContext.h>
#include <FEXCore/Utils/Allocator.h>
#include <FEXCore/Utils/CompilerDefs.h>
#include <FEXCore/Utils/EnumUtils.h>
@@ -33,7 +30,6 @@ $end_info$
#include "Interface/Core/Interpreter/InterpreterOps.h"
#include <sys/mman.h>
#include <stdio.h>
#include <unistd.h>
#include <string.h>
@@ -87,7 +83,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header const *IROp, IR::NodeID Node) {
} else {
switch(Info.ABI) {
case FABI_VOID_U16:{
SpillStaticRegs();
SpillStaticRegs(TMP1);
PushDynamicRegsAndLR(TMP1);
@@ -107,7 +103,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header const *IROp, IR::NodeID Node) {
break;
case FABI_F80_F32:{
SpillStaticRegs();
SpillStaticRegs(TMP1);
PushDynamicRegsAndLR(TMP1);
const auto Src1 = GetVReg(IROp->Args[0].ID());
@@ -131,7 +127,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header const *IROp, IR::NodeID Node) {
break;
case FABI_F80_F64:{
SpillStaticRegs();
SpillStaticRegs(TMP1);
PushDynamicRegsAndLR(TMP1);
@@ -157,7 +153,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header const *IROp, IR::NodeID Node) {
case FABI_F80_I16:
case FABI_F80_I32: {
SpillStaticRegs();
SpillStaticRegs(TMP1);
PushDynamicRegsAndLR(TMP1);
@@ -187,7 +183,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header const *IROp, IR::NodeID Node) {
break;
case FABI_F32_F80:{
SpillStaticRegs();
SpillStaticRegs(TMP1);
PushDynamicRegsAndLR(TMP1);
@@ -213,7 +209,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header const *IROp, IR::NodeID Node) {
break;
case FABI_F64_F80:{
SpillStaticRegs();
SpillStaticRegs(TMP1);
PushDynamicRegsAndLR(TMP1);
@@ -239,7 +235,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header const *IROp, IR::NodeID Node) {
break;
case FABI_F64_F64: {
SpillStaticRegs();
SpillStaticRegs(TMP1);
PushDynamicRegsAndLR(TMP1);
@@ -263,7 +259,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header const *IROp, IR::NodeID Node) {
break;
case FABI_F64_F64_F64: {
SpillStaticRegs();
SpillStaticRegs(TMP1);
PushDynamicRegsAndLR(TMP1);
@@ -289,7 +285,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header const *IROp, IR::NodeID Node) {
break;
case FABI_I16_F80:{
SpillStaticRegs();
SpillStaticRegs(TMP1);
PushDynamicRegsAndLR(TMP1);
@@ -314,7 +310,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header const *IROp, IR::NodeID Node) {
}
break;
case FABI_I32_F80:{
SpillStaticRegs();
SpillStaticRegs(TMP1);
PushDynamicRegsAndLR(TMP1);
@@ -339,7 +335,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header const *IROp, IR::NodeID Node) {
}
break;
case FABI_I64_F80:{
SpillStaticRegs();
SpillStaticRegs(TMP1);
PushDynamicRegsAndLR(TMP1);
@@ -364,7 +360,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header const *IROp, IR::NodeID Node) {
}
break;
case FABI_I64_F80_F80:{
SpillStaticRegs();
SpillStaticRegs(TMP1);
PushDynamicRegsAndLR(TMP1);
@@ -392,7 +388,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header const *IROp, IR::NodeID Node) {
}
break;
case FABI_F80_F80:{
SpillStaticRegs();
SpillStaticRegs(TMP1);
PushDynamicRegsAndLR(TMP1);
@@ -419,7 +415,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header const *IROp, IR::NodeID Node) {
}
break;
case FABI_F80_F80_F80:{
SpillStaticRegs();
SpillStaticRegs(TMP1);
PushDynamicRegsAndLR(TMP1);
@@ -449,6 +445,78 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header const *IROp, IR::NodeID Node) {
ins(ARMEmitter::SubRegSize::i16Bit, Dst, 4, ARMEmitter::Reg::r1);
}
break;
case FABI_I32_I64_I64_I128_I128_I16: {
SpillStaticRegs(TMP1);
PushDynamicRegsAndLR(TMP1);
const auto Op = IROp->C<IR::IROp_VPCMPESTRX>();
const auto Is64Bit = Op->GPRSize == 8;
const auto Src1 = GetVReg(Op->LHS.ID());
const auto Src2 = GetVReg(Op->RHS.ID());
const auto SrcRAX = GetReg(Op->RAX.ID());
const auto SrcRDX = GetReg(Op->RDX.ID());
// We can be cheeky and encode the size at bit 8 to save a parameter
const auto Control = Op->Control | (uint16_t(Is64Bit) << 8);
mov(ARMEmitter::XReg::x0, SrcRAX.X());
mov(ARMEmitter::XReg::x1, SrcRDX.X());
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r2, Src1, 0);
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r3, Src1, 1);
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r4, Src2, 0);
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r5, Src2, 1);
movz(ARMEmitter::Size::i32Bit, ARMEmitter::Reg::r6, Control);
ldr(ARMEmitter::XReg::x7, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
#ifdef VIXL_SIMULATOR
GenerateIndirectRuntimeCall<uint32_t, uint64_t, uint64_t, uint64_t, uint64_t, uint64_t, uint64_t, uint16_t>(ARMEmitter::Reg::r7);
#else
blr(ARMEmitter::Reg::r7);
#endif
PopDynamicRegsAndLR();
FillStaticRegs();
const auto Dst = GetReg(Node);
mov(Dst.W(), ARMEmitter::WReg::w0);
break;
}
case FABI_I32_I128_I128_I16: {
SpillStaticRegs(TMP1);
PushDynamicRegsAndLR(TMP1);
const auto Op = IROp->C<IR::IROp_VPCMPISTRX>();
const auto Src1 = GetVReg(Op->LHS.ID());
const auto Src2 = GetVReg(Op->RHS.ID());
const auto Control = Op->Control;
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r0, Src1, 0);
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r1, Src1, 1);
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r2, Src2, 0);
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r3, Src2, 1);
movz(ARMEmitter::Size::i32Bit, ARMEmitter::Reg::r4, Control);
ldr(ARMEmitter::XReg::x5, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
#ifdef VIXL_SIMULATOR
GenerateIndirectRuntimeCall<uint32_t, uint64_t, uint64_t, uint64_t, uint64_t, uint16_t>(ARMEmitter::Reg::r5);
#else
blr(ARMEmitter::Reg::r5);
#endif
PopDynamicRegsAndLR();
FillStaticRegs();
const auto Dst = GetReg(Node);
mov(Dst.W(), ARMEmitter::WReg::w0);
break;
}
case FABI_UNKNOWN:
default:
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
@@ -484,7 +552,7 @@ static uint64_t Arm64JITCore_ExitFunctionLink(FEXCore::Core::CpuStateFrame *Fram
FEXCore::ARMEmitter::Emitter::ClearICache((void*)branch, 24);
// Add de-linking handler
Context::Context::ThreadAddBlockLink(Thread, GuestRip, (uintptr_t)record, [branch, LinkerAddress]{
Context::ContextImpl::ThreadAddBlockLink(Thread, GuestRip, (uintptr_t)record, [branch, LinkerAddress]{
FEXCore::ARMEmitter::Emitter emit((uint8_t*)(branch), 24);
FEXCore::ARMEmitter::ForwardLabel l_BranchHost;
emit.ldr(FEXCore::ARMEmitter::XReg::x0, &l_BranchHost);
@@ -498,7 +566,7 @@ static uint64_t Arm64JITCore_ExitFunctionLink(FEXCore::Core::CpuStateFrame *Fram
record[0] = HostCode;
// Add de-linking handler
Context::Context::ThreadAddBlockLink(Thread, GuestRip, (uintptr_t)record, [record, LinkerAddress]{
Context::ContextImpl::ThreadAddBlockLink(Thread, GuestRip, (uintptr_t)record, [record, LinkerAddress]{
record[0] = LinkerAddress;
});
}
@@ -509,7 +577,7 @@ static uint64_t Arm64JITCore_ExitFunctionLink(FEXCore::Core::CpuStateFrame *Fram
void Arm64JITCore::Op_NoOp(IR::IROp_Header const *IROp, IR::NodeID Node) {
}
Arm64JITCore::Arm64JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread)
Arm64JITCore::Arm64JITCore(FEXCore::Context::ContextImpl *ctx, FEXCore::Core::InternalThreadState *Thread)
: CPUBackend(Thread, INITIAL_CODE_SIZE, MAX_CODE_SIZE)
, Arm64Emitter(ctx, 0)
, HostSupportsSVE{ctx->HostFeatures.SupportsAVX}
@@ -517,20 +585,16 @@ Arm64JITCore::Arm64JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::Intern
RAPass = Thread->PassManager->GetPass<IR::RegisterAllocationPass>("RA");
uint32_t NumUsedGPRs = NumGPRs;
uint32_t NumUsedGPRPairs = NumGPRPairs;
uint32_t UsedRegisterCount = RegisterCount;
RAPass->AllocateRegisterSet(RegisterClasses);
RAPass->AllocateRegisterSet(UsedRegisterCount, RegisterClasses);
RAPass->AddRegisters(FEXCore::IR::GPRClass, NumUsedGPRs);
RAPass->AddRegisters(FEXCore::IR::GPRFixedClass, SRA64.size());
RAPass->AddRegisters(FEXCore::IR::FPRClass, NumFPRs);
RAPass->AddRegisters(FEXCore::IR::FPRFixedClass, SRAFPR.size() );
RAPass->AddRegisters(FEXCore::IR::GPRPairClass, NumUsedGPRPairs);
RAPass->AddRegisters(FEXCore::IR::GPRClass, ConfiguredGPRs);
RAPass->AddRegisters(FEXCore::IR::GPRFixedClass, ConfiguredSRAGPRs);
RAPass->AddRegisters(FEXCore::IR::FPRClass, ConfiguredFPRs);
RAPass->AddRegisters(FEXCore::IR::FPRFixedClass, ConfiguredSRAFPRs);
RAPass->AddRegisters(FEXCore::IR::GPRPairClass, ConfiguredGPRPairs);
RAPass->AddRegisters(FEXCore::IR::ComplexClass, 1);
for (uint32_t i = 0; i < NumUsedGPRPairs; ++i) {
for (uint32_t i = 0; i < ConfiguredGPRPairs; ++i) {
RAPass->AddRegisterConflict(FEXCore::IR::GPRClass, i * 2, FEXCore::IR::GPRPairClass, i);
RAPass->AddRegisterConflict(FEXCore::IR::GPRClass, i * 2 + 1, FEXCore::IR::GPRPairClass, i);
}
@@ -543,7 +607,7 @@ Arm64JITCore::Arm64JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::Intern
Common.PrintValue = reinterpret_cast<uint64_t>(PrintValue);
Common.PrintVectorValue = reinterpret_cast<uint64_t>(PrintVectorValue);
Common.ThreadRemoveCodeEntryFromJIT = reinterpret_cast<uintptr_t>(&Context::Context::ThreadRemoveCodeEntryFromJit);
Common.ThreadRemoveCodeEntryFromJIT = reinterpret_cast<uintptr_t>(&Context::ContextImpl::ThreadRemoveCodeEntryFromJit);
Common.CPUIDObj = reinterpret_cast<uint64_t>(&CTX->CPUID);
{
@@ -551,9 +615,14 @@ Arm64JITCore::Arm64JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::Intern
Common.CPUIDFunction = PMF.GetConvertedPointer();
}
{
FEXCore::Utils::MemberFunctionToPointerCast PMF(&FEXCore::CPUIDEmu::RunXCRFunction);
Common.XCRFunction = PMF.GetConvertedPointer();
}
Common.SyscallHandlerObj = reinterpret_cast<uint64_t>(CTX->SyscallHandler);
Common.SyscallHandlerFunc = reinterpret_cast<uint64_t>(FEXCore::Context::HandleSyscall);
Common.ExitFunctionLink = reinterpret_cast<uintptr_t>(&Context::Context::ThreadExitFunctionLink<Arm64JITCore_ExitFunctionLink>);
Common.ExitFunctionLink = reinterpret_cast<uintptr_t>(&Context::ContextImpl::ThreadExitFunctionLink<Arm64JITCore_ExitFunctionLink>);
// Fill in the fallback handlers
@@ -570,23 +639,25 @@ Arm64JITCore::Arm64JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::Intern
// Must be done after Dispatcher init
ClearCache();
}
void Arm64JITCore::InitializeSignalHandlers(FEXCore::Context::Context *CTX) {
CTX->SignalDelegation->RegisterHostSignalHandler(SIGILL, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
return Thread->CTX->Dispatcher->HandleSIGILL(Thread, Signal, info, ucontext);
}, true);
// Setup dynamic dispatch.
if (CTX->Dispatcher->GetConfig().StaticRegisterAllocation) {
RT_LoadRegister = &Arm64JITCore::Op_LoadRegisterSRA;
RT_StoreRegister = &Arm64JITCore::Op_StoreRegisterSRA;
}
else {
RT_LoadRegister = &Arm64JITCore::Op_LoadRegister;
RT_StoreRegister = &Arm64JITCore::Op_StoreRegister;
}
#ifdef _M_ARM_64
CTX->SignalDelegation->RegisterHostSignalHandler(SIGBUS, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
if (!Thread->CPUBackend->IsAddressInCodeBuffer(ArchHelpers::Context::GetPc(ucontext))) {
// Wasn't a sigbus in JIT code
return false;
}
return FEXCore::ArchHelpers::Arm64::HandleSIGBUS(Thread->CTX->Config.ParanoidTSO(), Signal, info, ucontext);
}, true);
#endif
if (ParanoidTSO()) {
RT_LoadMemTSO = &Arm64JITCore::Op_ParanoidLoadMemTSO;
RT_StoreMemTSO = &Arm64JITCore::Op_ParanoidStoreMemTSO;
}
else {
RT_LoadMemTSO = &Arm64JITCore::Op_LoadMemTSO;
RT_StoreMemTSO = &Arm64JITCore::Op_StoreMemTSO;
}
}
void Arm64JITCore::EmitDetectionString() {
@@ -656,7 +727,7 @@ bool Arm64JITCore::IsGPR(IR::NodeID Node) const {
return Class == IR::GPRClass || Class == IR::GPRFixedClass;
}
void *Arm64JITCore::CompileCode(uint64_t Entry,
CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry,
FEXCore::IR::IRListView const *IR,
FEXCore::Core::DebugData *DebugData,
FEXCore::IR::RegisterAllocationData *RAData,
@@ -669,6 +740,21 @@ void *Arm64JITCore::CompileCode(uint64_t Entry,
this->Entry = Entry;
this->RAData = RAData;
this->DebugData = DebugData;
this->IR = IR;
// Fairly excessive buffer range to make sure we don't overflow
uint32_t BufferRange = SSACount * 16 + GDBEnabled * Dispatcher::MaxGDBPauseCheckSize;
if ((GetCursorOffset() + BufferRange) > CurrentCodeBuffer->Size) {
CTX->ClearCodeCache(ThreadState);
}
CodeData.BlockBegin = GetCursorAddress<uint8_t*>();
// Put the code header at the start of the data block.
ARMEmitter::BackwardLabel JITCodeHeaderLabel{};
Bind(&JITCodeHeaderLabel);
JITCodeHeader *CodeHeader = GetCursorAddress<JITCodeHeader *>();
CursorIncrement(sizeof(JITCodeHeader));
#ifdef VIXL_DISASSEMBLER
const auto DisasmBegin = GetCursorAddress<const vixl::aarch64::Instruction*>();
@@ -678,14 +764,6 @@ void *Arm64JITCore::CompileCode(uint64_t Entry,
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, Entry);
#endif
this->IR = IR;
// Fairly excessive buffer range to make sure we don't overflow
uint32_t BufferRange = SSACount * 16 + GDBEnabled * Dispatcher::MaxGDBPauseCheckSize;
if ((GetCursorOffset() + BufferRange) > CurrentCodeBuffer->Size) {
CTX->ClearCodeCache(ThreadState);
}
// AAPCS64
// r30 = LR
// r29 = FP
@@ -706,10 +784,15 @@ void *Arm64JITCore::CompileCode(uint64_t Entry,
// X1-X3 = Temp
// X4-r18 = RA
GuestEntry = GetCursorAddress<uint8_t *>();
CodeData.BlockEntry = GetCursorAddress<uint8_t*>();
// Get the address of the JITCodeHeader and store in to the core state.
// Two instruction cost, each 1 cycle.
adr(TMP1, &JITCodeHeaderLabel);
str(TMP1, STATE, offsetof(FEXCore::Core::CPUState, InlineJITBlockHeader));
if (GDBEnabled) {
auto GDBSize = CTX->Dispatcher->GenerateGDBPauseCheck(GuestEntry, Entry);
auto GDBSize = CTX->Dispatcher->GenerateGDBPauseCheck(CodeData.BlockEntry, Entry);
CursorIncrement(GDBSize);
}
@@ -755,6 +838,7 @@ void *Arm64JITCore::CompileCode(uint64_t Entry,
for (auto [CodeNode, IROp] : IR->GetCode(BlockNode)) {
const auto ID = IR->GetID(CodeNode);
switch (IROp->Op) {
#define REGISTER_OP_RT(op, x) case FEXCore::IR::IROps::OP_##op: std::invoke(RT_##x, this, IROp, ID); break
#define REGISTER_OP(op, x) case FEXCore::IR::IROps::OP_##op: Op_##x(IROp, ID); break
// ALU ops
REGISTER_OP(TRUNCELEMENTPAIR, TruncElementPair);
@@ -822,7 +906,6 @@ void *Arm64JITCore::CompileCode(uint64_t Entry,
REGISTER_OP(ATOMICFETCHNEG, AtomicFetchNeg);
// Branch ops
REGISTER_OP(SIGNALRETURN, SignalReturn);
REGISTER_OP(CALLBACKRETURN, CallbackReturn);
REGISTER_OP(EXITFUNCTION, ExitFunction);
REGISTER_OP(JUMP, Jump);
@@ -833,10 +916,12 @@ void *Arm64JITCore::CompileCode(uint64_t Entry,
REGISTER_OP(VALIDATECODE, ValidateCode);
REGISTER_OP(THREADREMOVECODEENTRY, ThreadRemoveCodeEntry);
REGISTER_OP(CPUID, CPUID);
REGISTER_OP(XGETBV, XGETBV);
// Conversion ops
REGISTER_OP(VINSGPR, VInsGPR);
REGISTER_OP(VCASTFROMGPR, VCastFromGPR);
REGISTER_OP(VDUPFROMGPR, VDupFromGPR);
REGISTER_OP(FLOAT_FROMGPR_S, Float_FromGPR_S);
REGISTER_OP(FLOAT_FTOF, Float_FToF);
REGISTER_OP(VECTOR_STOF, Vector_SToF);
@@ -861,8 +946,8 @@ void *Arm64JITCore::CompileCode(uint64_t Entry,
// Memory ops
REGISTER_OP(LOADCONTEXT, LoadContext);
REGISTER_OP(STORECONTEXT, StoreContext);
REGISTER_OP(LOADREGISTER, LoadRegister);
REGISTER_OP(STOREREGISTER, StoreRegister);
REGISTER_OP_RT(LOADREGISTER, LoadRegister);
REGISTER_OP_RT(STOREREGISTER, StoreRegister);
REGISTER_OP(LOADCONTEXTINDEXED, LoadContextIndexed);
REGISTER_OP(STORECONTEXTINDEXED, StoreContextIndexed);
REGISTER_OP(SPILLREGISTER, SpillRegister);
@@ -871,22 +956,13 @@ void *Arm64JITCore::CompileCode(uint64_t Entry,
REGISTER_OP(STOREFLAG, StoreFlag);
REGISTER_OP(LOADMEM, LoadMem);
REGISTER_OP(STOREMEM, StoreMem);
case FEXCore::IR::IROps::OP_LOADMEMTSO:
if (ParanoidTSO()) {
Op_ParanoidLoadMemTSO(IROp, ID);
}
else {
Op_LoadMemTSO(IROp, ID);
}
break;
case FEXCore::IR::IROps::OP_STOREMEMTSO:
if (ParanoidTSO()) {
Op_ParanoidStoreMemTSO(IROp, ID);
}
else {
Op_StoreMemTSO(IROp, ID);
}
break;
REGISTER_OP_RT(LOADMEMTSO, LoadMemTSO);
REGISTER_OP_RT(STOREMEMTSO, StoreMemTSO);
REGISTER_OP(VLOADVECTORMASKED, VLoadVectorMasked);
REGISTER_OP(VSTOREVECTORMASKED, VStoreVectorMasked);
REGISTER_OP(MEMSET, MemSet);
REGISTER_OP(MEMCPY, MemCpy);
REGISTER_OP(CACHELINECLEAR, CacheLineClear);
REGISTER_OP(CACHELINECLEAN, CacheLineClean);
REGISTER_OP(CACHELINEZERO, CacheLineZero);
@@ -955,6 +1031,8 @@ void *Arm64JITCore::CompileCode(uint64_t Entry,
REGISTER_OP(VZIP2, VZip2);
REGISTER_OP(VUNZIP, VUnZip);
REGISTER_OP(VUNZIP2, VUnZip2);
REGISTER_OP(VTRN, VTrn);
REGISTER_OP(VTRN2, VTrn2);
REGISTER_OP(VBSL, VBSL);
REGISTER_OP(VCMPEQ, VCMPEQ);
REGISTER_OP(VCMPEQZ, VCMPEQZ);
@@ -1009,7 +1087,7 @@ void *Arm64JITCore::CompileCode(uint64_t Entry,
if (DebugData) {
DebugData->Subblocks.push_back({
static_cast<uint32_t>(BlockStartHostCode - GuestEntry),
static_cast<uint32_t>(BlockStartHostCode - CodeData.BlockEntry),
static_cast<uint32_t>(GetCursorAddress<uint8_t *>() - BlockStartHostCode)
});
}
@@ -1022,8 +1100,24 @@ void *Arm64JITCore::CompileCode(uint64_t Entry,
}
PendingTargetLabel = nullptr;
auto CodeEnd = GetCursorAddress<uint8_t *>();
ClearICache(GuestEntry, CodeEnd - GuestEntry);
// Add the JitCodeTail
auto JITBlockTailLocation = GetCursorAddress<uint8_t *>();
auto JITBlockTail = GetCursorAddress<JITCodeTail*>();
CursorIncrement(sizeof(JITCodeTail));
// Put the block's RIP entry in the tail.
// This will be used for RIP reconstruction in the future.
// TODO: This needs to be a data RIP relocation once code caching works.
// Current relocation code doesn't support this feature yet.
JITBlockTail->RIP = Entry;
CodeHeader->OffsetToBlockTail = JITBlockTailLocation - CodeData.BlockBegin;
CodeData.Size = GetCursorAddress<uint8_t *>() - CodeData.BlockBegin;
JITBlockTail->Size = CodeData.Size;
ClearICache(CodeData.BlockBegin, CodeData.Size);
#ifdef VIXL_DISASSEMBLER
const auto DisasmEnd = GetCursorAddress<const vixl::aarch64::Instruction*>();
@@ -1031,13 +1125,13 @@ void *Arm64JITCore::CompileCode(uint64_t Entry,
#endif
if (DebugData) {
DebugData->HostCodeSize = CodeEnd - GuestEntry;
DebugData->HostCodeSize = CodeData.Size;
DebugData->Relocations = &Relocations;
}
this->IR = nullptr;
return GuestEntry;
return CodeData;
}
void Arm64JITCore::ResetStack() {
@@ -1056,12 +1150,8 @@ void Arm64JITCore::ResetStack() {
}
}
std::unique_ptr<CPUBackend> CreateArm64JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread) {
return std::make_unique<Arm64JITCore>(ctx, Thread);
}
void InitializeArm64JITSignalHandlers(FEXCore::Context::Context *CTX) {
Arm64JITCore::InitializeSignalHandlers(CTX);
fextl::unique_ptr<CPUBackend> CreateArm64JITCore(FEXCore::Context::ContextImpl *ctx, FEXCore::Core::InternalThreadState *Thread) {
return fextl::make_unique<Arm64JITCore>(ctx, Thread);
}
CPUBackendFeatures GetArm64JITBackendFeatures() {
+31 -37
View File
@@ -6,7 +6,6 @@ $end_info$
#pragma once
#include <FEXCore/IR/RegisterAllocationData.h>
#include "Interface/Core/ArchHelpers/Arm64Emitter.h"
#include "Interface/Core/ArchHelpers/CodeEmitter/Emitter.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
@@ -14,15 +13,18 @@ $end_info$
#include <aarch64/assembler-aarch64.h>
#include <aarch64/disasm-aarch64.h>
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Core/CPUBackend.h>
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/IR/RegisterAllocationData.h>
#include <FEXCore/fextl/map.h>
#include <FEXCore/fextl/string.h>
#include <FEXCore/fextl/vector.h>
#include <array>
#include <cstdint>
#include <map>
#include <utility>
#include <vector>
namespace FEXCore::Core {
struct InternalThreadState;
@@ -31,13 +33,13 @@ namespace FEXCore::Core {
namespace FEXCore::CPU {
class Arm64JITCore final : public CPUBackend, public Arm64Emitter {
public:
explicit Arm64JITCore(FEXCore::Context::Context *ctx,
explicit Arm64JITCore(FEXCore::Context::ContextImpl *ctx,
FEXCore::Core::InternalThreadState *Thread);
~Arm64JITCore() override;
[[nodiscard]] std::string GetName() override { return "JIT"; }
[[nodiscard]] fextl::string GetName() override { return "JIT"; }
[[nodiscard]] void *CompileCode(uint64_t Entry,
[[nodiscard]] CPUBackend::CompiledCode CompileCode(uint64_t Entry,
FEXCore::IR::IRListView const *IR,
FEXCore::Core::DebugData *DebugData,
FEXCore::IR::RegisterAllocationData *RAData, bool GDBEnabled) override;
@@ -48,8 +50,6 @@ public:
void ClearCache() override;
static void InitializeSignalHandlers(FEXCore::Context::Context *CTX);
void ClearRelocations() override { Relocations.clear(); }
private:
@@ -57,32 +57,12 @@ private:
const bool HostSupportsSVE{};
ARMEmitter::BiDirectionalLabel *PendingTargetLabel;
FEXCore::Context::Context *CTX;
FEXCore::Context::ContextImpl *CTX;
FEXCore::IR::IRListView const *IR;
uint64_t Entry;
CPUBackend::CompiledCode CodeData{};
std::map<IR::NodeID, ARMEmitter::BiDirectionalLabel> JumpTargets;
/**
* @name Register Allocation
* @{ */
constexpr static uint32_t NumGPRs = RA64.size();
constexpr static uint32_t NumFPRs = RAFPR.size();
constexpr static uint32_t NumGPRPairs = RA64Pair.size();
constexpr static uint32_t NumCalleeGPRs = 10;
constexpr static uint32_t NumCalleeGPRPairs = 5;
constexpr static uint32_t RegisterCount = NumGPRs + NumFPRs + NumGPRPairs;
constexpr static uint32_t RegisterClasses = 6;
constexpr static uint64_t GPRBase = (0ULL << 32);
constexpr static uint64_t FPRBase = (1ULL << 32);
constexpr static uint64_t GPRPairBase = (2ULL << 32);
/** @} */
constexpr static uint8_t RA_32 = 0;
constexpr static uint8_t RA_64 = 1;
constexpr static uint8_t RA_FPR = 2;
fextl::map<IR::NodeID, ARMEmitter::BiDirectionalLabel> JumpTargets;
[[nodiscard]] FEXCore::ARMEmitter::Register GetReg(IR::NodeID Node) const {
const auto Reg = GetPhys(Node);
@@ -222,7 +202,7 @@ private:
*/
void PlaceNamedSymbolLiteral(NamedSymbolLiteralPair &Lit);
std::vector<FEXCore::CPU::Relocation> Relocations;
fextl::vector<FEXCore::CPU::Relocation> Relocations;
///< Relocation code loading
bool ApplyRelocations(uint64_t GuestEntry, uint64_t CodeEntry, uint64_t CursorEntry, size_t NumRelocations, const char* EntryRelocations);
@@ -230,10 +210,15 @@ private:
/** @} */
uint32_t SpillSlots{};
/**
* @brief Current guest RIP entrypoint
*/
uint8_t *GuestEntry{};
using OpType = void (Arm64JITCore::*)(IR::IROp_Header const *IROp, IR::NodeID Node);
// Runtime selection;
// Load and store register style.
OpType RT_LoadRegister;
OpType RT_StoreRegister;
// Load and store TSO memory style
OpType RT_LoadMemTSO;
OpType RT_StoreMemTSO;
#define DEF_OP(x) void Op_##x(IR::IROp_Header const *IROp, IR::NodeID Node)
@@ -312,7 +297,6 @@ private:
DEF_OP(AtomicFetchNeg);
///< Branch ops
DEF_OP(SignalReturn);
DEF_OP(CallbackReturn);
DEF_OP(ExitFunction);
DEF_OP(Jump);
@@ -323,10 +307,12 @@ private:
DEF_OP(ValidateCode);
DEF_OP(ThreadRemoveCodeEntry);
DEF_OP(CPUID);
DEF_OP(XGETBV);
///< Conversion ops
DEF_OP(VInsGPR);
DEF_OP(VCastFromGPR);
DEF_OP(VDupFromGPR);
DEF_OP(Float_FromGPR_S);
DEF_OP(Float_FToF);
DEF_OP(Vector_SToF);
@@ -343,6 +329,8 @@ private:
DEF_OP(StoreContext);
DEF_OP(LoadRegister);
DEF_OP(StoreRegister);
DEF_OP(LoadRegisterSRA);
DEF_OP(StoreRegisterSRA);
DEF_OP(LoadContextIndexed);
DEF_OP(StoreContextIndexed);
DEF_OP(SpillRegister);
@@ -353,6 +341,10 @@ private:
DEF_OP(StoreMem);
DEF_OP(LoadMemTSO);
DEF_OP(StoreMemTSO);
DEF_OP(VLoadVectorMasked);
DEF_OP(VStoreVectorMasked);
DEF_OP(MemSet);
DEF_OP(MemCpy);
DEF_OP(ParanoidLoadMemTSO);
DEF_OP(ParanoidStoreMemTSO);
DEF_OP(CacheLineClear);
@@ -417,6 +409,8 @@ private:
DEF_OP(VZip2);
DEF_OP(VUnZip);
DEF_OP(VUnZip2);
DEF_OP(VTrn);
DEF_OP(VTrn2);
DEF_OP(VBSL);
DEF_OP(VCMPEQ);
DEF_OP(VCMPEQZ);
+709 -24
View File
@@ -128,6 +128,170 @@ DEF_OP(LoadRegister) {
const auto Op = IROp->C<IR::IROp_LoadRegister>();
const auto OpSize = IROp->Size;
if (Op->Class == IR::GPRClass) {
[[maybe_unused]] const auto regId = (Op->Offset / Core::CPUState::GPR_REG_SIZE) - 1;
const auto regOffs = Op->Offset & 7;
LOGMAN_THROW_A_FMT(regId < SRA64.size(), "out of range regId");
switch (OpSize) {
case 1:
LOGMAN_THROW_AA_FMT(regOffs == 0 || regOffs == 1, "unexpected regOffs");
ldrb(GetReg(Node), STATE, Op->Offset);
break;
case 2:
LOGMAN_THROW_AA_FMT(regOffs == 0, "unexpected regOffs");
ldrh(GetReg(Node), STATE, Op->Offset);
break;
case 4:
LOGMAN_THROW_AA_FMT(regOffs == 0, "unexpected regOffs");
ldr(GetReg(Node).W(), STATE, Op->Offset);
break;
case 8:
LOGMAN_THROW_AA_FMT(regOffs == 0, "unexpected regOffs");
ldr(GetReg(Node).X(), STATE, Op->Offset);
break;
default:
LOGMAN_MSG_A_FMT("Unhandled LoadRegister GPR size: {}", OpSize);
break;
}
}
else if (Op->Class == IR::FPRClass) {
const auto regSize = HostSupportsSVE ? Core::CPUState::XMM_AVX_REG_SIZE
: Core::CPUState::XMM_SSE_REG_SIZE;
[[maybe_unused]] const auto regId = (Op->Offset - offsetof(Core::CpuStateFrame, State.xmm.avx.data[0][0])) / regSize;
LOGMAN_THROW_A_FMT(HostSupportsSVE, "Unsupported code path!");
LOGMAN_THROW_A_FMT(regId < SRAFPR.size(), "out of range regId");
const auto host = GetVReg(Node);
const auto regOffs = Op->Offset & 15;
switch (OpSize) {
case 1: {
LOGMAN_THROW_AA_FMT(regOffs == 0, "unexpected regOffs: {}", regOffs);
ldrb(host, STATE, Op->Offset);
break;
}
case 2: {
LOGMAN_THROW_AA_FMT(regOffs == 0, "unexpected regOffs: {}", regOffs);
ldrh(host, STATE, Op->Offset);
break;
}
case 4: {
LOGMAN_THROW_AA_FMT((regOffs & 3) == 0, "unexpected regOffs: {}", regOffs);
ldr(host.S(), STATE, Op->Offset);
break;
}
case 8: {
LOGMAN_THROW_AA_FMT((regOffs & 7) == 0, "unexpected regOffs: {}", regOffs);
ldr(host.D(), STATE, Op->Offset);
break;
}
case 16: {
LOGMAN_THROW_AA_FMT(regOffs == 0, "unexpected regOffs: {}", regOffs);
ldr(host.Q(), STATE, Op->Offset);
break;
}
}
} else {
LOGMAN_THROW_AA_FMT(false, "Unhandled Op->Class {}", Op->Class);
}
}
DEF_OP(StoreRegister) {
const auto Op = IROp->C<IR::IROp_StoreRegister>();
const auto OpSize = IROp->Size;
if (Op->Class == IR::GPRClass) {
[[maybe_unused]] const auto regId = (Op->Offset / Core::CPUState::GPR_REG_SIZE) - 1;
const auto regOffs = Op->Offset & 7;
LOGMAN_THROW_A_FMT(regId < SRA64.size(), "out of range regId");
const auto Src = GetReg(Op->Value.ID());
switch (OpSize) {
case 1:
LOGMAN_THROW_AA_FMT(regOffs == 0 || regOffs == 1, "unexpected regOffs");
strb(Src, STATE, Op->Offset);
break;
case 2:
LOGMAN_THROW_AA_FMT(regOffs == 0, "unexpected regOffs");
strh(Src, STATE, Op->Offset);
break;
case 4:
LOGMAN_THROW_AA_FMT(regOffs == 0, "unexpected regOffs");
str(Src.W(), STATE, Op->Offset);
break;
case 8:
LOGMAN_THROW_AA_FMT(regOffs == 0, "unexpected regOffs");
str(Src.X(), STATE, Op->Offset);
break;
default:
LOGMAN_MSG_A_FMT("Unhandled StoreRegister GPR size: {}", OpSize);
break;
}
} else if (Op->Class == IR::FPRClass) {
const auto regSize = HostSupportsSVE ? Core::CPUState::XMM_AVX_REG_SIZE
: Core::CPUState::XMM_SSE_REG_SIZE;
[[maybe_unused]] const auto regId = (Op->Offset - offsetof(Core::CpuStateFrame, State.xmm.avx.data[0][0])) / regSize;
LOGMAN_THROW_A_FMT(HostSupportsSVE, "Unsupported code path!");
LOGMAN_THROW_A_FMT(regId < SRAFPR.size(), "regId out of range");
const auto host = GetVReg(Op->Value.ID());
const auto regOffs = Op->Offset & 15;
switch (OpSize) {
case 1:
strb(host, STATE, Op->Offset);
break;
case 2:
LOGMAN_THROW_AA_FMT((regOffs & 1) == 0, "unexpected regOffs: {}", regOffs);
strh(host, STATE, Op->Offset);
break;
case 4:
LOGMAN_THROW_AA_FMT((regOffs & 3) == 0, "unexpected regOffs: {}", regOffs);
str(host.S(), STATE, Op->Offset);
break;
case 8:
LOGMAN_THROW_AA_FMT((regOffs & 7) == 0, "unexpected regOffs: {}", regOffs);
str(host.D(), STATE, Op->Offset);
break;
case 16:
LOGMAN_THROW_AA_FMT(regOffs == 0, "unexpected regOffs: {}", regOffs);
str(host.Q(), STATE, Op->Offset);
break;
default:
LOGMAN_MSG_A_FMT("Unhandled StoreRegister FPR size: {}", OpSize);
break;
}
} else {
LOGMAN_THROW_AA_FMT(false, "Unhandled Op->Class {}", Op->Class);
}
}
DEF_OP(LoadRegisterSRA) {
const auto Op = IROp->C<IR::IROp_LoadRegister>();
const auto OpSize = IROp->Size;
if (Op->Class == IR::GPRClass) {
const auto regId = (Op->Offset - offsetof(Core::CpuStateFrame, State.gregs[0])) / Core::CPUState::GPR_REG_SIZE;
const auto regOffs = Op->Offset & 7;
@@ -312,7 +476,7 @@ DEF_OP(LoadRegister) {
}
}
DEF_OP(StoreRegister) {
DEF_OP(StoreRegisterSRA) {
const auto Op = IROp->C<IR::IROp_StoreRegister>();
const auto OpSize = IROp->Size;
@@ -500,7 +664,6 @@ DEF_OP(StoreRegister) {
}
}
DEF_OP(LoadContextIndexed) {
const auto Op = IROp->C<IR::IROp_LoadContextIndexed>();
const auto OpSize = IROp->Size;
@@ -703,7 +866,7 @@ DEF_OP(SpillRegister) {
const auto Src = GetReg(Op->Value.ID());
switch (OpSize) {
case 1: {
if (SlotOffset > 4095) {
if (SlotOffset > LSByteMaxUnsignedOffset) {
LoadConstant(ARMEmitter::Size::i64Bit, TMP1, SlotOffset);
strb(Src, ARMEmitter::Reg::rsp, TMP1.R(), ARMEmitter::ExtendedType::LSL_64, 0);
}
@@ -713,7 +876,7 @@ DEF_OP(SpillRegister) {
break;
}
case 2: {
if (SlotOffset > 8190) {
if (SlotOffset > LSHalfMaxUnsignedOffset) {
LoadConstant(ARMEmitter::Size::i64Bit, TMP1, SlotOffset);
strh(Src, ARMEmitter::Reg::rsp, TMP1.R(), ARMEmitter::ExtendedType::LSL_64, 0);
}
@@ -723,7 +886,7 @@ DEF_OP(SpillRegister) {
break;
}
case 4: {
if (SlotOffset > 16380) {
if (SlotOffset > LSWordMaxUnsignedOffset) {
LoadConstant(ARMEmitter::Size::i64Bit, TMP1, SlotOffset);
str(Src.W(), ARMEmitter::Reg::rsp, TMP1.R(), ARMEmitter::ExtendedType::LSL_64, 0);
}
@@ -733,7 +896,7 @@ DEF_OP(SpillRegister) {
break;
}
case 8: {
if (SlotOffset > 32760) {
if (SlotOffset > LSDWordMaxUnsignedOffset) {
LoadConstant(ARMEmitter::Size::i64Bit, TMP1, SlotOffset);
str(Src.X(), ARMEmitter::Reg::rsp, TMP1.R(), ARMEmitter::ExtendedType::LSL_64, 0);
}
@@ -751,7 +914,7 @@ DEF_OP(SpillRegister) {
switch (OpSize) {
case 4: {
if (SlotOffset > 16380) {
if (SlotOffset > LSWordMaxUnsignedOffset) {
LoadConstant(ARMEmitter::Size::i64Bit, TMP1, SlotOffset);
str(Src.S(), ARMEmitter::Reg::rsp, TMP1.R(), ARMEmitter::ExtendedType::LSL_64, 0);
}
@@ -761,7 +924,7 @@ DEF_OP(SpillRegister) {
break;
}
case 8: {
if (SlotOffset > 32760) {
if (SlotOffset > LSDWordMaxUnsignedOffset) {
LoadConstant(ARMEmitter::Size::i64Bit, TMP1, SlotOffset);
str(Src.D(), ARMEmitter::Reg::rsp, TMP1.R(), ARMEmitter::ExtendedType::LSL_64, 0);
}
@@ -771,7 +934,7 @@ DEF_OP(SpillRegister) {
break;
}
case 16: {
if (SlotOffset > 65520) {
if (SlotOffset > LSQWordMaxUnsignedOffset) {
LoadConstant(ARMEmitter::Size::i64Bit, TMP1, SlotOffset);
str(Src.Q(), ARMEmitter::Reg::rsp, TMP1.R(), ARMEmitter::ExtendedType::LSL_64, 0);
}
@@ -803,7 +966,7 @@ DEF_OP(FillRegister) {
const auto Dst = GetReg(Node);
switch (OpSize) {
case 1: {
if (SlotOffset > 4095) {
if (SlotOffset > LSByteMaxUnsignedOffset) {
LoadConstant(ARMEmitter::Size::i64Bit, TMP1, SlotOffset);
ldrb(Dst, ARMEmitter::Reg::rsp, TMP1.R(), ARMEmitter::ExtendedType::LSL_64, 0);
}
@@ -813,7 +976,7 @@ DEF_OP(FillRegister) {
break;
}
case 2: {
if (SlotOffset > 8190) {
if (SlotOffset > LSHalfMaxUnsignedOffset) {
LoadConstant(ARMEmitter::Size::i64Bit, TMP1, SlotOffset);
ldrh(Dst, ARMEmitter::Reg::rsp, TMP1.R(), ARMEmitter::ExtendedType::LSL_64, 0);
}
@@ -823,7 +986,7 @@ DEF_OP(FillRegister) {
break;
}
case 4: {
if (SlotOffset > 16380) {
if (SlotOffset > LSWordMaxUnsignedOffset) {
LoadConstant(ARMEmitter::Size::i64Bit, TMP1, SlotOffset);
ldr(Dst.W(), ARMEmitter::Reg::rsp, TMP1.R(), ARMEmitter::ExtendedType::LSL_64, 0);
}
@@ -833,7 +996,7 @@ DEF_OP(FillRegister) {
break;
}
case 8: {
if (SlotOffset > 32760) {
if (SlotOffset > LSDWordMaxUnsignedOffset) {
LoadConstant(ARMEmitter::Size::i64Bit, TMP1, SlotOffset);
ldr(Dst.X(), ARMEmitter::Reg::rsp, TMP1.R(), ARMEmitter::ExtendedType::LSL_64, 0);
}
@@ -851,7 +1014,7 @@ DEF_OP(FillRegister) {
switch (OpSize) {
case 4: {
if (SlotOffset > 16380) {
if (SlotOffset > LSWordMaxUnsignedOffset) {
LoadConstant(ARMEmitter::Size::i64Bit, TMP1, SlotOffset);
ldr(Dst.S(), ARMEmitter::Reg::rsp, TMP1.R(), ARMEmitter::ExtendedType::LSL_64, 0);
}
@@ -861,7 +1024,7 @@ DEF_OP(FillRegister) {
break;
}
case 8: {
if (SlotOffset > 32760) {
if (SlotOffset > LSDWordMaxUnsignedOffset) {
LoadConstant(ARMEmitter::Size::i64Bit, TMP1, SlotOffset);
ldr(Dst.D(), ARMEmitter::Reg::rsp, TMP1.R(), ARMEmitter::ExtendedType::LSL_64, 0);
}
@@ -871,7 +1034,7 @@ DEF_OP(FillRegister) {
break;
}
case 16: {
if (SlotOffset > 65520) {
if (SlotOffset > LSQWordMaxUnsignedOffset) {
LoadConstant(ARMEmitter::Size::i64Bit, TMP1, SlotOffset);
ldr(Dst.Q(), ARMEmitter::Reg::rsp, TMP1.R(), ARMEmitter::ExtendedType::LSL_64, 0);
}
@@ -911,20 +1074,20 @@ FEXCore::ARMEmitter::ExtendedMemOperand Arm64JITCore::GenerateMemOperand(uint8_t
IR::MemOffsetType OffsetType,
uint8_t OffsetScale) {
if (Offset.IsInvalid()) {
return FEXCore::ARMEmitter::ExtendedMemOperand(Base, ARMEmitter::IndexType::OFFSET, 0);
return ARMEmitter::ExtendedMemOperand(Base.X(), ARMEmitter::IndexType::OFFSET, 0);
} else {
if (OffsetScale != 1 && OffsetScale != AccessSize) {
LOGMAN_MSG_A_FMT("Unhandled GenerateMemOperand OffsetScale: {}", OffsetScale);
}
uint64_t Const;
if (IsInlineConstant(Offset, &Const)) {
return FEXCore::ARMEmitter::ExtendedMemOperand(Base, ARMEmitter::IndexType::OFFSET, Const);
return ARMEmitter::ExtendedMemOperand(Base.X(), ARMEmitter::IndexType::OFFSET, Const);
} else {
auto RegOffset = GetReg(Offset.ID());
switch(OffsetType.Val) {
case IR::MEM_OFFSET_SXTX.Val: return FEXCore::ARMEmitter::ExtendedMemOperand(Base, RegOffset, FEXCore::ARMEmitter::ExtendedType::SXTX, (int)std::log2(OffsetScale) );
case IR::MEM_OFFSET_UXTW.Val: return FEXCore::ARMEmitter::ExtendedMemOperand(Base, RegOffset, FEXCore::ARMEmitter::ExtendedType::UXTW, (int)std::log2(OffsetScale) );
case IR::MEM_OFFSET_SXTW.Val: return FEXCore::ARMEmitter::ExtendedMemOperand(Base, RegOffset, FEXCore::ARMEmitter::ExtendedType::SXTW, (int)std::log2(OffsetScale) );
case IR::MEM_OFFSET_SXTX.Val: return ARMEmitter::ExtendedMemOperand(Base.X(), RegOffset.X(), ARMEmitter::ExtendedType::SXTX, (int)std::log2(OffsetScale) );
case IR::MEM_OFFSET_UXTW.Val: return ARMEmitter::ExtendedMemOperand(Base.X(), RegOffset.X(), ARMEmitter::ExtendedType::UXTW, (int)std::log2(OffsetScale) );
case IR::MEM_OFFSET_SXTW.Val: return ARMEmitter::ExtendedMemOperand(Base.X(), RegOffset.X(), ARMEmitter::ExtendedType::SXTW, (int)std::log2(OffsetScale) );
default: LOGMAN_MSG_A_FMT("Unhandled GenerateMemOperand OffsetType: {}", OffsetType.Val); break;
}
}
@@ -1101,14 +1264,14 @@ DEF_OP(LoadMemTSO) {
const auto Dst = GetReg(Node);
if (OpSize == 1) {
// 8bit load is always aligned to natural alignment
ldaprb(Dst, MemReg);
ldaprb(Dst.W(), MemReg);
}
else {
// Aligned
nop();
switch (OpSize) {
case 2:
ldaprh(Dst, MemReg);
ldaprh(Dst.W(), MemReg);
break;
case 4:
ldapr(Dst.W(), MemReg);
@@ -1182,6 +1345,106 @@ DEF_OP(LoadMemTSO) {
}
}
DEF_OP(VLoadVectorMasked) {
LOGMAN_THROW_A_FMT(HostSupportsSVE, "Need SVE support in order to use VLoadVectorMasked");
const auto Op = IROp->C<IR::IROp_VLoadVectorMasked>();
const auto OpSize = IROp->Size;
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto ElementSize = IROp->ElementSize;
const auto CMPPredicate = ARMEmitter::PReg::p0;
const auto GoverningPredicate = Is256Bit ? PRED_TMP_32B : PRED_TMP_16B;
const auto Dst = GetVReg(Node);
const auto MaskReg = GetVReg(Op->Mask.ID());
const auto MemReg = GetReg(Op->Addr.ID());
const auto MemSrc = GenerateSVEMemOperand(OpSize, MemReg, Op->Offset, Op->OffsetType, Op->OffsetScale);
LOGMAN_THROW_AA_FMT(ElementSize == 1 || ElementSize == 2 || ElementSize == 4 || ElementSize == 8, "Invalid size");
const auto SubRegSize =
ElementSize == 1 ? ARMEmitter::SubRegSize::i8Bit :
ElementSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
ElementSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
ElementSize == 8 ? ARMEmitter::SubRegSize::i64Bit : ARMEmitter::SubRegSize::i8Bit;
// Check if the sign bit is set for the given element size.
cmplt(SubRegSize, CMPPredicate, GoverningPredicate.Zeroing(), MaskReg.Z(), 0);
switch (ElementSize) {
case 1: {
ld1b<ARMEmitter::SubRegSize::i8Bit>(Dst.Z(), CMPPredicate.Zeroing(), MemSrc);
break;
}
case 2: {
ld1h<ARMEmitter::SubRegSize::i16Bit>(Dst.Z(), CMPPredicate.Zeroing(), MemSrc);
break;
}
case 4: {
ld1w<ARMEmitter::SubRegSize::i32Bit>(Dst.Z(), CMPPredicate.Zeroing(), MemSrc);
break;
}
case 8: {
ld1d(Dst.Z(), CMPPredicate.Zeroing(), MemSrc);
break;
}
default:
LOGMAN_MSG_A_FMT("Unhandled VLoadVectorMasked size: {}", ElementSize);
break;
}
}
DEF_OP(VStoreVectorMasked) {
LOGMAN_THROW_A_FMT(HostSupportsSVE, "Need SVE support in order to use VStoreVectorMasked");
const auto Op = IROp->C<IR::IROp_VStoreVectorMasked>();
const auto OpSize = IROp->Size;
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto ElementSize = IROp->ElementSize;
const auto CMPPredicate = ARMEmitter::PReg::p0;
const auto GoverningPredicate = Is256Bit ? PRED_TMP_32B : PRED_TMP_16B;
const auto RegData = GetVReg(Op->Data.ID());
const auto MaskReg = GetVReg(Op->Mask.ID());
const auto MemReg = GetReg(Op->Addr.ID());
const auto MemDst = GenerateSVEMemOperand(OpSize, MemReg, Op->Offset, Op->OffsetType, Op->OffsetScale);
LOGMAN_THROW_AA_FMT(ElementSize == 1 || ElementSize == 2 || ElementSize == 4 || ElementSize == 8, "Invalid size");
const auto SubRegSize =
ElementSize == 1 ? ARMEmitter::SubRegSize::i8Bit :
ElementSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
ElementSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
ElementSize == 8 ? ARMEmitter::SubRegSize::i64Bit : ARMEmitter::SubRegSize::i8Bit;
// Check if the sign bit is set for the given element size.
cmplt(SubRegSize, CMPPredicate, GoverningPredicate.Zeroing(), MaskReg.Z(), 0);
switch (ElementSize) {
case 1: {
st1b<ARMEmitter::SubRegSize::i8Bit>(RegData.Z(), CMPPredicate.Zeroing(), MemDst);
break;
}
case 2: {
st1h<ARMEmitter::SubRegSize::i16Bit>(RegData.Z(), CMPPredicate.Zeroing(), MemDst);
break;
}
case 4: {
st1w<ARMEmitter::SubRegSize::i32Bit>(RegData.Z(), CMPPredicate.Zeroing(), MemDst);
break;
}
case 8: {
st1d(RegData.Z(), CMPPredicate.Zeroing(), MemDst);
break;
}
default:
LOGMAN_MSG_A_FMT("Unhandled VStoreVectorMasked size: {}", ElementSize);
break;
}
}
DEF_OP(StoreMem) {
const auto Op = IROp->C<IR::IROp_StoreMem>();
const auto OpSize = IROp->Size;
@@ -1340,6 +1603,428 @@ DEF_OP(StoreMemTSO) {
}
}
DEF_OP(MemSet) {
// TODO: A future looking task would be to support this with ARM's MOPS instructions.
// The 8-bit non-atomic forward path directly matches ARM's SETP/SETM/SETE instruction,
// while the backward version needs some fixup to convert it to a forward direction.
//
// Assuming non-atomicity and non-faulting behaviour, this can accelerate this implementation.
// Additionally: This is commonly used as a memset to zero. If we know up-front with an inline constant
// that the value is zero, we can optimize any operation larger than 8-bit down to 8-bit to use the MOPS implementation.
const auto Op = IROp->C<IR::IROp_MemSet>();
const int32_t Size = Op->Size;
const auto MemReg = GetReg(Op->Addr.ID());
const auto Value = GetReg(Op->Value.ID());
const auto Length = GetReg(Op->Length.ID());
const auto Direction = GetReg(Op->Direction.ID());
const auto Dst = GetReg(Node);
// If Direction == 0 then:
// MemReg is incremented (by size)
// else:
// MemReg is decremented (by size)
//
// Counter is decremented regardless.
ARMEmitter::ForwardLabel BackwardImpl{};
ARMEmitter::ForwardLabel Done{};
mov(TMP1, Length.X());
if (Op->Prefix.IsInvalid()) {
mov(TMP2, MemReg.X());
}
else {
const auto Prefix = GetReg(Op->Prefix.ID());
add(TMP2, Prefix.X(), MemReg.X());
}
// Backward or forwards implementation depends on flag
cbnz(ARMEmitter::Size::i64Bit, Direction, &BackwardImpl);
auto MemStore = [this](auto Value, uint32_t OpSize, int32_t Size) {
switch (OpSize) {
case 1:
strb<ARMEmitter::IndexType::POST>(Value.W(), TMP2, Size);
break;
case 2:
strh<ARMEmitter::IndexType::POST>(Value.W(), TMP2, Size);
break;
case 4:
str<ARMEmitter::IndexType::POST>(Value.W(), TMP2, Size);
break;
case 8:
str<ARMEmitter::IndexType::POST>(Value.X(), TMP2, Size);
break;
default:
LOGMAN_MSG_A_FMT("Unhandled {} size: {}", __func__, Size);
break;
}
};
auto MemStoreTSO = [this](auto Value, uint32_t OpSize, int32_t Size) {
if (OpSize == 1) {
// 8bit load is always aligned to natural alignment
stlrb(Value.W(), TMP2);
}
else {
nop();
switch (OpSize) {
case 2:
stlrh(Value.W(), TMP2);
break;
case 4:
stlr(Value.W(), TMP2);
break;
case 8:
stlr(Value.X(), TMP2);
break;
default:
LOGMAN_MSG_A_FMT("Unhandled {} size: {}", __func__, Size);
break;
}
nop();
}
if (Size >= 0) {
add(ARMEmitter::Size::i64Bit, TMP2, TMP2, OpSize);
}
else {
sub(ARMEmitter::Size::i64Bit, TMP2, TMP2, OpSize);
}
};
// Emit forward direction memset then backward direction memset.
for (int32_t Direction : { 1, -1 }) {
const int32_t OpSize = Size;
const int32_t SizeDirection = Size * Direction;
ARMEmitter::BackwardLabel AgainInternal{};
ARMEmitter::ForwardLabel DoneInternal{};
// Early exit if zero count.
cbz(ARMEmitter::Size::i64Bit, TMP1, &DoneInternal);
Bind(&AgainInternal);
if (Op->IsAtomic) {
MemStoreTSO(Value, OpSize, SizeDirection);
}
else {
MemStore(Value, OpSize, SizeDirection);
}
sub(ARMEmitter::Size::i64Bit, TMP1, TMP1, 1);
cbnz(ARMEmitter::Size::i64Bit, TMP1, &AgainInternal);
Bind(&DoneInternal);
if (SizeDirection >= 0) {
switch (OpSize) {
case 1:
add(Dst.X(), MemReg.X(), Length.X());
break;
case 2:
add(Dst.X(), MemReg.X(), Length.X(), ARMEmitter::ShiftType::LSL, 1);
break;
case 4:
add(Dst.X(), MemReg.X(), Length.X(), ARMEmitter::ShiftType::LSL, 2);
break;
case 8:
add(Dst.X(), MemReg.X(), Length.X(), ARMEmitter::ShiftType::LSL, 3);
break;
default:
LOGMAN_MSG_A_FMT("Unhandled {} size: {}", __func__, OpSize);
break;
}
}
else {
switch (OpSize) {
case 1:
sub(Dst.X(), MemReg.X(), Length.X());
break;
case 2:
sub(Dst.X(), MemReg.X(), Length.X(), ARMEmitter::ShiftType::LSL, 1);
break;
case 4:
sub(Dst.X(), MemReg.X(), Length.X(), ARMEmitter::ShiftType::LSL, 2);
break;
case 8:
sub(Dst.X(), MemReg.X(), Length.X(), ARMEmitter::ShiftType::LSL, 3);
break;
default:
LOGMAN_MSG_A_FMT("Unhandled {} size: {}", __func__, OpSize);
break;
}
}
if (Direction == 1) {
b(&Done);
Bind(&BackwardImpl);
}
}
Bind(&Done);
// Destination already set to the final pointer.
}
DEF_OP(MemCpy) {
// TODO: A future looking task would be to support this with ARM's MOPS instructions.
// The 8-bit non-atomic path directly matches ARM's CPYP/CPYM/CPYE instruction,
//
// Assuming non-atomicity and non-faulting behaviour, this can accelerate this implementation.
const auto Op = IROp->C<IR::IROp_MemCpy>();
const int32_t Size = Op->Size;
const auto MemRegDest = GetReg(Op->AddrDest.ID());
const auto MemRegSrc = GetReg(Op->AddrSrc.ID());
const auto Length = GetReg(Op->Length.ID());
const auto Direction = GetReg(Op->Direction.ID());
auto Dst = GetRegPair(Node);
// If Direction == 0 then:
// MemRegDest is incremented (by size)
// MemRegSrc is incremented (by size)
// else:
// MemRegDest is decremented (by size)
// MemRegSrc is decremented (by size)
//
// Counter is decremented regardless.
ARMEmitter::ForwardLabel BackwardImpl{};
ARMEmitter::ForwardLabel Done{};
mov(TMP1, Length.X());
if (Op->PrefixDest.IsInvalid()) {
mov(TMP2, MemRegDest.X());
}
else {
const auto Prefix = GetReg(Op->PrefixDest.ID());
add(TMP2, Prefix.X(), MemRegDest.X());
}
if (Op->PrefixSrc.IsInvalid()) {
mov(TMP3, MemRegSrc.X());
}
else {
const auto Prefix = GetReg(Op->PrefixSrc.ID());
add(TMP3, Prefix.X(), MemRegSrc.X());
}
// TMP1 = Length
// TMP2 = Dest
// TMP3 = Src
// TMP4 = load+store temp value
// Backward or forwards implementation depends on flag
cbnz(ARMEmitter::Size::i64Bit, Direction, &BackwardImpl);
auto MemCpy = [this](uint32_t OpSize, int32_t Size) {
switch (OpSize) {
case 1:
ldrb<ARMEmitter::IndexType::POST>(TMP4.W(), TMP3, Size);
strb<ARMEmitter::IndexType::POST>(TMP4.W(), TMP2, Size);
break;
case 2:
ldrh<ARMEmitter::IndexType::POST>(TMP4.W(), TMP3, Size);
strh<ARMEmitter::IndexType::POST>(TMP4.W(), TMP2, Size);
break;
case 4:
ldr<ARMEmitter::IndexType::POST>(TMP4.W(), TMP3, Size);
str<ARMEmitter::IndexType::POST>(TMP4.W(), TMP2, Size);
break;
case 8:
ldr<ARMEmitter::IndexType::POST>(TMP4, TMP3, Size);
str<ARMEmitter::IndexType::POST>(TMP4, TMP2, Size);
break;
default:
LOGMAN_MSG_A_FMT("Unhandled {} size: {}", __func__, Size);
break;
}
};
auto MemCpyTSO = [this](uint32_t OpSize, int32_t Size) {
if (CTX->HostFeatures.SupportsRCPC) {
if (OpSize == 1) {
// 8bit load is always aligned to natural alignment
ldaprb(TMP4.W(), TMP3);
stlrb(TMP4.W(), TMP2);
}
else {
nop();
switch (OpSize) {
case 2:
ldaprh(TMP4.W(), TMP3);
break;
case 4:
ldapr(TMP4.W(), TMP3);
break;
case 8:
ldapr(TMP4, TMP3);
break;
default:
LOGMAN_MSG_A_FMT("Unhandled {} size: {}", __func__, Size);
break;
}
nop();
nop();
switch (OpSize) {
case 2:
stlrh(TMP4.W(), TMP2);
break;
case 4:
stlr(TMP4.W(), TMP2);
break;
case 8:
stlr(TMP4, TMP2);
break;
default:
LOGMAN_MSG_A_FMT("Unhandled {} size: {}", __func__, Size);
break;
}
nop();
}
}
else {
if (OpSize == 1) {
// 8bit load is always aligned to natural alignment
ldarb(TMP4.W(), TMP3);
stlrb(TMP4.W(), TMP2);
}
else {
nop();
switch (OpSize) {
case 2:
ldarh(TMP4.W(), TMP3);
break;
case 4:
ldar(TMP4.W(), TMP3);
break;
case 8:
ldar(TMP4, TMP3);
break;
default:
LOGMAN_MSG_A_FMT("Unhandled {} size: {}", __func__, Size);
break;
}
nop();
nop();
switch (OpSize) {
case 2:
stlrh(TMP4.W(), TMP2);
break;
case 4:
stlr(TMP4.W(), TMP2);
break;
case 8:
stlr(TMP4, TMP2);
break;
default:
LOGMAN_MSG_A_FMT("Unhandled {} size: {}", __func__, Size);
break;
}
nop();
}
}
if (Size >= 0) {
add(ARMEmitter::Size::i64Bit, TMP2, TMP2, OpSize);
add(ARMEmitter::Size::i64Bit, TMP3, TMP3, OpSize);
}
else {
sub(ARMEmitter::Size::i64Bit, TMP2, TMP2, OpSize);
sub(ARMEmitter::Size::i64Bit, TMP3, TMP3, OpSize);
}
};
// Emit forward direction memset then backward direction memset.
for (int32_t Direction : { 1, -1 }) {
const int32_t OpSize = Size;
const int32_t SizeDirection = Size * Direction;
ARMEmitter::BackwardLabel AgainInternal{};
ARMEmitter::ForwardLabel DoneInternal{};
// Early exit if zero count.
cbz(ARMEmitter::Size::i64Bit, TMP1, &DoneInternal);
Bind(&AgainInternal);
if (Op->IsAtomic) {
MemCpyTSO(OpSize, SizeDirection);
}
else {
MemCpy(OpSize, SizeDirection);
}
sub(ARMEmitter::Size::i64Bit, TMP1, TMP1, 1);
cbnz(ARMEmitter::Size::i64Bit, TMP1, &AgainInternal);
Bind(&DoneInternal);
// Needs to use temporaries just in case of overwrite
mov(TMP1, MemRegDest.X());
mov(TMP2, MemRegSrc.X());
mov(TMP3, Length.X());
if (SizeDirection >= 0) {
switch (OpSize) {
case 1:
add(Dst.first.X(), TMP1, TMP3);
add(Dst.second.X(), TMP2, TMP3);
break;
case 2:
add(Dst.first.X(), TMP1, TMP3, ARMEmitter::ShiftType::LSL, 1);
add(Dst.second.X(), TMP2, TMP3, ARMEmitter::ShiftType::LSL, 1);
break;
case 4:
add(Dst.first.X(), TMP1, TMP3, ARMEmitter::ShiftType::LSL, 2);
add(Dst.second.X(), TMP2, TMP3, ARMEmitter::ShiftType::LSL, 2);
break;
case 8:
add(Dst.first.X(), TMP1, TMP3, ARMEmitter::ShiftType::LSL, 3);
add(Dst.second.X(), TMP2, TMP3, ARMEmitter::ShiftType::LSL, 3);
break;
default:
LOGMAN_MSG_A_FMT("Unhandled {} size: {}", __func__, OpSize);
break;
}
}
else {
switch (OpSize) {
case 1:
sub(Dst.first.X(), TMP1, TMP3);
sub(Dst.second.X(), TMP2, TMP3);
break;
case 2:
sub(Dst.first.X(), TMP1, TMP3, ARMEmitter::ShiftType::LSL, 1);
sub(Dst.second.X(), TMP2, TMP3, ARMEmitter::ShiftType::LSL, 1);
break;
case 4:
sub(Dst.first.X(), TMP1, TMP3, ARMEmitter::ShiftType::LSL, 2);
sub(Dst.second.X(), TMP2, TMP3, ARMEmitter::ShiftType::LSL, 2);
break;
case 8:
sub(Dst.first.X(), TMP1, TMP3, ARMEmitter::ShiftType::LSL, 3);
sub(Dst.second.X(), TMP2, TMP3, ARMEmitter::ShiftType::LSL, 3);
break;
default:
LOGMAN_MSG_A_FMT("Unhandled {} size: {}", __func__, OpSize);
break;
}
}
if (Direction == 1) {
b(&Done);
Bind(&BackwardImpl);
}
}
Bind(&Done);
// Destination already set to the final pointer.
}
DEF_OP(ParanoidLoadMemTSO) {
const auto Op = IROp->C<IR::IROp_LoadMemTSO>();
const auto OpSize = IROp->Size;
@@ -1473,7 +2158,7 @@ DEF_OP(ParanoidStoreMemTSO) {
}
case 32: {
dmb(FEXCore::ARMEmitter::BarrierScope::ISH);
st1b<ARMEmitter::SubRegSize::i8Bit>(Src, PRED_TMP_32B, Addr, 0);
st1b<ARMEmitter::SubRegSize::i8Bit>(Src.Z(), PRED_TMP_32B, Addr, 0);
dmb(FEXCore::ARMEmitter::BarrierScope::ISH);
break;
}
@@ -4,18 +4,23 @@ tags: backend|arm64
$end_info$
*/
#ifndef _WIN32
#include <syscall.h>
#endif
#include "Interface/Core/ArchHelpers/CodeEmitter/Emitter.h"
#include "Interface/Core/JIT/Arm64/JITClass.h"
#include "FEXCore/Debug/InternalThreadState.h"
#include <FEXCore/Core/SignalDelegator.h>
namespace FEXCore::CPU {
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const *IROp, IR::NodeID Node)
DEF_OP(GuestOpcode) {
auto Op = IROp->C<IR::IROp_GuestOpcode>();
// metadata
DebugData->GuestOpcodes.push_back({Op->GuestEntryOffset, GetCursorAddress<uint8_t*>() - GuestEntry});
DebugData->GuestOpcodes.push_back({Op->GuestEntryOffset, GetCursorAddress<uint8_t*>() - CodeData.BlockBegin});
}
DEF_OP(Fence) {
@@ -55,15 +60,15 @@ DEF_OP(Break) {
str(ARMEmitter::XReg::x1, STATE, offsetof(FEXCore::Core::CpuStateFrame, SynchronousFaultData));
switch (Op->Reason.Signal) {
case SIGILL:
case Core::FAULT_SIGILL:
ldr(TMP1, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.GuestSignal_SIGILL));
br(TMP1);
break;
case SIGTRAP:
case Core::FAULT_SIGTRAP:
ldr(TMP1, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.GuestSignal_SIGTRAP));
br(TMP1);
break;
case SIGSEGV:
case Core::FAULT_SIGSEGV:
ldr(TMP1, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.GuestSignal_SIGSEGV));
br(TMP1);
break;
@@ -139,7 +144,7 @@ DEF_OP(Print) {
auto Op = IROp->C<IR::IROp_Print>();
PushDynamicRegsAndLR(TMP1);
SpillStaticRegs();
SpillStaticRegs(TMP1);
if (IsGPR(Op->Value.ID())) {
mov(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, GetReg(Op->Value.ID()));
@@ -157,6 +162,7 @@ DEF_OP(Print) {
PopDynamicRegsAndLR();
}
#ifndef _WIN32
DEF_OP(ProcessorID) {
// We always need to spill x8 since we can't know if it is live at this SSA location
uint32_t SpillMask = 1U << 8;
@@ -164,7 +170,7 @@ DEF_OP(ProcessorID) {
// Ordering is incredibly important here
// We must spill any overlapping registers first THEN claim we are in a syscall without invalidating state at all
// Only spill the registers that intersect with our usage
SpillStaticRegs(false, SpillMask);
SpillStaticRegs(TMP1, false, SpillMask);
// Now that we are spilled, store in the state that we are in a syscall
// Still without overwriting registers that matter
@@ -207,6 +213,11 @@ DEF_OP(ProcessorID) {
// Node is in w1
orr(ARMEmitter::Size::i64Bit, GetReg(Node), ARMEmitter::Reg::r0, ARMEmitter::Reg::r1, ARMEmitter::ShiftType::LSL, 12);
}
#else
DEF_OP(ProcessorID) {
ERROR_AND_DIE_FMT("Unsupported");
}
#endif
DEF_OP(RDRAND) {
auto Op = IROp->C<IR::IROp_RDRAND>();
File diff suppressed because it is too large. Load diff
+5 -6
View File
@@ -1,9 +1,10 @@
#pragma once
#include <memory>
#include <FEXCore/Core/CPUBackend.h>
#include <FEXCore/fextl/memory.h>
namespace FEXCore::Context {
struct Context;
class ContextImpl;
}
namespace FEXCore::Core {
@@ -13,14 +14,12 @@ struct InternalThreadState;
namespace FEXCore::CPU {
class CPUBackend;
[[nodiscard]] std::unique_ptr<CPUBackend> CreateX86JITCore(FEXCore::Context::Context *ctx,
[[nodiscard]] fextl::unique_ptr<CPUBackend> CreateX86JITCore(FEXCore::Context::ContextImpl *ctx,
FEXCore::Core::InternalThreadState *Thread);
void InitializeX86JITSignalHandlers(FEXCore::Context::Context *CTX);
CPUBackendFeatures GetX86JITBackendFeatures();
[[nodiscard]] std::unique_ptr<CPUBackend> CreateArm64JITCore(FEXCore::Context::Context *ctx,
[[nodiscard]] fextl::unique_ptr<CPUBackend> CreateArm64JITCore(FEXCore::Context::ContextImpl *ctx,
FEXCore::Core::InternalThreadState *Thread);
void InitializeArm64JITSignalHandlers(FEXCore::Context::Context *CTX);
CPUBackendFeatures GetArm64JITBackendFeatures();
} // namespace FEXCore::CPU
@@ -5,6 +5,7 @@ $end_info$
*/
#include "Interface/Core/JIT/x86_64/JITClass.h"
#include "Interface/Core/Dispatcher/X86Dispatcher.h"
#include <FEXCore/IR/IR.h>
#include <FEXCore/Utils/LogManager.h>
@@ -12,7 +13,6 @@ $end_info$
#include <array>
#include <stdint.h>
#include <utility>
#include <xbyak/xbyak.h>
namespace FEXCore::CPU {
@@ -5,6 +5,7 @@ $end_info$
*/
#include "Interface/Core/JIT/x86_64/JITClass.h"
#include "Interface/Core/Dispatcher/X86Dispatcher.h"
#include <FEXCore/IR/IR.h>
#include <FEXCore/Utils/LogManager.h>
@@ -12,7 +13,6 @@ $end_info$
#include <array>
#include <stdint.h>
#include <utility>
#include <xbyak/xbyak.h>
namespace FEXCore::CPU {
#define DEF_OP(x) void X86JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
@@ -7,6 +7,7 @@ $end_info$
#include "Interface/Context/Context.h"
#include "Interface/Core/CPUID.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Core/Dispatcher/X86Dispatcher.h"
#include "Interface/Core/LookupCache.h"
#include "Interface/Core/JIT/x86_64/JITClass.h"
#include "Interface/HLE/Thunks/Thunks.h"
@@ -25,27 +26,10 @@ $end_info$
#include <stdint.h>
#include <unordered_map>
#include <utility>
#include <xbyak/xbyak.h>
namespace FEXCore::CPU {
#define DEF_OP(x) void X86JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(SignalReturn) {
auto Op = IROp->C<IR::IROp_SignalReturn>();
// Adjust the stack first for a regular return
if (SpillSlots) {
add(rsp, SpillSlots * MaxSpillSlotSize); // + 8 to consume return address
}
if (Op->IsRT) {
jmp(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.SignalReturnHandlerRT)]);
}
else {
jmp(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.SignalReturnHandler)]);
}
}
DEF_OP(CallbackReturn) {
// Adjust the stack first for a regular return
if (SpillSlots) {
@@ -162,6 +146,7 @@ DEF_OP(Syscall) {
auto Op = IROp->C<IR::IROp_Syscall>();
// XXX: This is very terrible, but I don't care for right now
FEXCore::IR::SyscallFlags Flags = Op->Flags;
auto NumPush = RA64.size();
for (auto &Reg : RA64)
@@ -202,7 +187,11 @@ DEF_OP(Syscall) {
for (uint32_t i = RA64.size(); i > 0; --i)
pop(RA64[i - 1]);
mov (GetDst<RA_64>(Node), rax);
if ((Flags & FEXCore::IR::SyscallFlags::NORETURNEDRESULT) != FEXCore::IR::SyscallFlags::NORETURNEDRESULT) {
// Move result to its destination register.
// Only if `NORETURNEDRESULT` wasn't set, otherwise we might overwrite the CPUState refilled with `FillStaticRegs`
mov (GetDst<RA_64>(Node), rax);
}
}
DEF_OP(Thunk) {
@@ -218,7 +207,7 @@ DEF_OP(Thunk) {
mov(rdi, GetSrc<RA_64>(Op->ArgPtr.ID()));
auto thunkFn = ThreadState->CTX->ThunkHandler->LookupThunk(Op->ThunkNameHash);
auto thunkFn = static_cast<Context::ContextImpl*>(ThreadState->CTX)->ThunkHandler->LookupThunk(Op->ThunkNameHash);
mov(rax, reinterpret_cast<uintptr_t>(thunkFn));
call(rax);
@@ -318,10 +307,45 @@ DEF_OP(CPUID) {
mov(Dst.second, rdx);
}
DEF_OP(XGETBV) {
auto Op = IROp->C<IR::IROp_XGetBV>();
for (auto &Reg : RA64)
push(Reg);
// CPUID ABI
// this: rdi
// Function: rsi
//
// Result: RAX, RDX. 4xi32
// rsi can be in the source registers, so copy argument to edx first
mov (esi, GetSrc<RA_32>(Op->Function.ID()));
mov (rdi, qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.CPUIDObj)]);
auto NumPush = RA64.size();
if (NumPush & 1)
sub(rsp, 8); // Align
// {rdi, rsi, rdx}
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.XCRFunction)]);
if (NumPush & 1)
add(rsp, 8); // Align
for (uint32_t i = RA64.size(); i > 0; --i)
pop(RA64[i - 1]);
auto Dst = GetSrcPair<RA_64>(Node);
mov(Dst.first.cvt32(), eax);
mov(Dst.second, rax);
shr(Dst.second, 32);
}
#undef DEF_OP
void X86JITCore::RegisterBranchHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &X86JITCore::Op_##x
REGISTER_OP(SIGNALRETURN, SignalReturn);
REGISTER_OP(CALLBACKRETURN, CallbackReturn);
REGISTER_OP(EXITFUNCTION, ExitFunction);
REGISTER_OP(JUMP, Jump);
@@ -331,6 +355,7 @@ void X86JITCore::RegisterBranchHandlers() {
REGISTER_OP(VALIDATECODE, ValidateCode);
REGISTER_OP(THREADREMOVECODEENTRY, ThreadRemoveCodeEntry);
REGISTER_OP(CPUID, CPUID);
REGISTER_OP(XGETBV, XGETBV);
#undef REGISTER_OP
}
}
@@ -5,13 +5,12 @@ $end_info$
*/
#include "Interface/Core/JIT/x86_64/JITClass.h"
#include "Interface/Core/Dispatcher/X86Dispatcher.h"
#include <FEXCore/IR/IR.h>
#include <FEXCore/Utils/LogManager.h>
#include <array>
#include <stdint.h>
#include <xbyak/xbyak.h>
namespace FEXCore::CPU {
@@ -111,6 +110,53 @@ DEF_OP(VCastFromGPR) {
}
}
DEF_OP(VDupFromGPR) {
const auto Op = IROp->C<IR::IROp_VDupFromGPR>();
const auto OpSize = IROp->Size;
const auto ElementSize = IROp->ElementSize;
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto Dst = GetDst(Node);
const auto Src = GetSrc<RA_64>(Op->Src.ID()).cvt64();
vmovq(Dst, Src);
switch (ElementSize) {
case 1:
if (Is256Bit) {
vpbroadcastb(ToYMM(Dst), Dst);
} else {
vpbroadcastb(Dst, Dst);
}
break;
case 2:
if (Is256Bit) {
vpbroadcastw(ToYMM(Dst), Dst);
} else {
vpbroadcastw(Dst, Dst);
}
break;
case 4:
if (Is256Bit) {
vpbroadcastd(ToYMM(Dst), Dst);
} else {
vpbroadcastd(Dst, Dst);
}
break;
case 8:
if (Is256Bit) {
vpbroadcastq(ToYMM(Dst), Dst);
} else {
vpbroadcastq(Dst, Dst);
}
break;
default:
LOGMAN_MSG_A_FMT("Unhandled element size: {}", ElementSize);
return;
}
}
DEF_OP(Float_FromGPR_S) {
const auto Op = IROp->C<IR::IROp_Float_FromGPR_S>();
@@ -357,6 +403,7 @@ void X86JITCore::RegisterConversionHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &X86JITCore::Op_##x
REGISTER_OP(VINSGPR, VInsGPR);
REGISTER_OP(VCASTFROMGPR, VCastFromGPR);
REGISTER_OP(VDUPFROMGPR, VDupFromGPR);
REGISTER_OP(FLOAT_FROMGPR_S, Float_FromGPR_S);
REGISTER_OP(FLOAT_FTOF, Float_FToF);
REGISTER_OP(VECTOR_STOF, Vector_SToF);
@@ -5,12 +5,11 @@ $end_info$
*/
#include "Interface/Core/JIT/x86_64/JITClass.h"
#include "Interface/Core/Dispatcher/X86Dispatcher.h"
#include <FEXCore/IR/IR.h>
#include <array>
#include <stdint.h>
#include <xbyak/xbyak.h>
namespace FEXCore::CPU {
#define DEF_OP(x) void X86JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
@@ -21,23 +20,67 @@ DEF_OP(AESImc) {
}
DEF_OP(AESEnc) {
auto Op = IROp->C<IR::IROp_VAESEnc>();
vaesenc(GetDst(Node), GetSrc(Op->State.ID()), GetSrc(Op->Key.ID()));
const auto Op = IROp->C<IR::IROp_VAESEnc>();
const auto OpSize = IROp->Size;
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto Dst = GetDst(Node);
const auto Key = GetSrc(Op->Key.ID());
const auto State = GetSrc(Op->State.ID());
if (Is256Bit) {
vaesenc(ToYMM(Dst), ToYMM(State), ToYMM(Key));
} else {
vaesenc(Dst, State, Key);
}
}
DEF_OP(AESEncLast) {
auto Op = IROp->C<IR::IROp_VAESEncLast>();
vaesenclast(GetDst(Node), GetSrc(Op->State.ID()), GetSrc(Op->Key.ID()));
const auto Op = IROp->C<IR::IROp_VAESEncLast>();
const auto OpSize = IROp->Size;
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto Dst = GetDst(Node);
const auto Key = GetSrc(Op->Key.ID());
const auto State = GetSrc(Op->State.ID());
if (Is256Bit) {
vaesenclast(ToYMM(Dst), ToYMM(State), ToYMM(Key));
} else {
vaesenclast(Dst, State, Key);
}
}
DEF_OP(AESDec) {
auto Op = IROp->C<IR::IROp_VAESDec>();
vaesdec(GetDst(Node), GetSrc(Op->State.ID()), GetSrc(Op->Key.ID()));
const auto Op = IROp->C<IR::IROp_VAESDec>();
const auto OpSize = IROp->Size;
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto Dst = GetDst(Node);
const auto Key = GetSrc(Op->Key.ID());
const auto State = GetSrc(Op->State.ID());
if (Is256Bit) {
vaesdec(ToYMM(Dst), ToYMM(State), ToYMM(Key));
} else {
vaesdec(Dst, State, Key);
}
}
DEF_OP(AESDecLast) {
auto Op = IROp->C<IR::IROp_VAESDecLast>();
vaesdeclast(GetDst(Node), GetSrc(Op->State.ID()), GetSrc(Op->Key.ID()));
const auto Op = IROp->C<IR::IROp_VAESDecLast>();
const auto OpSize = IROp->Size;
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto Dst = GetDst(Node);
const auto Key = GetSrc(Op->Key.ID());
const auto State = GetSrc(Op->State.ID());
if (Is256Bit) {
vaesdeclast(ToYMM(Dst), ToYMM(State), ToYMM(Key));
} else {
vaesdeclast(Dst, State, Key);
}
}
DEF_OP(AESKeyGenAssist) {
@@ -76,18 +119,24 @@ DEF_OP(CRC32) {
}
DEF_OP(PCLMUL) {
auto Op = IROp->C<IR::IROp_PCLMUL>();
const auto Op = IROp->C<IR::IROp_PCLMUL>();
const auto OpSize = IROp->Size;
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
auto Dst = GetDst(Node);
auto Src1 = GetSrc(Op->Src1.ID());
auto Src2 = GetSrc(Op->Src2.ID());
const auto Dst = GetDst(Node);
const auto Src1 = GetSrc(Op->Src1.ID());
const auto Src2 = GetSrc(Op->Src2.ID());
switch (Op->Selector) {
case 0b00000000:
case 0b00000001:
case 0b00010000:
case 0b00010001:
vpclmulqdq(Dst, Src1, Src2, Op->Selector);
if (Is256Bit) {
vpclmulqdq(ToYMM(Dst), ToYMM(Src1), ToYMM(Src2), Op->Selector);
} else {
vpclmulqdq(Dst, Src1, Src2, Op->Selector);
}
break;
default:
LOGMAN_MSG_A_FMT("Unknown PCLMUL selector: {}", Op->Selector);
@@ -5,12 +5,12 @@ $end_info$
*/
#include "Interface/Core/JIT/x86_64/JITClass.h"
#include "Interface/Core/Dispatcher/X86Dispatcher.h"
#include <FEXCore/IR/IR.h>
#include <array>
#include <stdint.h>
#include <xbyak/xbyak.h>
namespace FEXCore::CPU {
+112 -29
View File
@@ -19,7 +19,6 @@ $end_info$
#include <FEXCore/Core/CPUBackend.h>
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Core/SignalDelegator.h>
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IntrusiveIRList.h>
@@ -28,6 +27,7 @@ $end_info$
#include <FEXCore/Utils/EnumUtils.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/Profiler.h>
#include <FEXCore/fextl/sstream.h>
#include <algorithm>
#include <array>
@@ -35,12 +35,9 @@ $end_info$
#include <stddef.h>
#include <stdint.h>
#include <signal.h>
#include <sys/mman.h>
#include <tuple>
#include <unordered_map>
#include <utility>
#include <vector>
#include <xbyak/xbyak.h>
// #define DEBUG_RA 1
// #define DEBUG_CYCLES
@@ -307,6 +304,66 @@ void X86JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
}
break;
case FABI_I32_I64_I64_I128_I128_I16: {
PushRegs();
const auto Op = IROp->C<IR::IROp_VPCMPESTRX>();
const auto Is64Bit = Op->GPRSize == 8;
const auto LHS = GetSrc(Op->LHS.ID());
const auto RHS = GetSrc(Op->RHS.ID());
const auto SrcRAX = GetSrc<RA_64>(Op->RAX.ID());
const auto SrcRDX = GetSrc<RA_64>(Op->RDX.ID());
// Encode the size check into the 8th bit to save a parameter
const auto Control = Op->Control | (uint16_t(Is64Bit) << 8);
mov(rdi, SrcRAX);
mov(rsi, SrcRDX);
movq(rdx, LHS);
pextrq(rcx, LHS, 1);
movq(r8, RHS);
pextrq(r9, RHS, 1);
sub(rsp, 16);
mov(dword [rsp], Control);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])]);
add(rsp, 16);
PopRegs();
mov(GetDst<RA_32>(Node), rax);
break;
}
case FABI_I32_I128_I128_I16: {
PushRegs();
const auto Op = IROp->C<IR::IROp_VPCMPISTRX>();
const auto LHS = GetSrc(Op->LHS.ID());
const auto RHS = GetSrc(Op->RHS.ID());
const auto Control = Op->Control;
movq(rdi, LHS);
pextrq(rsi, LHS, 1);
movq(rdx, RHS);
pextrq(rcx, RHS, 1);
mov(r8, Control);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])]);
PopRegs();
mov(GetDst<RA_32>(Node), rax);
break;
}
case FABI_UNKNOWN:
default:
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
@@ -330,7 +387,7 @@ static uint64_t X86JITCore_ExitFunctionLink(FEXCore::Core::CpuStateFrame *Frame,
}
auto LinkerAddress = Frame->Pointers.Common.ExitFunctionLinker;
Context::Context::ThreadAddBlockLink(Thread, GuestRip, (uintptr_t)record, [record, LinkerAddress]{
Context::ContextImpl::ThreadAddBlockLink(Thread, GuestRip, (uintptr_t)record, [record, LinkerAddress]{
// undo the link
record[0] = LinkerAddress;
});
@@ -342,14 +399,14 @@ static uint64_t X86JITCore_ExitFunctionLink(FEXCore::Core::CpuStateFrame *Frame,
void X86JITCore::Op_NoOp(IR::IROp_Header *IROp, IR::NodeID Node) {
}
X86JITCore::X86JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread)
X86JITCore::X86JITCore(FEXCore::Context::ContextImpl *ctx, FEXCore::Core::InternalThreadState *Thread)
: CPUBackend(Thread, INITIAL_CODE_SIZE, MAX_CODE_SIZE)
, CodeGenerator(0, this, nullptr) // this is not used here
, CTX {ctx} {
RAPass = Thread->PassManager->GetPass<IR::RegisterAllocationPass>("RA");
RAPass->AllocateRegisterSet(RegisterCount, RegisterClasses);
RAPass->AllocateRegisterSet(RegisterClasses);
RAPass->AddRegisters(FEXCore::IR::GPRClass, NumGPRs);
RAPass->AddRegisters(FEXCore::IR::FPRClass, NumXMMs);
RAPass->AddRegisters(FEXCore::IR::GPRPairClass, NumGPRPairs);
@@ -379,7 +436,7 @@ X86JITCore::X86JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalTh
Common.PrintValue = reinterpret_cast<uint64_t>(PrintValue);
Common.PrintVectorValue = reinterpret_cast<uint64_t>(PrintVectorValue);
Common.ThreadRemoveCodeEntryFromJIT = reinterpret_cast<uintptr_t>(&Context::Context::ThreadRemoveCodeEntryFromJit);
Common.ThreadRemoveCodeEntryFromJIT = reinterpret_cast<uintptr_t>(&Context::ContextImpl::ThreadRemoveCodeEntryFromJit);
Common.CPUIDObj = reinterpret_cast<uint64_t>(&CTX->CPUID);
{
@@ -387,9 +444,14 @@ X86JITCore::X86JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalTh
Common.CPUIDFunction = PMF.GetConvertedPointer();
}
{
FEXCore::Utils::MemberFunctionToPointerCast PMF(&FEXCore::CPUIDEmu::RunXCRFunction);
Common.XCRFunction = PMF.GetConvertedPointer();
}
Common.SyscallHandlerObj = reinterpret_cast<uint64_t>(CTX->SyscallHandler);
Common.SyscallHandlerFunc = reinterpret_cast<uint64_t>(FEXCore::Context::HandleSyscall);
Common.ExitFunctionLink = reinterpret_cast<uintptr_t>(&Context::Context::ThreadExitFunctionLink<X86JITCore_ExitFunctionLink>);
Common.ExitFunctionLink = reinterpret_cast<uintptr_t>(&Context::ContextImpl::ThreadExitFunctionLink<X86JITCore_ExitFunctionLink>);
// Fill in the fallback handlers
InterpreterOps::FillFallbackIndexPointers(Common.FallbackHandlerPointers);
@@ -399,12 +461,6 @@ X86JITCore::X86JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalTh
ClearCache();
}
void X86JITCore::InitializeSignalHandlers(FEXCore::Context::Context *CTX) {
CTX->SignalDelegation->RegisterHostSignalHandler(SIGILL, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
return Thread->CTX->Dispatcher->HandleSIGILL(Thread, Signal, info, ucontext);
}, true);
}
X86JITCore::~X86JITCore() {
}
@@ -587,7 +643,7 @@ std::tuple<X86JITCore::SetCC, X86JITCore::CMovCC, X86JITCore::JCC> X86JITCore::G
return { &CodeGenerator::sete , &CodeGenerator::cmove , &CodeGenerator::je };
}
void *X86JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IRListView const *IR, [[maybe_unused]] FEXCore::Core::DebugData *DebugData, FEXCore::IR::RegisterAllocationData *RAData, bool GDBEnabled) {
CPUBackend::CompiledCode X86JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IRListView const *IR, [[maybe_unused]] FEXCore::Core::DebugData *DebugData, FEXCore::IR::RegisterAllocationData *RAData, bool GDBEnabled) {
FEXCORE_PROFILE_SCOPED("x86::CompileCode");
JumpTargets.clear();
@@ -603,12 +659,27 @@ void *X86JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IRLi
CTX->ClearCodeCache(ThreadState);
}
GuestEntry = getCurr<uint8_t*>();
CodeData.BlockBegin = getCurr<uint8_t*>();
// Put the code header at the start of the data block.
Label JITCodeHeaderLabel{};
L(JITCodeHeaderLabel);
JITCodeHeader *CodeHeader = getCurr<JITCodeHeader *>();
setSize(getSize() + sizeof(JITCodeHeader));
CodeData.BlockEntry = getCurr<uint8_t*>();
// Get the address of the JITCodeHeader and store in to the core state.
// Only two instructions, so very low overhead.
lea(TMP1, ptr [rip + JITCodeHeaderLabel]);
mov(qword [STATE + offsetof(FEXCore::Core::CPUState, InlineJITBlockHeader)], TMP1);
CursorEntry = getSize();
this->IR = IR;
if (GDBEnabled) {
auto GDBSize = CTX->Dispatcher->GenerateGDBPauseCheck(GuestEntry, Entry);
auto GDBSize = CTX->Dispatcher->GenerateGDBPauseCheck(CodeData.BlockBegin, Entry);
setSize(getSize() + GDBSize);
}
@@ -691,7 +762,7 @@ void *X86JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IRLi
if (IROp->Op != IR::OP_BEGINBLOCK &&
IROp->Op != IR::OP_CONDJUMP &&
IROp->Op != IR::OP_JUMP) {
std::stringstream Inst;
fextl::stringstream Inst;
auto Name = FEXCore::IR::GetName(IROp->Op);
if (IROp->HasDest) {
@@ -731,7 +802,7 @@ void *X86JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IRLi
if (DebugData) {
DebugData->Subblocks.push_back({
static_cast<uint32_t>(BlockStartHostCode - GuestEntry),
static_cast<uint32_t>(BlockStartHostCode - CodeData.BlockBegin),
static_cast<uint32_t>(getCurr<uint8_t *>() - BlockStartHostCode)
});
}
@@ -744,29 +815,41 @@ void *X86JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IRLi
}
PendingTargetLabel = nullptr;
void *GuestExit = getCurr<void*>();
// Add the JitCodeTail
auto JITBlockTailLocation = getCurr<uint8_t *>();
auto JITBlockTail = getCurr<JITCodeTail*>();
setSize(getSize() + sizeof(JITCodeTail));
// Put the block's RIP entry in the tail.
// This will be used for RIP reconstruction in the future.
// TODO: This needs to be a data RIP relocation once code caching works.
// Current relocation code doesn't support this feature yet.
JITBlockTail->RIP = Entry;
CodeHeader->OffsetToBlockTail = JITBlockTailLocation - CodeData.BlockBegin;
CodeData.Size = getCurr<uint8_t*>() - CodeData.BlockBegin;
JITBlockTail->Size = CodeData.Size;
this->IR = nullptr;
ready();
if (DebugData) {
DebugData->HostCodeSize = reinterpret_cast<uintptr_t>(GuestExit) - reinterpret_cast<uintptr_t>(GuestEntry);
DebugData->HostCodeSize = CodeData.Size;
DebugData->Relocations = &Relocations;
}
return GuestEntry;
return CodeData;
}
std::unique_ptr<CPUBackend> CreateX86JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread) {
return std::make_unique<X86JITCore>(ctx, Thread);
fextl::unique_ptr<CPUBackend> CreateX86JITCore(FEXCore::Context::ContextImpl *ctx, FEXCore::Core::InternalThreadState *Thread) {
return fextl::make_unique<X86JITCore>(ctx, Thread);
}
CPUBackendFeatures GetX86JITBackendFeatures() {
return CPUBackendFeatures { };
}
void InitializeX86JITSignalHandlers(FEXCore::Context::Context *CTX) {
X86JITCore::InitializeSignalHandlers(CTX);
}
}
+21 -18
View File
@@ -9,18 +9,20 @@ $end_info$
#include <FEXCore/IR/RegisterAllocationData.h>
#include "Interface/Core/BlockSamplingData.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Core/Dispatcher/X86Dispatcher.h"
#include "Interface/Core/ObjectCache/Relocations.h"
#define XBYAK64
#include <xbyak/xbyak.h>
#include <xbyak/xbyak_util.h>
using namespace Xbyak;
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Core/CPUBackend.h>
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/Utils/MathUtils.h>
#include <FEXCore/fextl/string.h>
#include <FEXCore/fextl/unordered_map.h>
#include <FEXCore/fextl/vector.h>
#include "Interface/IR/Passes/RegisterAllocationPass.h"
#include <tuple>
@@ -51,13 +53,13 @@ const std::array<Xbyak::Xmm, 11> RAXMM_x = { xmm1, xmm2, xmm3, xmm4, xmm5, xmm6
class X86JITCore final : public CPUBackend, public Xbyak::CodeGenerator {
public:
explicit X86JITCore(FEXCore::Context::Context *ctx,
explicit X86JITCore(FEXCore::Context::ContextImpl *ctx,
FEXCore::Core::InternalThreadState *Thread);
~X86JITCore() override;
[[nodiscard]] std::string GetName() override { return "JIT"; }
[[nodiscard]] fextl::string GetName() override { return "JIT"; }
[[nodiscard]] void *CompileCode(uint64_t Entry,
[[nodiscard]] CPUBackend::CompiledCode CompileCode(uint64_t Entry,
FEXCore::IR::IRListView const *IR,
FEXCore::Core::DebugData *DebugData,
FEXCore::IR::RegisterAllocationData *RAData, bool GDBEnabled) override;
@@ -68,8 +70,6 @@ public:
void ClearCache() override;
static void InitializeSignalHandlers(FEXCore::Context::Context *CTX);
void ClearRelocations() override { Relocations.clear(); }
private:
@@ -123,7 +123,7 @@ private:
void PlaceNamedSymbolLiteral(NamedSymbolLiteralPair &Lit);
std::vector<FEXCore::CPU::Relocation> Relocations;
fextl::vector<FEXCore::CPU::Relocation> Relocations;
///< Relocation code loading
bool ApplyRelocations(uint64_t GuestEntry, uint64_t CodeEntry, uint64_t CursorEntry, size_t NumRelocations, const char* EntryRelocations);
@@ -135,11 +135,12 @@ private:
/** @} */
Label* PendingTargetLabel{};
FEXCore::Context::Context *CTX;
FEXCore::Context::ContextImpl *CTX;
FEXCore::IR::IRListView const *IR;
uint64_t Entry;
CPUBackend::CompiledCode CodeData{};
std::unordered_map<IR::NodeID, Label> JumpTargets;
fextl::unordered_map<IR::NodeID, Label> JumpTargets;
Xbyak::util::Cpu Features{};
bool MemoryDebug = false;
@@ -150,7 +151,6 @@ private:
constexpr static uint32_t NumGPRs = RA64.size(); // 4 is the minimum required for GPR ops
constexpr static uint32_t NumXMMs = RAXMM.size();
constexpr static uint32_t NumGPRPairs = RA64Pair.size();
constexpr static uint32_t RegisterCount = NumGPRs + NumXMMs + NumGPRPairs;
constexpr static uint32_t RegisterClasses = 6;
constexpr static uint64_t GPRBase = (0ULL << 32);
@@ -205,10 +205,6 @@ private:
void EmitDetectionString();
uint32_t SpillSlots{};
/**
* @brief Current guest RIP entrypoint
*/
uint8_t *GuestEntry{};
using SetCC = void (X86JITCore::*)(const Operand& op);
using CMovCC = void (X86JITCore::*)(const Reg& reg, const Operand& op);
@@ -308,7 +304,6 @@ private:
DEF_OP(AtomicFetchNeg);
///< Branch ops
DEF_OP(SignalReturn);
DEF_OP(CallbackReturn);
DEF_OP(ExitFunction);
DEF_OP(Jump);
@@ -318,10 +313,12 @@ private:
DEF_OP(ValidateCode);
DEF_OP(ThreadRemoveCodeEntry);
DEF_OP(CPUID);
DEF_OP(XGETBV);
///< Conversion ops
DEF_OP(VInsGPR);
DEF_OP(VCastFromGPR);
DEF_OP(VDupFromGPR);
DEF_OP(Float_FromGPR_S);
DEF_OP(Float_FToF);
DEF_OP(Vector_UToF);
@@ -347,6 +344,10 @@ private:
DEF_OP(StoreFlag);
DEF_OP(LoadMem);
DEF_OP(StoreMem);
DEF_OP(VLoadVectorMasked);
DEF_OP(VStoreVectorMasked);
DEF_OP(MemSet);
DEF_OP(MemCpy);
DEF_OP(CacheLineClear);
DEF_OP(CacheLineClean);
DEF_OP(CacheLineZero);
@@ -409,6 +410,8 @@ private:
DEF_OP(VZip2);
DEF_OP(VUnZip);
DEF_OP(VUnZip2);
DEF_OP(VTrn);
DEF_OP(VTrn2);
DEF_OP(VBSL);
DEF_OP(VCMPEQ);
DEF_OP(VCMPEQZ);
@@ -6,7 +6,7 @@ $end_info$
#include "Interface/Core/CPUID.h"
#include "Interface/Core/JIT/x86_64/JITClass.h"
#include "Interface/Core/Dispatcher/X86Dispatcher.h"
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/IR/IR.h>
#include <FEXCore/Utils/LogManager.h>
@@ -14,7 +14,6 @@ $end_info$
#include <array>
#include <stddef.h>
#include <stdint.h>
#include <xbyak/xbyak.h>
namespace FEXCore::CPU {
@@ -766,6 +765,302 @@ DEF_OP(StoreMem) {
}
}
DEF_OP(VLoadVectorMasked) {
const auto Op = IROp->C<IR::IROp_VLoadVectorMasked>();
const auto OpSize = IROp->Size;
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto ElementSize = IROp->ElementSize;
const auto Dst = GetDst(Node);
const auto Mask = GetSrc(Op->Mask.ID());
const Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Addr.ID());
const auto MemPtr = GenerateModRM(MemReg, Op->Offset, Op->OffsetType, Op->OffsetScale);
switch (ElementSize) {
case 4: {
if (Is256Bit) {
vmaskmovps(ToYMM(Dst), ToYMM(Mask), yword [MemPtr]);
} else {
vmaskmovps(Dst, Mask, xword [MemPtr]);
}
return;
}
case 8: {
if (Is256Bit) {
vmaskmovpd(ToYMM(Dst), ToYMM(Mask), yword [MemPtr]);
} else {
vmaskmovpd(Dst, Mask, xword [MemPtr]);
}
return;
}
default:
LOGMAN_MSG_A_FMT("Unhandled VLoadVectorMasked element size: {}", ElementSize);
return;
}
}
DEF_OP(VStoreVectorMasked) {
const auto Op = IROp->C<IR::IROp_VStoreVectorMasked>();
const auto OpSize = IROp->Size;
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto ElementSize = IROp->ElementSize;
const auto Data = GetDst(Op->Data.ID());
const auto Mask = GetSrc(Op->Mask.ID());
const Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Addr.ID());
const auto MemPtr = GenerateModRM(MemReg, Op->Offset, Op->OffsetType, Op->OffsetScale);
switch (ElementSize) {
case 4: {
if (Is256Bit) {
vmaskmovps(yword [MemPtr], ToYMM(Mask), ToYMM(Data));
} else {
vmaskmovps(xword [MemPtr], Mask, Data);
}
return;
}
case 8: {
if (Is256Bit) {
vmaskmovpd(yword [MemPtr], ToYMM(Mask), ToYMM(Data));
} else {
vmaskmovpd(xword [MemPtr], Mask, Data);
}
return;
}
default:
LOGMAN_MSG_A_FMT("Unhandled VStoreVectorMasked element size: {}", ElementSize);
return;
}
}
DEF_OP(MemSet) {
const auto Op = IROp->C<IR::IROp_MemSet>();
const int32_t Size = Op->Size;
const auto MemReg = GetSrc<RA_64>(Op->Addr.ID());
const auto Value = GetSrc<RA_64>(Op->Value.ID());
const auto Length = GetSrc<RA_64>(Op->Length.ID());
const auto Direction = GetSrc<RA_64>(Op->Direction.ID());
const auto Dst = GetSrc<RA_64>(Node);
// If Direction == 0 then:
// MemReg is incremented (by size)
// else:
// MemReg is decremented (by size)
//
// Counter is decremented regardless.
// TMP1 = rax
// TMP2 = rcx
// TMP4 = rdi
// That leaves us with TMP3 and TMP5
mov(rax, Value);
mov(rcx, Length);
mov(rdi, MemReg);
if (!Op->Prefix.IsInvalid()) {
add(rdi, GetSrc<RA_64>(Op->Prefix.ID()));
}
{
mov(TMP3, Length);
auto CalculateDest = [&]() {
mov(Dst, MemReg);
switch (Size) {
case 1:
break;
case 2:
shl(TMP3, 1);
break;
case 4:
shl(TMP3, 2);
break;
case 8:
shl(TMP3, 3);
break;
default:
LOGMAN_MSG_A_FMT("Unhandled {} size: {}", __func__, Size);
break;
}
};
Label AfterDir;
Label BackwardDir;
cmp(Direction, 0);
jne(BackwardDir);
// Incrementing DF flag.
cld();
CalculateDest();
add(Dst, TMP3);
jmp(AfterDir);
L(BackwardDir);
// Decrementing DF flag.
std();
CalculateDest();
sub(Dst, TMP3);
L(AfterDir);
}
switch (Size) {
case 1:
rep(); stosb();
break;
case 2:
rep(); stosw();
break;
case 4:
rep(); stosd();
break;
case 8:
rep(); stosq();
break;
default:
LOGMAN_MSG_A_FMT("Unhandled {} size: {}", __func__, Size);
break;
}
// Ensure we set DF back to zero. Required by the ABI.
cld();
}
DEF_OP(MemCpy) {
const auto Op = IROp->C<IR::IROp_MemCpy>();
const int32_t Size = Op->Size;
const auto MemRegDest = GetSrc<RA_64>(Op->AddrDest.ID());
const auto MemRegSrc = GetSrc<RA_64>(Op->AddrSrc.ID());
const auto Length = GetSrc<RA_64>(Op->Length.ID());
const auto Direction = GetSrc<RA_64>(Op->Direction.ID());
// If Direction == 0 then:
// MemRegDest is incremented (by size)
// MemRegSrc is incremented (by size)
// else:
// MemRegDest is decremented (by size)
// MemRegSrc is decremented (by size)
//
// Counter is decremented regardless.
// TMP1 = Length
// TMP2 = Dest
// TMP3 = Src
// TMP4 = Temp value
mov(TMP1, Length);
mov(TMP2, MemRegDest);
mov(TMP3, MemRegSrc);
if (!Op->PrefixDest.IsInvalid()) {
add(TMP2, GetSrc<RA_64>(Op->PrefixDest.ID()));
}
if (!Op->PrefixSrc.IsInvalid()) {
add(TMP3, GetSrc<RA_64>(Op->PrefixSrc.ID()));
}
auto Dst = GetSrcPair<RA_64>(Node);
Label Done;
Label BackwardImpl;
cmp(Direction, 0);
jne(BackwardImpl);
// Emit forward direction memcpy then backward direction memcpy.
for (int32_t Direction : { 1, -1 }) {
Label DoneInternal;
Label AgainInternal;
L(AgainInternal);
cmp(TMP1, 0);
je(DoneInternal);
{
switch (Size) {
case 1:
movzx(TMP4, byte [TMP3]);
mov(byte [TMP2], TMP4.cvt8());
break;
case 2:
movzx(TMP4, word [TMP3]);
mov(word [TMP2], TMP4.cvt16());
break;
case 4:
mov(TMP4.cvt32(), dword [TMP3]);
mov(dword [TMP2], TMP4.cvt32());
break;
case 8:
mov(TMP4, qword [TMP3]);
mov(qword [TMP2], TMP4);
break;
default:
LOGMAN_MSG_A_FMT("Unhandled {} size: {}", __func__, Size);
break;
}
}
if (Direction == 1) {
// Incrementing pointers
add(TMP2, Size);
add(TMP3, Size);
}
else {
// Decrementing pointers
sub(TMP2, Size);
sub(TMP3, Size);
}
// Decrement counter by one
sub(TMP1, 1);
jmp(AgainInternal);
L(DoneInternal);
// Pointer math using source pointers and length.
mov(TMP3, Length);
switch (Size) {
case 1:
break;
case 2:
shl(TMP3, 1);
break;
case 4:
shl(TMP3, 2);
break;
case 8:
shl(TMP3, 3);
break;
default:
LOGMAN_MSG_A_FMT("Unhandled {} size: {}", __func__, Size);
break;
}
// Needs to use temporaries just in case of overwrite
mov(TMP1, MemRegDest);
mov(TMP2, MemRegSrc);
mov(Dst.first, TMP1);
mov(Dst.second, TMP2);
if (Direction == 1) {
// Incrementing pointers
add(Dst.first, TMP3);
add(Dst.second, TMP3);
jmp(Done);
L(BackwardImpl);
}
else {
// Decrementing pointers
sub(Dst.first, TMP3);
sub(Dst.second, TMP3);
}
}
L(Done);
}
DEF_OP(CacheLineClear) {
auto Op = IROp->C<IR::IROp_CacheLineClear>();
@@ -820,6 +1115,10 @@ void X86JITCore::RegisterMemoryHandlers() {
REGISTER_OP(STOREMEM, StoreMem);
REGISTER_OP(LOADMEMTSO, LoadMem);
REGISTER_OP(STOREMEMTSO, StoreMem);
REGISTER_OP(VLOADVECTORMASKED, VLoadVectorMasked);
REGISTER_OP(VSTOREVECTORMASKED, VStoreVectorMasked);
REGISTER_OP(MEMSET, MemSet);
REGISTER_OP(MEMCPY, MemCpy);
REGISTER_OP(CACHELINECLEAR, CacheLineClear);
REGISTER_OP(CACHELINECLEAN, CacheLineClean);
REGISTER_OP(CACHELINEZERO, CacheLineZero);
@@ -6,6 +6,7 @@ $end_info$
#include "Interface/Context/Context.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Core/Dispatcher/X86Dispatcher.h"
#include "Interface/Core/JIT/x86_64/JITClass.h"
#include "FEXCore/Debug/InternalThreadState.h"
@@ -16,7 +17,6 @@ $end_info$
#include <array>
#include <stddef.h>
#include <stdint.h>
#include <xbyak/xbyak.h>
namespace FEXCore::CPU {
#define DEF_OP(x) void X86JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
@@ -24,7 +24,7 @@ namespace FEXCore::CPU {
DEF_OP(GuestOpcode) {
auto Op = IROp->C<IR::IROp_GuestOpcode>();
// metadata
DebugData->GuestOpcodes.push_back({Op->GuestEntryOffset, getCurr<uint8_t*>() - GuestEntry});
DebugData->GuestOpcodes.push_back({Op->GuestEntryOffset, getCurr<uint8_t*>() - CodeData.BlockBegin});
}
DEF_OP(Fence) {
@@ -43,6 +43,7 @@ DEF_OP(Fence) {
}
}
#ifndef _WIN32
DEF_OP(Break) {
auto Op = IROp->C<IR::IROp_Break>();
@@ -79,6 +80,11 @@ DEF_OP(Break) {
break;
}
}
#else
DEF_OP(Break) {
ERROR_AND_DIE_FMT("Unsupported");
}
#endif
DEF_OP(GetRoundingMode) {
auto Dst = GetDst<RA_32>(Node);
@@ -5,7 +5,7 @@ $end_info$
*/
#include "Interface/Core/JIT/x86_64/JITClass.h"
#include "Interface/Core/Dispatcher/X86Dispatcher.h"
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/IR/IR.h>
@@ -5,14 +5,13 @@ $end_info$
*/
#include "Interface/Core/JIT/x86_64/JITClass.h"
#include "Interface/Core/Dispatcher/X86Dispatcher.h"
#include <FEXCore/IR/IR.h>
#include <FEXCore/Utils/LogManager.h>
#include <array>
#include <stddef.h>
#include <stdint.h>
#include <xbyak/xbyak.h>
namespace FEXCore::CPU {
@@ -1944,7 +1943,7 @@ DEF_OP(VUnZip2) {
}
case 8: {
if (Is256Bit) {
vshufpd(ToYMM(Dst), ToYMM(VectorLower), ToYMM(VectorUpper), 0b1'1);
vshufpd(ToYMM(Dst), ToYMM(VectorLower), ToYMM(VectorUpper), 0b11'11);
vpermq(ToYMM(Dst), ToYMM(Dst), 0b11'01'10'00);
} else {
vshufpd(Dst, VectorLower, VectorUpper, 0b1'1);
@@ -1958,6 +1957,191 @@ DEF_OP(VUnZip2) {
}
}
DEF_OP(VTrn) {
const auto Op = IROp->C<IR::IROp_VTrn>();
const auto OpSize = IROp->Size;
const auto ElementSize = Op->Header.ElementSize;
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto Dst = GetDst(Node);
const auto VectorLower = GetSrc(Op->VectorLower.ID());
const auto VectorUpper = GetSrc(Op->VectorUpper.ID());
const auto LoadPshufbReg = [&](Xbyak::Xmm reg, uint64_t lower) {
mov(rax, lower);
mov(rcx, 0x80'80'80'80'80'80'80'80);
vmovq(reg, rax);
pinsrq(reg, rcx, 1);
};
switch (ElementSize) {
case 1: {
LoadPshufbReg(xmm15, 0x0E'0C'0A'08'06'04'02'00);
if (Is256Bit) {
vinserti128(ymm15, ymm15, xmm15, 1);
vpshufb(ymm14, ToYMM(VectorLower), ymm15);
vpshufb(ymm13, ToYMM(VectorUpper), ymm15);
vpunpcklbw(ToYMM(Dst), ymm14, ymm13);
} else {
vpshufb(xmm14, VectorLower, xmm15);
vpshufb(xmm13, VectorUpper, xmm15);
vpunpcklbw(Dst, xmm14, xmm13);
}
break;
}
case 2: {
LoadPshufbReg(xmm15, 0x0D'0C'09'08'05'04'01'00);
if (Is256Bit) {
vinserti128(ymm15, ymm15, xmm15, 1);
vpshufb(ymm14, ToYMM(VectorLower), ymm15);
vpshufb(ymm13, ToYMM(VectorUpper), ymm15);
vpunpcklwd(ToYMM(Dst), ymm14, ymm13);
} else {
vpshufb(xmm14, VectorLower, xmm15);
vpshufb(xmm13, VectorUpper, xmm15);
vpunpcklwd(Dst, xmm14, xmm13);
}
break;
}
case 4: {
LoadPshufbReg(xmm15, 0x0B'0A'09'08'03'02'01'00);
if (Is256Bit) {
vinserti128(ymm15, ymm15, xmm15, 1);
vpshufb(ymm14, ToYMM(VectorLower), ymm15);
vpshufb(ymm13, ToYMM(VectorUpper), ymm15);
vpunpckldq(ToYMM(Dst), ymm14, ymm13);
} else {
vpshufb(xmm14, VectorLower, xmm15);
vpshufb(xmm13, VectorUpper, xmm15);
vpunpckldq(Dst, xmm14, xmm13);
}
break;
}
case 8: {
LoadPshufbReg(xmm15, 0x07'06'05'04'03'02'01'00);
if (Is256Bit) {
vinserti128(ymm15, ymm15, xmm15, 1);
vpshufb(ymm14, ToYMM(VectorLower), ymm15);
vpshufb(ymm13, ToYMM(VectorUpper), ymm15);
vpunpcklqdq(ToYMM(Dst), ymm14, ymm13);
} else {
vpshufb(xmm14, VectorLower, xmm15);
vpshufb(xmm13, VectorUpper, xmm15);
vpunpcklqdq(Dst, xmm14, xmm13);
}
break;
}
default:
LOGMAN_MSG_A_FMT("Unknown Element Size: {}", ElementSize);
return;
}
}
DEF_OP(VTrn2) {
const auto Op = IROp->C<IR::IROp_VTrn2>();
const auto OpSize = IROp->Size;
const auto ElementSize = Op->Header.ElementSize;
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto Dst = GetDst(Node);
const auto VectorLower = GetSrc(Op->VectorLower.ID());
const auto VectorUpper = GetSrc(Op->VectorUpper.ID());
const auto LoadPshufbReg = [&](Xbyak::Xmm reg, uint64_t lower) {
mov(rax, lower);
mov(rcx, 0x80'80'80'80'80'80'80'80);
vmovq(reg, rax);
pinsrq(reg, rcx, 1);
};
switch (ElementSize) {
case 1: {
LoadPshufbReg(xmm15, 0x0F'0D'0B'09'07'05'03'01);
if (Is256Bit) {
vinserti128(ymm15, ymm15, xmm15, 1);
vpshufb(ymm14, ToYMM(VectorLower), ymm15);
vpshufb(ymm13, ToYMM(VectorUpper), ymm15);
vpunpcklbw(ToYMM(Dst), ymm14, ymm13);
} else {
vpshufb(xmm14, VectorLower, xmm15);
vpshufb(xmm13, VectorUpper, xmm15);
vpunpcklbw(Dst, xmm14, xmm13);
}
break;
}
case 2: {
LoadPshufbReg(xmm15, 0x0F'0E'0B'0A'07'06'03'02);
if (Is256Bit) {
vinserti128(ymm15, ymm15, xmm15, 1);
vpshufb(ymm14, ToYMM(VectorLower), ymm15);
vpshufb(ymm13, ToYMM(VectorUpper), ymm15);
vpunpcklwd(ToYMM(Dst), ymm14, ymm13);
} else {
vpshufb(xmm14, VectorLower, xmm15);
vpshufb(xmm13, VectorUpper, xmm15);
vpunpcklwd(Dst, xmm14, xmm13);
}
break;
}
case 4: {
LoadPshufbReg(xmm15, 0x0F'0E'0D'0C'07'06'05'04);
if (Is256Bit) {
vinserti128(ymm15, ymm15, xmm15, 1);
vpshufb(ymm14, ToYMM(VectorLower), ymm15);
vpshufb(ymm13, ToYMM(VectorUpper), ymm15);
vpunpckldq(ToYMM(Dst), ymm14, ymm13);
} else {
vpshufb(xmm14, VectorLower, xmm15);
vpshufb(xmm13, VectorUpper, xmm15);
vpunpckldq(Dst, xmm14, xmm13);
}
break;
}
case 8: {
LoadPshufbReg(xmm15, 0x0F'0E'0D'0C'0B'0A'09'08);
if (Is256Bit) {
vinserti128(ymm15, ymm15, xmm15, 1);
vpshufb(ymm14, ToYMM(VectorLower), ymm15);
vpshufb(ymm13, ToYMM(VectorUpper), ymm15);
vpunpcklqdq(ToYMM(Dst), ymm14, ymm13);
} else {
vpshufb(xmm14, VectorLower, xmm15);
vpshufb(xmm13, VectorUpper, xmm15);
vpunpcklqdq(Dst, xmm14, xmm13);
}
break;
}
default:
LOGMAN_MSG_A_FMT("Unknown Element Size: {}", ElementSize);
return;
}
}
DEF_OP(VBSL) {
const auto Op = IROp->C<IR::IROp_VBSL>();
@@ -2541,11 +2725,71 @@ DEF_OP(VFCMPUNO) {
}
DEF_OP(VUShl) {
LOGMAN_MSG_A_FMT("Unimplemented");
const auto Op = IROp->C<IR::IROp_VUShl>();
const auto OpSize = IROp->Size;
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto ElementSize = IROp->ElementSize;
LOGMAN_THROW_AA_FMT(ElementSize == 4 || ElementSize == 8,
"VUShl only supports 32-bit and 64-bit elements");
const auto Dst = GetDst(Node);
const auto ShiftVector = GetSrc(Op->ShiftVector.ID());
const auto Vector = GetSrc(Op->Vector.ID());
switch (ElementSize) {
case 4:
if (Is256Bit) {
vpsllvd(ToYMM(Dst), ToYMM(Vector), ToYMM(ShiftVector));
} else {
vpsllvd(Dst, Vector, ShiftVector);
}
return;
case 8:
if (Is256Bit) {
vpsllvq(ToYMM(Dst), ToYMM(Vector), ToYMM(ShiftVector));
} else {
vpsllvq(Dst, Vector, ShiftVector);
}
return;
default:
LOGMAN_MSG_A_FMT("Unknown Element Size: {}", ElementSize);
return;
}
}
DEF_OP(VUShr) {
LOGMAN_MSG_A_FMT("Unimplemented");
const auto Op = IROp->C<IR::IROp_VUShr>();
const auto OpSize = IROp->Size;
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto ElementSize = IROp->ElementSize;
LOGMAN_THROW_AA_FMT(ElementSize == 4 || ElementSize == 8,
"VUShr only supports 32-bit and 64-bit elements");
const auto Dst = GetDst(Node);
const auto ShiftVector = GetSrc(Op->ShiftVector.ID());
const auto Vector = GetSrc(Op->Vector.ID());
switch (ElementSize) {
case 4:
if (Is256Bit) {
vpsrlvd(ToYMM(Dst), ToYMM(Vector), ToYMM(ShiftVector));
} else {
vpsrlvd(Dst, Vector, ShiftVector);
}
return;
case 8:
if (Is256Bit) {
vpsrlvq(ToYMM(Dst), ToYMM(Vector), ToYMM(ShiftVector));
} else {
vpsrlvq(Dst, Vector, ShiftVector);
}
return;
default:
LOGMAN_MSG_A_FMT("Unknown Element Size: {}", ElementSize);
return;
}
}
DEF_OP(VSShr) {
@@ -3147,6 +3391,23 @@ DEF_OP(VShlI) {
const auto Vector = GetSrc(Op->Vector.ID());
switch (ElementSize) {
case 1: {
const auto Mask = 0xFFU >> BitShift;
mov(rax, Mask);
vmovq(xmm15, rax);
if (Is256Bit) {
vpsllw(ToYMM(Dst), ToYMM(Vector), BitShift);
vpbroadcastb(ymm15, xmm15);
vpand(ToYMM(Dst), ToYMM(Dst), ymm15);
} else {
vpsllw(Dst, Vector, BitShift);
vpbroadcastb(xmm15, xmm15);
vpand(Dst, Dst, ymm15);
}
break;
}
case 2: {
if (Is256Bit) {
vpsllw(ToYMM(Dst), ToYMM(Vector), BitShift);
@@ -4304,6 +4565,8 @@ void X86JITCore::RegisterVectorHandlers() {
REGISTER_OP(VZIP2, VZip2);
REGISTER_OP(VUNZIP, VUnZip);
REGISTER_OP(VUNZIP2, VUnZip2);
REGISTER_OP(VTRN, VTrn);
REGISTER_OP(VTRN2, VTrn2);
REGISTER_OP(VBSL, VBSL);
REGISTER_OP(VCMPEQ, VCMPEQ);
REGISTER_OP(VCMPEQZ, VCMPEQZ);
+10 -12
View File
@@ -11,15 +11,15 @@ $end_info$
#include "Interface/Context/Context.h"
#include "Interface/Core/LookupCache.h"
#include <sys/mman.h>
namespace FEXCore {
LookupCache::LookupCache(FEXCore::Context::Context *CTX)
: ctx {CTX} {
LookupCache::LookupCache(FEXCore::Context::ContextImpl *CTX)
: BlockLinks_mbr { fextl::pmr::get_default_resource() }
, ctx {CTX} {
TotalCacheSize = ctx->Config.VirtualMemSize / 4096 * 8 + CODE_SIZE + L1_SIZE;
BlockLinks_pma = fextl::make_unique<std::pmr::polymorphic_allocator<std::byte>>(&BlockLinks_mbr);
// Setup our PMR map.
BlockLinks = BlockLinks_pma.new_object<BlockLinksMapType>();
BlockLinks = BlockLinks_pma->new_object<BlockLinksMapType>();
// Block cache ends up looking like this
// PageMemoryMap[VirtualMemoryRegion >> 12]
@@ -33,7 +33,7 @@ LookupCache::LookupCache(FEXCore::Context::Context *CTX)
// Allocate a region of memory that we can use to back our block pointers
// We need one pointer per page of virtual memory
// At 64GB of virtual memory this will allocate 128MB of virtual memory space
PagePointer = reinterpret_cast<uintptr_t>(FEXCore::Allocator::mmap(nullptr, TotalCacheSize, PROT_READ | PROT_WRITE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0));
PagePointer = reinterpret_cast<uintptr_t>(FEXCore::Allocator::VirtualAlloc(TotalCacheSize));
// Allocate our memory backing our pages
// We need 32KB per guest page (One pointer per byte)
@@ -52,7 +52,7 @@ LookupCache::LookupCache(FEXCore::Context::Context *CTX)
LookupCache::~LookupCache() {
const size_t TotalCacheSize = ctx->Config.VirtualMemSize / 4096 * 8 + CODE_SIZE + L1_SIZE;
FEXCore::Allocator::munmap(reinterpret_cast<void*>(PagePointer), TotalCacheSize);
FEXCore::Allocator::VirtualFree(reinterpret_cast<void*>(PagePointer), TotalCacheSize);
// No need to free BlockLinks map.
// These will get freed when their memory allocators are deallocated.
@@ -62,7 +62,7 @@ void LookupCache::ClearL2Cache() {
std::lock_guard<std::recursive_mutex> lk(WriteLock);
// Clear out the page memory
// PagePointer and PageMemory are sequential with each other. Clear both at once.
madvise(reinterpret_cast<void*>(PagePointer), ctx->Config.VirtualMemSize / 4096 * 8 + CODE_SIZE, MADV_DONTNEED);
FEXCore::Allocator::VirtualDontNeed(reinterpret_cast<void*>(PagePointer), ctx->Config.VirtualMemSize / 4096 * 8 + CODE_SIZE);
AllocateOffset = 0;
}
@@ -70,11 +70,9 @@ void LookupCache::ClearCache() {
std::lock_guard<std::recursive_mutex> lk(WriteLock);
// Clear L1 and L2 by clearing the full cache.
madvise(reinterpret_cast<void*>(PagePointer), TotalCacheSize, MADV_DONTNEED);
// Clear the BlockLinks allocator which frees the BlockLinks map implicitly.
BlockLinks_mbr.release();
FEXCore::Allocator::VirtualDontNeed(reinterpret_cast<void*>(PagePointer), TotalCacheSize);
// Allocate a new pointer from the BlockLinks pma again.
BlockLinks = BlockLinks_pma.new_object<BlockLinksMapType>();
BlockLinks = BlockLinks_pma->new_object<BlockLinksMapType>();
// All code is gone, clear the block list
BlockList.clear();
}
+11 -15
View File
@@ -1,30 +1,28 @@
#pragma once
#include "Interface/Context/Context.h"
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/fextl/map.h>
#include <FEXCore/fextl/memory_resource.h>
#include <FEXCore/fextl/robin_map.h>
#include <FEXCore/fextl/vector.h>
#include <FEXCore/fextl/memory_resource.h>
#include <cstdint>
#include <functional>
#include <map>
#include <memory_resource>
#include <stddef.h>
#include <utility>
#include <vector>
#include <mutex>
#include <tsl/robin_map.h>
namespace FEXCore {
namespace Context {
struct Context;
}
class LookupCache {
public:
struct LookupCacheEntry {
uintptr_t HostCode;
uintptr_t GuestCode;
};
LookupCache(FEXCore::Context::Context *CTX);
LookupCache(FEXCore::Context::ContextImpl *CTX);
~LookupCache();
uintptr_t FindBlock(uint64_t Address) {
@@ -69,7 +67,7 @@ public:
return 0;
}
std::map<uint64_t, std::vector<uint64_t>> CodePages;
fextl::map<uint64_t, fextl::vector<uint64_t>> CodePages;
// Appends Block {Address} to CodePages [Start, Start + Length)
// Returns true if new pages are marked as containing code
@@ -170,8 +168,6 @@ public:
private:
void CacheBlockMapping(uint64_t Address, uintptr_t HostCode) {
std::lock_guard<std::recursive_mutex> lk(WriteLock);
// Do L1
auto &L1Entry = reinterpret_cast<LookupCacheEntry*>(L1Pointer)[Address & L1_ENTRIES_MASK];
L1Entry.GuestCode = Address;
@@ -247,10 +243,10 @@ private:
// This makes `BlockLinks` look like a raw pointer that could memory leak, but since it is backed by the MBR, it won't.
std::pmr::monotonic_buffer_resource BlockLinks_mbr;
using BlockLinksMapType = std::pmr::map<BlockLinkTag, std::function<void()>>;
std::pmr::polymorphic_allocator<std::byte> BlockLinks_pma {&BlockLinks_mbr};
fextl::unique_ptr<std::pmr::polymorphic_allocator<std::byte>> BlockLinks_pma;
BlockLinksMapType *BlockLinks;
tsl::robin_map<uint64_t, uint64_t> BlockList;
fextl::robin_map<uint64_t, uint64_t> BlockList;
size_t TotalCacheSize;
@@ -260,7 +256,7 @@ private:
size_t AllocateOffset {};
FEXCore::Context::Context *ctx;
FEXCore::Context::ContextImpl *ctx;
uint64_t VirtualMemSize{};
};
}
@@ -1,12 +1,16 @@
#pragma once
#include <FEXCore/Utils/CompilerDefs.h>
#include <cstdint>
namespace FEXCore::CodeSerialize {
// If any of the config options mismatch on load then the cache won't be used
// Any of these will result in codegen changes
struct CodeObjectSerializationConfig {
// Cookie in the header of the file, isn't part of the config hash
struct
FEX_PACKED
CodeObjectSerializationConfig {
// Cookie in the header of the file, isn't part of the config hash
uint64_t Cookie{};
// Instructions per block configuration
@@ -16,41 +20,45 @@ namespace FEXCore::CodeSerialize {
unsigned Arch : 4;
// Multiblock enabled
bool MultiBlock : 1;
unsigned MultiBlock : 1;
// Hardware TSO enabled
unsigned HardwareTSOEnabled : 1;
// TSO enabled
bool TSOEnabled : 1;
unsigned TSOEnabled : 1;
// ABI local flag unsafe optimization
bool ABILocalFlags : 1;
unsigned ABILocalFlags : 1;
// ABI no PF unsafe optimization
bool ABINoPF : 1;
unsigned ABINoPF : 1;
// Static register allocation enabled
bool SRA : 1;
unsigned SRA : 1;
// Paranoid TSO mode enabled
bool ParanoidTSO : 1;
unsigned ParanoidTSO : 1;
// Guest code execution mode (We don't support live mode switch)
bool Is64BitMode : 1;
unsigned Is64BitMode : 1;
// SMC checks style
unsigned SMCChecks : 2;
// x87 reduced precision
bool x87ReducedPrecision : 1;
unsigned x87ReducedPrecision : 1;
// Padding to remove uninitialized data warning from asan
// Shows remaining amount of bits available for config
unsigned _Pad : 18;
unsigned _Pad : 17;
bool operator==(CodeObjectSerializationConfig const &other) const {
return Cookie == other.Cookie &&
MaxInstPerBlock == other.MaxInstPerBlock &&
Arch == other.Arch &&
MultiBlock == other.MultiBlock &&
HardwareTSOEnabled == other.HardwareTSOEnabled &&
TSOEnabled == other.TSOEnabled &&
ABILocalFlags == other.ABILocalFlags &&
ABINoPF == other.ABINoPF &&
@@ -67,6 +75,7 @@ namespace FEXCore::CodeSerialize {
Hash <<= 32; Hash |= other.MaxInstPerBlock;
Hash <<= 1; Hash |= other.Arch;
Hash <<= 1; Hash |= other.MultiBlock;
Hash <<= 1; Hash |= other.HardwareTSOEnabled;
Hash <<= 1; Hash |= other.TSOEnabled;
Hash <<= 1; Hash |= other.ABILocalFlags;
Hash <<= 1; Hash |= other.ABINoPF;
@@ -79,6 +88,6 @@ namespace FEXCore::CodeSerialize {
}
};
static_assert(sizeof(CodeObjectSerializationConfig) == 16, "Size changed");
static_assert(sizeof(CodeObjectSerializationConfig) == 16, "Size changed");
static_assert((sizeof(CodeObjectSerializationConfig) - sizeof(uint64_t)) == 8, "Config size exceeded 64its. Need to change how the hash is generated!");
}
@@ -2,25 +2,24 @@
#include "Interface/Core/ObjectCache/ObjectCacheService.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/fextl/memory.h>
#include <FEXCore/fextl/string.h>
#include <FEXHeaderUtils/Filesystem.h>
#include <fcntl.h>
#include <filesystem>
#include <memory>
#include <string>
#include <sys/uio.h>
#include <sys/mman.h>
#include <xxhash.h>
namespace FEXCore::CodeSerialize {
void AsyncJobHandler::AsyncAddNamedRegionJob(uintptr_t Base, uintptr_t Size, uintptr_t Offset, const std::string &filename) {
void AsyncJobHandler::AsyncAddNamedRegionJob(uintptr_t Base, uintptr_t Size, uintptr_t Offset, const fextl::string &filename) {
#ifndef _WIN32
// This function adds a named region *JOB* to our named region handler
// This needs to be as fast as possible to keep out of the way of the JIT
auto BaseFilename = std::filesystem::path(filename).filename().string();
const fextl::string BaseFilename = FHU::Filesystem::GetFilename(filename);
if (!BaseFilename.empty()) {
// Create a new entry that once set up will be put in to our section object map
auto Entry = std::make_unique<CodeRegionEntry>(
auto Entry = fextl::make_unique<CodeRegionEntry>(
Base,
Size,
Offset,
@@ -77,12 +76,14 @@ namespace FEXCore::CodeSerialize {
// Tell the async thread that it has work to do
CodeObjectCacheService->NotifyWork();
}
#endif
}
void AsyncJobHandler::AsyncRemoveNamedRegionJob(uintptr_t Base, uintptr_t Size) {
#ifndef _WIN32
// Removing a named region through the job system
// We need to find the entry that we are deleting first
std::unique_ptr<CodeRegionEntry> EntryPointer;
fextl::unique_ptr<CodeRegionEntry> EntryPointer;
{
std::unique_lock lk {CodeObjectCacheService->GetEntryMapMutex()};
@@ -119,9 +120,10 @@ namespace FEXCore::CodeSerialize {
// Tell the async thread that it has work to do
CodeObjectCacheService->NotifyWork();
}
#endif
}
void AsyncJobHandler::AsyncAddSerializationJob(std::unique_ptr<SerializationJobData> Data) {
void AsyncJobHandler::AsyncAddSerializationJob(fextl::unique_ptr<SerializationJobData> Data) {
// XXX: Actually add serialization job
}
}
@@ -1,10 +1,12 @@
#include "Interface/Context/Context.h"
#include "Interface/Core/ObjectCache/ObjectCacheService.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/fextl/memory.h>
#include <FEXCore/fextl/string.h>
namespace FEXCore::CodeSerialize {
NamedRegionObjectHandler::NamedRegionObjectHandler(FEXCore::Context::Context *ctx) {
NamedRegionObjectHandler::NamedRegionObjectHandler(FEXCore::Context::ContextImpl *ctx) {
DefaultSerializationConfig.Cookie = CODE_COOKIE;
// Initialize the Arch from CPUID
@@ -23,14 +25,14 @@ namespace FEXCore::CodeSerialize {
DefaultSerializationConfig.x87ReducedPrecision = ctx->Config.x87ReducedPrecision;
}
void NamedRegionObjectHandler::AddNamedRegionObject(CodeRegionMapType::iterator Entry, const std::string &base_filename, const std::string &filename, bool Executable) {
void NamedRegionObjectHandler::AddNamedRegionObject(CodeRegionMapType::iterator Entry, const fextl::string &base_filename, const fextl::string &filename, bool Executable) {
// XXX: Add named region objects
// XXX: Until entry loading is complete just claim it is loaded
Entry->second->NamedJobRefCountMutex.unlock();
}
void NamedRegionObjectHandler::RemoveNamedRegionObject(uintptr_t Base, uintptr_t Size, std::unique_ptr<CodeRegionEntry> Entry) {
void NamedRegionObjectHandler::RemoveNamedRegionObject(uintptr_t Base, uintptr_t Size, fextl::unique_ptr<CodeRegionEntry> Entry) {
// XXX: Remove named region objects
// XXX: Until entry loading is complete just claim it is loaded
@@ -40,7 +42,7 @@ namespace FEXCore::CodeSerialize {
void NamedRegionObjectHandler::HandleNamedRegionObjectJobs() {
// Walk through all of our jobs sequentially until the work queue is empty
while (NamedWorkQueueJobs.load()) {
std::unique_ptr<AsyncJobHandler::NamedRegionWorkItem> WorkItem;
fextl::unique_ptr<AsyncJobHandler::NamedRegionWorkItem> WorkItem;
{
// Lock the work queue mutex for a short moment and grab an item from the list
@@ -1,8 +1,8 @@
#include "Interface/Core/ObjectCache/ObjectCacheService.h"
#include <FEXCore/Config/Config.h>
#include <memory>
#include <FEXCore/fextl/memory.h>
#include <FEXCore/Utils/Threads.h>
namespace {
static void* ThreadHandler(void *Arg) {
@@ -13,7 +13,7 @@ namespace {
}
namespace FEXCore::CodeSerialize {
CodeObjectSerializeService::CodeObjectSerializeService(FEXCore::Context::Context *ctx)
CodeObjectSerializeService::CodeObjectSerializeService(FEXCore::Context::ContextImpl *ctx)
: CTX {ctx}
, AsyncHandler { &NamedRegionHandler , this }
, NamedRegionHandler { ctx } {
@@ -38,13 +38,13 @@ namespace FEXCore::CodeSerialize {
void CodeObjectSerializeService::Initialize() {
// Add a canary so we don't crash on empty map iterator handling
auto it = AddressToEntryMap.insert_or_assign(~0ULL, std::make_unique<CodeRegionEntry>());
UnrelocatedAddressToEntryMap.insert_or_assign(~0ULL, it.first->second.get());
auto it = AddressToEntryMap.insert_or_assign(~0ULL, fextl::make_unique<CodeRegionEntry>());
UnrelocatedAddressToEntryMap.insert_or_assign(~0ULL, it.first->second.get());
uint64_t OldMask = FEXCore::Threads::SetSignalMask(~0ULL);
WorkerThread = FEXCore::Threads::Thread::Create(ThreadHandler, this);
uint64_t OldMask = FEXCore::Threads::SetSignalMask(~0ULL);
WorkerThread = FEXCore::Threads::Thread::Create(ThreadHandler, this);
FEXCore::Threads::SetSignalMask(OldMask);
}
}
void CodeObjectSerializeService::DoCodeRegionClosure(uint64_t Base, CodeRegionEntry *it) {
if (Base == ~0ULL) {
@@ -61,8 +61,7 @@ namespace FEXCore::CodeSerialize {
void CodeObjectSerializeService::ExecutionThread() {
// Set our thread name so we can see its relation
char ThreadName[16] = "ObjectCodeSeri\0";
pthread_setname_np(pthread_self(), ThreadName);
FEXCore::Threads::SetThreadName("ObjectCodeSeri\0");
while (WorkerThreadShuttingDown.load() != true) {
// Wait for work
WorkAvailable.Wait();
@@ -81,5 +80,5 @@ namespace FEXCore::CodeSerialize {
// Safely clear our maps now
AddressToEntryMap.clear();
UnrelocatedAddressToEntryMap.clear();
}
}
}
@@ -6,13 +6,14 @@
#include <FEXCore/Utils/Event.h>
#include <FEXCore/Utils/Threads.h>
#include <FEXCore/fextl/map.h>
#include <FEXCore/fextl/memory.h>
#include <FEXCore/fextl/queue.h>
#include <FEXCore/fextl/robin_map.h>
#include <FEXCore/fextl/string.h>
#include <FEXCore/fextl/vector.h>
#include <map>
#include <memory>
#include <shared_mutex>
#include <string>
#include <vector>
#include <tsl/robin_map.h>
namespace FEXCore::CodeSerialize {
// XXX: Does this need to be signal safe?
@@ -66,13 +67,13 @@ namespace FEXCore::CodeSerialize {
uint64_t Offset{};
// Filename of the object
std::string Filename{};
fextl::string Filename{};
CodeObjectSerializationHeader EntryHeader{};
/** @} */
// The filename of the object cache for this entry
std::string ObjectEntrySourceFilename{};
fextl::string ObjectEntrySourceFilename{};
// In the case of file corruption that we can detect, we can disable serialization early for an entry
// We should be resiliant to corruption but things happen
@@ -105,12 +106,12 @@ namespace FEXCore::CodeSerialize {
char *CodeData{};
size_t FileSize{};
std::vector<CodeObjectFileSection> FileCodeSections;
fextl::vector<CodeObjectFileSection> FileCodeSections;
/** @} */
// This per section map takes the most time to load and needs to be quick
// This is the map of all code segments for this entry
tsl::robin_map<uint64_t, CodeObjectFileSection*> SectionLookupMap{};
fextl::robin_map<uint64_t, CodeObjectFileSection*> SectionLookupMap{};
/** @} */
// Default initialization
@@ -120,7 +121,7 @@ namespace FEXCore::CodeSerialize {
CodeRegionEntry(uint64_t Base,
uint64_t Size,
uint64_t Offset,
std::string const &Filename,
fextl::string const &Filename,
CodeObjectSerializationHeader const &DefaultHeader)
: Base {Base}
, Size {Size}
@@ -131,8 +132,8 @@ namespace FEXCore::CodeSerialize {
};
// Map type must use an interator that isn't invalidation on erase/insert
using CodeRegionMapType = std::map<uint64_t, std::unique_ptr<CodeRegionEntry>>;
using CodeRegionPtrMapType = std::map<uint64_t, CodeRegionEntry*>;
using CodeRegionMapType = fextl::map<uint64_t, fextl::unique_ptr<CodeRegionEntry>>;
using CodeRegionPtrMapType = fextl::map<uint64_t, CodeRegionEntry*>;
class NamedRegionObjectHandler;
class CodeObjectSerializeService;
@@ -160,7 +161,7 @@ namespace FEXCore::CodeSerialize {
// These are the reolocations for this serialization job
// Relatively small number of entries most of the time
std::vector<FEXCore::CPU::Relocation> Relocations;
fextl::vector<FEXCore::CPU::Relocation> Relocations;
/**
* @name Objects filled in from the Code Object Serialization service when a job is added
@@ -186,9 +187,9 @@ namespace FEXCore::CodeSerialize {
/**
* @name Async job submission functions
* @{ */
void AsyncAddNamedRegionJob(uintptr_t Base, uintptr_t Size, uintptr_t Offset, const std::string &filename);
void AsyncAddNamedRegionJob(uintptr_t Base, uintptr_t Size, uintptr_t Offset, const fextl::string &filename);
void AsyncRemoveNamedRegionJob(uintptr_t Base, uintptr_t Size);
void AsyncAddSerializationJob(std::unique_ptr<SerializationJobData> Data);
void AsyncAddSerializationJob(fextl::unique_ptr<SerializationJobData> Data);
/** @} */
/**
@@ -219,22 +220,22 @@ namespace FEXCore::CodeSerialize {
class WorkItemAddNamedRegion : public NamedRegionWorkItem {
public:
WorkItemAddNamedRegion(const std::string &base, const std::string &filename, bool executable, CodeRegionMapType::iterator entry)
WorkItemAddNamedRegion(const fextl::string &base, const fextl::string &filename, bool executable, CodeRegionMapType::iterator entry)
: NamedRegionWorkItem {NamedRegionJobType::JOB_ADD_NAMED_REGION}
, BaseFilename {base}
, Filename {filename}
, Executable {executable}
, Entry {entry}
{}
const std::string BaseFilename;
const std::string Filename;
const fextl::string BaseFilename;
const fextl::string Filename;
bool Executable;
CodeRegionMapType::iterator Entry;
};
class WorkItemRemoveNamedRegion : public NamedRegionWorkItem {
public:
WorkItemRemoveNamedRegion(uint64_t base, uint64_t size, std::unique_ptr<CodeRegionEntry> entry)
WorkItemRemoveNamedRegion(uint64_t base, uint64_t size, fextl::unique_ptr<CodeRegionEntry> entry)
: NamedRegionWorkItem {NamedRegionJobType::JOB_REMOVE_NAMED_REGION}
, Base {base}
, Size {size}
@@ -242,7 +243,7 @@ namespace FEXCore::CodeSerialize {
uint64_t Base;
uint64_t Size;
std::unique_ptr<CodeRegionEntry> Entry;
fextl::unique_ptr<CodeRegionEntry> Entry;
};
/** @} */
@@ -253,7 +254,7 @@ namespace FEXCore::CodeSerialize {
class NamedRegionObjectHandler final {
public:
NamedRegionObjectHandler(FEXCore::Context::Context *ctx);
NamedRegionObjectHandler(FEXCore::Context::ContextImpl *ctx);
void HandleNamedRegionObjectJobs();
@@ -281,9 +282,9 @@ namespace FEXCore::CodeSerialize {
*
* This adds the job that will do the loading of file resources and data tracking.
*/
void AsyncAddNamedRegionWorkItem(const std::string &base, const std::string &filename, bool executable, CodeRegionMapType::iterator entry) {
void AsyncAddNamedRegionWorkItem(const fextl::string &base, const fextl::string &filename, bool executable, CodeRegionMapType::iterator entry) {
std::unique_lock lk {NamedWorkQueueMutex};
WorkQueue.emplace(std::make_unique<AsyncJobHandler::WorkItemAddNamedRegion> (
WorkQueue.emplace(fextl::make_unique<AsyncJobHandler::WorkItemAddNamedRegion> (
base,
filename,
executable,
@@ -292,9 +293,9 @@ namespace FEXCore::CodeSerialize {
++NamedWorkQueueJobs;
}
void AsyncRemoveNamedRegionWorkItem(uint64_t Base, uint64_t Size, std::unique_ptr<CodeRegionEntry> Entry) {
void AsyncRemoveNamedRegionWorkItem(uint64_t Base, uint64_t Size, fextl::unique_ptr<CodeRegionEntry> Entry) {
std::unique_lock lk {NamedWorkQueueMutex};
WorkQueue.emplace(std::make_unique<AsyncJobHandler::WorkItemRemoveNamedRegion> (
WorkQueue.emplace(fextl::make_unique<AsyncJobHandler::WorkItemRemoveNamedRegion> (
Base,
Size,
std::move(Entry)
@@ -321,13 +322,13 @@ namespace FEXCore::CodeSerialize {
// The job queue itself
// Jobs get consumed as a FIFO
// Jobs always get appended to the end
std::queue<std::unique_ptr<AsyncJobHandler::NamedRegionWorkItem>> WorkQueue{};
fextl::queue<fextl::unique_ptr<AsyncJobHandler::NamedRegionWorkItem>> WorkQueue{};
/**
* @name Named Region object handling
* @{ */
void AddNamedRegionObject(CodeRegionMapType::iterator Entry, const std::string &base_filename, const std::string &filename, bool Executable);
void RemoveNamedRegionObject(uintptr_t Base, uintptr_t Size, std::unique_ptr<CodeRegionEntry> Entry);
void AddNamedRegionObject(CodeRegionMapType::iterator Entry, const fextl::string &base_filename, const fextl::string &filename, bool Executable);
void RemoveNamedRegionObject(uintptr_t Base, uintptr_t Size, fextl::unique_ptr<CodeRegionEntry> Entry);
/** @} */
};
@@ -338,7 +339,7 @@ namespace FEXCore::CodeSerialize {
*/
class CodeObjectSerializeService final {
public:
CodeObjectSerializeService(FEXCore::Context::Context *ctx);
CodeObjectSerializeService(FEXCore::Context::ContextImpl *ctx);
/**
* @brief Initialize the internal interface
@@ -365,7 +366,7 @@ namespace FEXCore::CodeSerialize {
* @param Offset - The offset from the file
* @param filename - The filename itself
*/
void AsyncAddNamedRegionJob(uintptr_t Base, uintptr_t Size, uintptr_t Offset, const std::string &filename) {
void AsyncAddNamedRegionJob(uintptr_t Base, uintptr_t Size, uintptr_t Offset, const fextl::string &filename) {
AsyncHandler.AsyncAddNamedRegionJob(Base, Size, Offset, filename);
}
@@ -385,7 +386,7 @@ namespace FEXCore::CodeSerialize {
*
* @param Data - A fully filled out struct containing all the code serialization
*/
void AsyncAddSerializationJob(std::unique_ptr<AsyncJobHandler::SerializationJobData> Data) {
void AsyncAddSerializationJob(fextl::unique_ptr<AsyncJobHandler::SerializationJobData> Data) {
AsyncHandler.AsyncAddSerializationJob(std::move(Data));
}
/** @} */
@@ -440,10 +441,10 @@ namespace FEXCore::CodeSerialize {
void NotifyWork() { WorkAvailable.NotifyOne(); }
private:
FEXCore::Context::Context *CTX;
FEXCore::Context::ContextImpl *CTX;
Event WorkAvailable{};
std::unique_ptr<FEXCore::Threads::Thread> WorkerThread;
fextl::unique_ptr<FEXCore::Threads::Thread> WorkerThread;
std::atomic_bool WorkerThreadShuttingDown {false};
AsyncJobHandler AsyncHandler;
NamedRegionObjectHandler NamedRegionHandler;
+232 -306
View File
@@ -7,12 +7,12 @@ $end_info$
#include "Interface/Context/Context.h"
#include "Interface/Core/OpcodeDispatcher.h"
#include "Interface/Core/X86Tables/X86Tables.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/Context.h>
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Core/X86Enums.h>
#include <FEXCore/Debug/X86Tables.h>
#include <FEXCore/HLE/SyscallHandler.h>
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IREmitter.h>
@@ -35,6 +35,7 @@ void OpDispatchBuilder::SyscallOp(OpcodeArgs) {
constexpr size_t SyscallArgs = 7;
using SyscallArray = std::array<uint64_t, SyscallArgs>;
size_t NumArguments{};
const SyscallArray *GPRIndexes {};
static constexpr SyscallArray GPRIndexes_64 = {
FEXCore::X86State::REG_RAX,
@@ -54,13 +55,26 @@ void OpDispatchBuilder::SyscallOp(OpcodeArgs) {
FEXCore::X86State::REG_RDI,
FEXCore::X86State::REG_RBP,
};
static_assert(GPRIndexes_64.size() == GPRIndexes_32.size());
static std::array<uint64_t, SyscallArgs> GPRIndexes_Hangover = {
static constexpr SyscallArray GPRIndexes_Hangover = {
FEXCore::X86State::REG_RCX,
};
size_t NumArguments{};
static constexpr SyscallArray GPRIndexes_Win64 = {
FEXCore::X86State::REG_RAX,
FEXCore::X86State::REG_R10,
FEXCore::X86State::REG_RDX,
FEXCore::X86State::REG_R8,
FEXCore::X86State::REG_R9,
FEXCore::X86State::REG_RSP,
};
static constexpr SyscallArray GPRIndexes_Win32 = {
FEXCore::X86State::REG_RAX,
FEXCore::X86State::REG_RSP,
};
SyscallFlags DefaultSyscallFlags = FEXCore::IR::SyscallFlags::DEFAULT;
const auto OSABI = CTX->SyscallHandler->GetOSABI();
if (OSABI == FEXCore::HLE::SyscallOSABI::OS_LINUX64) {
@@ -68,9 +82,19 @@ void OpDispatchBuilder::SyscallOp(OpcodeArgs) {
GPRIndexes = &GPRIndexes_64;
}
else if (OSABI == FEXCore::HLE::SyscallOSABI::OS_LINUX32) {
NumArguments = GPRIndexes_64.size();
NumArguments = GPRIndexes_32.size();
GPRIndexes = &GPRIndexes_32;
}
else if (OSABI == FEXCore::HLE::SyscallOSABI::OS_WIN64) {
NumArguments = 6;
GPRIndexes = &GPRIndexes_Win64;
DefaultSyscallFlags = FEXCore::IR::SyscallFlags::NORETURNEDRESULT;
}
else if (OSABI == FEXCore::HLE::SyscallOSABI::OS_WIN32) {
NumArguments = 2;
GPRIndexes = &GPRIndexes_Win32;
DefaultSyscallFlags = FEXCore::IR::SyscallFlags::NORETURNEDRESULT;
}
else if (OSABI == FEXCore::HLE::SyscallOSABI::OS_HANGOVER) {
NumArguments = 1;
GPRIndexes = &GPRIndexes_Hangover;
@@ -109,13 +133,20 @@ void OpDispatchBuilder::SyscallOp(OpcodeArgs) {
Arguments[4],
Arguments[5],
Arguments[6],
FEXCore::IR::SyscallFlags::DEFAULT);
DefaultSyscallFlags);
if (OSABI != FEXCore::HLE::SyscallOSABI::OS_HANGOVER) {
if (OSABI != FEXCore::HLE::SyscallOSABI::OS_HANGOVER &&
(DefaultSyscallFlags & FEXCore::IR::SyscallFlags::NORETURNEDRESULT) != FEXCore::IR::SyscallFlags::NORETURNEDRESULT) {
// Hangover doesn't want us returning a result here
// syscall is being abused as a thunk for now.
StoreGPRRegister(X86State::REG_RAX, SyscallOp);
}
if (Op->TableInfo->Flags & X86Tables::InstFlags::FLAGS_BLOCK_END) {
// RIP could have been updated after coming back from the Syscall.
NewRIP = _LoadContext(GPRSize, GPRClass, offsetof(FEXCore::Core::CPUState, rip));
_ExitFunction(NewRIP);
}
}
void OpDispatchBuilder::ThunkOp(OpcodeArgs) {
@@ -277,18 +308,6 @@ void OpDispatchBuilder::IRETOp(OpcodeArgs) {
BlockSetRIP = true;
}
void OpDispatchBuilder::SIGRETOp(OpcodeArgs) {
uint8_t Literal = Op->Src[0].Data.Literal.Value;
const uint8_t GPRSize = CTX->GetGPRSize();
// Store the new RIP
bool IsRT = CTX->Config.Is64BitMode() || Literal;
_SignalReturn(IsRT);
auto NewRIP = _LoadContext(GPRSize, GPRClass, offsetof(FEXCore::Core::CPUState, rip));
// This ExitFunction won't actually get hit but needs to exist
_ExitFunction(NewRIP);
BlockSetRIP = true;
}
void OpDispatchBuilder::CallbackReturnOp(OpcodeArgs) {
const uint8_t GPRSize = CTX->GetGPRSize();
// Store the new RIP
@@ -445,7 +464,7 @@ void OpDispatchBuilder::ADCOp(OpcodeArgs) {
GenerateFlags_ADC(Op, Result, Before, Src, CF);
}
template<uint32_t SrcIndex>
template<uint32_t SrcIndex, bool SetFlags>
void OpDispatchBuilder::SBBOp(OpcodeArgs) {
// Calculate flags early.
CalculateDeferredFlags();
@@ -471,10 +490,12 @@ void OpDispatchBuilder::SBBOp(OpcodeArgs) {
StoreResult(GPRClass, Op, Result, -1);
}
if (Size < 4) {
Result = _Bfe(Size, Size * 8, 0, Result);
if (SetFlags) {
if (Size < 4) {
Result = _Bfe(Size, Size * 8, 0, Result);
}
GenerateFlags_SBB(Op, Result, Before, Src, CF);
}
GenerateFlags_SBB(Op, Result, Before, Src, CF);
}
void OpDispatchBuilder::PUSHOp(OpcodeArgs) {
@@ -876,7 +897,7 @@ OrderedNode *OpDispatchBuilder::SelectCC(uint8_t OP, OrderedNode *TrueValue, Ord
case 0x7: { // JA - Jump if CF == 0 && ZF == 0
auto Flag1 = GetRFLAG(FEXCore::X86State::RFLAG_ZF_LOC);
auto Flag2 = GetRFLAG(FEXCore::X86State::RFLAG_CF_LOC);
auto Check = _Or(Flag1, _Lshl(Flag2, _Constant(1)));
auto Check = _Or(Flag1, Flag2);
SrcCond = _Select(FEXCore::IR::COND_EQ,
Check, ZeroConst, TrueValue, FalseValue);
break;
@@ -1523,7 +1544,7 @@ void OpDispatchBuilder::SAHFOp(OpcodeArgs) {
OrderedNode *Src = LoadGPRRegister(X86State::REG_RAX, 1, 8);
// Clear bits that aren't supposed to be set
Src = _And(Src, _Constant(~0b101000));
Src = _Andn(Src, _Constant(0b101000));
// Set the bit that is always set here
Src = _Or(Src, _Constant(0b10));
@@ -1758,6 +1779,18 @@ void OpDispatchBuilder::CPUIDOp(OpcodeArgs) {
StoreGPRRegister(X86State::REG_RCX, _Bfe(32, 0, Result_Upper));
}
void OpDispatchBuilder::XGetBVOp(OpcodeArgs) {
OrderedNode *Function = LoadGPRRegister(X86State::REG_RCX);
auto Res = _XGetBV(Function);
OrderedNode *Result_Lower = _ExtractElementPair(Res, 0);
OrderedNode *Result_Upper = _ExtractElementPair(Res, 1);
StoreGPRRegister(X86State::REG_RAX, Result_Lower);
StoreGPRRegister(X86State::REG_RDX, Result_Upper);
}
template<bool SHL1Bit>
void OpDispatchBuilder::SHLOp(OpcodeArgs) {
OrderedNode *Src{};
@@ -3805,73 +3838,21 @@ void OpDispatchBuilder::STOSOp(OpcodeArgs) {
StoreGPRRegister(X86State::REG_RDI, TailDest);
}
else {
// Calculate deffered flags.
// This block is ending and it needs flag status
CalculateDeferredFlags();
// FEX doesn't support partial faulting REP instructions.
// Converting this to a `MemSet` IR op optimizes this quite significantly in our codegen.
// If FEX is to gain support for faulting REP instructions, then this implementation needs to change significantly.
OrderedNode *Src = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *Dest = LoadGPRRegister(X86State::REG_RDI);
// Create all our blocks
auto LoopHead = CreateNewCodeBlockAfter(GetCurrentBlock());
auto LoopTail = CreateNewCodeBlockAfter(LoopHead);
auto LoopEnd = CreateNewCodeBlockAfter(LoopTail);
// Only ES prefix
auto Segment = GetSegment(0, FEXCore::X86Tables::DecodeFlags::FLAG_ES_PREFIX, true);
// At the time this was written, our RA can't handle accessing nodes across blocks.
// So we need to re-load and re-calculate essential values each iteration of the loop.
// First thing we need to do is finish this block and jump to the start of the loop.
// RA can now better allocate things, move these ops before the header, to avoid accessing
// DF on every iteration
auto SizeConst = _Constant(Size);
auto NegSizeConst = _Constant(-Size);
// Calculate direction.
OrderedNode *Counter = LoadGPRRegister(X86State::REG_RCX);
auto DF = GetRFLAG(FEXCore::X86State::RFLAG_DF_LOC);
auto PtrDir = _Select(FEXCore::IR::COND_EQ,
DF, _Constant(0),
SizeConst, NegSizeConst);
_Jump(LoopHead);
SetCurrentCodeBlock(LoopHead);
{
OrderedNode *Counter = LoadGPRRegister(X86State::REG_RCX);
// Can we end the block?
_CondJump(Counter, LoopEnd, LoopTail, {COND_EQ});
}
SetCurrentCodeBlock(LoopTail);
{
OrderedNode *Src = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *Dest = LoadGPRRegister(X86State::REG_RDI);
// Only ES prefix
Dest = AppendSegmentOffset(Dest, 0, FEXCore::X86Tables::DecodeFlags::FLAG_ES_PREFIX, true);
// Store to memory where RDI points
_StoreMemAutoTSO(GPRClass, Size, Dest, Src, Size);
OrderedNode *TailCounter = LoadGPRRegister(X86State::REG_RCX);
OrderedNode *TailDest = LoadGPRRegister(X86State::REG_RDI);
// Decrement counter
TailCounter = _Sub(TailCounter, _Constant(1));
// Store the counter so we don't have to deal with PHI here
StoreGPRRegister(X86State::REG_RCX, TailCounter);
// Offset the pointer
TailDest = _Add(TailDest, PtrDir);
StoreGPRRegister(X86State::REG_RDI, TailDest);
// Jump back to the start, we have more work to do
_Jump(LoopHead);
}
// Make sure to start a new block after ending this one
SetCurrentCodeBlock(LoopEnd);
auto Result = _MemSet(CTX->IsAtomicTSOEnabled(), Size, Segment ?: InvalidNode, Dest, Src, Counter, DF);
StoreGPRRegister(X86State::REG_RCX, _Constant(0));
StoreGPRRegister(X86State::REG_RDI, Result);
}
}
@@ -3884,75 +3865,35 @@ void OpDispatchBuilder::MOVSOp(OpcodeArgs) {
// RA now can handle these to be here, to avoid DF accesses
const auto Size = GetSrcSize(Op);
auto SizeConst = _Constant(Size);
auto NegSizeConst = _Constant(-Size);
// Calculate direction.
auto DF = GetRFLAG(FEXCore::X86State::RFLAG_DF_LOC);
auto PtrDir = _Select(FEXCore::IR::COND_EQ, DF, _Constant(0), SizeConst, NegSizeConst);
if (Op->Flags & (FEXCore::X86Tables::DecodeFlags::FLAG_REP_PREFIX | FEXCore::X86Tables::DecodeFlags::FLAG_REPNE_PREFIX)) {
// Calculate flags early. because end of block
CalculateDeferredFlags();
auto SrcAddr = LoadGPRRegister(X86State::REG_RSI);
auto DstAddr = LoadGPRRegister(X86State::REG_RDI);
auto Counter = LoadGPRRegister(X86State::REG_RCX);
// Create all our blocks
auto LoopHead = CreateNewCodeBlockAfter(GetCurrentBlock());
auto LoopTail = CreateNewCodeBlockAfter(LoopHead);
auto LoopEnd = CreateNewCodeBlockAfter(LoopTail);
auto DstSegment = GetSegment(0, FEXCore::X86Tables::DecodeFlags::FLAG_ES_PREFIX, true);
auto SrcSegment = GetSegment(Op->Flags, FEXCore::X86Tables::DecodeFlags::FLAG_DS_PREFIX);
auto Result = _MemCpy(CTX->IsAtomicTSOEnabled(), Size,
DstSegment ?: InvalidNode,
SrcSegment ?: InvalidNode,
DstAddr, SrcAddr, Counter, DF);
// At the time this was written, our RA can't handle accessing nodes across blocks.
// So we need to re-load and re-calculate essential values each iteration of the loop.
OrderedNode *Result_Dst = _ExtractElementPair(Result, 0);
OrderedNode *Result_Src = _ExtractElementPair(Result, 1);
// First thing we need to do is finish this block and jump to the start of the loop.
_Jump(LoopHead);
SetCurrentCodeBlock(LoopHead);
{
OrderedNode *Counter = LoadGPRRegister(X86State::REG_RCX);
_CondJump(Counter, LoopEnd, LoopTail, {COND_EQ});
}
SetCurrentCodeBlock(LoopTail);
{
OrderedNode *Src = LoadGPRRegister(X86State::REG_RSI);
OrderedNode *Dest = LoadGPRRegister(X86State::REG_RDI);
Dest = AppendSegmentOffset(Dest, 0, FEXCore::X86Tables::DecodeFlags::FLAG_ES_PREFIX, true);
Src = AppendSegmentOffset(Src, Op->Flags, FEXCore::X86Tables::DecodeFlags::FLAG_DS_PREFIX);
Src = _LoadMemAutoTSO(GPRClass, Size, Src, Size);
// Store to memory where RDI points
_StoreMemAutoTSO(GPRClass, Size, Dest, Src, Size);
OrderedNode *TailCounter = LoadGPRRegister(X86State::REG_RCX);
// Decrement counter
TailCounter = _Sub(TailCounter, _Constant(1));
// Store the counter so we don't have to deal with PHI here
StoreGPRRegister(X86State::REG_RCX, TailCounter);
// Offset the pointer
OrderedNode *TailSrc = LoadGPRRegister(X86State::REG_RSI);
OrderedNode *TailDest = LoadGPRRegister(X86State::REG_RDI);
TailSrc = _Add(TailSrc, PtrDir);
TailDest = _Add(TailDest, PtrDir);
StoreGPRRegister(X86State::REG_RSI, TailSrc);
StoreGPRRegister(X86State::REG_RDI, TailDest);
// Jump back to the start, we have more work to do
_Jump(LoopHead);
}
// Make sure to start a new block after ending this one
SetCurrentCodeBlock(LoopEnd);
StoreGPRRegister(X86State::REG_RCX, _Constant(0));
StoreGPRRegister(X86State::REG_RDI, Result_Dst);
StoreGPRRegister(X86State::REG_RSI, Result_Src);
}
else {
auto SizeConst = _Constant(Size);
auto NegSizeConst = _Constant(-Size);
auto PtrDir = _Select(FEXCore::IR::COND_EQ, DF, _Constant(0), SizeConst, NegSizeConst);
OrderedNode *RSI = LoadGPRRegister(X86State::REG_RSI);
OrderedNode *RDI = LoadGPRRegister(X86State::REG_RDI);
RDI= AppendSegmentOffset(RDI, 0, FEXCore::X86Tables::DecodeFlags::FLAG_ES_PREFIX, true);
@@ -4748,7 +4689,7 @@ void OpDispatchBuilder::CMPXCHGPairOp(OpcodeArgs) {
SetCurrentCodeBlock(NextJumpTarget);
}
void OpDispatchBuilder::CreateJumpBlocks(std::vector<FEXCore::Frontend::Decoder::DecodedBlocks> const *Blocks) {
void OpDispatchBuilder::CreateJumpBlocks(fextl::vector<FEXCore::Frontend::Decoder::DecodedBlocks> const *Blocks) {
OrderedNode *PrevCodeBlock{};
for (auto &Target : *Blocks) {
auto CodeNode = CreateCodeNode();
@@ -4763,7 +4704,7 @@ void OpDispatchBuilder::CreateJumpBlocks(std::vector<FEXCore::Frontend::Decoder:
}
}
void OpDispatchBuilder::BeginFunction(uint64_t RIP, std::vector<FEXCore::Frontend::Decoder::DecodedBlocks> const *Blocks) {
void OpDispatchBuilder::BeginFunction(uint64_t RIP, fextl::vector<FEXCore::Frontend::Decoder::DecodedBlocks> const *Blocks) {
Entry = RIP;
auto IRHeader = _IRHeader(InvalidNode, 0);
Current_Header = IRHeader.first;
@@ -4844,20 +4785,19 @@ uint32_t OpDispatchBuilder::GetDstBitSize(X86Tables::DecodedOp Op) const {
return GetDstSize(Op) * 8;
}
OrderedNode *OpDispatchBuilder::AppendSegmentOffset(OrderedNode *Value, uint32_t Flags, uint32_t DefaultPrefix, bool Override) {
OrderedNode *OpDispatchBuilder::GetSegment(uint32_t Flags, uint32_t DefaultPrefix, bool Override) {
const uint8_t GPRSize = CTX->GetGPRSize();
if (CTX->Config.Is64BitMode) {
if (Flags & FEXCore::X86Tables::DecodeFlags::FLAG_FS_PREFIX) {
Value = _Add(Value, _LoadContext(GPRSize, GPRClass, offsetof(FEXCore::Core::CPUState, fs_cached)));
return _LoadContext(GPRSize, GPRClass, offsetof(FEXCore::Core::CPUState, fs_cached));
}
else if (Flags & FEXCore::X86Tables::DecodeFlags::FLAG_GS_PREFIX) {
Value = _Add(Value, _LoadContext(GPRSize, GPRClass, offsetof(FEXCore::Core::CPUState, gs_cached)));
return _LoadContext(GPRSize, GPRClass, offsetof(FEXCore::Core::CPUState, gs_cached));
}
// If there was any other segment in 64bit then it is ignored
}
else {
OrderedNode *Segment{};
uint32_t Prefix = Flags & FEXCore::X86Tables::DecodeFlags::FLAG_SEGMENTS;
if (!Prefix || Override) {
// If there was no prefix then use the default one if available
@@ -4867,29 +4807,28 @@ OrderedNode *OpDispatchBuilder::AppendSegmentOffset(OrderedNode *Value, uint32_t
// With the segment register optimization we store the GDT bases directly in the segment register to remove indexed loads
switch (Prefix) {
case FEXCore::X86Tables::DecodeFlags::FLAG_ES_PREFIX:
Segment = _LoadContext(GPRSize, GPRClass, offsetof(FEXCore::Core::CPUState, es_cached));
break;
return _LoadContext(GPRSize, GPRClass, offsetof(FEXCore::Core::CPUState, es_cached));
case FEXCore::X86Tables::DecodeFlags::FLAG_CS_PREFIX:
Segment = _LoadContext(GPRSize, GPRClass, offsetof(FEXCore::Core::CPUState, cs_cached));
break;
return _LoadContext(GPRSize, GPRClass, offsetof(FEXCore::Core::CPUState, cs_cached));
case FEXCore::X86Tables::DecodeFlags::FLAG_SS_PREFIX:
Segment = _LoadContext(GPRSize, GPRClass, offsetof(FEXCore::Core::CPUState, ss_cached));
break;
return _LoadContext(GPRSize, GPRClass, offsetof(FEXCore::Core::CPUState, ss_cached));
case FEXCore::X86Tables::DecodeFlags::FLAG_DS_PREFIX:
Segment = _LoadContext(GPRSize, GPRClass, offsetof(FEXCore::Core::CPUState, ds_cached));
break;
return _LoadContext(GPRSize, GPRClass, offsetof(FEXCore::Core::CPUState, ds_cached));
case FEXCore::X86Tables::DecodeFlags::FLAG_FS_PREFIX:
Segment = _LoadContext(GPRSize, GPRClass, offsetof(FEXCore::Core::CPUState, fs_cached));
break;
return _LoadContext(GPRSize, GPRClass, offsetof(FEXCore::Core::CPUState, fs_cached));
case FEXCore::X86Tables::DecodeFlags::FLAG_GS_PREFIX:
Segment = _LoadContext(GPRSize, GPRClass, offsetof(FEXCore::Core::CPUState, gs_cached));
break;
default: break; // Do nothing
return _LoadContext(GPRSize, GPRClass, offsetof(FEXCore::Core::CPUState, gs_cached));
default:
break; // Do nothing
}
}
return nullptr;
}
if (Segment) {
Value = _Add(Value, Segment);
}
OrderedNode *OpDispatchBuilder::AppendSegmentOffset(OrderedNode *Value, uint32_t Flags, uint32_t DefaultPrefix, bool Override) {
auto Segment = GetSegment(Flags, DefaultPrefix, Override);
if (Segment) {
Value = _Add(Value, Segment);
}
return Value;
@@ -4957,21 +4896,9 @@ OrderedNode *OpDispatchBuilder::LoadSource_WithOpSize(FEXCore::IR::RegisterClass
else if (gpr >= FEXCore::X86State::REG_XMM_0) {
const auto gprIndex = gpr - X86State::REG_XMM_0;
const auto regSize = CTX->HostFeatures.SupportsAVX ?
Core::CPUState::XMM_AVX_REG_SIZE :
Core::CPUState::XMM_SSE_REG_SIZE;
// Load the full register size if it is a XMM register source.
Src = LoadXMMRegister(gprIndex);
// If we are wanting a high-index then we need to extract an element from the upper half of the reg.
// We can only extract an element size here.
// TODO: Have the instruction doing this load do the extract instead of here.
// We don't have enough information here to know if we can avoid this dup.
if (highIndex && OpSize < Core::CPUState::XMM_SSE_REG_SIZE) {
Src = _VDupElement(regSize, OpSize, Src, 1);
}
// Now extract the subregister if it was a partial load /smaller/ than SSE size
// TODO: Instead of doing the VMov implicitly on load, hunt down all use cases that require partial loads and do it after load.
// We don't have information here to know if the operation needs zero upper bits or can contain data.
@@ -5173,28 +5100,23 @@ void OpDispatchBuilder::StoreResult_WithOpSize(FEXCore::IR::RegisterClassType Cl
}
else if (gpr >= FEXCore::X86State::REG_XMM_0) {
const auto gprIndex = gpr - X86State::REG_XMM_0;
const auto highIndex = Operand.Data.GPR.HighBits ? 1 : 0;
const auto VectorSize = CTX->HostFeatures.SupportsAVX ? 32 : 16;
auto Result = Src;
if (highIndex || OpSize != VectorSize) {
// Partial writes can come from GPR or FPR.
if (OpSize != VectorSize) {
// Partial writes can come from FPRs.
// TODO: Fix the instructions doing partial writes rather than dealing with it here.
auto SrcVector = LoadXMMRegister(gprIndex);
if (Class == IR::GPRClass) {
Result = _VInsGPR(VectorSize, OpSize, highIndex, SrcVector, Src);
}
else {
// OpSize of 16 is special in that it is expected to zero the upper bits of the 256-bit operation.
// TODO: Longer term we should enforce the difference between zero and insert.
if (VectorSize == Core::CPUState::XMM_AVX_REG_SIZE && OpSize == Core::CPUState::XMM_SSE_REG_SIZE) {
Result = _VMov(OpSize, Src);
}
else {
Result = _VInsElement(VectorSize, OpSize, highIndex, 0, SrcVector, Src);
}
LOGMAN_THROW_A_FMT(Class != IR::GPRClass, "Partial writes from GPR not allowed. Instruction: {}",
Op->TableInfo->Name);
// OpSize of 16 is special in that it is expected to zero the upper bits of the 256-bit operation.
// TODO: Longer term we should enforce the difference between zero and insert.
if (VectorSize == Core::CPUState::XMM_AVX_REG_SIZE && OpSize == Core::CPUState::XMM_SSE_REG_SIZE) {
Result = _VMov(OpSize, Src);
} else {
Result = _VInsElement(VectorSize, OpSize, 0, 0, SrcVector, Src);
}
}
@@ -5330,7 +5252,7 @@ void OpDispatchBuilder::StoreResult(FEXCore::IR::RegisterClassType Class, FEXCor
StoreResult(Class, Op, Op->Dest, Src, Align, AccessType);
}
OpDispatchBuilder::OpDispatchBuilder(FEXCore::Context::Context *ctx)
OpDispatchBuilder::OpDispatchBuilder(FEXCore::Context::ContextImpl *ctx)
: IREmitter {ctx->OpDispatcherAllocator}
, CTX {ctx} {
ResetWorkingList();
@@ -5365,58 +5287,7 @@ void OpDispatchBuilder::MOVGPRNTOp(OpcodeArgs) {
StoreResult(GPRClass, Op, Src, 1, MemoryAccessType::ACCESS_STREAM);
}
void OpDispatchBuilder::ALUOp(OpcodeArgs) {
bool RequiresMask = false;
FEXCore::IR::IROps IROp;
switch (Op->OP) {
case 0x0:
case 0x1:
case 0x2:
case 0x3:
case 0x4:
case 0x5:
IROp = FEXCore::IR::IROps::OP_ADD;
RequiresMask = true;
break;
case 0x8:
case 0x9:
case 0xA:
case 0xB:
case 0xC:
case 0xD:
IROp = FEXCore::IR::IROps::OP_OR;
break;
case 0x20:
case 0x21:
case 0x22:
case 0x23:
case 0x24:
case 0x25:
IROp = FEXCore::IR::IROps::OP_AND;
break;
case 0x28:
case 0x29:
case 0x2A:
case 0x2B:
case 0x2C:
case 0x2D:
IROp = FEXCore::IR::IROps::OP_SUB;
RequiresMask = true;
break;
case 0x30:
case 0x31:
case 0x32:
case 0x33:
case 0x34:
case 0x35:
IROp = FEXCore::IR::IROps::OP_XOR;
break;
default:
IROp = FEXCore::IR::IROps::OP_LAST;
LOGMAN_MSG_A_FMT("Unknown ALU Op: 0x{:x}", Op->OP);
break;
}
void OpDispatchBuilder::ALUOpImpl(OpcodeArgs, FEXCore::IR::IROps ALUIROp, FEXCore::IR::IROps AtomicFetchOp, bool RequiresMask) {
auto Size = GetDstSize(Op);
// X86 basic ALU ops just do the operation between the destination and a single source
@@ -5429,43 +5300,24 @@ void OpDispatchBuilder::ALUOp(OpcodeArgs) {
HandledLock = true;
OrderedNode *DestMem = LoadSource(GPRClass, Op, Op->Dest, Op->Flags, -1, false);
DestMem = AppendSegmentOffset(DestMem, Op->Flags);
switch (IROp) {
case FEXCore::IR::IROps::OP_ADD: {
Dest = _AtomicFetchAdd(Size, Src, DestMem);
Result = _Add(Dest, Src);
break;
}
case FEXCore::IR::IROps::OP_SUB: {
Dest = _AtomicFetchSub(Size, Src, DestMem);
Result = _Sub(Dest, Src);
break;
}
case FEXCore::IR::IROps::OP_OR: {
Dest = _AtomicFetchOr(Size, Src, DestMem);
Result = _Or(Dest, Src);
break;
}
case FEXCore::IR::IROps::OP_AND: {
Dest = _AtomicFetchAnd(Size, Src, DestMem);
Result = _And(Dest, Src);
break;
}
case FEXCore::IR::IROps::OP_XOR: {
Dest = _AtomicFetchXor(Size, Src, DestMem);
Result = _Xor(Dest, Src);
break;
}
default:
LOGMAN_MSG_A_FMT("Unknown Atomic IR Op: {}", ToUnderlying(IROp));
break;
}
auto FetchOp = _AtomicFetchAdd(Size, Src, DestMem);
// Overwrite our atomic op type
FetchOp.first->Header.Op = AtomicFetchOp;
Dest = FetchOp;
auto ALUOp = _Add(Dest, Src);
// Overwrite our IR's op type
ALUOp.first->Header.Op = ALUIROp;
Result = ALUOp;
}
else {
Dest = LoadSource(GPRClass, Op, Op->Dest, Op->Flags, -1);
auto ALUOp = _Add(Dest, Src);
// Overwrite our IR's op type
ALUOp.first->Header.Op = IROp;
ALUOp.first->Header.Op = ALUIROp;
Result = ALUOp;
StoreResult(GPRClass, Op, Result, -1);
@@ -5477,7 +5329,7 @@ void OpDispatchBuilder::ALUOp(OpcodeArgs) {
// Flags set
{
switch (IROp) {
switch (ALUIROp) {
case FEXCore::IR::IROps::OP_ADD:
GenerateFlags_ADD(Op, Result, Dest, Src);
break;
@@ -5495,6 +5347,11 @@ void OpDispatchBuilder::ALUOp(OpcodeArgs) {
}
}
template<FEXCore::IR::IROps ALUIROp, FEXCore::IR::IROps AtomicFetchOp, bool RequiresMask>
void OpDispatchBuilder::ALUOp(OpcodeArgs) {
ALUOpImpl(Op, ALUIROp, AtomicFetchOp, RequiresMask);
}
void OpDispatchBuilder::INTOp(OpcodeArgs) {
IR::BreakDefinition Reason;
bool SetRIPToNext = false;
@@ -5503,14 +5360,19 @@ void OpDispatchBuilder::INTOp(OpcodeArgs) {
case 0xCD: { // INT imm8
uint8_t Literal = Op->Src[0].Data.Literal.Value;
if (Literal == 0x80) {
#ifndef _WIN32
constexpr uint8_t SYSCALL_LITERAL = 0x80;
#else
constexpr uint8_t SYSCALL_LITERAL = 0x2E;
#endif
if (Literal == SYSCALL_LITERAL) {
// Syscall on linux
SyscallOp(Op);
return;
}
Reason.ErrorRegister = Literal << 3 | (0b010);
Reason.Signal = SIGSEGV;
Reason.Signal = Core::FAULT_SIGSEGV;
// GP is raised when task-gate isn't setup to be valid
Reason.TrapNumber = X86State::X86_TRAPNO_GP;
Reason.si_code = 0x80;
@@ -5518,33 +5380,33 @@ void OpDispatchBuilder::INTOp(OpcodeArgs) {
}
case 0xCE: // INTO
Reason.ErrorRegister = 0;
Reason.Signal = SIGSEGV;
Reason.Signal = Core::FAULT_SIGSEGV;
Reason.TrapNumber = X86State::X86_TRAPNO_OF;
Reason.si_code = 0x80;
break;
case 0xF1: // INT1
Reason.ErrorRegister = 0;
Reason.Signal = SIGTRAP;
Reason.Signal = Core::FAULT_SIGTRAP;
Reason.TrapNumber = X86State::X86_TRAPNO_DB;
Reason.si_code = 1;
SetRIPToNext = true;
break;
case 0xF4: { // HLT
Reason.ErrorRegister = 0;
Reason.Signal = SIGSEGV;
Reason.Signal = Core::FAULT_SIGSEGV;
Reason.TrapNumber = X86State::X86_TRAPNO_GP;
Reason.si_code = 0x80;
break;
}
case 0x0B: // UD2
Reason.ErrorRegister = 0;
Reason.Signal = SIGILL;
Reason.Signal = Core::FAULT_SIGILL;
Reason.TrapNumber = X86State::X86_TRAPNO_UD;
Reason.si_code = 2;
break;
case 0xCC: // INT3
Reason.ErrorRegister = 0;
Reason.Signal = SIGTRAP;
Reason.Signal = Core::FAULT_SIGTRAP;
Reason.TrapNumber = X86State::X86_TRAPNO_BP;
Reason.si_code = 0x80;
SetRIPToNext = true;
@@ -5826,8 +5688,12 @@ void OpDispatchBuilder::InstallHostSpecificOpcodeHandlers() {
static constexpr std::tuple<uint16_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> AVXTable[] = {
{OPD(1, 0b00, 0x10), 1, &OpDispatchBuilder::VMOVUPS_VMOVUPD_Op},
{OPD(1, 0b01, 0x10), 1, &OpDispatchBuilder::VMOVUPS_VMOVUPD_Op},
{OPD(1, 0b10, 0x10), 1, &OpDispatchBuilder::VMOVSSOp},
{OPD(1, 0b11, 0x10), 1, &OpDispatchBuilder::VMOVSDOp},
{OPD(1, 0b00, 0x11), 1, &OpDispatchBuilder::VMOVUPS_VMOVUPD_Op},
{OPD(1, 0b01, 0x11), 1, &OpDispatchBuilder::VMOVUPS_VMOVUPD_Op},
{OPD(1, 0b10, 0x11), 1, &OpDispatchBuilder::VMOVSSOp},
{OPD(1, 0b11, 0x11), 1, &OpDispatchBuilder::VMOVSDOp},
{OPD(1, 0b00, 0x12), 1, &OpDispatchBuilder::VMOVLPOp},
{OPD(1, 0b01, 0x12), 1, &OpDispatchBuilder::VMOVLPOp},
@@ -5853,6 +5719,9 @@ void OpDispatchBuilder::InstallHostSpecificOpcodeHandlers() {
{OPD(1, 0b00, 0x29), 1, &OpDispatchBuilder::VMOVAPS_VMOVAPD_Op},
{OPD(1, 0b01, 0x29), 1, &OpDispatchBuilder::VMOVAPS_VMOVAPD_Op},
{OPD(1, 0b10, 0x2A), 1, &OpDispatchBuilder::AVXCVTGPR_To_FPR<4>},
{OPD(1, 0b11, 0x2A), 1, &OpDispatchBuilder::AVXCVTGPR_To_FPR<8>},
{OPD(1, 0b00, 0x2B), 1, &OpDispatchBuilder::VMOVVectorNTOp},
{OPD(1, 0b01, 0x2B), 1, &OpDispatchBuilder::VMOVVectorNTOp},
@@ -5903,6 +5772,11 @@ void OpDispatchBuilder::InstallHostSpecificOpcodeHandlers() {
{OPD(1, 0b10, 0x59), 1, &OpDispatchBuilder::AVXVectorScalarALUOp<IR::OP_VFMUL, 4>},
{OPD(1, 0b11, 0x59), 1, &OpDispatchBuilder::AVXVectorScalarALUOp<IR::OP_VFMUL, 8>},
{OPD(1, 0b00, 0x5A), 1, &OpDispatchBuilder::Vector_CVT_Float_To_Float<8, 4, true>},
{OPD(1, 0b01, 0x5A), 1, &OpDispatchBuilder::Vector_CVT_Float_To_Float<4, 8, true>},
{OPD(1, 0b10, 0x5A), 1, &OpDispatchBuilder::AVXScalar_CVT_Float_To_Float<8, 4>},
{OPD(1, 0b11, 0x5A), 1, &OpDispatchBuilder::AVXScalar_CVT_Float_To_Float<4, 8>},
{OPD(1, 0b00, 0x5B), 1, &OpDispatchBuilder::AVXVector_CVT_Int_To_Float<4, false>},
{OPD(1, 0b01, 0x5B), 1, &OpDispatchBuilder::AVXVector_CVT_Float_To_Int<4, false, true>},
{OPD(1, 0b10, 0x5B), 1, &OpDispatchBuilder::AVXVector_CVT_Float_To_Int<4, false, false>},
@@ -5946,6 +5820,10 @@ void OpDispatchBuilder::InstallHostSpecificOpcodeHandlers() {
{OPD(1, 0b01, 0x6F), 1, &OpDispatchBuilder::VMOVAPS_VMOVAPD_Op},
{OPD(1, 0b10, 0x6F), 1, &OpDispatchBuilder::VMOVUPS_VMOVUPD_Op},
{OPD(1, 0b01, 0x70), 1, &OpDispatchBuilder::VPSHUFWOp<4, true>},
{OPD(1, 0b10, 0x70), 1, &OpDispatchBuilder::VPSHUFWOp<2, false>},
{OPD(1, 0b11, 0x70), 1, &OpDispatchBuilder::VPSHUFWOp<2, true>},
{OPD(1, 0b01, 0x74), 1, &OpDispatchBuilder::AVXVectorALUOp<IR::OP_VCMPEQ, 1>},
{OPD(1, 0b01, 0x75), 1, &OpDispatchBuilder::AVXVectorALUOp<IR::OP_VCMPEQ, 2>},
{OPD(1, 0b01, 0x76), 1, &OpDispatchBuilder::AVXVectorALUOp<IR::OP_VCMPEQ, 4>},
@@ -5954,6 +5832,8 @@ void OpDispatchBuilder::InstallHostSpecificOpcodeHandlers() {
{OPD(1, 0b01, 0x7C), 1, &OpDispatchBuilder::VHADDPOp<IR::OP_VFADDP, 8>},
{OPD(1, 0b11, 0x7C), 1, &OpDispatchBuilder::VHADDPOp<IR::OP_VFADDP, 4>},
{OPD(1, 0b01, 0x7D), 1, &OpDispatchBuilder::VHSUBPOp<8>},
{OPD(1, 0b11, 0x7D), 1, &OpDispatchBuilder::VHSUBPOp<4>},
{OPD(1, 0b01, 0x7E), 1, &OpDispatchBuilder::MOVBetweenGPR_FPR},
{OPD(1, 0b10, 0x7E), 1, &OpDispatchBuilder::MOVQOp},
@@ -5966,8 +5846,12 @@ void OpDispatchBuilder::InstallHostSpecificOpcodeHandlers() {
{OPD(1, 0b10, 0xC2), 1, &OpDispatchBuilder::AVXVFCMPOp<4, true>},
{OPD(1, 0b11, 0xC2), 1, &OpDispatchBuilder::AVXVFCMPOp<8, true>},
{OPD(1, 0b01, 0xC4), 1, &OpDispatchBuilder::VPINSRWOp},
{OPD(1, 0b01, 0xC5), 1, &OpDispatchBuilder::PExtrOp<2>},
{OPD(1, 0b00, 0xC6), 1, &OpDispatchBuilder::VSHUFOp<4>},
{OPD(1, 0b01, 0xC6), 1, &OpDispatchBuilder::VSHUFOp<8>},
{OPD(1, 0b01, 0xD0), 1, &OpDispatchBuilder::VADDSUBPOp<8>},
{OPD(1, 0b11, 0xD0), 1, &OpDispatchBuilder::VADDSUBPOp<4>},
@@ -5977,7 +5861,7 @@ void OpDispatchBuilder::InstallHostSpecificOpcodeHandlers() {
{OPD(1, 0b01, 0xD4), 1, &OpDispatchBuilder::AVXVectorALUOp<IR::OP_VADD, 8>},
{OPD(1, 0b01, 0xD5), 1, &OpDispatchBuilder::AVXVectorALUOp<IR::OP_VSMUL, 2>},
{OPD(1, 0b01, 0xD6), 1, &OpDispatchBuilder::MOVQOp},
{OPD(1, 0b01, 0xD7), 1, &OpDispatchBuilder::UnimplementedOp},
{OPD(1, 0b01, 0xD7), 1, &OpDispatchBuilder::MOVMSKOpOne},
{OPD(1, 0b01, 0xD8), 1, &OpDispatchBuilder::AVXVectorALUOp<IR::OP_VUQSUB, 1>},
{OPD(1, 0b01, 0xD9), 1, &OpDispatchBuilder::AVXVectorALUOp<IR::OP_VUQSUB, 2>},
@@ -6015,6 +5899,8 @@ void OpDispatchBuilder::InstallHostSpecificOpcodeHandlers() {
{OPD(1, 0b01, 0xF2), 1, &OpDispatchBuilder::VPSLLOp<4>},
{OPD(1, 0b01, 0xF3), 1, &OpDispatchBuilder::VPSLLOp<8>},
{OPD(1, 0b01, 0xF4), 1, &OpDispatchBuilder::VPMULLOp<4, false>},
{OPD(1, 0b01, 0xF5), 1, &OpDispatchBuilder::VPMADDWDOp},
{OPD(1, 0b01, 0xF6), 1, &OpDispatchBuilder::VPSADBWOp},
{OPD(1, 0b01, 0xF7), 1, &OpDispatchBuilder::MASKMOVOp},
{OPD(1, 0b01, 0xF8), 1, &OpDispatchBuilder::AVXVectorALUOp<IR::OP_VSUB, 1>},
@@ -6025,17 +5911,27 @@ void OpDispatchBuilder::InstallHostSpecificOpcodeHandlers() {
{OPD(1, 0b01, 0xFD), 1, &OpDispatchBuilder::AVXVectorALUOp<IR::OP_VADD, 2>},
{OPD(1, 0b01, 0xFE), 1, &OpDispatchBuilder::AVXVectorALUOp<IR::OP_VADD, 4>},
{OPD(2, 0b01, 0x00), 1, &OpDispatchBuilder::VPSHUFBOp},
{OPD(2, 0b01, 0x01), 1, &OpDispatchBuilder::VHADDPOp<IR::OP_VADDP, 2>},
{OPD(2, 0b01, 0x02), 1, &OpDispatchBuilder::VHADDPOp<IR::OP_VADDP, 4>},
{OPD(2, 0b01, 0x03), 1, &OpDispatchBuilder::VPHADDSWOp},
{OPD(2, 0b01, 0x04), 1, &OpDispatchBuilder::VPMADDUBSWOp},
{OPD(2, 0b01, 0x05), 1, &OpDispatchBuilder::VPHSUBOp<2>},
{OPD(2, 0b01, 0x06), 1, &OpDispatchBuilder::VPHSUBOp<4>},
{OPD(2, 0b01, 0x07), 1, &OpDispatchBuilder::VPHSUBSWOp},
{OPD(2, 0b01, 0x08), 1, &OpDispatchBuilder::VPSIGN<1>},
{OPD(2, 0b01, 0x09), 1, &OpDispatchBuilder::VPSIGN<2>},
{OPD(2, 0b01, 0x0A), 1, &OpDispatchBuilder::VPSIGN<4>},
{OPD(2, 0b01, 0x0B), 1, &OpDispatchBuilder::VPMULHRSWOp},
{OPD(2, 0b01, 0x0C), 1, &OpDispatchBuilder::VPERMILRegOp<4>},
{OPD(2, 0b01, 0x0D), 1, &OpDispatchBuilder::VPERMILRegOp<8>},
{OPD(2, 0b01, 0x0E), 1, &OpDispatchBuilder::VTESTPOp<4>},
{OPD(2, 0b01, 0x0F), 1, &OpDispatchBuilder::VTESTPOp<8>},
{OPD(2, 0b01, 0x16), 1, &OpDispatchBuilder::VPERMDOp},
{OPD(2, 0b01, 0x17), 1, &OpDispatchBuilder::PTestOp},
{OPD(2, 0b01, 0x18), 1, &OpDispatchBuilder::VBROADCASTOp<4>},
{OPD(2, 0b01, 0x19), 1, &OpDispatchBuilder::VBROADCASTOp<8>},
{OPD(2, 0b01, 0x1A), 1, &OpDispatchBuilder::VBROADCASTOp<16>},
@@ -6054,6 +5950,10 @@ void OpDispatchBuilder::InstallHostSpecificOpcodeHandlers() {
{OPD(2, 0b01, 0x29), 1, &OpDispatchBuilder::AVXVectorALUOp<IR::OP_VCMPEQ, 8>},
{OPD(2, 0b01, 0x2A), 1, &OpDispatchBuilder::VMOVVectorNTOp},
{OPD(2, 0b01, 0x2B), 1, &OpDispatchBuilder::VPACKUSOp<4>},
{OPD(2, 0b01, 0x2C), 1, &OpDispatchBuilder::VMASKMOVOp<4, false>},
{OPD(2, 0b01, 0x2D), 1, &OpDispatchBuilder::VMASKMOVOp<8, false>},
{OPD(2, 0b01, 0x2E), 1, &OpDispatchBuilder::VMASKMOVOp<4, true>},
{OPD(2, 0b01, 0x2F), 1, &OpDispatchBuilder::VMASKMOVOp<8, true>},
{OPD(2, 0b01, 0x30), 1, &OpDispatchBuilder::AVXExtendVectorElements<1, 2, false>},
{OPD(2, 0b01, 0x31), 1, &OpDispatchBuilder::AVXExtendVectorElements<1, 4, false>},
@@ -6061,6 +5961,7 @@ void OpDispatchBuilder::InstallHostSpecificOpcodeHandlers() {
{OPD(2, 0b01, 0x33), 1, &OpDispatchBuilder::AVXExtendVectorElements<2, 4, false>},
{OPD(2, 0b01, 0x34), 1, &OpDispatchBuilder::AVXExtendVectorElements<2, 8, false>},
{OPD(2, 0b01, 0x35), 1, &OpDispatchBuilder::AVXExtendVectorElements<4, 8, false>},
{OPD(2, 0b01, 0x36), 1, &OpDispatchBuilder::VPERMDOp},
{OPD(2, 0b01, 0x37), 1, &OpDispatchBuilder::AVXVectorALUOp<IR::OP_VCMPGT, 8>},
{OPD(2, 0b01, 0x38), 1, &OpDispatchBuilder::AVXVectorALUOp<IR::OP_VSMIN, 1>},
@@ -6074,7 +5975,9 @@ void OpDispatchBuilder::InstallHostSpecificOpcodeHandlers() {
{OPD(2, 0b01, 0x40), 1, &OpDispatchBuilder::AVXVectorALUOp<IR::OP_VSMUL, 4>},
{OPD(2, 0b01, 0x41), 1, &OpDispatchBuilder::VPHMINPOSUWOp},
{OPD(2, 0b01, 0x45), 1, &OpDispatchBuilder::VPSRLVOp},
{OPD(2, 0b01, 0x46), 1, &OpDispatchBuilder::VPSRAVDOp},
{OPD(2, 0b01, 0x47), 1, &OpDispatchBuilder::VPSLLVOp},
{OPD(2, 0b01, 0x58), 1, &OpDispatchBuilder::VBROADCASTOp<4>},
{OPD(2, 0b01, 0x59), 1, &OpDispatchBuilder::VBROADCASTOp<8>},
@@ -6083,6 +5986,9 @@ void OpDispatchBuilder::InstallHostSpecificOpcodeHandlers() {
{OPD(2, 0b01, 0x78), 1, &OpDispatchBuilder::VBROADCASTOp<1>},
{OPD(2, 0b01, 0x79), 1, &OpDispatchBuilder::VBROADCASTOp<2>},
{OPD(2, 0b01, 0x8C), 1, &OpDispatchBuilder::VPMASKMOVOp<false>},
{OPD(2, 0b01, 0x8E), 1, &OpDispatchBuilder::VPMASKMOVOp<true>},
{OPD(2, 0b01, 0xDB), 1, &OpDispatchBuilder::VAESIMCOp},
{OPD(2, 0b01, 0xDC), 1, &OpDispatchBuilder::VAESEncOp},
{OPD(2, 0b01, 0xDD), 1, &OpDispatchBuilder::VAESEncLastOp},
@@ -6100,6 +6006,9 @@ void OpDispatchBuilder::InstallHostSpecificOpcodeHandlers() {
{OPD(3, 0b01, 0x0A), 1, &OpDispatchBuilder::AVXVectorRound<4, true>},
{OPD(3, 0b01, 0x0B), 1, &OpDispatchBuilder::AVXVectorRound<8, true>},
{OPD(3, 0b01, 0x0C), 1, &OpDispatchBuilder::VPBLENDDOp},
{OPD(3, 0b01, 0x0D), 1, &OpDispatchBuilder::VBLENDPDOp},
{OPD(3, 0b01, 0x0E), 1, &OpDispatchBuilder::VPBLENDWOp},
{OPD(3, 0b01, 0x0F), 1, &OpDispatchBuilder::VPALIGNROp},
{OPD(3, 0b01, 0x14), 1, &OpDispatchBuilder::PExtrOp<1>},
{OPD(3, 0b01, 0x15), 1, &OpDispatchBuilder::PExtrOp<2>},
@@ -6107,16 +6016,29 @@ void OpDispatchBuilder::InstallHostSpecificOpcodeHandlers() {
{OPD(3, 0b01, 0x17), 1, &OpDispatchBuilder::PExtrOp<4>},
{OPD(3, 0b01, 0x18), 1, &OpDispatchBuilder::VINSERTOp},
{OPD(3, 0b01, 0x19), 1, &OpDispatchBuilder::VEXTRACT128Op},
{OPD(3, 0b01, 0x20), 1, &OpDispatchBuilder::VPINSRBOp},
{OPD(3, 0b01, 0x21), 1, &OpDispatchBuilder::VINSERTPSOp},
{OPD(3, 0b01, 0x22), 1, &OpDispatchBuilder::VPINSRDQOp},
{OPD(3, 0b01, 0x38), 1, &OpDispatchBuilder::VINSERTOp},
{OPD(3, 0b01, 0x39), 1, &OpDispatchBuilder::VEXTRACT128Op},
{OPD(3, 0b01, 0x40), 1, &OpDispatchBuilder::VDPPOp<4>},
{OPD(3, 0b01, 0x41), 1, &OpDispatchBuilder::VDPPOp<8>},
{OPD(3, 0b01, 0x42), 1, &OpDispatchBuilder::VMPSADBWOp},
{OPD(3, 0b01, 0x46), 1, &OpDispatchBuilder::VPERM2Op},
{OPD(3, 0b01, 0x4A), 1, &OpDispatchBuilder::AVXVectorVariableBlend<4>},
{OPD(3, 0b01, 0x4B), 1, &OpDispatchBuilder::AVXVectorVariableBlend<8>},
{OPD(3, 0b01, 0x4C), 1, &OpDispatchBuilder::AVXVectorVariableBlend<1>},
{OPD(3, 0b01, 0x60), 1, &OpDispatchBuilder::VPCMPESTRMOp},
{OPD(3, 0b01, 0x61), 1, &OpDispatchBuilder::VPCMPESTRIOp},
{OPD(3, 0b01, 0x62), 1, &OpDispatchBuilder::VPCMPISTRMOp},
{OPD(3, 0b01, 0x63), 1, &OpDispatchBuilder::VPCMPISTRIOp},
{OPD(3, 0b01, 0xDF), 1, &OpDispatchBuilder::VAESKeyGenAssistOp},
};
#undef OPD
@@ -6186,19 +6108,19 @@ void OpDispatchBuilder::InstallHostSpecificOpcodeHandlers() {
void InstallOpcodeHandlers(Context::OperatingMode Mode) {
constexpr std::tuple<uint8_t, uint8_t, X86Tables::OpDispatchPtr> BaseOpTable[] = {
// Instructions
{0x00, 6, &OpDispatchBuilder::ALUOp},
{0x00, 6, &OpDispatchBuilder::ALUOp<FEXCore::IR::IROps::OP_ADD, FEXCore::IR::IROps::OP_ATOMICFETCHADD, true>},
{0x08, 6, &OpDispatchBuilder::ALUOp},
{0x08, 6, &OpDispatchBuilder::ALUOp<FEXCore::IR::IROps::OP_OR, FEXCore::IR::IROps::OP_ATOMICFETCHOR, false>},
{0x10, 6, &OpDispatchBuilder::ADCOp<0>},
{0x18, 6, &OpDispatchBuilder::SBBOp<0>},
{0x18, 6, &OpDispatchBuilder::SBBOp<0, true>},
{0x20, 6, &OpDispatchBuilder::ALUOp},
{0x20, 6, &OpDispatchBuilder::ALUOp<FEXCore::IR::IROps::OP_AND, FEXCore::IR::IROps::OP_ATOMICFETCHAND, false>},
{0x28, 6, &OpDispatchBuilder::ALUOp},
{0x28, 6, &OpDispatchBuilder::ALUOp<FEXCore::IR::IROps::OP_SUB, FEXCore::IR::IROps::OP_ATOMICFETCHSUB, true>},
{0x30, 6, &OpDispatchBuilder::ALUOp},
{0x30, 6, &OpDispatchBuilder::ALUOp<FEXCore::IR::IROps::OP_XOR, FEXCore::IR::IROps::OP_ATOMICFETCHXOR, false>},
{0x38, 6, &OpDispatchBuilder::CMPOp<0>},
{0x50, 8, &OpDispatchBuilder::PUSHREGOp},
@@ -6274,6 +6196,7 @@ void InstallOpcodeHandlers(Context::OperatingMode Mode) {
{0xCE, 1, &OpDispatchBuilder::INTOp},
{0xD4, 1, &OpDispatchBuilder::AAMOp},
{0xD5, 1, &OpDispatchBuilder::AADOp},
{0xD6, 1, &OpDispatchBuilder::SBBOp<0, false>},
};
constexpr std::tuple<uint8_t, uint8_t, X86Tables::OpDispatchPtr> BaseOpTable_64[] = {
@@ -6327,9 +6250,8 @@ void InstallOpcodeHandlers(Context::OperatingMode Mode) {
{0x12, 2, &OpDispatchBuilder::MOVLPOp},
{0x14, 1, &OpDispatchBuilder::PUNPCKLOp<4>},
{0x15, 1, &OpDispatchBuilder::PUNPCKHOp<4>},
{0x16, 1, &OpDispatchBuilder::MOVLHPSOp},
{0x17, 1, &OpDispatchBuilder::MOVUPSOp},
{0x28, 2, &OpDispatchBuilder::MOVUPSOp},
{0x16, 2, &OpDispatchBuilder::MOVHPDOp},
{0x28, 2, &OpDispatchBuilder::MOVAPSOp},
{0x2A, 1, &OpDispatchBuilder::MMX_To_XMM_Vector_CVT_Int_To_Float<4, false>},
{0x2B, 1, &OpDispatchBuilder::MOVVectorNTOp},
{0x2C, 1, &OpDispatchBuilder::Vector_CVT_Float_To_Int<4, false, false>},
@@ -6345,7 +6267,7 @@ void InstallOpcodeHandlers(Context::OperatingMode Mode) {
{0x57, 1, &OpDispatchBuilder::VectorALUOp<IR::OP_VXOR, 16>},
{0x58, 1, &OpDispatchBuilder::VectorALUOp<IR::OP_VFADD, 4>},
{0x59, 1, &OpDispatchBuilder::VectorALUOp<IR::OP_VFMUL, 4>},
{0x5A, 1, &OpDispatchBuilder::Vector_CVT_Float_To_Float<8, 4>},
{0x5A, 1, &OpDispatchBuilder::Vector_CVT_Float_To_Float<8, 4, false>},
{0x5B, 1, &OpDispatchBuilder::Vector_CVT_Int_To_Float<4, false>},
{0x5C, 1, &OpDispatchBuilder::VectorALUOp<IR::OP_VFSUB, 4>},
{0x5D, 1, &OpDispatchBuilder::VectorALUOp<IR::OP_VFMIN, 4>},
@@ -6419,7 +6341,6 @@ void InstallOpcodeHandlers(Context::OperatingMode Mode) {
{0xFE, 1, &OpDispatchBuilder::VectorALUOp<IR::OP_VADD, 4>},
// FEX reserved instructions
{0x36, 1, &OpDispatchBuilder::SIGRETOp},
{0x37, 1, &OpDispatchBuilder::CallbackReturnOp},
};
@@ -6437,7 +6358,7 @@ void InstallOpcodeHandlers(Context::OperatingMode Mode) {
{OPD(FEXCore::X86Tables::TYPE_GROUP_1, OpToIndex(0x80), 0), 1, &OpDispatchBuilder::SecondaryALUOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_1, OpToIndex(0x80), 1), 1, &OpDispatchBuilder::SecondaryALUOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_1, OpToIndex(0x80), 2), 1, &OpDispatchBuilder::ADCOp<1>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_1, OpToIndex(0x80), 3), 1, &OpDispatchBuilder::SBBOp<1>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_1, OpToIndex(0x80), 3), 1, &OpDispatchBuilder::SBBOp<1, true>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_1, OpToIndex(0x80), 4), 1, &OpDispatchBuilder::SecondaryALUOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_1, OpToIndex(0x80), 5), 1, &OpDispatchBuilder::SecondaryALUOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_1, OpToIndex(0x80), 6), 1, &OpDispatchBuilder::SecondaryALUOp},
@@ -6446,7 +6367,7 @@ void InstallOpcodeHandlers(Context::OperatingMode Mode) {
{OPD(FEXCore::X86Tables::TYPE_GROUP_1, OpToIndex(0x81), 0), 1, &OpDispatchBuilder::SecondaryALUOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_1, OpToIndex(0x81), 1), 1, &OpDispatchBuilder::SecondaryALUOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_1, OpToIndex(0x81), 2), 1, &OpDispatchBuilder::ADCOp<1>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_1, OpToIndex(0x81), 3), 1, &OpDispatchBuilder::SBBOp<1>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_1, OpToIndex(0x81), 3), 1, &OpDispatchBuilder::SBBOp<1, true>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_1, OpToIndex(0x81), 4), 1, &OpDispatchBuilder::SecondaryALUOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_1, OpToIndex(0x81), 5), 1, &OpDispatchBuilder::SecondaryALUOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_1, OpToIndex(0x81), 6), 1, &OpDispatchBuilder::SecondaryALUOp},
@@ -6455,7 +6376,7 @@ void InstallOpcodeHandlers(Context::OperatingMode Mode) {
{OPD(FEXCore::X86Tables::TYPE_GROUP_1, OpToIndex(0x83), 0), 1, &OpDispatchBuilder::SecondaryALUOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_1, OpToIndex(0x83), 1), 1, &OpDispatchBuilder::SecondaryALUOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_1, OpToIndex(0x83), 2), 1, &OpDispatchBuilder::ADCOp<1>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_1, OpToIndex(0x83), 3), 1, &OpDispatchBuilder::SBBOp<1>}, // Unit tests find this setting flags incorrectly
{OPD(FEXCore::X86Tables::TYPE_GROUP_1, OpToIndex(0x83), 3), 1, &OpDispatchBuilder::SBBOp<1, true>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_1, OpToIndex(0x83), 4), 1, &OpDispatchBuilder::SecondaryALUOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_1, OpToIndex(0x83), 5), 1, &OpDispatchBuilder::SecondaryALUOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_1, OpToIndex(0x83), 6), 1, &OpDispatchBuilder::SecondaryALUOp},
@@ -6635,7 +6556,7 @@ void InstallOpcodeHandlers(Context::OperatingMode Mode) {
{0x57, 1, &OpDispatchBuilder::VectorALUOp<IR::OP_VXOR, 16>},
{0x58, 1, &OpDispatchBuilder::VectorALUOp<IR::OP_VFADD, 8>},
{0x59, 1, &OpDispatchBuilder::VectorALUOp<IR::OP_VFMUL, 8>},
{0x5A, 1, &OpDispatchBuilder::Vector_CVT_Float_To_Float<4, 8>},
{0x5A, 1, &OpDispatchBuilder::Vector_CVT_Float_To_Float<4, 8, false>},
{0x5B, 1, &OpDispatchBuilder::Vector_CVT_Float_To_Int<4, false, true>},
{0x5C, 1, &OpDispatchBuilder::VectorALUOp<IR::OP_VFSUB, 8>},
{0x5D, 1, &OpDispatchBuilder::VectorALUOp<IR::OP_VFMIN, 8>},
@@ -6827,7 +6748,7 @@ constexpr uint16_t PF_F2 = 3;
constexpr std::tuple<uint8_t, uint8_t, FEXCore::X86Tables::OpDispatchPtr> SecondaryModRMExtensionOpTable[] = {
// REG /2
{((1 << 3) | 0), 1, &OpDispatchBuilder::UnimplementedOp},
{((1 << 3) | 0), 1, &OpDispatchBuilder::XGetBVOp},
// REG /7
{((3 << 3) | 1), 1, &OpDispatchBuilder::RDTSCPOp},
@@ -7406,6 +7327,11 @@ constexpr uint16_t PF_F2 = 3;
{OPD(0, PF_3A_66, 0x41), 1, &OpDispatchBuilder::DPPOp<8>},
{OPD(0, PF_3A_66, 0x42), 1, &OpDispatchBuilder::MPSADBWOp},
{OPD(0, PF_3A_66, 0x60), 1, &OpDispatchBuilder::VPCMPESTRMOp},
{OPD(0, PF_3A_66, 0x61), 1, &OpDispatchBuilder::VPCMPESTRIOp},
{OPD(0, PF_3A_66, 0x62), 1, &OpDispatchBuilder::VPCMPISTRMOp},
{OPD(0, PF_3A_66, 0x63), 1, &OpDispatchBuilder::VPCMPISTRIOp},
{OPD(0, PF_3A_NONE, 0xCC), 1, &OpDispatchBuilder::SHA1RNDS4Op},
};
#undef PF_3A_NONE
+142 -15
View File
@@ -1,24 +1,24 @@
#pragma once
#include "Interface/Core/Frontend.h"
#include "Interface/Core/X86Tables/X86Tables.h"
#include "Interface/Context/Context.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/Context.h>
#include <FEXCore/Core/X86Enums.h>
#include <FEXCore/Debug/X86Tables.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IREmitter.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/fextl/map.h>
#include <FEXCore/fextl/vector.h>
#include <cstdint>
#include <fmt/format.h>
#include <map>
#include <stddef.h>
#include <utility>
#include <vector>
namespace FEXCore::IR {
class Pass;
@@ -75,7 +75,7 @@ public:
OrderedNode* flagsOpDestSigned{};
OrderedNode* flagsOpSrcSigned{};
FEXCore::Context::Context *CTX{};
FEXCore::Context::ContextImpl *CTX{};
// Used during new op bringup
bool ShouldDump {false};
@@ -85,7 +85,7 @@ public:
bool HaveEmitted;
};
std::map<uint64_t, JumpTargetInfo> JumpTargets;
fextl::map<uint64_t, JumpTargetInfo> JumpTargets;
OrderedNode* GetNewJumpBlock(uint64_t RIP) {
auto it = JumpTargets.find(RIP);
@@ -149,7 +149,7 @@ public:
return false;
}
OpDispatchBuilder(FEXCore::Context::Context *ctx);
OpDispatchBuilder(FEXCore::Context::ContextImpl *ctx);
OpDispatchBuilder(FEXCore::Utils::IntrusivePooledAllocator &Allocator);
void ResetWorkingList();
@@ -157,7 +157,7 @@ public:
bool HadDecodeFailure() const { return DecodeFailure; }
bool NeedsBlockEnder() const { return NeedsBlockEnd; }
void BeginFunction(uint64_t RIP, std::vector<FEXCore::Frontend::Decoder::DecodedBlocks> const *Blocks);
void BeginFunction(uint64_t RIP, fextl::vector<FEXCore::Frontend::Decoder::DecodedBlocks> const *Blocks);
void Finalize();
// Dispatch builder functions
@@ -168,6 +168,7 @@ public:
void MOVGPRNTOp(OpcodeArgs);
void MOVVectorOp(OpcodeArgs);
void MOVVectorNTOp(OpcodeArgs);
template<FEXCore::IR::IROps ALUIROp, FEXCore::IR::IROps AtomicFetchOp, bool RequiresMask>
void ALUOp(OpcodeArgs);
void INTOp(OpcodeArgs);
void SyscallOp(OpcodeArgs);
@@ -176,12 +177,11 @@ public:
void NOPOp(OpcodeArgs);
void RETOp(OpcodeArgs);
void IRETOp(OpcodeArgs);
void SIGRETOp(OpcodeArgs);
void CallbackReturnOp(OpcodeArgs);
void SecondaryALUOp(OpcodeArgs);
template<uint32_t SrcIndex>
void ADCOp(OpcodeArgs);
template<uint32_t SrcIndex>
template<uint32_t SrcIndex, bool SetFlags>
void SBBOp(OpcodeArgs);
void PUSHOp(OpcodeArgs);
void PUSHREGOp(OpcodeArgs);
@@ -219,6 +219,7 @@ public:
void MOVOffsetOp(OpcodeArgs);
void CMOVOp(OpcodeArgs);
void CPUIDOp(OpcodeArgs);
void XGetBVOp(OpcodeArgs);
template<bool SHL1Bit>
void SHLOp(OpcodeArgs);
void SHLImmediateOp(OpcodeArgs);
@@ -304,7 +305,6 @@ public:
// SSE
void MOVAPSOp(OpcodeArgs);
void MOVUPSOp(OpcodeArgs);
void MOVLHPSOp(OpcodeArgs);
void MOVLPOp(OpcodeArgs);
void MOVSHDUPOp(OpcodeArgs);
void MOVSLDUPOp(OpcodeArgs);
@@ -356,7 +356,7 @@ public:
void Vector_CVT_Int_To_Float(OpcodeArgs);
template<size_t DstElementSize, size_t SrcElementSize>
void Scalar_CVT_Float_To_Float(OpcodeArgs);
template<size_t DstElementSize, size_t SrcElementSize>
template<size_t DstElementSize, size_t SrcElementSize, bool IsAVX>
void Vector_CVT_Float_To_Float(OpcodeArgs);
template<size_t SrcElementSize, bool Narrow, bool HostRoundingMode>
void Vector_CVT_Float_To_Int(OpcodeArgs);
@@ -416,12 +416,18 @@ public:
template <size_t ElementSize, bool Scalar>
void AVXVectorRound(OpcodeArgs);
template <size_t DstElementSize, size_t SrcElementSize>
void AVXScalar_CVT_Float_To_Float(OpcodeArgs);
template <size_t SrcElementSize, bool Narrow, bool HostRoundingMode>
void AVXVector_CVT_Float_To_Int(OpcodeArgs);
template <size_t SrcElementSize, bool Widen>
void AVXVector_CVT_Int_To_Float(OpcodeArgs);
template <size_t DstElementSize>
void AVXCVTGPR_To_FPR(OpcodeArgs);
template <size_t ElementSize, bool Scalar>
void AVXVFCMPOp(OpcodeArgs);
@@ -437,18 +443,29 @@ public:
void VANDNOp(OpcodeArgs);
void VBLENDPDOp(OpcodeArgs);
void VPBLENDDOp(OpcodeArgs);
void VPBLENDWOp(OpcodeArgs);
template <size_t ElementSize>
void VBROADCASTOp(OpcodeArgs);
template <size_t ElementSize>
void VDPPOp(OpcodeArgs);
void VEXTRACT128Op(OpcodeArgs);
template <IROps IROp, size_t ElementSize>
void VHADDPOp(OpcodeArgs);
template <size_t ElementSize>
void VHSUBPOp(OpcodeArgs);
void VINSERTOp(OpcodeArgs);
void VINSERTPSOp(OpcodeArgs);
template <size_t ElementSize, bool IsStore>
void VMASKMOVOp(OpcodeArgs);
void VMOVAPS_VMOVAPD_Op(OpcodeArgs);
void VMOVUPS_VMOVUPD_Op(OpcodeArgs);
@@ -459,26 +476,52 @@ public:
void VMOVSHDUPOp(OpcodeArgs);
void VMOVSLDUPOp(OpcodeArgs);
void VMOVSDOp(OpcodeArgs);
void VMOVSSOp(OpcodeArgs);
void VMOVVectorNTOp(OpcodeArgs);
void VMPSADBWOp(OpcodeArgs);
template <size_t ElementSize>
void VPACKSSOp(OpcodeArgs);
template <size_t ElementSize>
void VPACKUSOp(OpcodeArgs);
void VPBLENDDOp(OpcodeArgs);
void VPALIGNROp(OpcodeArgs);
void VPCMPESTRIOp(OpcodeArgs);
void VPCMPESTRMOp(OpcodeArgs);
void VPCMPISTRIOp(OpcodeArgs);
void VPCMPISTRMOp(OpcodeArgs);
void VPERM2Op(OpcodeArgs);
void VPERMDOp(OpcodeArgs);
void VPERMQOp(OpcodeArgs);
template <size_t ElementSize>
void VPERMILImmOp(OpcodeArgs);
template <size_t ElementSize>
void VPERMILRegOp(OpcodeArgs);
void VPHADDSWOp(OpcodeArgs);
void VPHMINPOSUWOp(OpcodeArgs);
template <size_t ElementSize>
void VPHSUBOp(OpcodeArgs);
void VPHSUBSWOp(OpcodeArgs);
void VPINSRBOp(OpcodeArgs);
void VPINSRDQOp(OpcodeArgs);
void VPINSRWOp(OpcodeArgs);
void VPMADDUBSWOp(OpcodeArgs);
void VPMADDWDOp(OpcodeArgs);
template <bool IsStore>
void VPMASKMOVOp(OpcodeArgs);
void VPMULHRSWOp(OpcodeArgs);
@@ -488,11 +531,19 @@ public:
template <size_t ElementSize, bool Signed>
void VPMULLOp(OpcodeArgs);
void VPSADBWOp(OpcodeArgs);
void VPSHUFBOp(OpcodeArgs);
template <size_t ElementSize, bool Low>
void VPSHUFWOp(OpcodeArgs);
template <size_t ElementSize>
void VPSLLOp(OpcodeArgs);
void VPSLLDQOp(OpcodeArgs);
template <size_t ElementSize>
void VPSLLIOp(OpcodeArgs);
void VPSLLVOp(OpcodeArgs);
template <size_t ElementSize>
void VPSRAOp(OpcodeArgs);
@@ -501,6 +552,7 @@ public:
void VPSRAIOp(OpcodeArgs);
void VPSRAVDOp(OpcodeArgs);
void VPSRLVOp(OpcodeArgs);
template <size_t ElementSize>
void VPSRLDOp(OpcodeArgs);
@@ -515,6 +567,12 @@ public:
template <size_t ElementSize>
void VPSRLIOp(OpcodeArgs);
template <size_t ElementSize>
void VSHUFOp(OpcodeArgs);
template <size_t ElementSize>
void VTESTPOp(OpcodeArgs);
void VZEROOp(OpcodeArgs);
// X87 Ops
@@ -755,6 +813,8 @@ private:
FEXCore::IR::IROp_IRHeader *Current_Header{};
OrderedNode *Current_HeaderNode{};
void ALUOpImpl(OpcodeArgs, FEXCore::IR::IROps ALUIROp, FEXCore::IR::IROps AtomicFetchOp, bool RequiresMask);
// Opcode helpers for generalizing behavior across VEX and non-VEX variants.
OrderedNode* ADDSUBPOpImpl(OpcodeArgs, size_t ElementSize,
@@ -764,9 +824,18 @@ private:
void AVXVectorScalarALUOpImpl(OpcodeArgs, IROps IROp, size_t ElementSize);
void AVXVectorUnaryOpImpl(OpcodeArgs, IROps IROp, size_t ElementSize, bool Scalar);
template <size_t ElementSize>
void AVXVectorVariableBlend(OpcodeArgs);
void AVXVariableShiftImpl(OpcodeArgs, IROps IROp);
OrderedNode* AESKeyGenAssistImpl(OpcodeArgs);
OrderedNode* AESIMCImpl(OpcodeArgs);
OrderedNode* CVTGPR_To_FPRImpl(OpcodeArgs, size_t DstElementSize,
const X86Tables::DecodedOperand& Src1Op,
const X86Tables::DecodedOperand& Src2Op);
OrderedNode* DPPOpImpl(OpcodeArgs, const X86Tables::DecodedOperand& Src1,
const X86Tables::DecodedOperand& Src2,
const X86Tables::DecodedOperand& Imm, size_t ElementSize);
@@ -774,21 +843,52 @@ private:
OrderedNode* ExtendVectorElementsImpl(OpcodeArgs, size_t ElementSize,
size_t DstElementSize, bool Signed);
OrderedNode* HSUBPOpImpl(OpcodeArgs, size_t ElementSize,
const X86Tables::DecodedOperand& Src1Op,
const X86Tables::DecodedOperand& Src2Op);
OrderedNode* InsertPSOpImpl(OpcodeArgs, const X86Tables::DecodedOperand& Src1,
const X86Tables::DecodedOperand& Src2,
const X86Tables::DecodedOperand& Imm);
OrderedNode* MPSADBWOpImpl(OpcodeArgs, const X86Tables::DecodedOperand& Src1Op,
const X86Tables::DecodedOperand& Src2Op,
const X86Tables::DecodedOperand& ImmOp);
OrderedNode* PACKSSOpImpl(OpcodeArgs, size_t ElementSize,
OrderedNode *Src1, OrderedNode *Src2);
OrderedNode* PACKUSOpImpl(OpcodeArgs, size_t ElementSize,
OrderedNode *Src1, OrderedNode *Src2);
OrderedNode* PALIGNROpImpl(OpcodeArgs, const X86Tables::DecodedOperand& Src1,
const X86Tables::DecodedOperand& Src2,
const X86Tables::DecodedOperand& Imm);
void PCMPXSTRXOpImpl(OpcodeArgs, bool IsExplicit, bool IsMask);
OrderedNode* PHADDSOpImpl(OpcodeArgs, const X86Tables::DecodedOperand& Src1,
const X86Tables::DecodedOperand& Src2);
OrderedNode* PHMINPOSUWOpImpl(OpcodeArgs);
OrderedNode* PHSUBOpImpl(OpcodeArgs, const X86Tables::DecodedOperand& Src1,
const X86Tables::DecodedOperand& Src2, size_t ElementSize);
OrderedNode* PHSUBSOpImpl(OpcodeArgs, const X86Tables::DecodedOperand& Src1Op,
const X86Tables::DecodedOperand& Src2Op);
OrderedNode* PINSROpImpl(OpcodeArgs, size_t ElementSize,
const X86Tables::DecodedOperand& Src1Op,
const X86Tables::DecodedOperand& Src2Op,
const X86Tables::DecodedOperand& Imm);
OrderedNode* PMADDWDOpImpl(OpcodeArgs, const X86Tables::DecodedOperand& Src1,
const X86Tables::DecodedOperand& Src2);
OrderedNode* PMADDUBSWOpImpl(OpcodeArgs, const X86Tables::DecodedOperand& Src1Op,
const X86Tables::DecodedOperand& Src2Op);
OrderedNode* PMULHRSWOpImpl(OpcodeArgs, OrderedNode *Src1, OrderedNode *Src2);
OrderedNode* PMULHWOpImpl(OpcodeArgs, bool Signed,
@@ -797,6 +897,12 @@ private:
OrderedNode* PMULLOpImpl(OpcodeArgs, size_t ElementSize, bool Signed,
OrderedNode *Src1, OrderedNode *Src2);
OrderedNode* PSADBWOpImpl(OpcodeArgs, const X86Tables::DecodedOperand& Src1Op,
const X86Tables::DecodedOperand& Src2Op);
OrderedNode* PSHUFBOpImpl(OpcodeArgs, const X86Tables::DecodedOperand& Src1,
const X86Tables::DecodedOperand& Src2);
OrderedNode* PSIGNImpl(OpcodeArgs, size_t ElementSize,
OrderedNode *Src1, OrderedNode *Src2);
@@ -812,9 +918,22 @@ private:
OrderedNode* PSRLDOpImpl(OpcodeArgs, size_t ElementSize,
OrderedNode *Src, OrderedNode *ShiftVec);
OrderedNode* SHUFOpImpl(OpcodeArgs, size_t ElementSize,
const X86Tables::DecodedOperand& Src1,
const X86Tables::DecodedOperand& Src2,
const X86Tables::DecodedOperand& Imm);
void VMASKMOVOpImpl(OpcodeArgs, size_t ElementSize, size_t DataSize, bool IsStore,
const X86Tables::DecodedOperand& MaskOp,
const X86Tables::DecodedOperand& DataOp);
void VMOVScalarOpImpl(OpcodeArgs, size_t ElementSize);
OrderedNode* VFCMPOpImpl(OpcodeArgs, size_t ElementSize, bool Scalar,
OrderedNode *Src1, OrderedNode *Src2, uint8_t CompType);
void VTESTOpImpl(OpcodeArgs, size_t ElementSize);
void VectorALUOpImpl(OpcodeArgs, IROps IROp, size_t ElementSize);
void VectorALUROpImpl(OpcodeArgs, IROps IROp, size_t ElementSize);
void VectorScalarALUOpImpl(OpcodeArgs, IROps IROp, size_t ElementSize);
@@ -824,6 +943,12 @@ private:
OrderedNode* VectorRoundImpl(OpcodeArgs, size_t ElementSize,
OrderedNode *Src, uint64_t Mode);
OrderedNode* Scalar_CVT_Float_To_FloatImpl(OpcodeArgs, size_t DstElementSize, size_t SrcElementSize,
const X86Tables::DecodedOperand& Src1Op,
const X86Tables::DecodedOperand& Src2Op);
void Vector_CVT_Float_To_FloatImpl(OpcodeArgs, size_t DstElementSize, size_t SrcElementSize, bool IsAVX);
OrderedNode* Vector_CVT_Float_To_IntImpl(OpcodeArgs, size_t SrcElementSize, bool Narrow, bool HostRoundingMode);
OrderedNode* Vector_CVT_Int_To_FloatImpl(OpcodeArgs, size_t SrcElementSize, bool Widen);
@@ -831,6 +956,8 @@ private:
#undef OpcodeArgs
OrderedNode *AppendSegmentOffset(OrderedNode *Value, uint32_t Flags, uint32_t DefaultPrefix = 0, bool Override = false);
OrderedNode *GetSegment(uint32_t Flags, uint32_t DefaultPrefix = 0, bool Override = false);
void UpdatePrefixFromSegment(OrderedNode *Segment, uint32_t SegmentReg);
enum class MemoryAccessType {
@@ -1424,21 +1551,21 @@ private:
return !Op->Dest.IsGPR();
}
void CreateJumpBlocks(std::vector<FEXCore::Frontend::Decoder::DecodedBlocks> const *Blocks);
void CreateJumpBlocks(fextl::vector<FEXCore::Frontend::Decoder::DecodedBlocks> const *Blocks);
bool BlockSetRIP {false};
bool Multiblock{};
uint64_t Entry;
OrderedNode* _StoreMemAutoTSO(FEXCore::IR::RegisterClassType Class, uint8_t Size, OrderedNode *Addr, OrderedNode *Value, uint8_t Align = 1) {
if (CTX->IsTSOEnabled())
if (CTX->IsAtomicTSOEnabled())
return _StoreMemTSO(Class, Size, Value, Addr, Invalid(), Align, MEM_OFFSET_SXTX, 1);
else
return _StoreMem(Class, Size, Value, Addr, Invalid(), Align, MEM_OFFSET_SXTX, 1);
}
OrderedNode* _LoadMemAutoTSO(FEXCore::IR::RegisterClassType Class, uint8_t Size, OrderedNode *ssa0, uint8_t Align = 1) {
if (CTX->IsTSOEnabled())
if (CTX->IsAtomicTSOEnabled())
return _LoadMemTSO(Class, Size, ssa0, Invalid(), Align, MEM_OFFSET_SXTX, 1);
else
return _LoadMem(Class, Size, ssa0, Invalid(), Align, MEM_OFFSET_SXTX, 1);
@@ -5,7 +5,8 @@ desc: Handles x86/64 Crypto instructions to IR
$end_info$
*/
#include <FEXCore/Debug/X86Tables.h>
#include "Interface/Core/X86Tables/X86Tables.h"
#include <FEXCore/IR/IREmitter.h>
#include <FEXCore/Utils/LogManager.h>
#include "Interface/Core/OpcodeDispatcher.h"
@@ -77,7 +78,7 @@ void OpDispatchBuilder::SHA1RNDS4Op(OpcodeArgs) {
using FnType = OrderedNode* (*)(OpDispatchBuilder&, OrderedNode*, OrderedNode*, OrderedNode*);
const auto f0 = [](OpDispatchBuilder &Self, OrderedNode *B, OrderedNode *C, OrderedNode *D) -> OrderedNode* {
return Self._Xor(Self._And(B, C), Self._And(Self._Not(B), D));
return Self._Xor(Self._And(B, C), Self._Andn(D, B));
};
const auto f1 = [](OpDispatchBuilder &Self, OrderedNode *B, OrderedNode *C, OrderedNode *D) -> OrderedNode* {
return Self._Xor(Self._Xor(B, C), D);
@@ -204,7 +205,7 @@ void OpDispatchBuilder::SHA256MSG2Op(OpcodeArgs) {
void OpDispatchBuilder::SHA256RNDS2Op(OpcodeArgs) {
const auto Ch = [this](OrderedNode *E, OrderedNode *F, OrderedNode *G) -> OrderedNode* {
return _Xor(_And(E, F), _And(_Not(E), G));
return _Xor(_And(E, F), _Andn(G, E));
};
const auto Major = [this](OrderedNode *A, OrderedNode *B, OrderedNode *C) -> OrderedNode* {
return _Xor(_Xor(_And(A, B), _And(A, C)), _And(B, C));
@@ -280,7 +281,7 @@ void OpDispatchBuilder::VAESIMCOp(OpcodeArgs) {
void OpDispatchBuilder::AESEncOp(OpcodeArgs) {
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *Result = _VAESEnc(Dest, Src);
OrderedNode *Result = _VAESEnc(16, Dest, Src);
StoreResult(FPRClass, Op, Result, -1);
}
@@ -293,7 +294,7 @@ void OpDispatchBuilder::VAESEncOp(OpcodeArgs) {
OrderedNode *State = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *Key = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags, -1);
OrderedNode *Result = _VAESEnc(State, Key);
OrderedNode *Result = _VAESEnc(DstSize, State, Key);
if (Is128Bit) {
Result = _VMov(16, Result);
@@ -304,7 +305,7 @@ void OpDispatchBuilder::VAESEncOp(OpcodeArgs) {
void OpDispatchBuilder::AESEncLastOp(OpcodeArgs) {
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *Result = _VAESEncLast(Dest, Src);
OrderedNode *Result = _VAESEncLast(16, Dest, Src);
StoreResult(FPRClass, Op, Result, -1);
}
@@ -317,7 +318,7 @@ void OpDispatchBuilder::VAESEncLastOp(OpcodeArgs) {
OrderedNode *State = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *Key = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags, -1);
OrderedNode *Result = _VAESEncLast(State, Key);
OrderedNode *Result = _VAESEncLast(DstSize, State, Key);
if (Is128Bit) {
Result = _VMov(16, Result);
@@ -328,7 +329,7 @@ void OpDispatchBuilder::VAESEncLastOp(OpcodeArgs) {
void OpDispatchBuilder::AESDecOp(OpcodeArgs) {
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *Result = _VAESDec(Dest, Src);
OrderedNode *Result = _VAESDec(16, Dest, Src);
StoreResult(FPRClass, Op, Result, -1);
}
@@ -341,7 +342,7 @@ void OpDispatchBuilder::VAESDecOp(OpcodeArgs) {
OrderedNode *State = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *Key = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags, -1);
OrderedNode *Result = _VAESDec(State, Key);
OrderedNode *Result = _VAESDec(DstSize, State, Key);
if (Is128Bit) {
Result = _VMov(16, Result);
@@ -352,7 +353,7 @@ void OpDispatchBuilder::VAESDecOp(OpcodeArgs) {
void OpDispatchBuilder::AESDecLastOp(OpcodeArgs) {
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *Result = _VAESDecLast(Dest, Src);
OrderedNode *Result = _VAESDecLast(16, Dest, Src);
StoreResult(FPRClass, Op, Result, -1);
}
@@ -365,7 +366,7 @@ void OpDispatchBuilder::VAESDecLastOp(OpcodeArgs) {
OrderedNode *State = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *Key = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags, -1);
OrderedNode *Result = _VAESDecLast(State, Key);
OrderedNode *Result = _VAESDecLast(DstSize, State, Key);
if (Is128Bit) {
Result = _VMov(16, Result);
@@ -399,7 +400,7 @@ void OpDispatchBuilder::PCLMULQDQOp(OpcodeArgs) {
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
const auto Selector = static_cast<uint8_t>(Op->Src[1].Data.Literal.Value);
auto Res = _PCLMUL(Dest, Src, Selector);
auto Res = _PCLMUL(16, Dest, Src, Selector);
StoreResult(FPRClass, Op, Res, -1);
}
@@ -413,7 +414,7 @@ void OpDispatchBuilder::VPCLMULQDQOp(OpcodeArgs) {
OrderedNode *Src2 = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags, -1);
const auto Selector = static_cast<uint8_t>(Op->Src[2].Data.Literal.Value);
OrderedNode *Res = _PCLMUL(Src1, Src2, Selector);
OrderedNode *Res = _PCLMUL(DstSize, Src1, Src2, Selector);
if (Is128Bit) {
Res = _VMov(16, Res);
}
@@ -7,10 +7,10 @@ $end_info$
#include "Interface/Context/Context.h"
#include "Interface/Core/OpcodeDispatcher.h"
#include "Interface/Core/X86Tables/X86Tables.h"
#include <FEXCore/Core/X86Enums.h>
#include <FEXCore/Config/Config.h>
#include <FEXCore/Debug/X86Tables.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/IR/IR.h>
@@ -52,10 +52,9 @@ void OpDispatchBuilder::SetPackedRFLAG(bool Lower8, OrderedNode *Src) {
InvalidateDeferredFlags();
}
auto OneConst = _Constant(1);
for (size_t i = 0; i < NumFlags; ++i) {
const auto FlagOffset = FlagOffsets[i];
auto Tmp = _And(_Lshr(Src, _Constant(FlagOffset)), OneConst);
auto Tmp = _Bfe(4, 1, FlagOffset, Src);
SetRFLAG(Tmp, FlagOffset);
}
}
@@ -271,10 +270,8 @@ void OpDispatchBuilder::CalculcateFlags_ADC(uint8_t SrcSize, OrderedNode *Res, O
// SF
{
auto SignBitConst = _Constant(Size - 1);
auto LshrOp = _Lshr(Res, SignBitConst);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(LshrOp);
auto SignOp = _Bfe(1, SrcSize * 8 - 1, Res);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(SignOp);
}
// PF
@@ -306,28 +303,11 @@ void OpDispatchBuilder::CalculcateFlags_ADC(uint8_t SrcSize, OrderedNode *Res, O
// OF
// Signed
{
auto NegOne = _Constant(~0ULL);
auto XorOp1 = _Xor(_Xor(Src1, Src2), NegOne);
auto XorOp1 = _Not(_Xor(Src1, Src2));
auto XorOp2 = _Xor(Res, Src1);
OrderedNode *AndOp1 = _And(XorOp1, XorOp2);
switch (Size) {
case 8:
AndOp1 = _Bfe(1, 7, AndOp1);
break;
case 16:
AndOp1 = _Bfe(1, 15, AndOp1);
break;
case 32:
AndOp1 = _Bfe(1, 31, AndOp1);
break;
case 64:
AndOp1 = _Bfe(1, 63, AndOp1);
break;
default:
LOGMAN_MSG_A_FMT("Unknown BFE size: {}", Size);
break;
}
AndOp1 = _Bfe(1, Size - 1, AndOp1);
SetRFLAG<FEXCore::X86State::RFLAG_OF_LOC>(AndOp1);
}
}
@@ -342,10 +322,8 @@ void OpDispatchBuilder::CalculcateFlags_SBB(uint8_t SrcSize, OrderedNode *Res, O
// SF
{
auto SignBitConst = _Constant(SrcSize * 8 - 1);
auto LshrOp = _Lshr(Res, SignBitConst);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(LshrOp);
auto SignOp = _Bfe(1, SrcSize * 8 - 1, Res);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(SignOp);
}
// PF
@@ -380,24 +358,7 @@ void OpDispatchBuilder::CalculcateFlags_SBB(uint8_t SrcSize, OrderedNode *Res, O
auto XorOp1 = _Xor(Src1, Src2);
auto XorOp2 = _Xor(Res, Src1);
OrderedNode *AndOp1 = _And(XorOp1, XorOp2);
switch (SrcSize) {
case 1:
AndOp1 = _Bfe(1, 7, AndOp1);
break;
case 2:
AndOp1 = _Bfe(1, 15, AndOp1);
break;
case 4:
AndOp1 = _Bfe(1, 31, AndOp1);
break;
case 8:
AndOp1 = _Bfe(1, 63, AndOp1);
break;
default:
LOGMAN_MSG_A_FMT("Unknown BFE size: {}", SrcSize);
break;
}
AndOp1 = _Bfe(1, SrcSize * 8 - 1, AndOp1);
SetRFLAG<FEXCore::X86State::RFLAG_OF_LOC>(AndOp1);
}
}
@@ -412,10 +373,8 @@ void OpDispatchBuilder::CalculcateFlags_SUB(uint8_t SrcSize, OrderedNode *Res, O
// SF
{
auto SignBitConst = _Constant(SrcSize * 8 - 1);
auto LshrOp = _Lshr(Res, SignBitConst);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(LshrOp);
auto SignOp = _Bfe(1, SrcSize * 8 - 1, Res);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(SignOp);
}
// PF
@@ -469,10 +428,8 @@ void OpDispatchBuilder::CalculcateFlags_ADD(uint8_t SrcSize, OrderedNode *Res, O
// SF
{
auto SignBitConst = _Constant(SrcSize * 8 - 1);
auto LshrOp = _Lshr(Res, SignBitConst);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(LshrOp);
auto SignOp = _Bfe(1, SrcSize * 8 - 1, Res);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(SignOp);
}
// PF
@@ -500,29 +457,12 @@ void OpDispatchBuilder::CalculcateFlags_ADD(uint8_t SrcSize, OrderedNode *Res, O
// OF
{
auto NegOne = _Constant(~0ULL);
auto XorOp1 = _Xor(_Xor(Src1, Src2), NegOne);
auto XorOp1 = _Not(_Xor(Src1, Src2));
auto XorOp2 = _Xor(Res, Src1);
OrderedNode *AndOp1 = _And(XorOp1, XorOp2);
switch (SrcSize) {
case 1:
AndOp1 = _Bfe(1, 7, AndOp1);
break;
case 2:
AndOp1 = _Bfe(1, 15, AndOp1);
break;
case 4:
AndOp1 = _Bfe(1, 31, AndOp1);
break;
case 8:
AndOp1 = _Bfe(1, 63, AndOp1);
break;
default:
LOGMAN_MSG_A_FMT("Unknown BFE size: {}", SrcSize);
break;
}
AndOp1 = _Bfe(1, SrcSize * 8 - 1, AndOp1);
SetRFLAG<FEXCore::X86State::RFLAG_OF_LOC>(AndOp1);
}
}
@@ -583,10 +523,8 @@ void OpDispatchBuilder::CalculcateFlags_Logical(uint8_t SrcSize, OrderedNode *Re
// SF
{
auto SignBitConst = _Constant(SrcSize * 8 - 1);
auto LshrOp = _Lshr(Res, SignBitConst);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(LshrOp);
auto SignOp = _Bfe(1, SrcSize * 8 - 1, Res);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(SignOp);
}
// PF
@@ -750,10 +688,8 @@ void OpDispatchBuilder::CalculcateFlags_SignShiftRight(uint8_t SrcSize, OrderedN
// SF
{
auto SignBitConst = _Constant(SrcSize * 8 - 1);
auto LshrOp = _Lshr(Res, SignBitConst);
COND_FLAG_SET(Src2, RFLAG_SF_LOC, LshrOp);
auto SignBitOp = _Bfe(1, SrcSize * 8 - 1, Res);
COND_FLAG_SET(Src2, RFLAG_SF_LOC, SignBitOp);
}
// OF
@@ -802,15 +738,14 @@ void OpDispatchBuilder::CalculcateFlags_ShiftLeftImmediate(uint8_t SrcSize, Orde
// SF
{
auto LshrOp = _Bfe(1, SrcSize * 8 - 1, Res);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(LshrOp);
auto SignOp = _Bfe(1, SrcSize * 8 - 1, Res);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(SignOp);
// OF
// In the case of left shift. OF is only set from the result of <Top Source Bit> XOR <Top Result Bit>
if (Shift == 1) {
auto SourceBit = _Bfe(1, SrcSize * 8 - 1, Src1);
SetRFLAG<FEXCore::X86State::RFLAG_OF_LOC>(_Xor(SourceBit, LshrOp));
SetRFLAG<FEXCore::X86State::RFLAG_OF_LOC>(_Xor(SourceBit, SignOp));
}
}
}
@@ -851,10 +786,8 @@ void OpDispatchBuilder::CalculcateFlags_SignShiftRightImmediate(uint8_t SrcSize,
// SF
{
auto SignBitConst = _Constant(SrcSize * 8 - 1);
auto LshrOp = _Lshr(Res, SignBitConst);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(LshrOp);
auto SignBitOp = _Bfe(1, SrcSize * 8 - 1, Res);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(SignBitOp);
// OF
// Only defined when Shift is 1 else undefined
@@ -902,10 +835,8 @@ void OpDispatchBuilder::CalculcateFlags_ShiftRightImmediate(uint8_t SrcSize, Ord
// SF
{
auto SignBitConst = _Constant(SrcSize * 8 - 1);
auto LshrOp = _Lshr(Res, SignBitConst);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(LshrOp);
auto SignBitOp = _Bfe(1, SrcSize * 8 - 1, Res);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(SignBitOp);
}
// OF
@@ -1115,10 +1046,8 @@ void OpDispatchBuilder::CalculcateFlags_BLSI(uint8_t SrcSize, OrderedNode *Src)
// SF
{
auto SignBit = _Constant(SrcSize * 8 - 1);
auto SFOp = _Lshr(Src, SignBit);
SetRFLAG<X86State::RFLAG_SF_LOC>(SFOp);
auto SignOp = _Bfe(1, SrcSize * 8 - 1, Src);
SetRFLAG<X86State::RFLAG_SF_LOC>(SignOp);
}
}
@@ -1174,10 +1103,8 @@ void OpDispatchBuilder::CalculcateFlags_BLSR(uint8_t SrcSize, OrderedNode *Resul
// SF
{
auto SignBit = _Constant(SrcSize * 8 - 1);
auto SFOp = _Lshr(Result, SignBit);
SetRFLAG<X86State::RFLAG_SF_LOC>(SFOp);
auto SignOp = _Bfe(1, SrcSize * 8 - 1, Result);
SetRFLAG<X86State::RFLAG_SF_LOC>(SignOp);
}
}
@@ -1230,9 +1157,8 @@ void OpDispatchBuilder::CalculcateFlags_BZHI(uint8_t SrcSize, OrderedNode *Resul
// SF
{
auto SFOp = _Lshr(Result, Bounds);
SetRFLAG<X86State::RFLAG_SF_LOC>(SFOp);
auto SignOp = _Bfe(1, SrcSize * 8 - 1, Result);
SetRFLAG<X86State::RFLAG_SF_LOC>(SignOp);
}
}
File diff suppressed because it is too large. Load diff
@@ -6,10 +6,10 @@ $end_info$
*/
#include "Interface/Core/OpcodeDispatcher.h"
#include "Interface/Core/X86Tables/X86Tables.h"
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Core/X86Enums.h>
#include <FEXCore/Debug/X86Tables.h>
#include <FEXCore/Utils/EnumUtils.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/IR/IREmitter.h>
@@ -1374,8 +1374,7 @@ void OpDispatchBuilder::X87FCMOV(OpcodeArgs) {
SrcCond = _Sbfe(1, 0, SrcCond);
OrderedNode *VecCond = _VCastFromGPR(16, 8, SrcCond);
VecCond = _VInsGPR(16, 8, 1, VecCond, SrcCond);
OrderedNode *VecCond = _VDupFromGPR(16, 8, SrcCond);
auto top = GetX87Top();
OrderedNode* arg;
@@ -1388,7 +1387,7 @@ void OpDispatchBuilder::X87FCMOV(OpcodeArgs) {
auto a = _LoadContextIndexed(top, 16, MMBaseOffset(), 16, FPRClass);
auto b = _LoadContextIndexed(arg, 16, MMBaseOffset(), 16, FPRClass);
auto Result = _VBSL(VecCond, b, a);
auto Result = _VBSL(16, VecCond, b, a);
// Write to ST[TOP]
_StoreContextIndexed(Result, top, 16, MMBaseOffset(), 16, FPRClass);
@@ -6,10 +6,10 @@ $end_info$
*/
#include "Interface/Core/OpcodeDispatcher.h"
#include "Interface/Core/X86Tables/X86Tables.h"
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Core/X86Enums.h>
#include <FEXCore/Debug/X86Tables.h>
#include <FEXCore/Utils/EnumUtils.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/IR/IREmitter.h>
Loaded 100 of 587 files, more files were not shown because too many files have changed in this diff. Show more